Johanna Cordova
2026
What LID Systems Say About Dialectal Variation. The Case of Yiddish, Quechua and Mande
Johanna Cordova | Eric Jordan | Valentina Fedchenko
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Johanna Cordova | Eric Jordan | Valentina Fedchenko
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
This study investigates the ability of speech-based language identification (LID) systems to handle dialectal variation in low-resource settings and explores whether classification outcomes correspond to phonetic proximity and can serve as an exploratory tool for dataset quality. We collected corpora for three macrolanguages Mande, Quechuan, and Yiddish, each presenting distinct internal variation, and evaluated three types of models: GMM, Whisper, and Wav2vec2-based architectures. Models were tested both within language families and across the entire multilingual dataset to assess generalization. Layer-wise classifiers built on wav2vec2-XLSR embeddings were used to identify the layers most sensitive to phonetic or phonological features. Results show that simple GMM models can generalize well in small, highly similar datasets, while Whisper-based classifiers tend to overfit, particularly on closely related dialects. Wav2vec2-XLSR (layer 12 + MLP) captures better fine phonetic and prosodic distinctions, suggesting that embeddings encode nuanced pronunciation cues. For datasets with more diverse sources like Quechua, Whisper demonstrates better generalization. Overall, LID classifiers can both reveal linguistic patterns and highlight dataset quality issues, with model architecture and layer-specific representations shaping performance.
Building Community-Centred NLP Resources for Puno Quechua
Elwin Huaman | Adrian Gamarra Lafuente | Johanna Cordova | Anna Korhonen
Proceedings of the Sixth Workshop on NLP for Indigenous Languages of the Americas (AmericasNLP)
Elwin Huaman | Adrian Gamarra Lafuente | Johanna Cordova | Anna Korhonen
Proceedings of the Sixth Workshop on NLP for Indigenous Languages of the Americas (AmericasNLP)
The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): (1) the largest speech corpus for any single Quechua variety, consisting in 66 hours of recordings for scripted and spontaneous speech (including 36 hours of manually transcribed and validated data), collected via a participatory design campaign; (2) the first systematic ASR benchmark for Puno Quechua, evaluating state-of-the-art models and fine-tuning Whisper-base, wav2vec2-base, and XLS-R-300M, with and without continued pre-training (CPT); (3) an open release of all datasets and fine-tuned models.
2024
Towards Universal Dependencies for Ancash Quechua
Johanna Cordova
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Johanna Cordova
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
This paper presents a brief description of some morphosyntactic features of Ancash Quechua, the majority variety of the Central Quechua language family (QI), for the purpose of building a corpus annotated according to the Universal Dependencies (UD) schema. The creation of such a corpus has two objectives: for Quechua linguistics, it opens up the possibility of more systematic linguistic studies and comparisons with other languages. It also enables the development of a syntactic parser, which would be the first NLP tool for a Quechua language of this family. For the UD project, adding Quechua, an agglutinative language with a rich morphology, makes it possible to point out some possible shortcomings of the universal annotation schema, and to fuel the discussion to adapt this schema to the specific features of the languages with a similar typology. The first step towards this work was first to gather and digitise the available linguistic resources, thus creating the first bilingual and sentence-aligned digital corpus in Ancash Quechua and Spanish. After identifying some linguistic features not fully described in the UD schema, we proposed annotation solutions, and built an initial corpus of around twenty sentences, which we are making freely available.
2021
Toward Creation of Ancash Lexical Resources from OCR
Johanna Cordova | Damien Nouvel
Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas
Johanna Cordova | Damien Nouvel
Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas
The Quechua linguistic family has a limited number of NLP resources, most of them being dedicated to Southern Quechua, whereas the varieties of Central Quechua have, to the best of our knowledge, no specific resources (software, lexicon or corpus). Our work addresses this issue by producing two resources for the Ancash Quechua: a full digital version of a dictionary, and an OCR model adapted to the considered variety. In this paper, we describe the steps towards this goal: we first measure performances of existing models for the task of digitising a Quechua dictionary, then adapt a model for the Ancash variety, and finally create a reliable resource for NLP in XML-TEI format. We hope that this work will be a basis for initiating NLP projects for Central Quechua, and that it will encourage digitisation initiatives for under-resourced languages.