Kellen Parker van Dam
2026
Automated Quality Control for Language Documentation: Detecting Phonotactic Inconsistencies in a Kokborok Wordlist
Kellen Parker van Dam | Abishek Stephen
Proceedings of the Fifth Workshop on NLP Applications to Field Linguistics
Kellen Parker van Dam | Abishek Stephen
Proceedings of the Fifth Workshop on NLP Applications to Field Linguistics
Lexical data collection in language documentation often contains transcription errors and borrowings that can mislead linguistic analysis. We present unsupervised methods to identify phonotactic inconsistencies in wordlists, applying them to a multilingual dataset of Kokborok varieties with Bangla. Using phoneme-level and syllable-level n-gram language models, our approach identifies potential transcription errors and borrowings. We evaluate our methods using hand annotated gold standard and rank the phonotactic outliers using precision and recall at K metric. The ranking approach provides field linguists with a method to flag entries requiring verification, supporting data quality improvement in low-resourced language documentation.
Wancho Dialectometry: Community-created data and the Living Dictionaries project
Kellen Parker van Dam
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Kellen Parker van Dam
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Community-created lexical resources for under-documented languages represent an underexplored data source for computational dialectology. This study evaluates the viability of such data for dialectometric analysis, using the Wancho (Glottocode: wanc1238) LivingDictionaries project as a case study. Wancho is a Tibeto-Burman language of the Southwestern Patkaian branch, spoken primarily in Longding District, Arunachal Pradesh, India. The dictionary is notable for being entirely community-built and speaker-facing, and uniquely among resources for Northeast India, it incorporates dialect-specific forms spanning village-level geolects and clanlects. We extract dialectal data via automated web scraping and apply a series of preprocessing steps to address inconsistencies in transcription, language labelling, and concept assignment. Pairwise linguistic distances are then computed using Sound Class Alignment (SCA, List 2010), which captures phonological similarity more accurately than raw edit distance by incorporating articulatory feature structure. The resulting distance matrix is analysed through UPGMA hierarchical clustering and NeighborNet split network inference. Despite the dataset’s uneven dialect coverage and absence of systematic cognate coding, SCA-based distances recover the traditional Upper/Lower/Middle Wancho distinction and correctly situate transitional varieties. These results hold even for dialects with as few as a dozen attested forms. We show that unlike Bayesian phylogenetic inference which is poorly suited to data of this density and distribution, SCA proves to be a reliable metric. Our findings suggest that SCA distance is robust to the kinds of noise and sparsity characteristic of community-generated lexical data, and that such resources constitute a viable, if imperfect, input for automated dialectometric workflows — particularly in contexts where fieldwork-based data collection is not currently feasible.
2025
Annotating and Inferring Compositional Structures in Numeral Systems Across Languages
Arne Rubehn | Christoph Rzymski | Luca Ciucci | Katja Bocklage | Alžběta Kučerová | David Snee | Abishek Stephen | Kellen Parker van Dam | Johann-Mattis List
Proceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP
Arne Rubehn | Christoph Rzymski | Luca Ciucci | Katja Bocklage | Alžběta Kučerová | David Snee | Abishek Stephen | Kellen Parker van Dam | Johann-Mattis List
Proceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP
Numeral systems across the world’s languages vary in fascinating ways, both regarding their synchronic structure and the diachronic processes that determined how they evolved in their current shape. For a proper comparison of numeral systems across different languages, however, it is important to code them in a standardized form that allows for the comparison of basic properties. Here, we present a simple but effective coding scheme for numeral annotation, along with a workflow that helps to code numeral systems in a computer-assisted manner, providing sample data for numerals from 1 to 40 in 25 typologically diverse languages. We perform a thorough analysis of the sample, focusing on the systematic comparison between the underlying and the surface morphological structure. We further experiment with automated models for morpheme segmentation, where we find allomorphy as the major reason for segmentation errors. Finally, we show that subword tokenization algorithms are not viable for discovering morphemes in low-resource scenarios.
Unstable Grounds for Beautiful Trees? Testing the Robustness of Concept Translations in the Compilation of Multilingual Wordlists
David Snee | Luca Ciucci | Arne Rubehn | Kellen Parker van Dam | Johann-Mattis List
Proceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP
David Snee | Luca Ciucci | Arne Rubehn | Kellen Parker van Dam | Johann-Mattis List
Proceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP
Multilingual wordlists play a crucial role in comparative linguistics. While many studies have been carried out to test the power of computational methods for language subgrouping or divergence time estimation, few studies have put the data upon which these studies are based to a rigorous test. Here, we conduct a first experiment that tests the robustness of concept translation as an integral part of the compilation of multilingual wordlists. Investigating the variation in concept translations in independently compiled wordlists from 10 dataset pairs covering 9 different language families, we find that on average, only 83% of all translations yield the same word form, while identical forms in terms of phonetic transcriptions can only be found in 23% of all cases. Our findings can prove important when trying to assess the uncertainty of phylogenetic studies and the conclusions derived from them.