Glossed Data in Northern Interior Salish

Anna Stacey


Abstract
The Northern Interior subgroup of the Salish language family, spoken in the Pacific Northwest of North America, comprises three languages: St’át’imcets, nɬeʔkepmxcín, and Secwepemctsín. Each has a small number of first-language (L1) speakers remaining due to the effects of colonization, though language revitalization efforts are ongoing. This work introduces the first compiled and cleaned language datasets in these languages, useable in natural language processing (NLP) projects. This data is in glossed format, with transcriptions in the language, translations into English, and linguistic segmentations and glosses that provide a detailed breakdown of meaning. In order to achieve consistently formatted data within and across each language, extensive data cleaning was conducted. This paper provides the glossed data standards that were developed and recounts the cleaning process. Scripts that help to automate parts of the data preparation processes are included. Finally, this work strives to keep the interconnectedness of language and community as a central consideration.
Anthology ID:
2026.lrec-1.278
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3490–3495
Language:
External URL:
https://lrec.elra.info/lrec2026-main-278
DOI:
10.63317/2isngefy6ags
Bibkey:
Cite (ACL):
Anna Stacey. 2026. Glossed Data in Northern Interior Salish. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3490–3495, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Glossed Data in Northern Interior Salish (Stacey, LREC 2026)
Copy Citation: