BILinMID: A Spanish-English Corpus of the US Midwest

Irati Hurtado


Abstract
This paper describes the Bilinguals in the Midwest (BILinMID) Corpus, a comparable text corpus of the Spanish and English spoken in the US Midwest by various types of bilinguals. Unlike other areas within the US where language contact has been widely documented (e.g., the Southwest), Spanish-English bilingualism in the Midwest has been understudied despite an increase in its Hispanic population. The BILinMID Corpus contains short stories narrated in Spanish and in English by 72 speakers representing different types of bilinguals: early simultaneous bilinguals, early sequential bilinguals, and late second language learners. All stories have been transcribed and annotated using various natural language processing tools. Additionally, a user interface has also been created to facilitate searching for specific patterns in the corpus as well as to filter out results according to specified criteria. Guidelines and procedures followed to create the corpus and the user interface are described in detail in the paper. The corpus is fully available online and it might be particularly interesting for researchers working on language variation and contact.
Anthology ID:
2022.lrec-1.590
Volume:
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Month:
June
Year:
2022
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
5511–5516
Language:
URL:
https://aclanthology.org/2022.lrec-1.590
DOI:
Bibkey:
Cite (ACL):
Irati Hurtado. 2022. BILinMID: A Spanish-English Corpus of the US Midwest. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 5511–5516, Marseille, France. European Language Resources Association.
Cite (Informal):
BILinMID: A Spanish-English Corpus of the US Midwest (Hurtado, LREC 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.lrec-1.590.pdf
Data
Universal Dependencies