Automatic Extraction of the Romanian Academic Word List: Data and Methods

Ana-Maria Bucur, Andreea Dincă, Madalina Chitez, Roxana Rogobete


Abstract
This paper presents the methodology and data used for the automatic extraction of the Romanian Academic Word List (Ro-AWL). Academic Word Lists are useful in both L2 and L1 teaching contexts. For the Romanian language, no such resource exists so far. Ro-AWL has been generated by combining methods from corpus and computational linguistics with L2 academic writing approaches. We use two types of data: (a) existing data, such as the Romanian Frequency List based on the ROMBAC corpus, and (b) self-compiled data, such as the expert academic writing corpus EXPRES. For constructing the academic word list, we follow the methodology for building the Academic Vocabulary List for the English language. The distribution of Ro-AWL features (general distribution, POS distribution) into four disciplinary datasets is in line with previous research. Ro-AWL is freely available and can be used for teaching, research and NLP applications.
Anthology ID:
2023.ranlp-1.26
Volume:
Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing
Month:
September
Year:
2023
Address:
Varna, Bulgaria
Editors:
Ruslan Mitkov, Galia Angelova
Venue:
RANLP
SIG:
Publisher:
INCOMA Ltd., Shoumen, Bulgaria
Note:
Pages:
234–241
Language:
URL:
https://aclanthology.org/2023.ranlp-1.26
DOI:
Bibkey:
Cite (ACL):
Ana-Maria Bucur, Andreea Dincă, Madalina Chitez, and Roxana Rogobete. 2023. Automatic Extraction of the Romanian Academic Word List: Data and Methods. In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, pages 234–241, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Cite (Informal):
Automatic Extraction of the Romanian Academic Word List: Data and Methods (Bucur et al., RANLP 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.ranlp-1.26.pdf