A Learning-Based Dependency to Constituency Conversion Algorithm for the Turkish Language

Büşra Marşan, Oğuz K. Yıldız, Aslı Kuzgun, Neslihan Cesur, Arife B. Yenice, Ezgi Sanıyar, Oğuzhan Kuyrukçu, Bilge N. Arıcan, Olcay Taner Yıldız


Abstract
This study aims to create the very first dependency-to-constituency conversion algorithm optimised for Turkish language. For this purpose, a state-of-the-art morphologic analyser and a feature-based machine learning model was used. In order to enhance the performance of the conversion algorithm, bootstrap aggregating meta-algorithm was integrated. While creating the conversation algorithm, typological properties of Turkish were carefully considered. A comprehensive and manually annotated UD-style dependency treebank was the input, and constituency trees were the output of the conversion algorithm. A team of linguists manually annotated a set of constituency trees. These manually annotated trees were used as the gold standard to assess the performance of the algorithm. The conversion process yielded more than 8000 constituency trees whose UD-style dependency trees are also available on GitHub. In addition to its contribution to Turkish treebank resources, this study also offers a viable and easy-to-implement conversion algorithm that can be used to generate new constituency treebanks and training data for NLP resources like constituency parsers.
Anthology ID:
2022.lrec-1.540
Volume:
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Month:
June
Year:
2022
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
5054–5062
Language:
URL:
https://aclanthology.org/2022.lrec-1.540
DOI:
Bibkey:
Cite (ACL):
Büşra Marşan, Oğuz K. Yıldız, Aslı Kuzgun, Neslihan Cesur, Arife B. Yenice, Ezgi Sanıyar, Oğuzhan Kuyrukçu, Bilge N. Arıcan, and Olcay Taner Yıldız. 2022. A Learning-Based Dependency to Constituency Conversion Algorithm for the Turkish Language. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 5054–5062, Marseille, France. European Language Resources Association.
Cite (Informal):
A Learning-Based Dependency to Constituency Conversion Algorithm for the Turkish Language (Marşan et al., LREC 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.lrec-1.540.pdf