Querying a Dozen Corpora and a Thousand Years with Fintan

Christian Chiarcos, Christian Fäth, Maxim Ionov


Abstract
Large-scale diachronic corpus studies covering longer time periods are difficult if more than one corpus are to be consulted and, as a result, different formats and annotation schemas need to be processed and queried in a uniform, comparable and replicable manner. We describes the application of the Flexible Integrated Transformation and Annotation eNgineering (Fintan) platform for studying word order in German using syntactically annotated corpora that represent its entire written history. Focusing on nominal dative and accusative arguments, this study hints at two major phases in the development of scrambling in modern German. Against more recent assumptions, it supports the traditional view that word order flexibility decreased over time, but it also indicates that this was a relatively sharp transition in Early New High German. The successful case study demonstrates the potential of Fintan and the underlying LLOD technology for historical linguistics, linguistic typology and corpus linguistics. The technological contribution of this paper is to demonstrate the applicability of Fintan for querying across heterogeneously annotated corpora, as previously, it had only been applied for transformation tasks. With its focus on quantitative analysis, Fintan is a natural complement for existing multi-layer technologies that focus on query and exploration.
Anthology ID:
2022.lrec-1.427
Volume:
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Month:
June
Year:
2022
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
4011–4021
Language:
URL:
https://aclanthology.org/2022.lrec-1.427
DOI:
Bibkey:
Cite (ACL):
Christian Chiarcos, Christian Fäth, and Maxim Ionov. 2022. Querying a Dozen Corpora and a Thousand Years with Fintan. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4011–4021, Marseille, France. European Language Resources Association.
Cite (Informal):
Querying a Dozen Corpora and a Thousand Years with Fintan (Chiarcos et al., LREC 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.lrec-1.427.pdf