Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization

Yiming Liang, Fang Zhao


Abstract
Recent years have seen growing interest in applying neural networks and contextualized word embeddings to the parsing of historical languages. However, most advances have focused on dependency parsing, while constituency parsing for low-resource historical languages like Middle Dutch has received little attention. In this paper, we adapt a transformer-based constituency parser to Middle Dutch, a highly heterogeneous and low-resource language, and investigate methods to improve both its in-domain and cross-domain performance. We show that joint training with higher-resource auxiliary languages increases F1 scores by up to 0.73, with the greatest gains achieved from languages that are geographically and temporally closer to Middle Dutch. We further evaluate strategies for leveraging newly annotated data from additional domains, finding that fine-tuning and data combination yield comparable improvements, and our neural parser consistently outperforms the currently used PCFG-based parser for Middle Dutch. We further explore feature-separation techniques for domain adaptation and demonstrate that a minimum threshold of approximately 200 examples per domain is needed to effectively enhance cross-domain performance.
Anthology ID:
2026.lrec-1.807
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
10276–10290
Language:
External URL:
https://lrec.elra.info/lrec2026-main-807
DOI:
10.63317/3bg3jhcbj8in
Bibkey:
Cite (ACL):
Yiming Liang and Fang Zhao. 2026. Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 10276–10290, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization (Liang & Zhao, LREC 2026)
Copy Citation: