Xabier Irastortza-Urbieta
2026
Language Mixture to Develop Accurate Galician Dependency Parsers: An Exploration of Its Effects
Xabier Irastortza-Urbieta | José M. García-Miguel | Marcos Garcia
Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects
Xabier Irastortza-Urbieta | José M. García-Miguel | Marcos Garcia
Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects
The development of accurate syntactic parsers remains a challenge for low-resource languages. To overcome it, the literature has proposed leveraging syntactic annotations from typologically related languages. This work investigates the viability and adequacy of this approach for Galician, evaluating the use of annotations from major Romance languages as source data. Our methodology extends beyond standard automatic evaluation to incorporate a detailed error analysis, which precisely quantifies the effects of multilingual training and assesses the practical scalability of the method. The results establish the necessity of embedding models for effective cross-lingual transfer and demonstrate that even languages not particularly close can yield adequate parsers. This work confirms the benefits of cross-lingual data augmentation while delineating its scalability limits. Furthermore, the error analysis identifies specific, typologically conditioned grammatical dependencies that remain persistent challenges for accurate dependency parsing.
HiTZ-IXA at ArchEHR-QA 2026: Evidence Alignment Through Self-Consistency and Prompt Curation in Memory-Constrained Environments
Xabier Irastortza-Urbieta | Maite Oronoz | Alicia Pérez
Proceedings of the Third Workshop on Patient-Oriented Language Processing (CL4Health) @ LREC 2026
Xabier Irastortza-Urbieta | Maite Oronoz | Alicia Pérez
Proceedings of the Third Workshop on Patient-Oriented Language Processing (CL4Health) @ LREC 2026
The development of question-answering systems capable of grounding their answers in Electronic Health Records could provide patients with faithful assistance while reducing the clinical workload. The ArchEHR-QA 2026 Shared Task was organized to advance progress in this context. In this paper, we present our strategies for addressing this shared task, which are focused primarily on evidence alignment and, to a lesser extent, on evidence identification. Our approaches rely exclusively on open-source models with up to 8 billion parameters, aiming to produce systems suitable for environments with memory constraints. We experimented with methods based on embedding models, prompt curation, self-consistency, and combination of LLMs. We concluded that prompt curation together with an effective post-processing step was crucial for creating stable systems, while self-consistency yielded considerable gains in performance. The results of our approaches suggest that small LLMs can substantially improve their accuracy in the evidence alignment task via simple and affordable techniques.