Gabriela Nicole Gonzalez Saez
Also published as: Gabriela nicole Gonzalez Saez
2026
COME-ALPs: Coreference Annotation with MErging Heuristics Using ALignment-based Projection in Parallel Corpora
Gabriela Nicole Gonzalez Saez | Mariam Nakhle | Illia Kholosha | Rachel Atherly | Marco Dinarelli
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Gabriela Nicole Gonzalez Saez | Mariam Nakhle | Illia Kholosha | Rachel Atherly | Marco Dinarelli
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Multi-lingual, parallel datasets annotated with discourse phenomena like coreferences are a rare resource. These datasets are useful and informative to evaluate models for NLP tasks taking long contextual information into account, as proved by the large literature published in the last couple of years on e.g. Context-Aware Neural Machine Translation (CA-NMT). Inspired by resources published in previous work, in this paper we propose an automated procedure to annotate multi-lingual, parallel data with coreferences. Through the use of accurate alignment and coreference annotation tools, we project the annotation from English data, where tools are most often more accurate, to one or more target languages. We apply some consistency constraints to obtain more accurate annotations on both source and target side. Using our procedure we generated two new resources that can be used for evaluating CA-NMT models. One starting from the well-known TED Talk’s data released for the IWSLT17 shared task, where we project the annotation from English to target languages as diverse as French, German and Chinese. The second resource is derived from the WMT24 shared task, consisting of news domain data in the same set of target languages. We release these resources, as well as the code framework for applying our annotation procedure, to the community.
Flipper: An Extended Document-Level Financial Dataset for Training and Evaluation with Annotated Discourse Phenomena
Mariam Nakhlé | Rachel Atherly | Gabriela nicole Gonzalez Saez | Marco Dinarelli | Raheel Qader | Hervé Blanchon
The 7th Financial Narrative Processing Workshop
Mariam Nakhlé | Rachel Atherly | Gabriela nicole Gonzalez Saez | Marco Dinarelli | Raheel Qader | Hervé Blanchon
The 7th Financial Narrative Processing Workshop
We present a new resource for Machine Translation (MT), namely a training and evaluation dataset containing parallel sections issued from authentic documents in the financial domain. We cover five language pairs: English-French, English-Spanish, English-German, English-Italian and French-Spanish. The total number of parallel sections is 122k and the number of tokens is 118M (source and target combined). MT has improved greatly in recent years, but certain phenomena still cause errors, particularly when context spans beyond a single sentence. Errors can lead to mistranslated pronouns, incorrect gender or number agreement, and inconsistent terminology, which can be especially problematic in high-stakes domains like finance. We therefore construct the dataset at document level (rather than sentence-level alignment) and also produce fine-grained annotations of context-sensitive phenomena. The annotation was performed using preexisting tools and custom scripts. The annotated phenomena are: formality, gender, terminology consistency, verb form and sentence reordering. This aims to improve document-level evaluation of MT models by enabling evaluation solely on texts containing a particular phenomenon of interest. Our primary contribution is the creation and public release of Flipper, a multilingual document-level parallel dataset in the financial domain, designed to support both training and targeted evaluation of context-sensitive machine translation.