Mark Andrade
2026
Quality and Appropriateness of Large Text Datasets for Irish NLP
Abigail Walsh | Mark Andrade | Jane Lauren Adkins | Ornait O’Connell | Éanna O’Connor | Ellen Rushe | Brian Davis
Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages
Abigail Walsh | Mark Andrade | Jane Lauren Adkins | Ornait O’Connell | Éanna O’Connor | Ellen Rushe | Brian Davis
Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages
The value of high-quality datasets for training essential language tools has long been recognised for NLP research. Despite the importance of such datasets, most language data available for training consists of large, automatically curated corpora, often scraped from web content. The quality of such datasets is often an unknown factor. This presents a problem for already low-resourced languages (such as Irish), as existing datasets may not provide adequate, representative language data for training effective models. This paper examines existing monolingual and parallel Irish text corpora to evaluate the quality of the language data, through manual review, automatic metrics, and LLMs as judges.
LLMs as Assistants for Data Annotation: Addressing Disagreement and Supporting Expert Processes
Mark Andrade | Bláithín Heffernan | Abigail Walsh | Sheila Castilho
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Mark Andrade | Bláithín Heffernan | Abigail Walsh | Sheila Castilho
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
This paper investigates the potential of Large Language Models to assist human annotation pipelines, with a particular focus on supporting the development of expert-informed annotation guidelines for document-level content categorisation. We present three experiments exploring distinct roles for LLMs in annotation: as annotators, as domain experts assisting in disagreement resolution, and as analysts of annotator discussions. Using GPT-4.5 and Claude Sonnet 4, we evaluate LLM-generated annotation guidelines for a document-level classification tasks in terms of coverage, applicability, and usefulness. Preliminary results are mixed-to-positive, with evidence that LLMs can provide useful support across different stages of the annotation pipeline, particularly when supplied with rich contextual information such as prior human annotations and annotator discussions. However, their effectiveness remains sensitive to prompting strategies and input configuration.
LLM Multi-Agent Systems for Long Triple Set Data-to-Text Generation
Chinonso Cynthia Osuji | Simon Mille | Mark Andrade | Jane Adkins | Ornait O’Connell | Elaine Uí Dhonnchadha | Bláithín Heffernan | Fírinne Nic an tSaoir | Anya Belz | Thiago Castro Ferreira | Brian Davis
Findings of the Association for Computational Linguistics: ACL 2026
Chinonso Cynthia Osuji | Simon Mille | Mark Andrade | Jane Adkins | Ornait O’Connell | Elaine Uí Dhonnchadha | Bláithín Heffernan | Fírinne Nic an tSaoir | Anya Belz | Thiago Castro Ferreira | Brian Davis
Findings of the Association for Computational Linguistics: ACL 2026
Generating coherent, semantically accurate text from large structured inputs remains a persistent challenge in data-to-text generation, as single-step LLM mappings from data-to-text limit control over discourse structuring and amplify hallucinations and omissions as input size grows. We introduce a new dataset of extended DBpedia triple sets (up to 199 triples per input), and a modular multi-agent framework: specialised LLM agents handle content ordering, text structuring, and surface realisation under the supervision of an orchestrator and guardrail control loop. The system generates multi-paragraph outputs in English and Irish (low-resource). We compare a three-worker multi-agent configuration against a single-worker multi-task variant and a strong end-to-end baseline. Quality is assessed via human evaluation and LLM-as-a-judge (with truncation-based sanity checks). Results show slightly superior coherence for the multi-agent approach in both languages, with statistically significant inter-rater correlation over all criteria for English and no statistically significant correlation for Irish. Human-LLM alignment is very weak overall, thus exposing key limits in scalable NLG evaluation.
2025
eSTÓR: Curating Irish Datasets for Machine Translation
Abigail Walsh | Órla Ní Loinsigh | Jane Adkins | Ornait O’Connell | Mark Andrade | Teresa Clifford | Federico Gaspari | Jane Dunne | Brian Davis
Proceedings of Machine Translation Summit XX: Volume 2
Abigail Walsh | Órla Ní Loinsigh | Jane Adkins | Ornait O’Connell | Mark Andrade | Teresa Clifford | Federico Gaspari | Jane Dunne | Brian Davis
Proceedings of Machine Translation Summit XX: Volume 2
Minority languages such as Irish are massively under-resourced, particularly in terms of high-quality domain-relevant data, limiting the capabilities of machine translation (MT) engines, even those integrating large language models (LLMs). The eSTÓR project, described in this paper, focuses on the collection and curation of high-quality Irish text data for diverse domains.
Are Multi-Agents the new Pipeline Architecture for Data-to-Text Systems?
Chinonso Cynthia Osuji | Brian Timoney | Mark Andrade | Thiago Castro Ferreira | Brian Davis
Proceedings of the 18th International Natural Language Generation Conference
Chinonso Cynthia Osuji | Brian Timoney | Mark Andrade | Thiago Castro Ferreira | Brian Davis
Proceedings of the 18th International Natural Language Generation Conference
Large Language Models (LLMs) have achieved remarkable results in natural language generation, yet challenges remain in data-to-text (D2T) tasks, particularly in controlling output, ensuring transparency, and maintaining factual consistency with the input. We introduce the first LLM-based multi-agent framework for D2T generation, coordinating specialized agents to produce high-quality, interpretable outputs. Our system combines the reasoning and acting abilities of ReAct agents, the self-correction of Reflexion agents, and the quality assurance of Guardrail agents, all directed by an Orchestrator agent that assigns tasks to three specialists—content ordering, text structuring, and surface realization—and iteratively refines outputs based on Guardrail feedback. This closed-loop design enables precise control and dynamic optimization, yielding text that is coherent, accurate, and grounded in the input data. On a relatively simple dataset like WebNLG, our framework performs competitively with end-to-end systems, highlighting its promise for more complex D2T scenarios.