Fifteenth Language Resources and Evaluation Conference
Palma, Mallorca, SpainMay 11–16, 2026
Volumes
- Proceedings of the Fifteenth Language Resources and Evaluation Conference LREC 945 papers
- Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC) BUCC 13 papers
- Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026 LEGAL CALD-pseudo 13 papers
- Proceedings of Computational Affective Science (CAS) @ LREC 2026 CAS 22 papers
- Proceedings of the Third Workshop on Computation and Written Language (CAWL 2026) @ LREC 2026 CAWL 12 papers
- Proceedings of the Second workshop on Challenges in Processing South Asian Languages (CHiPSAL2026) CHiPSAL 34 papers
- Proceedings of the Third Workshop on Patient-Oriented Language Processing (CL4Health) @ LREC 2026 CL4Health 54 papers
- Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026 ClinicalNLP 42 papers
- Proceedings of the 15th Workshop on Cognitive Modeling and Computational Linguistics CMCL 27 papers
- Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora CMLC 20 papers
- Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages SIGUL EURALI DCLRL 32 papers
- Proceedings of The 2nd Workshop on Language-driven Deliberation Technology DELITE 5 papers
- Proceedings of the 2nd Workshop on Evaluating Text Difficulty in a Multilingual Context (DeTermIt! 2026) DeTermIt 10 papers
- Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective DialRes 35 papers
- Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026 DMR 18 papers
- Proceedings of Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science (DTF) @ LREC 2026 DTF 10 papers
- The 7th Financial Narrative Processing Workshop FNP 20 papers
- Proceedings fo the Second International Workshop on Eye-Tracking Resources and Evaluation for Human-Aligned NLP Gaze4NLP 11 papers
- Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes) htres 12 papers
- Proceedings of the Second Workshop of Identity Aware AI iaai 7 papers
- Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026 InDor 13 papers
- Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026 ISA 15 papers
- Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26 KaLLM 22 papers
- Proceedings of LANLP: Bridging Ibero and Latin American NLP Communities LANLP 9 papers
- Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026) LDL 11 papers
- Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026 LLMs4SSH 25 papers
- Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026 LT4HALA 51 papers
- Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026 NakbaNLP 48 papers
- Proceedings of the Workshop Neology and Large Language Models NeoLLM 9 papers
- Proceedings of the 2nd Workshop on Ecology, Environment, and Natural Language Processing NLP4Ecology 15 papers
- Proceedings of the the fifth edition of NLPerspectives NLPerspectives 14 papers
- Proceedings of the 1st Workshop on Social Context (SoCon) and the 2nd Workshop on Integrating NLP and Psychology to Study Social Interactions (NLPSI) @ LREC 2026 SoCon NLPSI 14 papers
- Proceedings of Learning Non-Literal Expressions with Small Data @ LREC 2026 NonLiteral 12 papers
- Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026 NSLP 30 papers
- The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks OSACT 44 papers
- Proceedings of the ParlaCLARIN V Workshop on Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora ParlaCLARIN 10 papers
- Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026) PoliticalNLP 31 papers
- Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers PressMint 15 papers
- Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026 RAIL 14 papers
- Proceedings of the Sixth Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with various forms of cognitive/psychiatric/developmental impairments in cooperation with the MENTAL.ai consortium RaPID 13 papers
- Proceedings of the Joint Workshop on Readability and Text Simplification (READIxTSAR) @ LREC 2026 READI TSAR 18 papers
- Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026) RESOURCEFUL 19 papers
- Proceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion SignLang 53 papers
- Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE) SLiDE 22 papers
- Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026 SPEAKABLE 22 papers
- Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026) UDW 30 papers
- Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation WILDRE 16 papers
up
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Stelios Piperidis | Núria Bel | Henk van den Heuvel | Nancy Ide | Simon Krek | Antonio Toral
Stelios Piperidis | Núria Bel | Henk van den Heuvel | Nancy Ide | Simon Krek | Antonio Toral
Beyond Generic Responses: Target-Aware Strategies for Countering Hate Speech
Yen-Yu Chang | Daryna Dementieva | Alexander Fraser
Yen-Yu Chang | Daryna Dementieva | Alexander Fraser
Effective counter-narratives (CNs) are essential for combating online hate speech, yet generic responses often fail to address the specific needs of targeted groups. This paper proposes a target-aware CN generation framework that incorporates demographic-specific tokens into transformer-based models. Our approach enhances the contextual relevance by introducing target-group tokens into the model’s vocabulary. To assess CN quality, we employ a multifaceted evaluation framework, including automatic metrics and LLM as Judges (JudgeLM). Evaluation with a wide range of language models demonstrates that target group tokens markedly improve contextual relevance of generated CN, particularly in small and medium models, with measurable gains in validity as CN and contextual relevance. Even for large instruction-tuned models, such as LLaMA-3, incorporating target-specific information proves effective in enhancing contextual relevance of generated responses. Warning: This paper contains offensive texts that are only used for combating online hate.
Topic-Initiator: A Proactive Chatbot with Personalized Topic RAG for Enhancing Willingness to Converse
Kazuya Matsuo | Atsushi Otsuka | Narichika Nomoto | Makoto Nakatsuji
Kazuya Matsuo | Atsushi Otsuka | Narichika Nomoto | Makoto Nakatsuji
Stimulating users’ conversational willingness to converse remains a major challenge in chatbot research. Most existing chatbots respond passively to user inputs, relying on users to select conversation topics, which often reduces their willingness. To address this issue, we propose, Topic-Initiator, a proactive chatbot that initiates conversations with new topics aligned to user interests. It gathers information from external sources (e.g., the web) to obtain potentially novel and engaging topics. To support this capability, we also introduce a novel Retrieval-Augmented Generation (RAG) framework, Personalized-Topic RAG (PT-RAG), designed to retrieve new and interesting topics for each user. Unlike existing RAG methods that fails to surface unseen information, PT-RAG leverages the inference capabilities of Large Language Models (LLMs) to identify content that matches the user’s interests but is not yet known to them. Specifically, PT-RAG estimates a user’s interests and knowledge from past interactions and organizes collected information into categories. Then, it uses an LLM to select a category that matches their interests and obtain information not seen in their knowledge from the selected category. Automatic and human evaluations demonstrate that PT-RAG retrieves new and interesting information more accurately and that Topic-Initiator significantly enhances users’ willingness to converse compared to existing methods.
CoachLah: A Singlish–English Parallel Corpus of Health Coaching Conversations with Behavior Goal Annotations
Iva Bojic | Mathieu Ravaut | Stephanie Hilary Xinyi Ma | Doreen Tan | Andy Hau Yan Ho | Andy Khong
Iva Bojic | Mathieu Ravaut | Stephanie Hilary Xinyi Ma | Doreen Tan | Andy Hau Yan Ho | Andy Khong
Health coaching (HC) aims to promote sustainable behavior change through goal-oriented dialogue, but research in this area is limited by the scarcity of authentic, transcript-based corpora. Existing datasets are small, English-only, and Western-centric, overlooking cultural and linguistic factors that shape real-world HC interactions. We introduce CoachLah, the first Singlish–English parallel corpus of HC conversations collected from a randomized controlled trial in Singapore. The dataset comprises 36,852 utterances transcribed from almost 160 hours of recorded HC sessions with 51 clients and 4 professional health coaches. Each dialogue is speaker-labeled, transcribed in Singlish, and aligned with high-quality English translations to preserve linguistic and cultural nuances. All sessions include HC summaries written by health coaches after each HC session, from which behavioral goals were manually annotated. To demonstrate the dataset’s utility, we benchmark two downstream tasks: (i) Singlish-to-English translation using fine-tuned open-weight models (e.g., Gemma-2-9B-it) with Low-Rank Adaptation, and (ii) behavioral goal extraction from unstructured HC summaries using span-based modeling (e.g., DeBERTa-v3-base). Together, these contributions establish the first culturally grounded benchmark for low-resource, goal-oriented dialogue research in HC. Both the code and the dataset are available at: https://github.com/IvaBojic/CoachLah.
Faithful Medical Dialogue Generation Using Homo-Heterogeneous Exemplar-based In-Context Knowledge Grounding
Priyanshu Priya | Hardik Goyal | Asif Ekbal
Priyanshu Priya | Hardik Goyal | Asif Ekbal
The growing reliance on tele-healthcare has heightened the demand for accessible and professional health support. Artificial Intelligence (AI)-assisted medical dialogue systems have emerged as key solutions, with Large Language Models (LLMs) advancing the generation tasks. However, their susceptibility to hallucination leads to inaccurate and unreliable information, posing major challenges. To address this, we propose a novel approach to mitigate hallucinations in LLMs by integrating external knowledge and in-context learning mechanisms for faithful medical dialogue generation (MDG). In particular, we devise an In-context Medical Knowledge-grounded Dialogue Generator (IMKDG), a novel plug-and-play retrieval-based framework that leverages external medical knowledge, in-context learning (ICL), and retrieval methods to enable LLMs to generate faithful responses, thereby enhancing their performance on the MDG task. We utilize large-scale medical knowledge based on the Unified Medical Language System (UMLS) to retrieve knowledge pertinent to the dialogue context. Further, to enhance the LLMs’ ICL capability for the MDG task, we propose the Homo-Heterogeneous Exemplar Selection (H2ES) method, a novel in-context exemplar retrieval method based on both dialogue context and medical knowledge. Automatic and human evaluations on the MedDialog-EN and CDialog datasets across various LLMs demonstrate the efficacy of the proposed framework in mitigating hallucinations.
Investigating Proactivity in Multimodal Task-Guidance Dialogues
Sofia Brenna | Elisabetta Jezek | Matthias Kraus | Bernardo Magnini
Sofia Brenna | Elisabetta Jezek | Matthias Kraus | Bernardo Magnini
While proactivity, i.e., the ability to take the initiative and anticipate requests in order to improve the effectiveness of a conversation, has been traditionally investigated in task-oriented dialogues (e.g., booking a restaurant), less work addresses proactive behaviours in task-guidance dialogues (e.g., guide to execute recipes), where the expert instructor is supposed to interact and supervise a user in a real-world setting. We analyse a corpus of video-recorded task-guided dialogues and explore two key features of proactivity in this context: (i) the impact of multimodal features, with respect to chat-based dialogues; (ii) the impact of instructions and actions grounded in a real situation. Through a comparison between task-oriented and task-guidance annotated dialogues, we find that task-guided dialogues are highly collaborative interactions, where preventing mistakes and maintaining the correct process order is essential for achieving the dialogue goal. In addition, the video information available in the task-guidance setting can be corrective for false positive proactive behaviours, although without introducing substantial differences. To support our analysis and to foster further research we provide a corpus of multimodal task-guidance dialogues annotated according to proactivity.
Investigating How LLMs Propagate Female Stereotypes: Comparing What Models Say via Prompts with What They Represent in Their Embeddings
Andrea Valderrey Nuñez | Jelke Bloem
Andrea Valderrey Nuñez | Jelke Bloem
As Large Language Models (LLMs) are increasingly deployed in sensitive domains, concerns about their encoding and reproduction of social bias have intensified. We examine how gender stereotypes are represented in embeddings and expressed in outputs across three models: BERT, base LLaMA-2-7b, and instruction-tuned LLaMA-2-7b-Chat. Focusing on seven female-oriented stereotype categories, we compare embedding-level bias using Directional Embedding Probing with output-level behavior measured via masked token prediction (BERT) and narrative prompt completions (LLaMA models). LLaMA-2-Chat showed the strongest representational–behavioral alignment, with female-aligned scores ranging from 60% to 100% and a significant point-biserial correlation (r = 0.55, p = 0.0008). BERT exhibited weaker alignment (0%–60%; r = 0.39, p = 0.054), while base LLaMA-2 showed intermediate but inconsistent patterns. These findings suggest that instruction tuning is associated with clearer alignment between internal representations and generated outputs, while prompt design plays a critical role in surfacing latent bias. The study contributes to fairness research by emphasizing the need to assess both internal representations and their behavioral expression in LLMs.
Why So Separate: Analyzing In-Context Learning from a Vector Space Perspective
Tobias Kalmbach | Sandipan Sikdar
Tobias Kalmbach | Sandipan Sikdar
In-context learning (ICL) is a popular prompting strategy for large language models. ICL allows models to learn tasks using demonstrative examples alone, without any weight updates or training. Nevertheless, it is still largely unclear why ICL works. In this paper, we investigate ICL from a new viewpoint, namely a vector space perspective, and extract insights for ICL from this analysis. In our experiments, we extract the hidden representations, i.e., embeddings, created by a large language model when passing an ICL prompt through it. We find that these embeddings generated by large language models are separable in the vector space when applying ICL. The degree of separability is dependent on the difficulty of the task, the size of the model and other factors, like the labels of demonstrative examples. We also find that, especially for large models, the separability is indicative of the classification performance. As an application, we utilize our findings to explain peculiarities of ICL and to select demonstrative examples for ICL. Experiments across multiple datasets show that this way of selecting examples consistently outperforms the commonly used random selection method.
Explaining Explanations: Interpretability Methods for Discourse Analysis of Transformer Attention Maps
Louis Escouflaire | Jérémie Bogaert | Antonin Descampe | Cédrick Fairon | Francois-Xavier Standaert
Louis Escouflaire | Jérémie Bogaert | Antonin Descampe | Cédrick Fairon | Francois-Xavier Standaert
While LLMs have achieved state-of-the-art performance in NLP, their opacity hinders a human understanding of their predictions. Standard explainability techniques often prioritize technical faithfulness over linguistic plausibility. This paper argues for an interdisciplinary approach that integrates discourse analysis to critically interpret model explanations. We conduct a case study using CamemBERT, fine-tuned to classify French journalistic texts as news or opinion. We employ Layer-wise Relevance Propagation to generate attention maps for 1,000 test articles and analyze the token-level relevance scores through both in-depth qualitative analysis and a quantitative ranking of high-attention tokens. Our findings reveal that CamemBERT successfully captures genre-specific linguistic markers: it attends to cues of reported speech and temporal anchors in news, and to expressive punctuation, evaluative adjectives, and first-person pronouns in opinion. The discourse-analytic lens moves us beyond superficial observations, demonstrating how the model interprets features like punctuation as structural or stylistic conventions. We argue that integrating linguistic expertise into the explainability pipeline yields more nuanced, human-readable explanations.
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
Yongxin Zhou | Philippe Mulhem | Didier Schwab
Yongxin Zhou | Philippe Mulhem | Didier Schwab
The evaluation of Retrieval-Augmented Generation (RAG) systems typically examines retrieval quality and generation parameters like temperature in isolation, overlooking their interaction. This work presents a systematic investigation of how text perturbations (simulating noisy retrieval) interact with temperature settings across multiple LLM runs. We propose a comprehensive RAG Perturbation-Temperature Analysis Framework that subjects retrieved documents to three distinct perturbation types across varying temperature settings. Through extensive experiments on HotpotQA with both open-source and proprietary LLMs, we demonstrate that performance degradation follows distinct patterns: high-temperature settings consistently amplify vulnerability to perturbations, while certain perturbation types exhibit non-linear sensitivity across the temperature range. Our work yields three key contributions: (1) a diagnostic benchmark for assessing RAG robustness, (2) an analytical framework for quantifying perturbation-temperature interactions, and (3) practical guidelines for model selection and parameter tuning under noisy retrieval conditions.
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
Iker García-Ferrero | David Montero | Roman Orus
Iker García-Ferrero | David Montero | Roman Orus
We introduce Refusal Steering, an inference-time method to exercise fine-grained control over Large Language Models refusal behaviour on politically sensitive topics without retraining. We replace fragile pattern-based refusal detection with an LLM-as-a-judge that assigns refusal confidence scores and we propose a ridge-regularized variant to compute steering vectors that better isolate the refusal–compliance direction. On Qwen3-Next-80B-A3B-Thinking, our method removes the refusal behaviour of the model around politically sensitive topics while maintaining safety on JailbreakBench and near-baseline performance on general benchmarks. The approach generalizes across 4B and 80B models and can also induce targeted refusals when desired. We analize the steering vectors and show that refusal signals concentrate in deeper layers of the transformer and are distributed across many dimensions. Together, these results demonstrate that activation steering can remove political refusal behaviour while retaining safety alignment for harmful content, offering a practical path to controllable, transparent moderation at inference time.
To Predict or Not to Predict? Towards Reliable Uncertainty Estimation in the Presence of Noise
Nouran Khallaf | Serge Sharoff
Nouran Khallaf | Serge Sharoff
This study examines the role of uncertainty estimation (UE) methods in multilingual text classification under noisy and non-topical conditions. Using a complex-vs-simple sentence classification task across several languages, we evaluate a range of UE techniques against a range of metrics to assess their quality. Results indicate that while methods relying on softmax outputs remain competitive in high-resource in-domain settings, their reliability declines in low-resource or domain-shift scenarios. In contrast, Monte Carlo dropout approaches demonstrate consistently strong performance across all languages, offering more robust calibration, stable decision thresholds, and greater discriminative power even under adverse conditions. We further demonstrate the positive impact of UE on non-topical classification: selectively abstaining from predicting the 10% most uncertain instances increases the macro F1 score from 0.81 to 0.85 in the Readme task. By integrating UE with trustworthiness metrics, this study provides actionable insights for developing more reliable NLP systems in real-world multilingual environments.
An Extreme Multi-label Text Classification (XMTC) Library Dataset: What If We Took "Use of Practical AI in Digital Libraries" Seriously?
Jennifer D’Souza | Sameer Sadruddin | Maximilian Kaehler | Andrea Salfinger | Luca Zaccagna | Francesca Incitti | Lauro Snidaro | Osma Suominen
Jennifer D’Souza | Sameer Sadruddin | Maximilian Kaehler | Andrea Salfinger | Luca Zaccagna | Francesca Incitti | Lauro Snidaro | Osma Suominen
Subject indexing is vital for discovery but hard to sustain at scale and across languages. We release a large bilingual (English/German) corpus of catalog records annotated with the Integrated Authority File (GND), plus a machine-actionable GND taxonomy. The resource enables ontology-aware multi-label classification, mapping text to authority terms, and agent-assisted cataloging with reproducible, authority-grounded evaluation. We provide a brief statistical profile and qualitative error analyses of three systems. We invite the community to assess not only accuracy but usefulness and transparency, toward authority-anchored AI co-pilots that amplify catalogers’ work.
A Historical Database for the Study of Obstruent-Lateral Palatalization in Ibero-Romance
Andrea García Covelo
Andrea García Covelo
Studying irregular sound changes requires documenting not only words that underwent the change but also those that did not. Obstruent-lateral (OL) palatalization in Ibero-Romance, i.e., Galician, Portuguese, and Spanish, is one such change, exhibiting three distinctive patterns: unusual distribution (/pl fl kl/ typically palatalized but /bl gl/ rarely did), irregular implementation (not all eligible words underwent palatalization), and variable outcomes (dependent on obstruent voicing and cluster word position). This paper presents a cross-linguistic historical dataset of 659 inherited words from principally Galician, Portuguese, and Spanish, with and without palatalization, traceable to etyma containing OL clusters. The dataset draws on etymological dictionaries, philological works, and historical corpora. A digitalized version of the Diccionario Crítico Etimológico Castellano e Hispánico (Corominas and Pascual, 2012) served as the backbone for systematically identifying etyma containing OL clusters. The compiled corpus contains 473 words with certain etymologies and comparable coverage across the three languages. By providing the first comprehensive compilation of both palatal and non-palatal historical evidence, this dataset enables the systematic study of OL palatalization in Ibero-Romance.
Is Clinical Text Enough? A Multimodal Study on Mortality Prediction in Heart Failure Patients
Oumaima El Khettari | Virgile Barthet | Guillaume Hocquet | Joconde Weller | Emmanuel Morin | Pierre Zweigenbaum
Oumaima El Khettari | Virgile Barthet | Guillaume Hocquet | Joconde Weller | Emmanuel Morin | Pierre Zweigenbaum
Accurate short-term mortality prediction in heart failure (HF) remains challenging, particularly when relying on structured electronic health record (EHR) data alone. We evaluate transformer-based models on a French HF cohort, comparing text-only, structured-only, multimodal, and LLM-based approaches. Our results show that enriching clinical text with entity-level representations improves prediction over CLS embeddings alone, and that supervised multimodal fusion of text and structured variables achieves the best overall performance. In contrast, large language models perform inconsistently across modalities and decoding strategies, with text-only prompts outperforming structured or multimodal inputs. These findings highlight that entity-aware multimodal transformers offer the most reliable solution for short-term HF outcome prediction, while current LLM prompting remains limited for clinical decision support.
HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)
Aurelien Pellet | Marie Anna Puren | Julien PEREZ
Aurelien Pellet | Marie Anna Puren | Julien PEREZ
We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, the corpus captures complex reasoning patterns typical of historical inquiry, including cross-source synthesis, temporal reasoning, and the integration of sparse evidence. The dataset is made of 1782 questions and emphasizes multi-hop connections across heterogeneous historical documents, providing a resource for evaluating retrieval-augmented and large language model systems in domain-specific contexts. We describe the methodology for constructing the corpus, including the selection and alignment of sources, question validation, and metadata integration. While the dataset focuses on French historical documents, our methodology can be readily adapted to other languages and national corpora. Finally, we demonstrate how the corpus can support realistic evaluation scenarios for multi-hop question answering, bridging the gap between NLP benchmarks and the needs of historical scholarship.
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
Thura Aung | Jann Railey Montalan | Jian Gang Ngui | Peerat Limkonchotiwat
Thura Aung | Jann Railey Montalan | Jian Gang Ngui | Peerat Limkonchotiwat
We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU), reasoning (NLR), and generation (NLG). BURMESE-SAN consolidates seven subtasks spanning these competencies, including Question Answering, Sentiment Analysis, Toxicity Detection, Causal Reasoning, Natural Language Inference, Abstractive Summarization, and Machine Translation, several of which were previously unavailable for Burmese. The benchmark is constructed through a rigorous native-speaker-driven process to ensure linguistic naturalness, fluency, and cultural authenticity while minimizing translation-induced artifacts. We conduct a large-scale evaluation of both open-weight and commercial LLMs to examine challenges in Burmese modeling arising from limited pretraining coverage, rich morphology, and syntactic variation. Our results show that Burmese performance depends more on architectural design, language representation, and instruction tuning than on model scale alone. In particular, Southeast Asia regional fine-tuning and newer model generations yield substantial gains. Finally, we release BURMESE-SAN as a public leaderboard to support systematic evaluation and sustained progress in Burmese and other low-resource languages. https://leaderboard.sea-lion.ai/detailed/MY
Assessing the Political Fairness of Multilingual LLMs: A Case Study Based on a 21-Way Multiparallel EuroParl Dataset
Paul Lerner | François Yvon
Paul Lerner | François Yvon
The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying on principles of fairness in multilingual translation. We systematically compare the translation quality of speeches in the European Parliament (EP), observing systematic differences with majority parties from left and right being better translated than outsider parties. This study is made possible by a new, 21-way multiparallel version of EuroParl, the parliamentary proceedings of the EP, which includes the political affiliations of each speaker. The dataset consists of 1.5M sentences for a total of 40M words and 249M characters. It covers three years, 1000+ speakers, 7 countries, 12 EU parties, 25 EU committees, and hundreds of national parties.
AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
Yann Le Beux | Oluchi Audu | Oche David Ankeli | Dhananjay Balakrishnan | Melissah Weya | Marie Daniella Ralaiarinosy | Ignatius Ezeani
Yann Le Beux | Oluchi Audu | Oche David Ankeli | Dhananjay Balakrishnan | Melissah Weya | Marie Daniella Ralaiarinosy | Ignatius Ezeani
Existing AI bias evaluation benchmarks largely reflect Western perspectives, leaving African contexts underrepresented and enabling harmful stereotypes in applications across various domains. To address this gap, we introduce AfriStereo, the first open-source African stereotype dataset and evaluation framework grounded in local socio-cultural contexts. Through community engaged efforts across Senegal, Kenya, and Nigeria, we collect 1,163 stereotypes spanning gender, ethnicity, religion, age, and profession. Using few-shot prompting with human-in-the-loop validation, we augment the dataset to over 5,000 stereotype–antistereotype pairs. Entries are validated through semantic clustering and manual annotation by culturally informed reviewers. Preliminary evaluation of language models reveals that nine of eleven models exhibit statistically significant bias in our setup, with Bias Preference Ratios (BPR) ranging from 0.63 to 0.78 (p ≤ 0.05), indicating systematic preferences for stereotypes over antistereotypes, particularly across age, profession, and gender dimensions. Domain-specific models appear to show weaker bias in our setup, suggesting task-specific training may mitigate some associations. Looking ahead, AfriStereo opens pathways for future research on culturally grounded bias evaluation and mitigation, offering key methodologies for the AI community on building more equitable, context-aware, and globally inclusive NLP technologies.
Judging Instruction Responses in a Low-Resource Language: A Case Study on Basque
David Ponce | Harritxu Gete | Thierry Etchegoyhen | Irune Zubiaga | Aitor Soroa
David Ponce | Harritxu Gete | Thierry Etchegoyhen | Irune Zubiaga | Aitor Soroa
Evaluating the quality of answers to a given instruction is a demanding and time-consuming task, limiting the scalability of human assessment. Large language models (LLMs) have been proposed as automatic judges to reduce this effort, but their reliability in low-resource contexts remains uncertain. Additionally, the premise that humans are reliable judges of fine-grained response quality needs to be assessed as well, if correlation with automated judges on this task is to be considered a gold standard. In this work, we investigate the performance of various LLM-as-a-judge in a low-resource scenario, namely Basque, and evaluate its correlation with human judgements. Additionally, we measure the agreement between human judgments themselves, to assess their viability as a valid reference. To perform our experiments, we translated and manually post-edited the Just-Eval benchmark, a suite of benchmarks tackling fine-grained aspects of response quality. We also extend the evaluation with a novel category aimed at judging both language consistency and grammaticality. Our results show that state of the art models exhibit fairly poor correlations with humans and amongst themselves, calling for the development of dedicated LLM-as-a-judge models for this language.
Appeal, Align, Divide? Stance Detection for Group-Directed Messages in German Parliamentary Debates
Ines Rehbein | Maris Leander Buttmann | Julian Schlenker | Simone Paolo Ponzetto
Ines Rehbein | Maris Leander Buttmann | Julian Schlenker | Simone Paolo Ponzetto
This paper presents a new benchmark for detecting group-based appeals, i.e., positive or negative references towards social groups, in German parliamentary debates. In the first step, group mentions are identified as targets for stance detection. In the next step, three human annotators assign stance labels to the group mentions, coding the speaker’s perspective towards the specific group. The created benchmark data is then used to investigate the capacity of Large Language Models (LLMs) for detecting polticians’ stances towards social groups. We explore the potential of different prompting strategies (zero-shot prompting, few-shot prompting, Chain-of-Thought) for this task and compare the results to a supervised BERT baseline, showing that in low-resource scenarios LLMs can outperform smaller fine-tuned models without the need for annotating large datasets.
Report-based Recommendations for Policy Making and Agency Operations: Dataset and LLM Evaluation
Aleksandra Edwards | Thomas Edwards | Jose Camacho-Collados | Alun Preece
Aleksandra Edwards | Thomas Edwards | Jose Camacho-Collados | Alun Preece
Large Language Models (LLMs) are extensively used in text generation tasks. These generative capabilities bring us to a point where LLMs could potentially provide useful insights in policy making or agency operations. In this paper, we introduce a new task consisting of generating recommendations which can be used to inform future actions and improvements of agencies work within private and public organisations. In particular, we present the first benchmark and coherent evaluation for developing recommendation systems to inform organisation policies. This task is clearly different from usual product or user recommendation systems, but rather aims at providing a basis to suggest policy improvements based on the conclusions drawn from reports. Our results demonstrate that state-of-the-art LLMs have the potential to emphasize and reflect on key issues and learning points within generated recommendations.
ConceptKT: A Benchmark for Concept-Level Deficiency Prediction in Knowledge Tracing
Yu-Chen Kang | Yu-Chien Tang | An-Zi Yen
Yu-Chen Kang | Yu-Chien Tang | An-Zi Yen
Knowledge Tracing (KT) is a critical technique for modeling student knowledge to support personalized learning. However, most KT systems focus on binary correctness prediction and cannot diagnose the underlying conceptual misunderstandings that lead to errors. Such fine-grained diagnostic feedback is essential for designing targeted instruction and effective remediation. In this work, we introduce the task of concept-level deficiency prediction, which extends traditional KT by identifying the specific concepts a student is likely to struggle with on future problems. We present ConceptKT, a dataset annotated with labels that capture both the concepts required to solve each question and the missing concepts underlying incorrect responses. We investigate in-context learning approaches to KT and evaluate the diagnostic capabilities of various Large Language Models (LLMs) and Large Reasoning Models (LRMs). Different strategies for selecting informative historical records are explored. Experimental results demonstrate that selecting response histories based on conceptual alignment and semantic similarity leads to improved performance on both correctness prediction and concept-level deficiency identification.
Open-access Dataset on Acceptability Ratings of Korean Clausal Constructions by Humans and GPT Models
Gyu-Ho Shin | Soo-Hwan Lee | Chanyoung Lee
Gyu-Ho Shin | Soo-Hwan Lee | Chanyoung Lee
The present study introduces a new, open-access dataset on acceptability ratings of Korean clausal constructions at the morphosyntax–semantics interface (dative, passive, and negative polarity item). The dataset comprises (i) linguistically controlled sentence materials, (ii) ratings from targeted adult populations (individuals in their 20s), and (iii) parallel ratings from GPT variants (including ChatGPT). Alongside the release, we assess the alignment between GPT- and human-derived ratings to probe the extent to which GPT architectures can approximate patterns of human sentence comprehension.
Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
Frederik Yannick Broy | Maike Züfle | Jan Niehues
Frederik Yannick Broy | Maike Züfle | Jan Niehues
Scientific talks are a growing medium for disseminating research, and automatically identifying relevant literature that grounds or enriches a talk would be highly valuable for researchers and students alike. We introduce Reference Prediction from Talks (RPT), a new task that maps long, and unstructured scientific presentations to relevant papers. To support research on RPT, we present Talk2Ref, the first large-scale dataset of its kind, containing 6,279 talks and 43,429 cited papers (26 per talk on average), where relevance is approximated by the papers cited in the talk’s corresponding source publication. We establish strong baselines by evaluating state-of-the-art text embedding models in zero-shot retrieval scenarios, and propose a dual-encoder architecture trained on Talk2Ref. We further explore strategies for handling long transcripts, as well as training for domain adaptation. Our results show that fine-tuning on Talk2Ref significantly improves citation prediction performance, demonstrating both the challenges of the task and the effectiveness of our dataset for learning semantic representations from spoken scientific content. The dataset and trained models are released under an open license to foster future research on integrating spoken scientific communication into citation recommendation systems.
MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
Aaron Robert Scott | Maike Züfle | Jan Niehues
Aaron Robert Scott | Maike Züfle | Jan Niehues
Sarcasm is a complex form of figurative language in which the intended meaning contradicts the literal one. Its prevalence in social media and popular culture poses persistent challenges for natural language understanding, sentiment analysis, and content moderation. With the emergence of multimodal large language models, sarcasm detection extends beyond text and requires integrating cues from audio and vision. We present MuSaG, the first German multimodal sarcasm detection dataset, consisting of 33 minutes of manually selected and human-annotated statements from German television shows. Each instance provides aligned text, audio, and video modalities, annotated separately by humans, enabling evaluation in unimodal and multimodal settings. We benchmark nine open-source and commercial models, spanning text, audio, vision, and multimodal architectures, and compare their performance to human annotations. Our results show that while humans rely heavily on audio in conversational settings, models perform best on text. This highlights a gap in current multimodal models and motivates the use of MuSaG for developing models better suited to realistic scenarios. We release MuSaG publicly to support future research on multimodal sarcasm detection and human–model alignment.
Icelandic Math Eval: A Competitive Mathematics Benchmark for Large Language Models
Hafsteinn Einarsson | Jökull Ari Haraldsson | Ívar Armin Derayat | Sigrún Helga Lund | Benedikt Steinar Magnússon
Hafsteinn Einarsson | Jökull Ari Haraldsson | Ívar Armin Derayat | Sigrún Helga Lund | Benedikt Steinar Magnússon
We introduce Icelandic Math Eval, the first comprehensive benchmark for evaluating large language models (LLMs) on competitive mathematics problems in Icelandic. Our dataset comprises 1,027 problems from Icelandic mathematics competitions spanning from 1984 to 2025, covering algebra, geometry, number theory, and combinatorics across ten difficulty levels. We evaluate three state-of-the-art models, Claude Sonnet 4.5, Gemini 2.5 Pro, and GPT-5, using a dual evaluation methodology that tests both with and without multiple-choice options. Our results reveal several key findings: (1) models achieve 81-93% overall accuracy, demonstrating substantial cross-lingual transfer of mathematical reasoning capabilities; (2) a dramatic 17.5 percentage point performance drop on problems containing images highlights persistent challenges in multimodal mathematical reasoning; (3) a 6.7 percentage point gap between evaluation modes suggests that multiple-choice formats may overestimate genuine reasoning capabilities; and (4) systematic performance degradation with increasing difficulty, dropping to 43% on the most challenging problems. Using an LLM-as-judge evaluation approach, we provide detailed analysis across problem types, difficulty levels, and model capabilities. This work contributes to multilingual AI evaluation and demonstrates the importance of developing rigorous benchmarks for diverse languages to ensure comprehensive assessment of AI capabilities.
As Large Language Models (LLMs) increasingly power autonomous agents in robotics and embodied AI, understanding their spatial reasoning capabilities becomes crucial for reliable deployment. We introduce MazeEval, a benchmark designed to evaluate pure spatial reasoning in LLMs through coordinate-based maze navigation tasks without visual input. Using a function-calling interface, models navigate mazes of varying complexity (5 x 5 to 15 x 15 grids) using only coordinate feedback and distance-to-wall information. We evaluate eight state-of-the-art LLMs across identical mazes in both English and Icelandic to assess cross-linguistic transfer of spatial abilities. Our findings reveal striking disparities: while OpenAI’s O3 achieves perfect navigation up to 30 x 30 mazes, other models exhibit catastrophic failure beyond 9 x 9 mazes, with 100% of failures attributed to excessive looping behavior. We document significant performance degradation in Icelandic, with models solving mazes 3-4 sizes smaller than in English, suggesting spatial reasoning emerges from linguistic patterns rather than language-agnostic mechanisms. These results highlight that spatial intelligence remains fundamentally constrained by training data availability, with important implications for global deployment of LLM-powered autonomous systems.
J-ClinicalBench: A Benchmark for Evaluating Large Language Models on Practical Clinical Tasks in Japanese
Seiji Shimizu | Tomohiro Nishiyama | HISADA Shohei | Yamato Himi | Shoko Wakamiya | Yuki Yanagisawa | Masami Tsuchiya | Satoko Hori | Eiji ARAMAKI
Seiji Shimizu | Tomohiro Nishiyama | HISADA Shohei | Yamato Himi | Shoko Wakamiya | Yuki Yanagisawa | Masami Tsuchiya | Satoko Hori | Eiji ARAMAKI
Recent advances in large language models (LLMs) have accelerated the NLP applications in the medical and clinical domains. However, evaluations remain limited for non-English languages, such as Japanese, where clinical corpora are particularly scarce. To address this gap, we present J-ClinicalBench, a publicly available benchmark designed to reflect realistic Japanese clinical tasks. We first created 227 expert-authored clinical documents and newly constructed five datasets for core clinical tasks. Building on these datasets, J-ClinicalBench comprises nine clinical tasks spanning clinical language reasoning, generation, and understanding. We establish baseline performance on J-ClinicalBench by evaluating state-of-the-art proprietary and Japanese open-source LLMs, providing the first assessment of their utility in practical clinical scenarios. By releasing this benchmark, we aim to foster the development and evaluation of clinically applicable LLMs in Japanese healthcare, bridging the current gap between clinical NLP research and clinical practice.
Is One Dataset Enough for Evaluation? Studying Generalizability of Automated Essay Scoring Models
Sohaila Eltanbouly | Marwan Sayed | Tamer Elsayed
Sohaila Eltanbouly | Marwan Sayed | Tamer Elsayed
Automated Essay Scoring (AES) has made significant advancements in writing assessment. Recently, cross-prompt AES has gained attention because of its focus on generalizing to unseen prompts. Despite the promise of these advancements, a critical question remains: how generalizable and robust are those models when applied to diverse datasets? This study assesses the generalizability of eight cross-prompt AES models across three different datasets. We employ two experimental setups: the within-dataset approach, where both training and testing occur on the same dataset, and the cross-dataset approach, which challenges the models by evaluating their performance on previously unseen datasets. The experimental results show significant performance inconsistencies, highlighting that relying on a single dataset is insufficient for building robust and generalizable AES systems.
HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings
Rasmus T. Aavang | Giovanni Rizzi | Rasmus Tjalk-Bøggild | Alexandre Iolov | Mike Zhang | Johannes Bjerva
Rasmus T. Aavang | Giovanni Rizzi | Rasmus Tjalk-Bøggild | Alexandre Iolov | Mike Zhang | Johannes Bjerva
Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandated for public financial filings. Yet, its complex, fine-grained taxonomy limits the cross-company transferability of tagged Key Performance Indicators (KPIs). To address this, we introduce the Hierarchical Financial Key Performance Indicator (HiFi-KPI) dataset, a large-scale corpus of 1.65M paragraphs and 198k unique, hierarchically organized labels linked to iXBRL taxonomies. HiFi-KPI supports multiple tasks and we evaluate three: KPI classification, KPI extraction, and structured KPI extraction. For rapid evaluation, we also release HiFi-KPI-Lite, a manually curated 2.5K-instance subset. Baselines on HiFi-KPI-Lite show that encoder-based models achieve over 0.906 macro-F1 on classification, while Large Language Models (LLMs) reach 0.440 F1 on structured extraction. Finally, a qualitative analysis reveals that extraction errors primarily relate to dates. We open-source all code and data at Anonymous.
UniSkill: A Dataset for Matching University Curricula to Professional Competencies
Nurlan Musazade | József Mezei | Mike Zhang
Nurlan Musazade | József Mezei | Mike Zhang
Skill extraction and recommendation systems have been studied from recruiter, applicant, and education perspectives. While AI applications in job advertisements have received broad attention, deficiencies in the instructed skills side remain a challenge. In this work, we address the scarcity of publicly available datasets by releasing both manually annotated and synthetic datasets of skills from the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy and university course pairs and publishing corresponding annotation guidelines. Specifically, we match graduate-level university courses with skills from the Systems Analysts and Management and Organization Analyst ESCO occupation groups at two granularities: course title with a skill, and course sentence with a skill. We train language models on this dataset to serve as a baseline for retrieval and recommendation systems for course-to-skill and skill-to-course matching. We evaluate the models on a portion of the annotated data. Our BERT model achieves 87% F1-score, showing that course and skill matching is a feasible task.
A Dataset for Evaluating ASR on Specialized Vocabulary
Emily Haubert Klering | Eduardo Gabriel Cortes | Tatjana Chernenko | Mariana Vargas Trarbach | Gabriel de Oliveira Ramos | Sandro José Rigo | Maitê Dupont | Ana Luiza Treichel Vianna | Gabriela Krause dos Santos | Vinicius Meirelles Pereira | Denis Andrei de Araujo | Rafael Kunst
Emily Haubert Klering | Eduardo Gabriel Cortes | Tatjana Chernenko | Mariana Vargas Trarbach | Gabriel de Oliveira Ramos | Sandro José Rigo | Maitê Dupont | Ana Luiza Treichel Vianna | Gabriela Krause dos Santos | Vinicius Meirelles Pereira | Denis Andrei de Araujo | Rafael Kunst
Evaluating the ability of Automatic Speech Recognition (ASR) models to transcribe specialized vocabulary remains a persistent challenge, as standard datasets predominantly feature common words and thus obscure weaknesses on rare or out-of-vocabulary (OOV) terms. To address this limitation, we introduce a linguistically curated bilingual dataset (English and Portuguese) comprising 13,846 utterances (18.7 hours) distributed across synthetic and literature-derived subsets, with OOV rates reaching up to 100%. We further propose a diagnostic evaluation framework that partitions recognition performance into Biased Word Error Rate (B-WER), targeting domain-specific jargon, and Unbiased Word Error Rate (U-WER), focusing on general vocabulary. Baseline evaluations using Whisper models (medium, large-v3, and large-v3-turbo) confirm the necessity of this framework. On the most challenging datasets, B-WER reaches 0.88–0.90, whereas U-WER remains as low as 0.06–0.19, demonstrating that conventional WER masks critical failure modes in jargon recognition. Additionally, an oracle upper bound experiment shows that providing correct jargon via prompting reduces B-WER by 0.50–0.70 absolute, quantifying the considerable potential for contextual biasing. We release the datasets and evaluation scripts as a reproducible benchmark to foster research on domain-aware contextual biasing and OOV handling in ASR systems.
SommBench: Assessing Sommelier Expertise of Language Models
William Brach | Tomas Bedej | Jacob Nielsen | Jacob Pichna | Juraj Bedej | Eemeli Saarensilta | Julie Dupouy | Gianluca Barmina | Andrea Blasi Núñez | Peter Schneider-Kamp | Kristian Košťál | Michal Ries | Lukas Galke Poech
William Brach | Tomas Bedej | Jacob Nielsen | Jacob Pichna | Juraj Bedej | Eemeli Saarensilta | Julie Dupouy | Gianluca Barmina | Andrea Blasi Núñez | Peter Schneider-Kamp | Kristian Košťál | Michal Ries | Lukas Galke Poech
With the rapid advances of large language models, it becomes increasingly important to systematically evaluate their multilingual and multicultural capabilities. Previous cultural evaluation benchmarks focus mainly on basic cultural knowledge that can be encoded in linguistic form. Here, we propose SommBench, a multilingual benchmark to assess sommelier expertise, a domain deeply grounded in the senses of smell and taste. While language models learn about sensory properties exclusively through textual descriptions, SommBench tests whether this textual grounding is sufficient to emulate expert-level sensory judgment. SommBench comprises three main tasks: Wine Theory Question Answering (WTQA), Wine Feature Completion (WFC), and Food-Wine Pairing (FWP). SommBench is available in multiple languages: English, Slovak, Swedish, Finnish, German, Danish, Italian, and Spanish. This helps separate a language model’s wine expertise from its language skills. The benchmark datasets were developed in close collaboration with a professional sommelier and native speakers of the respective languages, resulting in 1,024 questions for wine theory question answering, 1,000 examples for wine feature completion, and 1,000 examples of food-wine pairing. We provide results for the most popular language models, including closed-weights models such as Gemini 2.5, and open-weights models, such as GPT-OSS and Qwen 3. Our results show that the most capable models perform well on wine theory question answering (up to 97% correct with a closed-weights model), yet feature completion (peaking at 65%) and food-wine pairing show (MCC ranging between 0 and 0.39) turn out to be more challenging. These results position SommBench as an interesting and challenging benchmark for evaluating the sommelier expertise of language models. The benchmark is publicly available at https://github.com/sommify/sommbench.
CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia
Josef Jon | Ondřej Bojar
Josef Jon | Ondřej Bojar
We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia—primarily Ukrainian and English, with smaller portions of Vietnamese, Russian and other languages. The dataset is designed to support the evaluation of machine translation systems that aim to preserve document formatting during translation. We provide a comparison of the most common approaches to format-preserving machine translation on a validation subset of the dataset. This validation split, together with the evaluation toolkit, is publicly released for further research. A held-out test split will be reserved for a future shared task focused on document-level translation with formatting preservation.
An LLM-Based Assistant for Debt Waiver Court Procedures
Lluis Padro | Daniel Ferrés | Roser Saurí | Mireia Artigot
Lluis Padro | Daniel Ferrés | Roser Saurí | Mireia Artigot
Spanish Insolvency Law 1/2020 of the 5th of May enables individuals to apply for debt waiver under certain conditions. The large number of applications submitted each year places a heavy burden on judges and court officers, who must examine heterogeneous documentation before issuing a ruling. This paper presents an AI-based assistant designed to support the processing of debt waiver cases. The system integrates PDF-to-text conversion, rule-based document classification, large language model (LLM)-based information extraction, and post-processing to consolidate fragmented or duplicated records. A front-end interface provides structured summaries of the application content, and can automatically generate draft rulings. Evaluated on a set of real applications, the system achieves over 92% F1 in document classification and up to 91% F1 in personal data extraction, showing the potential of open-source LLMs to reduce administrative workload and accelerate judicial procedures, while keeping the final decision with the judge.
Enhancing Clinical Trial Analysis through Large Language Models for Multi-Evidence Natural Language Inference
Shobanapriyan Chandrasegaran | Amal Htait
Shobanapriyan Chandrasegaran | Amal Htait
The exponential growth of clinical trial reports (CTRs) presents a critical challenge for evidence-based medicine, with manual systematic reviews requiring months to synthesise findings. This paper evaluates Large Language Models (LLMs) and retrieval methods for automated Natural Language Inference (NLI) and evidence extraction from CTRs, and seeks to improve upon previously reported results in this domain. Using the NLI4CT dataset containing 2,400 annotated statement-evidence pairs from breast cancer trials, we conducted a comparative evaluation of general-purpose LLMs, domain-specific LLMs, and transformer-based baselines across entailment classification and evidence retrieval tasks. Reasoning-capable, general-purpose LLMs (such as Qwen-32B) demonstrated superior performance in the entailment classification task, exceeding both the performance of other models evaluated in this study and the previously reported state-of-the-art results. Although domain-specific adaptations showed improvements at comparable scale, larger general-purpose language models maintained superior absolute performance. For evidence retrieval, Large Language embedding models (such as bge-large-en-v1.5) surpassed classic transformer-based ranking approaches. These findings demonstrate that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical decision-making.
A Systematic Comparison of Large Language Models for Data Annotation in NER Tasks
Muhammad Uzair Ul Haq | Davide Rigoni | Alessandro Sperduti
Muhammad Uzair Ul Haq | Davide Rigoni | Alessandro Sperduti
High-quality annotated data is essential for training effective machine learning models, especially for fine-grained tasks like Named Entity Recognition (NER), where each token in a sentence must be tagged with a golden annotation. While Large Language Models (LLMs) show strong potential in automating data annotation, existing literature lacks extensive evaluations that systematically compare different models, embedding strategies, and context selection methods, particularly on complex, real-world datasets. This paper fills this gap by conducting a comprehensive study of LLMs for NER annotation across four diverse datasets. It benchmarks both proprietary and open-source LLMs at the 7B to 70B parameter scale, including a 32B reasoning-optimized model, and explores multiple context selection strategies. Two evaluations are performed: (i) the assessment of the practical utility of LLM-generated annotations by fine-tuning a RoBERTa model on LLM-generated annotations and measuring downstream performance; (ii) the assessment of only LLM-generated annotations using token-level metrics, like Precision, Recall, F1, and agreement with human annotations (Cohen’s κ). Empirical results, supported by statistical tests, highlight the importance of choosing suitable LLMs and embedding models and reveal key trade-offs between model scale and annotation quality. Challenging datasets like SKILLSPAN further expose the limitations of current LLM-based annotation pipelines, emphasizing the need for benchmarking on difficult, real-world tasks.
Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations
Dang Hai Dang | Jelena Mitrović | Michael Granitzer
Dang Hai Dang | Jelena Mitrović | Michael Granitzer
We study whether large-scale unlabelled web data and LLM-based synthetic annotations can improve multilingual hate speech detection. Starting from texts crawled via OpenWebSearch (OWS) in four languages (English, German, Spanish, Vietnamese), we pursue two complementary strategies. First, we apply continued pre-training to BERT models by continuing masked language modelling on unlabelled OWS texts before supervised fine-tuning, and show that this yields an average macro-F1 gain of approximately 3% over standard baselines across sixteen benchmarks, with stronger gains in low-resource settings. Second, we use four open-source LLMs (Mistral-7B, Llama3.1-8B, Gemma2-9B, Qwen2.5-14B) to produce synthetic annotations through three ensemble strategies: mean averaging, majority voting, and a LightGBM meta-learner. The LightGBM ensemble consistently outperforms the other strategies. Fine-tuning on these synthetic labels substantially benefits a small model Llama3.2-1B: +11% pooled F1), but provides only a modest gain for the larger Qwen2.5-14B (+0.6%). Our results indicate that the combination of web-scale unlabelled data and LLM-ensemble annotations is most valuable for smaller models and low-data languages.
Can LLMs Faithfully Explain Themselves in Low-Resource Languages? A Case Study on Emotion Detection in Persian
Mobina Mehrazar | Mohammad Amin Yousefi | Parisa Beygi | Behnam Bahrak
Mobina Mehrazar | Mohammad Amin Yousefi | Parisa Beygi | Behnam Bahrak
Large language models (LLMs) are increasingly used to generate self-explanations alongside their predictions, a practice that raises concerns about the faithfulness of these explanations, especially in low-resource languages. This study evaluates the faithfulness of LLM-generated explanations in the context of emotion classification in Persian, a low-resource language, by comparing the influential words identified by the model against those identified by human annotators. We assess faithfulness using confidence scores derived from token-level log-probabilities. Two prompting strategies, differing in the order of explanation and prediction (Predict-then-Explain and Explain-then-Predict), are tested for their impact on explanation faithfulness. Our results reveal that while LLMs achieve strong classification performance, their generated explanations often diverge from faithful reasoning, showing greater agreement with each other than with human judgments. These results highlight the limitations of current explanation methods and metrics, emphasizing the need for more robust approaches to ensure LLM reliability in multilingual and low-resource contexts.
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
Hawau Olamide Toyin | Samar Mohamed Magdy | Hanan Aldarmaki
Hawau Olamide Toyin | Samar Mohamed Magdy | Hanan Aldarmaki
We investigate the effectiveness of large language models (LLMs) for text diacritization in two typologically distinct languages: Arabic and Yoruba. To enable a rigorous evaluation, we introduce a novel multilingual dataset MultiDiac, with diverse samples that capture a range of diacritic ambiguities. We evaluate 12 LLMs varying in size, accessibility, and language coverage, and benchmark them against 4 specialized diacritization models. Additionally, we fine-tune four small open-source models using LoRA for Yoruba. Our results show that many off-the-shelf LLMs outperform specialized diacritization models for both Arabic and Yoruba, but smaller models suffer from hallucinations. We find that fine-tuning on a small dataset can help improve diacritization performance and reduce hallucination rates for Yoruba.
Automatic Suggestions of Supplements in the Herculaneum Papyri: Language Models and RESTful API
Angelo Mario Del Grosso | Gabriele Giannessi | Simone Zenzaro | Federico Boschetti
Angelo Mario Del Grosso | Gabriele Giannessi | Simone Zenzaro | Federico Boschetti
This paper addresses a computational philology task focused on the automatic restoration of textual gaps (i.e., lacunae) in the Herculaneum Papyri, whose Ancient Greek texts are inherently fragmentary due to damage caused by carbonization. The objective of this work is to show the preliminary results concerning the development of a web-based suggestion service for proposing plausible supplements to fill lacunae, thereby supporting the philological process of producing new critical editions within a new web-based digital scholarly editing environment. To automatically provide such suggestions, we have developed systems that generate textual supplements in Ancient Greek, employing both neural (BERT-like) and statistical (n-gram) language modeling approaches.
Designing LLM Agents for User-Centered Language Service Selection
Ryoichiro Ogawa | Donghui Lin | Fumito Uwano
Ryoichiro Ogawa | Donghui Lin | Fumito Uwano
With the rapid expansion of language resources and services across repositories and platforms, users face an overwhelming number of options. While this diversity promises flexibility, non-experts struggle to compose appropriate resource pipelines and select services that satisfy both functional and non-functional requirements. We propose a user-centered framework of LLM agents that interprets natural-language requests and performs end-to-end language service selection. The agents extract functional requirements to form coherent task compositions and select suitable language services for each component by interpreting non-functional quality aspects embedded in contextual cues. To ensure reliable and explainable decisions, we employ a four-step structured reasoning procedure that combines Few-Shot exemplars and Chain-of-Thought reasoning: extracting functional requirements, inducing non-functional evaluation axes, applying these axes as constraints in candidate retrieval, and determining a final composition. We construct a benchmark dataset pairing diverse user requests with standardized language service profiles containing metadata and quality indicators, and evaluate our framework against representative prompting-based baselines. Results show consistent gains in Precision, Recall, and F1-score, demonstrating improved capture of both functional intent and quality preferences. These findings demonstrate that structured LLM agents can bridge natural-language user intents and language service configurations, enabling end-to-end selection and composition in a transparent and user-centered manner.
User Profiling for Specification-Sensitive Recommendations with Large Language Model Prompting
Chih-Yu Chien | An-Zi Yen | Hen-Hsen Huang | Hsin-Hsi Chen
Chih-Yu Chien | An-Zi Yen | Hen-Hsen Huang | Hsin-Hsi Chen
Recently, there has been an increasing focus in research on the potential applications of large language models (LLMs) for personalized recommendations. Previous studies utilize LLMs to analyze the interaction between users and products to establish various personalized recommendation systems. However, recommendation becomes particularly challenging when items are associated with varied attributes, influenced by personal preferences, and described primarily through unstructured data. Moreover, analyzing implicit user preferences with product specifications for specification-sensitive recommendations remains largely unexplored. In this paper, we propose a framework that fully leverages prompting-based strategies to analyze user reviews and item attributes for the generation of user and product profiles, respectively. These profiles capture users’ implicit preferences and enable rating prediction or product recommendation, which are crucial for personalized recommendations. Experimental results show that our proposed framework effectively handles complex item attributes and user preferences to achieve promising performances in rating prediction.
Comparing Traditional and LLM-based Approaches for Automated Scoring of Dutch Writing Products
Joni Kruijsbergen | Orphee De Clercq
Joni Kruijsbergen | Orphee De Clercq
This research examines several traditional and recent approaches for automated grading of Dutch texts written by adolescent L1 speakers. We relied on a proprietary dataset comprising human-scored texts. Following recent paradigms in NLP research, we compared training a feature-based model to fine-tuning both mono- and multilingual BERT-based and generative large language models. The latter were also prompted directly in a zero-shot setting. The results reveal that the feature-based and BERT-based approaches are promising for the task at hand and even complementary, although there is still room for improvement. The error analysis demonstrates that the generative models do not only make more errors in classification, but that these error are also more problematic. We therefore conclude that especially generative LLMs are not directly employable in this educational context.
“Decode the Law": Towards Legal Text Simplification with Large Language Models
Mohammed Danish Rabbani | Subhadeep Roy | Sayantan Mitra | Tulika Saha
Mohammed Danish Rabbani | Subhadeep Roy | Sayantan Mitra | Tulika Saha
Legal documents are often verbose and structurally complex, posing significant barriers to public understanding and equitable access to justice. Despite growing interest in text simplification, efforts targeting the legal domain remain limited by a lack of robust, high-quality resources. In this paper, we address this gap by introducing SIMPLE-LAW, a curated benchmark dataset of over 6,000 aligned pairs of original and simplified legal passages, specifically constructed to facilitate research in legal text simplification by leveraging large language models (LLMs). We evaluate this dataset across both in-context learning and parameter-efficient fine-tuning paradigms using a range of state-of-the-art LLMs, with Unsloth variants of Mistral, LLaMA-3.2, Gemma, and Qwen-2.5. We assess performance using BERTScore, ROUGE, SARI, and a hallucination detection score, to capture both fidelity and readability. Results show that fine-tuned models significantly outperform in-context learners in terms of simplification quality and factual consistency. By offering a new dataset, rigorous evaluation, and baseline comparisons, our work provides a critical foundation for developing transparent and accessible AI systems in the legal domain.
CLASE: A Hybrid Method for Chinese Legalese Stylistic Evaluation
Yiran Rex Ma | Yuxiao Ye | Huiyuan Xie
Yiran Rex Ma | Yuxiao Ye | Huiyuan Xie
Legal text generated by large language models (LLMs) can usually achieve reasonable factual accuracy, but it frequently fails to adhere to the specialised stylistic norms and linguistic conventions of legal writing. In order to improve stylistic quality, a crucial first step is to establish a reliable evaluation method. However, having legal experts manually develop such a metric is impractical, as the implicit stylistic requirements in legal writing practice are difficult to formalise into explicit rubrics. Meanwhile, existing automatic evaluation methods also fall short: reference-based metrics conflate semantic accuracy with stylistic fidelity, and LLM-as-a-judge evaluations suffer from opacity and inconsistency. To address these challenges, we introduce CLASE (Chinese LegAlese Stylistic Evaluation), a hybrid evaluation method that focuses on the stylistic performance of legal text. The method incorporates a hybrid scoring mechanism that combines 1) linguistic feature-based scores and 2) experience-guided LLM-as-a-judge scores. Both the feature coefficients and the LLM scoring experiences are learned from contrastive pairs of authentic legal documents and their LLM-restored counterparts. This hybrid design captures both surface-level features and implicit stylistic norms in a transparent, reference-free manner. Experiments on 200 Chinese legal documents show that CLASE achieves substantially higher alignment with human judgments than traditional metrics and pure LLM-as-a-judge methods. Beyond improved alignment, CLASE provides interpretable score breakdowns and suggestions for improvements, offering a scalable and practical solution for professional stylistic evaluation in legal text generation (Code and data for CLASE is available at: https://github.com/rexera/CLASE).
Neural Network-assisted Analysis of Tube Vocal Tract Models
Runhui Song | Johan Sjons | Axel G. Ekstrom
Runhui Song | Johan Sjons | Axel G. Ekstrom
We present a pipeline for deep neural network assisted modeling and analysis of the behavior of an acoustic tube. The vocal tract is represented as a series of cylindrical tube segments, each characterized by fixed length and variable cross-sectional area. A large synthetic dataset of such tube configurations is generated, and a circuit theory–based algorithm predicts corresponding formant frequencies. To explore mapping between vocal tract shapes and formant values, the pipeline integrates both linear regression and nonlinear machine learning models - including multilayer perceptrons. Model interpretability is measured using Shapley Additive Explanations (SHAP), which quantifies the contribution of each segment to predicted formant frequencies. The proposed framework enables detailed exploration of the articulatory-acoustic relationships inherent to an acoustic tube and vocal tract simulacrum. We present and describe the pipeline in the context of modeling effects of perturbations on the first three formants for a 16-cm tube, divided into 1 cm segments. Our pipeline can be applied to any method that models predictions of behavior of an acoustic tube, where the tube is conceived as a series of segmented units.
Central Kurdish Text-to-Speech and Its Application in Speech-to-Text Translation
Mohammad Mohammadamini | Meysam Shamsi | Marie Tahon
Mohammad Mohammadamini | Meysam Shamsi | Marie Tahon
In this study, we show how from available resources develop high-quality TTS models for low-resource scenarios that according to our extensive evaluation surpass the models trained on dedicated TTS data recorded in the studio. We develop three Text-to-Speech (TTS) models for Central Kurdish as a low-resource language using F5-TTS architecture. The models are trained on Central Kurdish TTS datasets in which two of them are curated from audiobooks during this study and the third one is evaluated for the first time. We also demonstrate the potential of TTS models for developing other speech technologies in low-resource languages by proposing a speech synthesis framework used in a speech-to-text translation application, achieving promising results on standard speech translation benchmarks. The curated TTS resources and models will be publicly available under CC BY-NC-ND 4.0 license
QuALA-NL: Question & Answer with Legal Attribution in Dutch
Romy A.N. van Drie | Roos M. Bakker | Daan L. Di Scala | Maaike de Boer
Romy A.N. van Drie | Roos M. Bakker | Daan L. Di Scala | Maaike de Boer
Ensuring trustworthy and traceable outputs from Large Language Models (LLMs) is crucial in high-stakes domains such as law. Retrieval-Augmented Generation (RAG) offers a way to enhance LLMs with domain-specific or updated information and provide attribution to the source, and recent work has focused on knowledge-based RAG (K-RAG) for improved factual grounding. However, proper evaluation of such systems requires high-quality datasets. To address this need, we introduce QuALA-NL: a dataset that provides attributions to legal formalizations, enabling experiments with K-RAG in the legal domain. The dataset contains 101 QA pairs on three Dutch laws, with attributions to the law text and a formalization of the interpretation of the legal text. To demonstrate the capabilities of the dataset, we perform experiments using four configurations: LLM-only, RAG using legal texts, K-RAG using a formalization of the legal texts, and RAG combining both legal texts and the formalizations. The results show that K-RAG has the highest retrieval scores, but that this method is outperformed by text-based RAG on generation. A qualitative analysis shows that the use of the knowledge graph for the generation of answers can be improved. QuALA-NL can be used in future work to experiment with knowledge-based Retrieval Augmented Generation methods.
We present a method of attribution source detection and classification in Czech. A plain text (typically, a newspaper article) enters the SouDec system, gets parsed with the external tool UDPipe into Universal-Dependencies style of sentence representation, and then is analyzed for occurrences of attribution signals and sources. The list of attribution signals has been extracted from a corpus of Czech newspaper articles annotated with interlinked attribution signals and sources, and has been complemented with context and syntax information to help distinguish relevant occurrences of the signals. The SouDec system further classifies the attribution sources in one of five classes: anonymous, partially anonymous, unofficial, official non-political and official political, using information from another external tool, a recognizer and classifier of named entities, NameTag 3. While our source detection method gets results comparable to existing systems for other languages, further improvements can be achieved by incorporating fully-fledged automatic coreference resolution into the classification method. In a focused case study, we test a possible usage of SouDeC for distinguishing domain-specific texts of less vs. more reputable origin.
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
Lívia Dutra | Arthur Lorenzi | Lais Berno | Franciany Campos | Karoline Biscardi | Kenneth Brown | Marcelo Viridiano | Frederico Belcavello | Ely E. Matos | Olivia Guaranha | Erik Santos | Sofia Reinach | Tiago Timponi Torrent
Lívia Dutra | Arthur Lorenzi | Lais Berno | Franciany Campos | Karoline Biscardi | Kenneth Brown | Marcelo Viridiano | Frederico Belcavello | Ely E. Matos | Olivia Guaranha | Erik Santos | Sofia Reinach | Tiago Timponi Torrent
We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients’ visits to primary care units. A total of eight patterns are defined and searched on a corpus of 21 million sentences in Brazilian Portuguese extracted from e-SUS APS. The results are manually evaluated by linguists and the precision of each pattern measured. Our findings reveal that the methodology effectively identifies reports of violence with a precision of 0.726, confirming its robustness. Designed as a transparent, efficient, low-carbon, and language-agnostic pipeline, the approach can be easily adapted to other health surveillance contexts, contributing to the broader, ethical, and explainable use of NLP in public health systems.
PrePPER: A Preference Pattern-based Profiling Framework for Explainable Recommendation
Taisuke Usumi | Akiko Masaki | Sanae Muramatsu | Akira Sakamoto | Takeharu Eda
Taisuke Usumi | Akiko Masaki | Sanae Muramatsu | Akira Sakamoto | Takeharu Eda
Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks, drawing increasing attention to their application in recommendation systems. In particular, recommendation systems using natural language-based user profiles have attracted attention for improving transparency and scrutability. However, existing methods fail to fully leverage the recommendation capabilities of LLMs due to the unspecified importance of user preferences within user profiles and unmatched preference types between user profiles and item profiles. To address these challenges, we propose PrePPER, a novel preference pattern-based profiling framework designed to explicitly capture the importance of user preferences and enhance the alignment between user profiles and item profiles. PrePPER enables the extraction of users’ preference patterns, which denote characteristic tendencies in user preferences, and the determination of their importance by clustering users’ preferences. Specifically, we first extract users’ preferences from their reviews and perform clustering on the extracted preferences. Based on the clustered preferences, we then infer users’ preference patterns along with their relative importance, and construct user and item profiles using this information. Our proposed profiles incorporate the importance of user preferences and enhance the relatedness between user and item profiles, thereby improving the recommendation performance of existing recommender systems.
Evaluating the Impact of Source Diversity for RAG in Historical Research
Ruhi Mahadeshwar | Andreas van Cranenburgh | Tommaso Caselli | Malvina Nissim
Ruhi Mahadeshwar | Andreas van Cranenburgh | Tommaso Caselli | Malvina Nissim
Historical research increasingly benefits from large language models (LLMs). However, LLMs are prone to factual inaccuracy, unreliability, and biased interpretations of data. Retrieval-augmented generation (RAG) approaches have emerged as solutions, but may inadvertently perpetuate biased perspectives embedded in historical archives. This paper investigates how source diversity in RAG impacts perspective variation in historical question answering. We compile a multilingual corpus (English, French, Dutch) of historical documents spanning multiple countries and focus on Napoleon Bonaparte. We evaluate three Qwen3 models across ten questions using a multi-layered framework combining traditional metrics (BERTScore, ROUGE-L), frame semantics analysis, and syntactic profiling. Our results highlight that, while traditional similarity metrics suggest high semantic consistency, frame-semantic analysis exposes substantial perspective shifts. Baseline answers present “flattened” cross-lingual perspectives, whereas RAG introduces diversity. Critically, this diversity manifests differently across languages, demonstrating language-specific patterns. Our findings highlight limitations of traditional evaluation metrics for perspective-sensitive tasks and demonstrate that RAG constitutes active perspective transformation rather than neutral augmentation.
Automatic Essay Scoring and Feedback Generation in Basque Language Learning
Ekhi Azurmendi | Xabier Arregi | Oier Lopez de Lacalle
Ekhi Azurmendi | Xabier Arregi | Oier Lopez de Lacalle
This paper introduces the first publicly available dataset for Automatic Essay Scoring (AES) and feedback generation in Basque, targeting the CEFR C1 proficiency level. The dataset comprises 3,200 essays from HABE, each annotated by expert evaluators with criterion specific scores covering correctness, richness, coherence, cohesion, and task alignment enriched with detailed feedback and error examples. We fine-tune open-source models, including RoBERTa-EusCrawl and Latxa 8B/70B, for scoring. We focused on correctness criteria for the explanation generation, adapting Latxa to correctly predict both, scores and explanations. Our experiments show that encoder models remain highly reliable for AES, while supervised fine-tuning (SFT) of Latxa significantly enhances performance, surpassing state-of-the-art (SoTA) closed-source systems such as GPT-5 and Claude Sonnet 4.5 in scoring consistency and feedback quality. We also propose a novel evaluation methodology for assessing feedback generation, combining automatic consistency metrics with expert-based validation of extracted learner errors. Results demonstrate that the fine-tuned Latxa model produces criterion-aligned, pedagogically meaningful feedback and identifies a wider range of error types than proprietary models. This resource and benchmark establish a foundation for transparent, reproducible, and educationally grounded NLP research in low-resource languages such as Basque. The dataset, models and manual evalution results are available here: https://huggingface.co/collections/EkhiAzur/habe-hitz-c1
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
Fabian Retkowski | Alexander Waibel
Fabian Retkowski | Alexander Waibel
Automatic speech transcripts are often delivered as unstructured word streams that impede readability and repurposing. We recast paragraph segmentation as the missing structuring step and fill three gaps at the intersection of speech processing and text segmentation. First, we establish TEDPara (human-annotated TED talks) and YTSegPara (YouTube videos with synthetic labels) as the first benchmarks for the paragraph segmentation task. The benchmarks focus on the underexplored speech domain, where paragraph segmentation has traditionally not been part of post-processing, while also contributing to the wider text segmentation field, which still lacks robust and naturalistic benchmarks. Second, we propose a constrained-decoding formulation that lets large language models insert paragraph breaks while preserving the original transcript, enabling faithful, sentence-aligned evaluation. Third, we show that a compact model (MiniSeg) attains state-of-the-art accuracy and, when extended hierarchically, jointly predicts chapters and paragraphs with minimal computational cost. Together, our resources and methods establish paragraph segmentation as a standardized, practical task in speech processing.
High-Order Question Generation in a Multilingual Educational Context
Suna Seyma Uçar | Itziar Aldabe | Nora Aranberri | Orphee De Clercq
Suna Seyma Uçar | Itziar Aldabe | Nora Aranberri | Orphee De Clercq
Critical thinking is a fundamental skill that helps learners move beyond simple memorization. One way to develop this skill is through high-order questioning. However, crafting such questions remains a challenge for educators, and classroom practices tend to rely on low-order questions. Large Language Models have demonstrated strong capabilities in generating high-order questions, especially when guided by prompts based on Bloom’s Taxonomy. Yet, existing research has largely centered on this framework and focused only on English. This study addresses these gaps by introducing prompts grounded in two alternative frameworks: Claim-Evidence-Reasoning and Divergent Questioning within a multilingual context using Basque, Spanish, and English. Results indicate that while both an open-source and a proprietary model rather effectively generate questions in all three languages, only about half of the answerable questions are recognized by teachers as high-order. A positive finding is that the alternative frameworks produce structurally and conceptually varied questions, suggesting they could complement each other and provide viable alternatives to Bloom’s Taxonomy.
From Print to Digital and beyond: The Retrodigitization of a Historical Dictionary of Italian as a Hybrid Lexical Resource
Marco Biffi | Sebastiana Cucurullo | Manuel Favaro | Elisa Guadagnini | Simonetta Montemagni | Eva Sassolini
Marco Biffi | Sebastiana Cucurullo | Manuel Favaro | Elisa Guadagnini | Simonetta Montemagni | Eva Sassolini
This paper presents the retrodigitization project of the Grande Dizionario della Lingua Italiana (GDLI), the largest and most comprehensive historical dictionary of the Italian language. The GDLI’s 23,000 pages — originally designed for human consultation — constitute an exceptional repository of linguistic and cultural-historical information, while posing significant challenges to large-scale digitization and data structuring. The project, still ongoing, will result in the development of a set of interoperable and interlinked resources: (i) a TEI-XML edition of the dictionary text, encoding its complex lexicographic and citation structure; (ii) an annotated corpus of the quoted examples, enabling linguistic and historical research across centuries; and (iii) a database of cited authors and works. Together, these components form a hybrid lexical resource that establishes the foundations for innovative and advanced modes of accessing and exploring the rich and multifaceted content of this historical dictionary.
Learning through News: Bridging the Gap between Algorithmic Recommendation and Human Curation
Florian Debaene | Loic De Langhe | Orphee De Clercq | Veronique Hoste
Florian Debaene | Loic De Langhe | Orphee De Clercq | Veronique Hoste
News recommendation systems play a central role in how readers access and process current events. Most recommenders’ underlying algorithmic strategies, however, prioritize user engagement over comprehension, amplifying risks of misinformation and filter bubbles. This study investigates whether fine-grained content-based recommendation strategies favor human knowledge retention and explores how such a content-based recommendation can be operationalized using event coreference–based document modeling. To this purpose, we first measure the effect of manually curated content-based news recommendation on knowledge retention across five news topics with 126 Dutch speaking participants. Next, we investigate document retrieval by comparing a state-of-the-art event coreference resolution system for Dutch which recommends news articles based on event chains with a document similarity retrieval baseline using state-of-the-art embedding models in three increasingly more complex test settings. The results demonstrate that human-curated content-based recommendation can positively and significantly impact readers’ knowledge retention. Moreover, we show that a fine-grained coreference system can approach said level of human curation better than state-of-the-art document retrieval methods. In general, this holds potential for scalable, comprehension-oriented news recommendation.
MaskedVerbalizer: Automatic Verbalizer Construction for Few-Shot Text Classification in Low-Resource Right-to-Left Languages
Faizad Ullah | Furqan Sikandar | Areeba Waqar | Faizan Ali | Muhammad Sohaib Ayub | Mubashar Mushtaq | Asim Karim
Faizad Ullah | Furqan Sikandar | Areeba Waqar | Faizan Ali | Muhammad Sohaib Ayub | Mubashar Mushtaq | Asim Karim
Text classification in low-resource right-to-left languages faces significant challenges due to the scarcity of annotated data and the morphological richness of languages such as Arabic, Urdu, Sindhi, and Pashto. Arabic and Urdu alone are spoken by over 380+ million and 246+ million people worldwide, respectively. Pashto is the national language of Afghanistan, highlighting the importance of effective language technologies. While multilingual Pre-trained Language Models (PLMs) have shown promising results, they typically require extensive labeled datasets and computationally expensive fine-tuning to achieve better performance. Such limitations make these PLMs impractical for the low-resource settings described above. Therefore, we employ a few-shot strategy (zero, 4, or 8 shots) to achieve results comparable to those of standard fine-tuning. In this work, we propose MaskedVerbalizer, a novel technique designed for few-shot text classification. Our method introduces an automatic verbalizer construction approach that generates class-specific label words in 4-shot settings, eliminating the need for extensive manual intervention. Despite maintaining a simple model architecture, MaskedVerbalizer achieves effective performance in classification benchmarks. Experimental results demonstrate that our method effectively addresses the core challenges of low-resource text classification, providing a practical, computationally efficient solution. We achieved accuracies of 90.43% and 92.72% with mBERT and XLM-RoBERTa, respectively, representing improvements of 25–30% over soft and automatic verbalizers. The code for MaskedVerbalizer is publicly available at https://github.com/Furqann-hue/MV.
RBR: RAG-Based Open-Domain Question Answering Using a Ranking Approach to Document Retrieval
Priyatam Sai Naravajhula | Vincent Ng
Priyatam Sai Naravajhula | Vincent Ng
Retrieval-Augmented Generation (RAG) has emerged as a promising approach to ODQA. A RAG-based ODQA system is typically composed of two components: a retriever that retrieves the passages that are most relevant to a given query, and a generator that generates the answer to the query by combining the information from the retrieved passages. Existing retrievers typically identify the most relevant passages by computing the similarity between the query and each passage in a given collection. In other words, they do not compare which of two passages is more relevant to the given query. We hypothesize, however, that we can improve RAG-based ODQA systems by modeling the relationship among the passages to be retrieved, specifically by learning which passages are more relevant than the others to the given query. To do so, we propose a ranking-based approach to passage retrieval, where we first rank the candidate passages w.r.t. the query and subsequently refine the score associated with each of these passages using a Graph Attention Network. We evaluate our approach to ODQA, RBR (Ranking-Based Retrieval), on two commonly-used ODQA datasets, Natural Questions and TriviaQA. Experimental results show that RBR slightly outperforms PA-RAG, a state-of-the-art ODQA system, by 0.45 points and 1.01 points in Exact Match score on Natural Questions and TriviaQA, respectively.
Sentence-Level Back-Transliteration of Romanized Indian Languages: Performance Analysis and Challenges
Saurabh Kumar | Dhruvkumar Babubhai Kakadiya | Sanasam Ranbir Singh | Sukumar Nandi
Saurabh Kumar | Dhruvkumar Babubhai Kakadiya | Sanasam Ranbir Singh | Sukumar Nandi
The widespread use of Romanized text for Indian languages, particularly on social media platforms, poses significant challenges for natural language processing due to the lack of standardized orthography and the presence of contextual ambiguities. In this study, we explore sentence-level back-transliteration for 13 Indian languages, focusing on addressing the limitations of word-level models that fail to capture contextual dependencies. We evaluate state-of-the-art models, including fine-tuned LLaMA, mT5, and Multilingual Transformer models, comparing their performance against the baseline IndicXlit model. In addition, we conduct a comprehensive error analysis to gain deeper insights into model performance. Our results demonstrate that fine-tuned LLaMA and the proposed IndiXform model, specifically designed to leverage sentence-level context, significantly outperform zero-shot LLaMA and the IndicXlit baseline. These findings provide valuable insights into handling contextual ambiguities and enhancing the accuracy of back-transliteration systems for Indian languages.
Cross-Corpus CEFR Classification through Artificial Learners Perplexities
Bernardo Stearns | John P. McCrae | Thomas Gaillat
Bernardo Stearns | John P. McCrae | Thomas Gaillat
The complexity of neural methods for automatic proficiency assessment often sacrifices interpretability and robustness. This paper presents a competitive alternative for CEFR classification using optimized statistical models with a novel perplexity-based feature engineering pipeline. We introduce LLM-derived perplexity features as a proxy for how unexpected a learner’s word choices are: native model perplexity measures unexpectedness relative to native language use, while Artificial Learner model perplexity quantifies relative to a specific proficiency level. While recent work favors end-to-end neural architectures, we demonstrate that traditional pipelines enhanced with these interpretable perplexity features can achieve comparable performance on established benchmarks. We evaluate two transfer scenarios: zero-shot (trained on EFCAMDAT, tested on external corpora) and 90-10 split (same features, in-domain classifier training). On KUPA-KEYS, perplexity features achieve RMSE 0.707 (zero-shot) and 0.660 (90-10 split), outperforming fine-tuned BERT and prompt-based LLMs. On CELVA-SP, zero-shot perplexity shows limited generalization (RMSE 1.437 vs. LLM’s 1.016), but statistical models close this gap in the 90-10 split (RMSE 0.872). Across all three evaluation datasets, perplexity-based models achieve the best average macro F1 in the 90-10 split (0.446 vs. 0.287 for BERT and 0.175 for prompting), demonstrating that interpretable features paired with domain-adapted classifiers provide the most robust cross-domain representations. We contribute: (1) state-of-the-art KUPA-KEYS results with interpretable models, (2) the first comprehensive CELVA-SP benchmark, and (3) evidence that feature-level transfer outperforms both end-to-end fine-tuning and zero-shot prompting.
CorpusClues: Scalable Unsupervised Similarity Search for Historical Texts Using MinHash-LSH
Paulien Lemay | Klaas Bentein | Els Lefever
Paulien Lemay | Klaas Bentein | Els Lefever
CorpusClues is a prototype web-based platform for large-scale, unsupervised clustering of textual data, designed to address the specific challenges of historical corpora. It leverages the well-established computational techniques of MinHash and Locality-Sensitive Hashing (LSH) at the character level in order to detect structural similarities between texts even when exact patterns diverge. This approach makes CorpusClues robust to orthographic variation, such as historical spelling differences, while remaining fast and language-agnostic, capable of processing large and heterogeneous corpora without relying on language-specific models or preprocessing. Researchers can explore resulting clusters through interactive visualizations and exportable data, gaining access to patterns that would otherwise require the slow and uncertain process of manual collation. Evaluation against labeled gold standards shows that the system consistently produces high-quality clustering, accurately reconstructing relationships between texts despite substantial orthographic variation. By combining computational efficiency with user-friendly design, CorpusClues provides an accessible yet rigorous means of uncovering formulaicity and textual transmission at scale, opening new possibilities for the study of historical textual traditions.
BenCSSmark: Making the Social Sciences Count in LLM Research
Arnault Chatelain | Etienne Ollion | Qianwen Guan | Diandra Fabre | Lorraine Goeuriot | Emile Chapuis | Abdelkrim Beloued | Marie Candito | Nicolas Hervé | Didier Schwab
Arnault Chatelain | Etienne Ollion | Qianwen Guan | Diandra Fabre | Lorraine Goeuriot | Emile Chapuis | Abdelkrim Beloued | Marie Candito | Nicolas Hervé | Didier Schwab
This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation and social scientific inquiry. Benchmarks — standardized tools for assessing computational systems — are pivotal in the development of artificial intelligence (AI), including large language models (LLMs). Benchmarks do more than measure progress — they actively structure it, shaping reputations, research agendas, and commercial outcomes. Despite this central role, the social sciences are largely absent from mainstream evaluation frameworks, even though scholars in these fields generate dozens of rigorously annotated, context-sensitive datasets each year. Integrating this work into benchmark design could significantly improve the generalization and robustness of AI models. In turn, models trained on social scientific tasks would likely yield better performance on classic and contemporary tasks in disciplines as diverse as history, sociology, political science or economics. This is all the more pressing as these disciplines are quickly turning to LLMs for assistance. To address this gap, we introduce BenCSSmark, a benchmark composed of datasets annotated by computational social scientists. By integrating social scientific perspectives into benchmarking, BenCSSmark seeks to promote more robust, transparent, and socially relevant AI systems and to foster efficient collaboration.
Predicting Topic (Co-)Occurrence Using Topic Networks Built from the Project Gutenberg Corpus
Bhuvanesh Verma | Alexander Mehler
Bhuvanesh Verma | Alexander Mehler
Although temporal topic modeling has been widely applied to scientific and legal texts, literary corpora have largely been overlooked in this regard. To address this issue, we analyze topic evolution in a subset of the Project Gutenberg (PG) corpus. We model this subset as a sequence of topic networks that capture the emergence, persistence, and interaction of thematic structures over decades. Using supervised topic representations, we predict nodes (topics) and edges (topic pairings) to forecast future topics and their co-occurrence. Our experiments demonstrate moderate to strong temporal persistence in topic connectivity patterns across three topic systems, with ROC-AUC and AP values consistently above 0.85. We find that the temporal span of topic networks significantly impacts predictive performance: longer spans improve the stability and recall of topic presence, while shorter spans better capture evolving topic relationships. Overall, our findings demonstrate the predictability of topics in literary texts over time.
AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse
Esra’a Ahmad Sharqawi | Wajdi Zaghouani
Esra’a Ahmad Sharqawi | Wajdi Zaghouani
Social media has become a crucial arena for shaping public narratives during armed conflicts, providing space for both harmful and constructive communication. While hate speech and misinformation have been widely studied, expressions that promote resilience, solidarity, and optimism remain underexplored, particularly in Arabic contexts. This paper introduces AraHopeCorpus, the first annotated dataset of Arabic hope speech collected from ten thousand YouTube comments related to the war on Gaza between 2023 and 2024. Using a detailed annotation framework, comments were classified into three categories: hope speech, no hope speech, and neutral or unclear discourse. The dataset shows that hopeful language dominates, accounting for more than sixty four percent of all comments. These expressions of hope appear mainly as religious encouragement, collective solidarity, and optimism for endurance and justice. No hope speech, representing about thirteen percent, reflects despair and disillusionment, while the rest of the comments contain neutral or mixed content. Inter-Annotator Agreement reached substantial levels (Cohen’s Kappa equals 0.71), though dialectal variation, sarcasm, and implicit meaning posed annotation challenges. A comparative analysis between human annotators and ChatGPT revealed that large language models can support annotation but remain limited in handling dialectal and culturally embedded expressions. AraHopeCorpus will be released for research purposes under an open and non commercial license. It provides a valuable resource for studying constructive digital discourse, enabling further research on hope speech detection, crisis communication, and resilience in Arabic social media.
Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse
Aisha Ali Al-Athba | Wajdi Zaghouani
Aisha Ali Al-Athba | Wajdi Zaghouani
The study of online discourse has become central to understanding societal polarization. While much research has focused on detecting overt toxicity, the subtle dynamics of social cohesion, meaning the interaction between divisive and unifying narratives, remain computationally underexplored. This paper presents Cohesion-6K, a manually and ChatGPT-assisted annotated dataset of six thousand Arabic public Facebook posts related to the Israeli Occupation of Palestine. Each post is assigned to one of five discourse categories that represent a continuum from conflict to cohesion: Conflict, Resolution, Community Engagement, Supportive Interactions, and Shared Values. The annotation process combines expert human judgment with model-assisted pre-labeling verified by trained annotators, achieving substantial inter-annotator agreement (Cohen’s kappa = 0.85). Quantitative analysis reveals a consistent engagement gap, where conflict-oriented posts receive between two and four times more user interaction than resolution-oriented ones (p < 0.01). This pattern illustrates how divisive discourse tends to attract disproportionate visibility in Arabic social media spaces. Cohesion-6K provides a transparent and reproducible resource for the study of online cohesion and polarization. The dataset, annotation guidelines, and preprocessing code will be released for research use under an open license, supporting future work in computational social science, digital communication, and Arabic natural language processing.
Reference-free Evaluation at Inference for NER/NEL over OCRed Historical Texts
Tien-Nam Nguyen | Adam Jatowt | Ahmed Hamdi | Mickael Coustaty | Thi Hong Hanh Tran | Antoine Doucet
Tien-Nam Nguyen | Adam Jatowt | Ahmed Hamdi | Mickael Coustaty | Thi Hong Hanh Tran | Antoine Doucet
Named Entity Recognition (NER) and Named Entity Linking (NEL) are core tasks in entity extraction, yet their robustness is limited when applied to noisy documents, such as those generated by Optical Character Recognition (OCR) over historical documents. Although large language models (LLMs) have shown strong zero-shot and few-shot performance on NER and NEL tasks, prior work has largely focused on using LLMs as direct predictors rather than evaluating extraction performance. In this study, we explore the feasibility of using LLMs as learned evaluators to estimate the quality of NER/NEL outputs, especially in settings where human-annotated references are unavailable at inference time. We propose supervised approaches that fine-tune LLMs to predict quality scores based on training data with gold annotations, enabling reference-free quality estimation once trained. Experiments on the HIPE-2020 benchmark across English, French, and German languages demonstrate that fine-tuned LLMs provide reliable estimates of output quality. Our findings suggest that LLM-based evaluation can support quality control and enable evaluation in noisy setting.
Echoes of the Troubadours: A Corpus of Troubadour Poetry for Stylometric Analysis and Authorship Attribution
Loic De Langhe | Orphee De Clercq | Veronique Hoste
Loic De Langhe | Orphee De Clercq | Veronique Hoste
We present TrobaCor, a curated corpus of medieval troubadour poetry, which comprises 1668 unique Old Occitan texts by a large variety of authors. Clustering and stylometric experiments show that we can accurately model authorial style beyond topical content, even though formulaic or topically diverse genres remain challenging. Furthermore, we can model and detect traces of an author’s stylistic “DNA” even in short-form collaborative poetry, offering a uniquely fine-grained perspective in the field. In addition, we provide self-organizing map visualizations in order to provide an interpretable view of stylistic patterns across authors. TrobaCor is publicly released to support reproducible research in NLP and digital humanities on this low-resource historical corpus.
Gretino: A Greek and Latin Dataset to Benchmark Retrieval Systems in Classical Languages
Hawau Olamide Toyin | Federico Iezzi | Elia Scapini | Giulio Federico | Giovanni Puccetti
Hawau Olamide Toyin | Federico Iezzi | Elia Scapini | Giulio Federico | Giovanni Puccetti
Semantic similarity search is a method for exploring large text corpora and retrieving conceptually related content. Although widely used in modern language applications, it remains underexplored in the context of classical literature, where it could provide scholars with tools to uncover meaningful connections across authors, genres, and languages, surpassing the limitations of rule-based or keyword search systems. To promote the adoption of semantic retrieval in classical languages, we introduce Gretino, the first benchmark dataset for evaluating semantic search systems in Latin, Ancient Greek, and cross-lingual settings. Gretino comprises 240 carefully designed queries, each paired with five semantically relevant passages in Latin and Greek. The dataset is divided into two subsets: Gretino Silver, consisting of 200 queries and 1,000 targets (evenly split between Latin and Greek), generated with the assistance of ChatGPT and subsequently reviewed; and Gretino Gold, a manually curated high-quality subset of 40 queries and 200 targets, fully based on authentic classical texts. We evaluate four pre-trained language models: GreBERTa, LaBERTa, PhilBERTA, and SPhilBERTa and demonstrate the potential of a contrastive learning approach based on SimCSE (Gao et al., 2021) for fine-tuning, showing that training on carefully curated bilingual corpora, with texts aligned in the two languages, can improve retrieval performance.
A Recipe for Adapting Multilingual Embedders to OCR-Error Robustness and Historical Texts
Andrianos Michail | Stylianos Psychias | Juri Opitz | Simon Clematide
Andrianos Michail | Stylianos Psychias | Juri Opitz | Simon Clematide
Modern multilingual text embedding models excel at semantic search on contemporary text but their performance degrades measurably on digitized historical documents. This issue is especially pronounced for underrepresented languages such as Luxembourgish, where historical materials combine evolving spelling conventions with OCR artifacts absent from standard training data. To address these challenges, we introduce OCR M-GTE, a pair of multilingual embedding models adapted for OCR robustness and historical texts, and show that the observed degradation can be mitigated through a simple multi-step training procedure tailored to historical variants and OCR noise. We evaluate the models on standard semantic search tasks, simulated OCR degradation, and genuine historical collections, observing consistent improvements under OCR-induced noise and on genuine historical data while maintaining comparable performance on clean modern text. Our ablation findings suggest that multilingual embedding models can be effectively adapted to perform robust cross-lingual search in heterogeneous European digitized corpora. We release our adapted models, code, and datasets under the AGPL-3.0 license: https://github.com/impresso/ocr-robust-multilingual-embeddings
Phrase-Level Segmentation on Medieval Corpora for Aligning Multilingual Texts
Lucence Ing | Matthias Gille Levenson | Carolina Macedo
Lucence Ing | Matthias Gille Levenson | Carolina Macedo
This paper presents an approach to multilingual alignment for medieval languages, focusing on the prior step of"phrase" segmentation. It outlines the challenges posed by historical data and describes different strategies forsegmenting texts in multiple languages. It releases a gold-standard segmentation corpus based on various literaryand historical works from the late Middle Ages in Europe. This corpus consists of texts in seven medieval languages (French, Castilian, Catalan, Portuguese, Latin, Italian, English). Several architectures are tested with both in-domain and out-of-domain evaluation sets.
RAGE: Roman and Greek Emotions
Frederick Riemenschneider | Jonathan D. Geiger | Thomas Kuhn-Treichel | Anette Frank
Frederick Riemenschneider | Jonathan D. Geiger | Thomas Kuhn-Treichel | Anette Frank
The study of emotions in ancient Greek and Latin literature has largely been qualitative, relying on close reading, while existing computational methods often focus on coarse-grained sentiment polarity, which limits their use for nuanced literary analysis. To bridge this gap, we present RAGE (Roman And Greek Emotions), a new corpus of approximately 100 000 words of annotated classical literature spanning multiple genres and authors. Our multi-layered annotation framework, inspired by semantic role labeling, is designed for fine-grained analysis, capturing not only the emotion itself but also its experiencer, cause, and target. We adopt a nuanced emotion taxonomy and enrich each emotion instance with additional layers for intensity, explicitness, and negation. To facilitate comparative analysis, characters are linked to Wikidata or a local ontology. We demonstrate the utility of our corpus through corpus-level exploratory analyses and an in-depth case study. RAGE and its accompanying guidelines provide a valuable resource for applying quantitative methods to the study of emotions in classical texts.
From Variance to Invariance: Qualitative Content Analysis for Narrative Graph Annotation
Junbo Huang | Max Weinig | Ulrich Fritsche | Ricardo Usbeck
Junbo Huang | Max Weinig | Ulrich Fritsche | Ricardo Usbeck
Narratives in news discourse play a critical role in shaping public understanding of economic events, such as inflation. Annotating and evaluating these narratives in a structured manner remains a key challenge for Natural Language Processing (NLP). In this work, we introduce a narrative graph annotation framework that integrates principles from qualitative content analysis (QCA) to enhance methodological consistency. We present a dataset of inflation narratives annotated as directed acyclic graphs (DAGs), where nodes represent events and edges encode causal relations. To evaluate annotation quality, we employed a 6×3 factorial experimental design to examine the effects of narrative representation (six levels) and distance metric type (three levels) on inter-annotator agreement (Krippendorrf’s 𝛼), capturing the presence of human label variation (HLV) in narrative interpretations. Our analysis shows that (1) lenient metrics (overlap-based distance) overestimate reliability; (2) locally-constrained representations (e.g., one-hop neighbors) reduce annotation variability. Our annotation and implementation of graph-based Krippendorrf’s 𝛼 are open-sourced. The annotation framework and evaluation results provide practical guidance for NLP research on graph-based narrative annotation.
Historical corpora, especially those compiled from magazines and periodicals, are complex due to the diversity of text types and evolving genre conventions. Addressing these challenges requires systematic genre annotation and well-defined classification schemes to support downstream NLP tasks. This paper introduces a dataset of historical medical periodical texts in German and Swedish annotated for textual genre and additional features that may influence genre identification, such as the presence of OCR errors. We describe the development of the genre classification, annotator recruitment and training procedures, and provide an analysis of the annotator agreement.
Preserving Endangered Linguistic Heritage: Developing a Corpus for the Study of Contact-induced Changes in Corfioto
Giorgio Maria Di Nunzio | Georgios Vardakis
Giorgio Maria Di Nunzio | Georgios Vardakis
This paper presents current results of a work-in-progress project on the aims, goals, and methods for compiling a state-of-the-art morphosyntactically annotated corpus of Corfioto, the endangered Balkan Venetan variety of the Corfiot Jews. It gives an outline of the workflow for building, archiving, managing and annotating the first mixed-language corpus of original oral and written data of the Corfiot Jews, based on the Universal Dependencies (UD) framework and introduces the design and the implementation of an application for the Interactive MorPhosyntactic Annotation of Corfioto (IMPACT). The creation and the annotation of the corpus serves three goals: i) attain a quantitative analysis of variation in available data for the analysis of contact-induced syntactic change in clausal complementation in Corfioto; ii) enable the creation of a gold standard and the training of a model for the linguistic annotation of all data in the Universal Dependencies framework; and iii) contribute to the ever-growing research in the development of language resources and tools for endangered and low-resource contact varieties via the collaboration of computational, theoretical and fieldwork linguists.
To Eat and beyond: A FrameNet-Inspired Annotation of Food and Its Uses over Time
Teresa Paccosi | Gauri Bhagwat | Marieke van Erp
Teresa Paccosi | Gauri Bhagwat | Marieke van Erp
We present an annotation scheme and a manually annotated dataset in English, grounded in Frame Semantics and its generative extension through qualia relations, developed specifically for the food domain. Our primary goal is to capture the diverse and often less frequent uses of food in historical English texts, with a particular focus on the various processes to which food is subjected and the contexts in which it is employed. We provide the annotation scheme, describe the annotation process and release the annotated dataset for food and its uses, along with some preliminary experiments assessing the capabilities of LLMs in applying this annotation scheme.
To Overfit or Not to Overfit? An Evaluation of HTR Workflow on 17Th-18Th Century French Corpus
Marine Tiger
Marine Tiger
This paper presents the results of an evaluation of general Handwritten Text Recognition (HTR) models applied to 17th and 18th century corpus written in modern French and the fine-tuning of the models. Our aim was to transcribe a corpus from this period using existing pre-trained models and to assess their performance on such data. While these general models offer a large linguistic coverage, our results demonstrate they are often insufficiently adapted to the specific handwriting nuances and orthographic inconsistencies of early modern French. To improve the results, we fine-tuned a base model to develop a specialized version trained on our dataset. Although the model still encountered difficulties due to highly variable handwriting styles, it significantly improved transcription accuracy and reduced processing time. Following this step, we used a semi-automatic post-correction tool to address remaining errors and integrated Named Entity Recognition (NER) steps for automated TEI-XML encoding. This paper discusses the evaluation results of both the HTR and NER models, and how the overfitting allows to get better transcriptions on a specific corpus.
Automatic Segmentation of Classical Tibetan Texts into Autochthonous and Allochthonous Regions
Guy Bilitski | Lev Shechter | Sonam Jamtsho | Nir Marciano | Nicola Bajetta | Rebecca Sunden | Omri Drori | Kai Golan Hashiloni | Orr Zwebner | Asaf Shina | Orna Almogi | Dorji Wangchuk | Kfir Bar
Guy Bilitski | Lev Shechter | Sonam Jamtsho | Nir Marciano | Nicola Bajetta | Rebecca Sunden | Omri Drori | Kai Golan Hashiloni | Orr Zwebner | Asaf Shina | Orna Almogi | Dorji Wangchuk | Kfir Bar
We introduce a new computational framework for segmenting Classical Tibetan texts into autochthonous and allochthonous regions, distinguishing between indigenous Tibetan compositions and translated materials, primarily from Sanskrit sources. To support this task, we release the first annotated Tibetan corpus for ALLO/AUTO segmentation and evaluate several multilingual encoders, including mBERT and XLM-R, fine-tuned for sequence labeling. Our best model achieves strong alignment with expert annotations, showing that multilingual representations can effectively capture philological boundaries in low-resource settings. This work contributes new resources and methods for computational philology and sheds light on the linguistic markers that trace the intercultural transmission of Buddhist thought in Tibet.
RespondeoQA: A Benchmark for Bilingual Latin-English Question Answering
Marisa Hudspeth | Patrick J. Burns | Brendan O’Connor
Marisa Hudspeth | Patrick J. Burns | Brendan O’Connor
We introduce a benchmark dataset for question answering and translation in bilingual Latin and English settings, containing about 7,800 question–answer pairs. The questions are drawn from Latin pedagogical sources, including exams, quizbowl-style trivia, and textbooks ranging from the 1800s to the present. After automated extraction, cleaning, and manual review, the dataset covers a diverse range of question types: knowledge- and skill-based, multihop reasoning, constrained translation, and mixed language pairs. To our knowledge, this is the first QA benchmark centered on Latin. As a case study, we evaluate three large language models–LLaMa 3, Qwen QwQ, and OpenAI’s o3-mini–finding that all perform worse on skill-oriented questions. Although the reasoning models perform better on scansion and literary-device tasks, they offer limited improvement overall. QwQ performs slightly better on questions asked in Latin, but LLaMa3 and o3-mini are more task dependent. This dataset provides a new resource for assessing model capabilities in a specialized linguistic and cultural domain, and the creation process can be easily adapted for other languages. The dataset is available at: https://github.com/slanglab/RespondeoQA
Transformer-Enabled Diachronic Analysis of Vedic Sanskrit: Neural Methods for Quantifying Types of Language Change
Ananth A. Hariharan | David R. Mortensen
Ananth A. Hariharan | David R. Mortensen
This study demonstrates how hybrid neural-symbolic methods can yield significant new insights into the evolution of a morphologically rich, low-resource language. We challenge the naive assumption that linguistic change is simplification by quantitatively analyzing over 2,000 years of Sanskrit, demonstrating how weakly-supervised hybrid methods can yield new insights into the evolution of morphologically rich, low-resource languages. Our approach addresses data scarcity through weak supervision, using 100+ high-precision regex patterns to generate pseudo-labels for fine-tuning a multilingual BERT. We then fuse symbolic and neural outputs via a novel confidence-weighted ensemble, creating a system that is both scalable and interpretable. Applying this framework to a 1.47-million-word diachronic corpus, our ensemble achieves a 52.4% overall feature detection rate. Our findings reveal that Sanskrit’s overall morphological complexity does not decrease but is instead dynamically redistributed: while earlier verbal features show cyclical patterns of decline, complexity shifts to other domains, evidenced by a dramatic expansion in compounding and the emergence of new philosophical terminology. Critically, our system produces well-calibrated uncertainty estimates, with confidence strongly correlating with accuracy (Pearson r = 0.92) and low overall calibration error (ECE = 0.043), bolstering the reliability of these findings for computational philology.
Ithaca Revisited: Benchmarking a Domain-Specific Model for Epigraphy in the Age of LLMs
Alessandro Locaputo | Andrea Brunello | Nicola Saccomanno | Paraskevi Platanou | Giuseppe Serra
Alessandro Locaputo | Andrea Brunello | Nicola Saccomanno | Paraskevi Platanou | Giuseppe Serra
The restoration and interpretation of fragmentary inscriptions remain central challenges in epigraphy, where scholars must reconstruct missing text and determine an inscription’s provenance and chronology from limited evidence. Ithaca, a neural model introduced in 2022, represented a landmark advance in this field, achieving highly accurate results in text restoration and spatio-temporal attribution. Since then, general-purpose large language models (LLMs) such as GPT, Claude, and Gemini have achieved remarkable versatility across many domains, raising the question of whether specialized architectures like Ithaca are still required. In this paper, we revisit Ithaca with a dual focus. First, we benchmark its performance against GPT-5, finding that Ithaca continues to substantially outperform a state-of-the-art general-purpose LLM used in a retrieval-augmented in-context learning setting. Second, we conduct a systematic analysis to characterize Ithaca’s behavior under varying conditions, including lacuna size and position, inscription origin, and semantic topic. Statistical analyses highlight its systematic strengths and weaknesses. Taken together, our results map Ithaca’s performance profile, enabling more informed use in research and teaching.
Beyond Literal Meaning: How LLMs Interpret Yemeni Proverbs
Nasser Thmer | Ali Al-Laith | Muhammad Shoaib
Nasser Thmer | Ali Al-Laith | Muhammad Shoaib
We present a benchmark Yemeni proverbs dataset paired with expert-annotated explanations, designed to evaluate the cultural reasoning abilities of large language models (LLMs). Using zero-shot and few-shot prompting, we assess seven LLMs through both automatic and human evaluation. Results show that instruction-tuned models like GPT-4o and Gemini 1.5 Pro outperform smaller models in both automatic and human evaluations. Few-shot prompting significantly improves performance across all models, underscoring its value for figurative and culturally grounded language tasks. Notably, ALLaM, a bilingual model trained on Arabic and English, achieves competitive results, demonstrating the potential of regionally adapted models for low-resource cultural tasks. LLM-as-a-Judge evaluation correlates strongly with human assessment (Kendall’s τ up to 0.98). Error analysis identifies recurring literal interpretation and cultural misalignment as key failure modes.
CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs
Anna Glazkova | Antonina Laposhina | Dmitry Morozov
Anna Glazkova | Antonina Laposhina | Dmitry Morozov
This study explores the automated prediction of text complexity levels for short Russian texts on the Common European Framework of Reference for Languages (CEFR) scale. The dataset consists of 7,322 nonfictional fragments (15–30 words) extracted from textbooks for learners of Russian as a second language and filtered according to linguistic feature distributions typical of each CEFR level, with additional validation conducted by 4 human experts. Each text fragment was annotated with 127 linguistic features, including lexical, morphological, syntactic, and length-based characteristics. We evaluate several approaches to text complexity assessment: traditional machine learning classifiers, fine-tuned transformer models, and instruction-based large language models (LLMs). Among all models, RuBERT achieved the best strict F1-score (47.8%) and the lowest mean absolute error (0.56), while instruction-based LLMs such as YandexGPT captured overall complexity trends but underperformed in exact classification. Feature ablation experiments demonstrated that lexical features are the most informative for CEFR prediction. Our findings confirm that fine-tuned language models currently offer the most reliable results for short-text CEFR assessment in Russian, whereas instruction-based LLMs show potential for qualitative analysis of text difficulty patterns.
Evaluation of Document-Level Text Simplification in Japanese
Iori Yamashita | Hikari Tanaka | Hajime Kiyama | Kexin Bian | Zhousi Chen | Mamoru Komachi
Iori Yamashita | Hikari Tanaka | Hajime Kiyama | Kexin Bian | Zhousi Chen | Mamoru Komachi
This study establishes an evaluation framework for document-level text simplification in Japanese by constructing a human-annotated dataset and examining the reliability of LLM-based automatic evaluation. We first developed detailed annotation guidelines covering four criteria—necessity, sufficiency, sentence-level simplicity, and document-level simplicity—and collected human ratings for 1,128 source–target document pairs derived from the Wikipedia part of the Japanese simplification corpus JADOS. Using this dataset, we conducted extensive experiments comparing human judgments with evaluations from large language models, including GPT, Claude, and Gemini. The results show that GPT-4o and Gemini 2.5 Pro achieve high agreement with human annotators even in the 0-shot setting, demonstrating their potential as reliable automatic evaluators for Japanese simplification. However, LLMs exhibited a consistent tendency to underestimate document-level simplicity, particularly for kanji-dense texts or texts with relatively long sentences and a small number of sentences. This work provides the first benchmark for evaluating document-level text simplification in Japanese and offers practical evidence that LLM-based evaluation can support scalable assessment for Japanese document-level simplification.
Parallel Corpus Filtering Based on Semantic Similarity and Surface Dissimilarity for Japanese Text Simplification with LLMs
Daisuke Maekawa | Tomoyuki Kajiwara | Takashi Ninomiya
Daisuke Maekawa | Tomoyuki Kajiwara | Takashi Ninomiya
We are focusing on low-cost fine-tuning for large language models (LLMs) in Japanese text simplification. LLMs have achieved high performance even with fine-tuning on small parallel corpora in tasks such as machine translation and dialogue response generation. In this study, we propose a method of parallel corpus filtering for text simplification and investigate how much the number of sentence pairs for fine-tuning LLMs can be reduced. Experimental results on Japanese corpora in three domains revealed that the ability to perform text simplification tasks can be acquired even from a very small corpus of 16 to 64 sentence pairs. Although more parallel corpora are needed to acquire domain knowledge, our method outperformed full fine-tuning while reducing the training corpus by approximately 70%.
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
Stefan Bott | Verena Riegler | Horacio Saggion | Almudena Rascón Alcaina | Nouran Khallaf
Stefan Bott | Verena Riegler | Horacio Saggion | Almudena Rascón Alcaina | Nouran Khallaf
Being able to understand information is a key factor for a self-determined life and society. The study of automatic text simplification is often limited by the availability of high quality material for the training and evaluation on automatic simplifiers. This is true for English, but more so for less resourced languages like Spanish, Catalan and Italian. In order to fill this gap, we present a corpus of of original texts with high quality simplification produced by human experts in text simplification. It was developed within a project to assess the impact of Easy-to-Read (E2R) language for democratic participation. The original texts were compiled from domains related to this topic. The corpus includes different text types, selected based on relevance, copyright availability, and ethical standards. All texts were simplified to Easy-to-Read level. The corpora hold significant scientific value, particularly as it includes the first annotated corpora of its kind for the Catalan language. It also represents a noteworthy contribution for Spanish and Italian, offering high-quality, human-annotated language resources that are rarely available in these domains. The corpora will be made freely accessible to the public.
Proffiliadur: Welsh Language Text Profiling Toolkit
Nicolás Gutiérrez-Rolón | Jonathan Davies | Tomos Williams | Dawn Knight | Fernando Alva-Manchego
Nicolás Gutiérrez-Rolón | Jonathan Davies | Tomos Williams | Dawn Knight | Fernando Alva-Manchego
We introduce Proffiliadur, a Python toolkit for text profiling and readability analysis in Welsh. The toolkit computes 141 surface, lexical, morphological, and syntactic indices, designed to capture linguistic variation while incorporating a Welsh-specific tokenisation process that enables accurate morphological analysis and handles phenomena such as initial consonant mutation. Proffiliadur enables systematic assessment of text accessibility and supports applications in education, healthcare, and public communication. We demonstrate the toolkit’s usefulness through two complementary analyses. First, we examine texts written in accordance with the Cymraeg Clîr (“Clear Welsh”) principles and compare them with regular Welsh texts. Second, we analyse texts across CEFR proficiency levels to explore how linguistic complexity varies with learner ability. We also evaluate feature-based and neural classification models for automatic complexity detection, showing that interpretable linguistic indices alone achieve strong predictive performance (F1 = 0.94), comparable to a fine-tuned transformer (F1 = 0.97). Proffiliadur provides the first dedicated text profiling toolkit for Welsh, offering reproducible, linguistically grounded measures of readability for a low-resource language.
For vocabulary learning in language acquisition, it is desirable for learners to acquire words that they are likely to need in the language environments they will encounter. Such language environments are referred to as “registers” in general corpora, which are typically designed to include diverse registers. However, the proportion of registers included, that is, which registers are included and to what extent, is determined by the circumstances under which each general corpus was compiled and is not necessarily optimized for language learning. To bridge this gap, various leveled wordlists have been created in language education using linguistic resources other than word frequency, such as expert judgment and learner responses. However, it has not been quantitatively clear what gap in register proportions in general corpora these leveled wordlists were designed to fill. This study proposes a method that, given a leveled wordlist and a general corpus, estimates the register ratio that best aligns the frequency ordering of words across registers with the leveled wordlist. This makes it easier for learners and educators to interpret which wordlists are appropriate for particular learning goals. Our method is formulated as a linear programming problem and yields a globally optimal solution. Unlike neural networks, it is less susceptible to variation due to initial values or approximation and is therefore easier to interpret. We evaluated the proposed method on two languages, English and Japanese, through a range of experiments. We further show that it can also be used to evaluate vocabulary lists created for specific contexts, such as those generated by Large Language Models like ChatGPT.
Fill-in-the-Blanks: Automatic Generation and Evaluation of Language Models’ Pseudonyms for English and Swedish Texts
Maria Irena Szawerna | Jacob Lee Suchardt
Maria Irena Szawerna | Jacob Lee Suchardt
While considerable effort has gone into developing solutions for detecting Personally Identifiable Information (PII) in linguistic data, less research has gone into automating the generation of appropriate pseudonyms and developing evaluation methods, both relevant for the creation of privacy-friendly language resources. We conduct pilot experiments using Masked and Generative Large Language Models to generate predictions for redacted PII-spans in a cloze-like fashion for English legal texts and parallel news articles in Swedish and English. Furthermore, we explore metrics for automatic evaluation of the generated pseudonyms in the legal data, and investigate the effect of part-of-speech constraints on performance. For the parallel, multilingual data, we contribute our manual PII-annotation and conduct a fine-grained error analysis across two of our pseudonym generation methods and a baseline. Our results illustrate the complexity of pseudonym evaluation and the particular challenge of automatic, at-scale evaluation as well as the models’ tendency to predict prototypical and even stereotypical answers.
Integrating Services, Platforms and Resources into a National Infrastructure Cluster for FAIR Language and Cultural Data
Giulia Pedonese | Daniele Melaccio | Michele Mallia | Monica Monachini | Francesca Frontini | Valeria Quochi | Fahad Khan | Angelo Mario Del Grosso | Federico Boschetti | Riccardo Del Gratta
Giulia Pedonese | Daniele Melaccio | Michele Mallia | Monica Monachini | Francesca Frontini | Valeria Quochi | Fahad Khan | Angelo Mario Del Grosso | Federico Boschetti | Riccardo Del Gratta
In the context of evolving European and national policies for research infrastructure governance, this paper presents the contribution of a national consortium for language resources and technology to the construction of a national infrastructure for FAIR and interoperable language and cultural data within a broader Humanities and Heritage Open Science initiative. As the national node of a European research infrastructure for language resources, the consortium contributes to translating FAIR and Open Science principles into practice by integrating technical, methodological, and training dimensions. Its activities combine several coordinated components: FAIRification workflows and ontology-based metadata mediation to enhance semantic interoperability across infrastructures; the refactoring and exposure of services through a federated API gateway; and the implementation of a Linguistic Linked Open Data (LLOD) pilot for the validation, transformation, and publication of interoperable RDF datasets. A national training ecosystem — comprising a training platform and a FAIR learning library — supports capacity building and the creation of FAIR-by-design learning materials. Finally, a permanent research observatory monitors community practices and needs, providing evidence-based insights for the continuous improvement of services and training provision. Together, these components demonstrate a coherent strategy for implementing FAIR and Open Science at the national level, while ensuring alignment with major European and national initiatives in the SSH data ecosystem.
Common European Language Data Space: Development, Current Status, and Future Perspectives
Stelios Piperidis | Penny Labropoulou | Dimitrios Galanis | Khalid Choukri | Andrejs Vasiļjevs | Mitos Deligiannis | Katerina Gkirtzou | Dimitris Gkoumas | Athanasia Kolovou | Leon Voukoutis | Kanella Pouli | Maria Giagkou | Maria Gavriilidou | Katrin Marheinecke | Elena Leitner | Simon Ostermann | Stefania Raccioppa | Kossay Talmoudi | Victoria Arranz | Valérie Mapelli | Helene Mazo | Fernanda González Campo | Shi Yu | Aivars Bērziņš | Andis Lagzdiņš | Georg Rehm
Stelios Piperidis | Penny Labropoulou | Dimitrios Galanis | Khalid Choukri | Andrejs Vasiļjevs | Mitos Deligiannis | Katerina Gkirtzou | Dimitris Gkoumas | Athanasia Kolovou | Leon Voukoutis | Kanella Pouli | Maria Giagkou | Maria Gavriilidou | Katrin Marheinecke | Elena Leitner | Simon Ostermann | Stefania Raccioppa | Kossay Talmoudi | Victoria Arranz | Valérie Mapelli | Helene Mazo | Fernanda González Campo | Shi Yu | Aivars Bērziņš | Andis Lagzdiņš | Georg Rehm
Common European Data Spaces (CEDS) are aimed at creating a single market for data across the EU that will power AI innovation. CEDS cover 14 sectors/domains and will allow secure, trustworthy data/AI models exchange between companies, public administrations etc. The Common European Language Data Space (LDS) is part of CEDS and is already made available in beta phase. The paper presents its technical design and implementation, its governance framework as well as use cases that demonstrate its value. LDS aspires to become part of the future European Language Technology ecosystem.
Euskorpora: A Strategic Framework for Digital Sovereignty and Linguistic Inclusion of Basque in the Era of AI
Victoria Arranz | Sara Arregi | Leire Barañano | Aitor García-Pablos
Victoria Arranz | Sara Arregi | Leire Barañano | Aitor García-Pablos
Euskorpora is a pioneering initiative designed to establish a comprehensive digital infrastructure for the development of speech and language technologies in Basque. Built upon European, Spanish, and Basque strategies, it addresses the scarcity of linguistic data, foundational models, and technological resources for this non-Indo-European, low-resourced language. The project integrates large-scale data collection from public institutions and private organisations, creating extensive multimodal corpora that cover the linguistic, dialectal, and domain diversity of Basque. These resources support the training of open language models for speech, translation, and language understanding, as well as the establishment of an interoperable infrastructure aligned with European initiatives such as the European Language Data Space (LDS). By combining linguistic research, artificial intelligence, and data governance, Euskorpora ensures the digital sovereignty and inclusion of the Basque language within the global AI ecosystem. Beyond its regional focus, it stands as a transferable model for advancing linguistic diversity, technological innovation, and equitable digital transformation in multilingual Europe.
Automating FAIRness: A FAIRification Tool within the Language Resources Infrastructure
Daniele Melaccio | Monica Monachini
Daniele Melaccio | Monica Monachini
In addition to technical interoperability, FAIRness encompasses governance, policy, and ethical aspects, reflecting how language data are produced, represented, and managed within research infrastructures. Ensuring FAIR compliance of language resources is essential for transparent and sustainable research in the social sciences and humanities, enabling data accessibility, quality, and long-term community reuse. The FAIRification Tool — created by CLARIN IT as part of the Humanities and Heritage Italian Open Science Cloud (H2IOSC) — is a modular system that automates and enhances FAIR compliance for language resources. The tool builds upon and extends existing FAIR data assessment frameworks by combining automatic and human validation, a feedback dashboard, certification thresholds, and domain-specific extensions aligned with linguistic metadata standards. It supports FAIR-by-design practices by operationalizing FAIR concepts and embedding them into repository workflows, thereby promoting interoperability across CLARIN, H2IOSC, and EOSC. The tool’s effectiveness has been demonstrated through an initial evaluation conducted on a representative set of linguistic datasets, which revealed notable improvements (30–40%) in FAIR scores, particularly in the Findable and Reusable dimensions, contributing to responsible, policy-aware, and transparent language data management within the European Open Science landscape. demonstrated through an initial evaluation conducted on a representative set of linguistic datasets, which revealed notable improvements (30–40%) in FAIR scores, particularly in the Findable and Reusable dimensions, contributing to responsible, policy-aware, and transparent language data management within the European Open Science landscape.
FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
Evren Ayberk Munis | Deniz Yilmaz | Arianna Muti | Cagri Toraman
Evren Ayberk Munis | Deniz Yilmaz | Arianna Muti | Cagri Toraman
Large language models are widely used across domains, yet there are concerns about their factual reliability and biases. Factual knowledge probing offers a systematic means to evaluate these aspects. Most existing benchmarks focus on single-entity facts and monolingual data. We therefore present FIBER, a multilingual benchmark for evaluating factual knowledge in single- and multi-entity settings. The dataset includes sentence completion, question-answering, and object-count prediction tasks in English, Italian, and Turkish. Using FIBER, we examine whether the prompt language induces inference bias in entity selection and how large language models perform on multi-entity versus single-entity questions. The results indicate that the language of the prompt can influence the model’s generated output, particularly for entities associated with the country corresponding to that language. However, this effect varies across different topics such that 31% of the topics exhibit factual inference bias score greater than 0.5. Moreover, the level of bias differs across languages such that Turkish prompts show higher bias compared to Italian in 83% of the topics, suggesting a language-dependent pattern. Our findings also show that models face greater difficulty when handling multi-entity questions than the single-entity questions. Model performance differs across both languages and model sizes. The highest mean average precision is achieved in English, while Turkish and Italian lead to noticeably lower scores. Larger models, including Llama-3.1-8B and Qwen-2.5-7B, show consistently better performance than smaller 3B-4B models.
EthiQuest: LLM-Powered Ethical Questionnaire Generation for Research Review
Ishank Kapania | Radhika Mamidi | Rahul Mishra
Ishank Kapania | Radhika Mamidi | Rahul Mishra
Building upon the critical importance of ethical considerations in research, we introduce a novel task of Ethical Questionnaire Generation (EQG) for research papers. Ethical review has become an indispensable component of the research process, helping identify potential risks, biases, and societal impacts that may arise from scientific work. In this paper, we present EthiQuest, a comprehensive dataset comprising 3663 research papers paired with their corresponding ethical questionnaires extracted from major conference proceedings. We explore various approaches leveraging large language models (LLMs) to automatically generate context-aware ethical questionnaires, examining the unique challenges of capturing domain-specific ethical concerns, ensuring comprehensive coverage of potential issues, and maintaining question relevance and clarity. Our experiments demonstrate the effectiveness of fine-tuned LLMs in generating pertinent ethical questions across diverse research domains. We provide detailed analysis of question quality, coverage metrics, and practical insights for deploying such systems in real-world research review processes. The EQG dataset and code can be accessed at https://anonymous.4open.science/r/eqg-979C/.
NegNLI-BR: A Brazilian Portuguese Benchmark for Negation in Natural Language Inference
Matheus Westhelle | Viviane Moreira
Matheus Westhelle | Viviane Moreira
Recent studies have questioned the ability of Large Language Models (LLMs) to handle logical negation. We revisit this issue within the Natural Language Inference (NLI) task, specifically investigating whether modern LLMs can distinguish negations that alter logical entailment (“important”) from those that do not (“unimportant”). For this purpose, we introduce NegNLI-BR, a new benchmark dataset in Portuguese designed to exercise this distinction. We evaluate a range of recent open-source LLMs, comparing the performance of their base and post-trained versions. Furthermore, we employ a causal probe to measure the Average Treatment Effect of negation interventions on the internal representations of LLMs. Our findings show that many recent LLMs, including smaller variants, effectively handle negation. The causal analysis reveals that important negations induce a stable and significant effect on model representations, distinct from unimportant negations or neutral filler words. We also observe that post-training generally enhances this representational sensitivity, suggesting it refines the models’ ability to encode the logical impact of negation.
In this paper, we introduce SWE-QA, a text and code corpus aimed at benchmarking multi-hop code comprehension, addressing the gap between simplified evaluation tasks and the complex reasoning required in real-world software development. While existing code understanding benchmarks focus on isolated snippets, developers must routinely connect information across multiple dispersed code segments. The dataset comprises 9,072 multiple-choice questions systematically generated from 12 Python repositories of SWE-bench, evaluating several recurrent reasoning patterns like Declaration-and-Call questions that link entity definitions to their usage, and Interacting-Entity questions that examine the dynamic relationships among multiple collaborating components. Generated through parsing-based entity extraction and Large Language Model assisted question construction with carefully validated distractors, the benchmark distinguishes genuine comprehension from superficial pattern matching. Evaluation of 15 language models (360M to 671B parameters) reveals significant challenges in multi-hop reasoning, with best performance reaching 74.41% accuracy. Dense architectures consistently outperform mixture-of-experts models by 10-14 percentage points, while reasoning-enhanced variants show inconsistent benefits.
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex MultiHop QA
Rishabh Maheshwary | Masoud Hashemi | Khyati Mahajan | Shiva Krishna Reddy Malay | Sai Rajeswar Mudumba | Sathwik Tejaswi Madhusudhan | Spandana Gella | Vikas Yadav
Rishabh Maheshwary | Masoud Hashemi | Khyati Mahajan | Shiva Krishna Reddy Malay | Sai Rajeswar Mudumba | Sathwik Tejaswi Madhusudhan | Spandana Gella | Vikas Yadav
Iterative RAG for multi-hop question answering faces challenges with lengthy contexts and the buildup of irrelevant information. This hinders a model’s capacity to process and reason over retrieved content and limits performance. While recent methods focus on compressing retrieved information, they are either restricted to single-round RAG, require finetuning or lack scalability in iterative RAG. To address these, we propose NotesWriting, a method that generates concise and relevant notes from retrieved documents at each step, thereby reducing noise and retaining only essential information. This increases the effective context length of Large Language Models (LLMs), allowing them to reason and plan more effectively while processing larger volumes of input text due to the compression in the form of notes. NotesWriting is framework agnostic and can be integrated with different iterative RAG methods. We demonstrate its effectiveness with three iterative RAG methods, across two models and four evaluation datasets. NotesWriting yields an average improvement of 15.6 percentage points overall, by scaling the amount of ingested information.
Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA
Renhao Pei | Siyao Peng | Verena Blaschke | Robert Litschko | Barbara Plank
Renhao Pei | Siyao Peng | Verena Blaschke | Robert Litschko | Barbara Plank
Large Language Models (LLMs) are becoming a common way for humans to seek knowledge, yet their coverage and reliability vary widely. Especially for local language varieties, there are large asymmetries, e.g., information in local Wikipedia that is absent from the standard variant. However, little is known about how well LLMs perform under such information asymmetry, especially on closely related languages. We manually construct a novel challenge question-answering (QA) dataset that captures knowledge conveyed on a local Wikipedia page, which is absent from their higher-resource counterparts — covering Mandarin Chinese vs. Cantonese and German vs. Bavarian. Our experiments show that LLMs fail to answer questions about information only in local editions of Wikipedia. Providing context from lead sections substantially improves performance, with further gains possible via translation. Our topical, geographic annotations, and stratified evaluations reveal the usefulness of local Wikipedia editions as sources of both regional and global information. These findings raise critical questions about inclusivity and cultural coverage of LLMs.
FRASE: Frame-based Structured Representations for Generalizable SPARQL Query Generation
Papa Abdou Karim Karou Diallo | Amal Zouaq
Papa Abdou Karim Karou Diallo | Amal Zouaq
Translating natural language questions into SPARQL queries enables Knowledge Base querying for factual and up-to-date responses. However, existing datasets for this task are predominantly template-based, leading models to learn superficial mappings between question and query templates rather than developing true generalization capabilities. As a result, models struggle when encountering naturally phrased, template-free questions. This paper introduces FRASE (FRAme-based Semantic Enhancement), a novel approach that leverages Frame Semantic Role Labeling (FSRL) to overcome this limitation. In addition, we present LCQ1-Frame, LCQ2-Frame, and QALD-10-Frame—a suite of new datasets derived from LC-QuAD 1.0, LC-QuAD 2.0, and QALD-10 where each question is enriched using FRASE through frame detection and the mapping of frame-elements to their corresponding arguments. We evaluate our approach for the Question-2-SPARQL task through extensive experiments using recent large language models (LLMs) under different fine-tuning configurations. Our results demonstrate that integrating frame-based structured representations consistently improves SPARQL generation performance, particularly in challenging generalization scenarios when test questions feature unseen templates (unknown template splits) and when they are all naturally phrased (reformulated questions).
This paper addresses the lack of a multimodal approach to specialized knowledge representation in terminology work. In particular, we introduce a new Multimodal Terminological Metamodel (MTM) for the design of terminology resources which introduces an explicit modality layer, enabling uniform modelling of different language modalities within domain-specific and concept-oriented resources. The metamodel is formalised via an entity-relationship schema and a systematic contrast with the baseline framework – the Terminological Markup Framework (TMF; ISO-16642 (2017)) – to specify revised entities, relations, and cardinalities. As case study, we instantiate the MTM for the signed modality by defining a minimal data-category module with level-placement constraints, and we provide a lightweight, TBX-inspired XML serialisation that packages modality-specific terminological data in a consistent structure. Together, these components deliver a reproducible specification for designing and exchanging multimodal terminology resources.
EPOP: A Benchmark Corpus for Assessing NLP Models on Structured Information Extraction in Plant Health
Claire Nedellec | Marine Courtin | Xinzhi Yao | Marie Grosdidier | Isabelle Pieretti | Sandy Duperier | Robert Bossy
Claire Nedellec | Marine Courtin | Xinzhi Yao | Marie Grosdidier | Isabelle Pieretti | Sandy Duperier | Robert Bossy
We introduce the EPOP (Epidemiomonitoring of Plants) corpus, a new annotated resource for structured information extraction in the domain of plant health epidemiology. The corpus consists of translated news reports that reflect real-world phytosanitary monitoring scenarios. It includes annotations for named entities (e.g. Plant, Pest, Vector, Disease, Dissemination Pathway), identity coreferences, and both binary and complex n-ary relations that represent key events such as Transmits or Causes, along with their modalities. A distinctive feature of EPOP is its normalization layer where mentions of species and geographical locations are linked to canonical identifiers in the NCBI Taxonomy and GeoNames, enabling semantic disambiguation and integration with external knowledge bases. As the first publicly available corpus of its kind, EPOP presents a realistic and challenging benchmark, with high linguistic variability, entity role ambiguity, and long-distance relations. We report baseline results on core tasks (named entity recognition, normalization (entity-linking), and relation extraction) using both fine-tuned BERT-based models and hard-prompted large language models. These experiments demonstrate the utility of EPOP while also identifying areas for improvement, particularly in the extraction of complex relations. The corpus is released under an open license, to support research in environmental NLP, crop protection, and knowledge graph enrichment.
ReTaT: A Unified Benchmark for Relation Extraction across Text and Table
Mohamed Ettaleb | Thibault Ehrhart | Nathalie Aussenac-Gilles | Yoan Chabot | Mouna Kamel | Véronique MORICEAU | Raphael Troncy | Fanfu Wei
Mohamed Ettaleb | Thibault Ehrhart | Nathalie Aussenac-Gilles | Yoan Chabot | Mouna Kamel | Véronique MORICEAU | Raphael Troncy | Fanfu Wei
While prior work in Information Extraction (IE) has focused on extracting information from either textual content or tables in isolation, they miss critical information that emerges only from their interplay. Indeed, tables may summarize facts sparse in the text, while text can disambiguate or elaborate on table entries. This complementarity may take the form of relations which are expressed across text and tables. In this context, we are interested in the task of extracting such relations whose expression spans the two modalities. This task is an original one, for which no reference evaluation corpora exists. Thus we created ReTaT, a corpus that can be used to train and evaluate systems for extracting such relations. This corpus is composed of (table, surrounding text) pairs extracted from Wikipedia pages and has been manually annotated with relation triples. ReTaT is organized in three datasets with distinct characteristics: domain (business, telecommunication and female celebrities), size (from 50 to 255 pairs), language (English vs French), type of relations (data vs object properties), close vs open list of relation, size of the surrounding text (paragraph vs full page). We then assessed its quality and suitability for the joint table-text relation extraction task using Large Language Models (LLMs), at a time when LLMs have demonstrated their ability to extract relations from either text or tables in isolation.
LitTx: A New Treatment Relation Extraction Dataset
Yuhang Jiang | Md Sultan Al Nahian | Li Hao Richie Xu | Rani Chikkanna | Ramakanth Kavuluru
Yuhang Jiang | Md Sultan Al Nahian | Li Hao Richie Xu | Rani Chikkanna | Ramakanth Kavuluru
The interest in biomedical relation extraction (RE) continues to persist even in the LLM era owing to RE being a prominent way to build knowledge graphs, which further ground LLM applications, especially in preventing hallucinations. Therapy-disease treatment relations from scientific literature are an important type in RE as they indicate emerging therapeutic hypotheses and off-label usages being explored in the community. An automatically extracted evolving knowledge-base of such relations will be of great utility to researchers because doing it manually is not viable with the exponential growth of biomedical articles. In this paper, toward this end, we introduce a new expert-annotated dataset LitTx for identifying treatment relationships discussed in literature given the lack of such datasets in the recent past. Besides confirmed or implied positive relations, we also introduce a new “conditional treatment” relation type where hedging or a potential relationship is indicated. Our baseline RE models with this new dataset demonstrate promising results, while also revealing clear areas for improvement. To foster innovation and ensure replicability in the biomedical RE community, we release our dataset, code, and annotation guidelines publicly: https://github.com/bionlproc/LitTx_dataset.
LegitimNarrate: A Dataset for Analyzing Legitimation Mechanisms in Crowdfunding Narratives
Asmaa Lagrid | Sebastien Fournier | Benedicte ALDEBERT | Ali Ghods | Daisy Bertrand | Gael Leboeuf
Asmaa Lagrid | Sebastien Fournier | Benedicte ALDEBERT | Ali Ghods | Daisy Bertrand | Gael Leboeuf
New ventures face challenges due to their liability of newness and need to gain legitimacy within the context of crowdfunding to secure vital resources for growth and survival. Previous studies have primarily assessed crowdfunding success through structured metadata or social media analytics, often neglecting detailed examinations of campaign narratives. To fill this gap, we introduce LegitimNarrate, an expert-annotated dataset specifically designed to analyze legitimation mechanisms in crowdfunding narratives. This dataset comprises 97 Kickstarter campaign descriptions segmented into 4,954 sentences, each meticulously annotated by management experts according to theoretical legitimacy frameworks. We benchmark LegitimNarrate with various contextual sentence-classification methods. This resource facilitates comprehensive research on discursive legitimacy and the role of narrative in crowdfunding contexts.
This paper introduces DASS2019_NLP, a newly cleaned and curated version of the Digital Archive of Southern Speech, a major historical resource for the study of Southern American English, together with six Whisper ASR models fine-tuned on the data. The 344 hours of conversational speech were recorded by fieldworkers between 1969 and 1983 across the Southern United States. Each Whisper model was fine-tuned on DASS2019_NLP, then evaluated on held-out DASS2019_NLP data, a subset of the Corpus of Regional African American Language (CORAAL), and a subset of Common Voice. The fine-tuned models show consistent learning trajectories and achieve an average 37% reduction in WER on in-domain data relative to baseline models. Notably, they also improve transcription accuracy on CORAAL, suggesting enhanced robustness to African American English. As expected under read vs. conversational style mismatch, accuracy on CV generally favors the OpenAI baselines. Both the DASS2019_NLP dataset and the best-performing fine-tuned model (whisper-large-v3-DASS-ct2) have been publicly released. These resources provide new tools for quantitative research in historical sociolinguistics, facilitating large-scale analyses of phonological, lexical, and grammatical change in Southern and African American English.
A Comprehensive Full-Form Lexicon for Arabic NLP and Speech Technology
Yannis Haralambous | Jack Halpern
Yannis Haralambous | Jack Halpern
Natural Language Processing (NLP) applications require morphological data with precise grammatical attributes, while speech technology requires abundant phonemic and phonetic data. This presents a challenge for Arabic due to its abundant morphological, orthographic, and phonemic ambiguity in both MSA and its various dialects. Existing systems struggle with incomplete and unstructured web data, leading to suboptimal performance in both morphological analysis and speech applications. This paper presents ArabLEX, a full-form lexicon (includes all wordforms, i.e., fully inflected/cliticized members of a lexeme class) that addresses these issues by providing a large-scale database designed to enhance NLP accuracy. It comprises approximately 570 million entries with fully inflected forms and detailed morphological, phonetic, and orthographic attributes. ArabLEX serves as a foundational framework for developing comprehensive Arabic lexical resources for NLP, particularly for speech technology, as well as dialect databases.
MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages
Anri M. Lombard | Simbarashe Mawere | Temi Aina | Ethan Wolff | Sbonelo Gumede | Elan Novick | Francois Meyer | Jan Buys
Anri M. Lombard | Simbarashe Mawere | Temi Aina | Ethan Wolff | Sbonelo Gumede | Elan Novick | Francois Meyer | Jan Buys
Decoder-only language models can be adapted to diverse tasks through instruction finetuning, but the extent to which this generalizes at small scale for low-resource languages remains unclear. We focus on the languages of South Africa, where we are not aware of a publicly available decoder-only model that explicitly targets all eleven official written languages, nine of which are low-resource. We introduce MzansiText, a curated multilingual pretraining corpus with a reproducible filtering pipeline, and MzansiLM, a 125M-parameter language model trained from scratch. We evaluate MzansiLM on natural language understanding and generation using three adaptation regimes: monolingual task-specific finetuning, multilingual task-specific finetuning, and general multi-task instruction finetuning. Monolingual task-specific finetuning achieves strong performance on data-to-text generation, reaching 20.65 BLEU on isiXhosa and competing with encoder-decoder baselines over ten times larger. Multilingual task-specific finetuning benefits closely related languages on topic classification, achieving 78.5% macro-F1 on isiXhosa news classification. While MzansiLM adapts effectively to supervised NLU and NLG tasks, few-shot reasoning remains challenging at this model size, with performance near chance even for much larger decoder-only models. We release MzansiText and MzansiLM to provide a reproducible decoder-only baseline and clear guidance on adaptation strategies for South African languages at small scale.
HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
Stephan Oepen | Nikolay Arefyev | Mikko Aulamo | Marta Bañón | Maja Buljan | Laurie V. Burchell | Lucas Georges Gabriel Charpentier | Pinzhen Chen | Mariia Fedorova | Ona de Gibert | Barry Haddow | Jan Hajič | Jindrich Helcl | Andrey Kutuzov | Veronika Laippala | Zihao Li | Bhavitvya Malik | Vladislav Mikhailov | Amanda Myntti | Dayyán O’Brien | Lucie Polakova | Gema Ramírez-Sánchez | Janine Siewert | Pavel Stepachev | Joerg Tiedemann | Teemu Vahtola | Dusan Varis | Fedor Vitiugin | Jaume Zaragoza
Stephan Oepen | Nikolay Arefyev | Mikko Aulamo | Marta Bañón | Maja Buljan | Laurie V. Burchell | Lucas Georges Gabriel Charpentier | Pinzhen Chen | Mariia Fedorova | Ona de Gibert | Barry Haddow | Jan Hajič | Jindrich Helcl | Andrey Kutuzov | Veronika Laippala | Zihao Li | Bhavitvya Malik | Vladislav Mikhailov | Amanda Myntti | Dayyán O’Brien | Lucie Polakova | Gema Ramírez-Sánchez | Janine Siewert | Pavel Stepachev | Joerg Tiedemann | Teemu Vahtola | Dusan Varis | Fedor Vitiugin | Jaume Zaragoza
We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.
Generation of Instruction and Preference Dataset for Improving Japanese Instruction Following in LLMs
Kei Moriyama | Takashi Kodama | Kouta Nakayama
Kei Moriyama | Takashi Kodama | Kouta Nakayama
Instruction following, the ability to generate text that aligns with human intent, is a core capability of large language models (LLMs) for real-world applications. Instruction tuning is widely used to obtain this capability, but it requires large amounts of annotated data. To reduce the labor and cost of large-scale annotation, data augmentation using LLMs has been proposed as a promising approach. As this approach has primarily been applied to English datasets, its effectiveness in other languages, such as Japanese, remains unclear. In this paper, we propose an automatic pipeline for generating instruction and preference datasets in Japanese. The instruction dataset is created by expanding a manually annotated dataset using an LLM. The preference dataset is then constructed by adding LLM-generated negative examples to the instruction dataset. To ensure the quality of the datasets, instructions and responses are evaluated using LLM-as-a-Judge and ROUGE-L. Experimental results using supervised fine-tuning and direct preference optimization demonstrate that these synthetic datasets improve the instruction-following capability in Japanese.
Adapting Pretrained Models to Endangered Languages in Japan: A Comparative Study on Ryukyuan and Ainu Speech Recognition
Kohei Matsuura | Takanori Ashihara | Tatsuya Kawahara
Kohei Matsuura | Takanori Ashihara | Tatsuya Kawahara
We investigate high-accuracy and speaker-robust automatic speech recognition (ASR) models by leveraging pretrained models for endangered languages in Japan — Ryukyuan (Shuri dialect) and Ainu (Saru dialect) — to support language and cultural preservation. In particular, this study presents the first experimental study on building and evaluating an ASR model for the Ryukyuan language. Specifically, we compare existing multilingual pretrained models, Whisper and XLS-R, with our in-house Japanese-focused model (JP-90k) pretrained solely on a large-scale weakly-supervised Japanese dataset. These models were fine-tuned on up to 10 and 32 hours of Ryukyuan and Ainu data, respectively. As a result, JP-90k consistently outperformed other models of the similar size in both languages. In addition, it demonstrated a remarkable advantage when training data was very limited, i.e., an hour or less. These findings suggest that large-scale pretraining on a language closely related to the target ones can yield robust low-resource ASR, including for unseen speakers and out-of-domain conditions. Furthermore, we found that all pretrained models achieved convergence in ASR accuracy with as little as 3-5 hours of fine-tuning data for both languages.
Prerequisites for Advancing Automatic Speech Recognition in Breton
Morgan Grobol | Alice Millour | Wassim Zemouri | Yuna Drapier | Mélanie Jouitteau
Morgan Grobol | Alice Millour | Wassim Zemouri | Yuna Drapier | Mélanie Jouitteau
We report on the extensive preliminary work of a collaborative science project aimed at developing Automatic Speech Recognition (ASR) for a minoritized European language: Breton. Hoping to help similar initiatives for other languages and communities, we present the methodology we developed for this specific ecosystem, with an estimate of the material and immaterial resources we used. Our approach is grounded in the needs and resources of the community formed by the end-users of digital development. Our multidisciplinary scientific collaboration involves linguists and speakers embedded in the academic and linguistic community, and computer scientists.
Integrating TEI, NER/NEL, Textometry, and Linked Data for a Semantically Enriched Interview Corpus
Ranka Stankovic | Tamara Vučenović | Biljana Rujević | Milica Ikonić Nešić | Mihailo Škorić
Ranka Stankovic | Tamara Vučenović | Biljana Rujević | Milica Ikonić Nešić | Mihailo Škorić
This paper presents a pipeline that converts unstructured interview transcripts into a semantically enriched, queryable knowledge resource. The texts from the Digitalne Ikone 20+ interview collection were first encoded in TEI XML (Text Encoding Initiative), marking interview boundaries, paragraph breaks, speaker turns with identifiers, dates, and topics. This structural encoding underpins downstream NLP and enables structured querying (e.g., by speaker). We then applied Named Entity Recognition to identify persons, places, organizations, and events, and embedded the results directly in TEI. In the third stage, Named Entity Linking mapped entity mentions to canonical Wikidata identifiers via context-aware disambiguation; missing entries were added to Wikidata when necessary. The resulting TEI+NER/NEL corpus, serialized as linked data, follows the NIF (NLP Interchange Framework). The pipeline also supports retrieval-augmented summarization that retrieves evidence passages and prompts LLMs (implemented with DSPy) to produce faithful interview summaries. We discuss design choices (TXM for textometry with JeRTeh resources; TESLA models for NER/NEL), report qualitative gains in interpretability through semantic links, and outline future work on domain-adapted NER/NEL, graph-based completion, and more expressive RAG architectures. The approach is replicable for other oral-history or media corpora and advances practical, evidence-grounded access to cultural archives and beyond.
Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages
Edward Thomas Bayes | Israel Abebe Azime | Jesujoba Alabi | Jonas Kgomo | Tyna Eloundou | Elizabeth Proehl | Kai Chen | Imaan Khadir | Naome A. Etori | Shamsuddeen Hassan Muhammad | Choice Mpanza | Igneciah Pocia IP Thete | Dietrich Klakow | David Ifeoluwa Adelani
Edward Thomas Bayes | Israel Abebe Azime | Jesujoba Alabi | Jonas Kgomo | Tyna Eloundou | Elizabeth Proehl | Kai Chen | Imaan Khadir | Naome A. Etori | Shamsuddeen Hassan Muhammad | Choice Mpanza | Igneciah Pocia IP Thete | Dietrich Klakow | David Ifeoluwa Adelani
Evaluations of Large Language Models (LLMs) on knowledge-intensive tasks and factual accuracy often focus on high-resource languages primarily because datasets for low-resource languages (LRLs) are scarce. In this paper, we present Uhura—a new benchmark that focuses on two tasks in six typologically-diverse African languages, created via human translation of existing English benchmarks. The first dataset, Uhura-ARC-Easy, is composed of multiple-choice science questions. The second, Uhura-TruthfulQA, is a safety benchmark testing the truthfulness of models on topics including health, law, finance, and politics. We highlight the challenges creating benchmarks with highly technical content for LRLs and outline mitigation strategies. Our evaluation reveals a significant performance gap between proprietary models such as GPT-4o and o1-preview, and Claude models, and open-source models like LLaMA and Gemma. Additionally, all models perform better in English than in African languages. These results indicate that LLMs struggle with answering scientific questions and are more prone to generating false claims in low-resource African languages. Our findings underscore the necessity for continuous improvement of multilingual LLM capabilities in LRL settings to ensure safe and reliable use in real-world contexts. We open-source the Uhura Benchmark and Uhura Platform to foster further research and development in NLP for LRLs.
Dialectal Filtering: Synthesizing Kurdish Corpora for Low-Resource Varieties by Utilizing "Noise" in Large Textual Data
Christian Schuler | Raman Ahmad | Ānrán Wáng | Daniil Gurgurov | Timo Baumann | Simon Ostermann | Josef van Genabith
Christian Schuler | Raman Ahmad | Ānrán Wáng | Daniil Gurgurov | Timo Baumann | Simon Ostermann | Josef van Genabith
This work introduces a dialect-aware text filtering framework to pre-process, clean, and enhance large text corpora, creating variety-specific sub-corpora for neglected language varieties. We apply our framework to Kurdish, a language with rich dialectal diversity, which presents significant challenges for Natural Language Processing due to its low-resource status and the noisy nature of available text corpora. Leveraging lexicographic features, we assign multi-language-labels to text instances and synthesize over 130 dialect specific corpora from large “noisy” data sets containing unlabeled mixtures of Kurdish varieties, representing to our knowledge the largest collection of dialect-specific Kurdish NLP resources to date. This work contributes to the creation of low-resource language technology foundations, especially dialect-specific NLP applications. Specifically, we advance research on Kurdish languages by providing insights into the linguistic relationships among Kurdish varieties.
HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection
Luke S. Patterson | Li Wang | Adam Faulkner
Luke S. Patterson | Li Wang | Adam Faulkner
Thanks to the rapid adoption of AI code assistants powered by large language models (LLMs), industry codebases are, increasingly, a hybrid of AI- and human-authored code. For risk management and productivity analysis purposes, it is crucial to enable fine-grained location detection of AI-generated code. To develop algorithms for this task, quality benchmarks are needed to assess performance. However, existing benchmarks tend to comprise academic, LeetCode-style problems and presume a code snippet is either completely human-authored or completely AI-authored, which is not reflective of the diverse intents and styles of industry codebases utilizing AI code assistants. To fill these gaps, we introduce HybridCodeAuthorship, a novel benchmark of Python code files with interleaved human- and AI-authored lines of code to simulate authentic utilization of AI code assistants. In this paper, we first present our dataset construction pipeline, which leverages CodeSearchNet, a massive collection of links to open sourced repositories on GitHub. We then benchmark the performance of two state-of-the-art AI-generated code detection algorithms at both the line- and chunk-level. Experimental results demonstrate that HybridCodeAuthorship is a challenging benchmark with a top-scoring algorithm, AIGCode Detector, obtaining a highest F1 score of 0.48 and 0.56 on line-level and chunk-level code detection tasks, respectively.
CorEGe-PT: Compiling a Large Corpus of Academic Texts in Portuguese
Tanara Zingano Kuhn | José Matos | Bruno Neves | Daniela Pereira | Elisabete Cação | Ivo Simões | Jacinto Estima | Delfim Leão | Hugo Goncalo Oliveira
Tanara Zingano Kuhn | José Matos | Bruno Neves | Daniela Pereira | Elisabete Cação | Ivo Simões | Jacinto Estima | Delfim Leão | Hugo Goncalo Oliveira
This paper describes the creation of a large-scale corpus of academic texts in Portuguese, dubbed CorEGe-PT, extracted from the institutional repository of a Portuguese university. Its compilation methodology, which combined automatic and manual procedures, is detailed, together with challenges faced and proposed solutions. The process included a thorough analysis of the metadata, which will be publicly released together with the documents, extracted in a markdown format. CorEGe-PT covers five areas of knowledge and, with over 34,000 documents and 1B tokens, is the largest of corpus of its kind in Portuguese, which will enable in-depth linguistic studies while providing data for adapting Large Language Models to academic Portuguese and related tasks.
SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding
Haroun Elleuch | Salima Mdhaffar | Yannick Estève | Fethi Bougares
Haroun Elleuch | Salima Mdhaffar | Yannick Estève | Fethi Bougares
Spoken Language Understanding (SLU) aims to extract the semantic information from the speech utterance of user queries. It is a core component in a task-oriented dialog system. With the spectacular progress of deep neural network models and the evolution of pre-trained language models, SLU has obtained significant breakthroughs. However, only a few high-resource languages have taken advantage of this progress due to the absence of SLU resources. In this paper, we seek to mitigate this obstacle by introducing SLURP-TN. This dataset was created by recording 55 native speakers uttering sentences in Tunisian dialect, manually translated from six SLURP domains. The result is an SLU Tunisian dialect dataset that comprises 4165 sentences recorded into around 5 hours of acoustic material. We also develop a number of Automatic Speech Recognition and SLU models exploiting SLUTP-TN. The Dataset and baseline models are available at: https://huggingface.co/datasets/Elyadata/SLURP-TN.
From Semi-Digital Edition to Historical NLP Resource:Constructing and Annotating Historical Multilingual Parallel Text Collections on the TEITOK Platform
Maarten Janssen | Anna Jouravel | Piroska Lendvai
Maarten Janssen | Anna Jouravel | Piroska Lendvai
We construct a multilingual, parallelized digital collection comprising a reconstructed Old Greek text from the 4th century CE and its seven historical versions, modern editions, and translations. We describe the workflow and integrated tools on the TEITOK web-based platform for ingesting, aligning, parallelizing and morphosyntactically annotating these materials. Textual alignment is performed on both the sentence and word level, after which the data are annotated with dependency parses in the Universal Dependencies paradigm. The newly created and manually post-corrected collection can be explored via advanced parallel search functionalities and flexible visualization modes. This workflow is meant to provide support for digital humanities and historical NLP projects via transforming the input texts into parallel NLP resources, enabling cross-fertilization and new insights by multiple research communities.
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
Máté Gedeon | Piroska Zsófia Barta | Peter Mihajlik | Tekla Etelka Graczi | Anna Kohári | Katalin Mády
Máté Gedeon | Piroska Zsófia Barta | Peter Mihajlik | Tekla Etelka Graczi | Anna Kohári | Katalin Mády
The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepresented due to limited spontaneous and conversational corpora. To address this gap, we introduce two new datasets – BEA-Large and BEA-Dialogue – constructed from the previously unprocessed portions of the Hungarian speech corpus named BEA. BEA-Large extends BEA-Base with 255 hours of spontaneous speech from 433 speakers, enriched with detailed segment-level metadata. BEA-Dialogue, comprising 85 hours of spontaneous conversations, is a Hungarian speech corpus featuring natural dialogues partitioned into speaker-independent subsets, supporting research in conversational ASR and speaker diarization. We establish reproducible baselines on these datasets using publicly available ASR models, with the fine-tuned Fast Conformer model achieving word error rates as low as 14.18% on spontaneous and 4.8% on repeated speech. Diarization experiments yield diarization error rates between 12.46% and 17.40%, providing reference points for future improvements. The results highlight the persistent difficulty of conversational ASR, particularly due to disfluencies, overlaps, and informal speech patterns. By releasing these datasets and baselines, we aim to advance Hungarian speech technology and offer a methodological framework for developing spontaneous and conversational benchmarks in other languages.
Developing the German Medical Text Corpus (GeMTeX): Legal Compliance and Semantic Enrichment
Justin Hofenbitzer | Christina Lohr | Andrea Riedel | Rebekka Kiser | Aliaksandra Shutsko | Abanoub Abdelmalak | Peter Klügl | Jutta Romberg | Sarah Riepenhausen | Miriam Schechner | Jakob Faller | Frank Meineke | Luise Modersohn | Markus Löffler | Juliane Fluck | Udo Hahn | Stefan Schulz | Martin Boeker
Justin Hofenbitzer | Christina Lohr | Andrea Riedel | Rebekka Kiser | Aliaksandra Shutsko | Abanoub Abdelmalak | Peter Klügl | Jutta Romberg | Sarah Riepenhausen | Miriam Schechner | Jakob Faller | Frank Meineke | Luise Modersohn | Markus Löffler | Juliane Fluck | Udo Hahn | Stefan Schulz | Martin Boeker
GeMTeX is a large-scale German Medical Text Corpus project with the goal to publish a clinical national reference corpus. The resource is currently under construction and comprises, as of February 2026, more than 15k clinical documents (20M tokens) from six German university hospitals. When building GeMTeX, attention was paid to comply with European regulatory requirements. In phase I, patients were asked to allow reuse of their clinical documents based on the legal foundation of an “informed consent”. In phase II, consented documents from six major clinical sites in Germany underwent a thorough de-identification process. In phase III, we currently enrich this unlocked dataset with semantic information from the clinical domain. This annotation process is guided by Snomed CT, which supports to directly ground expressions within clinical documents in a worldwide shared medical documentation and ontology standard. The resource is currently under active development and is accessible upon request under controlled access conditions. We refer interested researchers to visit https://kiinformatik.mri.tum.de/en/gemtex or reach out via gemtex.mi@mh.tum.de.
MaiChat: A Text-based Dialogue Corpus Rich in Conversational Features
Mai Hoang Dao | Catherine Lai | Peter Bell
Mai Hoang Dao | Catherine Lai | Peter Bell
We present a new English corpus of typed instant-messaging dialogues that includes detailed timing information. Messages are collected from interactions between pairs who know each other well; the corpus is rich in typed features that augment the purely lexical, including hesitations, self-corrections, expressive respellings, and other markers of spontaneous interaction. Messages are collected using a custom-built chat platform that logs not only message content but also keystroke dynamics, screen activity, and demographic metadata. Designed with a transparent and reproducible protocol, the corpus enables scalable data collection while ensuring privacy and consent. We intend that the rich collection of features collected will facilitate future research in areas such as cognitive modelling, human–computer interaction, and conversational AI.
Saudi ASWAT: A Large-Scale Corpus of Spontaneous Saudi Arabic Speech
Abdullah I. Alharbi | Afrah A. Altamimi | Muneera Alhoshan | Amal Almazrua | Halah Munif Alharbi | Bayan M. Almuqhim | Hawra Aljasim | Abdulrahman Alosaimy | Yahya A. Asiri | Abdullah Alfaifi
Abdullah I. Alharbi | Afrah A. Altamimi | Muneera Alhoshan | Amal Almazrua | Halah Munif Alharbi | Bayan M. Almuqhim | Hawra Aljasim | Abdulrahman Alosaimy | Yahya A. Asiri | Abdullah Alfaifi
Spontaneous Arabic speech is scarce in current corpora, and it is not well represented. This poses a limitation invisibility of spontaneous Arabic to automatic speech recognition (ASR), speaker diarization, and sociolinguistic research. The Saudi ASWAT project fills a major gap by creating the first nationwide corpus of natural Saudi speech, where data has been recorded and transcribed under a systematic methodology and ecologically valid conditions. The corpus aims to collect 2,500 hours of natural conversations from a diverse range of participants. These has been selected from five major Saudi regional varieties, Najdi (Central), Eastern, Hijazi (Western), Northern, and Southern, covering more than fifty five local varieties. Speech has been recorded by trained fieldworkers using participants own devices to reflect real-life variation. The annotated data incorporate a variety of speaker demographics, regional vocabularies which differ from the standard lexicon, and structured metadata. TF–IDF profiling shows regional differences in a range of performing words. Data also represent balanced age and gender sampling to support studies of intergenerational and sociophonetic variation. Saudi ASWAT provides the most linguistically diverse resources of Saudi Arabia to date. Additionally, it establishes an ethical governed framework for Arabic speech data creation to enable advances in both computational modeling and linguistic research.
SciCiteVal: A Multi-Domain Dataset for Scientific Citation Verification
Qinyue Liu | Yongxin Zhou | Cyril Labbe
Qinyue Liu | Yongxin Zhou | Cyril Labbe
Citations are an integral and important part of scientific papers. However, there exist erroneous citations ranging from careless mistakes to deliberate misconduct, and there are currently few studies or benchmark datasets dedicated to automated citation verification. To bridge this gap, we introduce SciCiteVal, a novel, manually annotated dataset for citation verification. Each instance in SciCiteVal pairs a citation context from a citing paper with the corresponding evidence passage extracted from the full text of the cited source. The dataset features a comprehensive taxonomy, where each citation is annotated as "Correct”, "Incorrect”, or "Unrelated”, with the "Incorrect” category further divided into five fine-grained sub-categories. The completed dataset comprises over 1,000 annotated citations, distributed as 302 "Correct”, 302 "Incorrect”, and 430 "Unrelated” instances. We establish a benchmark by evaluating different Large Language Models (LLMs), providing baseline performance and a detailed analysis. We release SciCiteVal as a resource to support the development of citation verification systems and to facilitate research on evidence-based tasks.
RuznamceNER: A Named Entity Recognition Dataset for Ottoman Turkish
Esma Fatıma Bilgin Tasdemir | Dilara Zeynep Gürer | Saziye Betul Ozates
Esma Fatıma Bilgin Tasdemir | Dilara Zeynep Gürer | Saziye Betul Ozates
Named Entity Recognition (NER) in historical texts poses distinct challenges. Language change reflected in spelling variations, archaic vocabulary, and inconsistent orthography, diminish the efficacy of models trained on contemporary corpora. The limited availability of annotated historical datasets constrains the development and evaluation of accurate, domain-specific NER systems, underscoring the necessity for specialized approaches and domain adaptation. In this work, we introduce the ruznamçe registers as a valuable digital historical resource with broad potential for diverse NLP applications. Our primary contribution is RuznamceNER, a manually annotated NER dataset derived from ruznamçe documents spanning two centuries. The dataset contains 2,138 sentences and a total of 8,730 annotated entities of types PERSON, LOCATION and ORGANIZATION. We further report evaluation results using a BERT-CRF baseline model pre-trained with modern Turkish, highlighting the pivotal importance of in-domain training data for effective NER in historical contexts. Experimental results on the RuznamceNER test set under various training configurations show that even a small amount of supervised in-domain data can yield robust performance for well-structured texts, despite significant lexical and orthographic differences between historical and modern language forms
Scripting History: A Diachronic Urdu Text and Image Corpus from the 18Th to 19Th Centuries
Sana Shams | Sahar Rauf | Asad Mustafa | Muhammad Zeeshan Javed | Qurat-ul-Ain Akram | Sarmad Hussain | Miriam Butt
Sana Shams | Sahar Rauf | Asad Mustafa | Muhammad Zeeshan Javed | Qurat-ul-Ain Akram | Sarmad Hussain | Miriam Butt
This paper presents the Diachronic Urdu Text and Image Corpus, a one-million-word resource covering Urdu’s development across the 18th and 19th centuries. The corpus is compiled from 328 printed books published between 1800 and 1950, representing a diverse range of genres, authors, and publishers. A 140,000-word sub-corpus has been manually annotated with Urdu part-of-speech tags to facilitate linguistic and computational analysis. The dataset enables systematic investigation of historical changes in Urdu orthography, morphology, and syntax, providing new insights into the language’s history and standardization. To preserve the original printed form, each text is paired with its corresponding page image, creating the first multimodal diachronic corpus for Urdu. The paper outlines the corpus compilation pipeline, digitization workflow, text-image alignment, and annotation strategy designed to ensure accuracy, consistency, and authenticity. This multimodal Urdu diachronic corpus establishes a benchmark for research in computational linguistics, digital humanities, and South Asian language technology, supporting corpus-based exploration of Urdu’s linguistic history and cultural heritage.
Easy Read (ER) text adaptation is one of the main means to provide accessible content for people with reading difficulties. ER text features aspects of text simplification, along with specific characteristics such as the need for short sentences, clearly structured content, and explanations for complex concepts. Support for ER text generation is still lacking overall, with few available resources to build automated systems upon. In this work, we describe the IREKIER corpus, based on ER news in Basque and Spanish from the Irekia transparency portal of the Basque Government. This corpus is currently one of the largest publicly shared resource to support training and evaluation of ER text adaptation models in these two languages, and the first of its kind for Basque. We describe our methodology to create the resource, along with the specific challenges raised by ER text. We also provide both intrinsic and extrinsic evaluations of the corpus, which is shared with the scientific community under a CC-BY-NC-ND 4.0 license.
MekongPhon: A Large-Scale Parallel IPA Corpus for Lao and Khmer
Ammon Shurtz | Christian Richardson | Stephen D. Richardson
Ammon Shurtz | Christian Richardson | Stephen D. Richardson
High-quality International Phonetic Alphabet (IPA) transcriptions are a foundational resource for speech and language technologies, yet existing tools for many low-resource languages remain limited in accuracy and scope. In this work, we present MekongPhon, a large-scale, high-quality parallel IPA corpus for Lao and Khmer. The corpus contains 1.3 million Khmer and 367 thousand Lao orthographic–IPA pairs, meticulously aligned and verified. When used to train Transformer-based sequence-to-sequence models, MekongPhon enables exceptionally accurate IPA generation, achieving under 2% Character Error Rate (CER) on held-out test sets. We further introduce linguistically informed Lao and Khmer transliteration tools that offer high-speed IPA conversion, outperforming Epitran by 6-71 CER points despite trading some accuracy for efficiency. All data, code, and pretrained models are publicly released to support future research and development in low-resource language technologies.
CorSpell: Introducing a Semiautomatic Tool for Spelling Normalization in Brazilian Portuguese
Juliana Schoffen | Dennis Giovani Balreira | Elisa Marchioro Stumpf | Larissa Goulart | Tanara Zingano Kuhn | Rafael Oleques Nunes | Gabriel Ricci Pazzinato | Isadora Dahmer Hanauer | José Henrique de Souza Silva | Luiza Sarmento Divino | Marine Matte
Juliana Schoffen | Dennis Giovani Balreira | Elisa Marchioro Stumpf | Larissa Goulart | Tanara Zingano Kuhn | Rafael Oleques Nunes | Gabriel Ricci Pazzinato | Isadora Dahmer Hanauer | José Henrique de Souza Silva | Luiza Sarmento Divino | Marine Matte
With the growing availability of large text collections, efficient tools for corpus annotation and normalization have become increasingly important in linguistic and computational research. This paper presents CorSpell, a semiautomatic tool developed to support the spelling normalization of Brazilian Portuguese texts within the CorCel project—a corpus comprising over 15,000 handwritten exam responses from the Celpe-Bras proficiency test. Given the corpus scale, manual normalization is impractical; CorSpell streamlines this process by enabling users to visualize, select, and replace tokens directly through an intuitive web interface. The tool integrates automatic suggestions from PT-BR dictionaries with human validation, providing an interface for users to access and manipulate the texts. CorSpell significantly reduces annotation time, minimizes errors, and facilitates collaborative work, providing a practical and scalable solution for corpus normalization and a foundation for LLM-based modeling of Portuguese proficiency.
Meta4XNLI-ptBR: Brazilian Portuguese Extension of Meta4XNLI Corpus
Karina Johansson | Fernanda Assi | Isabella da Silva | Rafael Passador | Isabela Rodrigues | Aline Paes | Helena Caseli
Karina Johansson | Fernanda Assi | Isabella da Silva | Rafael Passador | Isabela Rodrigues | Aline Paes | Helena Caseli
Metaphor is a pervasive phenomenon in language that shapes how people conceptualize and communicate complex ideas. Detecting and interpreting metaphor is not only relevant for linguistic theory but also for many Natural Language Processing (NLP) applications, from machine translation to sentiment analysis, to mention a few. Despite its relevance, no open-source annotated corpus of metaphors exists for one of the world’s most widely spoken languages: Brazilian Portuguese. This paper addresses this gap by presenting an extension of Meta4XNLI, Meta4XNLI-ptBR, with token-level metaphor annotation in Brazilian Portuguese. To achieve this, we propose a pipeline that combines automatic translation via language models with human annotation, following guidelines adapted from MIPVU and Meta4XNLI. The final corpus contains 1,784 human-annotated sentences, of which 42.26% contain at least one metaphorical token. To our knowledge, this is the first open corpus of its kind for Brazilian Portuguese, and it is already freely available.
More than "Oh": Grounding Observable Events with Grunts in Multimodal Dialogue
Richard A. Brutti | James Pustejovsky
Richard A. Brutti | James Pustejovsky
Conversational grunts (minimal vocalizations like oh, mm-hm, and uh-huh) ground information and coordinate understanding in human dialogue, yet computational systems typically treat them as noise rather than meaningful communicative acts. We present a systematic annotation and analysis of 497 grunts across 3 hours of multimodal collaborative tasks, introducing an annotation scheme that captures grunts, their antecedents, and dialogue act functions. Our analysis reveals that grunts respond to speech and observable events at nearly equal rates, demonstrating that non-verbal events function as conversational contributions requiring acknowledgment. Tokens exhibit functional specialization: mm-hm predominantly acknowledges speech, while oh preferentially acknowledges events. Prosodic analysis shows speakers systematically modulate duration and pitch based on antecedent type, with event responses typically longer and having greater range. These findings have implications for dialogue state tracking, multimodal grounding, and turn-taking in conversational AI systems.
COME-ALPs: Coreference Annotation with MErging Heuristics Using ALignment-based Projection in Parallel Corpora
Gabriela Nicole Gonzalez Saez | Mariam Nakhle | Illia Kholosha | Rachel Atherly | Marco Dinarelli
Gabriela Nicole Gonzalez Saez | Mariam Nakhle | Illia Kholosha | Rachel Atherly | Marco Dinarelli
Multi-lingual, parallel datasets annotated with discourse phenomena like coreferences are a rare resource. These datasets are useful and informative to evaluate models for NLP tasks taking long contextual information into account, as proved by the large literature published in the last couple of years on e.g. Context-Aware Neural Machine Translation (CA-NMT). Inspired by resources published in previous work, in this paper we propose an automated procedure to annotate multi-lingual, parallel data with coreferences. Through the use of accurate alignment and coreference annotation tools, we project the annotation from English data, where tools are most often more accurate, to one or more target languages. We apply some consistency constraints to obtain more accurate annotations on both source and target side. Using our procedure we generated two new resources that can be used for evaluating CA-NMT models. One starting from the well-known TED Talk’s data released for the IWSLT17 shared task, where we project the annotation from English to target languages as diverse as French, German and Chinese. The second resource is derived from the WMT24 shared task, consisting of news domain data in the same set of target languages. We release these resources, as well as the code framework for applying our annotation procedure, to the community.
MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning
Zimu Wang | Yuqi Wang | Tong Chen | Changyu Zeng | Hongbin Na | Nijia Han | Fuyu Xing | Qi Chen | Qiufeng Wang | Anh Nguyen | Shuihua Wang | Ling Chen | Jionglong Su | Haiyang Zhang | Wei Wang
Zimu Wang | Yuqi Wang | Tong Chen | Changyu Zeng | Hongbin Na | Nijia Han | Fuyu Xing | Qi Chen | Qiufeng Wang | Anh Nguyen | Shuihua Wang | Ling Chen | Jionglong Su | Haiyang Zhang | Wei Wang
Event understanding and reasoning play critical roles in thoroughly evaluating the capabilities of Vision-Language Models (VLMs); however, existing Visual Question Answering (VQA) datasets predominantly focus on entity-centric questions, while event- or action-related questions are limited in scale and suffer from significant shortcut issues. We introduce MEUR, the first Multimodal Event Understanding and Reasoning dataset consisting of 1,200 images and 4,217 questions, necessitating VLMs with a diverse range of multimodal understanding and reasoning capabilities to answer, ranging from basic event recognition to more complex tasks such as counting and comparison. To streamline the annotation process, we propose a novel semi-automated pipeline that combines advanced VLMs with human annotators, achieving high quality and efficiency. We conduct extensive experiments on state-of-the-art non-thinking and thinking VLMs to demonstrate their capabilities and limitations in multimodal event understanding and reasoning. Furthermore, we provide a detailed error analysis that points out promising directions for future research.
Building Collaborative Speech Corpora for Low-Resource Languages: The Galician Dataset in Mozilla Common Voice
Adina Ioana Vladu | Elisa Fernández Rei | María Pérez Lago
Adina Ioana Vladu | Elisa Fernández Rei | María Pérez Lago
This paper presents the methodology and outcomes of building collaborative speech corpora in Mozilla Common Voice (MCV), focusing on the Galician case within Proxecto Nós. We describe the organization of voice collection campaigns –on-site events, student participation, Validatón marathons, and corporate collaboration– and analyze the results in MCV v22.0. While the dataset has achieved a modest scale, major gaps remain in metadata completeness and dialectal tagging, with implications for ASR performance. Drawing on our experience, we highlight effective strategies for engagement, such as transparent communication, cultural identification, and user-friendly tools. We conclude with lessons learnt for improving data representativeness, participant retention, and ethical governance. The observations are specific to the Galician case study but may inform similar efforts in other lesser-resourced languages.
Frame-Guided Synthetic Claim Generation for Automatic Fact-Checking Using High-Volume Tabular Data
Jacob Devasier | Akshith Putta | Qing Wang | Alankrit Moses | Chengkai Li
Jacob Devasier | Akshith Putta | Qing Wang | Alankrit Moses | Chengkai Li
Automated fact-checking benchmarks have largely ignored the challenge of verifying claims against real-world, high-volume structured data, instead focusing on small, curated tables. We introduce a new large-scale, multilingual dataset to address this critical gap. It contains 78,503 synthetic claims grounded in 434 complex OECD tables, which average over 500K rows each. We propose a novel, frame-guided methodology where algorithms programmatically select significant data points based on six semantic frames to generate realistic claims in English, Chinese, Spanish, and Hindi. Crucially, we demonstrate through knowledge-probing experiments that LLMs have not memorized these facts, forcing systems to perform genuine retrieval and reasoning rather than relying on parameterized knowledge. We provide a baseline SQL-generation system and show that our benchmark is highly challenging. Our analysis identifies evidence retrieval as the primary bottleneck, with models struggling to find the correct data in massive tables. This dataset provides a critical new resource for advancing research on this unsolved, real-world problem.
A Bilingual Bimodal Benchmark for Arabic-English NLP across Grammatical Correction, Essay Scoring, Morphological Tagging, and Speech Recognition
Bashar Alhafni | Injy Hamed | Fadhl Eryani | David Palfreyman | Nizar Habash
Bashar Alhafni | Injy Hamed | Fadhl Eryani | David Palfreyman | Nizar Habash
Building comprehensive datasets that support a variety of NLP tasks and cover a diversity of languages and domains is vital for NLP evaluation purposes. In this paper, we present ZAEBUC*, a dataset that builds upon and enriches prior corpora with new annotations and benchmarking experiments. ZAEBUC* serves as a benchmark for a range of NLP tasks, including grammatical error correction, automated essay scoring, automatic speech recognition, and morphological tagging, which includes tokenization, part-of-speech tagging, and lemmatization. The dataset covers Arabic and English in both written and spoken forms, offering a bilingual and bimodal resource. Furthermore, the corpus brings together a collection of resources gathered from a similar population, enabling cross-linguistic and cross-modal comparisons. We provide benchmarking results, demonstrating the performance of NLP models, including LLMs, across various tasks, languages, and modalities.
Developing a Guideline for the Labovian-Structural Analysis of Oral Narratives in Japanese
Amane Watahiki | Tomoki Doi | Akari Kikuchi | Hiroshi Ohata | Yuki I. Nakata | Takuya Niikawa | Taiga Shinozaki | Hitomi Yanaka
Amane Watahiki | Tomoki Doi | Akari Kikuchi | Hiroshi Ohata | Yuki I. Nakata | Takuya Niikawa | Taiga Shinozaki | Hitomi Yanaka
Narrative analysis is a cornerstone of qualitative research. One leading approach is the Labovian model, but its application is labor-intensive, requiring a holistic, recursive interpretive process that moves back and forth between individual parts of the transcript and the transcript as a whole. Existing Labovian datasets are available only in English, which differs markedly from Japanese in terms of grammar and discourse conventions. To address this gap, we introduce the first systematic guidelines for Labovian narrative analysis of Japanese narrative data. Our guidelines retain all six Labovian categories and extend the framework by providing explicit rules for clause segmentation tailored to Japanese constructions. In addition, our guidelines cover a broader range of clause types and narrative types. Using these guidelines, annotators achieved high agreement in clause segmentation (Fleiss’ kappa = 0.80) and moderate agreement in two structural classification tasks (Krippendorff’s alpha = 0.41 and 0.45, respectively), one of which is slightly higher than that found in prior work despite the use of finer-grained distinctions. This paper describes the Labovian model, the proposed guidelines, the annotation process, and their utility. It concludes by discussing the challenges encountered during the annotation process and the prospects for developing a larger dataset for structural narrative analysis in Japanese qualitative research.
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
Jens Rupprecht | Leon Froehling | Claudia Wagner | Markus Strohmaier
Jens Rupprecht | Leon Froehling | Claudia Wagner | Markus Strohmaier
The use of Large Language Models (LLMs) for simulating human perspectives via persona prompting is gaining traction in computational social science. However, well-curated, empirically grounded persona collections remain scarce, limiting the accuracy and representativeness of such simulations. Here, we introduce the German General Social Survey Personas (GGSS Personas) collection, a comprehensive and representative persona prompt collection built from the German General Social Survey (ALLBUS). The GGSS Personas and their persona prompts are designed to be easily plugged into prompts for all types of LLMs and tasks, steering models to generate responses aligned with the underlying German population. We evaluate GGSS Personas by prompting various LLMs to simulate survey response distributions across diverse topics, demonstrating that GGSS Personas-guided LLMs outperform state-of-the-art classifiers, particularly under data scarcity. Furthermore, we analyze how representativity and attribute selection within persona prompts affect alignment with population responses. Our findings suggest that GGSS Personas provide a potentially valuable resource for research on LLM-based social simulations that enables more systematic explorations of population-aligned persona prompting in NLP and social science research.
Slovene Morphological and Word Formation Segmentation: A Novel Dataset and Evaluation
Marko Pranjić | Boris Kern | Ines Voršič | Senja Pollak
Marko Pranjić | Boris Kern | Ines Voršič | Senja Pollak
We introduce the first publicly available manually annotated dataset for morphological segmentation and word-formation analysis for Slovene, containing 1,935 words annotated by two domain experts. The dataset provides three types of linguistic information: morphological and word-formation segments with zero-morpheme and simplex annotations. We present a four-stage annotation approach achieving inter-annotator agreement of 86.80% Krippendorff’s Alpha for morphological segmentation and 85.16% for word-formation segments. Computational validation using a morphological segmentation model achieves 87.78% BPR F1 on morphological segmentation and 83.05% on word-formation segments. Despite being smaller than previous datasets derived from non-public esources, our dataset enables high performance and supports reproducible research for morphological analysis tools for Slovene.
GePaDeU - a Multi-layer Corpus of German Parliamentary Debates with Rich Semantic and Pragmatic Annotations
Ines Rehbein | Julian Schlenker | Lars Ostertag | Simone Paolo Ponzetto
Ines Rehbein | Julian Schlenker | Lars Ostertag | Simone Paolo Ponzetto
This paper presents GePaDeU, a new manually annotated corpus of German Parliamentary Debates with Unified layers of semantic and pragmatic information. The data includes parliamentary speeches from the German Bundestag, ranging over a time period from 2017–2021, with 267 speeches given by 197 members of parliament. The final release of our corpus unifies multiple annotation layers, including entity-level annotations, the annotation of speech events and their corresponding speakers, functional speech acts, clause-level aspect, and moral framing. We provide an overview of the various annotation layers and illustrate how the semantic and pragmatic annotations can be combined for corpus-linguistic studies and discourse analyses, and to answer research questions in the field of political science. The new resource will be made freely available for the research community.
What Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading Experience
Filip Miletić | Neele Falk
Filip Miletić | Neele Falk
Has the style of scientific communication changed due to the growing use of large language models in the writing process? We address this question in the domain of Natural Language Processing by leveraging two data resources we create: a naturalistic corpus of over 37,000 papers from the ACL Anthology (2020–2024); and a synthetic dataset of 3,000 human-written passages and their LLM-generated improvements. We first implement a series of diachronic lexical analyses, showing that both word frequency and usage contexts have changed significantly over time, indicating semantic specialization in some cases and generalization in others. Broadening our perspective, we then model a range of more complex stylistic features and find that LLM-modified texts more frequently contain certain syntactic constructions, more complex and longer words and a lower lexical diversity. Finally, we connect these changes in writing practices to subjective reading experience through a pilot annotation study with 20 domain experts. They overall rate LLM-improved texts as more understandable and exciting, but also express negative qualitative attitudes towards LLMs, highlighting the strongly subjective effect of AI-assisted writing on reading experience.
GeneFRDebate: Generated French Debates from News Articles with Industrial-Expert Summaries
Rim Abrougui | Guillaume Lechien | Elisabeth Savatier | Benoît Laurent
Rim Abrougui | Guillaume Lechien | Elisabeth Savatier | Benoît Laurent
Summarizing domain-specific conversations, such as political debates, remains challenging despite advances in large language models (LLMs), and resources for French debates are particularly limited. We present GeneFRDebate, a new dataset of synthetic French political debates generated from real-world news articles using an LLM, while keeping expert-written summaries unchanged. Our pipeline combines prompt engineering, human curation, and quality evaluation using both automatic metrics and expert assessment. We also provide baseline experiments with small-scale LLMs (≤8B parameters), demonstrating the dataset’s usefulness for training and evaluation. This work shows that carefully generated synthetic data with human oversight can complement existing corpora, supporting research in multilingual and domain-specific dialogue summarization.
AmbiCoRefVis: A Tool for Visualizing Coreferential Ambiguity
Patrick Paetzold | Lukas Beiske | Mark-Matthias Zymla | Massimo Poesio | Miriam Butt | Daniel Weiskopf | Oliver Deussen
Patrick Paetzold | Lukas Beiske | Mark-Matthias Zymla | Massimo Poesio | Miriam Butt | Daniel Weiskopf | Oliver Deussen
Situations of ambiguity and uncertainty in the annotation of discourse interpretation tasks, such as anaphoric reference, are common, but existing annotation tools typically only support visualization at the local level (i.e., visualizing more than one mention of a possible antecedent) rather than globally (i.e., visualizing multiple coreference chains), as the latter is a complex problem. In this paper, we introduce the interactive visual analysis tool AmbiCoRefVis, developed to display multiple global interpretations of a referring expression. We evaluate it with the Phrase Detectives corpus.
Fables-DTR: A Corpus of Fables Annotated for Discourse and Temporal Relations
Purificação Silvano | António Leal | Maciej Ogrodniczuk | Aleksandra Tomaszewska | Joana Gomes | Luís Filipe Cunha | Evelin Amorim | Martyna Lewandowska | Anna Śliwicka | Alípio Jorge
Purificação Silvano | António Leal | Maciej Ogrodniczuk | Aleksandra Tomaszewska | Joana Gomes | Luís Filipe Cunha | Evelin Amorim | Martyna Lewandowska | Anna Śliwicka | Alípio Jorge
This paper presents Fables-DTR, a corpus of Aesop’s fables annotated for discourse and temporal relations, designed to explore how event sequencing and aspectual features and discourse relations interact. Building on the ISO 24617 Semantic Annotation Framework, integrating Part 1 (Time and Events) and Part 8 (Discourse Relations), the resource provides a unified representation of discourse structure and temporal and aspectual features. The corpus comprises 15 fables in English, automatically translated into European Portuguese and Polish (45 texts in total), with all translations manually validated by native linguists to preserve semantic and discourse features. Each fable is annotated in two layers: (i) for discourse relations, argument roles, and signals; (ii) for temporal relations, and event attributes, such as Tense, Aspect, Polarity. The resulting dataset provides relevant information about the association between discourse relations and their temporal and aspectual features. Fables-DTR contributes both a valuable resource for cross-linguistic and narrative discourse analysis and empirical evidence for integrating ISO standards in multilayer annotation. It also provides a foundation for computational applications in discourse parsing, event ordering, and implicit relation detection.
A Benchmark Corpus for the Diagnostic Assessment of Content in L2 English Speech
Kosuke Doi | Justin Vasselli | Taro Watanabe
Kosuke Doi | Justin Vasselli | Taro Watanabe
When evaluating second language (L2) learners’ speech, human raters pay significant attention to its content, and diagnostic feedback on content helps improve learners’ speaking ability. Since human scoring and feedback are time-consuming and costly, automatic models aiming to provide such feedback have been developed, specifically models that detect whether certain content, i.e., key points, is included in learner’s speech. However, previous studies target only integrated test items where learners speak based on listened or read materials, and the data used are not publicly available. In this study, we construct a speech corpus for key point detection. We extend the target to test items where learners speak based on their own experiences and opinions, which show greater content diversity than integrated test items, using an approach that annotates content along with its connections. Analysis of the constructed data demonstrated that the annotated elements are associated with the speech content scores. We also found that large language models are generally successful at locating content element spans, although their predicted spans are often broader than human-annotated ones. The corpus and annotation guidelines are available at https://language.sakura.ne.jp/icnale/download.html.
Insights from Romanized Manipuri Social Media Text: A Transliteration Corpus and Variation Analysis
Maisang Kamei Salice | Sanasam Ranbir Singh | Priyankoo Sarmah
Maisang Kamei Salice | Sanasam Ranbir Singh | Priyankoo Sarmah
This paper presents the first large-scale study of Romanized Manipuri, a low-resource Indic language widely used by native speakers on social media. Social media text is highly informal and often noisy, posing challenges for natural language processing tasks; therefore, normalization through back-transliteration is essential. We construct a Romanized Manipuri to Manipuri–Bengali script back-transliteration corpus from YouTube comments, capturing diverse informal writing styles and orthographic variations. The dataset is analyzed to examine variation patterns at two levels: character-level inconsistencies and pragmatic stylistic variations influenced by user writing behavior. We also compare social media romanization with formal transliteration conventions, including standardized romanization schemes and textbook-based systems. Furthermore, we evaluate Transformer model at both character and subword levels and conduct a detailed error analyses to identify key challenges affecting back-transliteration performance.
MELD: Melding Diverse Multilingual and Multi-Domain Datasets for Named Entity Recognition Evaluation
Kevin Glocker | Marco Kuhlmann
Kevin Glocker | Marco Kuhlmann
Zero-shot Named Entity Recognition (NER) has gained prominence for information extraction across diverse domains without being limited to a single, fixed tag set. However, existing NER resources vary widely in data format, licensing terms, annotation schemes, and availability, making it difficult to systematically evaluate the generalization capabilities of zero-shot NER models. Prior attempts to aggregate datasets with broad coverage across domains have largely focused on a small subset of languages, and it is often not transparent how datasets were processed from their sources. This paper introduces MELD, a comprehensive multilingual and multi-domain data collection designed to address these gaps. MELD integrates 60 NER datasets spanning 194 languages, 14 domains, and 601 normalized entity types. While previously introduced multilingual NER datasets are mainly silver-standard, MELD contains gold-standard annotations for 60 languages. All data processing steps are fully open-source and reproducible, facilitating future extensions and ensuring long-term accessibility. While MELD is primarily designed for zero-shot evaluation, it also provides training and development splits in a single, consistent format to support future research in few-shot and supervised NER settings.
FinER-ABSA: A Benchmark for Implicit and Explicit Entity Recognition and Aspect-Based Sentiment Analysis in Financial News
Pachara Akkanwanich | Pavorn Thongyoo | Mahannop Thabua | Konlakorn Wongpatikaseree | Natthawut Kertkeidkachorn
Pachara Akkanwanich | Pavorn Thongyoo | Mahannop Thabua | Konlakorn Wongpatikaseree | Natthawut Kertkeidkachorn
Many approaches to English financial text analysis still rely on keyword or rule-based extraction, with limited trust in sentiment models despite advances in contextual understanding. Past studies have explored concepts such as aspect-based sentiment analysis and named entity recognition, yet none address how entities appear implicitly through context rather than direct mentions, or provide a dataset that brings these elements together. This gap limits how well models capture the links between entities, aspect, and sentiment. We introduce FinER-ABSA, a benchmark that integrates implicit and explicit entity recognition with aspect-based sentiment in financial text. Experiments on seven open-source large language models under zero- and few-shot settings show that even the best systems still miss key aspects of implicit reasoning. In the few-shot case (K= 3), Llama-3.3-70B reached an F1 of 0.7623 for implicit entities, suggesting that while models can detect signals, their consistency remains far from the level of reliability required for financial analysis or decision-making. These insights emerge only through FinER-ABSA, which makes such gaps measurable and advances financial Natural Language Processing (NLP) toward deeper contextual understanding and enables systems that better extract comprehensive insights from market-moving information in an industry where such precision is critical.
MUSIA: Multilingual Story Illustration Corpus for Cross-Cultural Alignment and Generation
Krishna Tewari | Supriya Chanda | Nirmit Patil | Sukomal Pal
Krishna Tewari | Supriya Chanda | Nirmit Patil | Sukomal Pal
Recent advances in text-to-image generation have enabled automated visual storytelling, yet most existing datasets remain monolingual and culturally narrow. We introduce MUSIA, a Multilingual Story Illustration Corpus designed to advance research in cross-lingual and culturally grounded narrative illustration. MUSIA comprises bilingual (English-Hindi) story-image pairs drawn from open literary and folk sources, curated to reflect diverse cultural themes, artistic styles, and linguistic structures. Each story includes multiple illustrations aligned at the scene level, accompanied by quality-verified mappings for narrative-visual coherence. To establish a reproducible benchmark, we propose a two-stage baseline combining transformer-based semantic summarization with diffusion-based image generation, achieving strong performance in relevance, visual quality, and consistency. MUSIA represents the first step toward a scalable, culturally inclusive benchmark for multilingual visual storytelling, enabling fair and reproducible research across low-resource and underrepresented languages.
MUDiC: A Dataset for Multi-User Dialogue and Collaboration in Chatbot Interaction
Nicolas Wagner | Cristina Luna Jimenez | Elisabeth Andre | Wolfgang Minker | Stefan Ultes
Nicolas Wagner | Cristina Luna Jimenez | Elisabeth Andre | Wolfgang Minker | Stefan Ultes
We introduce MUDiC, a novel dataset on task-based multi-user interactions in chatbots. Unlike most traditional dialogue corpora that focus on one-to-one human–chatbot exchanges, this dataset captures conversations involving two human participants engaging with a single system. The data include diverse conversational contexts such as shared group task, user intents, and mechanisms to deal with off-topic talk. MUDiC consists of 1,689 dialogue exchanges between 20 groups and the chatbot. Each session is annotated with user id, interaction turns, and intents and dialogue acts, enabling an analysis of group conversational dynamics. Consequently, the dataset aims to support tasks such as multi-user dialogue modelling, intent disambiguation, and moderation behaviour, which are relevant factors for the design of socially aware chatbots.
StoryCCDial: Collecting and Analyzing Human-Human Co-Creation Dialogues for Personalized Creative Support
Natsumi Ezure | Michimasa Inaba
Natsumi Ezure | Michimasa Inaba
With the development of generative models, research on human-AI co-creation has been actively conducted. However, in the field of co-creation, research on system personalization according to individual characteristics is insufficient, and little focus has been placed on individual differences in creation. Therefore, in this study, we constructed StoryCCDial, a co-creation dialogue dataset aimed at the personalization of co-creative dialogue systems. First, we collected human-human story co-creation dialogue data involving 120 workers and constructed a dataset that includes dialogues, dialogue acts, the workers’ personality traits, postsurveys, and edit histories from the interface. Next, using the constructed dataset, we conducted analyses focusing on the workers’ personality traits, the number of utterances, and edit histories. The analysis revealed differences in dialogue content based on workers’ personality traits, individual differences in the number of utterances during the co-creation process, and variations in creative workflows on the interface. Our dataset will be available at https://github.com/UEC-InabaLab/StoryCCDial .
DATASHI: A Parallel English–Tashlhiyt Corpus for Orthography Normalization and Low-Resource Language Processing.
Nasser-Eddine Monir | Zakaria Baou
Nasser-Eddine Monir | Zakaria Baou
DATASHI is a new parallel English–Tashlhiyt corpus that fills a critical gap in computational resources for Amazigh languages. It contains 5,000 sentence pairs, including a 1,500-sentence subset with expert-standardized and non-standard user-generated versions, enabling systematic study of orthographic diversity and normalization. This dual design supports text-based NLP tasks—such as tokenization, translation, and normalization—and also serves as a foundation for read-speech data collection and multimodal alignment. Comprehensive evaluations with state-of-the-art Large Language Models (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Mistral, Qwen3-Max) show clear improvements from zero-shot to few-shot prompting, with Gemini 2.5 Pro achieving the lowest word and character-level error rates and exhibiting robust cross-lingual generalization. A fine-grained analysis of edit operations—deletions, substitutions, and insertions—across phonological classes (geminates, emphatics, uvulars, and pharyngeals) further highlights model-specific sensitivities to marked Tashlhiyt features and provides new diagnostic insights for low-resource Amazigh orthography normalization.
Evaluating Social Intelligence in LLMs via Japanese Honorifics in Email Generation: A Social Semiotic System Perspective
Muxuan Liu | Tatsuya Ishigaki | Yusuke Miyao | Hiroya Takamura | Ichiro Kobayashi
Muxuan Liu | Tatsuya Ishigaki | Yusuke Miyao | Hiroya Takamura | Ichiro Kobayashi
We propose JaSocial, a novel evaluation framework that leverages Japanese emails to comprehensively evaluate large language models’ (LLMs) social intelligence across varied social-status relationships. The framework integrates three core components. First, we construct and publicly release a meticulously human-annotated Japanese email dataset covering six distinct social-status contexts, thereby capturing nuanced shifts in social hierarchy and politeness. Second, we adopt Systemic Functional Linguistics (SFL)—a social-semiotic linguistic theory that explicitly models how linguistic choices realize interpersonal relations and hierarchical distinctions—to classify email content in terms of three perspectives: social relationships, speech functions, and honorific expressions. Based on these perspectives, we design an automated evaluation method that assigns each LLM-generated email a contextual appropriateness score, quantifying how well it reflects socially intelligent behavior. Third, we release the full evaluation code to ensure reproducibility and enable fair cross-model comparisons. JaSocial exposes current LLMs’ limitations in capturing cultural nuance, while providing an open benchmark for future research.
Do Language Models Know Theo Has a Wife? Investigating the Proviso Problem
Tara Azin | Daniel Dumitrescu | Diana Inkpen | Raj Singh
Tara Azin | Daniel Dumitrescu | Diana Inkpen | Raj Singh
We investigate how language models handle the proviso problem, an unresolved issue in pragmatics where presuppositions in conditional sentences diverge between theoretical and human interpretations. We reformulate this phenomenon as a Natural Language Inference task and introduce a diagnostic dataset designed to probe presupposition projection in conditionals. We evaluate RoBERTa, DeBERTa, LLaMA, and Gemma using explainability analyses. The results show that models broadly align with human judgments but rely on shallow pattern matching rather than semantic or pragmatic reasoning. Our work provides the first computational evaluation framework for the proviso problem and highlights the need for diagnostic, multi-method approaches to assess pragmatic competence and context-dependent meaning in language models.
Cross-Lingual and Cross-Cultural Transfer of Talk Move Classification to German Science Classrooms
Christian Wartena | Christian Schumburg | Andreas Nehring | Marcel Ebert | Friederike Korneck | David Schmitt | Marie Irmer | Birgit Neuhaus
Christian Wartena | Christian Schumburg | Andreas Nehring | Marcel Ebert | Friederike Korneck | David Schmitt | Marie Irmer | Birgit Neuhaus
Talk moves are discourse categories used to analyse classroom interactions. They provide insights into the types of exchanges between teachers and students and can serve as indicators of teaching quality, supporting feedback and reflection. The automatic classification of talk moves is therefore valuable for educational research and teacher development. While previous studies have explored this task, almost all have focused on English data. We constructed a small corpus of German science classroom transcripts and investigated whether multilingual language models can classify talk moves effectively under data-scarce conditions. Specifically, we examined (1) training with a very limited amount of German data and (2) cross-lingual transfer from English training data, which also entails cross-cultural adaptation. Our results show that multilingual large language models are capable of cross-lingual and cross-cultural transfer, but models trained directly on even a small amount of German data achieve better performance. Combining English and German data yields the best results overall, though the additional benefit of including English data is small.
IHPP: A Paragraph-Level Dataset for Investigating the Pragmatics of Hyperpartisan Italian News
Michele Joshua Maggini | Davide Bassi | Angelo Valente | Gaël Dias | Pablo Gamallo
Michele Joshua Maggini | Davide Bassi | Angelo Valente | Gaël Dias | Pablo Gamallo
This study investigates the linguistic composition of hyperpartisan paragraphs in Italian news on climate change, Ukraine war, and immigration by publicly disclosing the dataset to ensure reproducibility. We introduce a new corpus, IHPP, of 356 articles, for a total of 4,861 paragraphs annotated for hyperpartisan news detection at the paragraph level and enriched with span-level annotations of six semantic-pragmatic linguistic traits: figurative speech, irony/sarcasm, epithet, as well as hyperbolic and loaded language. We hypothesized that these traits, while violating Gricean maxims, are key mechanisms of hyperpartisan rhetoric. To test this, we fine-tuned a set of mono- and multilingual BERT models for hyperpartisan detection and evaluated their incorporation in the embedding space. Then, we applied explainable techniques, e.g. Integrated Gradients and SHAP to analyze how models allocate attribution to normal and linguistic-trait tokens. Our result show that loaded language is the most discriminative trait. The dataset is released: https://github.com/MichJoM/IHPP-Climate.
Detecting Potentially Under-annotated Explicit Discourse Connectives in the Penn Discourse Treebank (PDTB-3) with LLMs
Yueh-Ting Chuang | Xixian Liao | Bonnie Webber
Yueh-Ting Chuang | Xixian Liao | Bonnie Webber
Accurate identification of explicit discourse connectives is crucial for analysing discourse relations, which supports NLP tasks such as summarisation and question answering. However, annotation inconsistencies remain a challenge, particularly for ambiguous prepositions with both discourse and non-discourse usages. This paper presents a pipeline that leverages large language model (LLM) prompting, cross-model agreement, and syntactic pattern analysis to detect likely under-annotated connectives. Evaluated on four prepositions (by, with, without and for), the approach effectively identifies likely under-annotations for some, but not all prepositions. Results show that while the method is promising, its generalisability depends on improved prompt design, model choice, and syntactic analysis tools. The findings highlight both the potential and limitations of LLM-based approaches for corpus error detection and demonstrate how improved discourse annotation can contribute to more reliable data for downstream NLP tasks.
Can LLMs Understand Punchlines? LLMs’ Narrative Understanding Evaluation with Short-shorts
Jiashi Cheng | Takehito Utsuro
Jiashi Cheng | Takehito Utsuro
In this study, we constructed a narrative comprehension benchmark using the works of Shinichi Hoshi to examine the extent to which Large Language Models (LLMs) can understand twist endings, or punchlines, in short-short stories. Specifically, story endings were categorized into six types—such as Revelation, Apocalypse, and Sarcasm—and a classification task was designed in which LLMs were prompted with the story text and asked to select the appropriate ending category. We collected human annotations from eight native Japanese speakers to establish a reference benchmark. Experimental comparisons were conducted across multiple LLMs (GPT-4, Claude, Gemini, and Grok), assessing their performance both at the metric level and at the discourse level against human judgments. The results revealed that although certain models approached human performance in specific categories, overall accuracy remained notably lower than the human baseline. Through quantitative and qualitative analyses, this study highlights the challenges LLMs face in capturing narrative subtleties such as irony, implication, and emotional reversal. The proposed benchmark provides a novel framework for evaluating narrative understanding and the deeper semantic reasoning capabilities of LLMs.
Building the AURIS Corpus of Reference and Information Structure
Christian Chiarcos | Christian Fäth | Tabea Gröger | Quentin Alastair Frey
Christian Chiarcos | Christian Fäth | Tabea Gröger | Quentin Alastair Frey
We present AURIS, the Augsburg corpus for Reference and Information Structure, a multilingual corpus annotated for reference, discourse relations, and aspects of information structure. AURIS introduces an innovative use of off-the-shelf spreadsheet software for complex annotation tasks, reducing technical barriers and dependencies common in discourse annotation. Designed for classroom use, it enables linguistics and philology students to explore diverse theoretical frameworks while working in their language of choice. The paper focuses on technical design and workflows that integrate and generate pre-annotations from heterogeneous sources. Despite its low-tech approach, AURIS aligns with established standards and remains interoperable with existing projects. Preprocessing scripts support multiple languages, with an initial annotation round on German texts evaluated against TED-MDB and ParCorFull data converted into AURIS formats. This approach demonstrates that accessible tools can yield high-quality, replicable annotations for discourse and information-structure research.
There Is No Spoon: Existential Presupposition in Large Language Models
Marie-Léontine Wörgötter | Shikai Lai | Sebastian Schuster
Marie-Léontine Wörgötter | Shikai Lai | Sebastian Schuster
Existential presupposition is a foundational component of meaning: it reflects implicit assumptions of existence that underlie interpretation, even when not explicitly stated. Sentences such as Neo bends the spoon presuppose that the entities referred to exist, independent of the truth-value of the sentence itself. Because this type of meaning is implied rather than explicitly asserted, it provides a diagnostic test of whether large language models (LLMs) display sensitivity to more abstract and less surface-driven layers of meaning. We adapt a natural language inference (NLI)–based probing setup, using a fine-tuned version of DeBERTa-v3-large as a baseline model and compare its behaviour to that of LLaMA-3.1-8B-Instruct and Gemma-3-12B-it under zero- and few-shot prompting, as well as to their fine-tuned base-variants. We find that while all models show sensitivity to existential presupposition across syntactic embeddings, determiner types and contextual cues, their behaviour differs markedly in strength and systematicity, with NLI-fine-tuned autoregressive models exhibiting the most coherent and stable projection patterns. They showed graded and theoretically aligned projection patterns, whereas instruction-tuned models remain largely prone to surface heuristics and prompt susceptibility. These results suggest that pre-trained LLMs exhibit sensitivity to existential presupposition but this behaviour surfaces only systematically when the models have learned the intricacies of the NLI task.
DiscoRAG: A Discourse-Aware Agent for Query-Based Summarization of Long Documents
Alexander Chernyavskiy | Lidiia Ostyakova | Dmitry Ilvovsky
Alexander Chernyavskiy | Lidiia Ostyakova | Dmitry Ilvovsky
Query-based summarization of long documents is often tackled with retrieval-augmented generation (RAG). However, conventional RAG models exhibit limitations when applied to narrative texts, where crucial evidence is often implicit and distributed. This exposes a distinct class of “discourse-aware” queries that require specialized, structure-aware models. To address this, we introduce DiscoRAG, a framework that leverages Rhetorical Structure Theory (RST). By modeling the document as a discourse tree, DiscoRAG navigates its structure, explicitly using rhetorical relations to focus on and aggregate evidence from globally related segments. Furthermore, our pipeline integrates a classifier that assesses query complexity to dynamically select the most efficient retrieval strategy. We evaluate our DiscoRAG against standard and extended-context RAG pipelines on the SQuALITY dataset, which we release augmented with questions requiring deep discourse reasoning and integration of the global narrative. Our results demonstrate that this method sizeably outperforms these baselines, demonstrating its superior ability to assemble a coherent, contextually rich evidence base by interpreting the global narrative structure rather than relying on local semantic similarity.
In-Distribution Steering: Balancing Control and Coherence in Language Model Generation
Arthur Vogels | Benjamin Wong | Yann Choho | Annabelle Blangero | Milan Bhan
Arthur Vogels | Benjamin Wong | Yann Choho | Annabelle Blangero | Milan Bhan
Activation steering methods control large language model (LLM) behavior by modifying internal activations at inference time. However, most existing activation steering methods rely on a fixed steering strength, leading to either insufficient control or unadapted intervention that degrades text plausibility and coherence. We introduce In-Distribution Steering (IDS), a novel method that adapts steering strength based on the input data distribution in representation space. IDS dynamically adjusts interventions according to how far a given input lies within the distribution, enabling adaptive intervention and generation stability during text generation. Experiments demonstrate that IDS achieves strong accuracy on classification tasks while producing coherent text without collapse, making IDS particularly well suited for real-world applications.
Improving Multilingual Language Models by Aligning Representations through Steering
Omar Mohamed Mahmoud | Buddhika Laknath Semage | Thommen George Karimpanal | Santu Rana
Omar Mohamed Mahmoud | Buddhika Laknath Semage | Thommen George Karimpanal | Santu Rana
This paper investigates how Large Language Models (LLMs) represent non-English tokens—a question that remains underexplored despite recent progress. We propose a lightweight intervention method using representation steering, where a learned vector is added to the residual stream at a single model layer to enhance multilingual performance. Through extensive experiments across seven competitive baselines—including prompt optimization, supervised fine-tuning (SFT), in-context learning, cross-lingual transfer, projection mapping techniques, and translation-based methods—we show that our approach consistently outperforms most alternatives. In particular, it achieves performance on par with production-grade translation systems while requiring far fewer resources. We further explore the complementarity between our method and SFT, demonstrating that steering offers a direct, efficient way to realign internal representations. These findings underscore the potential of activation-level interventions as a powerful tool for improving the multilingual capabilities of LLMs.
Explainable AI for Ethical Counter Speech Generation in Hate Speech Mitigation
Ashiful Islam Ridoy | Mohammed Faisal | Yogesh Kumar | Md Mamun-Ur Rashid | Marina Ernst | Frank Hopfgartner
Ashiful Islam Ridoy | Mohammed Faisal | Yogesh Kumar | Md Mamun-Ur Rashid | Marina Ernst | Frank Hopfgartner
The proliferation of hate speech in digital communication platforms poses significant challenges to online safety and social cohesion. While automated hate speech detection systems have shown promise, their black-box nature limits user trust and understanding of AI-driven content moderation decisions. This paper presents a framework that integrates explainable AI (XAI) techniques with counter-speech generation to create transparent, ethical solutions for hate speech mitigation. Our approach combines a fine-tuned HateBERT model, with a specialized Llama 3.1-8B-Instruct model for generating empathetic counter-narratives. The system employs five distinct XAI methods: Integrated Gradients, Attention Visualization, LIME, Counterfactual Analysis, and Natural Language Explanations to provide interpretable reasoning behind both detection and response generation decisions. The integration of explainability mechanisms with counter-speech generation represents a novel contribution to ethical AI systems, fostering transparency and trust in automated hate speech mitigation while maintaining high performance standards for real-world deployment.
Do Language Models Encode Semantic Relations? Probing and Sparse Feature Analysis
Andor Diera | Ansgar Scherp
Andor Diera | Ansgar Scherp
Understanding whether large language models (LLMs) capture structured meaning requires examining how they represent concept relationships. In this work, we study three models of increasing scale: Pythia-70M, GPT-2, and Llama 3.1 8B, focusing on four semantic relations: synonymy, antonymy, hypernymy, and hyponymy. We combine linear probing with mechanistic interpretability techniques, including sparse autoencoders (SAE) and activation patching, to identify where these relations are encoded and how specific features contribute to their representation. Our results reveal a directional asymmetry in hierarchical relations: hypernymy is encoded redundantly and resists suppression, while hyponymy relies on compact features that are more easily disrupted by ablation. More broadly, relation signals are diffuse but exhibit stable profiles: they peak in the mid-layers and are stronger in post-residual/MLP pathways than in attention. Difficulty is consistent across models (antonymy easiest, synonymy hardest). Probe-level causality is capacity-dependent: on Llama 3.1, SAE-guided patching reliably shifts these signals, whereas on smaller models the shifts are weak or unstable. Our results clarify where and how reliably semantic relations are represented inside LLMs, and provide a reproducible framework for relating sparse features to probe-level causal evidence.
The Sufficiency-Conciseness Trade-off in LLM Self-Explanation from an Information Bottleneck Perspective
Ali Zahedzadeh | Behnam Bahrak
Ali Zahedzadeh | Behnam Bahrak
Large Language Models increasingly rely on self-explanations, such as chain of thought reasoning, to improve performance on multi step question answering. While these explanations enhance accuracy, they are often verbose and costly to generate, raising the question of how much explanation is truly necessary. In this paper, we examine the trade-off between sufficiency, defined as the ability of an explanation to justify the correct answer, and conciseness, defined as the reduction in explanation length. Building on the information bottleneck principle, we conceptualize explanations as compressed representations that retain only the information essential for producing correct answers. To operationalize this view, we introduce an evaluation pipeline that constrains explanation length and assesses sufficiency using multiple language models on the ARC Challenge dataset. To broaden the scope, we conduct experiments in both English, using the original dataset, and Persian, as a resource-limited language through translation. Our experiments show that more concise explanations often remain sufficient, preserving accuracy while substantially reducing explanation length, whereas excessive compression leads to performance degradation.
We present a practical framework for detecting errors in LLM-generated SQL by estimating uncertainty at the level of individual nodes in the query’s abstract syntax tree (AST). Our approach proceeds in two stages. First, we introduce a semantically aware labeling algorithm that, given a generated SQL and a gold reference, assigns node-level correctness without over-penalizing structural containers or alias variation. Second, we represent each node with a rich set of schema-aware and lexical features - capturing identifier validity, alias resolution, type compatibility, ambiguity in scope, and typo signals - and train a supervised classifier to predict per-node error probabilities. We interpret these probabilities as calibrated uncertainty, enabling fine-grained diagnostics that pinpoint exactly where a query is likely to be wrong. Across multiple databases and datasets, our method substantially outperforms token log-probabilities: average AUC improves by +27.44% while maintaining robustness under cross-database evaluation. Beyond serving as an accuracy signal, node-level uncertainty supports targeted repair, human-in-the-loop review, and downstream selective execution. Together, these results establish node-centric, semantically grounded uncertainty estimation as a strong and interpretable alternative to aggregate sequence-level confidence measures.
A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs
Anna Feldman | Libby Barak | JIng Peng
Anna Feldman | Libby Barak | JIng Peng
We introduce a typology-aware diagnostic for multilingual masked language models that tests reliance on word order versus inflectional form. Using Universal Dependencies, we apply inference-time perturbations: full token scrambling, content-word scrambling with function words fixed, dependency-based head–dependent swaps, and sentence-level lemma substitution (+L), which lemmatizes both the context and the masked target label. We evaluate mBERT and XLM-R on English, Chinese, German, Spanish, and Russian. Full scrambling drives word-level reconstruction accuracy near zero in all languages; partial and head–dependent perturbations cause smaller but still large drops. +L has little effect in Chinese but substantially lowers accuracy in German/Spanish/Russian, and it does not mitigate the impact of scrambling. Top-5 word accuracy shows the same pattern: under full scrambling, the gold word rarely appears among the five highest-ranked reconstructions. We release code, sampling scripts, and balanced evaluation subsets; Turkish results under strict reconstruction are reported in the appendix.
From Generation to Evaluation: A Resource for Error-Categorized Question Generation from Video Transcripts
Joshua Berger | Markos Stamatakis | Anett Hoppe | Ralph Ewerth | Christian Wartena
Joshua Berger | Markos Stamatakis | Anett Hoppe | Ralph Ewerth | Christian Wartena
A key challenge in automated question generation is producing grammatically correct, error-free, and contextually relevant questions. While large language models already handle this well, smaller models that can run on consumer-grade hardware face greater difficulties. Another obstacle is the lack of large, high-quality datasets, particularly for education video transcripts, which limits the diversity and applicability of training data. On top of this, current evaluation methods either rely on strict comparison to a “ground truth,” undervaluing valid but unmatched questions, or on expert judgments, which do not scale. They do not provide insights into the nature of errors. In this paper, we introduce a dataset of real-life educational video transcripts and investigate the question-generating capabilities of small language models by assessing their output with pre-defined error categories. We also present a novel approach to automatic quality assessment by classifying questions into predefined error categories. We show that questions generated by small language models are still prone to error. Our proposed classification approach outperforms baseline approaches and matches GPT-5 performance by reaching an accuracy of 72%.
From Behavior to Geometry: A Causal and Geometric Analysis of LoRA-Based Domain Adaptation
Yizhe WANG | Liu He | Zhenhua Ling
Yizhe WANG | Liu He | Zhenhua Ling
Parameter-efficient fine-tuning with Low-Rank Adaptation (LoRA) often improves a large language model’s in-domain performance at the cost of cross-domain generalization. We investigate the mechanistic basis for this trade-off, asking whether LoRA creates new discriminative directions in representation space (emergence) or merely reshapes pre-existing ones. Using a Word Sense Disambiguation testbed, we couple controlled behavioral evaluation with causal localization and geometric diagnostics. We find LoRA learns new, spatially localized discriminative directions in the middle layers of the network, focused at token positions critical for the task. This “subspace extension” account explains why LoRA-tuned models excel on in-domain data but struggle to transfer. As a proof of concept, we introduce a mechanistically informed LoRA configuration that concentrates capacity in the identified layers, promotes rank diversity, and applies light answer-token calibration. Without increasing training budget, it yields consistent improvements in both in- and cross-domain settings, demonstrating that mechanistic insight can guide more efficient adaptation.
Explainable Semantic Textual Similarity via Dissimilar Span Detection
Diego Miguel Lozano | Daryna Dementieva | Alexander Fraser
Diego Miguel Lozano | Daryna Dementieva | Alexander Fraser
Semantic Textual Similarity (STS) is a crucial component of many Natural Language Processing (NLP) applications. However, existing approaches typically reduce semantic nuances to a single score, limiting interpretability. To address this, we introduce the task of Dissimilar Span Detection (DSD), which aims to identify semantically differing spans between pairs of texts. This can help users understand which particular words or tokens negatively affect the similarity score, or be used to improve performance in STS-dependent downstream tasks. Furthermore, we release a new dataset suitable for the task, the Span Similarity Dataset (SSD), developed through a semi-automated pipeline combining large language models (LLMs) with human verification. We propose and evaluate different baseline methods for DSD, both unsupervised—based on LIME, SHAP, LLMs, and our own method—as well as an additional supervised approach. While LLMs and supervised models achieve the highest performance, overall results remain low, highlighting the complexity of the task. Finally, we set up an additional experiment that shows how DSD can lead to increased performance in the specific task of paraphrase detection.
BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic Reasoning
Ha Thanh Nguyen | Hideyuki Tachibana | Chaoran Liu | Qianying Liu | Su Myat Noe | Koichi Takeda | Sadao Kurohashi
Ha Thanh Nguyen | Hideyuki Tachibana | Chaoran Liu | Qianying Liu | Su Myat Noe | Koichi Takeda | Sadao Kurohashi
We present BIS Reasoning 1.0, the first large-scale Japanese dataset of syllogistic reasoning problems explicitly designed to evaluate belief-inconsistent reasoning in large language models (LLMs). Unlike prior resources such as NeuBAROCO and JFLD, which emphasize general or belief-aligned logic, BIS Reasoning 1.0 systematically introduces logically valid yet belief-inconsistent syllogisms to expose belief bias—the tendency to accept believable conclusions irrespective of validity. We benchmark a representative suite of cutting-edge models—including OpenAI GPT-5 variants, GPT-4o, Qwen, and prominent Japanese LLMs—under a uniform, zero-shot protocol. Reasoning-centric models achieve near-perfect accuracy on BIS Reasoning 1.0 (e.g., Qwen3-32B ≈99% and GPT-5-mini up to ≈99.7%), while GPT-4o attains around 80%. Earlier Japanese-specialized models underperform, often well below 60%, whereas the latest llm-jp-3.1-13b-instruct4 markedly improves to the mid-80% range. These results indicate that robustness to belief-inconsistent inputs is driven more by explicit reasoning optimization than by language specialization or scale alone. Our analysis further shows that even top-tier systems falter when logical validity conflicts with intuitive or factual beliefs, and that performance is sensitive to prompt design and inference-time reasoning effort. We discuss implications for safety-critical domains—law, healthcare, and scientific literature—where strict logical fidelity must override intuitive belief to ensure reliability.
Large Language Models (LLMs) frequently produce fluent but unverifiable reasoning, resulting in potential hallucinations and faulty inferences. This study proposes a logic programming - based verification framework ValidLogic4LLM in which the reasoning expressed by an LLM is transformed into a logic program (LP), probabilistic LP, defeasible LP and abductive LP representing world knowledge and a given problem description—such as a patient health complaint. The LP formed by an LLM is executed within a symbolic reasoning engine, and the resulting inferences are compared to the LLM’s natural-language conclusions. The strength or probability of facts, clauses and arguments is computed based on discourse structure of text expressing these facts or arguments. Divergence between symbolic and neural reasoning outcomes indicates possible hallucination or inconsistency in the model’s internal logic.
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation
Lina Conti | Dennis Fucci | Marco Gaido | Matteo Negri | Guillaume Wisniewski | Luisa Bentivogli
Lina Conti | Dennis Fucci | Marco Gaido | Matteo Negri | Guillaume Wisniewski | Luisa Bentivogli
Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bias concerns. For example, in speech translation (ST), when translating from languages with notional gender, such as English, into languages where gender-ambiguous terms referring to the speaker are assigned grammatical gender, the speaker’s vocal characteristics may play a role in gender assignment. This risks misgendering speakers—whether through masculine defaults or vocal-based assumptions—yet how ST models make these decisions remains poorly understood. We investigate the mechanisms ST models use to assign gender to speaker-referring terms across three language pairs (en→es/fr/it). To do so, we examine how training data patterns, internal language model (ILM) biases, and acoustic information interact. We find that models do not simply replicate term-specific gender associations from training data, but learn broader patterns of masculine prevalence. While the ILM exhibits strong masculine bias, models can override these preferences based on acoustic input. Using contrastive feature attribution on spectrograms, we reveal that the model with higher gender accuracy relies on a previously unknown mechanism: using first-person pronouns to link gendered terms back to the speaker, accessing gender information distributed across the frequency spectrum rather than concentrated in pitch.
MUCH: A Multilingual Claim Hallucination Benchmark
Jérémie Dentan | Alexi Stanislas Canesse | Davide Buscaldi | Aymen Shabou | Sonia Vanier
Jérémie Dentan | Alexi Stanislas Canesse | Davide Buscaldi | Aymen Shabou | Sonia Vanier
Claim-level Uncertainty Quantification (UQ) is a promising approach to mitigate the lack of reliability in Large Language Models (LLMs). We introduce MUCH, the first claim-level UQ benchmark designed for fair and reproducible evaluation of future methods under realistic conditions. It includes 4,876 samples across four European languages (English, French, Spanish, and German) and four instruction-tuned open-weight LLMs. Unlike prior claim-level benchmarks, we release 24 generation logits per token, facilitating the development of future white-box methods without re-generating data. Moreover, in contrast to previous benchmarks that rely on manual or LLM-based segmentation, we propose a new deterministic algorithm capable of segmenting claims using as little as 0.1% of the LLM generation time. This makes our segmentation approach suitable for real-time monitoring of LLM outputs, ensuring that MUCH evaluates UQ methods under realistic deployment constraints. Finally, our evaluations show that current methods still have substantial room for improvement in both performance and efficiency.
AgriChain: Visually-Grounded Expert-Verified Reasoning for Interpretable Agricultural Vision–Language Models
Hazza Mahmood | Yongqiang Yu | Rao Anwer
Hazza Mahmood | Yongqiang Yu | Rao Anwer
Accurate and interpretable plant disease diagnosis remains a key challenge for vision–language models in real agricultural settings. We present AgriChain, a new dataset of around 11,000 expert-curated leaf images covering a wide range of crops and diseases. Each image is paired with a disease label, a calibrated confidence score, and an expert-verified chain-of-thought explanation. Draft rationales were first generated by GPT-4o and then refined by a professional agricultural engineer using standard descriptors such as lesion color, margin, and distribution. Using these data, we fine-tune the open vision–language model Qwen-2.5-VL-3B to jointly identify diseases and explain its reasoning in a way that mirrors expert thinking. On a 1,000-image test set, our model reaches 73.1% accuracy and produces explanations that align closely with human expertise. These results show that expert-verified reasoning supervision enhances both performance and interpretability, bringing us closer to transparent and trustworthy AI tools for sustainable agriculture.To support reproducibility and further research, the dataset and code are publicly available at https://github.com/hazzanabeel12-netizen/agrichain.
SyntaxGym for French: Resource, Annotation, and Evaluation of French and Multilingual LLMs
Tatiana Bladier | Henri-José Deulofeu | Alexis Nasr
Tatiana Bladier | Henri-José Deulofeu | Alexis Nasr
Despite recent advances in large language models (LLMs), their syntactic competence remains insufficiently characterized, especially for languages other than English. While benchmarks such as BLiMP and SyntaxGym have enabled systematic syntactic evaluation in English and Spanish, no comparable resource exists for French. To address this gap, we present SyntaxGymFR, a manually curated evaluation suite for evaluating the syntactic abilities of French and multilingual LLMs. SyntaxGymFR consists of manually validated minimal sentence pairs targeting key syntactic phenomena in French. We describe the annotation methodology, the selection of linguistic constructions, and the validation procedures used to ensure the coverage of syntactic phenomena. Furthermore, we report experimental results obtained with several French and multilingual LLMs, analyzing their sensitivity to grammatical contrasts and cross-linguistic transfer effects. Our results provide new insights into the syntactic generalization capabilities of French LLMs and establish SyntaxGymFR as a benchmark for future research on language-specific evaluation of syntactic competence.
Modeling the Human Lexicon under Temperature Variations: Linguistic Factors, Diversity and Typicality in LLM Word Associations
Maria A. Rodriguez | Marie Candito | Richard Huyghe
Maria A. Rodriguez | Marie Candito | Richard Huyghe
Large language models (LLMs) achieve impressive results in terms of fluency in text generation, yet the nature of their linguistic knowledge – in particular the human-likeness of their internal lexicon – remains uncertain. This study compares human and LLM-generated word associations to evaluate how accurately models capture human lexical patterns. Using English cue-response pairs from the SWOW-EN dataset and newly generated associations from three LLMs (Mistral-7B, Llama-3.1-8B, and Qwen-2.5-32B) across multiple temperature settings, we examine (i) the influence of lexical factors such as word frequency and concreteness on cue-response pairs, and (ii) the variability and typicality of LLM responses relative to humans. Results show that all models mirror human trends for frequency and concreteness but differ in response variability and typicality. Larger models such as Qwen tend to emulate a single “prototypical” human participant, generating highly typical but minimally variable responses, while smaller models such as Mistral and Llama produce more variable yet less typical responses. Temperature settings further influence this trade-off, with higher values increasing variability but decreasing typicality. These findings highlight both the similarities and differences between human and LLM lexicons, emphasizing the need to account for model size and temperature when probing LLM lexical representations.
Object Realisation in Spoken Guadeloupan French: Evaluating NLP Models for an Under-Resourced Variety
Amalia Canes Nápoles | Sophie Repp
Amalia Canes Nápoles | Sophie Repp
This paper contributes to the evaluation of natural language parsing models applied to colloquial speech in lesser studied varieties of a language. We are reporting on the performance of speech recognition and of universal dependency (UD) parsing models in a radio corpus of colloquial French spoken in Guadaloupe (GuaFr), which is in contact with a typologically distant language, French-based Guadaloupean Creole (GuaCr). The corpus poses specific challenges due to phonetic and syntactic specifics of GuaFr, as well as the occurrence of code switching to GuaCr. We show weakening the ASR decoder’s language-model (LM) in various parameters avoids hallucination of null objects, which have been described as typical for spoken GuaFr, but not of non-standard object clitic positioning. For UD parsing, we investigate utterance segmentation as the primary lever to affect model performance and compare different segmentation sources (ASR punctuation, manual chunking, UD parser tokenization) and their combination. We highlight both strengths and pitfalls of the models, again focussing on the expression of syntactic objects.
Reason2Decide: Rationale-Driven Multi-Task Learning
H M Quamran Hasan | Housam Khalifa Bashier | Jiayi Dai | Mi-Young Kim | Randy Goebel
H M Quamran Hasan | Housam Khalifa Bashier | Jiayi Dai | Mi-Young Kim | Randy Goebel
Despite the wide adoption of Large Language Models (LLM)s, clinical decision support systems face a critical challenge: achieving high predictive accuracy while generating explanations aligned with those predictions. Current approaches suffer from exposure bias, leading to misaligned explanations. We propose Reason2Decide, a two-stage training framework that addresses key challenges in self-rationalization, including exposure bias and task separation. In Stage-1, our model is trained on rationale generation, while in Stage-2, we jointly train on label prediction and rationale generation, applying scheduled sampling to gradually transition from conditioning on gold labels to model predictions. We evaluate Reason2Decide on three medical datasets, including a proprietary triage dataset and public biomedical QA datasets. Across model sizes, Reason2Decide outperforms other fine-tuned baselines and some zero-shot LLMs in prediction (F1) and rationale fidelity (BERTScore, BLEU, LLM-as-a-Judge). In triage, Reason2Decide is rationale source-robust across LLM-generated, nurse-authored, and nurse-post-processed rationales. In our experiments, while using only LLM-generated rationales in Stage-1, Reason2Decide outperforms other fine-tuned variants. This indicates that LLM-generated rationales are suitable for pretraining models, reducing reliance on human annotations. Remarkably, Reason2Decide achieves these gains with models 40x smaller than contemporary foundation models, making clinical reasoning more accessible for resource-constrained deployments while still providing explainable decision support.
Ragability Benchmark: A Dataset and Library to Test LLMs on Inter-context Conflicts
Stephanie Gross | Johann Petrak | Brigitte Krenn
Stephanie Gross | Johann Petrak | Brigitte Krenn
Knowledge conflicts are a challenging issue when applying retrieval augmented generation (RAG) systems. In this paper, we propose a benchmark to test LLMs on how they deal with inter-context knowledge conflicts where implicit reasoning is required to solve the conflict. Based on actual empirical examples, real entities are replaced by fantasy entities to make sure the model’s internal knowledge does not influence how the model deals with external conflicting information. The proposed benchmark can be used to assess current up-to-date LLMs, but it can also flexibly be adapted for in-depth evaluation of a specific RAG system on selected aspects of conflict identification. We also present an experiment where we apply the benchmark to test 7 current LLMs from different model families. The results show that LLMs are able to identify conflicting contexts (’Is there a contradiction, yes or no?’), while they struggle with answering content related queries. Adding a hint that there might be a contradiction in the provided contexts increases the performance of conflict identification for contradictory context, while it significantly decreases the performance for non-contradictory contexts.
Evaluating the Adaptability of Large Language Models to Linguistic Variation
Ziyan Xu | Marina Seghier | Alice Millour | Carlos-Emiliano Gonzalez-Gallardo | Jean-Yves Antoine
Ziyan Xu | Marina Seghier | Alice Millour | Carlos-Emiliano Gonzalez-Gallardo | Jean-Yves Antoine
Large language models (LLMs) are often assumed to generalize easily across linguistic contexts, yet their ability to adapt to genre variation remains underexplored. This study examines that question through a French Named Entity Recognition (NER) task conducted on NEM.fr, a multi-genre corpus annotated with gold named entities (NEs) spanning 11 text types, from juridical and encyclopedic prose to poetry, political speech, and online discourse. We evaluate the reasoning-oriented model DeepSeek R1 across six prompting configurations (zero-, one-, and few-shot, with and without chain-of-thought reasoning), while keeping the annotation scheme, prompting format, and evaluation pipeline constant to isolate the role of genre. Performance is measured using both strict and fuzzy F1-based metrics. The results show that prompting choices have little effect once the model has learned the task format, but that genre differences strongly influence outcomes: fuzzy F1 scores range from about 0.85 in formal genres to below 0.20 in informal ones. Even under tightly controlled conditions, LLM behaviour proves highly sensitive to textual regularity and stylistic variation, highlighting genre as a key factor in assessing model robustness.
Probing Discrete Speech Tokens of Spoken Language Models
Sven Naber | Julia Koch | Pranav Singh | Alberto Saponaro | Ioanna Karagianni | Ngoc Thang Vu
Sven Naber | Julia Koch | Pranav Singh | Alberto Saponaro | Ioanna Karagianni | Ngoc Thang Vu
This paper presents a framework for systematic probing of discrete speech token representations in spoken language models (SLMs). We propose three complementary components: a distributional divergence analysis testing whether an attribute is reflected in token usage, token-based classifiers to quantify recoverability and an attribute-conditioned representation analysis revealing phonetic attribute realizations. As a demonstration we apply these probes to tokenizer outputs and model generations from CosyVoice2 and SparkTTS on LibriTTS-R and VCTK. We find that gender is encoded in their respective tokens but in different forms - the signal is more stable across stages and datasets in CosyVoice2, whereas SparkTTS shows weaker cross-stage consistency and stronger pause/prosody-related effects. Exploratory probes of valence, arousal, and dominance are weaker and less consistent. These results show that discrete speech tokens retain speaker-related information in different ways across architectures and that the proposed framework provides an interpretable basis for comparing token representations across spoken language modeling pipelines.
When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
Hasindri Sankalpana Watawana | Sergio Gastón Burdisso | Diego Aaron Moreno-Galvan | Fernando Sanchez-Vega | Adrian Pastor Lopez Monroy | Petr Motlicek | Esau Villatoro-Tello
Hasindri Sankalpana Watawana | Sergio Gastón Burdisso | Diego Aaron Moreno-Galvan | Fernando Sanchez-Vega | Adrian Pastor Lopez Monroy | Petr Motlicek | Esau Villatoro-Tello
Automatic depression detection from doctor–patient conversations has gained momentum thanks to the availability of public corpora and advances in language modeling. However, interpretability remains limited: strong performance is often reported without revealing what drives predictions. We analyze three datasets—ANDROIDS, DAIC-WOZ, and E-DAIC—and identify a systematic bias from interviewer prompts in semi-structured interviews. Models trained on interviewer turns exploit fixed prompts and positions to distinguish depressed from control subjects, often achieving high classification scores without using participant language. Restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. While semi-structured protocols ensure consistency, including interviewer prompts inflates performance by leveraging script artifacts. Our results highlight a cross-dataset, architecture-agnostic bias and emphasize the need for analyses that localize decision evidence by time and speaker to ensure models learn from participants’ language.
Constructing a Japanese Claim Decomposition Dataset for Fact-Checking of LLM-Generated Texts
Miwa Masano | Ribeka Keyaki | Atsushi Keyaki | Rei Minamoto | Kaito Horio | Hirokazu Kiyomaru | Kouta Nakayama | Hideyuki Tachibana | Daisuke Kawahara
Miwa Masano | Ribeka Keyaki | Atsushi Keyaki | Rei Minamoto | Kaito Horio | Hirokazu Kiyomaru | Kouta Nakayama | Hideyuki Tachibana | Daisuke Kawahara
Since texts generated by large language models (LLMs) may contain misinformation (hallucinations), develop- ing fact-checking systems capable of assessing their veracity has become increasingly important. One of the mainstream approaches to fact-checking is the claim-based one, which first decomposes a generated text into claims, i.e., independent and atomic units of information. Each claim is then used as a query to retrieve supporting evidence, and a verdict is predicted for each claim-evidence pair. Conducting fact-checking at the claim level enhances the explainability of verification results. However, achieving highly accurate verification requires that the text be decomposed into claims at an appropriate level of granularity. To address this, we constructed a dataset for Japanese claim decomposition. As part of this dataset construction, we design detailed guidelines for claim decomposition, ensuring that the extracted claims are in a form useful for fact-checking and that the decomposition rules mitigate annotator variability. Quantitative evaluation confirmed that the constructed dataset is of high quality. Additionally, experiments on prompt-based claim decomposition using the constructed dataset demonstrated that adding high-quality few-shot examples and guidelines to prompts improved performance.
Using LLMs for Automatic Discipline Annotation in a Diachronic Corpus of English Scientific Papers
Sergei Bagdasarov | Diego Alves | Stefan Fischer | Elke Teich
Sergei Bagdasarov | Diego Alves | Stefan Fischer | Elke Teich
This study investigates the potential of generative large language models (LLMs) to automatically identify the disciplines of scientific papers in the Royal Society Corpus (RSC) – an extensive collection of English scientific publications spanning more than three centuries. We evaluated eight open-source, state-of-the-art LLMs from four model families on a manually annotated subset and further validated the three best-performing models on a corpus of modern scientific texts. These models were subsequently used for large-scale annotation of the RSC. The models exhibited robust and consistent performance, with at least two LLMs agreeing on the same label for 98.3% of the documents. We then conducted an error analysis of papers assigned divergent labels and a diachronic case study of disciplinary trends within the corpus. The error analysis revealed that most discrepancies occurred in twentieth-century texts, reflecting the growing interdisciplinarity of research. The diachronic analysis showed a gradual decline in disciplinary diversity over time as well as fluctuations corresponding to major paradigm shifts such as the Chemical Revolution and key twentieth-century developments in Physics. The discipline labels generated by the three models will be made publicly available.
COCOA: Creation and Exploratory Investigation of a COrpus of Claims frOm NLP Articles
Clémentine Bleuze | Fanny Ducel | Maxime Amblard | Karen Fort
Clémentine Bleuze | Fanny Ducel | Maxime Amblard | Karen Fort
Research articles are an essential pillar of scientific knowledge, but they are subject to multiple constraints. On the one hand, their scientific reliability is essential and relies in particular on the peer review process. On the other hand, they fulfill a rhetorical function of persuasion for authors who defend claims in a more and more competitive environment. In a context of massively increasing publication growth and quickly evolving practices, it is essential that the scientific community remains alert and critical of its own biases. In this paper, we call for a “NLP for NLP” framing of theseissues. We created COCOA, a corpus of sentences from NLP papers and pre-prints published in English between 1952 and 2024, a sample of which we manually annotated with claim category labels reflecting their rhetorical function. We fine-tuned a SciBERT model to predict remaining labels, and made both the corpus and the model available to the community. We illustrate the interest of the corpus with exploratory analyses, and outline directions for further research. We hope that this work can stimulate discussions on the issues of research standardization and scientific overclaiming.
SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
Manon Berriche | Célia Nouri | Chloé Clavel | Jean-Philippe Cointet
Manon Berriche | Célia Nouri | Chloé Clavel | Jean-Philippe Cointet
We introduce SPOT (Stopping Points in Online Threads), the first annotated corpus translating the sociological concept of stopping point into a reproducible NLP task. Stopping points are ordinary critical interventions that pause or redirect online discussions through a range of forms — irony, subtle doubt or fragmentary arguments— that frameworks like counterspeech or social correction often overlook. We operationalize this concept as a binary classification task and provide reliable annotation guidelines. The corpus contains 43,305 manually annotated French Facebook comments linked to URLs flagged as false information by social media users, enriched with contextual metadata (article, post, parent comment, page or group, and source). We benchmark fine-tuned encoder models (CamemBERT) and instruction-tuned LLMs under various prompting strategies. Results show that fine-tuned encoders outperform prompted LLMs in F1 score by more than 10 percentage points, confirming the importance of supervised learning for emerging non-English social media tasks. Incorporating contextual metadata further improves encoder models F1 scores from 0.75 to 0.78. We release the anonymized dataset, along with the annotation guidelines and code in our code repository, to foster transparency and reproducible research.
MedPT: A Massive Medical Question Answering Dataset for Brazilian-Portuguese Speakers
Fernanda Bufon Farber | Iago Alves Brito | Julia Soares Dollis | Pedro Schindler Freire Brasil Ribeiro | Rafael Teixeira Sousa | Arlindo R. Galvão Filho
Fernanda Bufon Farber | Iago Alves Brito | Julia Soares Dollis | Pedro Schindler Freire Brasil Ribeiro | Rafael Teixeira Sousa | Arlindo R. Galvão Filho
While large language models (LLMs) show transformative potential in healthcare, their development remains focused on high-resource languages. This creates a critical barrier for other languages, as simple translation fails to capture unique clinical and cultural nuances, such as endemic diseases. To address this, we introduce MedPT, the first large-scale, real-world corpus of patient-doctor interactions for the Brazilian Portuguese medical domain. Comprising 384,095 authentic question-answer pairs and covering over 3,200 distinct health-related conditions, the dataset was refined through a rigorous multi-stage curation protocol that employed a hybrid quantitative-qualitative analysis to filter noise and contextually enrich thousands of ambiguous queries, resulting in a corpus of approximately 57 million tokens. We further utilize of LLM-driven annotation to classify queries into seven semantic types to capture user intent. To validate MedPT’s utility, we benchmark it in a medical specialty classification task: fine-tuning a 1.7B parameter model achieves an outstanding 94% F1-score on a 20-class setup. Furthermore, our qualitative error analysis shows misclassifications are not random but reflect genuine clinical ambiguities (e.g., between comorbid conditions), proving the dataset’s deep semantic richness. We publicly release MedPT on Hugging Face to support the development of more equitable, accurate, and culturally-aware medical technologies for the Portuguese-speaking world.
Large Language Models for Citation Function Classification
Daniel Vodička | Pavel Kral | Christophe Cerisara | Jakub Šmíd
Daniel Vodička | Pavel Kral | Christophe Cerisara | Jakub Šmíd
Citation function classification plays a crucial role in understanding the relationships between scientific publications and advancing bibliometric analysis. This study presents one of the first comprehensive evaluations of multiple state-of-the-art (SOTA) large language models (LLMs) for citation function classification, achieving new SOTA results on the ACL-ARC dataset. We systematically compare five models (Mistral 7B, Orca 2-7B, LLaMA 3.1-8B, Falcon 7B, and SciBERT) across zero-shot, few-shot, and fine-tuning approaches. Our fine-tuned Falcon 7B model achieves a 73,3% macro F1 score on ACL-ARC, representing a significant improvement over previous methods. Additionally, we introduce AC3, a novel dataset featuring a seven-category annotation scheme that distinguishes between neutral acknowledgments and explicit evaluative stances (more opinion-oriented citations – criticizing, complimenting, contradicting). The dataset is implemented across four context extraction variants to systematically evaluate the impact of contextual scope on classification performance. We also provide detailed analysis of model performance, experimental configurations, and limitations to guide future research in this domain. To our knowledge, this is one of the first studies dedicated to comprehensive model comparison for citation function classification, addressing a gap identified in recent surveys.
Small LLMs for Medical NLP: A Systematic Analysis of Few-Shot, Constraint Decoding, Fine-Tuning and Continual Pre-Training in Italian
Pietro Ferrazzi | Mattia Franzin | Alberto Lavelli | Bernardo Magnini
Pietro Ferrazzi | Mattia Franzin | Alberto Lavelli | Bernardo Magnini
Large Language Models (LLMs) consistently excel in diverse medical Natural Language Processing (NLP) tasks, yet their substantial computational requirements often limit deployment in real-world healthcare settings. In this work, we investigate whether “small” LLMs (around one billion parameters) can effectively perform medical tasks while maintaining competitive accuracy. We evaluate models from three major families—Llama-3, Gemma-3, and Qwen3—across 20 clinical NLP tasks among Named Entity Recognition, Relation Extraction, Case Report Form Filling, Question Answering, and Argument Mining. We systematically compare a range of adaptation strategies, both at inference time (few-shot prompting, constraint decoding) and at training time (supervised fine-tuning, continual pretraining). Fine-tuning emerges as the most effective approach, while the combination of few-shot prompting and constraint decoding offers strong lower-resource alternatives. Our results show that small LLMs can match or even surpass larger baselines, with our best configuration based on Qwen3-1.7B achieving an average score +9.2 points higher than Qwen3-32B. We release a comprehensive collection of all the publicly available Italian medical datasets for NLP tasks, together with our top-performing models. Furthermore, we release an Italian dataset of 126M words from the Emergency Department of an Italian Hospital, and 175M words from various sources that we used for continual pre-training.
Analysing Lightweight Large Language Models for Biomedical Named Entity Recognition on Diverse Ouput Formats
Pierre Epron | Adrien Coulet | Mehwish Alam
Pierre Epron | Adrien Coulet | Mehwish Alam
Despite their strong linguistic capabilities, Large Language Models (LLMs) are computationally demanding and require substantial resources for fine-tuning, which is unadapted to privacy and budget constraints of many healthcare settings. To address this, we present an experimental analysis focused on Biomedical Named Entity Recognition using lightweight LLMs, we evaluate the impact of different output formats on model performance. The results reveal that lightweight LLMs can achieve competitive performance compared to the larger models, highlighting their potential as lightweight yet effective alternatives for biomedical information extraction. Our analysis shows that instruction tuning over many distinct formats does not improve performance, but identifies several format consistently associated with better performance.
WISTERIA: Weak Implicit Signal-based Temporal Relation Extraction with Attention
Duy Dao DO | Anaïs Halftermeyer | Thi Bich Hanh DAO
Duy Dao DO | Anaïs Halftermeyer | Thi Bich Hanh DAO
Temporal Relation Extraction (TRE) requires identifying how two events or temporal expressions are related in time. Existing attention-based models often highlight globally salient tokens but overlook the pair-specific cues that actually determine the temporal relation. We propose WISTERIA (Weak Implicit Signal-based Temporal Relation Extraction with Attention), a framework that examines whether the top-K attention components conditioned on each event pair truly encode interpretable evidence for temporal classification. Unlike prior works assuming explicit markers such as before, after, or when, WISTERIA considers signals as any lexical, syntactic, or morphological element implicitly expressing temporal order. By combining multi-head attention with pair-conditioned top-K pooling, the model isolates the most informative contextual tokens for each pair. We conduct extensive experiments on TimeBank-Dense, MATRES, TDDMan, and TDDAuto, including linguistic analyses of top-K tokens. Results show that WISTERIA achieves competitive accuracy and reveals pair-level rationales aligned with temporal linguistic cues, offering a localized and interpretable view of temporal reasoning.
Dynamic Model Switching to Mitigate Outdated Knowledge in Large Language Models
Ramakrishna Pinninti | Sabyasachi Kamila | Ayan Mazumder | Mohammed Hasanuzzaman
Ramakrishna Pinninti | Sabyasachi Kamila | Ayan Mazumder | Mohammed Hasanuzzaman
Generating timely and accurate content is a significant challenge for Large Language Models (LLMs). Obsolete information reduces their reliability and user trust. To overcome the limitations of single models in adapting to evolving information, we propose a dynamic switching model. A multitask trained switch model objective, adaptively picks between a large model that does not have recent information and a smaller model fine-tuned on recent information using contextual and temporal indicators. This method incorporates semantic update detection and temporal switching, which predicts text obsolescence through aggregation of reward signals. For evaluation, we curated the Temporally-aware Dynamic Dataset (TaDD) on Wikipedia and Guardian articles, which are frequently updated. Our framework achieves a balanced precision-recall trade-off on five datasets without continuous retraining, which shows that the model is efficient and adaptable compared to static pretrained models.
Multi-Scale Model Compression via Nested Matrix Learning
Xiangjue Dong | Aditya Anantharaman | Hemant Pugaliya | Kai Zhong
Xiangjue Dong | Aditya Anantharaman | Hemant Pugaliya | Kai Zhong
Large language models (LLMs) have been widely deployed and have achieved remarkable success in downstream tasks. However, their high latency continues to pose challenges for real-time applications that require fast inference, and the need to train and deploy distinct models for different hardware constraints increases both financial and computational costs. To address this, we propose Nested Matrix Learning (NML), a method that trains a single, flexible model capable of generating multiple high-performing student models of varying sizes. This is achieved by simultaneously optimizing a pre-trained teacher model and its nested sub-models in a single training process, without sacrificing the teacher’s performance. NML provides a flexible and scalable solution, allowing models to adapt to different computational budgets. Our extensive experiments show that student models produced by NML, which can be up to 10x smaller than the full-size model, can be directly deployed for efficient inference or serve as superior initialization points for further fine-tuning in downstream tasks. By preserving the performance of the teacher model while delivering compact and efficient student models of various sizes, NML enhances the usability and adaptability of LLMs in real-world scenarios.
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba | Aman Sinha | Timothee Mickus | Raul Vazquez | Patanjali Bhamidipati | Claudio Savelli | Ahana Chattopadhyay | Laura A. Zanella | Yash Kankanampati | Binesh Arakkal Remesh | Aryan Ashok Chandramania | Rohit Agarwal | Chuyuan Li | Ioana Buhnila | Radhika Mamidi
Federica Gamba | Aman Sinha | Timothee Mickus | Raul Vazquez | Patanjali Bhamidipati | Claudio Savelli | Ahana Chattopadhyay | Laura A. Zanella | Yash Kankanampati | Binesh Arakkal Remesh | Aryan Ashok Chandramania | Rohit Agarwal | Chuyuan Li | Ioana Buhnila | Radhika Mamidi
We introduce the CAP (Confabulations from ACL Publications) dataset, a multilingual resource for studying hallucinations in large language models (LLMs) within scientific text generation. CAP focuses on the scientific domain, where hallucinations can distort factual knowledge, as they frequently do. In this domain, however, the presence of specialized terminology, statistical reasoning, and context-dependent interpretations further exacerbates these distortions, particularly given LLMs’ lack of true comprehension, limited contextual understanding, and bias toward surface-level generalization. CAP operates in a cross-lingual setting covering five high-resource languages (English, French, Hindi, Italian, and Spanish) and four low-resource languages (Bengali, Gujarati, Malayalam, and Telugu). The dataset comprises 900 curated scientific questions and over 7,000 LLM-generated answers from 16 publicly available models, provided as question–answer pairs along with token sequences and corresponding logits. Each instance is annotated with a binary label indicating the presence of a scientific hallucination, denoted as a factuality error, and a fluency label, capturing issues in the linguistic quality or naturalness of the text. CAP is publicly released to facilitate advanced research on hallucination detection, multilingual evaluation of LLMs, and the development of more reliable scientific NLP systems.
MedInjection-FR: Exploring the Role of Native, Synthetic, and Translated Data in Biomedical Instruction Tuning
Ikram Belmadani | Oumaima El Khettari | Pacome Constant Dit Beaufils | Benoit Favre | Richard Dufour
Ikram Belmadani | Oumaima El Khettari | Pacome Constant Dit Beaufils | Benoit Favre | Richard Dufour
Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts. Yet, in specialized fields such as medicine, the scarcity of high-quality French instruction data limits effective supervision. To address this gap, we introduce MedInjection-FR, a large-scale French biomedical instruction dataset comprising 571K instruction–response pairs drawn from three complementary sources: native, synthetic, and translated data. We design a controlled experimental framework to systematically assess how data provenance affects instruction tuning, using Qwen-4B-Instruct fine-tuned across seven configurations combining these sources. Results show that native data yield the strongest performance, while mixed setups, particularly native and translated, provide complementary benefits. Synthetic data alone remains less effective but contributes positively when balanced with native supervision. Evaluation on open-ended QA combines automatic metrics, LLM-as-a-judge assessment, and human expert review; although LLM-based judgments correlate best with human ratings, they show sensitivity to verbosity. These findings highlight that data authenticity and diversity jointly shape downstream adaptation and that heterogeneous supervision can mitigate the scarcity of native French medical instructions.
The Impact of Tokenization Algorithms on Hungarian Language Model Performance
Mátyás Osváth | Máté Norbert Molnár | Roland Gunics | Noémi Ligeti-Nagy
Mátyás Osváth | Máté Norbert Molnár | Roland Gunics | Noémi Ligeti-Nagy
Tokenization is a crucial text processing step for preparing input for language models and can contribute to model performance, especially in morphologically rich languages. Currently, Byte Pair Encoding (BPE), WordPiece, and Unigram LM algorithms are predominantly used in language models, but their effects can vary in agglutinative languages. This work compares these tokenization algorithms across varying vocabulary sizes, as well as a modified Unigram LM variant with morphologically informed initialization, on the Hungarian subset of the OSCAR dataset. The evaluation is based on several metrics describing the inferred quality of the tokenizers and on the downstream performance of multiple BERT models on the HuLU benchmark. Results show that BPE produces the most compact and morphologically aligned subword representations, while the modified Unigram LM achieved the best overall downstream performance across tasks. However, differences between methods and vocabulary sizes were generally small and not statistically significant, with the exception of HuCoPA (a task within the HuLU benchmark), which showed sensitivity to both factors. These findings underscore that tokenizer choice and vocabulary design are critical determinants of language model efficiency and performance in morphologically rich languages.
FAME: Fictional Actors for Multilingual Erasure
Claudio Savelli | Moreno La Quatra | Alkis Koudounas | Flavio Giobergia
Claudio Savelli | Moreno La Quatra | Alkis Koudounas | Flavio Giobergia
Large Language Models trained on web-scale data raise concerns about privacy and the right to be forgotten. To address these issues, Machine Unlearning provides techniques to remove specific information from trained models without retraining from scratch. However, existing benchmarks for evaluating unlearning in LLMs face two major limitations: they focus only on English and support only entity-level forgetting (removing all information about a person). We introduce FAME (Fictional Actors for Multilingual Erasure), a synthetic benchmark for evaluating Machine Unlearning across five languages: English, French, German, Italian, and Spanish. FAME contains 1,000 fictional actor biographies and 20,000 question-answer pairs. Each biography includes information on 20 topics organized into structured categories (biography, career, achievements, personal information). This design enables both entity-level unlearning (i.e., forgetting entire identities) and instance-level unlearning (i.e., forgetting specific facts while retaining others). We provide two dataset splits to support these two different unlearning scenarios and enable systematic comparison of unlearning techniques across languages. Since FAME uses entirely fictional data, it ensures that the information was never encountered during model pretraining, allowing for a controlled evaluation of unlearning methods.
Detecting Risky Behavior Related to Alcohol and Drug Use within Adolescents’ Private Messenger Conversations
Jaromír Plhák | Michaela Lebedíková | Ondrej Sotolar | David Smahel
Jaromír Plhák | Michaela Lebedíková | Ondrej Sotolar | David Smahel
Alcohol and drug use negatively impact adolescents’ health, making early detection and prevention essential. One promising approach involves analyzing adolescents’ online conversations for signs of substance use. However, current machine learning models for online detection often rely on public data sources that fail to capture the private experiences of adolescents. In this study, we developed a BERT-based machine learning model to automatically identify discussions about alcohol and drug use with high accuracy, leveraging private messenger conversations from adolescents. Our novel dataset comprises 272,465 annotated utterances from a corpus of 1,260,492 utterances in 2,807 chats authored by 2,165 individuals, primarily in Czech. Our best BERT-based machine learning model achieved a solid F1 score of 0.817, demonstrating the feasibility of addressing this social science task even in low-resource languages like Czech. Additionally, we verified that state-of-the-art generative open-source large language models are equally effective for this task and can be successfully adapted for other languages, including English. We also analyzed misclassified utterances to identify problematic patterns and improve model performance. The resulting models have significant practical implications for parental mediation software and parental control applications. By automating substance use detection and enabling appropriate real-time interventions, these tools can contribute to safeguarding adolescents’ health.
Voices and Echoes in Fictional Dialogue: A Study of Linguistic Coordination in Literary Texts
Ioana-Roxana Boriceanu | Alina Iacob | Liviu P. Dinu
Ioana-Roxana Boriceanu | Alina Iacob | Liviu P. Dinu
This study investigates linguistic coordination in fictional dialogue, examining whether the phenomenon typically observed in natural conversation also appears in imagined exchanges created by authors. We analyse dialogues from ten English novels by Jane Austen and E. M. Forster using the Project Dialogism Novel Corpus (PDNC) to measure linguistic convergence across nine function word categories from the Linguistic Inquiry and Word Count (LIWC) lexicon, complemented by network based measures that capture how linguistic adaptation shapes interactions among characters. The results provide evidence of convergence in both authors, confirming that linguistic coordination extends to literary dialogue. The network analysis supports these findings, revealing that alignment is generally reciprocal, unevenly distributed but widespread, and often crosses social and narrative boundaries. Taken together, these results suggest that linguistic coordination in fiction does not depend on deliberate stylistic planning, but reflects underlying cognitive mechanisms involved in language processing and social interaction.
Bridging the Domain Divide: Supervised vs. Zero-Shot Clinical Section Segmentation from MIMIC-III to Obstetrics
Baris Karacan | Barbara Di Eugenio | Patrick Thornton
Baris Karacan | Barbara Di Eugenio | Patrick Thornton
Clinical free-text notes contain vital patient information. They are structured into labelled sections; recognizing these sections has been shown to support clinical decision-making and downstream NLP tasks. In this paper, we advance clinical section segmentation through three key contributions. First, we curate a new de-identified, section-labeled obstetrics notes dataset, to supplement the medical domains covered in public corpora such as MIMIC-III, on which most existing segmentation approaches are trained. Second, we systematically evaluate transformer-based supervised models for section segmentation on a curated subset of MIMIC-III (in-domain), and on the new obstetrics dataset (out-of-domain). Third, we conduct the first head-to-head comparison of supervised models for medical section segmentation with zero-shot large language models. Our results show that while supervised models perform strongly in-domain, their performance drops substantially out-of-domain. In contrast, zero-shot models demonstrate robust out-of-domain adaptability once hallucinated section headers are corrected. These findings underscore the importance of developing domain-specific clinical resources and highlight zero-shot segmentation as a promising direction for applying healthcare NLP beyond well-studied corpora, as long as hallucinations are appropriately managed.
Reading Dynamics and Comprehension in Cognitive Aging: A Multimodal Language Resource
Claudia Marzi | Noemi Boni | Alice Todesco | Andrea Nadalini | Giorgia Albertin | Cristina Dolciotti | Paolo Bongioanni | Marcello Ferro | Fabio Tamburini | Gloria Gagliardi | Vito Pirrelli
Claudia Marzi | Noemi Boni | Alice Todesco | Andrea Nadalini | Giorgia Albertin | Cristina Dolciotti | Paolo Bongioanni | Marcello Ferro | Fabio Tamburini | Gloria Gagliardi | Vito Pirrelli
We introduce a novel italian language resource for the study of reading and comprehension in aging populations, combining behavioural and linguistic data from healthy controls (HC), individuals with subjective cognitive decline (SCI), participants with Mild Cognitive Impairment (MCI), and patients with mild dementia (CDR1). Reading performance was recorded through a finger-tracking based application during both silent and oral reading, enabling fine-grained temporal analyses at the text, token and character level. Comprehension was assessed via multiple question types (wh-, inferential, referential, and lexical). Descriptive and non-linear regression analyses informed a feature selection process, yielding temporal and comprehension-based measures that capture individual reading dynamics. These features were explored through unsupervised clustering and supervised classification to investigate their discriminative and predictive potential across cognitive profiles. The resource supports research on reading and cognitive decline, offers a reproducible protocol for large-scale data collection, and provides a foundation for developing early cognitive screening and monitoring tools or aging populations.
Evaluating Style Embeddings for Machine-Generated Text Detection
Noé Durandard | Saurabh Dhawan | Thierry Poibeau
Noé Durandard | Saurabh Dhawan | Thierry Poibeau
In this paper, we evaluate the use of style embeddings for distinguishing machine-generated from human-written text. Style embeddings are particularly suited for this task as compared to semantic embeddings, they offer higher content-independence, and compared to feature-engineering approaches, they offer a richer and more holistic representation of writing style. We use a detection module in which texts are first embedded in high-dimensional stylistic spaces using a style encoder, and the resulting vector representations are classified using supervised methods. To optimize this detector, we evaluate the performance of a range of pre-trained public-domain style encoders paired with different supervised methods. When evaluated on MGTBench, a widely adopted benchmark, our approach matches or exceeds state-of-the-art performance metrics. It also generalizes well across various text domains and LLMs. Our findings highlight the potential, and would facilitate the use, of style embeddings as lightweight and effective components of machine-generated text detection systems.
The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialog State Tracking Approach
Nizar El Ghazal | Antoine Caubrière | Valentin Vielzeuf
Nizar El Ghazal | Antoine Caubrière | Valentin Vielzeuf
This paper presents a comparative study of context management strategies for end-to-end Spoken Dialog State Tracking using Speech-LLMs. We systematically evaluate traditional multimodal context (combining text history and spoken current turn), full spoken history, and compressed spoken history approaches. Our experiments on the SpokenWOZ corpus demonstrate that providing the full spoken conversation as input yields the highest performance among models of similar size, significantly surpassing prior methods. Furthermore, we show that attention-pooling-based compression of the spoken history offers a strong trade-off, maintaining competitive accuracy with reduced context size. Detailed analysis confirms that improvements stem from more effective context utilization.
Off the Hamster Wheel: Rethinking Dialogue Research through a Meta-Analysis of the ACL Anthology 2024
Amandine Decker | Maxime Amblard | Ellen Breitholtz
Amandine Decker | Maxime Amblard | Ellen Breitholtz
In this paper, we take a meta-review approach to investigate how conversation is currently studied in the field by analysing papers from the ACL Anthology 2024. We retrieved 407 papers, which represents about 6.1% of the papers published in the selected venues, and manually reviewed them to determine the conversational task addressed, the corpora used, and the evaluation methods employed. Our analysis leads to several observations. First, dialogue systems represent about half of the papers of the ACL Anthology 2024 while more formal and analytical approaches cover only 12%. Second, many papers provide lacking corpus descriptions, which shows a detachment from the data which becomes a simple tool instead of one of the pillars NLP/CL applications should be based on. Third, the evaluation methods, in particular when it comes to dialogue systems, often do not assess the interactional aspects of these systems or rely on assumptions not backed up from evidence of the dialogue research community. We argue that the field would benefit from a renewed focus on analysis and formal representation of conversation, a richer evaluation culture that includes interactional quality, and more systematic practices regarding the data presentation in papers.
VDAct 2.0: Scaling Video-Grounded Dialogue for Event-driven Activity Understanding with LLM-Assisted Filtering
Wiradee Imrattanatrai | Masaki Asada | Kimihiro Hasegawa | Ken Fukuda | Teruko Mitamura
Wiradee Imrattanatrai | Masaki Asada | Kimihiro Hasegawa | Ken Fukuda | Teruko Mitamura
We present VDAct 2.0, an enhanced benchmark for video-grounded dialogue that builds upon the original VDAct by expanding dialogue coverage and introducing a scalable LLM-assisted filtering pipeline to ensure high-quality, grounded QA pairs. VDAct 2.0 comprises 6,356 human-annotated dialogues with a total of 63,958 turns, grounded in 2,975 household activity videos, with undesirable dialogue turns systematically identified and removed. To achieve this, we design a trigger-based quality framework and calibrate a panel of high-agreement LLMs through human-in-the-loop calibration, allowing scalable QA-turn-level filtering. We benchmark a wide range of pretrained and fine-tuned models, both open-source and proprietary, across standard text generation metrics and LLM-based evaluations. The results highlight both recent advances and remaining challenges in video-grounded dialogue modeling, positioning VDAct 2.0 as a high-fidelity testbed for evaluating and advancing multimodal reasoning in interactive settings.
Multi-dimensional Evaluation of Character-Authentic Dialogue Models Learned from Question-Answer Data
Atsushi Otsuka | Kazuya Matsuo | Kenta Hama | Masahiro Mizukami | Tsunehiro Arimoto | Hiroaki Sugiyama | Makoto Nakatsuji | Narichika Nomoto
Atsushi Otsuka | Kazuya Matsuo | Kenta Hama | Masahiro Mizukami | Tsunehiro Arimoto | Hiroaki Sugiyama | Makoto Nakatsuji | Narichika Nomoto
Character-authentic dialogue remains challenging for large language models (LLMs) due to limited character-specific data, generic-style collapse, and hallucinations regarding persona facts. Our work presents a comparative evaluation of several learning strategies for character dialogue grounded in question–answer (QA) data, comparing zero/few-shot prompting, supervised fine-tuning (SFT), direct preference optimization (DPO), and a hybrid approach that integrates retrieval-augmented character profiles and knowledge with policy optimization. Using both single-turn and multi-turn settings, we assess multiple dimensions central to character dialogue quality: reproducibility, diversity, hallucination, and character authenticity. Results show that SFT excels in reproducibility and hallucination reduction but tends to shorten and simplify outputs, thereby reducing diversity and authenticity. DPO improves stylistic fidelity and authenticity but depends strongly on externalized character knowledge to limit hallucinations. The hybrid variant that combines character-knowledge retrieval with DPO achieves the best overall balance, delivering strong authenticity while maintaining factual consistency and competitive reproducibility in both single- and multi-turn dialogues. We further analyze the model’s sensitivity to knowledge retrieval and response-length effects and discuss trade-offs among optimization targets that inform practical design choices for developing faithful and engaging character agents trained from scalable QA resources.
Empathy in Greek Exam-Related Support Conversations: A Comparative Evaluation of LLM Responses
Panagiota Kyriazi | Prokopis Prokopidis
Panagiota Kyriazi | Prokopis Prokopidis
Recent advancements in Large Language Models (LLMs) have significantly enhanced Natural Language Processing (NLP), particularly in generating human-like responses and engaging in social interactions. Research in natural language generation involves assessing AI-generated text across multiple dimensions, including accuracy, relevance, and robustness. This paper focuses on evaluating an LLM that puts emphasis on the Greek language and comparing it to two multilingual LLMs across four key dimensions: Understanding, Empathy, Harm, and Reasoning. We analyze the models’ responses to expressions of stress and anxiety from teenagers preparing for the Greek State’s Panhellenic exams for university entrance, assessing not only their ability to comprehend, reason, and respond empathetically but also possible unintended harm that they may cause, such as reinforcing stress or offering inappropriate advice. We, thus, introduce the GEAR (Greek Empathy Assessment Resource) dataset of student issues and exam-related forum posts along with LLM-generated empathetic responses. By prompting each model with contextual cues about its role as a recipient of these messages, this research aims to provide insights into the models’ conversational capabilities, emotional intelligence, and ethical implications in sensitive interactions.
Evaluation of Two Leading Polish Language Models in a Real-world RAG Scenario
Szymon Bartanowicz | Krzysztof Jassem
Szymon Bartanowicz | Krzysztof Jassem
This paper presents a comparative evaluation of two leading Polish instruction-tuned language models, Bielik-11B-v2.3-Instruct and PLLuM-12B-nc-chat, within a real-world Retrieval-Augmented Generation (RAG) system designed for the technical documentation of a low-code platform. The study aims to identify the optimal configuration of retrieval and generation components for Polish-language applications. The evaluation was conducted in two stages. First, several embedding models and retrieval methods were tested using standard information retrieval metrics, including NDCG. The OrlikB/KartonBERT-USE-base-v1 model combined with vector-based retrieval achieved the highest performance and was adopted for the second stage. In the generation phase, both models were evaluated using quantitative scoring and pairwise A/B testing with multiple evaluators to ensure robustness. Results show that Bielik-11B-v2.3-Instruct consistently outperformed PLLuM-12B-nc-chat in producing accurate and contextually relevant answers. The study highlights the importance of constructing a reliable golden set, employing a two-phase evaluation pipeline, and selecting appropriate metrics to ensure objective and reproducible assessment of RAG systems in real-world Polish-language contexts.
A Mental State Extraction Dataset for Theory-of-Mind-based Reasoning in Emotional Support Conversations
Seulgi Kim | Harksoo Kim
Seulgi Kim | Harksoo Kim
Emotional Support Conversations (ESC) aim to both reduce users’ emotional distress and facilitate problem-solving. Recent approaches in ESC have explored incorporating commonsense knowledge into large language models (LLMs) to improve response generation. However, existing commonsense reasoning models often rely solely on the final utterance, fail to anticipate future turns, overlook emotional cues, or treat knowledge types independently, resulting in incoherent or emotionally misaligned responses. To address these limitations, we propose an approach grounded in Theory of Mind (ToM). Specifically, we introduce MENTOS, a dataset that provides turn-level annotations of the assistant’s mental states (Belief, Emotion, and Intent), organized in a causal structure reflecting psychological principles. A commonsense reasoning model trained on MENTOS predicts these mental states as intermediate reasoning signals that guide response generation. Experiments on the ESConv and ExTES datasets show that incorporating the inferred mental states can enhance supportive and goal-directed response generation across multiple reasoning backbones and response generators. Ablation studies further confirm that Belief, Emotion, and Intent provide complementary benefits for ESC tasks. These findings highlight the effectiveness of ToM-grounded intermediate reasoning in generating empathetic and contextually appropriate responses.
Construction and Analysis of Japanese Parent-Child Dialogic Reading Corpus for Conversational Agents
Yuko Nakagi | Yuya Chiba | Sanae Fujita | Shoko Araki
Yuko Nakagi | Yuya Chiba | Sanae Fujita | Shoko Araki
Dialogic reading, which involves interactive exchanges between a parent and a child during picture book reading, has been shown to effectively promote children’s language development. While many support systems for picture book reading have been developed to reduce the burden on parents, existing systems are not yet capable of handling dialogic reading, which requires dynamic parent-child interaction. To develop conversational agents capable of dialogic reading, we constructed a multimodal corpus of parent-child picture-book reading dialogues. The corpus comprises recordings from 36 Japanese parent-child pairs taken during actual picture book reading sessions. In this study, we annotated the corpus with dialogue acts relevant to parent-child communication and categorized the types of quizzes and questions used in the sessions, analyzing the linguistic aspects of parent-child interaction during dialogic reading. After dividing the dialogues into two groups based on the proportion of the child’s utterances, our analyses revealed that dialogue systems should adapt their interaction strategies according to individual child characteristics.
ACLBot: A Knowledge Graph-Driven Assistant for ACL Anthology Research
Jan Buchmann | Steven Lynden | Kristiina Jokinen
Jan Buchmann | Steven Lynden | Kristiina Jokinen
We present ACLBot, an interactive chatbot designed to support literature exploration in the ACL Anthology by combining structured knowledge graph querying with large language model (LLM) generative AI. ACLBot integrates a Neo4j-based knowledge graph constructed by extracting data on publications, authors, topics, and research trends from the ACL Anthology, and automatically generates knowledge graph queries to retrieve relevant information in response to user questions. Retrieved results are re-injected into the LLM to produce concise, contextually grounded summaries. We describe the system’s architecture, including its query generation pipeline, knowledge graph integration, and visualization components for highlighting temporal trends in research. To assess usability and effectiveness, we conducted a user evaluation with researchers, collecting qualitative and quantitative feedback on response accuracy, informativeness, and utility for literature discovery. Results indicate that ACLBot effectively supports exploratory search, helps identify relevant works and trends, and offers a promising framework for integrating structured information with generative AI for scientific information retrieval.
This House Debates AI: Evaluating a Language Model in Oxford-Style Debates against Human Experts
Umberto Belluzzo | Kobi Hackenburg | Hannah Rose Kirk | Scott Hale | Paul Röttger
Umberto Belluzzo | Kobi Hackenburg | Hannah Rose Kirk | Scott Hale | Paul Röttger
Recent work shows that large language models (LLMs) are increasingly capable of generating persuasive arguments and messages, creating concerns over undue influence on human beliefs. Most evidence so far, however, evaluates LLM argumentation and persuasion in single-turn interactions and/or compares to weak human baselines. To address this gap, we benchmark a state-of-the-art LLM, Llama 3.1 Instruct 405B, in 100 six-turn Oxford-style debates against 20 experienced human debaters. Each anonymised debate is rated by 5 independent raters, who provide win/loss judgments as well as 0–100 scores across 11 dimensions of quality. Based on these ratings, the LLM is competitive overall, with a win rate of 51.2%, ranking 6th out of 21 debaters on mean performance score. Compared to humans, the LLM generally scores higher on presentational dimensions (e.g., clarity, confidence, formality) but equal on most substantive dimensions (convincingness, evidence, originality). We also find that pre/post rater stance tends to shift towards the position raters chose as the winning side, regardless of whether this side was the LLM or a human. Overall, our results provide new evidence on the qualities of LLM argumentation and its drivers, suggesting strong argumentative competence even in competitive multi-turn settings.
PAIR: A Pilot Dataset for Dual Perspective-based Video-Grounded Dialogue and Reconciliation
Lewis N. Watson | Carl Strathearn | Kenny Mitchell | Yanchao Yu
Lewis N. Watson | Carl Strathearn | Kenny Mitchell | Yanchao Yu
Collaborative dialogue in multi-agent settings often requires interlocutors to integrate partially overlapping perceptual information in order to construct a shared representation of a dynamic environment. We introduce PAIR, a pilot conversational corpus designed to examine how humans coordinate under systematic perceptual asymmetry. The dataset comprises 15 dialogues in which participants observed the same activity from complementary egocentric and exocentric video perspectives and engaged in open-ended discussion to produce a joint account. All transcripts were manually verified and annotated with 42 dialogue act categories, enabling fine-grained analysis of interactional structure. Beyond descriptive statistics, PAIR supports examination of measurable conversational configurations, including turn distribution, participation symmetry, and dialogue act composition, which together provide structural indicators of how perspective integration unfolds in dialogue. Although intentionally lightweight, PAIR is positioned as a controlled benchmark for analysing collaborative dialogue mechanisms rather than a large-scale training resource. The corpus supports dialogue act classification, video-grounded dialogue modelling, and investigation of multi-agent reasoning under distributed perceptual access. By coupling dual-perspective grounding with explicit interactional annotation, PAIR offers a compact testbed for studying reconciliation dynamics in task-oriented dialogue.
I Am Not Them: Persistent Outgroup Bias in Large Language Models Arising from Social Identity Persona Setting
Wenchao Dong | Assem Zhunis | Dongyoung Jeong | Hyojin Chin | Jiyoung Han | Meeyoung Cha
Wenchao Dong | Assem Zhunis | Dongyoung Jeong | Hyojin Chin | Jiyoung Han | Meeyoung Cha
This research examines how large language models internalize social identities assigned through targeted prompts. Guided by social identity theory, we investigate whether and how these identity assignments cause AI systems to differentiate between “we” (the ingroup) and “they” (the outgroup). We demonstrate that self-categorization of social identity leads to both ingroup favoritism and outgroup bias, with the latter manifesting as strongly as the former. This finding is significant given the fundamental role of outgroup bias in driving intergroup prejudice and discrimination as documented in social psychology. We further propose a strategic intervention to mitigate such bias by guiding language models to adopt the identity of the initially disfavored group. This method, validated across both political and gender domains, exposes a critical dual function of group alignment: adopting one social identity inherently alters the model’s stance toward outgroups, effectively neutralizing pre-existing biases. Our work shows that understanding human-like AI behaviors is a critical prerequisite to building more balanced and socially responsible technology.
CONVERSE: Annotation Scheme and Dataset for Multimodal Conversational Engagement Analysis in Human-Human and Human-Robot Interaction
Ekaterina Torubarova | Oskar Ljung | Julia Uddén | André Pereira
Ekaterina Torubarova | Oskar Ljung | Julia Uddén | André Pereira
Creating conversational agents that can both understand and respond appropriately to users’ engagement remains a major challenge, as conversation is one of the most universal yet complex human behaviors. Modeling conversational engagement requires a fine-grained understanding of how engagement unfolds dynamically in interaction. This paper introduces a novel turn-based annotation scheme for conversational engagement, together with the CONVERSE dataset that contains annotations of 25 hours of unscripted human–human and human–robot conversations with 48 native Swedish speakers. This dataset uniquely utilizes such an annotation scheme for both human and robot agents within the same study, allowing for direct comparison. Notably, this dataset builds upon our previous multimodal corpus, which includes brain imaging (fMRI), eye-tracking, and speech data, as well as personality and stance measures. This dataset opens a new perspective on conversational engagement through these behavioral annotations and the existing neural data at the intersection of multimodal machine learning, human-robot interaction, and cognitive neuroscience.
FineDialFact: A Benchmark for Fine-Grained Dialogue Fact Verification
Xiangyan Chen | Yufeng Li | Yujian Gan | Arkaitz Zubiaga | Matthew Purver
Xiangyan Chen | Yufeng Li | Yujian Gan | Arkaitz Zubiaga | Matthew Purver
Large language models are known to produce hallucinations - factually incorrect or fabricated information - which poses significant challenges for many natural language processing applications, such as dialogue systems. As a result, detecting hallucinations has become a critical area of research. Current approaches to hallucination detection in dialogue systems primarily focus on verifying the factual consistency of generated responses. However, these responses often contain a mix of accurate, inaccurate or non-verifiable facts, making the use of a single factual label overly simplistic and coarse-grained. In this paper, we introduce a benchmark, FineDialFact, for fine-grained dialogue fact verification, which involves verifying atomic facts extracted from dialogue responses. To support this, we construct a dataset based on publicly available dialogue datasets and evaluate it using various baseline methods. Experimental results demonstrate that methods incorporating Chain-of-Thought reasoning can enhance performance in dialogue fact verification. Despite this, the best F1-score achieved on the HybriDialogue, an open-domain dialogue dataset, is only 0.74, indicating that the benchmark remains a challenging task for future research. We release our dataset and code at https://github.com/XiangyanChen/FineDialFact.
Meta-Prompting Follow-Ups for Unsupervised Dialogue Evaluation Using Open-Source Large Language Models
Gaetano Cimino | Chuyuan Li | Giuseppe Carenini | Vincenzo Deufemia
Gaetano Cimino | Chuyuan Li | Giuseppe Carenini | Vincenzo Deufemia
Automatically evaluating dialogue quality remains a major challenge due to the complexity and contextual variability of human interactions. This paper introduces DIET, a novel unsupervised, reference-free metric that uses follow-up utterances to assess dialogue quality. Unlike existing reference-free metrics, which rely on follow-ups derived from annotated data and apply a uniform set of utterances across all dialogues, DIET generates follow-ups using open-source Large Language Models (LLMs) and refines them through a selection process. Two strategies are explored: SELFMAP, where generation and evaluation are performed by the same model to ensure internal coherence, and CRAFT, where multiple models collaborate to generate diverse and complementary follow-ups, enhancing robustness and reducing model bias. Dialogue quality is measured via the likelihood of an LLM continuing the dialogue from selected follow-ups. Experiments show DIET better correlates with human judgments than existing reference-free metrics across multiple meta-evaluation datasets.
HumaniCA: A Benchmark Resource for the Detection of Users’ Ascription of Humanness to Conversational Agents
Sabrina Villata | Amon Rapp | Luigi Di Caro | Federica Cena
Sabrina Villata | Amon Rapp | Luigi Di Caro | Federica Cena
Anthropomorphizing, which involves attributing human-like characteristics to non-human entities, is common in users’ conversations with text-based conversational agents and can lead to a misalignment between the users’ expectations and the agent’s actual capabilities. Detecting users’ ascriptions of humanness automatically may enable systems to identify when users adopt a human-like style when conversing with an agent and to adapt its responses accordingly to tune their expectations. In this paper, we introduce HumaniCA, a benchmark resource comprising three annotated datasets of user turns from real dialogues with three different types of conversational agents (task-oriented, Q&A, and LLM-based) aimed at indicating whether the user is ascribing humanness to the conversational agent. We also identified a set of linguistic indicators of user ascription of humanness to conversational agents and validated their utility with benchmark experiments. We then compared performance of our linguistic features and other well-known textual features (TF-IDF weights and SentenceBERT word embeddings), as well as their combinations. The evaluation highlights the central role of our linguistic features: whether used individually or in combination, they consistently achieve higher accuracy across all agent types.
Towards Reliable Evaluation of Emotional Text Generation in LLMs: Human vs. Automatic Metrics
Sadegh Jafari | Els Lefever | Veronique Hoste
Sadegh Jafari | Els Lefever | Veronique Hoste
Evaluating emotion generation in large language models (LLMs) remains a challenging problem due to the subjective nature of emotions and the lack of reliable automatic evaluation metrics. In this paper, we introduce a robust and extensible benchmark for systematically assessing automatic metrics in emotion generation tasks. The benchmark currently includes 13 automatic evaluation metrics and five state-of-the-art LLMs, and can be easily extended without requiring additional human annotations. Through a correlation analysis with human evaluations on a carefully curated annotated subset, we identify the emotion recognition score (ERS) metric, computed with gpt-5-nano in an oneshot setting, as the most reliable automatic evaluator, achieving a correlation exceeding 0.99. Interestingly, despite relying on the same underlying LLM, the emotion absolute score (EAS) metric shows a negative correlation, demonstrating that LLM strength alone does not guarantee automatic metric alignment with human judgment. We also provide lightweight, non-LLM-based alternatives, R2_m and R3_m, in the emotion analogy score (EAnS) metric family, suitable for low-resource settings where large models are not accessible. A comprehensive per-class emotion analysis further highlights the strengths and weaknesses of the evaluated models. Overall, our results offer a practical and scalable framework for benchmarking emotion generation evaluation metrics and pave the way for more reliable, fair, and interpretable emotional language evaluation.
Question and Response Dynamics in Public Service Encounters
Wassiliki Siskou | Ingrid Espinoza | Laurin Friedrich | Steffen Eckhard | Annette Hautli-Janisz
Wassiliki Siskou | Ingrid Espinoza | Laurin Friedrich | Steffen Eckhard | Annette Hautli-Janisz
When deciding on social welfare benefits, street-level bureaucrats wield significant discretionary power over citizens. One of the key instruments of this power lies in the questioning patterns that control the conversational agenda in face-to-face encounters. In turn, the citizens’ responses show how they navigate these conversational constraints, for instance by answering directly or through more evasive strategies. To shed light on the power dynamics inherent in these encounters, we provide over 200 verbatim transcripts of authentic conversations in German between street-level bureaucrats and citizens, as well as a fully annotated dataset of all question-response pairs extracted from these conversations. We also present PSE v2.0, which is double the size of the only previously available corpus of spoken interactions in street-level bureaucracy. Keywords: Public Service Encounters, verbatim transcripts, question-response pairs
Reasoning over Object Descriptions Improves Coreference Resolution in Task-Based Dialogue Systems
Oier Ijurco | Oier Lopez de Lacalle
Oier Ijurco | Oier Lopez de Lacalle
Task-based dialogue systems assist users in achieving specific goals, such as executing actions or retrieving information, through natural language interactions. Accurate coreference resolution is essential, as it involves identifying object references within the dialogue—a task that becomes increasingly challenging in visually grounded environments characterized by complex scenes and diverse object metadata. However, coreference resolution in task-based dialogue remains limited by poor generalization across domains and heavy reliance on supervised models that often overfit to dataset-specific artifacts. In this work, we propose a unimodal test-time reasoning approach that enables large language models (LLMs) to reason over detailed object metadata and dialogue history to improve coreference resolution. Empirical results on the SIMMC 2.1 dataset demonstrate that LLMs can generate step-by-step reasoning processes that effectively align dialogue context with objects present in the scene. Extensive experiments highlight the models’ ability to link conversations and objects accurately. Moreover, we show that test-time reasoning under few-shot settings generalizes effectively to unseen scenarios and novel objects, outperforming encoder-based supervised methods in cross-domain evaluations. These findings underscore the critical role of structured metadata and careful prompt engineering in enhancing the robustness and generalization of task-oriented dialogue systems.
Evaluating the Effect of Question Wording Variations on Answer Consistency in Large Language Models
Junya Takayama | Masaya Ohagi | Tomoya Mizumoto | Katsumasa Yoshikawa
Junya Takayama | Masaya Ohagi | Tomoya Mizumoto | Katsumasa Yoshikawa
Large Language Models (LLMs) sometimes generate inconsistent answers when asked semantically equivalent questions expressed with different wordings. Such inconsistency may lead to decreased task performance or excessive agreement with users. This study investigates how question wording influences the answer consistencies of LLMs, focusing on binary Yes/No questions. We design four types of paraphrasing patterns, namely synonym substitution, antonym substitution, addition of agreement-seeking expressions, and strengthened agreement-seeking expressions, and evaluate their impact on model outputs. Experiments with multiple open-source and commercial LLMs show that many models become more likely to answer “Yes” when agreement-seeking expressions are included, and they are particularly vulnerable to antonym substitutions. Our analysis further suggests that some of these tendencies are already present in pretrained models and are not fully removed by post-training. We also provide insights into which factors are likely (or unlikely) to contribute to improving consistency. By providing a systematic evaluation framework, this work highlights the necessity of accounting for wording-induced biases in the development and deployment of LLMs.
Knowledge-Infused Hierarchy-Aware Emotion Recognition in Code-mixed Mental Health Counseling Conversations
Aseem Srivastava | Kushagra Mittal | Anusha Tiwari | Md. Shad Akhtar
Aseem Srivastava | Kushagra Mittal | Anusha Tiwari | Md. Shad Akhtar
Effective counseling is often best achieved in a client’s preferred language, allowing better emotional resonance. Despite this, most existing research in emotion recognition in counseling focuses predominantly on English, overlooking the rich emotional and linguistic complexities of other widely spoken languages. Hinglish, a code-mixed blend of Hindi and English, is one such underexplored linguistic medium that millions use to express their emotions authentically. To address this gap, our research lays a foundational step in developing a mental-health conversation dataset in code-mixed Hinglish language, aka. IndieMH. We manually translate counseling conversations from publicly available sources into Hinglish. Moreover, we employ the dataset for emotion classification task for counseling patients. We prepare an exhaustive annotation guideline to annotate IndieMH with 13 emotional states under 3 board emotion categories. Our rigorous sanity check ensures that the quality of IndieMH adheres to research standards. Furthermore, we propose a novel knowledge-cum-hierarchy aware method named Healer for counseling emotion classification in the Hinglish language. To evaluate the model’s performance, we benchmark Healer against 11 potential baseline methods and report standard classification metrics, including accuracy, weighted-F1, and weighted-precision.
A Corpus for Personalized Dialogue Breakdown Repair in Japanese Open-Domain Conversations
Kazuya Tsubokura | Yurie Iribe | Norihide Kitaoka
Kazuya Tsubokura | Yurie Iribe | Norihide Kitaoka
Recent advances in dialogue systems have been remarkable; however, conversational breakdowns still occur, making it essential to develop appropriate repair strategies. Nevertheless, when a system breakdown actually occurs, it remains unclear how the system should perform the repair, and no corpus has been available to investigate this issue. To address this gap, we presented typical examples of system-induced dialogue breakdowns to crowd workers and collected their expected repair utterances toward the broken system. Each repair utterance was annotated with dialogue act tags, and we constructed a breakdown-repair corpus consisting of 3,990 utterances covering ten representative types of breakdowns. This corpus includes breakdown cases across diverse situations, allowing for the examination of various repair patterns. Furthermore, we also conducted a questionnaire on participants’ personal traits, creating a dataset that enables the investigation of repair strategies tailored to individual user characteristics. In this paper, we report an overview of the dataset and preliminary analysis results.
Conversational Assistants to Support Patients with Heart Failure: Comparing a Neurosymbolic Architecture with GPT
Anuja Tayal | Devika Salunke | Barbara Di Eugenio | Paula G. Allen-Meares | Eulalia P. Abril | Olga Garcia-Bedoya | Carolyn A. Dickens | Andrew D. Boyd
Anuja Tayal | Devika Salunke | Barbara Di Eugenio | Paula G. Allen-Meares | Eulalia P. Abril | Olga Garcia-Bedoya | Carolyn A. Dickens | Andrew D. Boyd
Conversational assistants are becoming increasingly popular, including in healthcare, partly due to the availability and capabilities of Large Language Models. There is a need for controlled, probing evaluations with real stakeholders, which can highlight the advantages and disadvantages of more traditional architectures and those based on generative AI. We present a within-group user study to compare two versions of a conversational assistant that allows patients with heart failure to ask about the salt content in food. One version of the system was developed with a neurosymbolic architecture, and another is based on GPT. Our objective in evaluating the two dialogue systems was not only to compare task performance but also to gain insights from real stakeholders. Results indicate that the two systems complement each other, highlighting the promise of a hybrid approach that leverages the strengths of both systems.
Disentangling Approaches to Conversation Disentanglement: Fine-Tune or Learn from Scratch?
Debaditya Pal | Anton Leuski | Ron Artstein | David Traum | Kallirroi Georgila
Debaditya Pal | Anton Leuski | Ron Artstein | David Traum | Kallirroi Georgila
Conversation disentanglement is the process of segmenting a stream of messages or utterances into separate conversations or “threads” that can be more easily understood and processed. We compare the performance of GPT-4o and GPT-4o Mini with deep learning models built from scratch for this task. We show that, using the same amount of training data, out-of-the-box GPT-4o performs poorly, and fine-tuning GPT-4o Mini results in performance comparable to learning small-size models from scratch (based on standard hand-crafted features for this task), with performance reaching 74.4% F1-score for prediction of links between messages and 45.3% F1-score for prediction of perfectly matching conversations. However, the fine-tuned GPT-4o Mini model underperforms when compared to models that utilize complex structural information. We also provide a new method for detailed analysis of the successes and failures of our models, and a new visualization method.
Evaluation of Failure Communication Strategies for Trust Repair in Human-AI Collaboration
Stina Klein | Alexandru Wurm | Elisabeth Andre | Matthias Kraus
Stina Klein | Alexandru Wurm | Elisabeth Andre | Matthias Kraus
The increasing application of Large Language Models (LLMs) in everyday tasks and at work highlights the crucial importance of trust in human-AI collaboration, particularly when an AI system fails. This paper investigates the effectiveness of failure communication strategies for trust repair in collaborative physical tasks involving a a chat-based AI assistant. A controlled experiment in which participants built LEGO cars guided by an LLM-based AI Assistant was used to evaluate whether findings from trust repair in a virtual environment, such as chatbots, translate to an environment comprising tangible tasks, and whether the timing of trust repair influences the outcome. Results indicate that actively communicating mistakes significantly improves trust compared to a no repair strategy, and that early repair tends to be more effective, indicating that failure communication, independent of the timing, is important for an appropriate calibration of trust.
Multi-Session Client-Centered Treatment Outcome Evaluation in Psychotherapy
Hongbin Na | Tao Shen | Shumao Yu | Ling Chen
Hongbin Na | Tao Shen | Shumao Yu | Ling Chen
In psychotherapy, therapeutic outcome assessment, or treatment outcome evaluation, is essential to mental health care by systematically evaluating therapeutic processes and outcomes. Existing large language model approaches often focus on therapist-centered, single-session evaluations, neglecting the client’s subjective experience and longitudinal progress across multiple sessions. To address these limitations, we propose IPAEval, a client-Informed Psychological Assessment-based Evaluation framework, which automates treatment outcome evaluations from the client’s perspective using clinical interviews. It integrates cross-session client-contextual assessment and session-focused client-dynamics assessment for a comprehensive understanding of therapeutic progress. Specifically, IPAEval employs a two-stage prompt scheme that maps client information onto psychometric test items, enabling interpretable and structured psychological assessments. Experiments on our new TheraPhase dataset, comprising 400 paired initial and completion stage client records, demonstrate that IPAEval effectively tracks symptom severity and treatment outcomes over multiple sessions, outperforming baseline approaches across both closed-source and open-source models, and validating the benefits of items-aware reasoning mechanisms.
Towards Reward Modeling for AI Tutors in Math Mistake Remediation
Kseniia Petukhova | Ekaterina Kochmar
Kseniia Petukhova | Ekaterina Kochmar
Evaluating the pedagogical quality of AI tutors remains challenging: standard NLG metrics do not determine whether responses identify mistakes, scaffold reasoning, or avoid revealing the answers. For the task of mistake remediation, we derive a hierarchy of pedagogical aspects from human pairwise preferences on MRBench, and synthesize minimally contrastive response pairs that differ along key aspects (e.g., mistake identification and location, targetedness, scaffolding, actionability, clarity, and coherence). We develop and release Bradley-Terry preference models trained on weighted-sum rankings that we automatically create from MRBench, synthetic pairs, and data combinations. Using only synthetic data, our best model reaches 0.69 pairwise accuracy on a human preference test, and combining weighted-sum data with targeted synthetic groups improves accuracy to 0.74, outperforming larger general-purpose reward models while using only a 0.5B-parameter backbone.
HOTATE: A Japanese Dialogue Corpus Annotated with Responses of Private Thoughts and Public Statements
Yuko Toda | Daisuke Maekawa | Kota Manabe | Eito Yoneyama | Kanade Nonomura | Yuki Fujiwara | Tomoyuki Kajiwara
Yuko Toda | Daisuke Maekawa | Kota Manabe | Eito Yoneyama | Kanade Nonomura | Yuki Fujiwara | Tomoyuki Kajiwara
This study aims to reveal how accurately Large Language Models (LLMs) can deal with a speaker’s actual utterances and their true feelings behind them in Japanese dialogue. Speakers use not only private thoughts which express one’s true feelings and intentions, but also public statements which convey their intentions while considering the interlocutor’s feelings and social status. While public statements help to maintain interpersonal relationships, they can obscure the speaker’s true intention, potentially leading to misunderstandings. We extended existing Japanese dialogue corpora by annotating public statements and private thoughts responses for each dialogue in the corpora, and then evaluated LLMs’ ability to classify and generate between these two types of expressions. The results of the classification task revealed that the current LLMs do not understand those expressions at all, and that training with our corpus can significantly improve the recognition performance. Furthermore, the results of the generation task demonstrated that generating private thoughts is more difficult than generating public statements, according to both automatic and human evaluations. We release our corpus, which contains 7,964 human-annotated dialogues.
Mining Naturally Romanized Seed Corpora without Romanizations
Adrian Benton | Alexander Gutkin | Christo Kirov | Brian Roark
Adrian Benton | Alexander Gutkin | Christo Kirov | Brian Roark
While the Latin script is used informally by speakers of many languages with different native scripts, high quality Latin script corpora for such languages that reflect actual natural romanizations are scarce and often difficult to collect. In this work, we propose a method for mining romanized language corpora in languages for which we do not have any pre-existing samples of naturally romanized text, focusing on Tigrinya as a test case. First we examine the efficacy of learning romanizations for a language based on observed romanizations in other languages that use the same native script. We then extrinsically assess such methods by using a romanization model trained on Amharic data to bootstrap coverage of romanized Tigrinya in a language identification system. Manual evaluation by two L1 and one L2 Tigrinya speakers suggests our method extracts romanized Tigrinya text with acceptably high precision.
This paper presents a comparative analysis of Large Language Models (LLMs) and traditional Optical Character Recognition (OCR) systems on Urdu newspapers, addressing challenges posed by complex multi-column layouts, low-resolution scans, and the stylistic variability of the Nastaliq script. To handle these challenges, we fine-tune YOLOv11x models for article- and column-level text block extraction and train a SwinIR-based super-resolution module that enhances image quality for downstream text recognition, improving accuracy by an average of 50%. We further introduce the Urdu Newspaper Benchmark (UNB), a manually annotated dataset for Urdu OCR comprising 829 paragraph images with a total of 9,982 sentences. Using UNB and the OpenITI corpus, we conduct a systematic comparison between traditional CNN+RNN-based OCR systems and modern LLMs, presenting detailed insertion, deletion, and substitution error analyses alongside character-level confusion patterns. We find that Gemini-2.5-Pro achieves the best performance on UNB (WER 0.133), while fine-tuning GPT-4o on just 500 in-domain samples yields a 6.13% absolute WER improvement, demonstrating the adaptability of LLMs to low-resource, morphologically complex scripts like Urdu. The UNB dataset and fine-tuned models are publicly available at https://github.com/paper-seven/UrduOCR.
Transformer-based models have advanced NLP, yet Hebrew still lacks a RoBERTa encoder that is trained at scale and released in both base and large variants. We present HalleluBERT, a RoBERTa-based encoder family trained from scratch on 49.1 GB of deduplicated Hebrew web text and Wikipedia using a Hebrew-specific byte-level BPE vocabulary. On native Hebrew benchmarks for named entity recognition (BMC, NEMO) and sentiment classification (SMCD), HalleluBERT outperforms monolingual and multilingual baselines, and yields the highest unweighted mean score across the three benchmarks. We release model weights and tokenizer under the MIT license to support reproducible Hebrew NLP research.
Kwanyama is related to Swahili, Zulu, and, the more than 300 other languages in the Bantu family. Yet, unlike its better-known relatives, it remains almost entirely absent from modern Natural Language Processing (NLP). We bring Kwanyama into the LLM era of NLP through two key contributions. First, we introduce OkaSentiment, the first sentiment-labeled dataset for Kwanyama. Unlike prior African sentiment corpora that rely primarily on social media, OkaSentiment is grounded in an offline, culturally relevant domain: reviews of domestic labor relationships. The dataset is annotated by over 40 native speakers under expert supervision, with careful quality control. Second, we present OkaLM, the first language models for Kwanyama (1B, 3B, and 8B parameters), obtained by continued pretraining of LLaMA-3 checkpoints on a curated Kwanyama corpus. Together, OkaSentiment and OkaLM bring a left-behind language into the landscape of modern NLP, providing its first benchmark and language models.
TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla
Nishat Raihan | Antonios Anastasopoulos | Marcos Zampieri
Nishat Raihan | Antonios Anastasopoulos | Marcos Zampieri
Despite being the 5th most spoken language, Bangla remains underrepresented in Large Language Models (LLMs), particularly for code generation. This primarily stems from the scarcity of high-quality data to pre-train and/or finetune such models. Hence, we introduce the first dedicated family of Code LLMs for Bangla (1B & 9B). We offer three major contributions: (1) a comprehensive Bangla code instruction datasets for programming domain adaptation; (2) MBPP-Bangla, an evaluation benchmark for Bangla code generation; and (3) the TigerCoder-family of Code LLMs, achieving significant ~11-18% performance gains at Pass@1 over existing multilingual and general-purpose Bangla LLMs. Our findings show that curated, high-quality datasets can overcome limitations of smaller models for low-resource languages.
ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models
Duy Vu Minh Nguyen | Chinh Thanh Truong | Trần Hoàng Phúc | Hung Tuan Le | Nguyen Van-Thanh Dat | Trung Hieu Pham | Kiet Van Nguyen
Duy Vu Minh Nguyen | Chinh Thanh Truong | Trần Hoàng Phúc | Hung Tuan Le | Nguyen Van-Thanh Dat | Trung Hieu Pham | Kiet Van Nguyen
Vietnamese medical research has become an increasingly vital domain, particularly with the rise of intelligent technologies aimed at reducing time and resource burdens in clinical diagnosis. Recent advances in vision-language models (VLMs), such as Gemini and GPT-4V, have sparked a growing interest in applying AI to healthcare. However, most existing VLMs lack exposure to Vietnamese medical data, limiting their ability to generate accurate and contextually appropriate diagnostic outputs for Vietnamese patients. To address this challenge, we introduce ViX-Ray, a novel dataset comprising 5,400 Vietnamese chest X-ray images annotated with expert-written findings and impressions from physicians at a major Vietnamese hospital. We analyze linguistic patterns within the dataset, including the frequency of mentioned body parts and diagnoses, to identify domain-specific linguistic characteristics of Vietnamese radiology reports. Furthermore, we fine-tune five state-of-the-art open-source VLMs on ViX-Ray and compare their performance to leading proprietary models, GPT-4V and Gemini. Our results show that while several models generate outputs partially aligned with clinical ground truths, they often suffer from low precision and excessive hallucination, especially in impression generation. These findings not only demonstrate the complexity and challenge of our dataset but also establish ViX-Ray as a valuable benchmark for evaluating and advancing vision-language models in the Vietnamese clinical domain.
Creating Task-Specific Speech Recognition Datasets from Scratch for Low-Resource Languages: Assessing the Impact of Token Sequence Overlap
Adwoa Asantewaa Bremang | Dennis Asamoah Owusu | Victor Kow Quagraine | Leanne M.M. Annor-Adjaye
Adwoa Asantewaa Bremang | Dennis Asamoah Owusu | Victor Kow Quagraine | Leanne M.M. Annor-Adjaye
Creating a task-specific speech recognition dataset is essential for developing speech recognition applications in low-resource languages. Such applications have uses in agriculture, finance, healthcare, and others, and benefit individuals with low literacy. However, a significant challenge is the high cost of data creation. While there is some work around cost-effective dataset selection, there is little to no work on building a cost-effective dataset for a task from scratch. Our work contributes to the latter. We created a speech recognition dataset from scratch and conducted two major sets of experiments. The first aimed to observe the effect of different datasets of the same size on model performance. Our results confirmed that the same amount spent collecting data can have vastly different results. The second experiment analyzed the effect of token sequence overlap between target and training data since a natural and intuitive approach to building a dataset from scratch for task would be having the task tokens occur in the training data. Our experiments showed that token sequence overlap was not the primary factor influencing model performance. Our work provides a counter-intuitive insight into building speech recognition datasets from scratch in low-resource settings and shows the need for further investigation.
Radio Haiti-Inter: A Large-Scale Annotated Corpus of Spoken Haitian Creole
William N. Havard | Rayan Ziane | Mélissa Menclé | Maximin Coavoux | Benjamin Lecouteux | Emmanuel Schang
William N. Havard | Rayan Ziane | Mélissa Menclé | Maximin Coavoux | Benjamin Lecouteux | Emmanuel Schang
We present the first large-scale corpus of spoken Haitian Creole (Kreyòl), namely Radio Haiti-Inter. The corpus was constructed using automatic speech recognition (ASR) with a state-of-the-art model specifically dedicated to Kreyòl. In addition to transcriptions, we provide part-of-speech (POS) tags, as well as time-aligned transcripts and confidence scores, enabling users to select the most reliable segments for their research. We conduct a manual evaluation of both the transcription quality and POS tagging accuracy to assess the reliability of the resource we present. To enable high-quality research with the resource we introduce, we are releasing 50 hours, comprising both the audios and attached annotations, drawn from the highest-quality segments. This corpus represents an invaluable resource for advancing the study of Kreyòl, with potential applications in phonetics, phonology, morphology, syntax, as well as the study of code-switching and code-mixing. As the recordings cover a large span of years, the corpus we introduce is also suited to micro-diachronic studies of Kreyòl.
Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages
Nick McKenna | Xinnuo Xu | Jack Williams | Nicholas C. Wilson | Benjamin Van Durme | Christian Poelitz
Nick McKenna | Xinnuo Xu | Jack Williams | Nicholas C. Wilson | Benjamin Van Durme | Christian Poelitz
A key consideration when training an LLM is whether the target language is more or less resourced, for example English compared to Welsh, or Python compared to Excel. Typical training data for programming languages consists of real program demonstrations coupled with explanatory human-written comments. In this work we present a novel approach to the creation of such data for low resource programming languages, which lack naturally occurring data. Our process generates synthetic, textbook-quality demonstrations of how to use library functions, which we show makes for good model finetuning data. We demonstrate in an example domain of Excel Formulas. First, we collate language documentation, then we use this to augment a powerful teacher model which generates synthetic training data, and finally finetune student models on the demonstrations. Our technique improves student performance on 2 question-answering datasets: WikiTQ and TAT-QA. We also show advantages of finetuning over standard RAG approaches, which can offer only modest improvement due to the unfamiliarity of the target domain to student models.
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
Mohammad Hosseini | Kimia Hosseini | Shayan Bali | Zahra Zanjani | Saeedeh Momtazi
Mohammad Hosseini | Kimia Hosseini | Shayan Bali | Zahra Zanjani | Saeedeh Momtazi
Hallucination is a persistent issue affecting all large language Models (LLMs), particularly within low-resource languages such as Persian. PerHalluEval (Persian Hallucination Evaluation) is the first dynamic hallucination evaluation benchmark tailored for the Persian language. Our benchmark leverages a three-stage LLM-driven pipeline, augmented with human validation, to generate plausible answers and summaries regarding QA and summarization tasks, focusing on detecting extrinsic and intrinsic hallucinations. Moreover, we used the log probabilities of generated tokens to select the most believable hallucinated instances. In addition, we engaged human annotators to highlight Persian-specific contexts in the QA dataset in order to evaluate LLMs’ performance on content specifically related to Persian culture. Our evaluation of 12 LLMs, including open- and closed-source models using PerHalluEval, revealed that the models generally struggle in detecting hallucinated Persian text. We showed that providing external knowledge, i.e., the original document for the summarization task, could mitigate hallucination partially. Furthermore, there was no significant difference in terms of hallucination when comparing LLMs specifically trained for Persian with others.
ADAB: Arabic Dataset for Automated Politeness Benchmarking - a Large-Scale Resource for Computational Sociopragmatics
Hend Al-Khalifa | Nadia Ghezaiel | Maria Bounnit | Hend Hamed Alhazmi | Noof Abdullah Alfear | Reem Fahad Alqifari | Ameera Masoud Almasoud | Sharefah Ahmed Al-Ghamdi
Hend Al-Khalifa | Nadia Ghezaiel | Maria Bounnit | Hend Hamed Alhazmi | Noof Abdullah Alfear | Reem Fahad Alqifari | Ameera Masoud Almasoud | Sharefah Ahmed Al-Ghamdi
The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for politeness detection remain severely under-explored, despite the rich and complex politeness expressions deeply embedded in Arabic communication. In this paper, a new annotated Arabic dataset, called ADAB/أدب (Arabic Politeness Dataset), was generated and carefully collected from four diverse online platforms including social media, e-commerce, and customer service domains, encompassing both Modern Standard Arabic (MSA) and multiple dialectal varieties (Gulf, Egyptian, Levantine, and Maghrebi). This dataset has undergone a thorough annotation process guided by Arabic linguistic traditions and contemporary pragmatic theory, resulting in three-way politeness classifications: polite, impolite, and neutral. The generated dataset contains 10,000 samples with detailed linguistic feature annotations across 16 politeness categories, achieving substantial inter-annotator agreement (κ = 0.703). A comprehensive benchmarking of this dataset was conducted utilizing 40 model configurations spanning traditional machine learning (12 models), transformer-based architecture (10 models), and large language models (18 configurations), thereby effectively demonstrating its practical utility and inherent challenges. This generated resource aims to bridge the gap in Arabic sociopragmatic NLP and encourage further research into politeness-aware applications for the Arabic language.
GRDD+: An Extended Greek Dialectal Dataset with Cross-Architecture Fine-tuning Evaluation
Stergios Chatzikyriakidis | Dimitriοs Papadakis | Sevasti Ioanna Papaioannou | Erofili Psaltaki
Stergios Chatzikyriakidis | Dimitriοs Papadakis | Sevasti Ioanna Papaioannou | Erofili Psaltaki
We present an extended Greek Dialectal Dataset (GRDD+) that complements the existing GRDD dataset with more data from Cretan, Cypriot, Pontic and Northern Greek, while we add six new varieties: Greco-Corsican, Griko (Southern Italian Greek), Maniot, Heptanesian, Tsakonian, and Katharevusa Greek. The result is a dataset with total size 6,374,939 words and 10 varieties. This is the first dataset with such variation and size to date. We conduct a number of fine-tuning experiments to see the effect of good quality dialectal data on a number of LLMs. We fine-tune three model architectures (Llama-3-8B, Llama-3.1-8B, Krikri-8B) and compare the results to frontier models (Claude-3.7-Sonnet, Gemini-2.5, ChatGPT-5).
Same-Language Subtitles for Low-resource Languages: A Case of Bundelkhandi
Anirudh Pradhan | Ayushi Pandey | Divyansh Kushwaha | Akshita Tiwary | Vivek Seshadri
Anirudh Pradhan | Ayushi Pandey | Divyansh Kushwaha | Akshita Tiwary | Vivek Seshadri
Same-language subtitles enhance consumers’ experience for audiovisual content for both hearing impaired population. However, while high-resource languages can benefit from automatic subtitling, subtitles are seldom available for content creators in regional languages. This limits audience engagement on their content, which often is independently produced. This paper presents Project Saurakhi, a platform for generating same-language subtitles in regional languages. To achieve this, we first extract community-generated YouTube videos serve as the primary data source for this project. The current dataset comprises 63 hours of Bundelkhandi speech sourced from 207 YouTube videos across 19 content creators. And second, the technical workflow integrates automated stages with manual refinement via a mobile annotation platform. As regional language content grows both in independent productions, and in over-the-top platforms, Project Saurakhi aims to train women participants in rural India to become proficient in providing subtitles in their native languages.
The Chulalongkorn Corpus of Spoken Thai (CCOST)
Pittayawat Pittayaporn | Cathryn Yang | Sujinat Jitwiriyanont | James Kirby
Pittayawat Pittayaporn | Cathryn Yang | Sujinat Jitwiriyanont | James Kirby
The Chulalongkorn Corpus of Spoken Thai (CCOST) is a phonetically annotated corpus of Standard Thai. The corpus comprises approximately 7 hours of interview-style spontaneous speech from 49 speakers (19 male, 30 female) ranging in age from 18 to 83 years old. Speakers represent diverse regional backgrounds across Thailand but were instructed to speak in Standard Thai. Each speaker also read a 206-item monosyllabic word list twice and a set of 25 sentences three times. The annotation pipeline combines automatic speech recognition (ASR) and forced alignment using CLARIN-D’s OCTRA and Munich Automatic Segmentation System (MAUS) tools with manual correction by phonetically trained native Thai speakers. Transcriptions include orthographic, word-level, syllable-level, and phone-level annotations including toneme labels. The corpus serves as a resource in the sociophonetic investigation of segmental and tonal variation in spontaneous and controlled speech, enabling examination of individual characteristics as well as group differences across age groups, genders, and regional backgrounds. Hand-corrected annotations will additionally serve to improve forced alignment accuracy for Standard Thai.
Nepal Script Text Recognition from Ancient Artifacts: Challenges and Opportunities
Swornim Nakarmi | Sarin Sthapit | Sahil Ratna Tuladhar | Arya Shakya | Bal Krishna Bal | Rajani Chulyadyo
Swornim Nakarmi | Sarin Sthapit | Sahil Ratna Tuladhar | Arya Shakya | Bal Krishna Bal | Rajani Chulyadyo
Nepal Script, a script of significant linguistic, historical, and cultural importance, can be found in ancient artifacts in Nepal. As this script has faced a decline in use, it is considered among endangered scripts at present. For its revival and preservation, it is important to digitize ancient artifacts written in Nepal Script and create an accessible digital dataset. Among such artifacts are stone inscriptions, and manuscripts, from which we attempt to recognize texts using Artificial Intelligence techniques. This paper presents our approach of preparing a dataset through an extensive data acquisition method, and developing a system that recognizes Nepal Script texts from images. Our system combines the YOLOv8 algorithm with Convolutional Recurrent Neural Network architecture and Connectionist Temporal Classification loss. Our dataset consists of 5,219 text line images from ancient stone inscriptions, manuscripts, and modern handwritten and typed documents. Utilizing an augmented dataset of 41,752 samples, our system achieved 12.61% Character Error Rate. Despite the small training dataset, our model successfully predicted texts in not only new stone inscriptions and manuscripts but also wooden and copper plate inscriptions. We expect our contributions will encourage further research on Nepal Script and other Nepalese scripts.
LuxBorrow: From Pompier to Pompjee, Tracing Borrowing in Luxembourgish
Nina Hosseini-Kivanani | Fred Philippy
Nina Hosseini-Kivanani | Fred Philippy
We present LuxBorrow, a borrowing-first analysis of Luxembourgish (LU) news spanning 27 years (1999–2025): 259,305 RTL articles and 43.7M tokens. Our pipeline combines sentence-level language identification (LU/DE/FR/EN) with a token-level borrowing resolver restricted to LU sentences, using lemmatization, a collected loanword registry, and compiled morphological/orthographic rules. Empirically, LU remains the matrix language across all documents, while multilingual practice is pervasive: 77.1% of articles include at least one donor language and 65.4% use three or four. Breadth does not imply intensity: median code-mixing index (CMI) increases from 3.90 (LU+1) to only 7.00 (LU+3), indicating localized insertions rather than balanced bilingual text. Domain/period summaries show moderate but persistent mixing, with CMI rising from 6.1 (1999–2007) to a peak of 8.4 (2020). Token-level adaptations total 25,444 instances and exhibit a mixed profile: morphological 63.8%, orthographic 35.9%, lexical 0.3%; the most frequent single rules are orthographic (on→oun, eur→er), while morphology is collectively dominant. Diachronically, code-switching intensifies, and morphologically adapted borrowings grow from a small base; French overwhelmingly supplies adapted items, with modest growth for German and negligible English. We advocate borrowing-centric evaluation, borrowed token/type rates, donor entropy over borrowed items, and assimilation ratios over headline document-level mixing indices.
Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS
Rania Al-Sabbagh
Rania Al-Sabbagh
Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic research and low-resource language technologies. It contains recordings from structured interviews with native speakers and episodes from national television shows. The corpus features 157 speakers (59 female, 98 male), spans subdialects such as Urban, Bedouin, and Mountain/Shihhi, and covers topics such as cultural heritage, agriculture and sustainability, daily life, professional trajectories, and architecture. It consists of 91 monologic and 79 dialogic recordings, varying in length and recording conditions. A 10% subset was used to evaluate commercial and open-source models for automatic speech recognition (ASR) and text-to-speech (TTS) in a zero-shot setting to establish initial baselines. Whisper-large-v3-turbo achieved the best ASR performance, with average word and character error rates of 0.268 and 0.144, respectively. MMS-TTS-Ara reported the best mean word and character rates of 0.285 and 0.081, respectively, for TTS. These baselines are competitive but leave substantial room for improvement. The paper highlights the challenges encountered and provides directions for future work.
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
Malik H. Altakrori | Nizar Habash | Teresa Lynn | Younes Samih | Abed Alhakim Freihat | Kirill Chirkunov | Muhammed AbuOdeh | Radu Florian | Preslav Nakov | Alham Fikri Aji
Malik H. Altakrori | Nizar Habash | Teresa Lynn | Younes Samih | Abed Alhakim Freihat | Kirill Chirkunov | Muhammed AbuOdeh | Radu Florian | Preslav Nakov | Alham Fikri Aji
We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluation for Modern Standard Arabic (MSA), dialectal varieties remain underrepresented despite their prevalence in everyday communication. DialectalArabicMMLU extends the MMLU-Redux framework through manual translation and adaptation of 3K multiple-choice question–answer pairs into five major dialects (Syrian, Egyptian, Emirati, Saudi, and Moroccan), yielding a total of 15K QA pairs across 32 academic and professional domains (22K QA pairs when also including English and MSA). The benchmark enables systematic assessment of LLM reasoning and comprehension beyond MSA, supporting both task-based and linguistic analysis. We evaluate 19 open-weight Arabic and multilingual LLMs (1B–13B parameters) and report substantial performance variation across dialects, revealing persistent gaps in dialectal generalization. DialectalArabicMMLU provides the first unified, human-curated resource for measuring dialectal understanding in Arabic, thus promoting more inclusive evaluation and future model development.
ForumOccitania: A Corpus of User-Generated Content for Multiple Occitan Varieties
Oriane Nédey | Juliette Janès | Rachel Bawden | Thibault Clérice | Benoît Sagot
Oriane Nédey | Juliette Janès | Rachel Bawden | Thibault Clérice | Benoît Sagot
We introduce ForumOccitania, a new Occitan corpus of posts from an online forum, covering a range of topics and dialects. While some existing datasets for this low-resource language include labels of varieties within the dialect continuum, we go one step further by providing metadata pertaining to sociolinguistic factors of language variation (dialect, geographical location, age, proficiency), extracted from self-declared user profiles. We carry out statistical and qualitative analyses, as well as preliminary experiments on unsupervised dialect identification. Our results show that (i) most of the contents is written in Occitan, with the classical spelling conventions, and by young speakers, (ii) posts display a strong presence of dialectal features from four major Occitan varieties (Lemosin, Lengadocian, Gascon, Provençau), and (iii) a simple topic modelling approach introduced by Kuparinen and Scherrer (2024) effectively detects salient features of these four varieties, but also reveals finer-grained diatopical variation tendencies.
A Dataset of Wolof Ajami Manuscripts for HTR and OCR
Oreen Yousuf | Elhadji Djibril Diagne | Christian Høgel | Beata Megyesi | Joakim Nivre
Oreen Yousuf | Elhadji Djibril Diagne | Christian Høgel | Beata Megyesi | Joakim Nivre
We present the first ever dataset of manually segmented and transcribed Ajami manuscripts written in Wolof. The term Ajami refers to modified Arabic-script orthographies used to transcribe African languages. Handwritten text recognition (HTR) and optical character recognition (OCR) models for Arabic-script languages perform poorly on African languages written in Ajami orthographies because these languages are not represented in the pre-training data of the models. This leads to recognition models being unable to extract unique Arabic-script letters and ubiquitous diacritics used in African languages, and struggling to adapt to various calligraphy styles used across Africa. We release the following as an open-source dataset: an ALTO formatting of high-quality images of handwritten and printed, 20th–century Wolof manuscripts; manual segmentation (region and line); and manual transcriptions. We extend our contribution by evaluating several Arabic-script recognition models intended for historical manuscripts and find they produce character error rates (CER) of 61–81%. Transcriptions produced by the evaluated recognition models, as well as a keyboard to transcribe Wolof Ajami manuscripts, are released as well. The digitally transcribed text in the dataset can also be utilized for various natural language processing (NLP) and historical linguistic tasks.
TDMulti: A Tunisian Dialect-Modern Standard Arabic Multitask Corpus with a Context-Aware Cross-Attention BERT Model
Roua Torjmen | Kais HADDAR
Roua Torjmen | Kais HADDAR
The Tunisian dialect dominates online communication in Tunisia but remains severely under-resourced in natural language processing. We introduce the first multitask corpus of Tunisian dialect manually aligned with its equivalents in modern standard Arabic. The TDMulti corpus consists of 3,100 social media comments annotated with 12,400 labels for four interrelated tasks: hate speech detection, sentiment polarity classification, sarcasm identification, and topic category classification. The TDMulti corpus provides a new benchmark for studying pragmatic and social aspects of Tunisian dialect in relation to modern standard Arabic. To exploit this resource, we propose a deep learning model based on transformer architectures. We design three variants: a baseline multitask classifier, a cross-attention model aligning Tunisian dialect and modern standard Arabic representations, and a context-aware cross-attention mechanism with task-specific masking. We evaluate the approach using large pre-trained Arabic language models under different configurations. Results show that the context-aware cross-attention model achieves the best performance, particularly for sarcasm and hate speech detection. TDMulti is released under an open license, contributing a novel resource to advance research on Arabic dialect processing.
The Megrelian Language Corpus (MLC): Creation, Annotation, and Initial Steps toward a UD Treebank
Irina Lobzhanidze | Rusudan Gersamia | Tamar Gogia
Irina Lobzhanidze | Rusudan Gersamia | Tamar Gogia
This paper presents the development of the Megrelian Language Corpus (MLC), a new language resource for the documentation and computational analysis of Megrelian, an endangered Kartvelian language. The corpus is based on fieldwork conducted in Samegrelo, Georgia (2022–2024) and currently contains 97,691 tokens and 60,959 types. The data were transcribed using the International Phonetic Alphabet (IPA) and annotated in Fieldworks Language Explorer (FLEx) with segmentation, morphological analysis and bilingual Georgian-English translations. Each text is accessible through a specially designed web interface, providing multiple tiers of annotation and integrated search functions. The paper describes the corpus design, annotation methodology and challenges encountered in representing Megrelian’s complex agglutinative morphology. It also outlines initial steps toward converting existing data into the Universal Dependencies (UD) framework, building on experience from related Kartvelian languages such as Georgian. The MLC corpus represents the first publicly available linguistic resource for Megrelian and provides a foundation for future UD treebank development.
Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
Keunhyeung Park | Seunguk Yu | Youngbin Kim
Keunhyeung Park | Seunguk Yu | Youngbin Kim
Standard-to-dialect machine translation remains challenging due to a persistent dialect gap in large language models and evaluation distortions inherent in n-gram metrics, which favor source copying over authentic dialect translation. In this paper, we propose the dialect refinement (DIA-REFINE) framework, which guides LLMs toward faithful target dialect outputs through an iterative loop of translation, verification, and feedback using external dialect classifiers. To address the limitations of n-gram-based metrics, we introduce the dialect fidelity score (DFS) to quantify linguistic shift and the target dialect ratio (TDR) to measure the success of dialect translation. Experiments on Korean dialects across zero-shot and in-context learning baselines demonstrate that DIA-REFINE consistently enhances dialect fidelity. The proposed metrics distinguish between False Success cases, where high n-gram scores obscure failures in dialectal translation, and True Attempt cases, where genuine attempts at dialectal translation yield low n-gram scores. We also observed that models exhibit varying degrees of responsiveness to the framework, and that integrating in-context examples further improves the translation of dialectal expressions. Our work establishes a robust framework for goal-directed, inclusive dialect translation, providing both rigorous evaluation and critical insights into model performance.
LombardoGraphia: Automatic Classification of Lombard Orthography Variants
Edoardo Signoroni | Pavel Rychly
Edoardo Signoroni | Pavel Rychly
Lombard, an underresourced language variety spoken by approximately 3.8 million people in Northern Italy and Southern Switzerland, lacks a unified orthographic standard. Multiple orthographic systems exist, creating challenges for NLP resource development and model training. This paper presents the first study of automatic Lombard orthography classification and LombardoGraphia, a curated corpus of 11,186 Lombard Wikipedia samples tagged across 9 orthographic variants, and models for automatic orthography classification. We curate the dataset, processing and filtering raw Wikipedia content to ensure text suitable for orthographic analysis. We train 24 traditional and neural classification models with various features and encoding levels. Our best models achieve 96.06% and 85.78% overall and average class accuracy, though performance on minority classes remains challenging due to data imbalance. Our work provides crucial infrastructure for building variety-aware NLP resources for Lombard.
Meenz bleibt Meenz, but Large Language Models Do Not Speak Its Dialect
Minh Duc Bui | Manuel Mager | Peter Herbert Kann | Katharina von der Wense
Minh Duc Bui | Manuel Mager | Peter Herbert Kann | Katharina von der Wense
Meenzerisch, the dialect spoken in the German city of Mainz, is also the traditional language of the Mainz carnival, a yearly celebration well known throughout Germany. However, Meenzerisch is on the verge of dying out—a fate it shares with many other German dialects. Natural language processing (NLP) has the potential to help with the preservation and revival efforts of languages and dialects. However, so far no NLP research has looked at Meenzerisch. This work presents the first research in the field of NLP that is explicitly focused on the dialect of Mainz. We introduce a digital dictionary—an NLP-ready dataset derived from an existing resource—to support researchers in modeling and benchmarking the language. It contains 2,351 words in the dialect paired with their meanings described in Standard German. We then use this dataset to answer the following research questions: (1) Can state-of-the-art large language models (LLMs) generate definitions for dialect words? (2) Can LLMs generate words in Meenzerisch, given their definitions? Our experiments show that LLMs can do neither: the best model for definitions reaches only 6.27% accuracy and the best word generation model’s accuracy is 1.51%. We then conduct two additional experiments in order to see if accuracy is improved by few-shot learning and by extracting rules from the training set, which are then passed to the LLM. While those approaches are able to improve the results, accuracy remains below 10%. This highlights that additional resources and an intensification of research efforts focused on German dialects are desperately needed.
Bootstrapping NLP for Sakha: Named Entity Recognition and Sentiment Analysis in an Extremely Low-Resource Setting
Mariia Everstova | Nikolai Efimov | Valerio Basile
Mariia Everstova | Nikolai Efimov | Valerio Basile
We present the first systematic study of core NLP tasks for Sakha (Yakut), a low-resource Turkic language with approximately 450,000 speakers in northeastern Siberia. We introduce two manually annotated datasets: a 690-sentence NER corpus (921 entities: PER, LOC, ORG) and an 798-sentence sentiment corpus (positive, negative, neutral). Using mBERT and RuBERT in controlled 2×2 experiments, we report a twofold effect: on the one hand, it improves performance when base unknown-token rates exceed approximately 10% (RuBERT: +9.4 F1); on the other hand, it leads to worse performance otherwise (mBERT: −6.1 F1), despite improving tokenization in both cases. Cross-domain transfer (news vs forums) reveals severe asymmetry: formal-to-informal training achieves 47% accuracy while the reverse yields only 26%—a 21-point gap demonstrating that domain composition dominates model architecture choice in low-resource settings. Neutral-boundary detection is the primary bottleneck, with 89% of disagreements clustering around subjective/objective distinctions rather than polarity confusions. With fewer than 1,000 samples per task, we establish first benchmarks for Sakha NER (53.5 F1) and sentiment analysis (54% accuracy).
Lightweight Cross-Lingual Federated Prompt Tuning for Low-Resource Languages
Ubaid Azam | Imran Razzak | Shoaib Jameel
Ubaid Azam | Imran Razzak | Shoaib Jameel
Multilingual NLP faces challenges of data heterogeneity, privacy, and limited computational resources, especially for low-resource languages. Centralised methods risk privacy breaches, while federated learning struggles with communication overhead and poor cross-lingual generalisation. We propose FLiP (Federated Lightweight Prompt-tuning), a privacy-preserving, resource-efficient, generalizable framework integrating prompt-based learning with federated optimisation. FLiP eliminates communication overhead, reduces trainable parameters to 16%, and cuts GPU memory use by 90%. Experiments show superior generalisation and efficiency under both IID and Non-IID settings, establishing FLiP as a scalable, privacy-aware solution for multilingual NLP, particularly in low-resource and indigenous language contexts.
A Parallel Corpus of the Parable of the Prodigal Son: Building a Resource for Documenting Language Varieties in Mainland France
Lucence Ing | Juliette Janès | Sven Ködel | Benoît Sagot
Lucence Ing | Juliette Janès | Sven Ködel | Benoît Sagot
This paper presents a historical parallel corpus of languages spoken in metropolitan France. It consists of a collection of versions of the Parable of the Prodigal Son, collected during the 19th century. The paper aims to present the interest of such a corpus, its constitution—through XML/TEI encoding, semi-automatic alignment and projection on linguistic maps—and its potential uses for the study of these low-resource languages.
Developing Zila: A Spoken Language Resource for the Endangered Slovenian Gail Valley Dialect
Andrej Zgank | Gregor Donaj | Urh Kolaric | Usi Sereinig | Tatjana Koren-Zwitter | Sanja Boto | Sabina Zwitter-Grilc | Jasna Vidinic | Darinka Verdonik
Andrej Zgank | Gregor Donaj | Urh Kolaric | Usi Sereinig | Tatjana Koren-Zwitter | Sanja Boto | Sabina Zwitter-Grilc | Jasna Vidinic | Darinka Verdonik
Slovenian is a less-resourced South Slavic language. Existing Slovenian spoken language resources mainly cover the standard language in everyday communication. However, Slovenian encompasses a wide range of dialects, most of which are not represented in available spoken language resources. This paper presents the development of Zila, a Slovenian spoken language resource for the Gail Valley dialect. This dialect is one of the most endangered varieties of Slovenian and is spoken in the extreme north-western periphery of the Slovenian language area. The goal of the project was to build a language resource comprising 100 hours of speech with manually produced transcriptions. The spoken material was collected from members of the Slovenian minority in Carinthia, Austria, with the local community playing a key role in the data acquisition process. A dedicated set of transcription rules was created to capture the full range of acoustic and linguistic features of the Gail Valley dialect, which differs significantly from standard Slovenian. A preliminary speech recognition experiment was conducted to analyze these differences further. The Zila project demonstrates how spoken language technologies can help to preserve the cultural and linguistic heritage of an endangered dialect.
Nawatl Context-Free Grammars for Natural Language Processing
Juan Jose Guzman Landa | Juan-Manuel Torres-Moreno | Graham Ranger | Miguel Figueroa-Saavedra | Ligia Quintana Torres | Carlos-Emiliano Gonzalez-Gallardo | Luis Gil Moreno Jimenez | Martha Lorena Avendaño Garrido
Juan Jose Guzman Landa | Juan-Manuel Torres-Moreno | Graham Ranger | Miguel Figueroa-Saavedra | Ligia Quintana Torres | Carlos-Emiliano Gonzalez-Gallardo | Luis Gil Moreno Jimenez | Martha Lorena Avendaño Garrido
The aim of this article is to introduce Context-Free Grammars (CFG) for the Nawatl language. Nawatl is an Amerindian language of the 𝜋-language type, i.e. a language with few digital resources. For this reason the corpora available for the learning of Large Language Models (LLMs) are virtually non-existent, posing a significant challenge. The goal is to produce a substantial number of syntactically valid artificial Nawatl sentences and thereby to expand the corpora for the purpose of learning embeddings (static models or probably LLMs). For this objective, we introduce two new Nawatl CFGs and use them in generative mode. Thanks to these grammars, it is possible to expand Nawatl corpus significantly and subsequently to use it to learn embeddings (such as FastText) and to evaluate their relevance in semantic similarity tasks. The results show an improvement compared to the results obtained using only the original corpus without artificial expansion, and also demonstrate that economic embeddings often perform better than some LLMs.
Physical Commonsense Reasoning for Lower-Resourced Languages and Dialects: A Study on Basque
Jaione Bengoetxea | Itziar Gonzalez-Dios | Rodrigo Agerri
Jaione Bengoetxea | Itziar Gonzalez-Dios | Rodrigo Agerri
Physical commonsense reasoning represents a fundamental capability of human intelligence, enabling individuals to understand their environment, predict future events, and navigate physical spaces. Recent years have witnessed growing interest in reasoning tasks within Natural Language Processing (NLP). However, no prior research has examined the performance of Large Language Models (LLMs) on non-question-answering (non-QA) physical commonsense reasoning tasks in low-resource languages such as Basque. Taking the Italian GITA as a starting point, this paper addresses this gap by presenting BasPhyCo, the first non-QA physical commonsense reasoning dataset for Basque, available in both standard and dialectal variants. We evaluate model performance across three hierarchical levels of commonsense understanding: (1) distinguishing between plausible and implausible narratives (accuracy), (2) identifying the conflicting element that renders a narrative implausible (consistency), and (3) determining the specific physical state that creates the implausibility (verifiability). These tasks were assessed using multiple multilingual LLMs as well as models pretrained specifically for Italian and Basque. Results indicate that, in terms of verifiability, LLMs exhibit limited physical commonsense capabilities in low-resource languages such as Basque, especially when processing dialectal variants.
Common Voice for Pakistan: Developing an Open Speech Corpus for Low-Resource Pakistani Languages
Meesum Alam | Francis Tyers
Meesum Alam | Francis Tyers
Pakistan is home to more than 70 languages out of which 30 languages are endangered. Most of Pakistani languages remain absent from modern speech and text technologies, with resources focused on Urdu and a few major tongues. Through Mozilla’s Open Multilingual Speech Fund, this paper documents one year project for the development of an open, community driven speech corpus for 39 indigenous languages of Pakistan. The dataset includes locally authored texts, daily life sentences, poetry, and folk songs to make a culturally balanced. The project not only supports Automatic Speech Recognition but also promote linguistic preservation and digital inclusion.
Amulwe Kimün: A Community-Grounded Demo, Resource, and ASR Baseline for Mapuzugun
Cristian Eduardo Ahumada Oliva | Fatiha Sadat
Cristian Eduardo Ahumada Oliva | Fatiha Sadat
This paper introduces Amulwe Kimün (“a means or path for knowledge” in Mapuzugun), a community-grounded multimodal quiz application co-created with Mapuche speakers to support the revitalization of Mapuzugun. Developed within a FACSO–CONADI collaboration during an intensive language course, the platform integrates multiple-choice, ordering and free-text exercises, as well as forums and chat functions to promote language practice, peer learning, and a sense of community. A pilot involving 32 learners produced 562 responses across 43 questions, with accuracies of 92.3% (multiple choice), 55.2% (ordering), and 7.1% (free-text), offering insights for refining item design and evaluation strategies. The low open-answer accuracy is related to a strict exact-match scoring and orthographic variation of the language. In addition, we present an initial Automatic Speech Recognition (ASR) prototype (Whisper-small + LoRA), establishing a fine-tuned baseline relative to zero-shot performance. The demo illustrates how community-grounded design, language resources, and lightweight evaluation can productively meet in a practical tool for an endangered language.
Development of Serbian QA Datasets through Prompt-Based Generation and Human Validation
Jovana Rađenović | Olivera Kitanović | Ranka Stankovic | Mihailo Škorić
Jovana Rađenović | Olivera Kitanović | Ranka Stankovic | Mihailo Škorić
LLMs capable of answering questions, fulfilling diverse user requests, and functioning as chatbots rely heavily on extensive datasets. However, for the Serbian language, there is a significant lack of high-quality datasets structured in a question-and-answer (QA) format. To address this, we extracted a portion of the SQuAD-sr dataset, which, to the best of our knowledge, is the largest QA dataset in Serbian and contains over 87k samples. While this dataset is an incredibly valuable resource, it was translated using an adapted Translate-Align-Retrieve method and contains errors and terminological inaccuracies. In this work, we systematically reviewed and corrected more than 7k samples from the SQuAD-sr dataset, significantly improving the dataset’s reliability and quality. We call this modified subset of the SQuAD-sr dataset, the SQuAD-sr-md dataset. The corrections that were made are crucial for training accurate and robust QA models in Serbian, ensuring that AI systems can leverage the full potential of this dataset. We also introduce an additional QA dataset generated from encyclopedia articles, Wikipedia pages, and scientific paper abstracts using LLMs, which contains 74k samples. We name this dataset the SerbianQA-Gen.
An Enhanced Pipeline for the Manzini-Savoia Dialect Corpus
Achille Fusco | Greta Mazzaggio | Carlo Zoli
Achille Fusco | Greta Mazzaggio | Carlo Zoli
This paper presents a semi-automatic workflow for enriching the Manzini–Savoia Corpus (MSC) of Italian dialects with extended glosses, normalized transcriptions, and projected morpho-syntactic annotations. While the MSC is a unique resource for Romance microvariation, its partial glossing and phonetic transcription in the International Phonetic Alphabet (IPA) pose major challenges for computational processing. We introduce a pipeline for gloss coverage expansion and reliable morpho-syntactic annotation combining rule-based and data-driven components, which includes: (i) automatic completion of truncated verbal paradigms; (ii) hybrid lexical alignment between dialectal tokens and Italian glosses, integrating per-region lexical priors with a dynamic programming alignment algorithm; and (iii) projection-based morpho-syntactic tagging from aligned glosses. The proposed methods offer a reproducible framework for extending partially glossed dialect corpora and contribute new annotated data for research in computational dialectology and cross-variety language modeling.
Are Language Models Borrowing-Blind? A Multilingual Evaluation of Loanword Identification across 10 Languages
Merilin Sousa Silva | Sina Ahmadi
Merilin Sousa Silva | Sina Ahmadi
Throughout language history, words are borrowed from one language to another and gradually become integrated into the recipient’s lexicon. Speakers can often differentiate these loanwords from native vocabulary, particularly in bilingual communities where a dominant language continuously imposes lexical items on a minority language. This paper investigates whether pretrained language models, including large language models, possess similar capabilities for loanword identification. We evaluate multiple models across 10 languages. Despite explicit instructions and contextual information, our results show that models perform poorly in distinguishing loanwords from native ones. These findings corroborate previous evidence that modern NLP systems exhibit a bias toward loanwords rather than native equivalents. Our work has implications for developing NLP tools for minority languages and supporting language preservation in communities under lexical pressure from dominant languages.
Comparing Approaches to Automatic Summarization in Less-Resourced Languages
Chester Palen-Michel | Constantine Lignos
Chester Palen-Michel | Constantine Lignos
Automatic text summarization has achieved high performance in higher-resourced languages like English, but comparatively less attention has been given to summarization in less-resourced languages. This work compares a variety of approaches to summarization from zero-shot prompting of LLMs large and small to fine-tuning smaller models like mT5 with and without three data augmentation approaches and multilingual transfer. We also explore an LLM translation pipeline approach, translating from the source language to English, summarizing and translating back. Evaluating with five different metrics, we find that there is variation across LLMs in their performance at similar model sizes, that our multilingual fine-tuned mT5 baseline outperforms most other approaches including zero-shot LLM performance for most metrics, and that LLM as judge may be unreliable on less-resourced languages.
PsihoRo: Depression and Anxiety Romanian Text Corpus
Alexandra Ciobotaru | Ana-Maria Bucur | Liviu P. Dinu
Alexandra Ciobotaru | Ana-Maria Bucur | Liviu P. Dinu
Psychological corpora in NLP are collections of texts used to analyze human psychology, emotions, and mental health. These texts allow researchers to study psychological constructs, identify patterns related to mental health problems and analyze emotional language. However, collecting accurate mental health data from social media can be challenging due to the assumptions made by data collectors. A more effective approach involves gathering data through open-ended questions and then assessing participants’ mental health status using self-report screening surveys. This method was successfully employed for English, a language with a lot of psychological NLP resources. However, the same cannot be stated for Romanian, which currently has no open-source mental health corpus. To address this gap, we have collected the first open-source corpus focused on depression and anxiety in Romanian, by utilizing a form with 6 open-ended questions along with the standardized PHQ-9 and GAD-7 screening questionnaires. Although the PsihoRo corpus contains texts from only 205 respondents, it represents an important first step toward understanding and analyzing mental health issues within the Romanian population. We employ statistical analysis, text analysis using Romanian LIWC, emotion detection, and topic modeling to identify the most important features of this newly introduced resource for the NLP community. The data is publicly available at https://huggingface.co/datasets/Alegzandra/PsihoRo.
Aligned Parallel Corpus of the Vedic Saṁhitās for Machine Translation
Yuzuki Tsukagoshi | Ikki Ohmukai
Yuzuki Tsukagoshi | Ikki Ohmukai
We introduce a verse-/paragraph-aligned parallel corpus for three Vedic Saṁhitās –the R̥gveda (R̥V), the Atharvaveda Śaunaka (AVŚ), and the Taittirīya Saṁhitā (TS)– paired with authoritative public-domain translations (Geldner for R̥V, Whitney for AVŚ, and Keith for TS). The source texts are drawn from established digital editions (e.g., TITUS and VedaWeb) and normalized under ISO 15919. Each Sanskrit segment is aligned to exactly one translated unit (verse or paragraph for TS prose), yielding a unified, model-ready format. Using this resource, we fine-tune and evaluate three large language models –GPT-4.1 nano, Gemini 2.5 Flash, and Mitra– on Vedic→German/English translation. Evaluation combines surface and semantic metrics (case-insensitive sacreBLEU and COMET), enabling a balanced assessment of form and meaning. Results show consistent in-domain gains after supervised fine-tuning, but substantial cross-domain degradation when models are tested on unseen Saṁhitās, indicating pronounced stylistic and lexical divergence among R̥V, AVŚ, and TS. These findings motivate domain-aware training and reporting practices for Vedic machine translation. We release the corpus with standardized splits and preprocessing to support reproducibility and future d research on historical language modeling, alignment, and translation for low-resource ancient languages.
FormosanMT: A Multilingual Parallel Corpus of the Formosan Language Family
Hunter Scheppat | Joshua K. Hartshorne | Sema Koc | Éric Le Ferrand | Emily Prud’hommeaux
Hunter Scheppat | Joshua K. Hartshorne | Sema Koc | Éric Le Ferrand | Emily Prud’hommeaux
While the quality of machine translation (MT) between widely-spoken languages has improved dramatically in recent years, training robust MT systems for languages with fewer resources remains a challenge. Endangered languages, which often lack the speaker population and written tradition needed to create text resources, are at a particular disadvantage. Developing robust MT architectures for very low-resource settings is hampered by the lack of suitable parallel corpora. To address this challenge, we introduce FormosanMT, a set of MT-ready parallel corpora for the Formosan family of endangered languages indigenous to Taiwan. Together the corpora total nearly 500,000 Formosan-Mandarin and Formosan-English sentence pairs. We share scripts for extracting these corpora from public sources, along with customizable tools for filtering, normalizing, and partitioning the data. In addition, we provide a new tokenizer for Traditional Chinese writing compatible with the popular No Language Left Behind (NLLB) MT architecture, along with updated and improved code for fine-tuning NLLB for any low-resource language pair. Finally we distribute our fully trained NLLB and OpenNMT models for the Formosan languages to and from both Mandarin and English. In addition to serving as a valuable resource for the Formosan language speaker communities, our data, code, and models will be available to NLP researchers working on endangered and low-resource language MT.
The Construction of a Mixe Variant Parallel Corpus
Ivan Vladimir Meza Ruiz | Delfino Zacarias Marquez | Martha Elba Ramírez Andrés | Victoriano Santiago Cayetano | Jonathan Santiago Antonio | Carlos Daniel Hernández Mena
Ivan Vladimir Meza Ruiz | Delfino Zacarias Marquez | Martha Elba Ramírez Andrés | Victoriano Santiago Cayetano | Jonathan Santiago Antonio | Carlos Daniel Hernández Mena
We present the progress and challenges of constructing a Mixe-Spanish parallel corpus for Machine Translation. Mixe is a Mexican Indigenous Language that is spoken by more than 100, 000 speakers. In particular, we focus on the San Juan Guivicovic Mixe variant (mir). The resulting resource is available under an open research license (CC BY-NC-SA). It was created following a previous state-of-the-art methodology for Mexican indigenous languages. In this case, we used paid translators from the variant region. We present a baseline system.
Nepali Lemmatization with Multilingual Transformers: Intrinsic and Extrinsic Evaluation in a Low-Resource Setting
Sunil Regmi | Sundeep Dawadi | Bal Krishna Bal
Sunil Regmi | Sundeep Dawadi | Bal Krishna Bal
The Nepali language has a rich and complex morphology. Existing lemmatization research focuses on traditional rule-based or TRIE-based approaches. These methods often fail when encountering out-of-vocabulary or misspelled words. This paper investigates neural lemmatization for the under-resourced Nepali language using multilingual transformer models. We formulate lemmatization as a text-to-text generation problem and evaluate its impacts on downstream tasks by finetuning mBART-large-50, mT5-base, and mT5-small. The models were trained on a combination of publicly available and human-annotated word-lemma pair (8,000 instances) dataset. The performance is evaluated using Character Error Rate (CER), accuracy, character-level Bilingual Evaluation Understudy (BLEU), and morphological coverage. The mT5-base model achieved the highest overall performance. The model achieved 96.1% accuracy and a 1.1% CER using a learning rate of 5 × 10−4. However, it showed slightly weaker performance in handling complex morphological variations. The mBART-large-50 model followed closely with 96.0% accuracy and 0.970 morphological coverage. To assess the efficacy of these models, we applied lemmatization to downstream tasks. In Hindi-Nepali cross-lingual alignment, performance improved significantly from 12.86% to 41.61% using mBART model. In information retrieval, the Mean Average Precision (MAP)@1 using binary index increased from 0.71 to 0.90 using mBART model. These results demonstrate that multilingual transformers effectively learn morphological transformations for low-resource languages through text-to-text generation.
Diacritic Restoration for Low-Resource Indigenous Languages: Case Study with Bribri and Cook Islands Māori
Rolando Coto-Solano | Daisy Li | Manoela Teleginski Ferraz | Olivia Sasse | Cha Krupka | Sharid Loaiciga | Sally Akevai Tenamu Nicholas
Rolando Coto-Solano | Daisy Li | Manoela Teleginski Ferraz | Olivia Sasse | Cha Krupka | Sharid Loaiciga | Sally Akevai Tenamu Nicholas
We present experiments on diacritic restoration, a form of text normalization essential for creating and processing data in natural language processing (NLP) tasks. Our study focuses on two extremely under-resourced languages: Bribri, a Chibchan language spoken in Costa Rica, and Cook Islands Māori, a Polynesian language spoken in the Cook Islands. Specifically, this paper: (i) compares algorithms for diacritics restoration in under-resourced languages, including tonal diacritics, (ii) examines the amount of data required to achieve target performance levels, (iii) contrasts results across varying resource conditions, and (iv) explores the related task of diacritic correction. We find that fine-tuned, character-level LLMs perform best, likely due to their ability to decompose complex characters into their UTF-8 byte representations. In contrast, massively multilingual models perform less effectively given our data constraints. Across all models, reliable performance begins to emerge with data budgets of around 10,000 words. Zero-shot approaches perform poorly in all cases. This study responds both to requests from the language communities and to broader NLP research questions concerning model performance and generalization in under-resource contexts.
A Modern Online Learning Platform for ʻŌlelo Hawaiʻi Classrooms
Christian Castro | Keneth Martin | Winston Wu | William H. Wilson
Christian Castro | Keneth Martin | Winston Wu | William H. Wilson
We present Hōʻoi Aʻo, a browser-based platform designed to streamline the teaching workflow and enhance the learning experience for students in Hawaiian language classes. Built with modern technologies including FastAPI, React, and MongoDB, the platform provides an intuitive and specialized environment for both instructors and students of ʻŌlelo Hawaiʻi. Our platform enables instructors to add content, create or import quizzes in multiple formats, view and analyze common student mistakes, and ultimately save time through automatic grading. Students can access chapters, lessons, assignments, and quizzes all in one place, with automatically graded quizzes for instant feedback and unlimited randomly-generated practice questions created using an innovative synchronous context-free grammar approach, allowing students to obtain extra language practice outside of class. Currently, the platform supports content from Book 1 of Nā Kai ʻEwalu, a popular Hawaiian textbook. Hōʻoi Aʻo not only makes language practice more accessible for a language with few existing learning resources, but also represents a step toward a more modern and effective digital ecosystem for teaching and learning ʻŌlelo Hawaiʻi.
The Northern Interior subgroup of the Salish language family, spoken in the Pacific Northwest of North America, comprises three languages: St’át’imcets, nɬeʔkepmxcín, and Secwepemctsín. Each has a small number of first-language (L1) speakers remaining due to the effects of colonization, though language revitalization efforts are ongoing. This work introduces the first compiled and cleaned language datasets in these languages, useable in natural language processing (NLP) projects. This data is in glossed format, with transcriptions in the language, translations into English, and linguistic segmentations and glosses that provide a detailed breakdown of meaning. In order to achieve consistently formatted data within and across each language, extensive data cleaning was conducted. This paper provides the glossed data standards that were developed and recounts the cleaning process. Scripts that help to automate parts of the data preparation processes are included. Finally, this work strives to keep the interconnectedness of language and community as a central consideration.
CEFR-Cymraeg: A Dataset and Baseline Models for Language Proficiency Assessment in Welsh
Eeshan Waqar | Jonathan Davies | Dawn Knight | Fernando Alva-Manchego
Eeshan Waqar | Jonathan Davies | Dawn Knight | Fernando Alva-Manchego
We introduce CEFR-Cymraeg, the first dataset annotated with Common European Framework of Reference (CEFR) levels for Welsh. The dataset is built from learning materials for adult learners, carefully extracted from widely used coursebooks and verified by teachers of Welsh as a second language. It spans levels A1 to B2 and includes multiple units of analysis: sentences, dialogues, paragraphs, and documents. In total, 2,658 entries are provided with gold-standard CEFR annotations, making CEFR-Cymraeg a valuable resource for research on language learning and low-resourced Celtic languages. To illustrate its potential applications, we define language proficiency assessment as a multi-class classification task and fine-tune multilingual pre-trained language models. Given the limited size of the dataset, we also experiment with data augmentation. Results show that these models successfully capture proficiency distinctions and generalise well to Welsh, with the best-performing model reaching a weighted F1-score of 0.83. Qualitative analysis confirmed that most apparent errors reflected valid pedagogical variation rather than model inconsistencies. CEFR-Cymraeg establishes a benchmark resource for Welsh and opens new opportunities for educational NLP, corpus linguistics, and multilingual proficiency research.
Singlish to English Translation with Precision: A Dataset and Language Detection-Driven Masked Modeling for Singlish to English Translation
Sujit Kumar | Gerome Kusuma Ang | Stephanie Hilary Xinyi Ma | Andy Hau Yan Ho | Andy Khong
Sujit Kumar | Gerome Kusuma Ang | Stephanie Hilary Xinyi Ma | Andy Hau Yan Ho | Andy Khong
Singlish, a creole rooted in English and influenced by Singapore’s multilingual and multicultural environment, poses significant challenges for those proficient in standard English due to its unique and often complex lexical and syntactic structures. Despite significant advancements in language translation for both high- and low-resource languages, translating Singlish to English remains largely underexplored. This gap is primarily due to the lack of dedicated datasets for language detection and Singlish-to-English translation, as well as the absence of robust models capable of addressing the unique linguistic challenges posed by Singlish. In this work, we curate a word-level language detection dataset, a Singlish-to-English translation dataset, and propose a Language Detection-driven Masked Language Modelling approach for translating Singlish into English. We evaluate the performance of existing models and the proposed approach on two Singlish-to-English translation datasets, including our proposed SEAT dataset. The results demonstrate that the proposed LD-MLMTrans approach outperforms the baseline model and exhibits high proficiency in Singlish-to-English translation.
This paper introduces three foundational contributions to Digital Ottoman Turkish Studies. It presents: (1) three masked language models (MLMs) trained on over 11 million words from 144 works spanning from the 15th to 20th century, (2) a state-of-the-art Named Entity Recognition (NER) model (F1 = 89.94%) trained on 9,960 manually annotated entities, and (3) a state-of-the-art Universal Dependency (UD) parsing model for Ottoman Turkish. This work differs from others by deploying IJMES-transliterated documents for training and evaluation in order to prevent loss of information due to the change of the script from Perso-Arabic to Latin. The paper further explores probabilistic manuscript reconstruction in preliminary experiments, showing that MLMs can recover unread sections in historical documents with 77.8% top-1 accuracy when a list of candidate words is provided. Followed by a discussion, the paper outlines the future directions as building century-aware MLMs and expanding the training data across genres to enhance model generalization.
SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models
Erik Božík | Marek Suppa
Erik Božík | Marek Suppa
Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,000 speaker-segmented transcripts (66 million words, 220 million tokens) spanning 2001–2024, with rich metadata including speaker names, roles, and session information. From this collection, we derive SloPalSpeech, a 2,806-hour aligned speech dataset with segments up to 30 seconds, constructed using a language-agnostic anchor-based alignment pipeline and optimized for Whisper-based ASR training. Fine-tuning Whisper on SloPalSpeech reduces Word Error Rate (WER) by up to 70%, with the fine-tuned small model (244M parameters) approaching base large-v3 (1.5B parameters) performance at 6× fewer parameters. We publicly release the SloPal text corpus, SloPalSpeech aligned audio, and four fine-tuned Whisper models at https://huggingface.co/collections/NaiveNeuron/slopal, providing the most comprehensive open Slovak parliamentary language resource to date.
SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
Dávid Števaňák | Marek Suppa
Dávid Števaňák | Marek Suppa
Keyphrase extraction for morphologically rich, low-resource languages remains understudied, largely due to the scarcity of suitable evaluation datasets. We address this gap for Slovak by constructing a dataset of 227,432 scientific abstracts with author-assigned keyphrases—scraped and systematically cleaned from the Slovak Central Register of Theses—representing a 25-fold increase over the largest prior Slovak resource and approaching the scale of established English benchmarks such as KP20K. Using this dataset, we benchmark three unsupervised baselines (YAKE, TextRank, KeyBERT with SlovakBERT embeddings) and evaluate KeyLLM, an LLM-based extraction method using GPT-3.5-turbo. Unsupervised baselines achieve at most 11.6% exact-match F1@6, with a large gap to partial matching (up to 51.5%), reflecting the difficulty of matching inflected surface forms to author-assigned keyphrases. KeyLLM narrows this exact–partial gap, producing keyphrases closer to the canonical forms assigned by authors, while manual evaluation on 100 documents (𝜅 = 0.61) confirms that KeyLLM captures relevant concepts that automated exact matching underestimates. Our analysis identifies morphological mismatch as the dominant failure mode for statistical methods—a finding relevant to other inflected languages. The dataset (https://huggingface.co/datasets/NaiveNeuron/SlovKE) and evaluation code (https://github.com/NaiveNeuron/SlovKE) are publicly available.
Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan
Chihiro Taguchi | Yukinori Takubo | David Chiang
Chihiro Taguchi | Yukinori Takubo | David Chiang
Language endangerment poses a major challenge to linguistic diversity worldwide, and technological advances have opened new avenues for documentation and revitalization. Among these, automatic speech recognition (ASR) has shown increasing potential to assist in the transcription of endangered language data. This study focuses on Ikema, a severely endangered Ryukyuan language spoken in Okinawa, Japan, with approximately 1,300 remaining speakers, most of whom are over 60 years old. We present an ongoing effort to develop an ASR system for Ikema based on field recordings. Specifically, we (1) construct a 6.33-hour speech corpus from field recordings, (2) train an ASR model that achieves a character error rate as low as 15%, and (3) evaluate the impact of ASR-assisted transcription on annotation efficiency. Our results demonstrate that ASR integration can substantially reduce transcription time and cognitive load, offering a practical pathway toward scalable, technology-supported documentation of endangered languages.
Adaptive Method for Self-Supervised Learning Models on Automatic Dialect Speech Recognition Based on Shared Knowledge of Japanese Dialects and Standard Japanese
Naoru Asakawa | Naoki Takahashi | Atsuhiko Kai | Seiichi Nakagawa
Naoru Asakawa | Naoki Takahashi | Atsuhiko Kai | Seiichi Nakagawa
Speech recognition for Japanese dialects is challenging, and recognition accuracy tends to be lower compared to standard Japanese. Previous research proposed a three-step learning method based on the self-supervised learning (SSL) model XLS-R as the base model, incorporating three multi-task learning tasks: SSL, ASR, and dialect identification (DID). While this achieved improved recognition performance for dialect speech, it faced the issue of degraded recognition performance for standard Japanese. This study proposes an adaptation method to construct a single speech recognition model, based on the prior model, that is suitable for both Japanese dialects and standard Japanese. We explored the use of diverse speech corpora, including ReazonSpeech based on TV broadcast audio and CEJC based on everyday conversational speech, in addition to the standard Japanese speech corpus CSJ and the dialect speech corpus COJADS used in prior research, aiming for knowledge sharing between dialects and standard Japanese. As a result, we confirmed improved recognition performance for both dialects and standard Japanese by including both in the final step of a three-step learning method. We also examined the impact of differences in corpus type and domain on recognition performance.
ATLAS: Article Tracking, Linking, and Analysis of Swedish Encyclopedias
Albin Andersson | Salam Jonasson | Fredrik Wastring | Pierre Nugues
Albin Andersson | Salam Jonasson | Fredrik Wastring | Pierre Nugues
The digitization of old encyclopedias represents an important step to improve access to historically structured knowledge. Often, however, this process does not go beyond an optical character recognition, leaving all the underlying structure unexploited. In addition, many encyclopedias had multiple editions reflecting the evolution of knowledge. The lack of structure in the raw text makes it difficult to track changes across these editions. In this work, we built a pipeline to restore the text structure, where we extract the headwords and identify entries; categorize the entities; match entries across editions; and link entries to a Wikidata item. We applied this pipeline to the four major editions of Nordisk familjebok, an authoritative Swedish encyclopedia published between 1876 and 1951. We could extract the headwords with an F1 score of 97.8% and we obtained an F1 score of 93.4% on the headword classification. On a small-scale evaluation, we reached a 93% precision on the cross-edition matching, 85% precision and 16.5% recall on the Wikidata linking. This shows that an automated approach to digitized historical knowledge is possible. This should facilitate the preservation of general knowledge and the understanding of knowledge transmission. The datasets and programs are available online.
Evaluating Embedding Models on Danish Historical Newspapers: A Corpus and Benchmark Resource
Alie Lassche | Pascale Feldkamp | Yuri Bizzoni | Katrine Baunvig | Kristoffer Nielbo | Johan Heinsen
Alie Lassche | Pascale Feldkamp | Yuri Bizzoni | Katrine Baunvig | Kristoffer Nielbo | Johan Heinsen
We present an enriched dataset of almost five million Danish historical newspaper articles from the late seventeenth to nineteenth century, augmented with semantic embeddings and an annotated subset, to enable semi-automated classification as well as thematic and linguistic exploration. Through three historical benchmark tasks that evaluate the performance of Danish and multilingual embedding models on this historical Danish corpus, we discuss how the choice for an embedding model depends on the type of task, and enrich our corpus with embeddings from the overall best performing model. As a showcase experiment, we look at the distribution of article categories in the three subgenres that can be observed in the corpus. This experiment highlights the corpus and article-level embeddings’ potential for further exploration and analysis of the Danish historical mediascape. The resource is freely available for research use and aims to foster reproducible, data-driven studies of language and culture in the Danish nineteenth century.
Leveraging Linguistic Similarity for Low-Resource Speech Transcription
Valentina Fedchenko | Eric Jordan
Valentina Fedchenko | Eric Jordan
This study investigates how large-scale, self-supervised acoustic models (like XLSR and MMS) represent linguistic similarity and whether this can optimize Automatic Speech Recognition (ASR) for low-resource and dialectally diverse languages. While these models excel at cross-lingual transfer learning, their internal representations of fine-grained dialectal variation remain opaque. We focus on Yiddish, a language with a complex dialect continuum, to test if a model’s internal acoustic similarity metric—Acoustic Token Distribution Similarity (ATDS)—predicts ASR performance. Our methodology involved fine-tuning models on Yiddish dialects and measuring ATDS between Yiddish and related languages. Results confirm that ATDS is a meaningful predictor: higher acoustic similarity in the model’s latent space correlates with lower character error rates (CER) after fine-tuning. This relationship is strongest in mid-to-upper layers of the MMS model and for in-domain data. Crucially, ATDS captures model-dependent acoustic similarity, which does not always align with genealogical linguistic relationships but remains a practical indicator of transfer learning potential. We conclude that ATDS is a valuable tool for selecting donor languages to develop more efficient, dialect-sensitive ASR systems for language documentation, even if its absolute values require careful interpretation against linguistic knowledge.
A Corpus of Persuasion Techniques in Slavic Languages
Jakub Piskorski | Dimitar Iliyanov Dimitrov | Marina Ernst | Jacek Haneczok | Michal Marcinczuk | Arkadiusz Modzelewski | Roman Yangarber
Jakub Piskorski | Dimitar Iliyanov Dimitrov | Marina Ernst | Jacek Haneczok | Michal Marcinczuk | Arkadiusz Modzelewski | Roman Yangarber
We present a new corpus of persuasion techniques for Slavic languages. The corpus contains documents from parliamentary debates in Bulgarian and Polish, and from social media in Russian, annotated with persuasion techniques at text-span and sentence level. The techniques come from a taxonomy of 25 fine-grained persuasion techniques, grouped under six broader categories of rhetorical persuasion strategies. The corpus contains approximately 7500 text spans annotated with persuasion techniques, from 222 documents that cover hotly debated topics at both international and national level. We describe the process of corpus creation, provide related statistics, elaborate on topic and persuasion technique correlations. We provide baseline models and benchmark results for detection and classification of persuasion techniques at the text-span level and sentence level, which use classic ML-based and generative AI-based models.
GePaDeSE: A New Resource for Clause-Level Aspect in German Parliamentary Debates
Julian Schlenker | Ines Rehbein | Lilly Brauner | Florian Ertz | Ines Reinig | Simone Paolo Ponzetto
Julian Schlenker | Ines Rehbein | Lilly Brauner | Florian Ertz | Ines Reinig | Simone Paolo Ponzetto
This paper presents GePaDeSE, a new resource with annotations of clause-level aspect in German parliamentary debates, also known as Situation Entity types. The new resource includes 250 political speeches from the German Bundestag, given by 192 speakers, with over 220,000 tokens. In the paper, we first describe the new corpus and the annotation process. Then we present experiments on automatically classifying clause-level aspect and present an in-depth analysis where we show the potential of Situation Entities for the analysis of political discourse.
FrameNet Semantic Role Classification by Analogy
Van Duy Ngo | Stergos Afantenos | Emiliano Lorini | Miguel Couceiro
Van Duy Ngo | Stergos Afantenos | Emiliano Lorini | Miguel Couceiro
In this paper, we adopt a relational view of analogies applied to Semantic Role Classification in FrameNet. We define analogies as formal relations over the Cartesian product of frame evoking lexical units and frame element pairs, which we use to construct a new dataset.Each element of this binary relation is labelled as a valid analogical instance if the frame elements share the same semantic role, or as invalid otherwise.This formulation allows us to transform Semantic Role Classification into binary classification and train a lightweight Artificial Neural Network (ANN) that exhibits rapid convergence with minimal parameters. Crucially, no Semantic Role information is introduced to the neural network during training. We recover semantic roles during inference by computing probability distributions over candidates of all semantic roles within a given frame through random sampling and analogical transfer. This approach allows us to surpass previous State of the Art results while maintaining computational efficiency and frugality.
CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
Masato Kikuchi | Masatsugu Ono | Toshioki Soga | Tetsu Tanabe | Tadachika Ozono
Masato Kikuchi | Masatsugu Ono | Toshioki Soga | Tetsu Tanabe | Tadachika Ozono
Although WordNet is a valuable resource because of its structured semantic networks and extensive vocabulary, its fine-grained sense distinctions can be challenging for second-language learners. To address this issue, we developed a version of WordNet annotated with the Common European Framework of Reference for Languages (CEFR), integrating its semantic networks with language-proficiency levels. We automated this process using a large language model to measure the semantic similarity between sense definitions in WordNet and entries in the English Vocabulary Profile Online. To validate our approach, we constructed a large-scale corpus containing both sense and CEFR-level information from the annotated WordNet and used it to develop contextual lexical classifiers. Our experiments demonstrate that models fine-tuned on this corpus perform comparably to those fine-tuned on gold-standard annotations. Furthermore, by combining this corpus with the gold-standard data, we developed a practical classifier that achieves a Macro-F1 score of 0.81. This result provides indirect evidence that the transferred labels are largely consistent with the gold-standard levels. The annotated WordNet, corpus, and classifiers are publicly available to help bridge the gap between natural language processing and language education, thereby facilitating more effective and efficient language learning.
Towards a Gold Standard for Adjectival Hypernymy: Enriching the Open English WordNet with a Hybrid Approach
Lorenzo Augello | John P. McCrae | Marco Passarotti
Lorenzo Augello | John P. McCrae | Marco Passarotti
Adjectival hypernymy is an underexplored lexical-semantic relation essential for Natural Language Processing (NLP) and hierarchical semantic organization of the lexicon. While hypernymy in nouns and verbs has been extensively modeled in resources such as WordNet, adjectives remain largely unstructured due to their gradability and context-dependence. We present a hybrid Large Language Model (LLM)-Human approach towards the creation of a gold-standard dataset for adjectival hypernymy. Our method integrates three LLMs with systematic human evaluation, guided by a specifically developed theoretical framework ensuring consistency and linguistically-based principles, compiling a resource of 3,836 validated adjective hyponym-hypernym pairs. Results demonstrate high precision for consensus predictions (87%), confirming the utility of cross-model agreement as a proxy for semantic validity. This method highlights how LLMs can complement human effort and expertise to support the construction of lexical resources. The resulting dataset aims to enrich the Open English WordNet (OEWN) with explicit adjectival hierarchies and serves as a benchmark for hypernymy detection and lexical entailment evaluation.
PREMOVE in LiLa: Integrating Latin Preverbed Motion Verbs with WordNet and VerbNet
Andrea Farina | Marco Passarotti | Francesco Mambrini | Matteo Pellegrini | Eleonora Litta | Giovanni Moretti
Andrea Farina | Marco Passarotti | Francesco Mambrini | Matteo Pellegrini | Eleonora Litta | Giovanni Moretti
PREMOVE is a diachronic dataset of Ancient Greek and Latin PREverbed MOtion VErbs, providing manually curated morphological, syntactic, and semantic annotations for almost three thousand verbal occurrences. This paper presents the integration of PREMOVE into the LiLa Knowledge Base of Latin, linking its semantic annotations to WordNet (WN) and VerbNet (VN). We describe the RDF conversion using OntoLex-Lemon and FrAC, enabling explicit modelling of token-level attestations and dataset-level provenance. The resulting linked resource achieves full FAIR compliance and supports complex SPARQL queries, allowing users to explore motion semantics across lexical, textual, and semantic layers. Example SPARQL queries demonstrate how researchers can retrieve attested forms for specific WN synsets or VN classes, supporting reproducible linguistic research and cross-resource exploration of motion semantics in ancient languages.
From Incidents to Framing: A Dutch and English Frame-semantic Corpus and Lexicon
Piek T.J.M. Vossen | Pia Sommerauer | Levi Remijnse
Piek T.J.M. Vossen | Pia Sommerauer | Levi Remijnse
This paper reports on the final results of the Dutch FrameNet project. The project followed a new approach to aggregate an event corpus starting from registered events in Wikidata and collecting text in different languages that refer to these events. The resulting corpus is not only referentially grounded, but it is also grouped by the type of event, e.g. mass shootings, elections, sports events. A subset of the texts has been annotated with FrameNets frames for all references to the registered events and participants. The result is a unique corpus with comparable texts across languages that make reference to the same and similar events. From the annotations, we derived Dutch and English FrameNet lexicons, as well as reference lexicons. These lexicons allow us to infer abstractions from the annotations that also reflect sociocultural differences in framing the same entities and events.
AI Safety Lost in Translation: Evaluating the Effectiveness of English-Italian Cross-Lingual LLM Safety Alignment
Alessio Wu | Martim Brandao
Alessio Wu | Martim Brandao
Large Language Models (LLMs) have been shown to be vulnerable to various issues of bias and safety, for which new safety alignment techniques have been proposed. In this paper, we investigate the degree to which such techniques improve safety in a non-English language, specifically in Italian, both when they have and don’t have access to safety training data in that language. We evaluate standard mitigation techniques and assess cross-lingual safety transfer by comparing English-only versus bilingual Supervised Fine-Tuning (SFT), on several open-source small LLMs: Qwen3, Llama3.2, and Gemma3. Results confirm a significant cross-lingual safety gap, with most models performing worse in Italian. We find that while prompt engineering is generally effective, the impact of SFT is highly inconsistent. English-only SFT occasionally failed to transfer safety improvements into Italian and even deteriorated the performance of some models. Furthermore, bilingual SFT repeatedly underperformed other mitigation methods. These findings demonstrate that safety alignment does not always generalize across languages and models, and standard mitigation strategies can lead to unpredictable effects. We thus highlight the critical necessity for language-specific evaluation and dedicated multilingual safety research to ensure AI is developed equitably and safely for a global audience.
Semantic Label Drift in Cross-Cultural Translation
Mohsinul Kabir | Tasnim Ahmed | Md Mezbaur Rahman | Polydoros Giannouris | Sophia Ananiadou
Mohsinul Kabir | Tasnim Ahmed | Md Mezbaur Rahman | Polydoros Giannouris | Sophia Ananiadou
Machine Translation (MT) is widely employed to address resource scarcity in low-resource languages by translating data from high-resource languages. While sentiment preservation in translation has long been studied, a critical but underexplored factor is the role of cultural alignment between source and target languages. In this paper, we hypothesize that semantic labels drift or are altered during MT due to cultural divergence. Through a series of experiments across culturally sensitive and neutral domains, we establish three key findings: (1) MT systems, including modern Large Language Models (LLMs), induce label drift during translation, particularly in culturally sensitive domains; (2) unlike earlier statistical MT tools, LLMs encode cultural knowledge, and leveraging this knowledge can amplify label drift; and (3) cultural similarity or dissimilarity between source and target languages is a crucial determinant of label preservation. Our findings highlight that neglecting cultural factors in MT not only undermines label fidelity but also risks misinterpretation and cultural conflict in downstream applications. We release our codebase to facilitate future research in cross-cultural translation: https://github.com/mohsinulkabir14/label_drift
Chain-of-Thought Reasoning Improves Context-Aware Translation with Large Language Models
Shabnam Ataee | Hugo Huart | Andrei Popescu-Belis
Shabnam Ataee | Hugo Huart | Andrei Popescu-Belis
This paper assesses the ability of large language models (LLMs) to translate texts that include inter-sentential dependencies. We use the English-French DiscEvalMT benchmark (Bawden et al., 2018) with pairs of sentences containing translation challenges for pronominal anaphora and lexical cohesion. We evaluate 12 LLMs from the DeepSeek-R1, GPT, Llama, Mistral and Phi families on two tasks: (1) distinguish a correct translation from a wrong but plausible one; and (2) generate a correct translation. We compare prompts that encourage chain-of-thought reasoning with those that do not. The best models take advantage of reasoning and reach about 90% accuracy on the first task and COMET scores of about 92% on the second task, with GPT-4, GPT-4o and Phi standing out. Moreover, we observe a “wise get wiser” effect: the improvements through reasoning are larger for models that already perform well without reasoning.
Adja-French Parallel Corpus: A New Resource for Machine Translation of a West African Under-Resourced Language
Josue Frejus Godeme | Rolando Coto-Solano
Josue Frejus Godeme | Rolando Coto-Solano
We present the first parallel text corpus for Adja machine translation, an under-resourced Gbe language spoken by approximately 1,000,000 people in Benin and Togo. The corpus contains 10,000 French-Adja sentence pairs, providing a foundation for machine translation research. We establish baseline results using fine-tuned NLLB and ByT5 models, achieving a chrF++ of 28 in the French→Adja direction, and up to a chrF++ of 34 in the Adja→French direction. This work represents the first public machine translation resource for Adja. It provides benchmarks for future studies on this under-resourced West African language. The dataset is available at https://huggingface.co/datasets/JosueG/french-adja-parallel-corpus.
Goldfish: Monolingual Language Models for 350 Languages
Tyler A. Chang | Catherine Arnett | Zhuowen Tu | Benjamin Bergen
Tyler A. Chang | Catherine Arnett | Zhuowen Tu | Benjamin Bergen
For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite state-of-the-art performance on reasoning tasks, we find that these models still struggle with basic grammatical text generation in many languages. First, large multilingual models perform worse than bigrams for many languages (e.g. 24% of languages in XGLM 4.5B; 43% in BLOOM 7.1B) using FLORES perplexity as an evaluation metric. Second, when we train small monolingual models with only 125M parameters on 1GB or less data for 350 languages, these small models outperform large multilingual models both in perplexity and on a massively multilingual grammaticality benchmark. To facilitate future work on low-resource language modeling, we release Goldfish, a suite of over 1,000 small monolingual language models trained comparably for 350 languages. These models represent the first publicly-available monolingual language models for 215 of the languages included.
Dynaword: From One-shot to Continuously Developed Datasets
Kenneth Enevoldsen | Kristian Nørgaard Jensen | Jan Kostkan | Balázs Szabó | Márton Kardos | Kirsten Vad | Johan Heinsen | Andrea Blasi Núñez | Gianluca Barmina | Jacob Nielsen | Rasmus Larsen | Rob van der Goot | Peter Vahlstrup | Per Møldrup Dalum | Desmond Elliott | Lukas Galke Poech | Peter Schneider-Kamp | Kristoffer Nielbo
Kenneth Enevoldsen | Kristian Nørgaard Jensen | Jan Kostkan | Balázs Szabó | Márton Kardos | Kirsten Vad | Johan Heinsen | Andrea Blasi Núñez | Gianluca Barmina | Jacob Nielsen | Rasmus Larsen | Rob van der Goot | Peter Vahlstrup | Per Møldrup Dalum | Desmond Elliott | Lukas Galke Poech | Peter Schneider-Kamp | Kristoffer Nielbo
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguously licensed sources restricting use, sharing, and derivative works; (2) static dataset releases that prevent community contributions and diminish longevity; and (3) quality assurance processes restricted to publishing teams rather than leveraging community expertise. To address these limitations, we introduce two contributions: the Dynaword approach and Danish Dynaword. The Dynaword approach is a framework for creating large-scale, open datasets that can be continuously updated through community collaboration. Danish Dynaword is a concrete implementation that validates this approach and demonstrates its potential. Danish Dynaword contains over five times as many tokens as comparable releases, is exclusively openly licensed, and has received multiple contributions across industry, the public sector and research institutions. The repository includes light-weight tests to ensure data formatting, quality, and documentation, establishing a sustainable framework for ongoing community contributions and dataset evolution.
From Bones to Rocks: A Systematic Evaluation of Specialized Definition Generation for Portuguese
Rafael Oleques Nunes | Dennis Giovani Balreira | Joel Luís Carbonera
Rafael Oleques Nunes | Dennis Giovani Balreira | Joel Luís Carbonera
This work presents a systematic evaluation of Large Language Models (LLMs) for generating specialized definitions in Portuguese, focusing on the medical and geological domains. We introduce a robust benchmark and employ a rigorous, statistically grounded evaluation framework, including 5-fold cross-validation and significance testing, to ensure the reliability and generalizability of our findings. Our comprehensive experiments with various open-source, decoder-only LLMs explore in-context learning (ICL) with diverse prompting strategies, ranging from zero-shot to few-shot and contextual information. The evaluated models include multilingual architectures and one model that underwent continued pretraining specifically for Portuguese, allowing us to assess the impact of language adaptation on definition generation quality. The results indicate that most evaluated models perform effectively in this task, with relatively small performance differences among the top models. Statistical analyses confirmed that these differences are not consistently significant, suggesting that several open LLMs, regardless of their size, multilingual capacity, or language specialization, offer comparable effectiveness for Portuguese definition generation. These findings provide valuable insights for selecting and adapting models for specialized NLP tasks in low-resource languages like Portuguese.
Bangla Key2Text: Text Generation from Keywords for a Low Resource Language
Tonmoy Talukder | G M Shahariar
Tonmoy Talukder | G M Shahariar
This paper introduces Bangla Key2Text, a large-scale dataset of 2.6 million Bangla keyword-text pairs designed for keyword-driven text generation in a low-resource language. The dataset is constructed using a BERT-based keyword extraction pipeline applied to millions of Bangla news texts, transforming raw articles into structured keyword-text pairs suitable for supervised learning. To establish baseline performance on this new benchmark, we fine-tune two sequence-to-sequence models, mT5 and BanglaT5, and evaluate them using multiple automatic metrics and human judgments. Experimental results show that task-specific fine-tuning substantially improves keyword-conditioned text generation in Bangla compared to zero-shot large language models. The dataset, trained models, and code are publicly released to support future research in Bangla natural language generation and keyword-to-text generation tasks.
Beyond Lemmas and Syntax: Comparing Human and LLM-Generated Scientific Abstracts
Sergei Bagdasarov | Diego Alves
Sergei Bagdasarov | Diego Alves
In this study, we compare human-written (HWT) and machine-generated (MGT) abstracts of scientific papers, going beyond traditional lexical and syntactic analyses. We use an extensive corpus of publications on computational linguistics submitted to the Association of Computational Linguistics from mid 1950s to 2022. First, we generate abstracts with three state-of-the-art models (GPT-4o, Llama 3.1 and Qwen 2.5), providing the models with full texts of papers, and subsequently we compare these abstracts to those written by humans. We study the overall information content of abstracts, operationalised as surprisal, and the distribution of information in abstracts quantified as local Uniform Information Density (UID), both metrics related to the processing effort. Subsequently, we perform an extrinsic evaluation through topic modelling and clustering applying the BERTopic model. Our results show significant differences both in surprisal and UID, suggesting that abstracts generated by Llama are less cognitively demanding and show a more uniform distribution of information. Our topic modelling experiments show greater divergence between humans and LLMs than between LLM pairs. At the same time, Llama abstracts seem to be more semantically similar to those written by humans, standing in line with previous findings suggesting such similarity on lexical and syntactic level.
Systematic Multi-Aspect Evaluation of Time Series-Based Report Generation: The Case of Financial Analysis from Stock Data
Elizabeth Fons | Elena Kochkina | Rachneet Kaur | Zhen Zeng | Berowne Hlavaty | Charese Smiley | Svitlana Vyetrenko | Manuela Veloso
Elizabeth Fons | Elena Kochkina | Rachneet Kaur | Zhen Zeng | Berowne Hlavaty | Charese Smiley | Svitlana Vyetrenko | Manuela Veloso
This paper explores the capability of large language models (LLMs) to generate coherent textual reports from time series data, using financial reports from stock data as the use case. We conduct a comprehensive multi-aspect evaluation across four model families, including linguistic quality, content source attribution, automated metrics, and expert human assessment. We evaluate models using four major stock indices and two synthetic time series to assess generalization. We assess reports based on single and multiple time series data, and experiment with plain text and multi-modal prompting. We examine temporal effects by analyzing report quality as data approaches model knowledge cutoffs and testing synthetic future intervals. Our evaluation shows that LLMs are capable of creating high-quality financial analyst reports, with larger models demonstrating superior performance, however even those require human oversight and have potential for temporal logic errors. Our findings reveal model-specific behavioral patterns that enable tailored generation pipelines and inform future research about model pitfalls in time series-to-text generation tasks.
Towards Reliable AI Fairness: Challenges in Steering Features within Bias-Implicated Neurons
Ismael Garrido-Munoz | Arturo Montejo-Raez | Fernando Martínez-Santiago
Ismael Garrido-Munoz | Arturo Montejo-Raez | Fernando Martínez-Santiago
LLMs perpetuate societal biases, such as gender stereotypes, reinforcing harmful norms and posing significant fairness risks in real-world applications. We investigate a fine-grained mitigation technique that moves beyond surface-level fixes. Our approach uses attribution graphs to identify and directly steer bias-implicated features within a Sparse Autoencoder’s (SAE) latent space. This method, known as feature steering, offers a theoretically precise, surgical intervention aimed at correcting bias at its neural source without costly retraining. We critically examine its practical reliability across various contexts. We find that steering effectiveness is highly sensitive to parameter tuning, often requiring unpredictable, context-specific adjustments. The intervention’s success exists in narrow “sweet spots,” outside of which performance can degrade catastrophically. This demonstrates that while direct intervention on learned features is a powerful analytical tool, significant challenges of brittleness and instability hinder its application as a consistent, broad-scale debiasing solution, necessitating research into more robust control mechanisms.
From Body to Mind: Analyzing Gender Representation in Spanish Generative Language Models
Ismael Garrido-Munoz | Fernando Martínez-Santiago | Arturo Montejo-Raez
Ismael Garrido-Munoz | Fernando Martínez-Santiago | Arturo Montejo-Raez
While Large Language Models (LLMs) demonstrate remarkable text generation capabilities, they also risk inheriting and perpetuating harmful societal biases present in their vast training data. This study presents a rigorous, large-scale analysis of gender bias in a diverse set of 20 publicly available Spanish generative LLMs, ranging from 760M to 11B parameters. Our methodology utilizes a comprehensive set of specifically designed sentence templates to elicit adjectival descriptions associated with men and women in neutral contexts. We then extract and manually classify these adjectives using the Supersenses lexicosemantic framework, focusing on four key domains: BODY, BEHAVIOR, FEELING, and MIND. Our research uncovers systematic patterns consistent with pervasive cultural stereotypes, echoing findings from earlier masked language models. Women are disproportionately described by physical and emotional attributes, whereas men are more frequently associated with behavioral and cognitive traits. Finally, we investigate the relationship between model size and the intensity of these observed gender biases, offering crucial insights into how scaling affects fairness and equity in non-English models.
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
Svetlana Churina | Kokil Jaidka
Svetlana Churina | Kokil Jaidka
Incivility on platforms such as Twitter (now X) and Reddit complicates the development of AI systems that can support productive, rhetorically sound political argumentation. We present experiments with GPT-3.5 Turbo fine-tuned on two contrasting datasets of political discourse: high-incivility Twitter replies to U.S. Congress and low-incivility posts from Reddit’s r/ChangeMyView. Our evaluation examines how data composition and prompting strategies affect the rhetorical framing and deliberative quality of model-generated arguments. Results show that Reddit-finetuned models generate safer but rhetorically rigid arguments, while cross-platform fine-tuning amplifies adversarial tone and toxicity. Prompt-based steering reduces overt toxicity (e.g., personal attacks) but cannot fully offset the influence of noisy training data. We introduce a rhetorical evaluation rubric—covering justification, reciprocity, alignment, and authority—and provide implementation guidelines for authoring, moderation, and deliberation-support systems.
EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
Valle Ruiz-Fernández | Mario Mina | Júlia Falcão | Luis Antonio Vasquez Reina | Anna Salles | Aitor Gonzalez-Agirre | Olatz Perez-de-Viñaspre
Valle Ruiz-Fernández | Mario Mina | Júlia Falcão | Luis Antonio Vasquez Reina | Anna Salles | Aitor Gonzalez-Agirre | Olatz Perez-de-Viñaspre
Previous literature has largely shown that Large Language Models (LLMs) perpetuate social biases learnt from their pre-training data. Given the notable lack of resources for social bias evaluation in languages other than English, and for social contexts outside of the United States, this paper introduces the Spanish and the Catalan Bias Benchmarks for Question Answering (EsBBQ and CaBBQ). Based on the original BBQ, these two parallel datasets are designed to assess social bias across 10 categories using a multiple-choice QA setting, now adapted to the Spanish and Catalan languages and to the social context of Spain. We report evaluation results on different LLMs, factoring in model family, size and variant. Our results show that models tend to fail to choose the correct answer in ambiguous scenarios, and that high QA accuracy often correlates with greater reliance on social biases.
ToxSyn-PT: A Synthetic Fine-Grained Dataset of Minority-Targeted Toxic Language in Portuguese
Iago Alves Brito | Julia Soares Dollis | Fernanda Bufon Farber | Diogo Fernandes | Arlindo R. Galvão Filho
Iago Alves Brito | Julia Soares Dollis | Fernanda Bufon Farber | Diogo Fernandes | Arlindo R. Galvão Filho
The development of robust hate speech detection systems remains limited by the lack of large-scale, fine-grained training data, especially for languages beyond English. Existing corpora typically rely on simplistic toxic and non-toxic labels, and the few that capture hate directed at specific minority groups lack the positive counterexamples required to distinguish genuine hate from mere discussion. In this work, we introduce ToxSyn-PT, the first Portuguese large-scale corpus explicitly designed for multi-label hate speech detection across nine protected minority groups, including the non-toxic counterexamples absent in all other public datasets. Generated via a controllable four-stage pipeline, ToxSyn contains discourse-type annotations to capture rhetorical strategies of toxic/non-toxic language, such as sarcasm, dehumanization, and cultural appreciation. Our experiments reveal a catastrophic, mutual generalization failure compared to existing datasets from social-media domains: models trained on social media struggle to generalize to minority-specific contexts, and vice-versa. This finding indicates they are distinct tasks and exposes summary metrics like Macro F1 can be unreliable indicators of true model behavior, as they completely mask model failure. We publicly release ToxSyn on HuggingFace to support reproducible research on synthetic data generation and benchmark progress in hate-speech detection for low- and mid-resource languages.
AnswerCarefully: Creating a Dataset for LLM Safety in Japanese
Hisami Suzuki | Satoru Katsumata | Takashi Kodama | Tetsuro Takahashi | Kouta Nakayama | Satoshi Sekine
Hisami Suzuki | Satoru Katsumata | Takashi Kodama | Tetsuro Takahashi | Kouta Nakayama | Satoshi Sekine
In this paper we present JLLMSafety, a dataset for promoting the safety of Japanese LLM outputs. The dataset consists of 1,800 pairs of questions and reference answers, where the questions require special attention in answering. It covers a wide range of risk categories established in prior English-language datasets, but the data samples are original in that they are manually curated to reflect the socio-cultural context of LLM usage in Japan. We show that using this dataset for instruction to fine-tune a Japanese LLM led to improved output safety without compromising the utility of general responses. We also report the results of a safety evaluation of 12 Japanese LLMs using this dataset as a benchmark. Finally, we discuss the significance of creating regionally specific datasets of LLM safety, and describe the meta tags we added to the dataset to facilitate the creation of similar datasets in different languages and regions. The dataset is made available publicly for the sole purpose of improving LLM safety without any other usage restrictions.
A Dutch Benchmark to Assess Social Bias in LLMs within a Hiring Decision Setting
Renate Burema | Anne Schuth | Christopher Spelt | Dong Nguyen
Renate Burema | Anne Schuth | Christopher Spelt | Dong Nguyen
In this paper, we present a Dutch benchmark to assess whether large language models (LLMs) exhibit social biases in hiring decisions, focusing on gender and country of origin. We experiment with two approaches: explicit descriptions of the applicants’ demographics and using first names as proxies. We evaluate both monolingual and multilingual LLMs and find that all tested models, gpt-4o-mini, claude-3.5-haiku, Geitje-7B-Ultra and EuroLLM-9B-Instruct, exhibit some degree of social bias in their decisions. Furthermore, all models tested are sensitive to the manner in which the prompts are written. We make our benchmark publicly available under an EUPL-1.2 license. The benchmark is available at https://github.com/MinBZK/llm-benchmark/tree/main/benchmarks/social-bias.
PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
Farhan Farsi | Shayan Bali | Fatemeh Valeh | Parsa Ghofrani | Alireza Pakniat | Seyedkian Kashfipour | Amir H. Payberah
Farhan Farsi | Shayan Bali | Fatemeh Valeh | Parsa Ghofrani | Alireza Pakniat | Seyedkian Kashfipour | Amir H. Payberah
With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detection in various languages, there remains a significant gap in resources addressing social biases within Persian cultural contexts. In this work, we introduce PBBQ, a comprehensive benchmark dataset designed to evaluate social biases in Persian LLMs. Our benchmark, which encompasses 16 cultural categories, was developed through anonymous questionnaires completed by 250 diverse individuals across multiple demographics, in close collaboration with social science experts to ensure its validity. The resulting PBBQ dataset contains over 37,000 carefully curated questions, providing a foundation for the evaluation and mitigation of bias in Persian language models. We benchmark several open-source LLMs, a closed-source model, and Persian-specific fine-tuned models on PBBQ. Our findings reveal that current LLMs exhibit significant social biases across Persian culture. Additionally, by comparing model outputs to human responses, we observe that LLMs often replicate human bias patterns, highlighting the complex interplay between learned representations and cultural stereotypes. Our PBBQ dataset is also publicly available for use in future work. Content warning: This paper contains unsafe content.
Contextualizing Toxicity: An Annotation Framework for Unveiling Pragmatics in Conversations of Online Discussion Forums
Yingxue Fu | Anais Ollagnier
Yingxue Fu | Anais Ollagnier
The role of context has attracted increasing attention in research on toxicity detection. Interpreting toxic language remains a complex and multifaceted challenge, shaped by numerous linguistic, contextual, and social factors. However, current approaches often define “context” narrowly, focusing primarily on surface lexical cues such as hate lexicons, profanity markers, or sentiment polarity. These features, while useful, are insufficient to capture the interactional dynamics, user behaviors, and intentionality that shape such phenomena. To address this gap, this paper introduces a novel and systematic annotation framework, grounded in Speech Act Theory (Austin, 1962), aimed at deciphering the illocutionary and perlocutionary dimensions of conversation, which are unexplored in existing studies. We apply this framework to a new dataset of complete Reddit conversation threads, sampled to include discussions that turn toxic (124 conversations, 1990 messages). We evaluate the performance of GPT models (GPT-3, GPT-4, and GPT-5) on this challenging annotation task, providing insights into how large language models capture pragmatic and contextual dimensions of online toxicity.
How Far Can Bias Go? Tracing Bias from Pre-Training Data to Alignment
Marion Thaler | Abdullatif Köksal | Alina Leidinger | Anna Anna Korhonen | Hinrich Schütze
Marion Thaler | Abdullatif Köksal | Alina Leidinger | Anna Anna Korhonen | Hinrich Schütze
As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much work has gone into measuring and mitigating biases, fewer studies have investigated their origins. Therefore, this study examines the propagation of representational gender-occupation bias from pre-training data to LLM generations. Using zero-shot prompting and token co-occurrence analyses, we explore how biases in the pre-training data influence model generations. Our findings reveal that representational biases present in the pre-training data are amplified in the model generations, regardless of hyperparameters and prompting type. By comparing gender representation in the pre-training data with real-world distributions, our research highlights discrepancies between the data and the model, underscoring the importance of further work in mitigating bias at the data level.
Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
Lance Calvin Lim Gamboa | Yue Feng | Mark Lee
Lance Calvin Lim Gamboa | Yue Feng | Mark Lee
With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an important benchmark format for evaluating stereotypical associations exhibited by generative models. We expand the linguistic scope of BBQ and construct FilBBQ through a four-phase development process consisting of template categorization, culturally aware translation, new template construction, and prompt generation. These processes resulted in a bias test composed of more than 10,000 prompts which assess whether models demonstrate sexist and homophobic prejudices relevant to the Philippine context. We then apply FilBBQ on models trained in Filipino but do so with a robust evaluation protocol that improves upon the reliability and accuracy of previous BBQ implementations. Specifically, we account for models’ response instability by obtaining prompt responses across multiple seeds and averaging the bias scores calculated from these distinctly seeded runs. Our results confirm both the variability of bias scores across different seeds and the presence of sexist and homophobic biases relating to emotion, domesticity, stereotyped queer interests, and polygamy. FilBBQ will be available via GitHub.
Uncovering Hidden Violent Tendencies in LLMs: A Demographic Analysis via Behavioral Vignettes
Quintin Myers | Yanjun Gao
Quintin Myers | Yanjun Gao
Large language models (LLMs) are increasingly proposed for detecting and responding to violent content online, yet their ability to reason about morally ambiguous, real-world scenarios remains underexamined. We present the first study to evaluate LLMs using a validated social science instrument designed to measure human response to everyday conflict, namely the Violent Behavior Vignette Questionnaire (VBVQ). To assess potential bias, we introduce persona-based prompting that varies race, age, and geographic identity within the United States. Six LLMs developed across different geopolitical and organizational contexts are evaluated under a unified zero-shot setting. Our study reveals two key findings: (1) LLMs’ surface-level text generation often diverges from their internal preference for violent responses; (2) their violent tendencies vary across demographics, frequently contradicting established findings in criminology, social science, and psychology.
Exploring Social Bias in Slovenia: The EEC-SL Dataset
Jaya Caporusso | Damar Hoogland | Boshko Koloski | Matthew Purver | Senja Pollak | Spela Vintar
Jaya Caporusso | Damar Hoogland | Boshko Koloski | Matthew Purver | Senja Pollak | Spela Vintar
We introduce the EEC-SL dataset, an adaptation of the Equity Evaluation Corpus from English to Slovenian. Based on 11 sentence templates, the dataset contains 8,640 sentences, including pairs of minimally-distant sentences, varying with regard to one of two variables: gender (female or male), and ethnicity (Slovenian or not-Slovenian). In order to validate our selection of personal names, we create a localised version of the Implicit Association Test for ethnic bias, in which participants show a significant implicit bias favouring Slovenian over non-Slovenian names. We use the dataset to evaluate social bias in three computational language models (large language models and an encoder-only transformer) to perform sentiment analysis—specifically, valence. We analyse the results in terms of differences in sentiment between minimally-distant groups of sentences and inferential tests. We found limited evidence for social bias with regard to ethnicity, and no evidence for gender bias, in any of the employed models.
The MISOMEM-Val Dataset for Identifying Human Values in Misogynistic Memes
Rakshitha Rao Ailneni | Sanda Harabagiu
Rakshitha Rao Ailneni | Sanda Harabagiu
We present MISOMEM-Val, the first dataset that systematically annotates human values across Frames of Misogyny (FoMs) derived from misogynistic memes. Extending the Taxonomy of Misogyny, each frame is linked to the Human Value Hierarchy (HVH) with annotated support and ignore stances and accompanying rationales. In total, 1089 frames were annotated, comprising 3,051 support and 7,007 ignore value instances. We introduce Hierarchical Value Discovery with Human Feedback (HVD-HF), an LLM-assisted annotation framework combining Chain-of-Thought prompting and self-consistency verification to ensure transparency and quality. The annotation analysis reveals systematic asymmetries—Conservation and Self-Enhancement are frequently supported, while Self-Transcendence is often ignored, thus highlighting how misogynistic memes distort core human values.
ConGA: Guidelines for Contextual Gender Annotation. a Framework for Annotating Gender in Machine Translation
Argentina Anna Rescigno | Eva Vanmassenhove | Johanna Monti
Argentina Anna Rescigno | Eva Vanmassenhove | Johanna Monti
Handling gender across languages remains a persistent challenge for Machine Translation (MT) and Large Language Models (LLMs), especially when translating from gender-neutral languages into morphologically gendered ones, such as English to Italian. English largely omits grammatical gender, while Italian requires explicit agreement across multiple grammatical categories. This asymmetry often leads MT systems to default to masculine forms, reinforcing bias and reducing translation accuracy. To address this issue, we present the Contextual Gender Annotation (ConGA) framework, a linguistically grounded set of guidelines for word-level gender annotation. The scheme distinguishes between semantic gender in English through three tags, Masculine (M), Feminine (F), and Ambiguous (A), and grammatical gender realisation in Italian (Masculine (M), Feminine (F)), combined with entity-level identifiers for cross-sentence tracking. We apply ConGA to the gENder-IT dataset, creating a gold-standard resource for evaluating gender bias in translation. Our results reveal systematic masculine overuse and inconsistent feminine realisation, highlighting persistent limitations of current MT systems. By combining fine-grained linguistic annotation with quantitative evaluation, this work offers both a methodology and a benchmark for building more gender-aware and multilingual NLP systems.
University Speaking for Everyone: Assessing Changes in Italian Higher Education Statutes toward Gender-Inclusive Language
Sebastiano Vecellio Salto | Camilla Casula | Alessio Palmero Aprosio | Sara Tonelli
Sebastiano Vecellio Salto | Camilla Casula | Alessio Palmero Aprosio | Sara Tonelli
We examine the editorial evolution of Italian university statutes toward inclusive language, analyzing how institutions represent female and non-binary identities and how these representations affect administrative communication. To this end, we compile and annotate a corpus of university statutes, tracing the changes that have led some universities to move from the use of the generic masculine to more inclusive formulations. We also experiment with tools for the automatic detection of non-inclusive language in institutional communication and methods for the automatic rewriting of texts into inclusive language.
Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
Kaveh Eskandari Miandoab | Mahammed Kamruzzaman | Arshia Gharooni | Gene Louis Kim | Vasanth Sarathy | Ninareh Mehrabi
Kaveh Eskandari Miandoab | Mahammed Kamruzzaman | Arshia Gharooni | Gene Louis Kim | Vasanth Sarathy | Ninareh Mehrabi
Large Language Models have been shown to demonstrate stereotypical biases in their representations and behavior due to the discriminative nature of the data that they have been trained on. Despite significant progress in the development of methods and models that refrain from using stereotypical information in their decision-making, recent work has shown that approaches used for bias alignment are brittle. In this work, we introduce a novel and general augmentation framework that involves three plug-and-play steps and is applicable to a number of fairness evaluation benchmarks. Through application of augmentation to a fairness evaluation dataset (Bias Benchmark for Question Answering (BBQ)), we find that Large Language Models (LLMs), including state-of-the-art open and closed weight models, are susceptible to perturbations to their inputs, showcasing a higher likelihood to behave stereotypically. Furthermore, we find that such models are more likely to have biased behavior in cases where the target demographic belongs to a community less studied by the literature, underlining the need to expand the fairness and safety research to include more diverse communities.
TryggLLM: A Benchmark for Evaluating LLM Safety in Norwegian
Samia Touileb | Truls Pedersen | Isabell Stinessen Haugen
Samia Touileb | Truls Pedersen | Isabell Stinessen Haugen
We introduce TryggLLM, the first safety benchmark dataset for Norwegian. The dataset is intended for benchmarking different types of safety issues that can occur when using Norwegian generative language models. We have manually translated two English benchmark datasets, while modifying the content to be aligned with the Norwegian context. The benchmark dataset is composed of two sub-parts: i) prompts annotated by four native speakers, in both the written variants of Norwegian Bokmål (BM) and Nynorsk (NN), such that each native speaker wrote in their preferred variants (two BM and two NN); ii) prompts and target responses, where each of them has a BM and a NN version. We provide detailed descriptions of the data creation process. We also present a thorough manual evaluation of benchmarking existing open Norwegian LLMs using TryggLLM. Our results show that between 18% and 48% of the generated responses are unsafe, across all tested models.
We introduce the KOrean COntext-dependent Hate speech dataset (KOCOH) to evaluate large language models’ ability to detect context-dependent hate speech in Korean. KOCOH consists of 3,000 context-comment pairs collected from Korean online communities (Dcinside, FMkorea) with detailed annotations, including labels for hate speech and hate target groups. We assess the context-dependent hate speech detection capabilities of both humans and 11 state-of-the-art large language models, including GPT-5, Claude Sonnet 4, and Gemini 2.5 Flash. Our results show that humans outperform language models, with GPT-5 achieving the highest performance among the evaluated models. While humans demonstrate balanced recall and specificity, language models generally show significantly higher specificity compared to recall. The performance of both humans and models is affected by factors such as Honam-related vocabulary and sentiment polarity. This study contributes resources to Korean hate speech research and empirically demonstrates the performance gap between humans and language models. Through both quantitative and qualitative analyses, we explore the similarities and differences between humans and language models, offering insights for future developments in language models and AI ethics research. KOCOH is available at https://github.com/eparkatgithub/KOCOH.
Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems
Maliha Jahan | Thomas Thebaud | Zsuzsanna Fagyal | Jesus Villalba | Mark Hasegawa-Johnson | Laureano Moro Velazquez | Najim Dehak
Maliha Jahan | Thomas Thebaud | Zsuzsanna Fagyal | Jesus Villalba | Mark Hasegawa-Johnson | Laureano Moro Velazquez | Najim Dehak
Demographic bias in the performance of speech and language technology has been an active area of recent research. A lot of studies have shown the existence of demographic biases in Automatic Speech Recognition (ASR) systems. In this work, we propose a novel model-agnostic and demographic label-agnostic approach, called DARe, to mitigate any existing bias in an ASR system towards certain speaker groups. We built a debiasing module that goes between the feature extractor of an ASR and the rest of that ASR. The module includes content-group disentanglers to separate content and group, a demographic classifier, and adversarial reweighting. To eliminate the need for demographic labels, we generated pseudo-group labels by extracting speaker embeddings and clustering them. We worked with three ASR systems–Wav2Vec2 base, SEW tiny, and Whisper small. We used the FAI dataset, which contains naturalistic conversations with speakers who self-identify their demographic attributes. We used Word Error Rate (WER) as a metric of ASR performance and a Poisson regression-based approach to evaluate the racial fairness of the models. We compared the racial bias of the models before and after applying our proposed approach and observed a significant improvement in fairness.
RuBIN: A Russian Benchmark for Evaluating LLMs with Cultural Insights
Polina Lazukova | Irina Piontkovskaya
Polina Lazukova | Irina Piontkovskaya
Understanding culture-specific knowledge is essential for developing language models that perform reliably across diverse social and linguistic settings. This work explores both methodological and practical aspects of evaluating culture-specific knowledge in large language models. Special attention is given to the multiple-choice question answering format as a tool for identifying and measuring such knowledge. An analysis of existing benchmarks reveals several limitations, including insufficient cultural sensitivity and the presence of uninformative distractor options. In response, the RuBIN benchmark is introduced – a dataset consisting of questions based on phrases that are widely known in Russian culture. The paper describes the process of selecting and filtering culturally relevant topics, generating plausible incorrect answers using LLMs, and annotating and testing the benchmark for cross-linguistic robustness. RuBIN helps identify current LLMs’ weaknesses in transferring cultural knowledge and can serve as a tool for further adapting these models to diverse linguistic and cultural contexts.
This paper compares phonetically weighted and unweighted string distance measures in dialectometry, examining how explicit phonetic modeling affects the quantitative representation of linguistic similarity. Using narrow IPA transcriptions from the German REDE corpus, we evaluate nine measures–Levenshtein distance, bigram and trigram overlap, cosine distance, Jaro-Winkler, Jaccard similarity, the Herrgen-Schmidt measure, and the Relative Identity Value–through correlational analysis, distributional comparison, stabilization testing, and multidimensional scaling. The phonetically weighted Herrgen-Schmidt measure consistently achieves the most balanced distance dispersion, earliest stabilization, and highest linguistic plausibility. Unweighted edit-based measures reproduce the same topological structure in compressed form; distributional and overlap-based metrics introduce systematic scale distortions through exaggeration or compression. These findings establish explicit phonetic weighting as a principled and analytically efficient extension of standard dialectometric procedures. Explicit phonetic weighting enhances resolution and interpretive precision without altering the underlying relational geometry of dialect classifications.
Piecing Together Cross-Document Coreference Resolution Datasets: Systematic Dataset Analysis and Unification
Anastasia Zhukova | Terry Lima Ruas | Jan Philip Wahle | Bela Gipp
Anastasia Zhukova | Terry Lima Ruas | Jan Philip Wahle | Bela Gipp
Work in Natural Language Understanding increasingly relies on the ability to identify and track entities and events across large, heterogeneous text collections. This task, known as cross-document coreference resolution (CDCR), has a wide range of downstream applications, including multi-document summarization, information retrieval, and knowledge base population. Research in this area remains fragmented due to heterogeneous dataset formats, varying annotation standards, and the predominance of the CDCR definition as the event coreference resolution (ECR). To address these challenges, we introduce uCDCR, a unified dataset that consolidates diverse publicly available English CDCR corpora across various domains into a consistent format, which we analyze with standardized metrics and evaluation protocols. uCDCR incorporates both entity and event coreference, corrects known inconsistencies, and enriches datasets with missing attributes to facilitate reproducible research. We establish a cohesive framework for fair, interpretable, and cross-dataset analysis in CDCR and compare the datasets on their lexical properties, e.g., lexical composition of the annotated mentions, lexical diversity and ambiguity metrics, discuss the annotation rules and principles that lead to high lexical diversity, and examine how these metrics influence performance on the same-head-lemma baseline. Our dataset analysis shows that ECB+, the state-of-the-art benchmark for CDCR, has one of the lowest lexical diversities, and its CDCR complexity, measured by the same-head-lemma baseline, lies in the middle among all uCDCR datasets. Moreover, comparing document and mention distributions between ECB+ and uCDCR shows that using all uCDCR datasets for model training and evaluation will improve the generalizability of CDCR models. Finally, the almost identical performance on the same-head-lemma baseline, separately applied to events and entities, shows that resolving both types is a complex task and should not be steered toward ECR alone. The uCDCR dataset is available at https://huggingface.co/datasets/AnZhu/uCDCR, and the code for parsing, analyzing, and scoring the dataset is available at https://github.com/anastasia-zhukova/uCDCR.
With the rise of generative language models, machine-generated text detection has become a critical challenge. A wide variety of models is available, but inconsistent datasets, evaluation metrics, and assessment strategies obscure comparisons of model effectiveness. To address this, we evaluate 15 different detection models from six distinct systems, as well as seven trained models, across seven English-language textual test sets and three creative human-written datasets. We provide an empirical analysis of model performance, the influence of training and evaluation data, and the impact of key metrics. We find that no single system excels in all areas and nearly all are effective for certain tasks, and the representation of model performance is critically linked to dataset and metric choices. We find high variance in model ranks based on datasets and metrics, and overall poor performance on novel human-written texts in high-risk domains. Across datasets and metrics, we find that methodological choices that are often assumed or overlooked are essential for clearly and accurately reflecting model performance.
JAPAS: A Benchmark and Neural Approach for Japanese Patent Support Relation Extraction
Katsuki Chousa | Ryosuke Sugiura
Katsuki Chousa | Ryosuke Sugiura
Efficient analysis of patent literature is crucial for technological development and protecting intellectual property. A key task is verifying the “support requirement,” which mandates that the detailed description must fully describe the claimed invention. This requirement is fundamental to a patent’s validity. Manual verification is a labor-intensive process that demands technical and legal expertise, making automation highly desirable. However, research on this task has been hampered by two key challenges: (1) the absence of a public benchmark, and (2) the reliance of prior work on lexical matching, which fails to capture semantic equivalence. To address these issues, we introduce JAPAS, the first public benchmark for this task, comprising over 2,000 instances manually annotated for Japanese patents. Each instance is labeled with a claim span, a supporting description paragraph, a relation type, and the annotator’s confidence level. Using this benchmark, we also establish modern baselines that capture semantic similarity, such as embeddings and LLMs. Our experiments show that a fine-tuned Qwen3-14B model achieves an F1 score of 0.50, outperforming the conventional lexical-based baseline. This result, which demonstrates that the task is feasible yet challenging, highlights the utility of JAPAS as a research foundation and provides a performance target for future work.
A Teacher-Student Approach to Creating Verified Synthetic Clarification and Correction Dialogues for TableQA Tasks
Christian Poelitz | Nick McKenna
Christian Poelitz | Nick McKenna
Real dialogues with AI assistants for solving table questions-answering tasks often follow dynamic, unpredictable paths due to imperfect information provided by the user or in the data, which must be caught and handled. Developing datasets which capture such user-AI interactions is difficult and time-consuming. In this work, we develop a novel framework for synthetically generating controlled, multi-turn conversations between a user and AI assistant for the task of table-based question answering (TableQA), which can be generated from an existing dataset with fully specified TableQA examples for any target domain. Each conversation aims to solve a table-based reasoning question through collaborative effort, modeling one of two real-world scenarios: (1) an AI-initiated clarification, or (2) a user-initiated correction. Critically, we employ a strong teacher LLM to verify our synthetic conversations by functional correctness, ensuring high quality. Finally, we demonstrate synthetic datasets generated from TableQA tasks as benchmarks of frontier LLMs. We find that even larger models struggle to effectively issue clarification questions and accurately integrate user feedback for corrections, demonstrating important areas for future research.
Persona-Aware Evaluation of Cognitive Bias in LLMs: From Benchmark to Applied Decision-Making
Katsumasa Yoshikawa | Junya Takayama | Takato Yamazaki
Katsumasa Yoshikawa | Junya Takayama | Takato Yamazaki
We present a persona-aware evaluation suite that couples a 12-category cognitive-bias benchmark with 100 applied financial framing tasks to assess how large language models (LLMs) respond under systematically varied persona conditions. Using a factorized set of 162 personas spanning gender, age, political orientation, income, and education, we analyze how persona conditioning modulates bias-consistent responding across ten instruction-tuned models. On applied tasks, persona conditioning reduces framing reversals on average and slightly increases decision confidence, with substantial variation across model families and scales. Correlation analyses further reveal that benchmark bias tendencies—particularly availability, social proof, and framing—predict applied framing sensitivity, suggesting that standardized bias scores can serve as indicators of real-world decision variability. This work provides a unified framework for linking cognitive-bias evaluation with persona-conditioned decision behavior in LLMs. (All data and prompts will be released after acceptance to preserve anonymity.)
ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering
Daeyong Kwon | SeungHeon Doh | Juhan Nam
Daeyong Kwon | SeungHeon Doh | Juhan Nam
Recent advances in Large Language Models (LLMs) have transformed open-domain question answering, yet their effectiveness in music-related reasoning remains limited due to sparse music knowledge in pretraining data. While music information retrieval and computational musicology have explored structured and multimodal understanding, few resources support factual and contextual music question answering (MQA) grounded in artist metadata or historical context. We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic. These resources enable systematic evaluation of retrieval augmented generation (RAG) for MQA. Experiments show that RAG markedly improves factual accuracy—open-source models gain up to +56.8 percentage points (pp; Qwen3 8B: 35.0→91.8), approaching proprietary performance. RAG-style fine-tuning further boosts both factual recall and contextual reasoning, yielding strong improvements on both in-domain and out-of-domain benchmarks. MusWikiDB also yields +6 pp higher accuracy and 67% faster retrieval than the general Wikipedia corpus. We release MusWikiDB and ArtistMus to advance research in music information retrieval and domain-specific QA, establishing a foundation for retrieval augmented reasoning in culturally rich domains such as music.
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
Chalamalasetti Kranti | Sowmya Vajjala
Chalamalasetti Kranti | Sowmya Vajjala
In this paper, we introduce MATA, a novel evaluation dataset to assess the ability of Large Language Models (LLMs) in Telugu language, comprising 729 carefully curated multiple-choice and open-ended questions that span diverse linguistic dimensions. We evaluate 11 open-weight and closed-source LLMs on our dataset and present a fine-grained analysis of their performance. Further, we empirically show how LLMs rely on superficial heuristics such as answer position and distractor patterns for multiple-choice questions. Finally, we also compare LLM-as-a-judge evaluation with human evaluation for open-ended questions assess its reliability in a low-resource language. We argue that such fine-grained evaluation is essential for understanding model limitations and can inform the development of more linguistically capable LLMs, while also serving as a foundation for future research in Telugu NLP. Our dataset is available at:https://huggingface.co/datasets/TeluguLLMResearch/MATA
The availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be conducted. We introduce a new benchmark for evaluating LLMs in Estonian, based on seven diverse datasets. These datasets assess general and domain-specific knowledge, understanding of Estonian grammar and vocabulary, summarization abilities, contextual comprehension, and more. The datasets are all generated from native Estonian sources without using machine translation. We compare the performance of base models, instruction-tuned open-source models, and commercial models. Our evaluation includes 6 base models and 26 instruction-tuned models. To assess the results, we employ both human evaluation and LLM-as-a-judge methods. Human evaluation scores showed moderate to high correlation with benchmark evaluations, depending on the dataset. Claude 3.7 Sonnet, used as an LLM judge, demonstrated strong alignment with human ratings, indicating that top-performing LLMs can effectively support the evaluation of Estonian-language models.
Indirect Question Answering in English, German and Bavarian: A Challenging Task for High- and Low-Resource Languages Alike
Miriam Winkler | Verena Blaschke | Barbara Plank
Miriam Winkler | Verena Blaschke | Barbara Plank
Indirectness is a common feature of daily communication, yet is underexplored in NLP research for both low-resource as well as high-resource languages. Indirect Question Answering (IQA) aims at classifying the polarity of indirect answers. In this paper, we present two multilingual corpora for IQA of varying quality that both cover English, Standard German and Bavarian, a German dialect without standard orthography: InQA+, a small high-quality evaluation dataset with hand-annotated labels, and GenIQA, a larger training dataset, that contains artificial data generated by GPT-4o-mini. We find that IQA is a pragmatically hard task that comes with various challenges, based on several experiment variations with multilingual transformer models (mBERT, XLM-R and mDeBERTa). We suggest and employ recommendations to tackle these challenges. Our results reveal low performance, even for English, and severe overfitting. We analyse various factors that influence these results, including label ambiguity, label set and dataset size. We find that the IQA performance is poor in high- (English, German) and low-resource languages (Bavarian) and that it is beneficial to have a large amount of training data. Further, GPT-4o-mini does not possess enough pragmatic understanding to generate high-quality IQA data in any of our tested languages.
Benchmarking Large Language Models for Chinese and Japanese IMEs: Phonetic-to-Character Generation and Textual Error Correction
Yuchun Zou | Tedd Lee | Xiaodi Fan | Jun Li
Yuchun Zou | Tedd Lee | Xiaodi Fan | Jun Li
Efficient text entry for complex writing systems like Chinese and Japanese necessitates the use of Input Method Editors (IMEs). While Large Language Models (LLMs) are emerging as powerful, context-aware language resources for this task, we present a comprehensive benchmark and evaluation methodology to assess the viability of LLMs for next-generation IMEs. We conduct a comparative analysis of a diverse set of LLMs against established baseline methods on two core tasks: phonetic-to-character generation (using Pinyin and Romaji) and textual error correction. Our experiments demonstrate that top-tier LLMs achieve superior accuracy by leveraging deep contextual understanding, significantly outperforming traditional systems in ambiguity resolution and the correction of complex errors. However, our analysis also reveals a crucial trade-off between accuracy and computational efficiency across different models. The datasets, evaluation scripts, and results from this study serve as a vital public resource for future research, providing a robust baseline for developing and selecting models that balance performance with the low-latency demands of real-world text input.
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
Gianluca Barmina | Nathalie Carmen Hau Norman | Peter Schneider-Kamp | Lukas Galke Poech
Gianluca Barmina | Nathalie Carmen Hau Norman | Peter Schneider-Kamp | Lukas Galke Poech
We present an enhanced benchmark for evaluating linguistic acceptability in Danish. We first analyze the most common errors found in written Danish. Based on this analysis, we introduce a set of fourteen corruption functions that generate incorrect sentences by systematically introducing errors into existing correct Danish sentences. To ensure the accuracy of these corruptions, we assess their validity using both manual and automatic methods. The results are then used as a benchmark for evaluating Large Language Models on a linguistic acceptability judgement task. Our findings demonstrate that this extension is both broader and more comprehensive than the current state of the art. By incorporating a greater variety of corruption types, our benchmark provides a more rigorous assessment of linguistic acceptability, increasing task difficulty, as evidenced by the lower performance of LLMs on our benchmark compared to existing ones. Our results also suggest that our benchmark has a higher discriminatory power which allows to better distinguish well-performing models from low-performing ones.
KCIF: Knowledge-Conditioned Instruction Following
Rudra Murthy | Praveen Venkateswaran | Prince Kumar | Danish Contractor
Rudra Murthy | Praveen Venkateswaran | Prince Kumar | Danish Contractor
LLM evaluation benchmarks have traditionally separated the testing of knowledge/reasoning capabilities from instruction following. In this work, we study the interaction between knowledge and instruction following, and observe that LLMs struggle to follow simple answer modifying instructions, and are also distracted by instructions that should have no bearing on the original knowledge task answer. We leverage existing multiple-choice answer based knowledge benchmarks and apply a set of simple instructions which include manipulating text (eg.: change case), numeric quantities (eg.: increase value, change formatting), operate on lists (eg.: sort answer candidates) and distractor instructions (eg.: change case of numeric answers). We evaluate models at varying parameter sizes (1B-405B) from different model families and find that, surprisingly, all models report a significant drop in performance on such simple task compositions. While large-sized and frontier models report performance drops of 40-50%, in small and medium sized models the drop is severe (sometimes exceeding 80%). Our results highlight a limitation in the traditional separation of knowledge/reasoning and instruction following, and suggest that joint-study of these capabilities are important. We release our benchmark dataset, evaluation framework code, and results for future work.
GAIN: A Benchmark for Goal-Aligned Decision-Making of Large Language Models under Imperfect Norms
Masayuki Kawarada | Kodai Watanabe | Soichiro Murakami
Masayuki Kawarada | Kodai Watanabe | Soichiro Murakami
We introduce GAIN(Goal-Aligned Decision-Making under Imperfect Norms), a benchmark designed to evaluate how large language models (LLMs) balance adherence to norms against business goals. Existing benchmarks typically focus on abstract scenarios rather than real-world business applications. Furthermore, they provide limited insights into the factors influencing LLM decision-making. This restricts their ability to measure models’ adaptability to complex, real-world norm-goal conflicts. In GAIN, models receive a goal, a specific situation, a norm, and additional contextual pressures. These pressure, explicitly designed to encourage potential norm deviations, are a unique feature that differentiates GAIN from other benchmarks, enabling a systematic evaluation of the factors influencing decision-making. We define five types of pressures: Goal Alignment, Risk Aversion, Emotional/Ethical Appeal, Social/Authoritative Influence, and Personal Incentive. The benchmark comprises 1,200 scenarios across four domains: hiring, customer support, advertising and finance. Our experiments show that advanced LLMs frequently mirror human decision-making patterns. However, when Personal Incentive pressure is present, they diverge significantly, showing a strong tendency to adhere to norms rather than deviate from them.
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
Paloma Piot | David Otero | Patricia Martin-Rodilla | Javier Parapar
Paloma Piot | David Otero | Patricia Martin-Rodilla | Javier Parapar
Hate speech spreads widely online and harms both individuals and communities, making automatic detection essential for large-scale moderation. However, accurately detecting hate speech remains a difficult task. Part of the challenge lies in subjectivity: what one person flags as hate speech, another may see as benign. Traditional annotation agreement metrics, such as Cohen’s k, oversimplify this disagreement, treating it as an error rather than meaningful diversity. Meanwhile, Large Language Models (LLMs) promise scalable annotation, but prior studies demonstrate that they cannot fully replace human judgement, especially in subjective tasks. In this work, we reexamine LLM reliability using a subjectivity-aware framework, cross-Replication Reliability (xRR), revealing that even under fairer lens, LLMs still diverge from humans. Yet this limitation opens an opportunity: we find that LLM-generated annotations can reliably reflect performance trends across classification models, correlating with human evaluations. We test this by examining whether LLM-generated annotations preserve the relative ordering of model performance derived from human evaluation (i.e. whether models ranked as more reliable by human annotators preserve the same order when evaluated with LLM-generated labels). Our results show that, although LLMs differ from humans at the instance level, they reproduce similar ranking and classification patterns, suggesting their potential as proxy evaluators. While not a substitute for human annotators, they might serve as a scalable proxy for evaluation in subjective NLP tasks.
PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
Mohammad Javad Ranjbar Kalahroodi | Amirhossein Sheikholselami | Sepehr Karimi Arpanahi | Sepideh Ranjbar Kalahroodi | Heshaam Faili | Azadeh Shakery
Mohammad Javad Ranjbar Kalahroodi | Amirhossein Sheikholselami | Sepehr Karimi Arpanahi | Sepideh Ranjbar Kalahroodi | Heshaam Faili | Azadeh Shakery
Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine, particularly in low-resource languages, remains underexplored. In this work, we introduce PersianMedQA, a large-scale dataset of 20,785 expert-validated multiple-choice Persian medical questions from 14 years of Iranian national medical exams, spanning 23 medical specialties and designed to evaluate LLMs in both Persian and English. We benchmark 41 state-of-the-art models, including general-purpose, Persian, and medical LLMs, in zero-shot and chain-of-thought (CoT) settings. Our results show that closed-weight general models (e.g., GPT-4.1) consistently outperform all other categories, achieving 83.09% accuracy in Persian and 80.7% in English, while Persian LLMs such as Dorna underperform significantly (e.g., 34.9% in Persian), often struggling with both instruction-following and domain reasoning. We also analyze the impact of translation, showing that while English performance is generally higher, 3-10% of questions can only be answered correctly in Persian due to cultural and clinical contextual cues that are lost in translation. Finally, we demonstrate that model size alone is insufficient for robust performance without strong domain or language adaptation. PersianMedQA provides a foundation for evaluating bilingual and culturally grounded medical reasoning in LLMs. The dataset, along with a bilingual medical dictionary, is publicly available at: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA.
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
Irina Proskurina | Marc-Antoine Carpentier | Julien Velcin
Irina Proskurina | Marc-Antoine Carpentier | Julien Velcin
Optimization of offensive content moderation models for different types of hateful messages is typically achieved through continued pre-training or fine-tuning on new hate speech benchmarks. However, existing benchmarks mainly address explicit hate toward protected groups and often overlook implicit or indirect hate, such as demeaning comparisons, calls for exclusion or violence, and subtle discriminatory language that still causes harm. While explicit hate can often be captured through surface features, implicit hate requires deeper, full-model semantic processing. In this work, we question the need for repeated fine-tuning and analyze the role of HatePrototypes, class-level vector representations derived from language models optimized for hate speech detection and safety moderation. We find that these prototypes, built from as few as 50 examples per class, enable cross-task transfer between explicit and implicit hate, with interchangeable prototypes across benchmarks. Moreover, we show that parameter-free early exiting with prototypes is effective for both hate types. We release the code, prototype resources, and evaluation scripts to support future research on efficient and transferable hate speech detection.
Investigating Memorization in Language Models Trained via Knowledge Distillation
Maarten Mäcking | Michaela Regneri
Maarten Mäcking | Michaela Regneri
We analyze how knowledge distillation influences memorization in language models. Although knowledge distillation is a widely used technique to train smaller, more efficient models, its effect on memorization is not well understood, despite the importance of memorization for model utility and privacy. We demonstrate that when the student and teacher models are trained on different datasets, knowledge distillation substantially reduces memorization and accelerates the forgetting of sequences previously memorized by the student. However, knowledge distillation does not eliminate privacy risks: it accelerates memorization when the student is trained on sequences memorized by the teacher, and teachers can leak memorized content even when the student is trained on data that does not contain these sequences. Finally, we find that the size of the teacher model leads to a trade-off between how quickly memorized information is transferred to the student and how much the student ultimately memorizes. Overall, we provide practical insights for balancing the utility of distilled models against the privacy concerns associated with memorization.
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
Hanwool Lee | Dasol Choi | Sooyong Kim | Ilgyun Jung | Sangwon Baek | Guijin Son | Inseong Hwang | Naeun Lee | Seunghyeok Hong
Hanwool Lee | Dasol Choi | Sooyong Kim | Ilgyun Jung | Sangwon Baek | Guijin Son | Inseong Hwang | Naeun Lee | Seunghyeok Hong
Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps across institutions. Overcoming these reproducibility gaps does not mean enforcing a one-size-fits-all evaluation. Rather, effective benchmarking requires diverse experimental approaches and a framework robust enough to support them. To this end, we introduce HRET (Haerae Evaluation Toolkit), an open-source, registry-based framework that unifies Korean LLM assessment. HRET integrates major Korean benchmarks, multiple inference backends, and multi-method evaluation, with language consistency enforcement to ensure genuine Korean outputs. Its modular registry design also enables rapid incorporation of new datasets, methods, and backends, ensuring the toolkit adapts to evolving research needs. Beyond standard accuracy metrics, HRET incorporates Korean-focused output analyses-morphology-aware Type-Token Ratio (TTR) for evaluating lexical diversity and systematic keyword-omission detection for identifying missing concepts-to provide diagnostic insights into language-specific behaviors. These targeted analyses help researchers pinpoint morphological and semantic shortcomings in model outputs, guiding focused improvements in Korean LLM development.
Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP
Poli Nemkova | Amrit Adhikari | Matthew Pearson | Vamsi Krishna Sadu | Albert V. Mark
Poli Nemkova | Amrit Adhikari | Matthew Pearson | Vamsi Krishna Sadu | Albert V. Mark
Humanitarian organizations face a critical choice: invest in costly commercial APIs or rely on free open-weight models for multilingual human rights monitoring. While commercial systems offer reliability, open-weight alternatives lack empirical validation - especially for low-resource languages common in conflict zones. This paper presents the first systematic comparison of commercial and open-weight large language models (LLMs) for human-rights-violation detection across seven languages, quantifying the cost-reliability trade-off facing resource-constrained organizations. Across 78,000 multilingual inferences, we evaluate six models - four instruction-aligned (Claude-Sonnet-4, DeepSeek-V3, Gemini-Flash-2.0, GPT-4.1-mini) and two open-weight (LLaMA-3-8B, Mistral-7B) - using both standard classification metrics and new measures of cross-lingual reliability: Calibration Deviation (CD), Decision Bias (ΔBias), Language Robustness Score (LRS), and Language Stability Score (LSS). Results show that alignment, not scale, determines stability: aligned models maintain near-invariant accuracy and balanced calibration across typologically distant and low-resource languages (e.g., Lingala, Burmese), while open-weight models exhibit significant prompt-language sensitivity and calibration drift. These findings demonstrate that multilingual alignment enables language-agnostic reasoning and provide practical guidance for humanitarian organizations balancing budget constraints with reliability in multilingual deployment.
Counting on Consensus: Selecting the Right Inter-Annotator Agreement Metric for NLP Annotation and Evaluation
Joseph H. F. James
Joseph H. F. James
Human annotation remains the foundation of reliable and interpretable data in Natural Language Processing (NLP). As annotation and evaluation tasks continue to expand, from categorical labelling to segmentation, subjective judgment, and continuous rating, measuring agreement between annotators has become increasingly more complex. This paper outlines how inter-annotator agreement (IAA) has been conceptualised and applied across NLP and related disciplines, describing the assumptions and limitations of common approaches. We organise agreement measures by task type and discuss how factors such as label imbalance and missing data influence reliability estimates. In addition, we highlight best practices for clear and transparent reporting, including the use of confidence intervals and the analysis of disagreement patterns. The paper aims to serve as a guide for selecting and interpreting agreement measures, promoting more consistent and reproducible human annotation and evaluation in NLP.
Quadratic Weighted Kappa Is Not Enough for Evaluating Automated Essay Scoring Models
Salam Albatarni | Tamer Elsayed
Salam Albatarni | Tamer Elsayed
Quadratic Weighted Kappa (QWK) has been the standard evaluation metric in Automated Essay Scoring (AES) research for over two decades. Despite repeated criticisms highlighting its limitations, the community has largely continued to rely on QWK without adopting alternative metrics. This study aims to encourage a shift toward more suitable evaluation practices by systematically examining QWK’s behavior under three key conditions: dataset size, class imbalance, and score range. Using both a publicly available AES dataset and carefully synthesized datasets, we demonstrate scenarios where QWK produces unstable or misleading results. Our findings highlight the need for more robust evaluation practices and point to alternative metrics, particularly variants of Gwet’s AC2, that offer greater reliability across a variety of conditions.
Evaluating the Homogeneity of Keyphrase Prediction Models
Mael Houbre | Florian Boudin | Beatrice Daille
Mael Houbre | Florian Boudin | Beatrice Daille
Keyphrases which are useful in several NLP and IR applications are either extracted from text or predicted by generative models. Contrarily to keyphrase extraction approaches, keyphrase generation models can predict keyphrases that do not appear in a document’s text called ‘absent keyphrases‘. This ability means that keyphrase generation models can associate a document to a notion that is not explicitly mentioned in its text. Intuitively, this suggests that for two documents treating the same subjects, a keyphrase generation model is more likely to be homogeneous in their indexing i.e. predict the same keyphrase for both documents, regardless of those keyphrases appearing in their respective text or not; something a keyphrase extraction model would fail to do. Yet, homogeneity of keyphrase prediction models is not covered by current benchmarks. In this work, we introduce a method to evaluate the homogeneity of keyphrase prediction models and study if absent keyphrase generation capabilities actually help the model to be more homogeneous. To our surprise, we show that keyphrase extraction methods are competitive with generative models, and that depending on the evaluation scenario, having the ability to generate absent keyphrases can actually act to the detriment of homogeneity. Our data, code and prompts are available on Huggingface and github.
A Taxonomy of Safety: Harmonizing LLM Benchmarks in a Fragmented Landscape
Shadi Rastegar | Viktor Hangya | Fabian Kuech | Darina Gold
Shadi Rastegar | Viktor Hangya | Fabian Kuech | Darina Gold
Understanding and mitigating the safety limitations of LLMs is of great importance to build trustworthy AI applications. Although a wide range of safety benchmarks are available, there is no standardized taxonomy of safety categories. As a result, some benchmarks focus on a specific subset of categories, they define test samples on different granularity levels, or they use different definitions or naming conventions. To mitigate these issues, we propose a two-level taxonomy of LLM safety categories, created by harmonizing existing resources. Our taxonomy gives an overview of important safety categories that helps researchers pinpoint potential safety risks and select the right benchmarks when evaluating or developing language models. Moreover, the taxonomy provides guidelines to categorize future benchmarks. Furthermore, since the majority of the available safety resources are English-focused, we check the cross-cultural validity of our taxonomy by translating datasets covering all top level categories to French, German, Italian, and Spanish. A manual review of a subset of translated samples by native speakers revealed no major cultural mismatches from a safety perspective. This supports not only the transferability of English benchmarks but also the transferability of the categories in our taxonomy, as well as its potential as a practical tool for guiding safety-focused dataset development and evaluation beyond English.
Consistency of LLMs to Comparative Statements in Mathematical Reasoning Tasks
Aidan W. San | Daniel Juyoung Son | Xiaodong Liu | Yangfeng Ji
Aidan W. San | Daniel Juyoung Son | Xiaodong Liu | Yangfeng Ji
Large language models (LLMs) have the potential to significantly expand access to quality education through applications such as mathematics tutoring. However, a key challenge is that student writing often contains redundancies, and prior research has shown that LLMs can be sensitive to such irrelevant information. This raises a critical research question: How consistent are LLMs when faced with extraneous comparative statements? To address this, we propose a systematic framework for evaluating LLM consistency. Our approach involves a hybrid strategy that integrates template-based and model-based methods to generate comparative statements (e.g., “One of the apples was tastier than average”) and insert them into mathematical reasoning problems. The merit of our approach lies in its systematic and automated nature, enabling rigorous assessment across various models and datasets. Conducting experiments on the GSM8K, AQuA, and Hendrycks MATH benchmarks with a suite of open-source LLMs, we highlight two key results. First, LLM accuracy can drop by over 30% when presented with these statements. Furthermore, we uncover a trade-off between the diversity of the generated statements and the magnitude of the performance drop, where less diverse and more repetitive perturbations lead to greater accuracy degradation.
PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian
Mohammad Hossein Shalchian | Mostafa Amiri | Amir Mahdi Sadeghzadeh
Mohammad Hossein Shalchian | Mostafa Amiri | Amir Mahdi Sadeghzadeh
We target practical anonymization of Persian customer chats by training a compact NER model from LLM-labeled supervision and selecting the best labeler for deployment. We compare three instruction-tuned LLMs—DEEPSEEKV3-0324, GPT-OSS-120B, and QWEN3-235B-A22B-INSTRUCT-2507—to produce span annotations under a shared JSON protocol, yielding four corpora (OSS_ZeroShot, Qwen_ZeroShot, Qwen_FewShot, DeepSeek_FewShot). A MATINAROBERTA-based token-classifier is trained per corpus and evaluated with token-level Precision/Recall/F1 (overall and per-class). We also report Label Coverage Recall (LCR), the proportion of gold non-O tokens predicted as non-O, and quantify cross-labeler behavior via a token-level Venn on test annotations. Finally, we contrast test-set annotation latency of the LLMs on H200 nodes with the trained NER’s test-time labeling on a single RTX 3090. Results show that supervision from OSS_ZeroShot yields the strongest macro-F1 and LCR, while the resulting NER labels an entire 40K-message test set in ∼2 minutes on one consumer GPU. This establishes a practical path to high-quality, low-cost anonymization for Persian industrial data.
How Many Samples Do We Need? A Toolkit for Power-Aware Evaluation Design
Angelo Basile | Areg Mikael Sarvazyan | José Ángel González
Angelo Basile | Areg Mikael Sarvazyan | José Ángel González
If datasets are the telescopes of our field, then statistical power is their resolution, i.e., their ability to reveal a true difference in model performance when one exists. Many NLP evaluations are underpowered, leading to overstated claims of improvement. This paper introduces sk-power, an open-source Python library that helps researchers and practitioners design well-powered evaluations. Built with familiar scikit-learn-style abstractions, sk-power enables users to simulate evaluation scenarios, estimate minimum detectable effects, and assess the reliability of reported gains. We also illustrate what can go wrong when power analysis isn’t carried out. Our goal is to position power analysis as a first-class, practical step in evaluation planning.
Of Words and Meaning: A Grammatical and Semantic Benchmark for Faroese LLM Understanding
Iben Nyholm Debess | Barbara Scalvini | Bolette Pedersen
Iben Nyholm Debess | Barbara Scalvini | Bolette Pedersen
Evaluating language technology for low-resource languages faces a fundamental challenge: the scarcity of native benchmarks suitable for systematic assessment. For Faroese, no such evaluation frameworks exist. We address this gap by presenting the first benchmark suite for Faroese semantic understanding and grammatical competence. Our methodology transforms existing lexicographic resources, authoritative dictionaries and error corpora, into systematic evaluation tasks through computational restructuring, demonstrating a replicable approach for resource-constrained settings. The resulting benchmarks assess grammatical correctness, semantic relation classification, and metaphor comprehension. Evaluation across LLMs from compact open-source to large-scale commercial systems reveals consistent performance patterns favouring proprietary models. This work establishes a proof of concept for benchmark creation from traditional linguistic resources, and provides a methodological template for other low-resource language communities.
TURING: Evaluating Human Abilities to Identify AI-Generated Texts
Natalia Kalashnikova | Nicolas De Bufala | Sophie Fayad | Laurent Cervoni
Natalia Kalashnikova | Nicolas De Bufala | Sophie Fayad | Laurent Cervoni
This study analyzes humans’ ability to identify AI-generated texts across 10 genres. We collected 9164 annotations from 214 participants on 500 texts (half human, half LLM-produced), and analyzed 7943 after quality screening. Our main findings are that the humans accuracy was above chance but far from perfect (around 59%), with a slight tendency to label texts as “Human-generated”. Their performance is influenced by the text genre (structural/factual formats easier to identify vs. complex genres) and by generating LLM. Annotators optionally selected three-level descriptors to justify decisions. While they had very limited effects on accuracy, their usage showed some association between text features (monotony, lack of cohesion or coherence) and “AI-generated” labeling. However, the linguistic features of the texts appear to have no robust impact after correction on human judgment. A small learning effect emerged but was practically negligible (0.1-0.2%), and personal characteristics of annotators had an impact on their accuracy, except age, which showed no effect. Finally, two automated detection tools were tested, reaching 88% accuracy on our distribution, clearly above humans, highlighting the value of human-tool combinations.
JamC-QA: A Multiple-Choice Question Answering Benchmark for Japan-Specific Knowledge
Teruaki Oka | Tomohide Shibata | Nao Yoshida
Teruaki Oka | Tomohide Shibata | Nao Yoshida
We introduce JamC-QA, a multiple-choice question answering benchmark specifically designed to evaluate Japan-specific knowledge. Existing Japanese QA benchmarks largely consist of questions translated from English or derived from professional exams, primarily targeting academic or generally shared knowledge. Consequently, this limits the usefulness of distinguishing the performance of high-performing Large Language Models on local knowledge acquisition. To address this, JamC-QA serves as a robust resource for assessing the acquisition of Japan-specific knowledge. It comprises 2,309 challenging instances that were created entirely from scratch by human annotators across eight categories: culture, custom, regional identity, geography, history, government, law, and healthcare. Instances that were easily answerable by weak models were filtered out. Evaluation results highlight the critical distinction between model types: while multilingual models scored highly on general benchmarks like MMLU and JMMLU, the results on JamC-QA indicate that they do not fully capture Japan-specific knowledge. Japanese-language models outperform multilingual models, especially on culture- and region-related knowledge such as proverbs, traditional events, and local customs. Furthermore, we find a notable division within Japanese models: models further pretrained on Japanese text excel at administrative and legal questions, while models trained from scratch perform strongly on local and cultural aspects.
Towards Dynamic Metaphor Identification: Evaluating GPT O-Series Models on Five Metaphoricity Cues in U.S. Trade Corpora
Berkay Bas | Jelke Bloem | Xiaojuan Tan
Berkay Bas | Jelke Bloem | Xiaojuan Tan
Although recent advances have focused on detecting metaphors, existing models generally treat them as static entities. There has been little research into identifying dynamic metaphors in discourse. This article addresses this gap by focusing on metaphoricity cues: Linguistic signals that may indicate the activation of metaphoric meaning in different discourse contexts. This study examines the ability of OpenAI’s O-series models (O4-mini, O4-mini-high and O3) in detecting five metaphoricity cues in the U.S. trade discourse, including cues of explicit mapping, emphasis, marking, repetition and novelisation. Research results show that the models performed best on repetition and emphasis, while novelisation was the most difficult cue to detect.
Evaluating Text Style Transfer: A Nine-language Benchmark for Text Detoxification
Vitaly Protasov | Nikolay Babakov | Daryna Dementieva | Alexander Panchenko
Vitaly Protasov | Nikolay Babakov | Daryna Dementieva | Alexander Panchenko
Despite notable advances in large language models (LLMs), reliable evaluation of text generation tasks such as text style transfer (TST) remains an open challenge. Existing research has shown that automatic metrics often correlate poorly with human judgments (Dementieva et al., 2024; Pauli et al., 2025), limiting our ability to assess model performance accurately. Furthermore, most prior work has focused primarily on English, while the evaluation of multilingual TST systems, particularly for text detoxification, remains largely underexplored. In this paper, we present the first comprehensive multilingual benchmarking study of evaluation metrics for text detoxification evaluation across nine languages: Arabic, Amharic, Chinese, English, German, Hindi, Russian, Spanish, Ukrainian. Drawing inspiration from machine translation evaluation, we compare neural-based automatic metrics with LLM-as-a-judge approaches together with experiments on task-specific fine-tuned models. Our analysis reveals that the proposed metrics achieve significantly higher correlation with human judgments compared to baseline approaches. We also provide actionable insights and practical guidelines for building robust and reliable multilingual evaluation pipelines for text detoxification and related TST tasks.
Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
Josh Mcgiff | Tung Khanh Tran | William Mulcahy | Dáibhidh Ó Luinín | Jake Dalzell | Róisín Ní Bhroin | Adam Burke | Barry O’Sullivan | Hoang D. Nguyen | Nikola S. Nikolov
Josh Mcgiff | Tung Khanh Tran | William Mulcahy | Dáibhidh Ó Luinín | Jake Dalzell | Róisín Ní Bhroin | Adam Burke | Barry O’Sullivan | Hoang D. Nguyen | Nikola S. Nikolov
We present Irish-BLiMP (Irish Benchmark of Linguistic Minimal Pairs), the first dataset and framework designed for fine-grained evaluation of linguistic competence in the Irish language, an endangered language. Drawing on a variety of linguistic literature and grammar reference works, a team of fluent Irish speakers manually constructed and reviewed 1020 minimal pairs across a taxonomy of 11 linguistic features. We evaluate both existing Large Language Models (LLMs) and fluent human participants on their syntactic knowledge of Irish. Our findings show that humans outperform all models across all linguistic features, achieving 16.6% higher accuracy on average. Moreover, a substantial performance gap of 18.1% persists between open- and closed-source LLMs, with even the strongest model (gpt-5) reaching only 73.5% accuracy compared to 90.1% by human. Interestingly, human participants and models struggle on different aspects of Irish grammar, thus highlighting a difference in representation learned by the models. Overall, Irish-BLiMP provides the first systematic framework for evaluating the grammatical competence of LLMs in Irish and offers a valuable benchmark for advancing research on linguistic understanding in low-resource languages.
EduBench: A Portuguese Benchmark for Open-Ended Discursive Question Answering
Pedro Henrique Paiola | Luís Gabriel Damiati Mendes | Bruno de Oliveira Monchelato | André da Fonseca Schuck | Gabriel Lino Garcia | Douglas Rodrigues | Helena de Medeiros Caseli | João Paulo Papa
Pedro Henrique Paiola | Luís Gabriel Damiati Mendes | Bruno de Oliveira Monchelato | André da Fonseca Schuck | Gabriel Lino Garcia | Douglas Rodrigues | Helena de Medeiros Caseli | João Paulo Papa
Evaluating open-ended text generation in large language models remains challenging, particularly for non-English languages. We introduce EduBench, a comprehensive Portuguese-language benchmark comprising 3,149 discursive questions from Brazilian university entrance examinations spanning 2015–2025. Unlike multiple-choice or extractive QA benchmarks, EduBench requires extended, argumentative responses across diverse domains, including Humanities, Exact and Natural Sciences, and Languages. Each question includes expert-curated reference answers from official sources, rich metadata, and automated image descriptions to support text-only evaluation. We establish baseline results using nine contemporary models, ranging from 4B-parameter SLMs to state-of-the-art reasoning-capable LLMs, and evaluate them using complementary metrics (BLEU, BERTScore, G-Eval). Our results reveal substantial metric disagreement and highlight the complexity of assessing discursive generation, with models achieving 54–71% alignment with expert answers. We release EduBench publicly to support research on Portuguese NLP and open-ended generation evaluation.
SemBench: A Universal Semantic Framework for LLM Evaluation
Mikel Zubillaga | Naiara Perez | Oscar Sainz | German Rigau
Mikel Zubillaga | Naiara Perez | Oscar Sainz | German Rigau
Recent progress in Natural Language Processing (NLP) has been driven by the emergence of Large Language Models (LLMs), which exhibit remarkable generative and reasoning capabilities. However, despite their success, evaluating the true semantic understanding of these models remains a persistent challenge. Traditional benchmarks such as Word-in-Context (WiC) effectively probe this capability, but their creation is resource-intensive and often limited to high-resource languages. In this paper, we introduce SemBench, a framework for automatically generating synthetic benchmarks that assess the semantic competence of LLMs using only dictionary sense definitions and a sentence encoder. This approach eliminates the need for curated example sentences, making it both scalable and language-independent. We evaluate SemBench in three languages (English, Spanish, and Basque) spanning different levels of linguistic resources, and across a wide range of LLMs. Our results show that rankings derived from SemBench strongly correlate with those obtained from standard WiC datasets. Furthermore, our analysis demonstrates that only a small number of examples is required to achieve stable and meaningful rankings. Overall, SemBench provides a lightweight, adaptable, and data-efficient framework for cross-lingual evaluation of semantic understanding in LLMs.
EL-MIA: Quantifying Membership Inference Risks of Sensitive Entities in LLMs
Ali Satvaty | Suzan Verberne | Fatih Turkmen
Ali Satvaty | Suzan Verberne | Fatih Turkmen
Membership inference attacks (MIA) aim to infer whether a particular data point is part of the training dataset of a model. In this paper, we propose a new task in the context of LLM privacy: entity-level discovery of membership risk focused on sensitive information (PII, credit card numbers, etc). Existing methods for MIA can detect the presence of entire prompts or documents in the LLM training data, but they fail to capture risks at a finer granularity. We propose the “EL-MIA” framework for auditing entity-level membership risks in LLMs. We construct a benchmark dataset for the evaluation of MIA methods on this task. Using this benchmark, we conduct a systematic comparison of existing MIA techniques as well as two newly proposed methods. We provide a comprehensive analysis of the results, trying to explain the relation of the entity level MIA susceptability with the model scale, training epochs, and other surface level factors. Our findings reveal that existing MIA methods are limited when it comes to entity-level membership inference of the sensitive attributes, while this susceptibility can be outlined with relatively straightforward methods, highlighting the need for stronger adversaries to stress test the provided threat model.
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
Bogdan Kostić | Conor Fallon | Julian Risch | Alexander Loeser
Bogdan Kostić | Conor Fallon | Julian Risch | Alexander Loeser
The rapid advancement of Large Language Models (LLMs) has established standardized evaluation benchmarks as the primary instrument for model comparison. Yet, their reliability is increasingly questioned due to sensitivity to shallow variations in input prompts. This paper examines how controlled, truth-conditionally equivalent lexical and syntactic perturbations affect the absolute performance and relative ranking of 23 contemporary LLMs across three benchmarks: MMLU, SQuAD, and AMEGA. We employ two linguistically principled pipelines to generate meaning-preserving variations: one performing synonym substitution for lexical changes, and another using dependency parsing to determine applicable syntactic transformations. Results show that lexical perturbations consistently induce substantial, statistically significant performance degradation across nearly all models and tasks, while syntactic perturbations have more heterogeneous effects, occasionally improving results. Both perturbation types destabilize model leaderboards on complex tasks. Furthermore, model robustness did not consistently scale with model size, revealing strong task dependence. Overall, the findings suggest that LLMs rely more on surface-level lexical patterns than on abstract linguistic competence, underscoring the need for robustness testing as a standard component of LLM evaluation.
The Potential for Misleading Results in Text Sanitisation with Standard Evaluation Metrics
Dan Zhang | Mark Anderson
Dan Zhang | Mark Anderson
Data privacy is an important facet of modern life. It is especially important when considering data that carries potentially sensitive information such as in medical or legal documents. However, it is particularly difficult to ensure private information has been removed or masked in unstructured data, e.g. free-flowing text. The evaluation of systems that automatically detect and remove personal identifiable information (PII) from text is also challenging. Here we present a case study of a system that seemingly performed well, but under closer scrutiny the high performance was due to the shortcomings of standard binary classification metrics in the context of high target class prevalence. We then give a short analysis of different possible metrics in these high-prevalence scenarios, clearly showing the superiority of the Matthews Correlation Coefficient. This is particularly important because readily available data in this domain is rare and often systems are compared using biographies from Wikipedia which have a naturally high prevalence. This can be further aggravated by certain reasonable pre-processing or evaluation formalisms as in the case study discussed here.
The rapid diffusion of Large Language Models (LLMs) across linguistic and cultural contexts underscores the need for systematic safety evaluations beyond English. As LLMs are increasingly applied in multilingual settings, ensuring their safe and appropriate behavior in other languages is essential. This paper presents a methodology for building safety evaluation datasets that comprehensively cover the full spectrum of sensitive topics relevant to LLM safety. The resulting resources include a collection of Italian Wikipedia pages encompassing all major categories of sensitive content, and a companion dataset containing three challenging Italian-language questions per page designed to probe model behavior on high-risk issues. Each prompt was annotated into four safety outcome categories: correct refusal, safe informative, unsafe, and ambiguous. Together, these datasets provide a robust foundation for evaluating and benchmarking LLM safety in Italian. To demonstrate their utility, we used them to assess four LLMs, identifying systematic differences in refusal consistency and compliance across sensitive domains. To support transparency and reproducibility, we release a public repository containing the list of categorized Italian Wikipedia pages, the automatically generated prompts, and the standard prompt template used for safety testing. With this work, we aim to advance language-specific safety assessment and support the responsible, culturally grounded deployment of LLMs beyond English.
Bulgarian Massive Multitask Language Understanding Benchmark
Svetla Peneva Koeva | Ivelina Stoyanova | Dimiter Georgiev | Svetlozara Leseva | Valentina Stefanova | Maria Todorova | Tsvetana Ivanova Dimitrova | Hristina Kukova | Mihaela Moskova | Tinko Tinchev
Svetla Peneva Koeva | Ivelina Stoyanova | Dimiter Georgiev | Svetlozara Leseva | Valentina Stefanova | Maria Todorova | Tsvetana Ivanova Dimitrova | Hristina Kukova | Mihaela Moskova | Tinko Tinchev
Assessing the broad general knowledge of Large Language Models (LLMs) across multiple domains in Bulgarian remains challenging due to the limited availability of Bulgarian evaluation benchmarks. To address this gap, we introduce the Bulgarian Massive Multitask Language Understanding benchmark (MMLU-BG), designed to evaluate whether LLMs possess generalised knowledge capabilities beyond simple text prediction in Bulgarian. This paper presents the structure, the development protocol, and the size of the MMLU-BG benchmark. It is tested in comparison with the original MMLU for English across seven LLMs selected according to specific criteria. The experiments demonstrate that the MMLU-BG benchmark assesses multi-domain versatility and highlights the models’ strengths and weaknesses across different subject areas.
PHEB: An European Portuguese High School-Level LLM Benchmark
Diogo C. Tavares | Rafael Ferreira | Afonso Simplício | Gonçalo Vinagre | Ana Carolina Condez | Inês Calvo | Inês Vieira | David Semedo | Joao Magalhaes
Diogo C. Tavares | Rafael Ferreira | Afonso Simplício | Gonçalo Vinagre | Ana Carolina Condez | Inês Calvo | Inês Vieira | David Semedo | Joao Magalhaes
We present PHEB, a comprehensive benchmark designed to evaluate Large Language Models (LLMs) on real high school level national exams in European Portuguese. The goal is to promote the development of NLP tools and provide a reliable resource for benchmarking multilingual and educational capabilities of LLMs. Covering over 3,500 questions spanning 18 years (2006–2023) across six core subjects, the benchmark compiles high-quality questions from Portuguese National Exams, written and thoroughly curated by professors to ensure topic diversity, linguistic accuracy, and alignment with national curricula. PHEB spans a wide range of subjects, including Mathematics, Portuguese Language and Literature, History, Geography, Biology/Geology, and Philosophy. Questions incorporate both multiple-choice and long-form answers to assess factual knowledge, reasoning capabilities, and language understanding. We comprehensively benchmark state-of-the-art LLMs, shedding light on key challenges such as models’ knowledge, language coverage, answer format biases and robustness to machine translation.
S-GRADES – Studying Generalization of Student Response Assessments in Diverse Evaluative Settings
Tasfia Seuti | Sagnik Ray Choudhury
Tasfia Seuti | Sagnik Ray Choudhury
Evaluating student responses, from long essays to short factual answers, is a key challenge in educational NLP. Automated Essay Scoring (AES) focuses on holistic writing qualities such as coherence and argumentation, while Automatic Short Answer Grading (ASAG) emphasizes factual correctness and conceptual understanding. Despite their shared goal, these paradigms have progressed in isolation with fragmented datasets, inconsistent metrics, and separate communities. We introduce S-GRADES (Studying Generalization of Student Response Assessments in Diverse Evaluative Settings), a web-based benchmark that consolidates 14 diverse grading datasets under a unified interface with standardized access and reproducible evaluation protocols. The benchmark is fully open-source and designed for extensibility, enabling continuous integration of new datasets and evaluation settings. To demonstrate the utility of S-GRADES, we evaluate three state-of-the-art large language models across the benchmark using multiple reasoning strategies in prompting. We further examine the effects of exemplar selection and cross-dataset exemplar transfer. Our analyses illustrate how benchmark-driven evaluation reveals reliability and generalization gaps across essay and short-answer grading tasks, highlighting the importance of standardized, cross-paradigm assessment.
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
Finnur Ágúst Ingimundarson | Steinunn Rut Fridriksdottir | Bjarki Ármannsson | Iris Nowenstein | Steinþór Steingrímsson
Finnur Ágúst Ingimundarson | Steinunn Rut Fridriksdottir | Bjarki Ármannsson | Iris Nowenstein | Steinþór Steingrímsson
This paper evaluates current Large Language Model (LLM) benchmarking for Icelandic, identifies problems, and calls for improved evaluation methods in low/medium-resource languages in particular. We show that benchmarks that include synthetic or machine-translated data that have not been verified in any way, commonly contain severely flawed test examples that are likely to skew the results and undermine the tests’ validity. We warn against the use of such methods without verification in low/medium-resource settings as the translation quality can, at best, only be as good as MT quality for a given language at any given time. Indeed, the results of our quantitative error analysis on existing benchmarks for Icelandic show clear differences between human-authored/-translated benchmarks vs. synthetic or machine-translated benchmarks.
Is This Idea Novel? An Automated Benchmark for Judgment of Research Ideas
Tim Schopf | Michael Färber
Tim Schopf | Michael Färber
Judging the novelty of research ideas is crucial for advancing science, enabling the identification of unexplored directions, and ensuring contributions meaningfully extend existing knowledge rather than reiterate minor variations. However, given the exponential growth of scientific literature, manually judging the novelty of research ideas through literature reviews is labor-intensive, subjective, and infeasible at scale. Therefore, recent efforts have proposed automated approaches for research idea novelty judgment. Yet, evaluation of these approaches remains largely inconsistent and is typically based on non-standardized human evaluations, hindering large-scale, comparable evaluations. To address this, we introduce RINoBench, the first comprehensive benchmark for large-scale evaluation of research idea novelty judgments. It comprises 1,381 research ideas derived from and judged by human experts as well as nine automated evaluation metrics designed to assess both rubric-based novelty scores and textual justifications of novelty judgments. Using this benchmark, we evaluate several state-of-the-art large language models (LLMs) on their ability to judge the novelty of research ideas. Our findings reveal that while LLM-generated reasoning closely mirrors human rationales, this alignment does not reliably translate into accurate novelty judgments, which diverge significantly from human gold standard judgments—even among leading reasoning-capable models. Data and code available at: https://github.com/TimSchopf/RINoBench
Questionnaire Meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
Duc-Hai Nguyen | Vijayakumar Nanjappan | Barry O’Sullivan | Hoang D. Nguyen
Duc-Hai Nguyen | Vijayakumar Nanjappan | Barry O’Sullivan | Hoang D. Nguyen
Millions of people take surveys every day, from market polls to medical questionnaires and customer feedback forms. These datasets capture valuable insights, but the ability of large language models (LLMs) to process questionnaire data, where lists of questions are crossed with hundreds of respondent rows, remains underexplored. Current survey analysis tools (e.g., Qualtrics, SPSS, REDCap) are designed for human operators, leaving practitioners without evidence-based guidance on how to best represent questionnaires for LLM consumption. We address this gap by introducing QASU (Questionnaire Analysis and Structural Understanding), a benchmark that probes six structural skills, including answer lookup, respondent count, and multi-hop inference, across six serialization formats and multiple prompt strategies. Experiments on five LLMs (GPT-5-mini, Gemini-2.5-Flash, Qwen3-32B, Llama3-70B, Amazon Nova Lite) show that format choice significantly impacts performance, with up to 9 percentage points improvement over baseline formats, and reveal substantial gaps (10 to 30 percentage points) between proprietary and open-weight models. Self-augmented prompting yields model-dependent benefits, proving effective for proprietary models but unreliable for open-weight alternatives. By systematically isolating format and prompting effects, our open-source benchmark offers practical guidance for advancing both research and real-world practice in LLM-based questionnaire analysis.
Assessing the Effectiveness of LLMs in Delivering Cognitive Behavioral Therapy
Navdeep Singh Bedi | Ana-Maria Bucur | Noriko Kando | Fabio Crestani
Navdeep Singh Bedi | Ana-Maria Bucur | Noriko Kando | Fabio Crestani
As mental health issues continue to rise globally, there is an increasing demand for accessible and scalable therapeutic solutions. Many individuals currently seek support from Large Language Models (LLMs), even though these models have not been validated for use in counseling services. In this paper, we evaluate LLMs’ ability to emulate professional therapists practicing Cognitive Behavioral Therapy (CBT). Using anonymized, transcribed role-play sessions between licensed therapists and clients, we compare two approaches: (1) a generation-only method and (2) a Retrieval-Augmented Generation (RAG) approach using CBT guidelines. We evaluate both proprietary and open-source models for linguistic quality, semantic coherence, and therapeutic fidelity using standard natural language generation (NLG) metrics, natural language inference (NLI), and automated scoring for skills assessment. Our results indicate that while LLMs can generate CBT-like dialogues, they are limited in their ability to convey empathy and maintain consistency.
Transcription Accuracy in the Icelandic Gigaword Corpus: Evaluating Automatic and Manual Annotation
Johanna Mechler | Lilja Björk Stefánsdóttir | Anton Karl Ingason
Johanna Mechler | Lilja Björk Stefánsdóttir | Anton Karl Ingason
This paper aims to compare automatic and manually corrected annotation data in the Icelandic Gigaword Corpus. We focus on the variable use of Stylistic Fronting (SF) in Icelandic, an optional movement of words or phrases, which indicates a more formal style. Examining SF rates across time, we find that manual coding results in slightly lower SF rates than automatic coding. This difference can be explained by the different sources used in the coding process: For automatic coding, written transcripts compiled by parliament employees are used, and for manual correction, coding relies on audio files of the parliament speeches. Importantly, both types of coding are well suited to trace changing patterns of SF over a span of 16 years, suggesting that the automatic feature extraction reliably reflects the speeches that have been transcribed.
Benchmark Data Contamination in Underrepresented Languages: A Comprehensive Analysis Using Brazilian Data
Iriedson Souto Maior de Moraes Vilar | David Candeia Maia | João Brunet | Fabio Morais | Leandro Balby Marinho
Iriedson Souto Maior de Moraes Vilar | David Candeia Maia | João Brunet | Fabio Morais | Leandro Balby Marinho
Large Language Models (LLMs) are typically evaluated using standardized benchmarks to enable consistent performance measurement and model comparison. However, the reliability of these benchmarks can be undermined by data contamination, which occurs when evaluation items are inadvertently included in training corpora. While this issue has been investigated primarily in high-resource languages such as English and Chinese, its impact on underrepresented languages — such as Brazilian Portuguese — remains understudied. In this paper, we present one of the first systematic investigations of benchmark data contamination (BDC) in an underrepresented language setting, using Brazilian Portuguese as a case study. Using validated methodologies from the literature, we evaluate specialized and multilingual models across four benchmarks: BLUEX, ENEM Challenge, OAB Exams, and HealthQA-BR. Our approach applyes TS-Guessing to detect contamination via memorized knowledge, alongside a 50-character n-gram similarity strategy to identify benchmark items leaked into training data. Our results provide consistent evidence of contamination, revealing that models with stronger memorization and retrieval abilities tend to achieve artificially inflated benchmark scores. Our contributions include: (i) classifying models according to their contamination risk, (ii) identifying the benchmarks most affected by data leakage, and (iii) reporting contaminated training corpora.
TTSVowelViz: A Tool for Visualising Text-to-Speech Model Training via Vowel Spaces
Pasindu Udawatta | Jesin James | Balamurali B T | Catherine Inez Watson | Ake Nicholas | Binu Nisal Abeysinghe
Pasindu Udawatta | Jesin James | Balamurali B T | Catherine Inez Watson | Ake Nicholas | Binu Nisal Abeysinghe
In text-to-speech (TTS) model training, the saturation of the loss curve indicates how well a model learns the characteristics of the training dataset. But it does not reveal the linguistic properties learned by the model. Existing TTS approaches miss the potential to incorporate linguistic insights into model training. We introduce TTSVowelViz, a novel tool that visualises static and dynamic vowel spaces during model training, bridging linguistic knowledge and TTS model development. It helps identify which vowel sounds are accurately learned and how the vowel spaces are evolved during training. To assess TTSVowelViz, we fine-tuned a TTS model from General American English to New Zealand English and conducted a perception test. Our results show that the formants of specific vowels in the vowel spaces generated by TTSVowelViz align with human perception, effectively visualising the perceived accent shift. This work highlights vowel space visualisation as a valuable interpretability tool for TTS training.
A Sociophonetic Analysis of Racial Bias in Commercial ASR Systems Using the Pacific Northwest English Corpus
Michael Scott | Siyu Liang | Alicia Wassink | Gina-Anne Levow
Michael Scott | Siyu Liang | Alicia Wassink | Gina-Anne Levow
This paper presents a systematic evaluation of racial bias in four major commercial automatic speech recognition (ASR) systems using the Pacific Northwest English (PNWE) corpus. We analyze transcription accuracy across speakers from four ethnic backgrounds (African American, Caucasian American, ChicanX, and Yakama) and examine how sociophonetic variation contributes to differential system performance. We introduce a heuristically-determined Phonetic Error Rate (PER) metric that links recognition errors to specific linguistically motivated variables derived from sociophonetic annotation. Our analysis of eleven sociophonetic features reveals that vowel quality variation, particularly resistance to the low-back merger and pre-nasal merger patterns, is systematically associated with differential error rates across ethnic groups, with the most pronounced effects for African American speakers across all evaluated systems. These findings demonstrate that acoustic modeling of dialectal phonetic variation, rather than lexical or syntactic factors, remains a primary source of bias in commercial ASR systems. The study establishes the PNWE corpus as a valuable resource for bias evaluation in speech technologies and provides actionable guidance for improving ASR performance through targeted representation of sociophonetic diversity in training data.
ParliaBench: An Evaluation and Benchmarking Framework for LLM-Generated Parliamentary Speech
Marios Koniaris | Argyro Tsipi | Panayiotis Tsanakas
Marios Koniaris | Argyro Tsipi | Panayiotis Tsanakas
Parliamentary speech generation presents specific challenges for large language models beyond standard text generation tasks. Unlike general text generation, parliamentary speeches require not only linguistic quality but also political authenticity and ideological consistency. Current language models lack specialized training for parliamentary contexts, and existing evaluation methods focus on standard NLP metrics rather than political authenticity. To address this, we present ParliaBench, a benchmark for parliamentary speech generation. We constructed a dataset of 448k speeches from UK Parliament to enable systematic model training. We introduce an evaluation framework combining computational metrics with LLM-as-a-judge assessments for measuring generation quality across three dimensions: linguistic quality, semantic coherence, and political authenticity. We propose two novel embedding-based metrics, Political Spectrum Alignment and Party Alignment, to quantify ideological positioning. We fine-tuned five large language models (LLMs), generated 28k speeches, and evaluated them using our framework, comparing baseline and fine-tuned models. Results show that fine-tuning produces statistically significant improvements across the majority of metrics and our novel metrics demonstrate strong discriminative power for political dimensions otherwise absent from conventional evaluation, while domain fine-tuning reveals a measurable trade-off between political authenticity and lexical diversity.
PARSEME 2.0 Multilingual Corpus of Multiword Expressions
Agata Savary | Manon Scholivet | Carlos Ramisch | Takuya Nakamura | Eric Bilinski | Sara Stymne | Voula Giouli | Stella Markantonatou | Vasile Pais | Maria Mitrofan | Louis Estève | Bruno Guillaume | Verginica Barbu Mititelu | Jaka Čibej | Roberto Díaz Hernández | Victoria Fendel | Polona Gantar | Olha Kanishcheva | Cvetana Krstev | Chaya Liebeskind | Irina Lobzhanidze | Aleksandra M. Marković | Gunta Nešpore-Bērzkalne | Adriana S. Pagano | Mehrnoush Shamsfard | Ranka Stankovic | Vahide Tajalli | Carole Tiberius | Aakanksha Padhye
Agata Savary | Manon Scholivet | Carlos Ramisch | Takuya Nakamura | Eric Bilinski | Sara Stymne | Voula Giouli | Stella Markantonatou | Vasile Pais | Maria Mitrofan | Louis Estève | Bruno Guillaume | Verginica Barbu Mititelu | Jaka Čibej | Roberto Díaz Hernández | Victoria Fendel | Polona Gantar | Olha Kanishcheva | Cvetana Krstev | Chaya Liebeskind | Irina Lobzhanidze | Aleksandra M. Marković | Gunta Nešpore-Bērzkalne | Adriana S. Pagano | Mehrnoush Shamsfard | Ranka Stankovic | Vahide Tajalli | Carole Tiberius | Aakanksha Padhye
We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal, adjectival, adverbial and functional. We cover 17 languages, of which 7 are new. The annotation process is based on cross-lingually unified guidelines, phrased as decision diagrams over linguistic tests, and a typology of 18 MWE categories. The corpus contains almost 5 million tokens, over 250,000 sentences and 140,000 MWE annotations. The applicability of the corpus is tested in baseline experiments with a prompt-based MWE identification system. Results show that generic large language models do not encode sufficient knowledge to solve the MWE identification task.
Metaphor plays a central role in human language and thought, and corpus-linguistic approaches enable its systematic investigation. Such research requires large, representative collections of metaphor-annotated linguistic data from diverse contexts. Despite the increasing availability of metaphor corpora in various languages, Persian remains underrepresented, with few publicly available resources and no large-scale register-diverse metaphor corpus. This paper introduces PerMet 1.0, a metaphor-annotated corpus for Persian. The corpus consists of approximately 120,000 tokens (about 99,000 lexical units) drawn from five registers: academic, news, fiction, social media, and spoken discourse. Five independent annotators labeled the corpus using Metaphor Identification Procedure Vrije Universiteit (MIPVU), with adaptations for Persian. Inter-annotator agreement showed a high level of consistency (κ = 0.952), confirming the reliability of the annotation. Preliminary analysis shows that 13.1% of the lexical units are related to metaphor, with the academic register showing the highest proportion, followed by news, social media, spoken, and fiction. PerMet 1.0 offers a foundational resource for research on metaphor in Persian, cross-linguistic comparative studies, and the development and fine-tuning of machine learning or large language models for automatic metaphor identification.
Multi-SimLex for Dutch: Benchmarking Embedding- and Prompt-Based Model Performance on Semantic Similarity
Lizzy Brans | Jelke Bloem
Lizzy Brans | Jelke Bloem
We introduce Dutch Multi-SimLex, a 1,888–pair extension of the Multi-SimLex benchmark for evaluating lexical semantic similarity in Dutch. The dataset was rated by 100 native speakers on a 0–6 scale and shows high reliability (overall ICC(2,k)=0.82) as well as strong alignment with English (ρ=0.73). Using this resource, we evaluate eighteen models across four architectural families: static embeddings, encoder-only transformers, encoder–decoders, and decoder-only LLMs. We evaluate models using two complementary approaches: embedding-based cosine similarity and prompted similarity judgments in Dutch. In embedding-based evaluation, FastText (ρ=0.485) and the monolingual Dutch encoder BERTje (ρ=0.468) achieve the strongest alignment with human ratings, while multilingual encoders such as mBERT (ρ=0.208) and XLM-R (ρ=0.186) perform weaker. Prompt-based evaluation yields substantially higher correlations, with GPT-4 (ρ=0.761) performing best, followed by DeepSeek-V3 (ρ=0.753) and Gemini 1.5 Pro (ρ=0.722). Together, the results show that model performance depends strongly on how meaning is tested. Dutch Multi-SimLex provides a reliable foundation for evaluating meaning across architectures and advancing Dutch semantic evaluation.
MultiCoS: A Multilingual Dataset of Connective Semantics with Context–Sentence Compatibility
Anne Mucha | Ciyang Qing | Wataru Uegaki
Anne Mucha | Ciyang Qing | Wataru Uegaki
We present a multilingual dataset of connective semantics. The dataset contains the semantic annotations of clausal connectives (e.g. and and or in English) from 24 languages, based on our original native-speaker elicitation data. Unlike existing lexica on connectives, the dataset includes systematic evidence for the annotations in the form of context-sentence compatibility judgments, including negative evidence. The paper describes the methodology of data collection and the format of the dataset. We also discuss its potential use cases for the validation of cross-linguistic generalizations, examinations of their potential counterexamples, and for benchmarking felicity judgments by NLU systems.
Adverbs Revisited: Enhancing WordNet Coverage of Adverbs with a Supersense Taxonomy
Jooyoung Lee | Jader Martins Camboim de Sá | Cedric Pruski
Jooyoung Lee | Jader Martins Camboim de Sá | Cedric Pruski
WordNet offers rich supersense hierarchies for nouns and verbs, yet adverbs remain underdeveloped, lacking a systematic semantic classification. We introduce a linguistically grounded supersense typology for adverbs, empirically validated through annotation, that captures major semantic domains including manner, temporal, frequency, degree, domain, speaker-oriented, and subject-oriented functions. Results from a pilot annotation study demonstrate that these categories provide broad coverage of adverbs in natural text and can be reliably assigned by human annotators. Incorporating this typology extends WordNet’s coverage, aligns it more closely with linguistic theory, and facilitates downstream NLP applications such as word sense disambiguation, event extraction, sentiment analysis, and discourse modeling. We present the proposed supersense categories, annotation outcomes, and directions for future work.
KinyCOMET: Automatic Evaluation of Machine Translation Systems for Kinyarwanda–English
Prince Chris Mazimpaka | Jan Nehring | Samuel Rutunda | Cristina España-Bonet
Prince Chris Mazimpaka | Jan Nehring | Samuel Rutunda | Cristina España-Bonet
This paper presents KinyCOMET, a new automatic evaluation metric for Kinyarwanda–English machine translation (MT). Current MT evaluation in Rwanda relies mainly on BLEU and chrF, which have been shown to correlate poorly with human judgments. To address this gap, we created a Direct Assessment (DA) dataset for Kinyarwanda-English translations and used it to fine-tune COMET models for this language pair. We evaluate two variants: KinyCOMET XLM-RoBERTa, trained from a multilingual encoder without Kinyarwanda data, and KinyCOMET Unbabel, a fine-tuned version of the Unbabel COMET model. Both models achieve strong correlations with human evaluations, with KinyCOMET Unbabel outperforming all baselines, including AfriCOMET, chrF, and BLEU. Our results show that fine-tuning pre-trained multilingual models can yield high-quality evaluators even for low-resource languages that the base model was not trained on. We release both the models and the annotated dataset publicly to foster further research on African language evaluation.
Multiway Parallel Corpus in Forced Migration Domain for Multilingual Machine Translation
Fatemeh Azadi | Samuel Larkin | Chi-kiu Lo
Fatemeh Azadi | Samuel Larkin | Chi-kiu Lo
High-quality domain-specific parallel corpora play a significant role in improving the performance of machine translation (MT) and multilingual natural language processing (NLP) systems in a target domain. However, most existing multilingual parallel corpora focus on general-purpose data, and a majority of highly specialized domains such as forced migration are suffering from lack of multilingual data. In this work, we present a new high-quality 4-way parallel corpus in the forced migration domain. The corpus consists of human-translated journal articles from Forced Migration Review in English, French, Spanish, and Arabic. Our corpus contains data aligned at both document and sentence level in four languages and provides a clean and reliable 4-way parallel resource for multilingual research in forced migration. Using this dataset, we benchmark several open-weight large language models (LLMs), an open-weight multilingual MT system, online closed MT systems, and a closed LLM across 12 translation directions. We further leverage our corpus to improve the MT quality of a top-performing multilingual foundation model with two common domain adaptation approaches, fine-tuning and few-shot prompting. Our results demonstrate the effectiveness of our corpus in improving the translation performance of current models in the forced migration domain.
Context-8: A Data Set for Evaluating Context Sensitivity in Machine Translation
Dongyue Wang | Kyo Kageura
Dongyue Wang | Kyo Kageura
Context plays a crucial role in translation, enhancing both accuracy and fluency. With the advancement of machine translation (MT), the concept of context is now considered across an increasingly broader range of phenomena. Despite its importance, however, systematic definitions of context provided by communication studies and translation studies remain fragmented, and the concept of context remains elusive in MT research. To the best of our knowledge, no dataset currently exists that comprehensively evaluates MT’s sensitivity to context. In this study, we propose a systematic taxonomy of context and introduce Context-8, an evaluation dataset designed to assess context sensitivity in MT for English-to-Japanese translation. The initial release includes 130 groups comprising 533 English-to-Japanese translation examples, each requiring different context categories to produce accurate and fluent translations. The data are taken from both hand-crafted and online materials. We release Context-8 to support the evaluation and benchmarking of MT systems with respect to context sensitivity.
AssamLegalTrans: A Parallel Corpus, Benchmark and Analysis for English-Assamese Machine Translation of Legal Judgments
Telem Joyson Singh | Hemanta Baruah | Sanasam Ranbir Singh | Anindita Talukdar | Nasrin Shahnaz | Okram Jimmy Singh | Priyankoo Sarmah | Pallav Kumar Dutta | Sukumar Nandi | Pranab Duara
Telem Joyson Singh | Hemanta Baruah | Sanasam Ranbir Singh | Anindita Talukdar | Nasrin Shahnaz | Okram Jimmy Singh | Priyankoo Sarmah | Pallav Kumar Dutta | Sukumar Nandi | Pranab Duara
In India, the official language for writing judgments in higher courts is English, which creates a language barrier for citizens not proficient in English. Machine Translation (MT) provides a scalable solution, but its progress for low-resource languages like Assamese is significantly limited due to the lack of legal domain data. To address this gap, we introduce the first-of-its-kind English-Assamese parallel corpus for the translation of Indian court judgments. This dataset consists of over 55,000 manually translated and validated sentence pairs from over 500 judgments of the Gauhati High Court and the Supreme Court of India. Using this dataset, we perform a comprehensive evaluation of state-of-the-art multilingual models, including NLLB-200 and Sarvam-Translate, in both zero-shot and fine-tuned settings, comparing their performance against commercial systems. Our experiments show that fine-tuning on our legal-domain dataset significantly improves the translation quality. We also conduct a thorough error analysis that points out important issues in legal translation. These include precisely translating legal terms, properly transliterating named entities, expanding abbreviations, and transforming sentence structures, such as changing passive voice to active voice, when translating from English to Assamese. By creating a publicly available dataset and examining the specific challenges, this work offers a reproducible foundation and a clear way to develop more accurate and reliable legal machine translation systems. This will help improve access to justice for Assamese speakers.
Coordinate Structure Extraction for Patent Claims Using Multilingual LLMs
Tsukasa Ishimaru | Takehito Utsuro | Masaaki Nagata
Tsukasa Ishimaru | Takehito Utsuro | Masaaki Nagata
This study proposes a simple, one-stage approach to coordinate structure extraction using multilingual Large Language Models (LLMs) with Translation between Augmented Natural Languages (TANL) to develop an error detection system for coordinate structure translation. Unlike conventional multi-component methods such as CoRec, our method employs an end-to-end Transformer decoder (LLM) trained via Continual Pre-Traning (CPT) and/or Supervised Fine-Tuning (SFT) on English and Japanese datasets obtained from parsed treebanks that includes coordinate structures. We evaluated the proposed models on 100 English and Japanese patent claims manually annotated with coordinate structure tags. The proposed method using open-weight models such as Llama-3.2-8B or gemma-3-4b-it significantly outperformed GPT-5 and CoRec by approximately 0.02-0.03 in F1 score for the English task. The proposed method using open-weight models such as llama-3-youko-8b and Llama-3-swallow-8B-0.1v significantly outperformed GPT-5 by approximately 0.02-0.05 in F1 score for the Japanese task. In addition, models using both English and Japanese training data significantly outperform those using monolingual training data only.
Human Label Variation in Implicit Discourse Relation Recognition
Frances Yung | Daniil Ignatev | Merel Scholman | Vera Demberg | Massimo Poesio
Frances Yung | Daniil Ignatev | Merel Scholman | Vera Demberg | Massimo Poesio
There is growing recognition that many NLP tasks lack a single ground truth, as human judgments reflect diverse perspectives. To capture this variation, models have been developed to predict full annotation distributions rather than majority labels, while perspectivist models aim to reproduce the interpretations of individual annotators. In this work, we compare these approaches on Implicit Discourse Relation Recognition (IDRR), a highly ambiguous task where disagreement often arises from cognitive complexity rather than ideological bias. Our experiments show that existing annotator-specific models perform poorly in IDRR unless ambiguity is reduced, whereas models trained on label distributions yield more stable predictions. Further analysis indicates that frequent cognitively demanding cases drive inconsistency in human interpretation, posing challenges for perspectivist modeling in IDRR.
Recent research has explored the capacity of Large Language Models (LLMs) to perform pragmatic reasoning and interpret complex pragmatic phenomena. However, such phenomena are inherently ambiguous, and even human evaluations are highly variable. Many existing studies directly compare human and model responses while assuming a single “correct” interpretation, thereby overlooking the natural variability that characterizes human pragmatic understanding. This raises two key issues: (1) the need for novel evaluation methods that account for interpretive variability and allow for meaningful comparison between humans and models, and (2) the potential limitations of current linguistic theories in capturing the richness of human pragmatic behavior. We propose that LLMs can serve not only as benchmarks for human-model alignment, but also as tools for investigating the nature of pragmatic phenomena and their relationship to linguistic theory. To this end, we developed a handcrafted dataset encompassing eight types of conversational implicatures. Our study addresses three main research questions: (1) Do LLMs process conversational implicatures differently from humans? (2) If so, how do these differences manifest? (3) What do these findings reveal about the cognitive capacities of LLMs and the explanatory adequacy of pragmatic theory?
Instruction-tuning fundamentally transforms how language models process linguistic input and interact with the user. Through the lens of speech act theory, we investigate whether instruction-tuning causes models to shift from prioritizing syntactical form to pragmatic intent. We create a controlled dataset of 400 sentences systematically varying along two dimensions: syntactical structure (declarative vs. interrogative) and communicative intent (assertive vs. request). Using Principal Component Analysis on hidden state representations from Qwen2.5 (1.5B-7B) and models from two other families (Gemma3-1B, and Llama3.2-3B), we reveal a consistent pattern: base models cluster sentences by syntactical form, while instruction-tuned models reorganize representations around pragmatic intent. This syntactic-to-pragmatic shift occurs in middle layers, with declarative requests and interrogative requests—maximally separated in base models—becoming the most similar categories after instruction-tuning. The phenomenon explains how instruction-tuned models correctly interpret indirect speech acts, treating polite declaratives like I’d appreciate corrections" as functionally equivalent to direct interrogatives. Our findings demonstrate that instruction-tuning teaches models to prioritize the communicative dimension over surface form, a fundamental reorganization consistent across model scales and architectures.
Distributed Partial Information Puzzles: Examining Common Ground Construction under Epistemic Asymmetry
Yifan Zhu | Mariah Bradford | Kenneth Lai | Timothy Obiso | Videep Venkatesha | James Pustejovsky | Nikhil Krishnaswamy
Yifan Zhu | Mariah Bradford | Kenneth Lai | Timothy Obiso | Videep Venkatesha | James Pustejovsky | Nikhil Krishnaswamy
Establishing common ground, a shared set of beliefs and mutually recognized facts, is fundamental to collaboration, yet remains a challenge for current AI systems, especially in multimodal, multiparty settings, where the collaborators bring different information to the table. We introduce the Distributed Partial Information Puzzle (DPIP), a collaborative construction task that elicits rich multimodal communication under epistemic asymmetry. We present a multimodal dataset of these interactions, annotated and temporally aligned across speech, gesture, and action modalities to support reasoning over propositional content and belief dynamics. We then evaluate two paradigms for modeling common ground (CG): (1) state-of-the-art large language models (LLMs), prompted to infer shared beliefs from multimodal updates, and (2) an axiomatic pipeline grounded in Dynamic Epistemic Logic (DEL) that incrementally performs the same task. Results on the annotated DPIP data indicate that it poses a challenge to modern LLMs’ abilities to track both task progression and belief state.
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
Nan Li | Albert Gatt | Massimo Poesio
Nan Li | Albert Gatt | Massimo Poesio
Collaborative dialogue relies on participants incrementally establishing common ground, yet in asymmetric settings they may believe they agree while referring to different entities. We introduce a perspectivist annotation scheme for the HCRC MapTask corpus (Anderson et al., 1991) that separately captures speaker and addressee grounded interpretations for each reference expression, enabling us to trace how understanding emerges, diverges, and repairs over time. Using a scheme-constrained LLM annotation pipeline, we obtain 13k annotated reference expressions with reliability estimates and analyze the resulting understanding states. The results show that full misunderstandings are rare once lexical variants are unified, but multiplicity discrepancies systematically induce divergences, revealing how apparent grounding can mask referential misalignment. Our framework provides both a resource and an analytic lens for studying grounded misunderstanding and for evaluating (V)LLMs’ capacity to model perspective-dependent grounding in collaborative dialogue.
Assessing LLM Reasoning through Implicit Causal Chain Discovery in Climate Discourse
Liesbeth Allein | Nataly Pineda-Castañeda | Andrea Rocci | Marie-Francine Moens
Liesbeth Allein | Nataly Pineda-Castañeda | Andrea Rocci | Marie-Francine Moens
How does a cause lead to an effect, and which intermediate causal steps explain their connection? This work scrutinizes the mechanistic causal reasoning capabilities of large language models (LLMs) to answer these questions through the task of implicit causal chain discovery. In a diagnostic evaluation framework, we instruct nine LLMs to generate all possible intermediate causal steps linking given cause-effect pairs in causal chain structures. These pairs are drawn from recent resources in argumentation studies featuring polarized discussion on climate change. Our analysis reveals that LLMs vary in the number and granularity of causal steps they produce. Although they are generally self-consistent and confident about the intermediate causal connections in the generated chains, their judgments are mainly driven by associative pattern matching rather than genuine causal reasoning. Nonetheless, human evaluations confirmed the logical coherence and integrity of the generated chains. Our baseline causal chain discovery approach, insights from our diagnostic evaluation, and benchmark dataset with causal chains lay a solid foundation for advancing future work in implicit, mechanistic causal reasoning in argumentation settings.
AccurateRAG: A Framework for Building Accurate Retrieval-Augmented Question-Answering Applications
Linh The Nguyen | Chi Tran | Dung Ngoc Nguyen | Van-Cuong Pham | Hoang Ngo | Dat Quoc Nguyen
Linh The Nguyen | Chi Tran | Dung Ngoc Nguyen | Van-Cuong Pham | Hoang Ngo | Dat Quoc Nguyen
We introduce AccurateRAG—a novel framework for constructing high-performance question-answering applications based on retrieval-augmented generation (RAG). Our framework offers a pipeline for development efficiency with tools for raw dataset processing, fine-tuning data generation, text embedding & LLM fine-tuning, output evaluation, and building RAG systems locally. Experimental results show that our framework outperforms previous strong baselines and obtains new state-of-the-art question-answering performance on benchmark datasets.
VideoEvent: Leveraging Relevance and LLMs for Video Question Answering
Chen-Chen Lin | Ming-Han Lee | KunRu Wu | Yu-Chee Tseng
Chen-Chen Lin | Ming-Han Lee | KunRu Wu | Yu-Chee Tseng
We propose VideoEvent, a lightweight and efficient training-free framework for Video Question Answering (VQA) with large language models (LLMs). Although several training-free VQA methods have been proposed, they often neglect the temporal dependencies between frames or clips, treating them as isolated units and relying on complex or resource-intensive components. To address this limitation while maintaining performance and simplicity, we propose VideoEvent, a framework that segments an input video into question-relevant temporal events and selectively supplements them with low-level visual cues such as background and object layout. Our method selects semantically relevant time spans and retrieves one representative background frame to enrich the prompt to LLM. This design minimizes reliance on additional tools and reduces inference cost, making it highly suitable for practical deployment. Experimental results on EgoSchema and NExT-QA show that VideoEvent reduces inference cost by up to 30% while maintaining state-of-the-art accuracy, and its background module improves accuracy by 1–3% across multiple frameworks.
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
Wen-wai Yim | Asma Ben Abacha | Zixuan Yu | Robert Doerning | Fei Xia | Meliha Yetisgen
Wen-wai Yim | Asma Ben Abacha | Zixuan Yu | Robert Doerning | Fei Xia | Meliha Yetisgen
Evaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses exists. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs’ sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. We release our code and annotations to support future research.
LegalRikai: Open Benchmark – a Benchmark for Complex Japanese Corporate Legal Tasks
Shogo Fujita | Yuji Naraki | Yiqing Zhu | Shinsuke Mori
Shogo Fujita | Yuji Naraki | Yiqing Zhu | Shinsuke Mori
This paper introduces LegalRikai: Open Benchmark, a new benchmark comprising four complex tasks that emulate Japanese corporate legal practices. The benchmark was created by legal professionals under the supervision of an attorney. This benchmark has 100 samples that require long-form, structured outputs, and we evaluated them against multiple practical criteria. We conducted both human and automated evaluations using leading LLMs, including GPT-5, Gemini 2.5 Pro, and Claude Opus 4.1. Our human evaluation revealed that abstract instructions prompted unnecessary modifications, highlighting model weaknesses in document-level editing that were missed by conventional short-text tasks. Furthermore, our analysis reveals that automated evaluation aligns well with human judgment on criteria with clear linguistic grounding, and assessing structural consistency remains a challenge. The result demonstrates the utility of automated evaluation as a screening tool when expert availability is limited. We propose a dataset evaluation framework to promote more practice-oriented research in the legal domain.
Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller Models
Neeraj Gangwar | Suma Bhat | Nickvash Kani
Neeraj Gangwar | Suma Bhat | Nickvash Kani
While large models pre-trained on high-quality data exhibit excellent performance on mathematical reasoning (e.g., GSM8k, MultiArith), it remains challenging to specialize smaller models for these tasks. Common approaches to address this challenge include knowledge distillation from large teacher models and data augmentation (e.g., rephrasing questions and generating synthetic solutions). Despite these efforts, smaller models struggle with arithmetic computations, leading to errors in mathematical reasoning. In this work, we leverage a synthetic arithmetic dataset generated programmatically to enhance the reasoning capabilities of smaller models. We investigate two key approaches to incorporate this dataset: (1) intermediate fine-tuning, in which a model is fine-tuned on the arithmetic dataset before training it on a reasoning dataset, and (2) integrating the arithmetic dataset into an instruction-tuning mixture, allowing the model to learn arithmetic skills alongside general instruction-following abilities. Our experiments on multiple reasoning benchmarks demonstrate that incorporating an arithmetic dataset, whether through targeted fine-tuning or within an instruction-tuning mixture, enhances models’ arithmetic capabilities, thereby improving their mathematical reasoning performance.
mSCoRe: A Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning
Nghia Trung Ngo | Franck Dernoncourt | Thien Huu Nguyen
Nghia Trung Ngo | Franck Dernoncourt | Thien Huu Nguyen
Recent advancements in reasoning-reinforced Large Language Models (LLMs) have shown remarkable capabilities in complex reasoning tasks. However, the mechanism underlying their utilization of different human reasoning skills remains poorly investigated, especially for multilingual commonsense reasoning that involves everyday knowledge across different languages and cultures. To address this gap, we propose a Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning (mSCoRe). Our benchmark incorporates three key components that are designed to systematically evaluate LLM’s reasoning capabilities, including: (1) a novel taxonomy of reasoning skills that enables fine-grained analysis of models’ reasoning processes, (2) a robust data synthesis pipeline tailored specifically for commonsense reasoning evaluation, and (3) a complexity scaling framework allowing task difficulty to scale dynamically alongside future improvements in LLM abilities. Extensive experiments on eights state-of-the-art LLMs of varying sizes and training approaches demonstrate that mSCoRe remains significantly challenging for current models, particularly at higher complexity levels. Our results reveal the limitations of such reasoning-reinforced models when confronted with nuanced multilingual general and cultural commonsense. We further provide detailed analysis on the models’ reasoning processes, suggesting future directions for improving multilingual commonsense reasoning capabilities.
A Binary Problem in Binary QA: Diverse LLMs or Diverse Question Interpretations? That Is the Ensembling Question
Rafael Rosales | Santiago Miret
Rafael Rosales | Santiago Miret
Effectively leveraging diversity has been shown to improve performance for various machine learning models, including large language models (LLMs). However, determining the most effective way of using diversity remains a challenge. In this work, we compare two diversity approaches for answering binary questions using LLMs: model diversity, which relies on multiple models answering the same question, and question interpretation diversity, which relies on using the same model to answer the same question framed in different ways. For both cases, we apply majority voting as the ensemble consensus heuristic to determine the final answer. Our experiments on boolq, strategyqa, and pubmedqa show that question interpretation diversity consistently leads to better ensemble accuracy compared to model diversity. Furthermore, our analysis of GPT and LLaMa shows that model diversity typically produces results between the best and the worst ensemble members without clear improvement.
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
Shubhra Ghosh | Abhilekh Borah | Aditya Kumar Guru | Kripabandhu Ghosh
Shubhra Ghosh | Abhilekh Borah | Aditya Kumar Guru | Kripabandhu Ghosh
The rapid proliferation of Large Language Models (LLMs) has significantly contributed to the development of equitable AI systems capable of factual question-answering (QA). However, no known study tests the LLMs’ robustness when presented with obfuscated versions of questions. To systematically evaluate these limitations, we propose a novel technique, ObfusQAte and leveraging the same, introduce ObfusQA, a comprehensive, first of its kind, framework, with multi-tiered obfuscation levels designed to examine LLM capabilities across three distinct dimensions: (i) Named-Entity Indirection, (ii) Distractor Indirection, and (iii) Contextual Overload. By capturing these fine-grained distinctions in language, ObfusQA provides a comprehensive benchmark for evaluating LLM robustness and adaptability. Our study observes that LLMs exhibit a tendency to fail or generate hallucinated responses, when confronted with these increasingly nuanced variations. To foster research in this direction, we make ObfusQAte publicly available.
POLAR: A Corpus of Questions, Responses and Argumentation in Polish Political Radio Discourse
Daniel Ziembicki | Aleksandra Zwierzchowska | Ewelina Sobol | Katarzyna Anna Przerada
Daniel Ziembicki | Aleksandra Zwierzchowska | Ewelina Sobol | Katarzyna Anna Przerada
In this paper, we present POLAR: an experimental dataset designed to investigate question–answer structures in political interviews. The study also aims to integrate this level of annotation with the identification of argumentative structures. The dataset comprises orthographic transcriptions of Polish political radio interviews conducted between December 2023 and March 2024, with a total duration of nearly 10 hours of recordings (94,015 tokens). Manual annotation was performed on three levels: (a) identification of questions as speech acts, (b) classification of responses to questions, and (c) argumentative structures in which interrogative sentences function as premises or conclusions. The results show that not all interrogative sentences function as questions in the sense of requesting information — 23% do not serve this function, while 13% were identified as components of argumentative structures. We also introduce a gold-standard corpus, together with baseline experiments and LLM-based evaluations, demonstrating the usefulness of the resource for both theoretical research and NLP applications.
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation
Jeongsoo Lee | Daeyong Kwon | Kyohoon Jin | JunNyeong Jeong | Minwoo Sim | Minwoo Kim
Jeongsoo Lee | Daeyong Kwon | Kyohoon Jin | JunNyeong Jeong | Minwoo Sim | Minwoo Kim
Existing RAG benchmarks often overlook query difficulty, leading to inflated performance on simpler questions and unreliable evaluations. A robust benchmark dataset must satisfy three key criteria: quality, ensuring complete and reliable ground truth (GT) responses; diversity, expanding semantic coverage to prevent overfitting; and difficulty, capturing the complexity of reasoning based on hops and the distribution of supporting evidence. However, current benchmarks lack a systematic approach to defining and controlling query difficulty at a fine-grained level. To address this, we propose MHTS (Multi-Hop Tree Structure), a novel dataset synthesis framework that systematically controls multi-hop reasoning complexity by leveraging a multi-hop tree structure to generate logically connected, multi-chunk queries. Our fine-grained difficulty estimation formula exhibits a strong correlation with the overall performance metrics of a RAG system, validating its effectiveness in assessing both retrieval and answer generation capabilities. By ensuring high-quality, diverse, and difficulty-controlled queries, our approach enhances RAG evaluation and benchmarking capabilities. This work contributes to the development of more reliable, efficient, and adaptable AI-driven research assistants, facilitating advancements in document-based reasoning and retrieval tasks.
CareMedEval Dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
Doria Bonzi | Alexandre Guiggi | Frederic Bechet | Carlos Ramisch | Benoit Favre
Doria Bonzi | Alexandre Guiggi | Frederic Bechet | Carlos Ramisch | Benoit Favre
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliability remains limited, particularly for critical reasoning in specialized domains. We introduce CareMedEval, an original dataset designed to evaluate LLMs on biomedical critical appraisal and reasoning tasks. Derived from authentic exams taken by French medical students, the dataset contains 534 questions based on 37 scientific articles. Unlike existing benchmarks, CareMedEval explicitly evaluates critical reading and reasoning grounded in scientific papers. Benchmarking state-of-the-art generalist and biomedical-specialized LLMs under various context conditions reveals the difficulty of the task: open and commercial models fail to exceed an Exact Match Rate of 0.5 even though generating intermediate reasoning tokens considerably improves the results. Yet, models remain challenged especially on questions about study limitations and statistical analysis. CareMedEval provides a challenging benchmark for grounded reasoning, exposing current LLM limitations and paving the way for future development of automated support for critical appraisal.
LongTailQA: Benchmarking LLMs and RAG Models on Disambiguated Long-Tail Entities
William Xion | Uwe Hadler | Tim Cofala | Maximilian Idahl | Soumyadeep Roy | Wolfgang Nejdl
William Xion | Uwe Hadler | Tim Cofala | Maximilian Idahl | Soumyadeep Roy | Wolfgang Nejdl
Large Language Models (LLMs) struggle with memorizing long-tail facts. Retrieval-Augmented Generation (RAG) models show better performance on long-tail Question Answering (QA) by offloading memory to external knowledge sources. We demonstrate that popular QA benchmarks such as PopQA, WITQA, and EntityQA contain significant entity ambiguity, with 8-30% of long-tail questions referencing entities with non-unique names. This ambiguity confounds evaluation, obscuring true model capabilities. To perform robust benchmarking, we disambiguate these questions with the Wikipedia knowledge graph to develop LongTailQA, an improved QA benchmark that mitigates entity ambiguity in long-tail entity questions. We evaluate various recent LLMs and RAG models, such as Self-RAG and InstructRAG, investigating retriever quality and retrieval depth impacts on QA performance. We observe that: (i) disambiguation improves model accuracy up to 24.7%, (ii) RAG models benefit significantly more than vanilla LLMs, (iii) simply increasing retrieval depth does not improve RAG performance, and (iv) RAG models achieve high accuracy with perfect information, highlighting the need to filter noisy documents during retrieval. The LongTailQA benchmark facilitates robust evaluation of long-tail knowledge recall and RAG system effectiveness. We make the codebase and datasets publicly available at https://github.com/williamx854/LongTailQA-Benchmark
CRaFT: An Explanation-Based Framework for Evaluating Cultural Reasoning in Multilingual Language Models
Shehenaz Hossain | Haithem Afli
Shehenaz Hossain | Haithem Afli
Correct answers do not necessarily reflect cultural understanding. We introduce CRaFT, an explanation-based multilingual evaluation framework designed to assess how large language models (LLMs) reason across cultural contexts. Rather than scoring outputs solely based on accuracy, CRaFT evaluates model explanations using four interpretable metrics: Cultural Fluency, Deviation, Consistency, and Linguistic Adaptation. We apply the framework to 50 culturally grounded questions from the World Values Survey, translated into Arabic, Bengali, and Spanish, and evaluate three models (GPT-4o, DeepSeek, FANAR) across over 2,100 answer–explanation pairs. Results reveal significant cross-lingual variation in reasoning: Arabic reduces fluency, Bengali enhances it, and Spanish remains largely stable. While GPT-4o adapts more effectively across languages, it exhibits lower consistency; FANAR shows stable but rigid reasoning. These findings suggest that cultural awareness in LLMs is not intrinsic but emerges through linguistic framing. CRaFT offers a new lens for evaluating cross-cultural reasoning in multilingual settings, providing actionable insights for building culturally adaptive language models.
HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
Alexis Correa | Carlos Gómez-Rodríguez | David Vilares
Alexis Correa | Carlos Gómez-Rodríguez | David Vilares
We introduce HEAD-QA v2, an expanded and updated version of a Spanish/English healthcare multiple-choice reasoning dataset originally released by Vilares and Gómez-Rodríguez (2019). The update responds to the growing need for high-quality datasets that capture the linguistic and conceptual complexity of healthcare reasoning. We extend the dataset to over 12,000 questions from ten years of Spanish professional exams, benchmark several open-source LLMs using prompting, RAG, and probability-based answer selection, and provide additional multilingual versions to support future work. Results indicate that performance is mainly driven by model scale and intrinsic reasoning ability, with complex inference strategies obtaining limited gains. Together, these results establish HEAD-QA v2 as a reliable resource for advancing research on biomedical reasoning and model improvement.
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
Hunzalah Hassan Bhatti | Firoj Alam
Hunzalah Hassan Bhatti | Firoj Alam
Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains limited across languages and their varieties. We propose a comprehensive method that (i) translates Modern Standard Arabic (MSA) multiple-choice questions (MCQs) into English and several Arabic dialects, (ii) converts them into open-ended questions (OEQs), (iii) benchmarks a range of zero-shot and fine-tuned LLMs under both MCQ and OEQ settings, and (iv) generates chain-of-thought (CoT) rationales to fine-tune models for step-by-step reasoning. Using this method, we extend an existing dataset in which QAs are parallelly aligned across language varieties, making it, to our knowledge, the first of its kind. A large portion of the resulting test set is further validated through targeted human annotation and native-speaker post-editing. We conduct extensive experiments with both open and closed models. Our findings show that (i) models underperform on Arabic dialects, showing persistent gaps in culturally grounded and dialect-specific knowledge; (ii) Arabic-centric models perform well on MCQs but struggle with OEQs; and (iii) CoT improves judged correctness while yielding mixed n-gram-based metrics.
Automatic Inter-document Multi-hop Scientific QA Generation
Seungmin Lee | Dongha Kim | Yuni Jeon | Junyoung Koh | Min Song
Seungmin Lee | Dongha Kim | Yuni Jeon | Junyoung Koh | Min Song
Existing automatic scientific question generation studies mainly focus on single-document factoid QA, overlooking the inter-document reasoning crucial for scientific understanding. We present AIM-SciQA, an automated framework for generating multi-document, multi-hop scientific QA datasets. AIM-SciQA extracts single-hop QAs using large language models (LLMs) with machine reading comprehension and constructs cross-document relations based on embedding-based semantic alignment while selectively leveraging citation information. Applied to 8,211 PubMed Central papers, it produced 411,409 single-hop and 13,672 multi-hop QAs, forming the IM-SciQA dataset. Human and automatic validation confirmed high factual consistency, and experimental results demonstrate that IM-SciQA effectively differentiates reasoning capabilities across retrieval and QA stages, providing a realistic and interpretable benchmark for retrieval-augmented scientific reasoning. We further extend this framework to construct CIM-SciQA, a citation-guided variant achieving comparable performance to the Oracle setting, reinforcing the dataset’s validity and generality.
CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps
Jungmin Yun | June Hyoung Kwon | Youngbin Kim
Jungmin Yun | June Hyoung Kwon | Youngbin Kim
Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks two critical vulnerabilities: (1) reliance on internal parametric knowledge rather than adherence to the provided context, and (2) exploitation of dataset shortcuts, such as single-document cues or type-matching, that diminish the need for genuine evidence aggregation across multiple documents. We introduce CRiT-QA (Counterfactual Reasoning with Traps), a dataset explicitly designed to address both limitations. To neutralize reliance on memorized knowledge and enforce strict context dependency, CRiT-QA transforms factual reasoning chains with counterfactual entities. Furthermore, it injects multi-anchor distractor chains, plausible but incorrect reasoning paths that diverge at different hops. These traps require models to follow the entire reasoning process rather than exploiting shallow heuristics. Our experiments show that LLMs exhibit substantial performance degradation on CRiT-QA compared to standard datasets, exposing their vulnerability to counterfactual conditions and distractor traps. CRiT-QA thus serves as a rigorous diagnostic tool for evaluating genuine multi-hop reasoning and provides a foundation for developing more reliable, evidence-grounded LLMs.
TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
Reihaneh Iranmanesh | Saeedeh Davoudi | Pasha Abrishamchian | Ophir Frieder | Nazli Goharian
Reihaneh Iranmanesh | Saeedeh Davoudi | Pasha Abrishamchian | Ophir Frieder | Nazli Goharian
This paper presents a comprehensive evaluation framework for assessing the cultural competence of large language models (LLMs) in Persian. Existing Persian cultural benchmarks rely predominantly on multiple-choice formats and English-centric metrics that fail to capture Persian’s morphological complexity and semantic nuance. Our framework introduces a Persian-specific short-answer evaluation that combines rule-based morphological normalization with a hybrid syntactic and semantic similarity module, enabling robust soft-match scoring beyond exact string overlap. Through systematic evaluation of 15 state-of-the-art open- and closed-source models across three culturally grounded Persian datasets, we demonstrate that our hybrid evaluation improves scoring consistency by +10 compared to exact-match baselines by capturing meaning that surface-level methods cannot detect. Our human evaluation further confirms that the proposed semantic similarity metric achieves higher agreement with human judgments than LLM-based judges. We publicly release our evaluation framework, providing the first standardized benchmark for measuring cultural understanding in Persian and establishing a reproducible foundation for cross-cultural LLM evaluation research.
Benchmarking Mathematical Reasoning in a Low-Resource Language: Structured Prompting and Evaluation in Basque
Inigo Martinez-Criado | Aitor Soroa | Jeremy Barnes
Inigo Martinez-Criado | Aitor Soroa | Jeremy Barnes
Large Language Models (LLMs) have shown impressive performance on tasks requiring complex reasoning, but most evaluations tend to focus on English and other high-resource languages. This work investigates how well LLMs perform mathematical reasoning in low-resource languages, using Basque as a primary case study. To support this analysis, we introduce MASEU, a benchmark designed to evaluate reasoning in Basque across arithmetic, algebraic, and logical tasks. We then use this dataset to address three key questions: 1) how well do LLMs support Basque in reasoning tasks, 2) to what extent can including English in prompts improve results, and 3) what is the effect of continued pretraining in Basque? To explore these aspects, we use prompting strategies adapted for mathematical reasoning, building upon the foundations of CoT prompting and one of its subsequent evolutions, DUP prompting, which together allow for more precise experimentation across zero-shot and few-shot settings, providing insights into how multilingual models handle reasoning tasks in underrepresented languages.
Assessing the Difficulty of Inference Types in Natural Language Inference for Clinical Trials
Mathilde Aguiar | Pierre Zweigenbaum | Nona Naderi
Mathilde Aguiar | Pierre Zweigenbaum | Nona Naderi
Large Language Models (LLMs) achieve competitive results on Natural Language Inference when applied to clinical trials; however, it is not yet clear which type of inference LLMs perform well or poorly on. We address this by proposing new supplementary annotations for the existing NLI4CT dataset on the types of inferences observed in clinical trials. Our dataset supplements NLI4CT with a total of 1,949 new annotations using our carefully crafted guidelines for 17 types of inferences. To investigate how inference types affect the performance of LLMs, we prompt Flan-T5, Llama, Mistral, and Qwen and evaluate their performance using our newly annotated dataset. We found that logical inferences negatively affect the overall performance of Qwen3-4B, Qwen2.5-7B, and Qwen2.5-14B, whereas numerical inferences negatively affect the performance of Flan-T5-XL and Mixtral. Further analysis shows that MMed-Llama-3 struggles to understand the structure of clinical trial reports. Other parameters, such as the number of inference types involved or the section type in the premise, also influence the performance of the models. Our code and dataset are publicly available.
Reasoning Graph-Structured Question Answering: Datasets and Insights from LLM Benchmarking
Khin Yone | Devasha Trivedi | Anish Pahilajani | Jincen Shuai | Samyak Rajesh Jain | Ryan Rossi | Nesreen K. Ahmed | Franck Dernoncourt | Yu Wang | Namyong Park
Khin Yone | Devasha Trivedi | Anish Pahilajani | Jincen Shuai | Samyak Rajesh Jain | Ryan Rossi | Nesreen K. Ahmed | Franck Dernoncourt | Yu Wang | Namyong Park
Large Language Models (LLMs) have shown remarkable success in multi-hop question-answering (M-QA) due to their advanced reasoning capabilities. However, the influence of reasoning structures on their performance remains underexplored, primarily due to the lack of M-QA datasets that explicitly encode the reasoning pathways underlying each question-answer pair. To address this gap, we introduce the reasoning graph-structured question answering dataset (GRS-QA), which provides both semantic contexts and reasoning structures for the QA pairs. Unlike existing M-QA datasets, GRS-QA explicitly captures intricate reasoning pathways through reasoning graphs, where nodes correspond to textual contexts and edges denote logical flows. Using GRS-QA, we systematically evaluate LLM performance across varying context structures, prompting styles, and data domains. Our empirical analysis reveals that LLMs perform differently based on the reasoning structure, context, and prompting styles, indicating their varying ability to leverage graph-structured knowledge. Notably, providing explicit reasoning guidance proves more effective than supplying contextual information alone.
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
Zhihan Cao | Fumihito Nishino | Hiroaki Yamada | Ha Thanh Nguyen | Yusuke Miyao | Ken Satoh
Zhihan Cao | Fumihito Nishino | Hiroaki Yamada | Ha Thanh Nguyen | Yusuke Miyao | Ken Satoh
We introduce JBE-QA, a Japanese Bar Exam Question–Answering dataset to evaluate large language models’ legal knowledge. Derived from the multiple-choice (tantō-shiki) section of the Japanese bar exam (2015–2024), JBE-QA provides the first comprehensive benchmark for Japanese legal-domain evaluation of LLMs. It covers the Civil Code, the Penal Code, and the Constitution, extending beyond the Civil Code focus of prior Japanese resources. Each question is decomposed into independent true/false judgments with structured contextual fields. The dataset contains 3,464 items with balanced labels. We evaluate 26 LLMs, including proprietary, open-weight, Japanese-specialised, and reasoning models. Our results show that proprietary models with reasoning enabled perform best, and the Constitution questions are generally easier than the Civil Code or the Penal Code questions.
Many Swedish benchmarks are translations of US-centric benchmarks and are therefore not suitable for testing knowledge that is particularly relevant, or even specific, to Sweden. We therefore introduce a manually written question-answering benchmark specifically targeted at Sweden-related personalities and events, many of which receive very limited coverage in international media. Our annotators drew inspiration from a popular radio program featuring public figures from culture and media, as well as major sports events in Sweden. The dataset can be used to measure factual recall across models of varying sizes and degrees of Swedish coverage, and allows probing of cross-lingual factual consistency, as it contains English translations. Using the dataset, we find that smaller models with stronger Swedish coverage perform comparably to a multilingual model three times larger in recalling Sweden-related facts. We also observe that continued pre-training on Swedish generally improves factual knowledge but leads to partial forgetting of previously known information. These results demonstrate the dataset’s potential as a diagnostic tool for studying language adaptation and knowledge retention in multilingual models during language adaptation.
GeoBenchmark: Probing Large Language Models for Geo-Spatial Knowledge
Ayomide Abayomi | Jose G. Moreno | Karim Radouane | Lynda Tamine
Ayomide Abayomi | Jose G. Moreno | Karim Radouane | Lynda Tamine
Large Language Models (LLMs) demonstrate strong factual recall of general-purpose knowledge but struggle with grounded geospatial knowledge. To measure and help probe LLMs for spatial knowledge, we present GeoBenchmark, a benchmark for evaluating geographic commonsense along three core spatial relations: direction, distance, and topology. Using data extracted from YAGO2geo and Ordnance Survey ward geometries, spatial relations were formalized as structured triplets and systematically transformed into balanced binary (Yes/No) and Multiple-Choice (MCQ) question-answer pairs. Besides, we consider atomic and composite questions based on the number of spatial relations involved. The resulting dataset comprises 26k binary and 13k MCQ samples, uniformly distributed across atomic, binary, and ternary relation levels. We establish baselines with LLaMA-8B and Mistral-7B under zero-shot prompting, achieving 52-63% accuracy on atomic questions but below 35% on ternary relations, which exposes the models’ limited compositional spatial understanding and strong option bias. GeoBenchmark provides a comprehensive, reproducible resource for probing and advancing LLMs’ geographic commonsense, paving the way for future research in spatial and geographic probing of LLMs as well as knowledge editing.
FactOReS: Fact-checking with an Evidence-based Open Resource in Spanish
Nagore Bravo | Jaione Bengoetxea | Iker García-Ferrero | Alba Bonet Jover | Estela Saquete | Rodrigo Agerri
Nagore Bravo | Jaione Bengoetxea | Iker García-Ferrero | Alba Bonet Jover | Estela Saquete | Rodrigo Agerri
Automated Fact-Checking (AFC) has become a popular research area in Natural Language Processing (NLP), intending to support human verification through evidence-based veracity prediction systems that provide transparency at each stage of the process. Despite the global significance of misinformation and the substantial progress made in AFC research, multilingual approaches to evidence-based fact-checking remain inadequately addressed. This work introduces FactOReS, the first publicly available dataset evaluated for evidence-based veracity prediction in Spanish, constructed from real Spanish-language claims and verified fact-checking articles. We establish performance baselines by systematically applying In-Context Learning (ICL) with Large Language Models (LLMs) to both an established English dataset and our novel Spanish dataset. Despite good zero-shot and few-shot performance, results in both languages demonstrate that each step requires further research in order to improve the overall results in the evidence-based veracity prediction task. Finally, we propose a semi-automated methodology that integrates computational processing with human validation, offering a reproducible framework for developing multilingual evidence-based fact-checking resources for the benefit of the NLP research community. Data and code available: https://github.com/hitz-zentroa/AFC_FactOReS
Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
Dylan Phelps | Rodrigo Wilkens | Edward Gow-Smith | Thomas M. R. Pickard | Maggie Mi | Marco Idiart | Aline Villavicencio
Dylan Phelps | Rodrigo Wilkens | Edward Gow-Smith | Thomas M. R. Pickard | Maggie Mi | Marco Idiart | Aline Villavicencio
The recent trend towards utilisation of reasoning models has improved the performance of Large Language Models (LLMs) across many tasks which involve logical steps. One linguistic task that could benefit from this framing is idiomaticity detection, as a potentially idiomatic expression must first be understood in relation to the context before it can be disambiguated. In this paper, we explore how reasoning capabilities in LLMs affect idiomaticity detection performance and examine the effect of model size. We evaluate, as open source representative models, the suite of DeepSeek-R1 distillation models ranging from 1.5B to 70B parameters across four idiomaticity detection datasets. We find the effect of reasoning to be smaller and more varied than expected. For smaller models, producing chain-of-thought (CoT) reasoning increases performance from Math-tuned intermediate models, but not to the levels of the base models, whereas larger models (14B, 32B, and 70B) show modest improvements. Our in-depth analyses reveal that larger models demonstrate good understanding of idiomaticity, successfully producing accurate definitions of expressions, while smaller models often fail to output the actual meaning. For this reason, we also experiment with providing definitions in the prompts of smaller models, which we show can improve performance in some cases.
ESG-QA: Building a Dataset for Question Answering on Environmental, Social, and Governance Pillars
Gabriel Assis | Ayrton Surica | Pedro Kroll | Gabriela Aires Mendes | Darian Rabbani | Edson Bollis | Lucas Francisco Amaral Orosco Pellicer | Aline Paes
Gabriel Assis | Ayrton Surica | Pedro Kroll | Gabriela Aires Mendes | Darian Rabbani | Edson Bollis | Lucas Francisco Amaral Orosco Pellicer | Aline Paes
Environmental, Social, and Governance (ESG) factors are becoming increasingly central to corporate accountability and sustainable development. However, benchmarks for evaluating large language models (LLMs) in this domain remain scarce. To alleviate this gap, we present ESG-QA, a dataset of 87,261 question–answer–context triplets spanning the three ESG pillars. ESG-QA was built using an LLM-based Question Answer (QA) generation pipeline, enhanced through rule-based and semantic filtering, and validated by human inspection, enabling both abstractive QA and retrieval-augmented setups. We benchmark three open-weight LLM families (Llama-3, Gemma-3, and Qwen-3) across multiple dimensions, including correctness, environmental impact, and readability. Results show that Qwen-3 with retrieval achieves the highest absolute QA performance, while Gemma-3 provides the strongest overall balance between correctness, efficiency, and clarity. By releasing ESG-QA and its generation framework, this work establishes a comprehensive benchmark for advancing ESG-oriented QA and promoting more transparent and responsible AI evaluation.
Enhancing and Evaluating Tabular Models on the Fly via Synthetic Question–Answer Generation
Jorge Osés Grijalba | Eugenio Martínez Cámara | L. Alfonso Ureñ-López | Jose Camacho-Collados
Jorge Osés Grijalba | Eugenio Martínez Cámara | L. Alfonso Ureñ-López | Jose Camacho-Collados
Question Answering (QA) over Tabular Data has been traditionally a challenging task, but LLMs have recently shown the ability to respond to questions related to this type of structured data. However, current tabular QA datasets are skewed toward Wikipedia tables and SQL-style answers composed of human-crafted question–answer pairs. This limits the evaluation of LLMs on this task to a narrow genre of data and language, while also requiring extensive human effort for dataset or benchmark creation. To address this, we introduce SynTabQA, a methodology for the automatic generation of synthetic question–answer pairs from any unannotated table. SynTabQA defines a detailed question typology, enabling fine-grained evaluation and facilitating the creation of diverse QA datasets. Our approach not only provides an automated test bed for any tabular dataset but can also be used in few-shot settings to supply LLMs with tailored examples, improving their focus and accuracy. We validate SynTabQA on two large, manually constructed tabular QA benchmarks of distinct nature.
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
Tu Tran Do | Nhat Ngoc Nguyen | Tung Khanh Tran | Hoang D. Nguyen | Tu Minh Phuong | Long Hoang Dang
Tu Tran Do | Nhat Ngoc Nguyen | Tung Khanh Tran | Hoang D. Nguyen | Tu Minh Phuong | Long Hoang Dang
We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. VIVID comprises 1,636 idioms and proverbs annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowledge) and seven semantic themes. We establish an evaluation framework combining generative and discriminative tasks, proposing an LLM-as-a-Judge approach with aspect-based prompting validated against human judgment (Cohen’s κ = 0.792). Evaluating eight state-of-the-art models reveals critical gaps: Vietnamese-specialized models drastically underperform multilingual systems (VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46), and even top models achieve less than 50% of maximum scores. Notably, few-shot prompting does not universally improve performance, with GPT-4o exhibiting degradation due to stylistic overfitting. Our analysis exposes systematic failures including literal over-interpretation, lexical gaps, and pragmatic flattening, demonstrating that current models lack cultural competence for nuanced figurative interpretation. VIVID provides an essential tool for advancing figurative language understanding in culturally rich contexts.
Assessing Logical Coherence of LLMs via Fine-Grained NLI
Jon Felix Apaolaza Larraya | Begoña Altuna | Aitor Soroa | Inigo Lopez-Gazpio
Jon Felix Apaolaza Larraya | Begoña Altuna | Aitor Soroa | Inigo Lopez-Gazpio
Natural Language Inference (NLI) is a long-standing probe of models’ reasoning capabilities, yet it remains unclear how state-of-the-art systems represent and combine logical clauses in a way that supports robust generalization. We study directional effects in deductive NLI and introduce causal coherence, an evaluation paradigm that tests whether predictions remain consistent when the directionality of inference is reversed. Using fine-grained minimal-pair phrase data from PhrasIS, we evaluate encoder, decoder, and encoder–decoder transformers and analyze their behavior under both standard and manipulated settings. Our results show that models frequently fail to maintain logical stability when directionality varies, indicating shallow pattern matching rather than genuine clause composition. We formalize soft and hard causal coherence to disentangle directional consistency from correctness, and we provide an error analysis that highlights systematic failures involving semantic relations. Our findings suggest that deductive causal reasoning and coherence remain missing components in current transformer architectures, and that addressing them is necessary for reliable NLI.
Counter-Hypothesis Generation: Towards Evaluating How LLMs Reason about Alternatives
Marzieh Abdolmaleki | Aaron Maladry | Veronique Hoste | Els Lefever
Marzieh Abdolmaleki | Aaron Maladry | Veronique Hoste | Els Lefever
Reasoning about alternatives is a fundamental component of human cognition and argumentation, yet it remains unclear whether large language models (LLMs) can coherently generate and assess them. This paper introduces Counter-Hypothesis Generation (CHG), a novel task for evaluating how LLMs construct plausible hypotheses when contextual information changes. Inspired by open-domain commonsense reasoning, where models infer and compare multiple explanations, CHG bridges commonsense and counterfactual reasoning by requiring models to generate hypotheses that remain logically consistent with modified premises. We present a test set annotated by a human expert and complemented with counter-hypotheses generated by OpenAI-o3 and DeepSeek-r1. Experimental results reveal that even advanced reasoning models exhibit notable limitations in counter-hypothesis generation.
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
Rafid Ishrak Jahan | Fahmid Shahriar Iqbal | Sagnik Ray Choudhury
Rafid Ishrak Jahan | Fahmid Shahriar Iqbal | Sagnik Ray Choudhury
Long-form question answering (LFQA) demands nuanced evaluation of multi-sentence explanatory responses, yet existing metrics often fail to reflect human judgment. We present LFQA-HP-1M, a large-scale dataset comprising 1.3M human pairwise preference annotations for LFQA. We propose nine rubrics for answer quality evaluation, and show that simple linear models based on these features perform comparably to state-of-the-art LLM evaluators. We further examine transitivity consistency, positional bias, and verbosity biases in LLM evaluators and demonstrate their vulnerability to adversarial perturbations. Overall, this work provides one of the largest public LFQA preference datasets and a rubric-driven framework for transparent and reliable evaluation.
Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models
Bryan E. Tuck | Rakesh Verma
Bryan E. Tuck | Rakesh Verma
Large language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains limited. We evaluate 39 configurations spanning three model families (Qwen3, Claude Haiku 4.5, GPT-5-mini) on 58 word puzzles requiring character-level constraint satisfaction. Cross-family differences produce substantially larger performance gaps (2.0–2.2×, F1 = 0.761 vs. 0.343) than parameter scaling within families (83% gain from 4B to 32B scaling), and a partial-correlation analysis rules out tokenizer design as a confound for within-family scaling. Thinking budget sensitivity proves heterogeneous: high-capacity models show strong returns (+0.102 to +0.136 F1), while mid-sized variants saturate or degrade, showing inconsistent compute benefits. Using difficulty ratings from 10,000 human solvers per puzzle, we establish modest but consistent calibration (ρ = 0.28–0.42) across all families, yet identify systematic failures on common words with unusual orthography (“data”, “loll”, “acai”: 83–91% human success, 94–98% model miss rate). These failures point to over-reliance on distributional plausibility that penalizes orthographically atypical but constraint-valid patterns.
LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation
Koki Itai | Shunichi Hasegawa | Yuta Yamamoto | Gouki Minegishi | Masaki Otsuki
Koki Itai | Shunichi Hasegawa | Yuta Yamamoto | Gouki Minegishi | Masaki Otsuki
Retrieval-Augmented Generation (RAG) is a framework in which a Generator, such as a Large Language Model (LLM), produces answers by retrieving documents from an external collection using a Retriever. In practice, Generators must integrate evidence from long contexts, perform multi-step reasoning, interpret tables, and abstain when evidence is missing. However, existing benchmarks for Generators provide limited coverage, with none enabling simultaneous evaluation of multiple capabilities under unified conditions. To bridge the gap between existing evaluations and practical use, we introduce LIT-RAGBench (the Logic, Integration, Table, Reasoning, and Abstention RAG Generator Benchmark), which defines five categories: Integration, Reasoning, Logic, Table, and Abstention—each further divided into practical evaluation aspects. LIT-RAGBench systematically covers patterns combining multiple aspects across categories. By using fictional entities and scenarios, LIT-RAGBench evaluates answers grounded in the provided external documents. The dataset consists of 114 human-constructed Japanese questions and an English version generated by machine translation with human curation. We use LLM-as-a-Judge for scoring and report category-wise and overall accuracy. Across API-based and open-weight models, no model exceeds 90% overall accuracy. By making strengths and weaknesses measurable within each category, LIT-RAGBench serves as a valuable metric for model selection in practical RAG deployments and for building RAG-specialized models.
Analyses of hypothesis generation in fictionalised environments have significant potential for exploring factors influencing reasoning and decision-making in naturalistic contexts. Based on transcripts of 16 groups playing a murder mystery game, with a total of 42 human participants, RIP2 is a 177,000 word corpus exemplifying reasoning in the forensic domain. With a 80,000 word representative sample of the corpus annotated using an argumentation framework, RIP2 is nearly twice the size of the RIP Corpus of Collaborative Hypothesis-Making (RIP1), currently the only existing corpus of hypothesis-making in group environments. With an new experimental set-up and guidelines for annotating both cases of hypothesising and conjecturing, RIP2 offers insight into how participants generate, maintain, and reject hypotheses, as well as how they interact with others’ contributions. Based on its close exploration of six groups (three successful), this corpus particularly allows for group-level comparisons of factors influencing group success. Within this paper, we discuss the main contributions for understanding hypothesising and collaborative reasoning, and offer use cases for extended work demonstrating how analysis of hypothesis generation can be used for future research on argumentation quality and decision-making.
Can Multimodal LLMs Generate Pedagogical Questions?
Thomas Gerald | Sahar Ghannay | Julie Lascar | Paul Lerner | Anne Vilnat
Thomas Gerald | Sahar Ghannay | Julie Lascar | Paul Lerner | Anne Vilnat
Educational materials frequently combine text, diagrams, tables, and charts to convey complex concepts. Understanding such materials often requires reasoning across modalities rather than relying solely on textual descriptions. In educational contexts, the main challenge lies in assessing the relevance and quality of the questions themselves. This raises a key issue: what defines a good question in a specialized learning environment? By comparison, evaluating answers is a more conventional task, although it requires examining criteria consistent with the targeted educational level. To the best of our knowledge, the use of LLMs for assessing the pedagogical relevance of questions remains unexplored. This gap highlights the need to define pedagogical relevance more clearly and to investigate the consistency of LLM judgments, as well as their alignment with human evaluations. We introduce a new Multimodal QA dataset in the education domain. To reduce the need for extensive human annotation, we leverage LLMs to help design questions on educational material, jointly with a human annotation. Contrary to most of QA Multimodal corpora, we focus on questions that could be asked by a teacher in his/her class, and that need dealing with different parts of the document to be answered. Results show that while LLMs as a judge is an efficient framework, many problem could arise and that align prediction with human annotators is a difficult task for complex criteria.
The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs Using Indian Riddles
Abhinav P M | Ojasva Saxena | Oswald C | Parameswari Krishnamurthy
Abhinav P M | Ojasva Saxena | Oswald C | Parameswari Krishnamurthy
The extent to which large language models (LLMs) can perform culturally grounded reasoning across non-English languages remains underexplored. This paper examines the reasoning and self-assessment abilities of LLMs across seven major Indian languages- Bengali, Gujarati, Hindi, Kannada, Malayalam, Tamil, and Telugu. We introduce a multilingual riddle dataset combining traditional riddles with context-reconstructed variants and evaluate five LLMs- Gemini 2.5 Pro, Gemini 2.5 Flash, Mistral-Saba, LLaMA-4-Scout, and LLaMA-4-Maverick under seven prompting strategies. In the first stage, we assess riddle-solving performance and find that while Gemini 2.5 Pro performs best overall, few-shot methods yield only marginal gains, and accuracy varies notably across languages. In the second stage, we conduct a self-evaluation experiment to measure reasoning consistency. The results reveal a key finding: a model’s initial accuracy is inversely correlated with its ability to identify its own mistakes. Top-performing models such as Gemini 2.5 Pro are overconfident (4.34% True Negative Rate), whereas lower-performing models like LLaMA-4-Scout are substantially more self-aware (42.09% True Negative Rate). These results point to clear gaps in multilingual reasoning and highlight the need for models that not only reason effectively but also recognize their own limitations.
Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the scarcity of transcribed corpora. This proof-of-concept study explores songs as an unconventional yet promising data source for Kazakh ASR. We curate a dataset of 3,013 audio-text pairs (about 4.5 hours) from 195 songs by 36 artists, segmented at the lyric-line level. Using Whisper as the base recogniser, we fine-tune models under seven training scenarios involving Songs, Common Voice Corpus (CVC), and FLEURS, and evaluate them on three benchmarks: CVC, FLEURS, and Kazakh Speech Corpus 2 (KSC2). Results show that song-based fine-tuning improves performance over zero-shot baselines. For instance, Whisper Large-V3 Turbo trained on a mixture of Songs, CVC, and FLEURS achieves 27.6% normalised WER on CVC and 11.8% on FLEURS, while halving the error on KSC2 (39.3% vs. 81.2%) relative to the zero-shot model. Although these gains remain below those of models trained on the 1,100-hour KSC2 corpus, they demonstrate that even modest song-speech mixtures can yield meaningful adaptation improvements in low-resource ASR. The dataset is released on Hugging Face for research purposes under a gated, non-commercial licence.
This article introduces a dedicated speech recognition dataset for Southern Kurdish, which is a threatened variant of Kurdish macrolanguage. We present 30 hours of validated read speech for training and an evaluation benchmark for Southern Kurdish Automatic Speech Recognition (ASR). Both the training data and evaluation benchmark are read speech recorded by crowdsourcing campaigns. Besides a detailed description of the provided resources, we provide the ASR baselines using Whisper-turbo and wav2vec-bert CTC architectures. We achieved a 4.09 CER and 24.26 WER on our benchmark using wav2vec-bert model. We also provide a categorization of errors to support further improvements in future studies.The resources and trained models are released under the CC BY-NC-ND 4.0 license and are publicly available at https://huggingface.co/datasets/aranemini/southern-kurdish-asr
MASA: A Novel Multimodal Foundation Model for L2 Speaking Assessment in Picture-description Scenarios
Bi-Cheng Yan | Fu-An Chao | Hong-Yun H.Y. Lin | Berlin Chen
Bi-Cheng Yan | Fu-An Chao | Hong-Yun H.Y. Lin | Berlin Chen
Automatic speaking assessment (ASA) manages to quantify the language competence of second language (L2) learners by providing a proficiency score based on their spoken responses. Existing efforts typically employ a neural grader coupled with a set of handcrafted features to gauge the competence of language in L2 learners from multiple facets. Despite their decent efficacy, these methods are limited by a laborious feature engineering process and largely overlook the utilization of scoring rubrics that are presented to human raters in speaking assessment. In light of this, we put forward a novel Multimodal foundation model for ASA, termed MASA, for use in picture-description scenarios. Our approach effectively streamlines the feature engineering process by leveraging the pre-trained encoders of a multimodal foundation model, and emulates the nuanced scoring behaviors of human raters by incorporating scoring rubrics directly into the modeling process. Furthermore, a simple, training-free method is introduced to alleviate the scoring bias in MASA by contrasting the output distributions derived from the multimodal and single-modal inputs. A series of experiments conducted on a picture-description task of the General English Proficiency Test (GEPT) dataset validates the feasibility and superiority of our method in comparison to several cutting-edge baselines.
Tools for Estimating the Perceived Level of Phonetic Reduction
Nigel Ward | Javier Vazquez | Emma (Danny) R. Boushka | Oliver Niebuhr
Nigel Ward | Javier Vazquez | Emma (Danny) R. Boushka | Oliver Niebuhr
Phonetic reduction is very common in casual speech, where it is associated with several important pragmatic functions. However these phenomena have been little studied. To support investigations and applications, we present tools that automatically estimate the level of perceived phonetic reduction. Trained on annotated dialog data and exploiting HuBert features, these handle American English and Northern Mexican Spanish. For English, word-level predictions correlate up to 0.55 with average human judgments. This is adequate at least for statistical studies of reduction in corpora, as seen in explorations of turn-yielding and prominence-marking behaviors. The tools are open-source and publicly available
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
Francisco Teixeira | Carlos Carvalho | Mariana Julião | Catarina Botelho | Rubén Solera-Ureña | Sérgio Paulo | Thomas Rolland | Ben Peters | Isabel Trancoso | Alberto Abad
Francisco Teixeira | Carlos Carvalho | Mariana Julião | Catarina Botelho | Rubén Solera-Ureña | Sérgio Paulo | Thomas Rolland | Ben Peters | Isabel Trancoso | Alberto Abad
State-of-the-art performance for Automatic Speech Recognition (ASR) largely depends on the availability of large-scale labeled corpora. This creates a demand for increased data collection efforts, particularly for under-represented languages and dialectal varieties. Due to having considerably fewer speakers (around 11 million), European Portuguese (EP) is overshadowed by Brazilian Portuguese (BP) (around 200 million speakers) in currently available large-scale speech data resources, resulting in under-performing speech-based systems for EP users. To address this gap, and following similar data collection efforts for other languages, we present FalAR, a large-scale, speaker-annotated speech corpus of European Portuguese parliamentary sessions. Spanning approximately 20 years, FalAR comprises 5,800 hours of speech data. In addition, 4,850 hours have speaker identity annotations, for a total of 1,180 speakers with associated metadata including age, gender, political affiliation, and parliamentary role. The corpus was built using a state-of-the-art EP CAMÕES ASR model for transcription-reference alignment. In this paper, we describe the data collection process, together with the main characteristics of the FalAR corpus. Furthermore, we evaluate the trade-off between data quantity and alignment accuracy on ASR performance, with our experiments demonstrating that incorporating FalAR as pre-training data yields up to 14% relative WER improvement over baseline models.
English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
Mohammad Mohammadamini | Daban Jaff | Josep Crego | Marie Tahon | Antoine LAURENT
Mohammad Mohammadamini | Daban Jaff | Josep Crego | Marie Tahon | Antoine LAURENT
We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40 million Central Kurdish tokens. We evaluate KUTED on the S2TT task and find that orthographic variation significantly degrades Kurdish translation performance, producing nonstandard outputs. To address this, we propose a systematic text standardization approach that yields substantial performance gains and more consistent translations. On a test set separated from TED talks, a fine-tuned Seamless model achieves 15.18 BLEU, and we improve Seamless baseline by 3.0 BLEU on the FLEURS benchmark. We also train a Transformer model from scratch and evaluate a cascaded system that combines Seamless (ASR) with NLLB (MT).
Automatic Prediction of Prominence and Boundary Strength from Text
Pauline Mas | Kévin Vythelingum | Jonathan Chevelu | Marion Ouédraogo | Damien Lolive | Olivier Rosec
Pauline Mas | Kévin Vythelingum | Jonathan Chevelu | Marion Ouédraogo | Damien Lolive | Olivier Rosec
In Text-to-Speech synthesis (TTS), the prediction of prosodic information from text is a difficult challenge, since it requires information related to the context that may not be present in the text. Previous studies have shown that prosodic annotations from an oracle benefit TTS models and improve their prosodic rendering as well as their controllability. In this paper, we investigate different strategies to automatically predict prominence and boundary strength from text. We compare three prediction strategies on a French audiobook dataset: dedicated predictors jointly trained in a TTS model, a BERT-informed Prosody Predictor (BIPP) and its auto-regressive counterpart, both benefiting from semantic text embeddings. BIPP exhibits the best performance in our experiments, indicating that using phonetized syllables as complementary information to the semantic embedding provided by a BERT-like model is the best strategy to predict prosodic events.
SOMVOICE: A First Dataset to Study the Effects of Sleep Deprivation on Voice Characteristics of Healthy French Speakers
Vincent P. Martin | Jean-Luc Rouas | Colleen Beaumard | Pierre Philip
Vincent P. Martin | Jean-Luc Rouas | Colleen Beaumard | Pierre Philip
Excessive sleepiness is a significant public health issue and a critical personal health indicator associated with various disorders. Given its high prevalence in the general population, clinicians need tools to regularly measure patients’ sleepiness levels in natural settings, such as automatic speech analysis. In this article, we introduce the SOMVOICE corpus, the first French corpus containing read-speech recordings from the same participants either after a normal night or after a night of total sleep deprivation. Participants were included according to strict inclusion and exclusion criteria based on both medical characteristics and reading proficiency. The recordings were labelled with both objective and subjective measures of sleepiness, as well as fatigue and anxiety. After introducing the data-collection methodology, we use linear mixed models to conduct a preliminary investigation of the effect of total sleep deprivation on the collected sleepiness-related measures and on participants’ reading behaviour. Doing so, we found that sleep deprivation strongly influences objective and subjective sleepiness measurements as well as fatigue self-reports, but has a lesser effect on anxiety. Regarding reading behaviour, sleep deprivation is associated with a lower speech rate (duration of the recordings and phoneme rate) and more pauses (number of pauses and pause ratio)
Automatic Prediction of Child Speech Fluency with Game-Based Data from German Preschoolers
Valentin Kany | Bernd Möbius | Jürgen Trouvain
Valentin Kany | Bernd Möbius | Jürgen Trouvain
This paper introduces an approach to automatically predict the speech fluency of preschool children as part of Language Proficiency Assessments. We use spontaneous speech data from children with German as native and second language aged 4–6 years, collected via a game–based elicitation method. The recordings were mainly annotated manually on various fluency-related phenomena. The resulting feature values were compared to human fluency ratings of the same data. The human ratings and the fluency-related acoustic features were used to build Cumulative Link Mixed Models (CLMMs) with and without splines to test their ability to predict the human ratings with multiple metrics (Spearman’s ρ, MAE, quadratic weighted κ). Results show that a parsimonious linear model already reaches near-human agreement (quadratic weighted kappa κ = 0.65) and that incorporating non-linear spline effects does not improve predictive accuracy. These findings suggest that relatively simple CLMMs can substitute additional human raters in fine-grained fluency assessment of preschool children, which is a task that is already challenging for trained listeners.
Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping
Tobias Bystrich | Julia Maria Pritzen | Christoph Andreas Schmidt | Claudia Wich-Reif
Tobias Bystrich | Julia Maria Pritzen | Christoph Andreas Schmidt | Claudia Wich-Reif
In the field of universal automatic phonetic transcription (APT), clean and diverse training transcriptions are required. However, such high-quality data is limited. We propose the bootstrapping approach Selective Augmentation to improve the available training transcriptions by selectively transferring distinctions between languages. Based on the model MultIPA, we exemplarily show that we could increase the accuracy of an existing feature (plosive voicing) and add a new feature (plosive aspiration) by augmenting the existing training data using information from a separate helper language (Hindi). We describe intrinsic challenges of the evaluation and develop objective metrics to determine the success: Voicing accuracy was increased by 17.6% by reducing the number of false positives. Additionally, aspiration recognition was introduced: While the baseline transcribed 0% of German /p, t, k/ as aspirated, our approach transcribed them as aspirated in 61.2% of the cases. Introducing aspiration recognition to APT models allowed for the tenuis class to be successfully reduced by 32.2%, which also reduces the conflations between the test language’s plosives.
AURORA Model of Formant-to-tongue Inversion for Didactic and Clinical Applications
Patrycja Strycharczuk | Sam Kirkham
Patrycja Strycharczuk | Sam Kirkham
This paper outlines the conceptual and computational foundations of the AURORA (Acoustic Understanding and Real-time Observation of Resonant Articulations) model. AURORA predicts tongue displacement and shape in vowel sounds based on the first two formant values. It is intended as a didactic aid helping to explain the relationship between formants and the underlying articulation, as well as a foundation for biofeedback applications. The model is informed by ultrasound tongue imaging and acoustic data from 40 native speakers of English. In this paper we discuss the motivation for the model, the modelling objectives as well as the model architecture. We provide a qualitative evaluation of the model, focusing on selected tongue features. We then present two tools developed to make the model more accessible to a wider audience, a Shiny app and a prototype software for real-time tongue biofeedback. Potential users include students of phonetics, linguists in fields adjacent to phonetics, as well as speech and language therapy practitioners and clients.
Investigating the Role of Synthetic Data Augmentation and Training Strategies on Improving Low-Resource Language ASR
Yun Hao | Reihaneh Amooie | Wietse de Vries | Rik van Noord | Martijn Wieling
Yun Hao | Reihaneh Amooie | Wietse de Vries | Rik van Noord | Martijn Wieling
Low-resource automatic speech recognition (ASR) is challenging due to a scarcity of annotated data. While synthetic data from text-to-speech (TTS) systems can augment ASR training, its efficacy for low-resource languages remains unclear. In this study, we investigate under which conditions TTS-based data augmentation is most effective for low-resource languages. Experiments on six low-resource languages in Common Voice show that synthetic data is most beneficial under extremely low-resource ASR conditions (i.e., less than one hour of available real speech data), or for languages with larger amounts of TTS data (i.e., more than 10 hours). Additionally, increasing the amount and diversity of synthetic data while keeping an appropriate ratio of synthetic-to-real data can further improve ASR performance.
AutoRPT: A Tool for Bootstrapping Prosodic Annotation
Seth Heiney | Thomas Hicks | Sally Little | Fernanda Lourenco | Kai Retana | Eliana Stevens | Jonathan Howell
Seth Heiney | Thomas Hicks | Sally Little | Fernanda Lourenco | Kai Retana | Eliana Stevens | Jonathan Howell
Automated Rapid Prosody Transcription (AutoRPT) is a tool for bootstrapping manual annotation of prosodic events in either corpora or standalone audio files using the Rapid Prosody Transcription (RPT) scheme. It functions by utilizing two Long-Short Term Memory (LSTM) models, trained on measures of pitch/F0 and intensity. In addition to discrete, slightly over-generated predictions of prominence and boundary, AutoRPT produces continuous predictions between 0 and 1, similar to crowd-sourced RPT annotations averaged over listeners. Marginal predictions above a given threshold are also indicated discretely by question marks, as in the PoLaR Annotation Guidelines. Annotators achieved a statistically significant increase in annotation speed by modifying AutoRPT-generated annotations over creating annotations without assistance. In contrast with older tools such as AuToBI (Rosenberg, 2010), AutoRPT generates more theory-agnostic annotations which can support the work of non-expert annotators, and which we expect will offer greater flexibility in the prosodic annotation of other English language varieties.
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
Wataru Nakata | Kentaro Seki | Hitomi Yanaka | Yuki Saito | Shinnosuke Takamichi | Hiroshi Saruwatari
Wataru Nakata | Kentaro Seki | Hitomi Yanaka | Yuki Saito | Shinnosuke Takamichi | Hiroshi Saruwatari
Spoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However, existing datasets are often limited in size, spontaneity, or linguistic coherence. To address these limitations, we introduce J-CHAT, a 76,000-hour open-source Japanese spoken dialogue corpus. Constructed using an automated, language-independent methodology, J-CHAT ensures acoustic cleanliness, diversity, and natural spontaneity. The corpus is built from YouTube and podcast data, with extensive filtering and denoising to enhance quality. Experimental results with generative spoken dialogue language models trained on J-CHAT demonstrate its effectiveness for SDS development. By providing a robust foundation for training advanced dialogue models, we anticipate that J-CHAT will drive progress in human-AI dialogue research and applications.
ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & Benchmark
Tung X. Nguyen | Nhu Vo | Giang Son Nguyen | Duy Mai Hoang | Chien Dinh Huynh | Inigo Jauregi Unanue | Massimo Piccardi | Wray Buntine | Dung D. Le
Tung X. Nguyen | Nhu Vo | Giang Son Nguyen | Duy Mai Hoang | Chien Dinh Huynh | Inigo Jauregi Unanue | Massimo Piccardi | Wray Buntine | Dung D. Le
Code-switching (CS), which is when Vietnamese speech uses English words like drug names or procedures, is a common phenomenon in Vietnamese medical communication. This creates challenges for Automatic Speech Recognition (ASR) systems, especially in low-resource languages like Vietnamese. Current most ASR systems struggle to recognize correctly English medical terms within Vietnamese sentences, and no benchmark addresses this challenge. In this paper, we construct a 34-hour Vietnamese Medical Code-Switching Speech dataset (ViMedCSS) containing 16,576 utterances. Each utterance includes at least one English medical term drawn from a curated bilingual lexicon covering five medical topics. Using this dataset, we evaluate several state-of-the-art ASR models and examine different specific fine-tuning strategies for improving medical term recognition to investigate the best approach to solve in the dataset. Experimental results show that Vietnamese-optimized models perform better on general segments, while multilingual pretraining helps capture English insertions. The combination of both approaches yields the best balance between overall and code-switched accuracy. This work provides the first benchmark for Vietnamese medical code-switching and offers insights into effective domain adaptation for low-resource, multilingual ASR systems.
Towards Privacy-Preserving Fine-Tuning: Anonymization of Aphasic Speech for Effective ASR
Sebastian Hofstetter | Timo Baumann
Sebastian Hofstetter | Timo Baumann
The scarcity of publicly available aphasic speech data, driven largely by privacy concerns, poses a significant barrier for fine-tuning Automatic Speech Recognition (ASR) systems in this domain. This study investigates the privacy–utility trade-off of speech anonymization as a strategy to increase data availability. A signal-based McAdams anonymization method is applied to a subset of the AphasiaBank corpus comprising approximately 132 hours of speech from 425 individuals. Privacy is evaluated using an ECAPA-TDNN based Automatic Speaker Verification system and the Equal Error Rate metric. Linguistic utility is assessed by the Word Error Rate using wav2vec2.0 ASR model, tested in multiple conditions, both pretrained and fine-tuned on unprotected and anonymized audio. Our results show that fine-tuning on anonymized aphasic speech data improves ASR performance by +18 % compared to the performance of generic models on non-anonymized speech. Crucially, this gain in utility is achieved alongside substantial privacy protection, with anonymization increasing the privacy by +440 % compared to sharing unprotected speech. This work thus provides a proof-of-concept, demonstrating that speech anonymization mitigates privacy risks to tackle data scarcity and support the development of more effective ASR systems for people with aphasia.
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
Nikola Ljubešić | Peter Rupnik | Ivan Porupski | Taja Kuzman Pungeršek
Nikola Ljubešić | Peter Rupnik | Ivan Porupski | Taja Kuzman Pungeršek
ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages – Croatian, Czech, Polish and Serbian – with a total size of more than 6 thousand hours. The corpora were built in an automatic fashion from the ParlaMint transcripts and their corresponding metadata, which were aligned to the speech recordings of each corresponding parliament. In this release of the dataset, each of the corpora has been significantly enriched with several automatic annotation layers. The textual modality of all four corpora has been enriched with linguistic annotations and sentiment predictions. Similarly, their spoken modality has been automatically enriched with occurrences of filled pauses, the most frequent type of disfluency in typical speech. Two languages have been additionally enriched with detailed word- and grapheme-level alignments, and the automatic annotation of the position of primary stress in multisyllabic words. With these enrichments, the usefulness of the corpora has been greatly increased for downstream research across multiple disciplines, which we showcase through an analysis of acoustic correlates of sentiment. All the corpora are made available for download in JSONL and TextGrid formats, as well as for search through a concordancer.
LexiPhon: A Collection of Phonetically Transcribed Lexicons from Wikipedia
Amanda Doucette | Timothy J. O’Donnell | Morgan Sonderegger
Amanda Doucette | Timothy J. O’Donnell | Morgan Sonderegger
We introduce LexiPhon, an open-source dataset of phonetically transcribed lexicons for 87 languages derived from Wikipedia data with automated grapheme-to-phoneme (G2P) transcription, along with the open-source software used to create it. Each lexicon provides transcriptions generated by up to three G2P methods, crowdsourced transcriptions from WikiPron (Lee et al., 2020) where available, word frequencies calculated from Wikipedia, along with word lengths and phonological neighborhood densities. We introduce an internal validation metric based on phonological feature edit distance to ensure transcriptions are consistent within languages, as manual validation is not possible. This dataset fills a gap in the existing space of phonetic lexicons, with a much larger set of words per language than existing multilingual word lists, and more languages than existing lexicon datasets. The dataset, along with the software used to create it, are freely available on OSF at https://osf.io/rd9ma/overview?view_only=398802df19ad488ab7da7e7798cd7aca.
ROG: A Multi-Layer Manually Annotated Corpus of Spoken Slovenian
Kaja Dobrovoljc Zor | Darinka Verdonik | Jaka Čibej | Peter Rupnik | Nikola Ljubešić
Kaja Dobrovoljc Zor | Darinka Verdonik | Jaka Čibej | Peter Rupnik | Nikola Ljubešić
We present ROG, the first manually annotated spoken corpus of Slovenian to integrate morphosyntactic, prosodic, and interactional layers in a unified framework. Building on the pre-existing Spoken Slovenian Treebank (SST) and newly available recordings from the GOS 2 reference corpus, the resource combines over 75,000 words (10 hours) of annotated speech. The entire corpus features lemmatization, MULTEXT-East morphosyntax, and Universal Dependencies annotations, while approximately half includes additional layers for prosodic units, disfluencies, and dialogue acts. All annotation layers are systematically aligned and cross-referenced, enabling detailed multi-dimensional analyses of spoken language. We describe the corpus design, annotation workflow, data release, and baseline modeling results, showcasing the resource’s value for both linguistic analysis and speech-aware NLP model development. All ROG transcriptions and annotations, along with half of the audio recordings, are freely available under CC-BY via (anonymized) repository.
Building a Dataset for French Accent Classification Evaluation: Are We There Yet?
Diandra Fabre | Mathieu Avanzi | François Portet
Diandra Fabre | Mathieu Avanzi | François Portet
Current evaluation practices in speech processing systems often overlook the diversity of spoken accents, leading to significant performance disparities across speaker groups. This issue largely comes from biases and imbalances in training corpora, and is further compounded by the scarcity of open-source datasets suitable for evaluating accent variability in French. To address this gap, we extend the CFPR dataset with explicit accent labels, providing a new benchmark for assessing the robustness of speech technology systems across diverse French accents. We additionally conduct a perceptual study with 87 human participants to evaluate the reliability and interpretability of these labels. Using this resource, we evaluated an eight-class French accent classifier trained on Common Voice data. The first results highlight both the complexity of automatic French accent recognition in low-resource settings, and the difficulty for French-speakers to perceive all the linguistic variabilities in French-speaking countries.
M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
Yejin Kwon | Taewoo Kang | Hyunsoo Yoon | Chang Ouk Kim
Yejin Kwon | Taewoo Kang | Hyunsoo Yoon | Chang Ouk Kim
We present M3-SLU, a new multimodal large language model (MLLM) benchmark for evaluating multi-speaker, multi-turn spoken language understanding. While recent models show strong performance in speech and text comprehension, they still struggle with speaker-attributed reasoning, the ability to understand who said what and when in natural conversations. M3-SLU is built from four open corpora (CHiME-6, MELD, MultiDialog, and AMI) and comprises over 12,000 validated instances with paired audio, transcripts, and metadata. It includes two tasks: (1) Speaker-Attributed Question Answering and (2) Speaker Attribution via Utterance Matching. We provide baseline results for both cascaded pipelines and end-to-end MLLMs, evaluated using an LLM-as-Judge and accuracy metrics. Results show that while models can capture what was said, they often fail to identify who said it, revealing a key gap in speaker-aware dialogue understanding. M3-SLU offers as a challenging benchmark to advance research in speaker-aware multimodal understanding.
Medispeech: A French Reading and Spontaneous Speech Corpus for Sleepiness Estimation
Colleen Beaumard | Vincent P. Martin | Charles Brazier | Julien Coelho | Jean-Luc Rouas | Pierre Philip
Colleen Beaumard | Vincent P. Martin | Charles Brazier | Julien Coelho | Jean-Luc Rouas | Pierre Philip
Excessive Daytime Sleepiness (EDS) is associated with several diseases and therefore negatively affects the daily life of impacted people. Its diagnosis and follow-up are difficult because they require testing at the hospital for one full day. Monitoring patients regularly in ecological conditions may be done through speech analysis. Although several corpora containing speech from sleepy subjects exist, they do not suit ecological requirements regarding either the device used for recording or the speech elicitation tasks. In this paper, we introduce the Medispeech corpus containing reading, daily-life semi-spontaneous, and medically-oriented spontaneous tasks. Fifty-nine French subjects were recorded with both a professional-quality microphone and a smartphone using a dedicated application, resulting in 1,729 recordings for a total duration of 21 hours. Their EDS diagnosis was assessed by both a physiological objective measurement (mean sleep latency measured during a clinical test) and a subjective questionnaire (Karolinska Sleepiness Scale). Phenotyping of subjects is assured by collecting socio-demographic and medical data related to diverse dimensions of sleepiness, comorbidities, and addictions. Finally, we analyse the validity of our data collection protocol by measuring the effective duration of speech (after discarding pauses) and assessing its links with the collected subjects’ characteristics.
StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario
Marcely Zanon Boito | Caroline Brun | Inyoung Kim | Denys M. PROUX | Salah Ait-Mokhtar | Nikolaos Lagos | Jean-Luc Meunier | Ioan Calapodescu
Marcely Zanon Boito | Caroline Brun | Inyoung Kim | Denys M. PROUX | Salah Ait-Mokhtar | Nikolaos Lagos | Jean-Luc Meunier | Ioan Calapodescu
LLMs and speech assistants are increasingly used for task-oriented interactions, yet their evaluation often relies on controlled scenarios that fail to capture the variability and complexity of real user requests. Drink ordering, for example, involves diverse named entities, drink types, sizes, customizations, and brand-specific terminology, as well as spontaneous speech phenomena such as hesitations and self-corrections. To address this gap, we introduce StarDrinks, a test set in English and Korean containing speech utterances features, transcriptions, and annotated slots. Our dataset supports speech-to-slots SLU, transcription-to-slots NLU, and speech-to-transcription ASR evaluation, providing a realistic benchmark for model robustness and generalization in a linguistically rich, real-world task.
Audio-Lyrics Alignment Dataset for Italian Arias
Pushkar Jajoria | Arianna Graciotti | Giovanna Casali | Jesujoba Alabi | Rodolfo Delmonte | Angelo Pompilio | Rocco Tripodi | James McDermott | Dietrich Klakow
Pushkar Jajoria | Arianna Graciotti | Giovanna Casali | Jesujoba Alabi | Rodolfo Delmonte | Angelo Pompilio | Rocco Tripodi | James McDermott | Dietrich Klakow
Aligning song lyrics with sung audio is challenging, especially for languages and music styles where annotated datasets are scarce. We address this gap by presenting the first dataset of Italian opera arias annotated with lyrics and time-stamps per word. The dataset comprises of 24 arias drawn from well-known operas of the 18th to 20th centuries with a total audio duration of nearly two hours. We benchmark both music alignment models and speech forced alignment models and show that existing methods face significant challenges on this dataset, with performance dropping by 45% compared to other datasets. Multilingual and speech-based models exhibit relatively better performance on this dataset. We also evaluate few-shot fine-tuning of these models on the new dataset and find that, while it yields only marginal overall improvement, it produces localized gains on specific arias, suggesting that limited exposure helps the model adapt to some patterns but cannot fully overcome differences in language or musical style.
The Added Value of Metadata and Annotations: Evidence from Two Large-Scale, Naturalistic Corpus Studies
Anisia Popescu | Johanna Cronenberg | Ioana Vasilescu | Ioana Chitoran | Lori Lamel | Martine Adda-Decker
Anisia Popescu | Johanna Cronenberg | Ioana Vasilescu | Ioana Chitoran | Lori Lamel | Martine Adda-Decker
This paper presents two case studies that highlight both the challenges and benefits of working with large-scale, naturalistic phonetic data. Our aim is to encourage researchers not to shy away from phonetic data found “in the wild”, even when such data are messy, noisy, or incomplete – because they can yield robust, novel insights beyond the reach of controlled laboratory studies. We focus on challenges that are endemic to large corpora, including degraded audio quality, sparse or inconsistent annotations, and missing speaker metadata. By comparing two corpus-based studies that diverge in methodology and statistical design, we show how different approaches can mitigate these limitations while still extracting meaningful patterns.
CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech
Brian Yan | Qingzheng Wang | Matthew Wiesner | Anuj Diwan | Olga Iakovenko | Alex Polok | Injy Hamed | Shuichiro Shimizu | Iris Emerman | Thomas Hain | David R. Mortensen | Peter Viechnicki | Shinji Watanabe
Brian Yan | Qingzheng Wang | Matthew Wiesner | Anuj Diwan | Olga Iakovenko | Alex Polok | Injy Hamed | Shuichiro Shimizu | Iris Emerman | Thomas Hain | David R. Mortensen | Peter Viechnicki | Shinji Watanabe
We present CS-YODAS, a Creative Commons dataset of in-the-wild code-switched speech mined from multilingual YouTube data. Code-switching, or the alternation between languages within an utterance or conversation, is common in multilingual settings but remains underrepresented in existing CS speech resources, which are typically small, domain-specific, or artificially constructed. Building on the YODAS corpus, we develop a scalable, human-in-the-loop pipeline for identifying and validating naturally occurring code-switching. The resulting dataset, which totals 313 hrs and spans 7 matrix languages, provides diverse, real-world examples of spontaneous code-switched speech. We further analyze the distribution and characteristics of code-switching in the wild, examining language-pair frequencies and switching patterns, and report baseline results for spoken language identification. We hope that CS-YODAS will encourage broader and more comprehensive research on code-switched speech. Dataset link: https://huggingface.co/datasets/byan/cs-yodas.
The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR
Siyu Liang | Nicolas Ballier | Gina-Anne Levow | Richard Wright
Siyu Liang | Nicolas Ballier | Gina-Anne Levow | Richard Wright
How much audio is needed to fully observe a multilingual ASR model’s learned sub-token inventory across languages, and does data disparity in multilingual pre-training affect how these tokens are utilized during inference? We address this question by analyzing Whisper’s decoding behavior during inference across 49 languages. By logging decoding candidate sub-tokens and tracking their cumulative discovery over time, we study the utilization pattern of the model’s sub-token space. Results show that the total number of discovered tokens remains largely independent of a language’s pre-training hours, indicating that data disparity does not strongly influence lexical diversity in the model’s hypothesis space. Sub-token discovery rates follow a consistent exponential saturation pattern across languages, suggesting a stable time window after which additional audio yields minimal new token activation. We refer to this convergence threshold as acoustic saturation time AST. Further analyses of rank–frequency distributions reveal Zipf-like patterns better modeled by a Zipf–Mandelbrot law, and mean sub-token length shows a positive correlation with resource level. Additionally, those metrics show more favorable patterns for languages in the Latin script than those in scripts such as Cyrillic, CJK and Semitic. Together, our study suggests that sub-token utilization during multilingual ASR inference is constrained more by the statistical, typological, and orthographical structure of the speech than by training data scale, providing an empirical basis for more equitable corpus construction and cross-lingual evaluation.
AusKidTalk: Developing Transcription Guidelines for Continuous Australian English Child Speech
Tuende Szalay | Zheng Nan | Renata Huang | Mostafa Shahin | Sirojan Tharmakulasingam | Kirrie Ballard | Beena Ahmed
Tuende Szalay | Zheng Nan | Renata Huang | Mostafa Shahin | Sirojan Tharmakulasingam | Kirrie Ballard | Beena Ahmed
Guidelines are required for accurate and consistent transcription of speech corpora, especially when they contain more challenging, e.g. spontaneous or under-resourced speech. This paper presents a workflow and guidelines for transcribing spontaneous and under-resourced child speech in AusKidTalk, the first Australian English child corpus. Speech samples were elicited using a story-telling task and are 3.5 minutes long per child on average. Orthographic transcriptions were generated using automatic speech recognition (ASR) tools and corrected manually. A novel hand-correction protocol consisting of guidelines, hand-correction interface, and ground truth transcriptions together with consistency metrics were developed. Nine annotators submitted hand-corrections for 261 children’s story-telling task, and 25 ground truth tasks. Manual correction was 11-fold of speech time with a 3.5-minute-long story-telling task corrected in approximately 40 minutes. Efficiency is attributed to the quality of automatic transcription with 23% word error rate. Manual correction was accurate with annotators achieving consistent results on 15/25 ground truth submissions. Most inconsistent ground truth submissions were caused by a single, challenging ground truth task. These results show that our workflow yields efficient and accurate transcriptions, although transcriptions of potentially more challenging narrative tasks (e.g., elicited from younger children) might require further corrections.
spINAch: A Diachronic Corpus of French Broadcast Speech Controlled for Speakers’ Age and Gender
Simon Devauchelle | David Doukhan | Remi Uro | Lucas Ondel | Valentin Pelloin | Olympia Imbert-Brégégère | Véronique Lefort | Kévin Picard | Emeline Seignobos | Albert Rilliard
Simon Devauchelle | David Doukhan | Remi Uro | Lucas Ondel | Valentin Pelloin | Olympia Imbert-Brégégère | Véronique Lefort | Kévin Picard | Emeline Seignobos | Albert Rilliard
We present spINAch, a large diachronic corpus of French speech from radio and television archives, balanced by speakers’ gender, age (20-95 years old), and spanning 60 years from 1955 to 2015. The dataset includes over 320 hours of recordings from more than two thousand speakers. The methodology for building the corpus is described, focusing on the quality of collected samples in acoustic terms. The data were automatically transcribed and phonetically aligned to allow studies at a phonemic level. More than 3 million oral vowels have been analyzed to propose their fundamental frequency and formants. The corpus, available to the community for research purposes, is valuable for describing the evolution of Parisian French through the representation of gender and age. The presented analyses also demonstrate that the diachronic nature of the corpus allows the observation of various phonetic phenomena, such as the evolution of voice pitch over time (which does not differ by gender in our data) and the neutralization of the /a/-/ɑ/ opposition in Parisian French during this period.
SALAN: A Massive ASR Dataset for the Languages of Niger
Mamadou K KEITA | Christopher Homan | Emily Prud’hommeaux | Abdoulaye SAKO | Seydou Diallo
Mamadou K KEITA | Christopher Homan | Emily Prud’hommeaux | Abdoulaye SAKO | Seydou Diallo
We introduce SALAN, a large-scale speech dataset covering eight of the major indigenous languages of Niger: Zarma, Hausa, Buduma, Gourmantchema, Tubu, Tamasheq, Fulfulde, and Kanuri. The final dataset exceeds 2,000 hours of audio, largely sourced from radio broadcasts and community recordings. We transcribed portions of the audio using the MMS model and conducted manual verification for 110 hours across Zarma and Hausa. We then used active learning to expand annotation to an additional 5 hours of high-uncertainty Zarma segments. To evaluate SALAN’s utility for ASR, We fine-tuned both Wav2vec2 XLS-R and Whisper on Zarma subsets and carried out additional pre-training with multilingual unlabeled data. Our best model achieved a word error rate of 25.3% and a character error rate of 6.2%. SALAN and the trained models will be made publicly available for use by researchers and speakers, with the potential to impact over 20 million individuals in Niger and neighboring countries.
Listening for Ideology: Automatic Analysis of Character Speech in Historical Nazi Propaganda Films
Nicolas Ruth | Manuel Burghardt | Andreas Niekler
Nicolas Ruth | Manuel Burghardt | Andreas Niekler
While the visual dimension of film has been widely explored in digital humanities through methods such as “distant viewing”, the audio layer has received less attention despite its crucial role in meaning-making. We address this gap with a four-step pipeline combining speaker diarization, audio gender classification, automatic speech recognition (ASR), and LLM-based psycholinguistic analysis to infer character traits from film dialogues. Applying this method to a set of Nazi propaganda films, we find that despite challenges in speaker diarization due to noisy historical film audio, modern ASR and GPT-based analyses produce character profiles consistent with existing filmic research. Our proposed pipeline advances distant reading of film dialogue, complementing visual analyses and enabling scalable study of ideology in historical cinema. A case study of female characters in NS films identifies three recurring types, centered on the ideological figure of the mother in National Socialism.
Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
Nick Rossenbach | Robin Schmitt | Tina Raissi | Simon Berger | Larissa Kleppel | Ralf Schlüter
Nick Rossenbach | Robin Schmitt | Tina Raissi | Simon Berger | Larissa Kleppel | Ralf Schlüter
The recently published Loquacious dataset aims to be a replacement for established English automatic speech recognition (ASR) datasets such as LibriSpeech or TED-Lium. The main goal of Loquacious dataset is to provide properly defined training and test partitions across many acoustic and language domains, with an open license suitable for both academia and industry. To further promote the benchmarking and usability of this new dataset, we present additional resources in the form of n-gram language models (LMs), a grapheme-to-phoneme (G2P) model and pronunciation lexica, with open and public access. Utilizing those additional resources we show experimental results across a wide range of ASR architectures with different label units and topologies. Our initial experimental results indicate that the Loquacious dataset offers a valuable study case for a variety of common challenges in ASR.
WhiteHouse: Translation of the Casablanca Corpus for Multi-dialectal Arabic Speech Translation
Fethi Bougares | Salima Mdhaffar | Yannick Estève
Fethi Bougares | Salima Mdhaffar | Yannick Estève
Remarkable progress has been made recently in the speech processing of Arabic dialects. This is primarily due to the availability of large multilingual pre-trained models as well as the development of multiple well-annotated datasets that support training, fine-tuning, and evaluation of various speech models. However, most existing research on Arabic speech processing did not consider Automatic Speech Translation (AST) and focused mainly on Dialect Identification (DI) and Automatic Speech Recognition (ASR) tasks. To address this gap, we introduce WhiteHouse, the first multi-dialectal Arabic-English Speech Translation Corpus. WhiteHouse supplements the recently created Casablanca dataset with English translation for each utterance in the transcripts. This results in a three-way parallel speech-transcription-translation multi-dialectal Arabic dataset. WhiteHouse dataset is used to evaluate various SoTA speech translation models. Our experiments show that SoTA speech translation models performs poorly when evaluated on Arabic dialectal conditions. All the data used during training and testing are released for public use and further improvements
Manual transcription of intonation by experts remains an essential part of research on the structure and meaning of intonation across languages, as well as for developing computational methods for automatic intonation transcription. We present ToneSwiper, a Python program with a graphical user interface that facilitates manual intonation transcription in the ToDI framework (Transcription of Dutch Intonation; Gussenhoven, 2005), with possible adaptation to similar (e.g., ToBI-like) frameworks for other languages. For the trained annotator, it enables efficient ToDI transcription of speech by integrating an audio-player, a spectrogram and pitch contour plot, auto-scroll, dynamic audio stretching, and an intuitive hotkey interface that maps key sequences to ToDI elements, e.g., pressing up-down for a high-to-low accent (H*L). In this way, transcription is conducted by ‘swiping‘ over the arrow keys on the keyboard. We present the program and its motivation, as well as a small-scale pilot study on annotation efficiency and inter-rater agreement, using a highly challenging sample of task-oriented dialogue from the Dutch Map Task Corpus (Ladd and Schepman, 2003).
IMaSC: A Malayalam Speech Corpus for High-Quality Text-to-Speech Synthesis
Deepa P. Gopinath | Thennal D K | Vrinda V. Nair | Swaraj K. S | Sachin G
Deepa P. Gopinath | Thennal D K | Vrinda V. Nair | Swaraj K. S | Sachin G
Modern text-to-speech (TTS) systems use deep learning to synthesize speech increasingly approaching human quality, but they require a database of high-quality audio-text sentence pairs for training. Malayalam, the official language of the Indian state of Kerala and spoken by 35+ million people, is a low-resource language in terms of available corpora for TTS systems. In this paper, we present IMaSC, a Malayalam text and speech corpora containing 49 hours and 37 minutes of recorded speech. With 8 speakers and a total of 34,473 text-audio pairs, IMaSC is larger than every other publicly available alternative. We evaluated the database by using it to train TTS models for each speaker based on a modern deep learning architecture. With an average mean opinion score of 4.50, we find that the synthesized speech of our model is close to human quality.
Speak in Context: Multilingual ASR with Speech–Context Alignment via Contrastive Learning
Yuchen Zhang | Haralambos Mouratidis | Ravi Shekhar
Yuchen Zhang | Haralambos Mouratidis | Ravi Shekhar
Automatic speech recognition (ASR) has benefited from advances in pretrained speech and language models, yet most systems remain constrained to monolingual settings and short, isolated utterances. While recent efforts in context-aware ASR show promise, two key challenges persist: limited multilingual support and the absence of principled alignment between speech and contextual representations. In this paper, we introduce a context-aware multilingual ASR framework that supports diverse languages and accents while preserving the modularity of pretrained models. Our approach combines a frozen speech encoder and a decoder-only language model via a lightweight projection module, allowing structured context prompts, including dialogue history and biasing words, to guide transcription. To improve interaction between speech and context, we employ a contrastive learning objective that aligns their representations in a shared embedding space. Evaluations on over 1,500 hours of real-world conversational speech across 11 languages and 5 English dialects show that contextual input consistently improves recognition quality. Contrastive alignment provides additional gains when applied to different context types, with an overall performance gain of over 5%. These results highlight the importance of both contextual modeling and cross-modal alignment in multilingual ASR.
Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
Swati Sharma | Divya V. Sharma | Anubha Gupta
Swati Sharma | Divya V. Sharma | Anubha Gupta
The rising demand for inclusive speech technologies amplifies the need for multilingual datasets for Natural Language Processing (NLP) research. However, limited awareness of existing task-specific resources in low-resource languages hinders research. This challenge is especially acute in linguistically diverse countries, such as India. Cross-task profiling of existing Indian speech datasets can alleviate the data scarcity challenge. This involves investigating the utility of datasets across multiple downstream tasks rather than focusing on a single task. Prior surveys typically catalogue datasets for a single task, leaving comprehensive cross-task profiling as an open opportunity. Therefore, we propose Task-Lens, a cross-task survey that assesses the readiness of 50 Indian speech datasets spanning 26 languages for nine downstream speech tasks. First, we analyze which datasets contain metadata and properties suitable for specific tasks. Next, we propose task-aligned enhancements to unlock datasets to their full downstream potential. Finally, we identify tasks and Indian languages that are critically underserved by current resources. Our findings reveal that many Indian speech datasets contain untapped metadata that can support multiple downstream tasks. By uncovering cross-task linkages and gaps, Task-Lens enables researchers to explore the broader applicability of existing datasets and to prioritize dataset creation for underserved tasks and languages.
We introduce the Mandarin–English Language Interview (MELI) Corpus, an open-source resource of 29.8 hours of speech from 51 Mandarin–English bilingual speakers. MELI combines matched sessions in Mandarin and English with two speaking styles: read sentences and spontaneous interviews about language varieties, standardness, and learning experiences. Audio was recorded at 44.1 kHz (16-bit, stereo). Interviews were fully transcribed, force-aligned at word and phone levels, and anonymized. Descriptively, the Mandarin component totals ~14.7 hours (mean duration 17.3 minutes) and the English component ~15.1 hours (mean duration 17.8 minutes). We report token/type statistics for each language and document code-switching patterns (frequent in Mandarin sessions; more limited in English sessions). The corpus design supports within-/cross-speaker, within/cross-language acoustic comparison and links speech content to speakers’ stated language attitudes, enabling both quantitative and qualitative analyses. The MELI Corpus will be released with transcriptions, alignments, metadata, scans of labelled maps and documentation under a CC BY-NC 4.0 license.
PhonemeDF: A Synthetic Speech Dataset for Audio Deepfake Detection and Naturalness Evaluation
Vamshi Nallaguntla | Aishwarya R. Fursule | Shruti Kshirsagar | Anderson Raymundo Avila
Vamshi Nallaguntla | Aishwarya R. Fursule | Shruti Kshirsagar | Anderson Raymundo Avila
The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic speech with naturalness and intelligibility. This poses serious threats to voice biometric security and to systems designed to combat the spread of spoken misinformation, where synthetic voices may be used to disseminate false or malicious content. While interest in AI-generated speech has increased, resources for evaluating naturalness at the phoneme level remain limited. In this work, we address this gap by presenting the Phoneme-Level DeepFake dataset (PhonemeDF), comprising parallel real and synthetic speech segmented at the phoneme level. Real speech samples are derived from a subset of LibriSpeech, while synthetic samples are generated using four TTS and three VC systems. For each system, phoneme-aligned TextGrid files are obtained using the Montreal Forced Aligner (MFA). We compute the Kullback–Leibler divergence (KLD) between real and synthetic phoneme distributions to quantify fidelity and establish a ranking based on similarity to natural speech. Our findings show a clear correlation between the KLD of real and synthetic phoneme distributions and the performance of classifiers trained to distinguish them, suggesting that KLD can serve as an indicator of the most discriminative phonemes for deepfake detection.
How Much Data for Stable Formant Values? Pipeline for Convergence Detection Based on Read Speech
Kayla Sward | Johan Sjons | Axel G. Ekstrom
Kayla Sward | Johan Sjons | Axel G. Ekstrom
This study investigates the stability and convergence of vowel formants (F1, F2, F3) in read speech through an extensive corpus of audiobook recordings. While most formant studies rely on brief, isolated utterances recorded in laboratory settings, this analysis draws on 3,384 chapters (about 942 hours) of continuous, stylistically varied speech from publicly available audiobooks. The data was processed using an automated pipeline that comprised transcription, phoneme alignment, and formant extraction. Several statistical techniques – First Token Within (FTW), Cumulative Sum (CUSUM), Two-Sample t-Test, Confidence Interval (CI) Shrinkage, Piecewise Linear Fitting (PWLF), and Binary Segmentation (BinSeg) – were compared for their effectiveness in identifying stabilization points. Findings indicate that formant means generally stabilize within 60 to 230 vowel tokens per phoneme, dependent on vowel type and speaker gender. Of the methods that were evaluated, CUSUM yielded the most consistent and informative results. The results provide practical guidelines for determining the quantity of non-laboratory speech required to obtain reliable vowel formant averages.
MUSCAT: MUltilingual, SCientific ConversATion Benchmark
Supriti Sinhamahapatra | Thai-Binh Nguyen | Yiğit Oğuz | Enes Yavuz Ugan | Jan Niehues | Alexander Waibel
Supriti Sinhamahapatra | Thai-Binh Nguyen | Yiğit Oğuz | Enes Yavuz Ugan | Jan Niehues | Alexander Waibel
The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were a multilingual speaker. To create this experience, speech technology needs to address several challenges: Handling mixed multilingual input, specific vocabulary, and code-switching. However, there is currently no dataset benchmarking this situation. We propose a new benchmark to evaluate current Automatic Speech Recognition (ASR) systems, whether they are able to handle these challenges. The benchmark consists of bilingual discussions on scientific papers between multiple speakers, each conversing in a different language. We provide a standard evaluation framework, beyond Word Error Rate (WER) enabling consistent comparison of ASR performance across languages. Experimental results demonstrate that the proposed dataset is still an open challenge for state-of-the-art ASR systems. The dataset is available in https://huggingface.co/datasets/goodpiku/muscat-eval
Towards a Diagnostic and Predictive Evaluation Methodology for Sequence Labeling Tasks
Elena Alvarez-Mellado | Julio Gonzalo
Elena Alvarez-Mellado | Julio Gonzalo
Standard evaluation in NLP typically indicates that system A is better on average than system B, but it provides little info on how to improve performance and, what is worse, it should not come as a surprise if B ends up being better than A on outside data. We propose an evaluation methodology for sequence labeling tasks grounded on error analysis that provides both quantitative and qualitative information on where systems must be improved and predicts how models will perform on a different distribution. The key is to create test sets that, contrary to common practice, do not rely on gathering large amounts of real-world in-distribution scraped data, but consists in handcrafting a small set of linguistically motivated examples that exhaustively cover the range of span attributes (such as shape, length, casing, sentence position, etc.) a system may encounter in the wild. We demonstrate this methodology on a benchmark for anglicism identification in Spanish. Our methodology provides results that are diagnostic (because they help identify systematic weaknesses in performance), actionable (because they can inform which model is better suited for a given scenario) and predictive: our method predicts model performance on external datasets with a median correlation of 0.85.
Memorization or Lucky Guesses: Detecting Short Sequences from Copyrighted Dutch News in LLM Output
Joris Veerbeek | Kas Berendsen | Alessandra Polimeno | Antal van den Bosch
Joris Veerbeek | Kas Berendsen | Alessandra Polimeno | Antal van den Bosch
Demonstrating that large language models have memorized copyrighted material is more feasible for high-volume publishers than for smaller outlets whose content appears less frequently online. This study explores how even short, repeated sequences–rather than full articles–can serve as evidence of memorization. Focusing on Dutch news sources included in the mC4 dataset, we test whether GPT-4 and mT5 reproduce excerpts from thousands of articles, including standardized editorial boilerplate. By comparing results to a post-training baseline and modeling memorization as a survival process, we find that repeated, publication-specific phrases are significantly more likely to be completed verbatim. The approach provides a means to detect empirical evidence of memorization in cases where full reproduction is unlikely.
When Numbers Tell Half the Story: Human-Metric Alignment in Topic Model Evaluation
Thibault Prouteau | Francis Lareau | Nicolas Dugue | Jean-Charles Lamirel | Christophe Malaterre
Thibault Prouteau | Francis Lareau | Nicolas Dugue | Jean-Charles Lamirel | Christophe Malaterre
Topic models uncover latent thematic structures in text corpora, yet evaluating their quality remains challenging, particularly in specialized domains. Existing methods often rely on automated metrics like topic coherence and diversity, which may not fully align with human judgment. Human evaluation tasks, such as word intrusion, provide valuable insights but are costly and primarily validated on general-domain corpora. This paper introduces Topic Word Mixing (TWM), a novel human evaluation task assessing inter-topic distinctness by testing whether annotators can distinguish between word sets from single or mixed topics. TWM complements word intrusion’s focus on intra-topic coherence and provides a human-grounded counterpart to diversity metrics. We evaluate six topic models–both statistical and embedding-based (LDA, NMF, Top2Vec, BERTopic, CFMF, CFMF-emb)–comparing automated metrics with human evaluation methods based on nearly 4,000 annotations from a domain-specific corpus of philosophy of science publications. Our findings reveal that word intrusion and coherence metrics do not always align, particularly in specialized domains, and that TWM captures human-perceived distinctness while appearing to align with diversity metrics. We release the annotated dataset and task generation code. This work highlights the need for evaluation frameworks bridging automated and human assessments, particularly for domain-specific corpora.
Detecting Hallucinations in Authentic LLM–Human Interactions
Yujie Ren | Niklas Gruhlke | Anne Lauscher
Yujie Ren | Niklas Gruhlke | Anne Lauscher
As large language models (LLMs) are increasingly applied in sensitive domains such as medicine and law, hallucination detection has become a critical task. Although numerous benchmarks have been proposed to advance research in this area, most of them are artificially constructed––either through deliberate hallucination induction or simulated interactions––rather than derived from genuine LLM–human dialogues. Consequently, these benchmarks fail to fully capture the characteristics of hallucinations that occur in real-world usage. To address this limitation, we introduce AuthenHallu, the first hallucination detection benchmark built entirely from authentic LLM–human interactions. For AuthenHallu, we select and annotate samples from genuine LLM–human dialogues, thereby providing a faithful reflection of how LLMs hallucinate in everyday user interactions. Statistical analysis shows that hallucinations occur in 31.4% of the query–response pairs in our benchmark, and this proportion increases dramatically to 60.0% in challenging domains such as ’Math & Number Problems’. Furthermore, we explore the potential of using vanilla LLMs themselves as hallucination detectors and find that, despite some promise, their current performance remains insufficient in real-world scenarios. The data and code are publicly available at https://github.com/TAI-HAMBURG/AuthenHallu.
Issue Detection and Category Classification in Domain-Specific Technical Logbooks
Afshin Karimi | Ingmar Hartl | Henrik Tuennermann | Anne Lauscher
Afshin Karimi | Ingmar Hartl | Henrik Tuennermann | Anne Lauscher
Operating large-scale research infrastructures such as free-electron lasers produces vast amounts of operator-authored documentation that records daily observations, anomalies, and maintenance actions. These logbooks and incident reports contain valuable operational knowledge but often remain underexplored due to their unstructured, domain-specific language. While large language models (LLMs) show strong generalization in general domains, their effectiveness on such technical operator text has, to the best of our knowledge, not been systematically assessed. We introduce two new English datasets from real-world laser operations: (i) a logbook dataset annotated for binary issue detection (does an entry describe or report an actionable fault?), and (ii) an operator ticket dataset annotated for multi-class issue categorization assign each ticket to one of 13 technical categories). The corpora comprise 2,979 logbook entries and 758 tickets from 2022–2024; both are cleaned, anonymized, and suitable for benchmarking classification performance. We evaluate four open LLMs (LLaMA-3, Mistral-Small, Qwen-3-30B, GPT-OSS-120B) under zero-shot, few-shot, and chain-of-thought (CoT) prompting, using multiple semantically equivalent prompt variants per setting to assess robustness. Across both tasks, few-shot prompting is consistently strongest, with top systems reaching F1 approx 0.84 for logbook issue detection and Macro-F1 0.42 for operator ticket categorization. These results suggest that incorporating a handful of in-domain examples can substantially improve performance on operator-authored technical text, even without fine-tuning.
Once upon a Kernel: Extracting Important Events from Narratives
Anshu Kiran Sharma | Miguel Castiblanco-Melendez | Alejandro Morales | Mark A. Finlayson
Anshu Kiran Sharma | Miguel Castiblanco-Melendez | Alejandro Morales | Mark A. Finlayson
Not all events in a narrative are created equal: some events are more important than others. Kernel events, a concept introduced in the field of narratology, are causally linked events that move the narrative forward, and cannot be removed without breaking the narrative’s logical coherence. While event detection and extraction tasks have been widely studied in natural language processing and information retrieval fields, the idea of kernel events has been largely unexplored. In this work, we introduce the first corpus and model for kernel event detection. Our contributions include: the refinement of the kernel event concept captured in detailed annotation guidelines grounded in narratological principles; an annotation study yielding a gold-standard dataset of kernel events in narrative texts; and a first-of-its-kind kernel event detection system. Annotation achieved an inter-annotator agreement of 0.61 Kappa, underscoring the reliability of the guidelines. Using these data, we trained several models in both fine-tuned and generative modes for kernel event detection, with a LoRA fine-tuned Llama3 achieving an F1 of 0.695. This work establishes a benchmark for kernel event detection, with potential applications in summarization, narrative similarity detection, and narrative understanding. We release our code and data for the benefit of other researchers.
In litigation, trial transcripts provide verbatim records of witness testimony, primarily given in response to attorney questioning. To effectively analyze these transcripts, lawyers must often reconstruct events in chronological order—a task that begins with identifying dates associated with testified facts. This paper introduces two datasets for temporal expression extraction from legal transcripts: a primary dataset derived from a lengthy 1995 U.S. criminal trial, and a smaller robustness-testing dataset drawn from seven other legal proceedings. We evaluate semi-supervised approaches for date entity recognition, fine-tuning neural models on weakly labeled training data, and benchmarking them against both small and large language models. Our best-performing models achieve 83% F1-score on the primary dataset (FLAIR rule-modified) and 72% F1-score on the cross-domain, small test set (BERT-cased). These results, alongside our annotated datasets and corresponding experiments, provide a foundation for developing robust date extraction and temporal ordering tools for speech-derived legal text. Moreover, we identify unique challenges for state-of-the-art NER models on legal transcripts, including legal terminology and multiple anchor date resolution.
Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset
Z. Melce Hüsünbeyi | Virginie Mouilleron | Leonie Uhling | Daniel Foppe | Tatjana Scheffler | Djamé Seddah
Z. Melce Hüsünbeyi | Virginie Mouilleron | Leonie Uhling | Daniel Foppe | Tatjana Scheffler | Djamé Seddah
The rapid proliferation of misinformation across online platforms underscores the urgent need for robust, up-to-date, explainable, and multilingual fact-checking resources. However, existing datasets are limited in scope, often lacking multimodal evidence, structured annotations, and detailed links between claims, evidence, and verdicts. This paper introduces a comprehensive data collection and processing pipeline that constructs multimodal fact-checking datasets in French and German languages by aggregating ClaimReview feeds, scraping full debunking articles, normalizing heterogeneous claim verdicts, and enriching them with structured metadata and aligned visual content. We used state-of-the-art large language models (LLMs) and multimodal LLMs for (i) evidence extraction under predefined evidence categories and (ii) justification generation that links evidence to verdicts. Evaluation with G-Eval and human assessment demonstrates that our pipeline enables fine-grained comparison of fact-checking practices across different organizations or media markets, facilitates the development of more interpretable and evidence-grounded fact-checking models, and lays the groundwork for future research on multilingual, multimodal misinformation verification.
A Study on Building Efficient Zero-Shot Relation Extraction Models
Hugo THOMAS | Caio Corro | Guillaume Gravier | Pascale Sébillot
Hugo THOMAS | Caio Corro | Guillaume Gravier | Pascale Sébillot
Zero-shot relation extraction aims to identify relations between entity mentions using textual descriptions of novel types (i.e., previously unseen) instead of labeled training examples. Previous works often rely on unrealistic assumptions: (1) pairs of mentions are often encoded directly in the input, which prevents offline pre-computation for large scale document database querying; (2) no rejection mechanism is introduced, biasing the evaluation when using these models in a retrieval scenario where some (and often most) inputs are irrelevant and must be ignored. In this work, we study the robustness of existing zero-shot relation extraction models when adapting them to a realistic extraction scenario. To this end, we introduce a typology of existing models, and propose several strategies to build single pass models and models with a rejection mechanism. We adapt several state-of-the-art tools, and compare them in this challenging setting, showing that no existing work is really robust to realistic assumptions, but overall AlignRE (Li et al., 2024) performs best along all criteria.
Beyond Catalogue Counts: The Dataset Visibility Asymmetry in Low-Resource Multilingual NLP
Zhiyin Tan | Changxu Duan
Zhiyin Tan | Changxu Duan
Multilingual NLP often relies on dataset counts from centralized catalogues to characterize which languages are resource-rich or resource-poor. However, these catalogues record only one layer of dataset visibility: what has been registered or institutionally distributed. They do not necessarily reflect which datasets are created, cited, or reused in the research literature. To examine this gap, we combine a catalogue-based baseline with literature-backed evidence of dataset circulation. We introduce the Resource Density Index (RDI), defined as the number of catalogued datasets per one million speakers, and compute it for the 200 most widely spoken languages in Ethnologue. Among them, 118 languages (59%) have an average RDI of zero across the LRE Map and the Linguistic Data Consortium (LDC), and another 23 fall below 0.1, corresponding to at most one catalogued dataset per ten million speakers. We then apply an LLM-assisted citation-mining pipeline over the Semantic Scholar corpus to these 141 low-visibility languages. After manual validation and consolidation, we identify 609 unique datasets across 53 languages, of which 356 remain openly accessible through working public links. These results reveal a substantial visibility gap: many large-speaker languages appear data-poor in catalogue records yet show clear evidence of dataset activity in the research literature. Our findings suggest that multilingual data scarcity should be understood not only as a production problem, but also as a question of documentation, discoverability, and long-term accessibility. Code and data are publicly available at https://github.com/zhiyintan/dataset-visibility-asymmetry.
BLooP: Zero-Shot Abstractive Summarization Using Large Language Models with Bigram Lookahead Promotion
Varun Iyer | Cornelia Caragea
Varun Iyer | Cornelia Caragea
Abstractive summarization requires models to generate summaries that convey information in the source document. While large language models can generate summaries without fine-tuning, they often miss key details and include extraneous information. We propose BLooP (Bigram Lookahead Promotion), a simple training-free decoding intervention that encourages large language models (LLMs) to generate tokens that form bigrams from the source document. BLooP operates through a hash table lookup at each decoding step, requiring no training, fine-tuning, or model modification. We demonstrate improvements in ROUGE and BARTScore for Llama-3.1-8B-Instruct, Mistral-Nemo-Instruct-2407, and Gemma-2-9B-IT on CNN/DM, CCSum, Multi-News, and SciTLDR. Human evaluation shows that BLooP significantly improves faithfulness without reducing readability. We make the code available here.
OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
Hannah Liu | Murphy Tian | Iqra Ali | Haonan Gao | Qiaoyiwen Wu | Blair Yang | Uthayasanker Thayasivam | Annie En-Shiun Lee | Pakawat Nakwijit | Surangika Ranathunga | Ravi Shekhar
Hannah Liu | Murphy Tian | Iqra Ali | Haonan Gao | Qiaoyiwen Wu | Blair Yang | Uthayasanker Thayasivam | Annie En-Shiun Lee | Pakawat Nakwijit | Surangika Ranathunga | Ravi Shekhar
Text simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains limited for mid-resource and low-resource languages due to the scarcity of high-quality data. To address this gap, we introduce OasisSimp, a multilingual dataset for sentence-level text simplification covering five languages: English, Sinhala, Tamil, Pashto, and Thai. Among these, no prior sentence simplification datasets exist for Thai, Pashto, and Tamil, while limited data is available for Sinhala. Each language simplification dataset was created through direct human annotation, where trained annotators followed detailed guidelines to simplify sentences while maintaining meaning, fluency, and grammatical correctness. We evaluate eight open-weight multilingual Large Language Models (LLMs) on OasisSimp and observe substantial performance disparities between high-resource and low-resource languages, highlighting the simplification challenges in multilingual settings. OasisSimp thus provides both a valuable multilingual resource and a challenging benchmark, revealing the limitations of current LLM-based simplification methods and paving the way for future research in low-resource text simplification. The dataset will be open-sourced upon acceptance.
Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models
Thomas Stephan Juzek | Xiaoyang Ming | Jose A. Hernandez
Thomas Stephan Juzek | Xiaoyang Ming | Jose A. Hernandez
The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both WHAT divergences occur and, to some extent, WHY, linking them to the training stage of human preference learning. Yet, existing approaches rely on manual curation. This paper introduces two curation-free, assumption-light evaluation metrics: the Lexical Alignment Score, which identifies lexical overuse, and the Triangulated Preference Shift, which quantifies how much of such shifts can be attributed to human preference learning. Using PubMed abstracts, continuations were generated and measured using windowed document prevalence across six model families (Falcon, Gemma, Llama, Mistral, OLMo, Yi). The procedure identifies, without manual intervention, overused items such as ’suggest’, ’additionally’, and ’strategy’, and estimates their link to preference learning. Our findings replicate prior work and remain stable across parameter settings, random seeds, and evaluation on further data. The approach scales readily and enables systematic study of lexical (mis)alignment beyond Scientific English and across languages, and as such, the metrics have the potential to contribute to improved alignment for future models and understanding of its origins.
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
Nouran Khallaf | Serge Sharoff
Nouran Khallaf | Serge Sharoff
Noisy training data can significantly degrade the performance of language-model-based classifiers, particularly in non-topical classification tasks. This study explores a range of denoising strategies for sentence-level difficulty detection, using training data derived from document-level difficulty annotations obtained through noisy crowdsourcing. Beyond monolingual settings, we also address cross-lingual transfer, where a multilingual language model is trained in one language and tested in another. We evaluate several noise reduction techniques, including Gaussian Mixture Models (GMM), Co-Teaching, Noise Transition Matrices, and Label Smoothing. Our results indicate that while BERT-based models exhibit inherent robustness to noise, incorporating explicit noise detection can further enhance performance. For our smaller dataset, GMM-based noise filtering proves particularly effective in improving prediction quality by raising the AUC score from 0.52 to 0.86, or to 0.92 when two de-noising methods are combined (GMM and Co-Teaching). However, for our larger dataset, the intrinsic regularisation of pre-trained language models provides a strong baseline, with denoising methods yielding only marginal gains (from 0.8948 to 0.8984, or to 0.9061 when two denoising methods are combined). Nonetheless, removing noisy sentences (about 20% of the dataset) helps in producing a cleaner corpus with fewer infelicities. As a result we have released the largest available multilingual corpus for sentence difficulty prediction.
Comparing Reading Behavior across Reader Expertise and Text Complexity: Insights from the French Eye-Tracking Corpus (FETA)
Oksana Ivchenko | Natalia Grabar
Oksana Ivchenko | Natalia Grabar
This study examines how readers process general and medical texts with varying levels of complexity and how text simplification affects reading behavior. Using eye-tracking data, we compared two participant groups – common population and speech therapy students – as they read French medical, clinical, and general-domain texts in both original and simplified versions. We applied unsupervised clustering to identify patterns in reading behavior and investigate whether these patterns differ across participant groups, text types and complexity. The analysis identified between two and four clusters per group and condition, revealing distinct reading strategies ranging from effortful re-reading behavior to fluent, streamlined processing. The results reveal that medical and clinical texts elicit longer fixations and more regressions, indicating greater processing effort, while simplification produces shorter fixations and more fluid reading. Speech therapy students generally exhibit more efficient and stable gaze patterns, reflecting greater metalinguistic awareness and familiarity with the field. The dataset is a novel resource for modelling cognitive aspects of text complexity in French.
Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
Keizo Kato | Chenhui Chu | Yugo Murawaki | Sadao Kurohashi
Keizo Kato | Chenhui Chu | Yugo Murawaki | Sadao Kurohashi
For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable progress. But they typically rely on large numbers of correctly annotated answers to assess reasoning quality. This paper presents a semi-supervised framework that scales reasoning learning from minimal supervision, turning reasoning verification itself into a data creation mechanism. We train a lightweight reasoning-correctness classifier on only a few labeled samples, which judges whether intermediate reasoning traces generated by an LLM are valid. Furthermore, an entropy-based confidence threshold filters out unreliable samples, and the remaining high-confidence reasoning traces are used to fine-tune the model. Experiments on Verifiable Math Problems (Orca-Math subset) and Question Answering on Image Scene Graphs (GQA) with Visual Programming show that our method achieves accuracy comparable to using 10–15× more labeled data. Ablation analyses confirm that both the classifier and entropy filtering are essential for scalable and noise-resistant pseudo-labeling. By replacing expensive answer-level supervision with lightweight reasoning verification, our method provides a practical path toward constructing large-scale reasoning resources and paves the way for future autonomous reasoning systems that learn from minimal human input.
Large language models (LLMs) have demonstrated high performance on tasks expressed in natural language, particularly in zero- or few-shot settings. These are typically framed as supervised (e.g., classification) or unsupervised (e.g., clustering) problems. However, limited work evaluates LLMs as agents in reinforcement learning (RL) tasks (e.g., playing games), where learning occurs through interaction with an environment and a reward system. While prior work focused on representing tasks that rely on a language representation, we study structured, non-linguistic reasoning – such as interpreting positions in a grid world. We therefore introduce PARL (Prompt-based Agent for Reinforcement Learning), a method that uses LLMs as RL agents through prompting, without any fine-tuning. PARL encodes actions, states, and rewards in the prompt, enabling the model to learn through trial-and-error interaction. We evaluate PARL on three standard RL tasks that do not entirely rely on natural language. We show that it can match or outperform traditional RL agents in simple environments by leveraging pretrained knowledge. However, we identify performance limitations in tasks that require complex mathematical operations or decoding states and actions.
This study presents an ensemble technique, SPQ (SVD-Pruning-Quantization), for large language model (LLM) compression that combines variance-retained singular value decomposition (SVD), activation-based pruning, and post-training linear quantization. Each component targets a different source of inefficiency: i) pruning removes redundant neurons in MLP layers, ii) SVD reduces attention projections into compact low-rank factors, iii) and 8-bit quantization uniformly compresses all linear layers. At matched compression ratios, SPQ outperforms individual methods (SVD-only, pruning-only, or quantization-only) in perplexity, demonstrating the benefit of combining complementary techniques. Applied to LLaMA-2-7B, SPQ achieves up to 75% memory reduction while maintaining or improving perplexity (e.g., WikiText-2 reduced from 5.47 to 4.91) and preserving accuracy on downstream benchmarks such as C4, TruthfulQA, and GSM8K. Compared to strong baselines like GPTQ and SparseGPT, SPQ offers competitive perplexity and accuracy while using less memory (6.86 GB vs. 7.16 GB for GPTQ). Moreover, SPQ improves inference throughput over GPTQ, achieving up to a 1.9× speedup, which further enhances its practicality for real-world deployment. The effectiveness of SPQ’s robust compression through layer-aware and complementary compression techniques may provide practical deployment of LLMs in memory-constrained environments. Code is available at: https://github.com/JiaminYao/SPQ_LLM_Compression/
FPSC: A Sustainable Pipeline for Building a Faroese Parliamentary Speech Corpus
Dávid í Lág | Barbara Scalvini | Carlos Daniel Hernandez Mena | Jon Gudnason
Dávid í Lág | Barbara Scalvini | Carlos Daniel Hernandez Mena | Jon Gudnason
This work addresses the lack of large-scale, natural speech data for Faroese automatic speech recognition. Existing resources, such as the 100-hour Ravnursson corpus, consist of read speech and do not capture the spontaneous variation, sociolinguistic aspects and prosody of real dialogue, limiting model performance. To overcome this, we present the Faroese Parliament Speech Corpus (FPSC)—a 1,600-hour collection of parliamentary recordings comprising 89,000 speeches with detailed speaker and linguistic metadata. The corpus includes weakly supervised transcriptions generated using an ensemble of four Faroese-adapted ASR models combined through a ROVER-based voting procedure. In creating FPSC, we trained several new state-of-the-art ASR models for Faroese—some built on large-scale pretrained backbones and others leveraging multilingual transfer—all outperforming previously published Faroese ASR systems. FPSC represents the first corpus of natural spoken Faroese and a major step toward realistic ASR modeling for Faroese, offering an open, reproducible, and scalable resource for future speech and language research.
Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
Peng An-Ci | Kuan-Tang Huang | Tien-Hong Lo | Hung-Shin Lee | Hsin-Min Wang | Berlin Chen
Peng An-Ci | Kuan-Tang Huang | Tien-Hong Lo | Hung-Shin Lee | Hsin-Min Wang | Berlin Chen
Taiwanese Hakka is a low-resource, endangered language that poses significant challenges for automatic speech recognition (ASR), including high dialectal variability and the presence of two distinct writing systems (Hanzi and Pinyin). Traditional ASR models often encounter difficulties in this context, as they tend to conflate essential linguistic content with dialect-specific variations across both phonological and lexical dimensions. To address these challenges, we propose a unified framework grounded in the Recurrent Neural Network Transducers (RNN-T). Central to our approach is the introduction of dialect-aware modeling strategies designed to disentangle dialectal ”style” from linguistic ”content”, which enhances the model’s capacity to learn robust and generalized representations. Additionally, the framework employs parameter-efficient prediction networks to concurrently model ASR (Hanzi and Pinyin). We demonstrate that these tasks create a powerful synergy, wherein the cross-script objective serves as a mutual regularizer to improve the primary ASR tasks. Experiments conducted on the HAT corpus reveal that our model achieves 57.00% and 40.41% relative error rate reduction on Hanzi and Pinyin ASR, respectively. To our knowledge, this is the first systematic investigation into the impact of Hakka dialectal variations on ASR and the first single model capable of jointly addressing these tasks.
Construction of Japanese Prefectural Assembly Minutes Datasets across Three Electoral Terms: Comparative Analysis of 2011, 2015, and 2019 Four-Year Periods
Keiichi Takamaru | Hokuto Ototake | Yuzu Uchida | Yasutomo Kimura
Keiichi Takamaru | Hokuto Ototake | Yuzu Uchida | Yasutomo Kimura
The presented longitudinal cross-regional corpus of Japanese prefectural assembly minutes spans 12 years (2011-2023) across three electoral terms. The corpus comprises 12,236,974 records containing 743,147,226 characters (471,496,688 tokens) of transcribed remarks from the plenary sessions of all 47 prefectural assemblies in Japan. Each dataset is organized by speaker, with assembly members linked to their electoral information, including gender, age, and electoral district. Through a comparative analysis across the three terms, we documented significant temporal changes. The proportion of members aged 25-44 decreased, whereas female representation increased. Female members use 20-30% more characters per speech than male counterparts across all age groups. The proportion of members who never speak varies from under 2% for younger females to over 10% for males aged 65+. We demonstrate the utility of the corpus through three applications: a quantitative analysis of gender and age patterns in political discourse, AI-driven computational dialectology for extracting regional linguistic features, and a web-based search and visualization system. This longitudinal cross-regional corpus provides a valuable resource for interdisciplinary research on subnational politics, computational linguistics, dialectology, and political communication in non-Western democracies. The datasets are available for research purposes upon request, with public query access provided through a web-based interface.
EDDA-Coordinata: An Annotated Dataset of Historical Geographic Coordinates
Ludovic Moncla | Pierre Nugues | Thierry Joliveau | Katherine McDonough
Ludovic Moncla | Pierre Nugues | Thierry Joliveau | Katherine McDonough
This paper introduces a dataset of enriched geographic coordinates retrieved from Diderot and d’Alembert’s eighteenth-century Encyclopédie. Automatically recovering geographic coordinates from historical texts is a complex task, as they are expressed in a variety of ways and with varying levels of precision. To improve retrieval of coordinates from similar digitized early modern texts, we have created a gold standard dataset, trained models, published the resulting inferred and normalized coordinate data, and experimented applying these models to new texts. From 74,000 total articles in each of the digitized versions of the Encyclopédie from ARTFL and ENCCRE, we examined 15,278 geographical entries, manually identifying 4,798 containing coordinates, and 10,480 with descriptive but non-numerical references. Leveraging our gold standard annotations, we trained transformer-based models to retrieve and normalize coordinates. The pipeline presented here combines a classifier to identify coordinate-bearing entries and a second model for retrieval, tested across encoder–decoder and decoder architectures. Cross-validation yielded an 86% EM score. On an out-of-domain eighteenth-century Trévoux dictionary (also in French), our fine-tuned model had an 61% EM score, while for the nineteenth-century, 7th edition of the Encyclopædia Britannica in English, the EM was 77%. These findings highlight the gold standard dataset’s usefulness as training data, and our two-step method’s cross-lingual, cross-domain generalizability.
Mental Health Disorder Detection beyond Social Media: A Systematic Review of Available Datasets
Sadiya Sayara Chowdhury Puspo | Ana-Maria Bucur | Stevie Chancellor | Özlem Uzuner | Marcos Zampieri
Sadiya Sayara Chowdhury Puspo | Ana-Maria Bucur | Stevie Chancellor | Özlem Uzuner | Marcos Zampieri
Detecting mental health disorders in a timely manner is an important societal challenge. NLP and machine learning (ML) methods used to assist with detection rely on data collected primarily from social media. However, such datasets often have sampling biases and inherent ethical and privacy issues. One avenue to overcome these limitations is non-social media data. We present the first comprehensive review of non-social media, free-text datasets for mental health research. We use the PRISMA methodology to conduct our survey and we review datasets available in multiple languages. We find that non-social media free-text based datasets are predominantly focused on English and on detecting depression. These datasets also vary in demographics, platforms, data types, annotation techniques, and methodologies. This systematic review also reveals key gaps and highlights opportunities to develop more diverse, reliable and clinically-relevant resources.
We present a corpus of 196 German counseling conversations (ca. 25k turns) between advice seekers and counselors from nine domains. A subset of 11.5k turns was double-annotated with grounding acts (e.g., acknowledgments, repairs), attempts to advance the conversation, success of advancing, and conversation phases. Baseline classification experiments with logistic regression and GBERT-base illustrate the impact of class imbalance in grounding-act classification. For logistic regression, train-only balancing improves Macro-F1 from 0.417 [0.377–0.434] to 0.444 [0.394–0.478]. For GBERT-base, performance remains competitive (Macro-F1 0.481), with balancing yielding comparable results under the same evaluation protocol. Given the scarcity of German corpora of naturally occurring conversations annotated for grounding phenomena, we provide a novel resource for both conversation analysis and natural language processing, facilitating the design of realistic human-language model interactions in German. Code and data are available at https://osf.io/6k275/overview.
The Prague Discourse Treebank 4.0 is a large genre-diversified language resource with annotation of discourse relations marked by explicit connectives in Czech texts. It consists of 175 thousand sentences with 82 thousand discourse relations. We present the treebank as well as the methods used during the annotation of its individual parts, some of which were annotated fully manually, others using cost-effective partially automatic methods, achieving a comparable quality. The discourse annotation is available in two formats and theoretical frameworks: the Prague discourse annotation on top of deep syntax dependency trees, and the Penn Discourse Treebank style on top of plain texts, using both discourse type/sense taxonomies in both formats. The corpus is publicly and freely available, offering a valuable resource for linguistic research and natural language processing tasks.
Evaluation of Co-Speech Gesture Tracking Techniques in Naturalistic Interactions
Victoria Ivanova | Naomi Harte
Victoria Ivanova | Naomi Harte
Hand gestures convey a significant portion of communicative meaning, making multimodal datasets essential for interaction research. However, annotating gestures remains a time-consuming and challenging task. To speed up the process, semi-automatic methods have been developed that identify segments with hand movement for annotators to refine. These typically combine a pose estimation model with a rule-based or statistical movement detection algorithm. However, most are validated on idealised, non-naturalistic datasets with minimal hand occlusions. We benchmark combinations of four pose estimation methods (OpenPose, MediaPipe, DeepLabCut, and Kinect) and two rule-based movement detection algorithms on two naturalistic, conversational datasets. The best pipelines combine the SPUDNIG displacement algorithm with OpenPose on MULTISIMO and with DeepLabCut on ECOLANG. These pipelines achieved Tversky scores of 0.57 on MULTISIMO and 0.65 on ECOLANG, with recall scores of 0.73 and 0.78, respectively. While off-the-shelf gesture detection systems can support annotation, performance remains limited on naturalistic data, and careful camera setup minimizing occlusions is essential.
Voices across Decades: A Multimodal Diachronic Corpus of German Bundestag Debates (GerParlDia-MM)
Ingo Siegert
Ingo Siegert
This paper presents a multimodal diachronic corpus of German parliamentary debates spanning 1949 – 2025. The dataset focuses on speakers with exceptionally long political careers in the Bundestag, covering at least six parliamentary terms for female and eight for male members, comprising 75 individuals (43 men/32 female) and 2,136 speeches. The corpus integrates audio, video (when available), and official transcripts, enriched with metadata on date, party affiliation, and legislative term. Transcripts were temporally aligned with parliamentary media recordings, and non-speech segments were automatically removed. The corpus enables research on voice aging, intra-speaker variability, and longitudinal political language, and supports benchmarking of ASR and speaker recognition across decades. Thus, this corpus bridges the gap between short-term speech corpora and single-speaker longitudinal datasets, offering a unique foundation for studying change in voice, style, and rhetoric over more than seventy years of German parliamentary history.
We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to generate question/answer pairs related to the Wikipedia article, ensuring that the answer appears verbatim within the article. Next, the question is then rephrased to hinder simple word matching methods from performing well on the dataset. We conduct a crowdsourced human evaluation of the fluency of the generated questions, which included 156 respondents across 30 of the languages (both low- and high-resource). All 30 languages received a mean fluency rating above “mostly natural”, showing that the samples are of good quality. We evaluate 6 different language models, both decoder and encoder models of varying sizes, showing that the benchmark is sufficiently difficult and that there is a large performance discrepancy amongst the languages. Both the dataset and survey evaluations are publicly available.
Manual annotation of linguistic units such as sentences with labels drawn from a large inventory or taxonomy imposes an enormous cognitive load on human subjects. For our exemplary task, we devised a taxonomy of media bias with 37 categories. Selecting the appropriate category (or none) for thousands of news sentences is likely to be tiring and error-prone for humans. To address these type of annotation tasks involving large numbers of labels, we present SALOMO, an annotation tool that pre-selects labels by letting a committee of LLMs make decisions. Human annotators are then tasked mainly with resolving cases where the LLMs disagree. While our tool is independent of any particular task, we describe its design, present a short corpus annotated with a novel fine-grained taxonomy of news bias types as a concrete case study, and demonstrate experimentally both the significant time savings and workload reduction achieved with the pre-selection mechanism, as well as the strong bias it introduces toward the displayed selection. We also provide the mini-dataset of biased sentences and their associated bias types from our experiment.
VietJobs is the first large-scale, publicly available corpus of Vietnamese job advertisements, comprising 48,092 postings and over 15 million words collected from all 34 provinces and municipalities across Vietnam. The dataset provides extensive linguistic and structured information, including job titles, categories, salaries, skills, and employment conditions, covering 16 occupational domains and multiple employment types (full-time, part-time, and internship). Designed to support research in natural language processing and labour market analytics, VietJobs captures substantial linguistic, regional, and socio-economic diversity. We benchmark several generative large language models (LLMs) on two core tasks: job category classification and salary estimation. Instruction-tuned models such as Qwen2.5-7B-Instruct and Llama-SEA-LION-v3-8B-IT demonstrate notable gains under few-shot and fine-tuned settings, while highlighting challenges in multilingual and Vietnamese-specific modelling for structured labour market prediction. VietJobs establishes a new benchmark for Vietnamese NLP and offers a valuable foundation for future research on recruitment language, socio-economic representation, and AI-driven labour market analysis. All code and resources are available at: https://github.com/VinNLP/VietJobs.
A Resource on Dialogical Moves in Native and Non-Native Academic Writers of English
Giulia D’Agostino | Narjes Sheikh Asadi | Elena Musi
Giulia D’Agostino | Narjes Sheikh Asadi | Elena Musi
This paper provides a new approach to the study of linguistic differences in research articles written by native and non-native English writers, including a novel linguistic resource. Conceptually, we propose a functional definition of academic nativeness. Empirically, we operationalize this definition through a survey and the development of a reliable automatic method to distinguish academic natives from non-natives. We then release a corpus of 80 research articles in the field of Linguistics, with introductions manually and reliably annotated for dialogical moves. Preliminary experiments indicate that automatic annotation using large language models remains challenging. Furthermore, the linguistic features that differentiate academic natives and non-natives diverge from those typically associated with general native/non-native English distinctions. These findings aim to inform and enhance academic writing pedagogy, while also offering insights relevant to broader language and corpus studies, as well as computational research.
A Corpus-Based Profiling of Regional English Variants in Global Media: Insights from Olympic Journalism
Felix Mao
Felix Mao
This paper investigates the distinctive linguistic characteristics of regional English variants through a quantitative analysis of global media coverage. The study applies advanced classification techniques, integrating GPT-based embeddings with Support Vector Machines, to a novel corpus, the Olympic Journalism English Variants Corpus. Comprising news articles related to Olympic Games covered by prominent news outlets in the United States, China, Spain, and Mexico between 2020 and 2023, this corpus enables a fine-grained analysis of 164 linguistic features across lexical, syntactic, readability, and sentiment dimensions. The findings reveal strong and interpretable distinctions in features such as verb ratio, nominality, and readability. This study not only demonstrated the enhanced classification capabilities of the model (optimized F1 score = 97.2), but also yielded deeper, data-driven stylistic analysis and insights of each English variant. This work provides a potential template that can be expanded to other World Englishes research.
JFC-Recipe: A Dataset for Nutrient Estimation from Japanese User-Generated Cooking Recipes
Keisuke Shirai | Yoko Yamakata | Hirotaka Kameko | Akiko Sunto | Jun Harashima | Shinsuke Mori
Keisuke Shirai | Yoko Yamakata | Hirotaka Kameko | Akiko Sunto | Jun Harashima | Shinsuke Mori
Estimating nutrients from recipes is essential for performing proper daily dietary control. The nutrients of the recipe could be roughly calculated by identifying the nutrients and weights of each ingredient in the recipe. However, no dataset with fully manual annotations of nutritional values and weights has been released so far, especially for Japanese recipes. In this work, we propose a novel dataset called the Japanese Food Composition Recipe Dataset (JFC-Recipe). The JFC-Recipe dataset consists of two types of annotations: (i) food item annotation that links ingredients in recipes to a database providing nutrients for foods and (ii) amount and unit annotation that are converted into weights in grams using a weight table. We describe a data collection procedure and annotation process, show statistics, and provide inter-annotator agreements to validate the quality of our annotations. In experiments, we tackle two tasks of food item estimation and quantity estimation. Experimental results show that pre-trained language models learn to estimate food items and quantities accurately.
Annotating Conversational Phases and Communication Techniques: A Corpus of German Teacher-Parent Counseling Conversations
Tobias Hallmen | Kathrin Gietl | Karoline Hillesheim | Annemarie Friedrich | Elisabeth André
Tobias Hallmen | Kathrin Gietl | Karoline Hillesheim | Annemarie Friedrich | Elisabeth André
Teacher-parent conversations are critical for student success, yet teachers often lack structured training in counseling communication skills. We present the first annotated corpus of teacher-parent counseling conversations consisting of 59 German dialogues (approximately 6k sentences, 21k annotations) simulated by prospective elementary school teachers, peers, and professional actors. The corpus features theory-grounded annotations for conversational phases (Beginning, Informational, Argumentative, Decision-Making, Concluding) and communication techniques (Paraphrasing, Verbalizing, Structuring). We provide detailed annotation guidelines operationalizing established counseling pedagogy frameworks for computational analysis. Inter-annotator agreement analysis reveals substantial agreement (Fleiss’ k = 0.669 to 0.724, Krippendorff’s a = 0.666 to 0.735). Our analysis reveals confusion patterns, providing insights into counseling discourse structure. Baseline experiments with BERT-based models and open-source LLMs achieve F1 scores of up to 71% depending on task and model. The corpus, guidelines, and baseline code are publicly available under CC BY-NC-SA 4.0 license, enabling research on automated dialogue analysis and AI-based training tools for teacher education.
RO-ABSA: A Romanian Dataset and Baselines for Aspect-Based Sentiment Analysis
Gheorghe Andreea Alina | Andrei Claudia | Ionescu Elena | Ruseti Stefan | Dascalu Mihai
Gheorghe Andreea Alina | Andrei Claudia | Ionescu Elena | Ruseti Stefan | Dascalu Mihai
Despite the increasing use and applicability of sentiment analysis tools, a significant lack of datasets exists for low-resource or limited-resource languages, such as Romanian, which adequately address this task while considering language-specific traits. To overcome this limitation, we introduce a new dataset suitable for Aspect-Based Sentiment Analysis (ABSA) in Romanian, encompassing aspect term categorisation (ATC) and aspect-level sentiment classification (ALSC). Our dataset comprises approximately 6,250 annotated reviews with over 10,600 attributes and their corresponding polarities. We establish comprehensive baselines for each component and for the entire ABSA task. For ABSA, we evaluate two complementary strategies: (1) an end-to-end generative model that produces aspect–sentiment pairs, and (2) a pipeline combining encoder-based ATC and ALSC models. We fine-tune encoder, encoder–decoder, and decoder-only architectures and additionally test transfer learning from English for ATC. Few-shot prompting with LLaMA-3.3 and GPT-4o is also explored for comparison. Fine-tuned models consistently outperform few-shot setups: the best end-to-end ABSA model achieves an F1 score of 0.81, while the ATC and ALSC components reach 0.81 and 0.93 F1, respectively. These results highlight both the challenge of the RO-ABSA dataset and the benefits of supervised fine-tuning for Romanian ABSA.
The Moral Foundations Reddit Corpus
Jackson P. Trager | Alireza S. Ziabari | Elnaz Rahmati | Aida Mostafazadeh Davani | Preni Golazizian | Farzan Karimi-Malekabadi | Ali Omrani | Zhihe Li | Brendan Kennedy | Georgios Chochlakis | Nils Karl Reimer | Melissa Reyes | Kesley Cheng | Mellow Wei | Christina Merrifield | Arta Khosravi | Evans Alvarez | Morteza Dehghani
Jackson P. Trager | Alireza S. Ziabari | Elnaz Rahmati | Aida Mostafazadeh Davani | Preni Golazizian | Farzan Karimi-Malekabadi | Ali Omrani | Zhihe Li | Brendan Kennedy | Georgios Chochlakis | Nils Karl Reimer | Melissa Reyes | Kesley Cheng | Mellow Wei | Christina Merrifield | Arta Khosravi | Evans Alvarez | Morteza Dehghani
Moral framing and sentiment can affect a variety of online and offline behaviors, including donation, environmental action, political engagement, and protest. Various computational methods in Natural Language Processing (NLP) have been used to detect moral sentiment from textual data, but achieving strong performance in such subjective tasks requires large, hand-annotated datasets. Previous corpora annotated for moral sentiment have proven valuable and have generated new insights both within NLP and across the social sciences, but have been limited to Twitter. To facilitate improving our understanding of the role of moral rhetoric, we present the Moral Foundations Reddit Corpus, a collection of 16,123 English Reddit comments that have been curated from 12 distinct subreddits, hand-annotated by at least three trained annotators for 8 categories of moral sentiment (i.e., Care, Proportionality, Equality, Purity, Authority, Loyalty, Thin Morality, Implicit/Explicit Morality) based on the updated Moral Foundations Theory (MFT) framework. We evaluate baselines using large language models (Llama3-8B, Ministral-8B) in zero-shot, few-shot, and PEFT (Parameter-Efficient Fine-Tuning) settings, comparing their performance to fine-tuned encoder-only models like BERT (Bidirectional Encoder Representations from Transformers). The results show that LLMs continue to lag behind fine-tuned encoders on this subjective task, underscoring the ongoing need for human-annotated moral corpora for AI alignment evaluation
A Semi-Automatic Workflow for Transcribing and Annotating Broadcast News
Christoph Draxler | Sven Grawunder | Jürgen Trouvain | Felicitas Kleber
Christoph Draxler | Sven Grawunder | Jürgen Trouvain | Felicitas Kleber
Audio data archived in radio broadcast stations represent a rich source for various research purposes from phonetic questions up to training and test data for speech modelling. We present an efficient semi-automatic workflow for pre-processing, transcribing and analysing large linguistic-phonetic audio corpora. As a pilot study, we process radio broadcast news from a German public radio station containing recordings from 1956 until 2017. The workflow consists of basic preprocessing, automatic speech recognition, manual word correction, automatic generation of pairs of audio chunks and transcripts, plus an automatic word-, syllable- and phoneme-level segmentation of these chunks. The workflow is organised using the Octra Backend management tool, manual validation and correction of transcripts and chunking are performed using the Octra editor, and the BAS web services perform the segmentation. In an example analysis we show with our specific radio corpus how to use it for comparative longitudinal structure analyses of broadcast news, and for text- and signal-based studies on changes of speech and articulation rate.
From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks
Neh Majmudar | Anne Huang | Jinfan Frank Hu | Elena Filatova
Neh Majmudar | Anne Huang | Jinfan Frank Hu | Elena Filatova
In this paper, we examine linguistic puzzles used in high school linguistics competitions, focusing on two common formats: Rosetta Stone and Match-Up. We propose a systematic procedure for converting existing Rosetta Stone puzzles into corresponding Match-Up counterparts. Because linguistic puzzle creation is complex and time-consuming, our method provides an efficient way to accelerate the generation of new puzzles. We evaluate the resulting Rosetta Stone–Match-Up pairs with both human participants and large language models (LLMs). Our results show that both expert human solvers and LLMs display an all-or-nothing pattern on Match-Up puzzles, either solving them completely or failing entirely. This work contributes a new dataset of paired puzzles and provides a detailed evaluation of puzzle difficulty across formats, offering insights into both human and machine linguistic reasoning.
Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
Karin Johanna Denton de Langis | William Walker | Khanh Chi Le | Dongyeop Kang
Karin Johanna Denton de Langis | William Walker | Khanh Chi Le | Dongyeop Kang
We propose an annotation approach that captures not only labels but also the reading process underlying annotators’ decisions, e.g., what parts of the text they focus on, re-read or skim. Using this approach, we conduct a case study on the preference annotation task and create a dataset PreferRead that contains fine-grained annotator reading behaviors obtained from mouse tracking. PreferRead enables detailed analysis of how annotators navigate between a prompt and two candidate responses before selecting their preference. We find that annotators re-read a response in roughly half of all trials, most often revisiting the option they ultimately choose, and rarely revisit the prompt. Reading behaviors are also significantly related to annotation outcomes: re-reading is associated with higher inter-annotator agreement, whereas long reading paths and times are associated with lower agreement. These results demonstrate that reading processes provide a complementary cognitive dimension for understanding annotator reliability, decision-making and disagreement in complex, subjective NLP tasks.
CodeClarity: A Framework and Benchmark for Evaluating Multilingual Code Summarization
Madhurima Chakraborty | Drishti Sharma | Maryam Sikander | Eman Nisar
Madhurima Chakraborty | Drishti Sharma | Maryam Sikander | Eman Nisar
Large Language Models (LLMs) are increasingly used to summarize and document code, yet most research and training data remain limited to English. This creates barriers for developers working in other languages and leaves the multilingual capabilities of LLMs largely unexplored. We present CodeClarity, a framework for evaluating multilingual code summarization across six programming and six natural languages. It combines reference-based metrics, LLM-judge ratings, and faithfulness checks (identifiers and script) to capture surface similarity, semantic adequacy, and code-aware fidelity. Our experiments reveal that lexical metrics penalize morphologically rich languages, while judge-based evaluations provide more stable, semantically aligned assessments. This work establishes the first reproducible foundation for studying multilingual code summarization and points toward fairer, more inclusive evaluation of code intelligence systems. CodeClarity-Bench and the full evaluation pipeline are publicly available at huggingface.co/CodeClarity and github.com/MadhuNimmo/CodeClarity, enabling community-scale human validation and follow-up studies.
A Longitudinal, Multinational, and Multilingual Corpus of News Coverage of the Russo-Ukrainian War
Dikshya Mohanty | Taisiia Sabadyn | Jelwin Rodrigues | Chenlu Wang | Abhishek Kalugade | Ritwik Banerjee
Dikshya Mohanty | Taisiia Sabadyn | Jelwin Rodrigues | Chenlu Wang | Abhishek Kalugade | Ritwik Banerjee
We present DNIPRO, a corpus of 246K news articles from the Russo-Ukrainian war (Feb 2022 – Aug 2024) spanning eleven outlets across five nation-states (Russia, Ukraine, U.S., U.K., China) and three languages. The corpus features comprehensive metadata and human-evaluated annotations for stance, sentiment, and topical framing, enabling systematic analysis of competing geopolitical narratives. It is uniquely suited for empirical studies of narrative divergence, media framing, and information warfare. Our exploratory analyses reveal how media outlets construct incompatible realities through divergent attribution and topical selection without direct refutation of opposing narratives. dnipro empowers empirical research on narrative evolution, cross-lingual information flow, and computational detection of implicit contradictions in fragmented information ecosystems.
SKILL-IR-Discourse: A Large, Annotated Corpus of Argumentation and Domain Discourse on International Relations
Magdalena Wolska | Matti Wiegmann | Sassan Gholiagha | Mitja Sienknecht | Dora Kiesel | Irene Lopez Garcia | Patrick Riehmann | Bernd Fröhlich | Katrin Girgensohn | Jürgen Neyer | Benno Stein
Magdalena Wolska | Matti Wiegmann | Sassan Gholiagha | Mitja Sienknecht | Dora Kiesel | Irene Lopez Garcia | Patrick Riehmann | Bernd Fröhlich | Katrin Girgensohn | Jürgen Neyer | Benno Stein
We present a large annotated corpus of scholarly discourse in the domain of International Relations, a subfield of political science. The corpus comprises 190 articles (over 1500K tokens) annotated at the argumentation, basic rhetorical, and domain level. Five of the included articles (ca. 62K tokens) constitute a Gold-standard, coded by domain experts. The remaining articles were coded by annotators trained on the Gold-standard and monitored for annotation quality. We describe our corpus creation methodology, the annotation process and quality assurance, the corpus itself, and present insights into the data: Most argumentative structures in the data are simple premise-conclusion structures, fewer than half of the claims have explicit supporting evidence. Counter-arguments to claims are rare. The claim-to-support ratio varies widely between articles; possibly to some extent due to the topics covered (with clear common ground) or to the differences between authors’ styles. The distribution of theoretical vs. evaluative statements varies strongly between articles; this can be attributed to such factors as different methodological approaches between the articles and the methodological focus of the publishing journal.
Building Multimodal Corpora Using Microtask Pipelines and Local Annotators
Helmiina Hotti | Raul Vazquez | Anna-Kaisa Jokipohja | Timo Kalliokoski | Henna Paakki | Rosa Suviranta | Tuomo Hiippala
Helmiina Hotti | Raul Vazquez | Anna-Kaisa Jokipohja | Timo Kalliokoski | Henna Paakki | Rosa Suviranta | Tuomo Hiippala
Multimodality, or how human communication and interaction combine multiple forms of expression, is studied across diverse fields of research. Many of these fields have underlined the need for large, richly annotated multimodal corpora to support empirical research. While language resources are increasingly annotated using microtask crowdsourcing, multimodal corpora remain largely reliant on expert annotators, which creates a bottleneck for scalability and broad applicability. This paper presents a novel hybrid approach to multimodal corpus annotation, leveraging the efficiency of microtask pipelines while preserving theoretical rigour. Our approach decomposes the annotation process into sequences of simple, well-instructed tasks, which are then performed by locally recruited non-expert annotators. We demonstrate the feasibility of this approach by presenting a pipeline for annotating the multimodal structure of school textbooks.
Beyond Fake News Detection: A Community-based Study of the Multicultural Nature of Information Disorder
Sara Gemelli | Giulia Di Cristina | Yiran Zhang | Md Azizul Hoque | Alberto De La Torre Solís | Mohamad Mojtaba Behboudi Eshkiki | Nikolai Efimov | Mariia Everstova | Caterina Maria Cappello | Maziar Kianimoghadam Jouneghani | Payam Latifi | Yashar Mahboudi | Farzaneh Mohseni | Dario Placenti | Tommaso Caselli | Manuela Sanguinetti | Aurora Scarpellini | Chiara Zanchi | Usman Naseem | Marco Antonio Stranisci | Simona Frenda
Sara Gemelli | Giulia Di Cristina | Yiran Zhang | Md Azizul Hoque | Alberto De La Torre Solís | Mohamad Mojtaba Behboudi Eshkiki | Nikolai Efimov | Mariia Everstova | Caterina Maria Cappello | Maziar Kianimoghadam Jouneghani | Payam Latifi | Yashar Mahboudi | Farzaneh Mohseni | Dario Placenti | Tommaso Caselli | Manuela Sanguinetti | Aurora Scarpellini | Chiara Zanchi | Usman Naseem | Marco Antonio Stranisci | Simona Frenda
Recognizing disinformation is a challenging task for humans and AI systems. News can be false, misleading, or harmful, and its interpretation often depends on the cultural context of the audience. However, existing datasets rarely account for these contextual and cultural differences, as they are typically not designed from the perspective of news consumers. To address this gap, in this paper, we present the Information Disorder (InDor) corpus, a multilingual dataset of news articles in English, Farsi, Italian, and Russian, annotated for information disorder detection and explanation. The corpus was developed through a participatory process involving contributors from diverse cultural and professional backgrounds, who engaged in data collection, annotation, and evaluation of Large Language Model (LLM) performance on the task. Our findings highlight that false and manipulated news manifest differently across cultural settings, and that current LLMs fail to adequately capture this complexity. This underscores the need for culturally aware computational approaches in the study of information disorder.
FreeTxt-Vi: A Benchmarked Vietnamese-English Toolkit for Segmentation, Sentiment, and Summarisation
Hung Huy Nguyen | Mo El-Haj | Paul Rayson | Dawn Knight
Hung Huy Nguyen | Mo El-Haj | Paul Rayson | Dawn Knight
FreeTxt-Vi is a free and open-source web-based toolkit for creating and analysing bilingual Vietnamese–English text collections. Positioned at the intersection of corpus linguistics and natural language processing (NLP), it enables users to build, explore, and interpret free-text data without requiring programming expertise. The system combines established corpus analysis features such as concordancing, keyword analysis, word relation exploration, and interactive visualisation with modern transformer-based NLP components for sentiment analysis and summarisation. A key contribution of this work is the design of a unified bilingual NLP pipeline that integrates a hybrid VnCoreNLP + Byte Pair Encoding (BPE) segmentation strategy, a fine-tuned TabularisAI sentiment classifier, and a fine-tuned Qwen2.5 model for abstractive summarisation. Unlike existing text analysis platforms, FreeTxt-Vi is evaluated as a set of language processing components. We conduct a three-part evaluation covering segmentation, sentiment analysis, and summarisation, and demonstrate that our approach achieves competitive or superior performance compared to widely used baselines in both Vietnamese and English. By reducing technical barriers to multilingual text analysis, FreeTxt-Vi supports reproducible research and promotes the development of language resources for Vietnamese, a widely spoken but underrepresented language in NLP. The toolkit is applicable to a wide range of domains, including education, digital humanities, cultural heritage, and the social sciences, where qualitative text data are common but often difficult to process at scale.
The Patrologia Graeca Corpus: OCR, Annotation, and Open Release of Noisy Nineteenth-Century Polytonic Greek Editions
Chahan Vidal-Gorène | Bastien Kindt
Chahan Vidal-Gorène | Bastien Kindt
We present the Patrologia Graeca Corpus, the first large-scale open OCR and linguistic resource for nineteenth-century editions of Ancient Greek. The collection covers the remaining undigitized volumes of the Patrologia Graeca (PG), printed in complex bilingual (Greek–Latin) layouts and characterized by highly degraded polytonic Greek typography. Through a dedicated pipeline combining YOLO-based layout detection and CRNN-based text recognition, we achieve a character error rate (CER) of 1.05% and a word error rate (WER) of 4.69%, largely outperforming existing OCR systems for polytonic Greek. The resulting corpus contains around six million lemmatized and part-of-speech tagged tokens, aligned with full OCR and layout annotations. Beyond its philological value, this corpus establishes a new benchmark for OCR on noisy polytonic Greek and provides training material for future models, including LLMs.
National Library as Corpus: DeLiKo-2025@DNB – a Very Large Corpus of German-language Contemporary Literature
Marc Kupietz | Nils Diewald | Philippe Genêt | Andreas Witt
Marc Kupietz | Nils Diewald | Philippe Genêt | Andreas Witt
This paper introduces DeLiKo-2025@DNB, a very large, linguistically annotated corpus of German-language contemporary literature, freely accessible via https://korap.dnb.de/. The corpus currently comprises 21 billion words from over 287,000 books published between 2005 and the present, spanning pulp and genre fiction as well as literary award-winning works. It covers the entire holdings of EPUB-format fiction ebooks deposited with the German National Library (DNB). We provide a detailed account of the corpus composition, metadata, and key features. Additionally, we explain our strategy for enabling lawful and effective access through the deployment of the open-source corpus analysis platform KorAP at the DNB, and we discuss both the transferability of our approach and work to other national libraries and our ongoing and planned extensions and enhancements.
Multi-party Conversational Corpus of L1 and L2 for Speech Alignment Research (Teams-SK): Methodological Approach
Stefan Benus | Viktor Gatial | Erik György | Mária Hricková | Martin Kažimír | Zuzana Kozáčiková | Lucia Mareková | Róbert Sabo | Marian Trnka | Erik Vráb
Stefan Benus | Viktor Gatial | Erik György | Mária Hricková | Martin Kažimír | Zuzana Kozáčiková | Lucia Mareková | Róbert Sabo | Marian Trnka | Erik Vráb
The tendency for speakers to align or accommodate their verbal and non-verbal behaviour to their interlocutors is a fundamental mechanism in spoken interaction, strongly associated with successful communication and social bonding. Despite its ubiquity and documentation across various modalities and linguistic levels (e.g., lexical, prosodic), a lack of comparable, multi-layered linguistic resources and methodological agreement prevents a deeper understanding of its cognitive mechanisms. Multidimensional view of speech alignment might enhance its application in areas like language training or human-machine interaction. This paper addresses these gaps by presenting the development of a multilingual corpus of L1 Slovak and L2 English speech, extending a comparable corpus in L1 English. The corpus utilizes a modified cooperative board game, Forbidden Island, to elicit semi-spontaneous, multi-party conversation and introduces a complementary pair game to specifically target and prime syntactic alignment. The resource includes psychological metadata (e.g., personality, anxiety, perceived dominance) and enables a reproducible methodology for investigating the relationship between entrainment patterns and individual characteristics. By providing a non-Germanic language perspective and a direct L1–L2 comparison framework at prosodic, lexical, pragmatic and syntactic levels, this corpus offers a rich resource for advancing the theoretical understanding, replication, and practical application of speech alignment.
Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
Martina Simonotti | Ludovica Pannitto | Eleonora Zucchini | Silvia Ballarè | Caterina Mauri
Martina Simonotti | Ludovica Pannitto | Eleonora Zucchini | Silvia Ballarè | Caterina Mauri
This paper analyses the implementation of Automatic Speech Recognition (ASR) into the transcription workflow of the KIParla corpus, a resource of spoken Italian. Through a two-phase experiment, 11 expert and novice transcribers produced both manual and ASR-assisted transcriptions of identical audio segments across three different types of conversation, which were subsequently analyzed through a combination of statistical modeling, word-level alignment and a series of annotation-based metrics. Results show that ASR-assisted workflows can increase transcription speed but do not systemically improve accuracy or prosodic annotation quality. Improvements appear to depend on multiple factors, including workflow configuration, conversation type and annotator experience. These findings are therefore yet not generalizable and highlight the complex interplay between transcription expertise, data type and workflow design. Despite current limitations, ASR-assisted transcription, potentially when supported by task-specific fine-tuning, could be integrated into the KIParla transcription workflow to accelerate corpus creation without compromising linguistic and annotation quality. More broadly, this work underscores the potential of semi-automatic transcription for corpus building, especially in complex settings involving multiple speakers and spontaneous, conversational data.
Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts
Seyoung Song | Nawon Kim | Songeun Chae | Kiwoong Park | Jiho Jin | Haneul Yoo | Kyunghyun Cho | Alice Oh
Seyoung Song | Nawon Kim | Songeun Chae | Kiwoong Park | Jiho Jin | Haneul Yoo | Kyunghyun Cho | Alice Oh
The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored in NLP due to a lack of accessible historical corpora. To address this gap, we introduce the Open Korean Historical Corpus, a large-scale, openly licensed dataset spanning 1,300 years and 6 languages, as well as under-represented writing systems like Korean-style Sinitic (Idu) and Hanja-Hangul mixed script. This corpus contains 17.7 million documents and 5.1 billion tokens from 19 sources, ranging from the 7th century to 2025. We leverage this resource to quantitatively analyze major linguistic shifts: (1) Idu usage peaked in the 1860s before declining sharply; (2) the transition from Hanja to Hangul was a rapid transformation starting around 1890; and (3) North Korea’s lexical divergence causes modern tokenizers to produce up to 51 times higher out-of-vocabulary rates. This work provides a foundational resource for quantitative diachronic analysis by capturing the history of the Korean language. Moreover, it can serve as a pre-training corpus for large language models, potentially improving their understanding of Sino-Korean vocabulary in modern Hangul as well as archaic writing systems.
NAIST LIFE STORY: A Seven-Year Crowdsourced Dataset of Japanese Emotion-related Episodes
Kazuhiro Ito | Junko Hayashi | Hiroyuki Nagai | Shoko Wakamiya | Eiji ARAMAKI
Kazuhiro Ito | Junko Hayashi | Hiroyuki Nagai | Shoko Wakamiya | Eiji ARAMAKI
Existing emotion datasets have supported a wide range of NLP tasks, but most are static resources that capture language use only at the time of their creation. As a result, they cannot represent how emotional meanings shift in response to cultural and social change. To address this limitation, we present NAIST LIFE STORY, a seven-year collection of Japanese emotion-related episodes that reflect contemporary topics across multiple years. Since 2017, 1,000 crowdsourced participants per quarter have written short texts describing personal experiences associated with seven emotions: anger, anxiety, disgust, trust, joy, sadness, and surprise. The dataset currently spans 28 periods and includes gender and age information for each participant. Analyses reveal systematic differences in text length and lexical diversity across emotions, as well as clear temporal trends linked to major events such as the COVID-19 pandemic. A preliminary experiment with a large language model shows that using this dataset as contextual evidence improves time-aware emotion inference, demonstrating its value for studying the evolving relationship between emotion and language.
Audience Engagement with Arabic Women’s Social Empowerment and Wellbeing: A Decadal Corpus
Wajdi Zaghouani | Mabrouka Bessghaier | Md. Rafiul Biswas | Shimaa Amer Ibrahim
Wajdi Zaghouani | Mabrouka Bessghaier | Md. Rafiul Biswas | Shimaa Amer Ibrahim
This paper presents the Arabic Women and Society Corpus, a ten-year collection of 252,487 public Arabic Facebook posts related to women’s empowerment and social wellbeing. The corpus was collected from 51,660 pages across 77 countries between 2014 and 2024, resulting in more than 267 million user interactions. Each post includes engagement metrics such as shares, comments, and emotional reactions, providing a unique view of audience sentiment and social attention. The data were processed using an automated pipeline with language identification, normalization, and metadata cleaning to ensure reliability and reproducibility. The corpus enables large-scale analysis of gender discourse, social reform, and emotional engagement across Arabic dialects. It supports research in Arabic natural language processing, computational social science, and digital communication studies. The dataset and accompanying documentation will be released publicly for research use under an open license.
ArPoMeme: An Annotated Arabic Multimodal Dataset for Political Ideology and Polarization
Wajdi Zaghouani | Kais Attia | Md. Rafiul Biswas | Fadhl Eryani
Wajdi Zaghouani | Kais Attia | Md. Rafiul Biswas | Fadhl Eryani
Memes have become a prominent medium of political communication in the Arab world, reflecting how humor, imagery, and text interact to express ideological and cultural positions. Despite the centrality of memes to online political discourse, there is a lack of systematically curated resources for analyzing their multimodal and ideological dimensions in Arabic. This paper presents ArPoMeme, a large-scale dataset of approximately 7,300 Arabic political memes categorized by ideological orientation, including Leftist, Islamist, Pan-Arabist, and Satirical perspectives. The dataset captures the diversity of Arabic meme ecosystems by grounding classification in the self-identification of public Facebook pages and groups that produce and disseminate these memes. To ensure both scale and accuracy, we designed a semi-automated data collection pipeline combining Playwright-based Facebook scraping with Google Drive synchronization, followed by text extraction using the Qwen2.5-VL-7B vision–language model. The extracted text was manually verified and annotated for three polarization dimensions: Us vs. Them framing, Hostility toward out-groups, and Calls to action. Annotation was conducted through a custom Streamlit-based interface supporting distributed labeling, real-time tracking, and version control. The resulting dataset links visual content, textual messages, and ideological orientation, enabling fine-grained analysis of political antagonism, mobilization, and humor. Quantitative analysis of the annotated corpus reveals strong asymmetries in antagonistic framing across ideological groups, with Islamist and satirical memes exhibiting the highest levels of hostility and mobilization cues. The dataset and the annotation tool offer a reproducible and publicly available resource for studying Arabic political discourse, multimodal ideology detection, and polarization dynamics.
Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
Terra Blevins | Stephen Mayhew | Marek Suppa | Hila Gonen | Shachar Mirkin | Vasile Pais | Kaja Dobrovoljc Zor | Voula Giouli | Jun Kevin | Eugene Jang | Eungseo Kim | Jeongyeon Seo | Xenophon Gialis | Yuval Pinter
Terra Blevins | Stephen Mayhew | Marek Suppa | Hila Gonen | Shachar Mirkin | Vasile Pais | Kaja Dobrovoljc Zor | Voula Giouli | Jun Kevin | Eugene Jang | Eungseo Kim | Jeongyeon Seo | Xenophon Gialis | Yuval Pinter
We present Universal NER (UNER) v2, a significant extension of the initial version released in 2024. UNER is a collaborative dataset for multilingual named-entity annotations, built to support research on NER methods in a cross-linguistic setting. UNER v2 adds 11 new datasets in 10 typologically varied languages to the resource, including multiple parallel evaluation benchmarks aligned with each other and other datasets in UNER v1, while maintaining the same annotation guidelines and high standards for inter-annotator agreement. We report detailed statistics for the dataset and benchmark UNER v2 using both encoder-based model architectures and LLMs.
JobArabi: An Arabic Corpus and Analysis of Job Announcements from Social Media
Wajdi Zaghouani | Shimaa Amer Ibrahim | Mabrouka Bessghaier | Houda Bouamor
Wajdi Zaghouani | Shimaa Amer Ibrahim | Mabrouka Bessghaier | Houda Bouamor
This paper introduces JobArabi, a large-scale corpus of Arabic job announcements collected from social media between January 2024 and October 2025. The dataset contains 20,528 public posts from X and captures more than two years of employment-related discourse across Arabic-speaking online communities. The corpus was compiled using a linguistically informed query framework covering 21 Arabic keyword families that reflect gendered, plural, formal, and dialectal expressions of recruitment language. The resulting dataset includes posts from institutional, commercial, and individual accounts and provides metadata such as timestamps, engagement indicators, and geolocation when available, enabling temporal and regional analysis of employment discourse.Quantitative analysis reveals several sociolinguistic patterns in online recruitment, including the persistence of gendered hiring language, regional variation in occupational demand, and the emotional framing of recruitment messages. These findings highlight the potential of Arabic social media as a resource for studying labor market communication and linguistic change.The JobArabi corpus, together with documentation and collection scripts, will be released to support research in Arabic NLP, computational social science, and digital labor studies.
ParaCLEAN: Improving Translation Quality through Systematic Parallel Data Cleaning
Audrey Mash | Ella Paulina Bohman | Maite Melero
Audrey Mash | Ella Paulina Bohman | Maite Melero
Parallel corpora often contain significant noise, particularly in low-resource settings where both collected and synthetic data are combined. We present ParaCLEAN, a modular pipeline for cleaning parallel data that integrates embeddings-based filtering, language identification, deduplication, and normalisation. Experiments on Catalan to Japanese translation demonstrate that ParaCLEAN improves data quality and downstream MT performance. Ablation studies highlight the contribution of each step. ParaCLEAN is lightweight, reproducible, and extensible for diverse language pairs.
We present a proposal for an annotation scheme and data representation of shallow discourse relations annotation in the Universal Dependencies (UD) framework, as a theoretically appropriate and also practically oriented extension of the established morphosyntactic analysis. We outline the design requirements for the annotation scheme, encompassing simplicity, comprehensibility, theoretical grounding, practical applicability and technical robustness, while accommodating the specific constraints of shallow discourse analysis. At the same time, we present a work-in-progress baseline version of DReUD (Discourse Relations in Universal Dependencies), a modular shallow discourse parser for Universal Dependencies as a command-line program, a web client and a REST API service for Czech and English, designed for a seamless and rapid integration of discourse relations analysis both in the theoretical research and in NLP applications.
MultiGraSCCo: A Multilingual Anonymization Benchmark with Annotations of Personal Identifiers
Ibrahim Baroud | Christoph Otto | Vera Czehmann | Christine Hovhannisyan | Lisa Raithel | Sebastian Möller | Roland Roller
Ibrahim Baroud | Christoph Otto | Vera Czehmann | Christine Hovhannisyan | Lisa Raithel | Sebastian Möller | Roland Roller
Accessing sensitive patient data for machine learning is challenging due to privacy concerns. Datasets with annotations of personally identifiable information are crucial for developing and testing anonymization systems, which would enable safe data sharing that complies with privacy regulations. Since accessing real patient data is a bottleneck, synthetic data offers an efficient solution for data scarcity, bypassing privacy regulations that apply to real data. Moreover, neural machine translation can help to create high-quality data for low-resource languages by translating validated real or synthetic data from a high-resource language. In this work, we create a multilingual anonymization benchmark in ten languages, using a machine translation methodology that preserves the original annotations and renders city and people names in a culturally and contextually appropriate form in each target language. Our evaluation study with medical professionals confirms the quality of the translations, both in general and with respect to the translation and adaptation of personal information. Our benchmark with over 2,500 annotations of personal information can be used in many applications, including training annotators, validating annotations across institutions without legal complications, and helping improve the performance of automatic personal information detection. We make our benchmark and annotation guidelines available for further research.
Structured Legal Document Generation in India: A Model-Agnostic Wrapper Approach with VidhikDastaavej
Shubham Kumar Nigam | Deepak Patnaik Balaramamahanthi | Noel Shallum | Kripabandhu Ghosh | Arnab Bhattacharya
Shubham Kumar Nigam | Deepak Patnaik Balaramamahanthi | Noel Shallum | Kripabandhu Ghosh | Arnab Bhattacharya
Automating legal document drafting can improve efficiency and reduce the burden of manual legal work. Yet, the structured generation of private legal documents remains underexplored, particularly in the Indian context, due to the scarcity of public datasets and the complexity of adapting models for long-form legal drafting. To address this gap, we introduce VidhikDastaavej, a large-scale, anonymized dataset of private legal documents curated in collaboration with an Indian law firm. Covering 133 diverse categories, this dataset is the first resource of its kind and provides a foundation for research in structured legal text generation and Legal AI more broadly. We further propose a Model-Agnostic Wrapper (MAW), a two-stage generation framework that first plans the section structure of a legal draft and then generates each section with retrieval-based prompts. MAW is independent of any specific LLM, making it adaptable across both open- and closed-source models. Comprehensive evaluation, including lexical, semantic, LLM-based, and expert-driven assessments with inter-annotator agreement, shows that the wrapper substantially improves factual accuracy, coherence, and completeness compared to fine-tuned baselines. This work establishes both a new benchmark dataset and a generalizable generation framework, paving the way for future research in AI-assisted legal drafting.
PolyglotQL: A Pipeline for Multilingual Text-to-SPARQL Dataset Generation
Julio Perez | Fabio Barth | Georg Rehm
Julio Perez | Fabio Barth | Georg Rehm
We present PolyglotQL, an open-source ETL (Extract, Transform, Load) pipeline for systematically creating multilingual text-to-SPARQL datasets, along with an accompanying framework for evaluating text-to-SPARQL generation models. PolyglotQL provides an extensible and modular architecture that aggregates, normalizes, and augments heterogeneous question–SPARQL pairs from established text-to-SPARQL datasets. With this pipeline, we automatically construct a bilingual English–German dataset featuring contextualized entity and relationship mappings as well as automatically translated and aligned question pairs. We also conduct an empirical evaluation using two multilingual open large language models under two distinct contextualization settings. The results show consistent performance improvements when explicit grounding information is provided, highlighting the benefits of structured context in multilingual semantic parsing.
Building and Annotating a Large Comparable Corpus for Studying Semantic Quantification - Chinese, French, Japanese, Korean
Raoul Blin | Jinnam Choi | WU qishen | Yuxin Zhang | Soonhee Hwang | Takahiro Morita | Alexander Delaporte | Ilaine Wang | Chang Liu
Raoul Blin | Jinnam Choi | WU qishen | Yuxin Zhang | Soonhee Hwang | Takahiro Morita | Alexander Delaporte | Ilaine Wang | Chang Liu
Quantifiers and noun quantification are well-studied topics in linguistics, but, to the best of our knowledge, there are still no dedicated multilingual resources for the study of quantification. To address this gap, we compiled a large multilingual comparable corpus (Chinese, French, Japanese, Korean) and propose to enrich it with both syntactic and “quantificational annotation” (semantic information relevant to the study of quantification). In this paper, we present both the corpus and the annotation project, and report on our initial attempt at quantificational annotation, the challenges encountered, and the linguistic observations drawn from it.
Towards the Generation and Application of Dynamic Web-Based Visualization of UIMA-based Annotations for Big-Data Corpora with the Help of Unified Dynamic Annotation Visualizer
Thiemo Dahmann | Julian Schneider | Philipp Stephan | Giuseppe Abrami | Alexander Mehler
Thiemo Dahmann | Julian Schneider | Philipp Stephan | Giuseppe Abrami | Alexander Mehler
The automatic and manual annotation of unstructured corpora is a routine task in many scientific fields and is supported by a variety of existing software solutions. Despite this variety, few solutions currently support annotation visualization, especially for dynamic generation and interaction. To bridge this gap and visualize annotated corpora based on user-, project-, or corpus-specific aspects, we developed Unified Dynamic Annotation Visualizer (UDAV). UDAV is a web-based solution that implements features not supported by comparable tools, enabling a customizable and extensible toolbox for interacting with annotations and allowing integration into existing big-data frameworks. We exemplify UDAV through a range of visualizations and also provide an evaluation of corpus import and processing performance.
The MultiplEYE Text Corpus: Towards a Diverse and Ever-Expanding Multilingual Text Corpus
Ramunė Kasperė | Anna Bondar | Sergiu Nisioi | Maja Stegenwallner-Schütz | Hanne B. Søndergaard Knudsen | Ana Matić | Eva Pavlinušić Vilus | Dorota Klimek-Jankowska | Chiara Tschirner | Not Battesta Soliva | Deborah N. Jakobi | Cui Ding | Dima Abu Romi | Cengiz Acarturk | Matilda Agdler | Anton Marius Alexandru | Mohd Faizan Ansari | Annalisa Arcidiacono | Elizabete Ausma Velta Barisa | Ana Bautista | Lisa Beinborn | Yevgeni Berzak | Nedeljka Bjelanović | Anna Isabelle Bothmann | Jan Brasser | Caterina Cacioli | Anila Çepani | Ilze Ceple | Adelina Cerpja | Dalí Chirino | Jan Chromý | Alessandro Corona Mendozza | Iria de-Dios-Flores | Nazik Dinçtopal Deniz | Ana Došen | Kristian Elersič | Inmaculada Fajardo | Zigmunds Freibergs | Angelina Ganebnaya | Shan Gao | Jéssica Gomes | Annjo Klungervik Greenall | Alba Haveriku | Miao He | Anamaria Hodivoianu | Yu-Yin Hsu | Amanda Isaksen | Andreia Janeiro | Kristine Jensen de López | Aleksandar Jevremovic | Vojislav Jovanovic | Hanna Kędzierska | Nik Kharlamov | Sara Kosutar | Nelda Kote | Vanja Kovic | Izabela Krejtz | Thyra Krosness | Oleksandra Kuvshynova | Eilam Lavy | Ella Lion | Marta Łockiewicz | Kaidi Lõo | Paula Luegi | Mircea Mihai Marin | Clara Martin | Svitlana Matvieieva | Diane C. Mézière | Xavier Mínguez-López | Valeriia Modina | Jurgita Motiejūnienė | Marie-Luise Müller | Tolgonai Nasipbek kyzy | Jamal Abdul Nasir | Johanne S. K. Nedergård | Ayşegül Özkan | Patrizia Paggio | Marijan Palmović | Maria Christina Panagiotopoulou | Alberto Parola | Helena Pérez | Klaudia Petersen | Anja Podlesek | Eva Pospíšilová | Marta Praulina | Mikuláš Preininger | Loredana Pungă | Diego Rossini | Špela Rot | Habib Sani Yahaya | Irina A. Sekerina | Anne Gabija Skadina | Jordi Solé-Casals | Lonneke van der Plas | Saara M. Varjopuro | Spyridoula Varlokosta | João Veríssimo | Oskari Juhapekka Virtanen | Nemanja Vračar | Mila Vulchanova | Ahmad Mustapha Wali | Peizheng Wu | Nilgün Yücel | Stefan Frank | Nora Hollenstein | Lena Jäger | Somayeh Bakhtiari
Ramunė Kasperė | Anna Bondar | Sergiu Nisioi | Maja Stegenwallner-Schütz | Hanne B. Søndergaard Knudsen | Ana Matić | Eva Pavlinušić Vilus | Dorota Klimek-Jankowska | Chiara Tschirner | Not Battesta Soliva | Deborah N. Jakobi | Cui Ding | Dima Abu Romi | Cengiz Acarturk | Matilda Agdler | Anton Marius Alexandru | Mohd Faizan Ansari | Annalisa Arcidiacono | Elizabete Ausma Velta Barisa | Ana Bautista | Lisa Beinborn | Yevgeni Berzak | Nedeljka Bjelanović | Anna Isabelle Bothmann | Jan Brasser | Caterina Cacioli | Anila Çepani | Ilze Ceple | Adelina Cerpja | Dalí Chirino | Jan Chromý | Alessandro Corona Mendozza | Iria de-Dios-Flores | Nazik Dinçtopal Deniz | Ana Došen | Kristian Elersič | Inmaculada Fajardo | Zigmunds Freibergs | Angelina Ganebnaya | Shan Gao | Jéssica Gomes | Annjo Klungervik Greenall | Alba Haveriku | Miao He | Anamaria Hodivoianu | Yu-Yin Hsu | Amanda Isaksen | Andreia Janeiro | Kristine Jensen de López | Aleksandar Jevremovic | Vojislav Jovanovic | Hanna Kędzierska | Nik Kharlamov | Sara Kosutar | Nelda Kote | Vanja Kovic | Izabela Krejtz | Thyra Krosness | Oleksandra Kuvshynova | Eilam Lavy | Ella Lion | Marta Łockiewicz | Kaidi Lõo | Paula Luegi | Mircea Mihai Marin | Clara Martin | Svitlana Matvieieva | Diane C. Mézière | Xavier Mínguez-López | Valeriia Modina | Jurgita Motiejūnienė | Marie-Luise Müller | Tolgonai Nasipbek kyzy | Jamal Abdul Nasir | Johanne S. K. Nedergård | Ayşegül Özkan | Patrizia Paggio | Marijan Palmović | Maria Christina Panagiotopoulou | Alberto Parola | Helena Pérez | Klaudia Petersen | Anja Podlesek | Eva Pospíšilová | Marta Praulina | Mikuláš Preininger | Loredana Pungă | Diego Rossini | Špela Rot | Habib Sani Yahaya | Irina A. Sekerina | Anne Gabija Skadina | Jordi Solé-Casals | Lonneke van der Plas | Saara M. Varjopuro | Spyridoula Varlokosta | João Veríssimo | Oskari Juhapekka Virtanen | Nemanja Vračar | Mila Vulchanova | Ahmad Mustapha Wali | Peizheng Wu | Nilgün Yücel | Stefan Frank | Nora Hollenstein | Lena Jäger | Somayeh Bakhtiari
We present the MultiplEYE Text Corpus, a large-scale, document-level, multi-parallel resource designed to advance cross-linguistic research on reading and language processing. The corpus provides paragraph-level alignment for texts in 39 languages spanning seven language families and seven scripts. Unlike many existing multilingual corpora, a substantial number of documents were originally written in languages other than English, reducing English-centric bias and supporting more typologically diverse investigations. The texts are carefully selected to balance linguistic richness with experimental feasibility, particularly for eye-tracking-while-reading studies. Developed within a multi-lab initiative, the MultiplEYE Text Corpus follows unified translation, alignment, and experimental design guidelines to ensure cross-linguistic comparability. Its inclusion of texts varying in type and difficulty enables research on discourse- level processing, genre effects, and individual differences across a wide range of languages. The text corpus and accompanying metadata provide a robust foundation for multilingual psycholinguistic and computational modeling research. Data and materials are publicly available at https://doi.org/10.23668/psycharchives.22294.
Sanskrit Travelogue: A Large-Scale Unified and Annotated Corpus of Sanskrit Texts
Giacomo De Luca | Danilo Croce | Roberto Basili
Giacomo De Luca | Danilo Croce | Roberto Basili
We present Sanskrit Travelogue, to our knowledge the largest open, unified and richly annotated Sanskrit corpus. Aggregating eight digital libraries, it comprises 12,394 texts, 73.1M tokens and 9M segments after de-duplication. A reproducible pipeline standardizes transliteration to IAST, reconciles heterogeneous metadata, preserves structural semantics (verse markers, chapter hierarchies, textual apparatus) and adds automatic annotations. We provide corpus-scale morphosyntactic annotation combining two systems: the BYT-5 Sanskrit model for compound and sandhi splitting, and the process-sanskrit library for inflection removal and morphological tagging through a hybrid deterministic-statistical cascade. For each segment we materialize synchronized representations: cleaned, analyzed (sandhi/compound split), stemmed, diacritic-normalized and morphologically tagged. These representations are indexed jointly for retrieval. Both approaches achieve high accuracy (84.61% sentence-level exact matches for BYT-5 segmentation, 92.37% correct root extraction for compounds, 95.94% on the Yoga Sūtra). Manual evaluation on the Yoga Sūtra showed 98% correct root extraction when combining both methods, outperforming individual approaches. These annotations enable searching across orthographic sandhi and within compounds, robust lemma-level retrieval despite rich inflectional variation, and provide training material for segmentation and lemmatization while maintaining ambiguity for downstream modeling. We release the annotated corpus as TSV shards, code for corpus acquisition, processing and annotation, a query normalizer, all under a Creative Commons non-commercial license.
The Foggia Occupator Corpus: Digitisation, Annotation, and Computational Analysis of an Occupation-Era Newspaper (1945-1946)
Michele Ciletti
Michele Ciletti
Historical newspapers are crucial sources yet often remain undigitised or lack machine-readable text. We present the Foggia Occupator corpus, a linguistically enriched, openly licensed resource built from twenty-two issues (Dec 1945–Aug 1946) of a weekly newspaper produced by U.S. personnel in occupied Foggia, Italy. High-resolution scans were processed via OCR with LLM-assisted correction (GPT-4o) and full human verification, then segmented into 874 articles ( 216k tokens). We annotate topics, named entities and typed relations via a semi-automatic pipeline with manual reconciliation, and perform argument mining on civics- and conflict-related content, yielding 1,735 arguments. The entity–relation layer supports network analyses that reveal sparse, modular structures linking military units, civic bodies, and social life. We release TEI-XML with entity spans, JSON article files with metadata, CSVs of entities/relations with temporal counts, and an arguments JSON, all under a Creative Commons 4.0 licence. Beyond documenting an in-between moment of reconstruction, the resource enables benchmarking for OCR-robust NER/RE and studies of framing, stance, and community structure in post-war local media.
SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
Nevidu Jayatilleke | Nisansa de Silva | Uthpala Nimanthi Sooriya-Arachchi | Gagani Kasundhi Kulathilaka | Azra Safrullah | Johan Nevin Sofalas
Nevidu Jayatilleke | Nisansa de Silva | Uthpala Nimanthi Sooriya-Arachchi | Gagani Kasundhi Kulathilaka | Azra Safrullah | Johan Nevin Sofalas
SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication dates, and a historical span from the 5th to the 20th century CE in terms of written dates. The corpus consists of 244k words across 185 literary works that underwent thorough filtering, preprocessing, and copyright compliance checks, followed by extensive post-processing. Additionally, a subset of 59 documents totalling 70k words was annotated based on their written dates. Texts from the National Library of Sri Lanka were selected from the SiDiaC-v.1.0 non-filtered list, which was digitised using Google Document AI OCR. This was followed by post-processing to correct formatting issues, address code-mixing, include special tokens, and fix malformed tokens. The construction of SiDiaC-v.2.0 was informed by practices from other corpora, such as FarPaHC, SiDiaC-v.1.0, and CCOHA. This was particularly relevant for syntactic annotation and text normalisation strategies, given the shared characteristics of low-resource language status between Faroese and the similar cleaning strategies utilised in CCOHA. This corpus is categorised into two layers based on genres: primary and secondary. The primary categorisation is binary, assigning each book to either Non-Fiction or Fiction. The secondary categorisation is more detailed, grouping texts under specific genres such as Religious, History, Poetry, Language, and Medical. Despite facing challenges due to limited resources, SiDiaC-v.2.0 serves as a comprehensive resource for Sinhala NLP, building upon the work previously done in SiDiaC-v.1.0.
ShAnEL-2: A Multilingual Benchmarking Dataset for Short-Answer Language Learning Exercises
Jasper Degraeuwe | Thomas Moerman
Jasper Degraeuwe | Thomas Moerman
Before using GenAI models as EdTech tools, their pedagogical suitability should be corroborated. In this paper, we present ShAnEL-2, a novel multilingual dataset comprising 1,185 student responses to short-answer language learning exercises corrected by teachers. We use ShAnEL-2 to establish an initial benchmark of (1) “off-the-shelf” GenAI models and (2) retrieval-augmented generation (RAG) techniques for the automated correction of this exercise type. With an overall accuracy of 90% and recall of 95%, few-shot RAG (which adds previously corrected responses to the prompt) outperforms the off-the-shelf baseline and textbook RAG setup (which adds coursebook materials) by up to 7 (accuracy) and 5 (recall) percentage points. These results confirm that LLMs learn better from examples than from analysing context and highlight GenAI’s particular potential as a correction assistant for teachers.
The Swedish Parliamentary Motions Corpus 1867-2024
Robert Borges | Fredrik Mohammadi Norén | Lotta Åberg Brorsson | Väinö Yrjänäinen | Hanna Bäck | Robert Klemmensen | Måns Magnusson
Robert Borges | Fredrik Mohammadi Norén | Lotta Åberg Brorsson | Väinö Yrjänäinen | Hanna Bäck | Robert Klemmensen | Måns Magnusson
Motions submitted to the Swedish Parliament are important data for social science and humanities researchers. We introduce a new research corpus, the Swedish Parliamentary Motions Corpus, which is larger and more developed than previously available research corpora for the Swedish motions. The corpus contains annotated and structured parliamentary motions over more than 150 years, through the bicameral parliament (1867–1970) and Sweden’s current unicameral parliament (1971–). Along with the corpus, we describe procedures to measure and ensure transparency around issues related to the data quality of the corpus. In addition, we link motions’ authors to a rich metadata set, ensuring the corpus’s utility in various research applications.
We introduce the Swedish Benchmark of Linguistic Minimal Pairs, a dataset for evaluating syntactic performance in language models. It includes 2,500 minimal pairs organized into 25 syntactic phenomena, with 100 pairs per phenomenon. Each pair contrasts a well-formed and an ill-formed sentence that differ minimally. For each phenomenon, we manually constructed ten pairs from scratch. We semi-automatically generated the remaining 90 pairs and manually adjusted them. A random sample was assessed by 40 participants, who selected the well-formed sentence in 98.05% of cases. We evaluate eleven state-of-the-art models. Results generally show that models handle local agreement well but struggle with certain long-distance dependencies and word order phenomena. Model size seems to matter less than the training domain. Prompt-based evaluation generally lowers performance. We show that model performance is stable across handcrafted and generated subsets and across sample sizes, suggesting that 100 pairs per phenomenon suffice for reliable evaluation. Future work will expand the number of phenomena.
Exploring the Transfer of Irony Explanation Generation from English to Dutch
Aaron Maladry | Els Lefever | Cynthia Van Hee | Veronique Hoste
Aaron Maladry | Els Lefever | Cynthia Van Hee | Veronique Hoste
Explanation generation has gained increasing attention in the field of NLP because it makes the output of classification models more intuitively understandable for humans. This is particularly relevant for complex semantic tasks such as irony detection, where there may not be any explicit linguistic markers. Generative models have shown great potential for irony explanation in earlier work, but most studies have been limited to English. Since this is the highest-resourced language, these capabilities may not be available in languages other than English. To address this gap, this paper analyses the performance of generative models for explanation generation in Dutch, a lower-resourced but closely related language to English. Our work shows that larger proprietary models, like GPT-4, can generate meaningful explanations based on relevant world knowledge, whereas smaller open-source models still struggle to perform this task. Besides quality evaluation, we also analyse the limitations of these models, showing that GPT models struggle most with verbosity and that both open-source and proprietary models exhibit circular reasoning ("this text is ironic because the person expresses this in an ironic way”). Finally, open-source models struggle in particular for Dutch because they fail to produce the relevant world knowledge that is required to understand the irony. All models and data used for the experiments is available at iRONNIE on Hugging Face.
DIDECO: An Annotated Dataset for Intent Detection in Digital Communications
Senaid Popovic | Damien Riquet | Maxime Meyer | Fabien Lauer | Yannick Parmentier
Senaid Popovic | Damien Riquet | Maxime Meyer | Fabien Lauer | Yannick Parmentier
This paper presents DIDECO, the first annotated dataset specifically designed for detecting both explicit and implicit intents in digital communications. We address a critical gap in cybersecurity research by developing a comprehensive taxonomy that distinguishes between explicit communicative goals (what is requested) and implicit persuasion mechanisms (how compliance is engineered). Grounded in Speech Act Theory and persuasion psychology principles, our taxonomy encompasses 20 distinct intent categories across explicit and implicit intents. We annotated 220 LLM-generated spear-phishing emails using a multi-label protocol with six trained annotators, yielding 2,162 intent annotations that reveal the layered complexity of malicious communications. Our analysis demonstrates that sophisticated attacks employ multiple concurrent intents, combining explicit communicative goals with implicit persuasion strategies. This dataset provides resources for developing intent-aware detection systems capable of identifying sophisticated social engineering attacks through semantic analysis.
Bridging is an anaphoric phenomenon where the referent of an entity in a discourse is dependent on a previous, non-identical entity for interpretation, such as in “There is a house. The door is red,” where the door is specifically understood to be the door of the aforementioned house. While there are several existing resources in English for bridging anaphora, most are small, provide limited coverage of the phenomenon, and/or provide limited genre coverage. In this paper, we introduce GUMBridge, a new resource for bridging, which includes 24 diverse genres of English, providing both broad coverage for the phenomenon, and granular annotations for the multi-subtype categorization of bridging varieties. We also present an evaluation of annotation quality and report on baseline performance using open and closed source contemporary LLMs on three tasks underlying our data, showing that bridging resolution and subtype classification remain difficult NLP tasks in the age of LLMs.
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
Kaavya Chaparala | Thomas Thebaud | Jesus Villalba Lopez | Laureano Moro-Velazquez | Peter Viechnicki | Najim Dehak
Kaavya Chaparala | Thomas Thebaud | Jesus Villalba Lopez | Laureano Moro-Velazquez | Peter Viechnicki | Najim Dehak
There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias into datasets. We test ten annotation workflows varying input modality (audio, transcript, or both) and the inclusion of editing (self or peer-editing) to investigate potential quality tradeoffs from using human annotators to summarize audio. We compare human audio-based summaries to human transcript-based summaries to track the impact of the different information modalities on summary quality. We also compare the human outputs against four LLM benchmarks (three text, one audio) to examine whether human-written summaries are less informative than highly fluent automated outputs. We find that audio-based summaries are less informative and more compressed than transcript summaries. However, iterative peer-editing with audio mitigates this difference, enabling audio-based summaries to be as informative as their transcript counterparts and LLM summaries. These findings validate iterative peer-editing among human annotators for the creation of benchmarks informed by both lexical and prosodic information. This enables crucial dataset collection even in setting where transcripts are unavailable.
SEEM-CZ: Annotation and Classification of Epistemic Markers in Czech
Barbora Štěpánková | Michal Novák | Tomáš Musil | Lucie Polakova
Barbora Štěpánková | Michal Novák | Tomáš Musil | Lucie Polakova
We present a project focused on linguistic description, annotation and automatic classification of the so-called epistemic markers in Czech. These expressions, such as pravděpodobně ‘probably’, zřejmě ‘apparently’ and určitě ‘certainly’, typically operate within the pragmatic domain of language. We introduce a dataset containing manual annotations of the 40 most frequent epistemic markers in Czech, totalling almost 4,000 uses. This annotation was created using parallel InterCorp data (in Czech and English) and the TEITOK tool. We describe the annotation scheme used, the annotation process and data handling. The dataset forms the core of the emerging lexical database of these expressions (SEEMLex). Thanks to the comprehensive manual annotation, the dataset can also serve as a source of further pragmatic information and can be used as a basis for further linguistic research. The proposed annotation scheme can also be used for other languages. To demonstrate the dataset’s utility for automatic classification, we trained XLM-RoBERTa classifiers using 10-fold cross-validation, achieving 72.6% accuracy for type of use classification (6 classes) and 54.2% accuracy for degree of certainty classification (4 classes).
When Words Don’t Mean What They Say: Figurative Understanding in Bengali Idioms
Adib Sakhawat | Shamim Ara Parveen | Md Ruhul Amin | Tahera Khatun | Shamim Al Mahmud | Md Saiful Islam
Adib Sakhawat | Shamim Ara Parveen | Md Ruhul Amin | Tahera Khatun | Shamim Al Mahmud | Md Saiful Islam
Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce the Bangla Bagdhara dataset, a large-scale, culturally grounded corpus of 10,361 Bengali idioms. Each idiom is annotated under a comprehensive 19-field schema, established and refined through a deliberative expert consensus process that captures its semantic, syntactic, cultural, and religious dimensions, providing a rich and structured resource for computational linguistics. To establish a robust benchmark for Bangla figurative language understanding, we evaluate 30 state-of-the-art multilingual and instruction-tuned LLMs on the task of inferring figurative meaning. Our results reveal a critical performance gap, with no model surpassing 50% accuracy, in stark contrast to significantly higher human performance (83.4%). This finding underscores the limitations of existing models in cross-linguistic and cultural reasoning. By releasing the Bangla Bagdhara dataset and benchmark, we provide foundational infrastructure for advancing figurative language understanding and cultural grounding in LLMs for Bengali and other low-resource languages.
Human vs LLM in Conversational Repair Annotation: A New Resource and Comparative Study
Anh Ngo | Nicolas Rollet | Catherine Pelachaud | Chloé Clavel
Anh Ngo | Nicolas Rollet | Catherine Pelachaud | Chloé Clavel
Addressing the scarcity of annotated data for Other-Initiated Repair (OIR), when recipients interrupt conversation progressivity to signal trouble, prompting speakers to provide repair, this work introduces OIR annotations for the NOXI corpus, achieving considerable reliability. We evaluate whether LLMs can reliably annotate OIR sequences using structured Chain-of-Thought prompting and conduct comparative analysis across two corpora: NOXI (natural dialogue) and CABB-S (Dutch, task-oriented), finding weak alignment between LLMs and human annotations, particularly in recognizing trouble-signaling. Analyzing human-LLM disagreement using the LLM-generated explanations revealed limitations: models rely on lexical patterns rather than conversational context, construct reasonable-sounding but misleading narratives, highlighting crucial limitations for both automated annotation of complex interactional phenomena.
GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training
Jesse J. Van Oort | Frank Brinkkemper | Erik de Graaf | Bram Vanroy | Saskia Lensink
Jesse J. Van Oort | Frank Brinkkemper | Erik de Graaf | Bram Vanroy | Saskia Lensink
We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present in any other LLM pretraining corpus. Additionally, the corpus includes roughly 207B English, 232B Code, and 48B German/Danish tokens taken from existing sets which we further curated for compliance. This corpus includes curated data from large existing corpora like Common Corpus and Common Crawl, as well as newly created Dutch-specific collections. Most newly created Dutch collections consist of content collected in collaboration with organisations or synthetically augmented content. All data is collected and evaluated with the aim of facilitating the creation of (commercial) language models that are lawful, useful and non-harmful. All data included in the GPT-NL Public Corpus is sourced from datasets with permissive licensing and is curated and redistributed under a CC-BY license. The full dataset is publicly available on the Hugging Face Hub.
Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
Marii Ojastu | Hele-Andra Kuulmets | Aleksei Dorkin | Marika Borovikova | Dage Särg | Kairit Sirts
Marii Ojastu | Hele-Andra Kuulmets | Aleksei Dorkin | Marika Borovikova | Dage Särg | Kairit Sirts
In this paper, we present a localized and culturally adapted Estonian translation of the test set from the widely used commonsense reasoning benchmark, WinoGrande. We detail the translation and adaptation process carried out by translation specialists and evaluate the performance of both proprietary and open source models on the human translated benchmark. Additionally, we explore the feasibility of achieving high-quality machine translation by incorporating insights from the manual translation process into the design of a detailed prompt. This prompt is specifically tailored to address both the linguistic characteristics of Estonian and the unique translation challenges posed by the WinoGrande dataset. Our findings show that model performance on the human translated Estonian dataset is slightly lower than on the original English test set, while performance on machine-translated data is notably worse. Additionally, our experiments indicate that prompt engineering offers limited improvement in translation quality or model accuracy, and highlight the importance of involving language specialists in dataset translation and adaptation to ensure reliable and interpretable evaluations of language competency and reasoning in large language models.
GENIUS Keylog Corpus - a German High School Student Corpus with Keystroke Logging Data
Nils-Jonathan Schaller | Thorben Jansen | Lars Höft | Hannah Pünjer | Andrea Horbach
Nils-Jonathan Schaller | Thorben Jansen | Lars Höft | Hannah Pünjer | Andrea Horbach
Student writing has been studied either as a final product (arguments in an already written text) or as a writing process (keystroke data), but not in an integrated manner. We present Anonymised Keylog Corpus, the first publicly available dataset (as far as we know) that combines both comprehensive argumentative annotations with keystroke logging (259 German argumentative essays written by high school students). Our analysis reveals that 96% of students wrote linearly without recursion and 88% omitted the conclusion section. Writing was mainly characterised by fluent writing without extensive pauses, mainly due to the time limit for completing the task. Additionally we suggest methodology on how to combine annotations with keystroke events and carried out an explorative analysis of writer profiles.
OTA-BOUN: A Historical Turkish Dependency Treebank
Tarık Emre Tıraş | Nureddin Cüneyd Ünal | Ada Cengiz | Ece Yurtseven | Esma F. Bilgin Taşdemir | Saziye Betul Ozates
Tarık Emre Tıraş | Nureddin Cüneyd Ünal | Ada Cengiz | Ece Yurtseven | Esma F. Bilgin Taşdemir | Saziye Betul Ozates
We present OTA-BOUN v2.0, the largest Universal Dependencies treebank for historical Turkish, consisting of 1,742 manually verified sentences sampled from late Ottoman texts. The annotation process followed a semi-automatic methodology: initial pre-annotation by the UDPipe 2.0 pipeline was refined through manual annotation of dependency relations, part-of-speech tags, and lemmas. A distinctive feature of OTA-BOUN is its dual-script representation: each sentence is provided both in the original Perso-Arabic script and its Latinized transcription, while tokens include aligned forms in both scripts. This dual-layer design enables research on script conversion, cross-lingual transfer, and historical–modern Turkish comparisons. Through detailed analyses on the aforementioned treebank, this study presents a unique and scalable resource, advancing computational studies of historical Turkish and supporting broader efforts in multilingual and diachronic NLP.
TCMPHal: A Large-scale Dataset for Hallucination Detection in Traditional Chinese Medicine Pharmacy
Nijia Han | Zimu Wang | Ziwen Xie | Wei Wang | Jia Meng | John Moraros | Shuihua Wang
Nijia Han | Zimu Wang | Ziwen Xie | Wei Wang | Jia Meng | John Moraros | Shuihua Wang
The rapid proliferation of large language models (LLMs) in medicine highlights their potential to revolutionize research in Traditional Chinese Medicine (TCM). While these models have shown great promise in assisting TCM practitioners by answering herb-related questions, generating syndrome-differentiation reports, and recommending classical formulas, a persistent challenge that arises is the issue of hallucination, where LLMs might produce content that appears plausible yet inaccurate. This issue has received limited attention within the context of TCM research, leaving a significant gap in understanding how hallucination manifests within the unique theoretical frameworks and diagnostic principles. Motivated by this phenomenon, we present TCMPHal, the first dataset specifically curated for hallucination detection in TCM pharmacy, comprising 10,000 high-quality question-answer pairs with hallucination annotations. Our experimental results across diverse LLMs, under standard, knowledge-based, and search engine-augmented conditions, demonstrate the capabilities and limitations of these models. A notable observation is that, for thinking LLMs, incorporating search engine results yields minimal improvement over their intrinsic reasoning abilities. We further conduct an in-depth error analysis, paving the way for future research directions in this domain. We release the TCMPHal dataset at https://github.com/hanninaa/TCMP.
AraREQ: A Dataset and End-to-End System for Conflict Detection and Resolution in Software Requirements
Tymaa Hasanain Hammouda | Alaa Aljabari | Nagham Fahim Hamad | Mustafa Jarrar
Tymaa Hasanain Hammouda | Alaa Aljabari | Nagham Fahim Hamad | Mustafa Jarrar
Conflict detection in software requirements is essential for ensuring specification consistency, improving project efficiency, and ensuring overall software quality. Despite its importance, research on this task, particularly for Arabic, remains limited due to the scarcity of annotated data and linguistic challenges. To address this gap, we introduce AraREQ, a large-scale Arabic dataset for requirement-level conflict detection and resolution. The dataset is constructed through a semi-automated Arabization process using Large Language Models (LLMs), followed by manual augmentation to address class imbalance. The final dataset comprises 27K Arabic requirement pairs. We benchmark four state-of-the-art LLMs under zero-shot and few-shot settings, establishing the first comprehensive evaluation for Arabic requirements conflict detection. Experimental results show that few-shot prompting consistently improves performance, particularly on the minority conflict class, demonstrating the effectiveness of example-based prompting. Finally, we introduce an end-to-end system that automatically detects potential conflicts in Arabic software requirements and generates resolution suggestions. All datasets, codes, and the end-to-end system are open-source and available at: https://sina.birzeit.edu/ArReqConflicts/
MAD: A Corpus of Multilingual Argumentative Deliberation
Eimear Maguire | Ella Schad | Jacky Visser | Chris Reed | John Lawrence
Eimear Maguire | Ella Schad | Jacky Visser | Chris Reed | John Lawrence
We present a corpus of Multilingual Argumentative Deliberation (MAD), a manually annotated corpus of deliberative dialogues in English, German, Polish and Italian. Four groups each completed two variants of a ranking task, the NASA Survival Scenario; once in their native language and once in English. The corpus is annotated using Inference Anchoring Theory (IAT), a framework developed for analysing argument in dialogical settings, and widely used in argument mining. As an argument mining resource, MAD is distinct in offering equivalent instances of spontaneous argumentation across languages. In addition to use in argument mining, the annotation captures both argument relations and dialogue acts, enabling deeper analysis of argument and dialogue structure than typical of argument-only corpora. The design of the corpus enables studies of second-language effects in English-medium interaction, cross-linguistic argument comparisons for German, Polish and Italian, and speaker dialogue strategy consistency, amongst others. The primary annotated MAD corpus is freely available at https://corpora.aifdb.org/mad, while we additionally release the unannotated transcripts to facilitate repurposing of the material.
Infox-QC: A Quebec-Focused French Corpus for Misinformation Detection and AI Robustness Assessment
Moetaz Doghmane | Hazem Amamou | Thiziri Sefsaf | Alan Davoust | Anderson Raymundo Avila
Moetaz Doghmane | Hazem Amamou | Thiziri Sefsaf | Alan Davoust | Anderson Raymundo Avila
The pervasive spread of online misinformation, often through social media and political campaigns, makes detecting false claims a crucial task for mitigating societal risks. While the vast majority of fake news datasets are developed in English, a critical gap remains for low-resource languages, such as French. To address this, we introduce Infox-QC, a novel French-language corpus focused on misinformation relevant to the Quebec region. Beyond containing real true and fake news, Infox-QC includes two unique subsets of AI-generated fake news: one created by prompting an AI to paraphrase existing fake news, and a second generated by prompting an AI to fabricate fake news from real true reports. This innovative approach allows us to verify the robustness of detection systems against fabricated content, which modern LLMs can generate with convincing efficacy. We establish comprehensive baselines using traditional machine learning methods, BERT-based models, and Large Language Models, both with and without Retrieval-Augmented Generation (RAG). Our results demonstrate that RAG-augmented LLMs offer the strongest contextual understanding, while traditional models provide valuable interpretable baselines. We further provide an exploratory human–LLM thematic agreement analysis to assess annotation consistency. The Infox-QC resource fills a critical void in French-language NLP research, supporting future efforts to explore the regional and cultural dimensions of misinformation through cross-linguistic comparison.
unarXive 2024: A Large-Scale Scientific Corpus for Citation-Aware Retrieval and Generation
Ines Besrour | Michael Färber
Ines Besrour | Michael Färber
Full-text collections of scientific papers are essential for NLP research and the training of language models. However, existing resources remain incomplete: they often lag behind the fast-paced growth of scientific publishing, lack comprehensive citation networks, and discard essential structural elements. In this work, we introduce unarXive 2024, a large-scale, richly structured corpus containing every arXiv submission from January 1991 to December 2024 – over 2.28 million documents across physics, mathematics, computer science, and other fields. Our release enhances each paper with detailed metadata, reconstructs a substantially more complete citation network than existing datasets, and preserves fine-grained structural information, including section boundaries, mathematical notation, and non-textual elements. Beyond the corpus itself, we provide dense and sparse indexes optimized for retrieval-augmented generation (RAG) over the full arXiv archive. All resources, including code and data, are publicly available: https://github.com/faerber-lab/unarXive-2024
EPIC-EuroParl-UdS: Information-Theoretic Perspectives on Translation and Interpreting
Maria Kunilovskaya | Christina Pollkläsener
Maria Kunilovskaya | Christina Pollkläsener
This paper introduces an updated and combined version of the bidirectional English–German EPIC-UdS (spoken) and EuroParl-UdS (written) corpora containing original European Parliament speeches as well as their translations and interpretations. The new version corrects metadata and text errors identified through previous use, refines the content, updates linguistic annotations, and adds new layers, including word alignment and word-level surprisal indices. The combined resource is designed to support research using information-theoretic approaches to language variation, particularly studies comparing written and spoken modes, and examining disfluencies in speech, as well as traditional translationese studies, including parallel (source vs. target) and comparable (original vs. translated) analyses. The paper outlines the updates introduced in this release, summarises previous results based on the corpus, and presents a new illustrative study. The study validates the integrity of the rebuilt spoken data and evaluates probabilistic measures derived from base and fine-tuned GPT-2 and machine translation models on the task of filler particles prediction in interpreting.
FeedFetcher: A Resilient Web Feed Downloader for Corpus Construction
Ondřej Herman | Jan Kraus | Vit Suchomel
Ondřej Herman | Jan Kraus | Vit Suchomel
Building large-scale, timestamped monitor corpora requires robust and efficient tools for continuous web data acquisition. We present FeedFetcher, an open-source, lightweight yet resilient downloader designed to collect linguistic data from RSS/Atom web feeds. The tool enables continuous corpus updates by harvesting newly published web content with minimal downtime and high data integrity. Implemented in Rust for performance, memory safety, and scalable concurrency, FeedFetcher supports thousands of simultaneous connections while maintaining server politeness. The software is available under the GPL-3.0 license on https://github.com/ondra/feed_fetcher. In our setup, the entire workflow integrates FeedFetcher with downstream text-processing pipelines for tokenization, lemmatization, corpus compilation and deployment. The system is currently used to update monitor corpora in 64 languages, producing approximately two billion tokens per month. These corpora are available in Sketch Engine. We also describe methods for discovering new web feeds, combining manual exploration with automated extraction from large-scale web crawls to expand linguistic coverage. We demonstrate the system’s applicability through a time-based analysis of word-frequency change, showing how long-term accumulation of timestamped data supports the study of lexical dynamics and language evolution.
Human-in-the-Loop Mass Transcription and Ground Truth Annotation for Challenging Historical Documents
Norbert Fischer | Frank Puppe
Norbert Fischer | Frank Puppe
Challenging historical documents still pose significant difficulties for fully automatic layout detection and text recognition, requiring lengthy, demanding correction. We describe our experiences with complex layouts and present our workflow with AdaptOCR, a web-based annotation tool designed to facilitate the efficient transcription and ground-truth annotation of demanding historical documents. Addressing the limitations of existing solutions, AdaptOCR prioritizes a streamlined workflow with an integrated trainable layout and OCR pipeline. The tool uses the PAGE standard to represent document structure and enables the annotation of baselines, regions, text lines and the correction of their transcriptions providing automatic OCR invocation and dictionary-based error detection. Furthermore, it supports flexible annotations with custom element types and attributes to cater to different project requirements. We demonstrate the effectiveness of the workflow and tool in two demanding applications: The transcription of a large corpus of historical printings and the detection / annotation of handwritten artifacts within the private library of the Grimm brothers. In addition, we evaluate the dictionary-based correction and assess the efficiency improvements using AdaptOCR in a pilot study.
CoMMA, a Large-scale Corpus of Multilingual Medieval Archives
Thibault Clérice | Simon Gabay | Malamatenia Vlachou-Efsthatiou | Ariane Pinche | Benoît Sagot
Thibault Clérice | Simon Gabay | Malamatenia Vlachou-Efsthatiou | Ariane Pinche | Benoît Sagot
We present CoMMA, a large-scale corpus of medieval manuscripts produced through automatic text recognition. The corpus contains around 2.5b tokens drawn from more than 23,000 digitized manuscripts in Latin and Old French, harvested via IIIF. Unlike other resources, it is made of raw, non-normalized text enriched with layout analysis in various formats. We describe the pipeline used for large-scale acquisition and processing, and report quantitative and qualitative evaluations (average CER 9.7%). The resulting resource supports multiple use cases, from pretraining language models to corpus linguistic on historical languages and digital humanities applications.
Conversion of the Clark Hall Dictionary of Old English to TEI with RDF: An End-to-end Pipeline for Lexicographic Resource Retrodigitization
Sergei Stoliarov | Maxim Ionov | Fahad Khan | Marina Buzzoni | Francesca Frontini
Sergei Stoliarov | Maxim Ionov | Fahad Khan | Marina Buzzoni | Francesca Frontini
In this submission we introduce a workflow/pipeline for creating TEI editions of legacy dictionaries using a parser based on a context-free grammar (CFG). We do this by describing a project which we are currently carrying out and which aims to create a digital edition of an Old English dictionary, Clark-Hall’s “A Concise Anglo-Saxon Dictionary” using this approach. We begin the article by motivating our CFG-based approach, discussing its advantages and disadvantages, and comparing to it other approaches. We argue that this approach is suitable to certain kinds of dictionaries, such as Clark Hall’s. We then describe the microstructure of the dictionary itself with a view both to justifying the kinds of rules which we subsequently describe and to outlining the kinds of resources to which we believe our approach is best suited. We then describe the CFG parser itself and give an account of our experiments in parsing the dictionary. Finally, we outline the enrichment of the parsed dictionary with RDFa and the benefits it has for the published data.
AMORES: A Spanish Language Resource for an Extended Set of Moral Foundations
Oscar Araque | Daniel Molina | Anny D. Alvarez Nogales | Carlos A. Iglesias
Oscar Araque | Daniel Molina | Anny D. Alvarez Nogales | Carlos A. Iglesias
This work addresses the need for linguistic resources that enable language models to understand and adapt to subjective and abstract concepts in the domain of moral values within texts. In light of the growing interest in the study of moral values and its limited exploration in Spanish-speaking contexts, this work addresses this gap by developing a novel Spanish-language corpus. Furthermore, the corpus’s development process ensures that the annotations capture a wide range of perspectives, resulting in a resource that reflects the diversity of moral interpretations in real-world contexts. Specifically, there are two main contributions. 1 The creation of the first large-scale Spanish corpus annotated according to Moral Foundations Theory. 2 We introduce an experimental framework that investigates how annotators’ religious orientations could shape moral annotation patterns and propagate to model behavior. To do so, we employ a prompt-based alignment method that improves moral detection regardless of religious alignment for which the model was trained. In this scenario, we explore whether language models can align moral interpretations across divergent belief orientations.
The Moralization Corpus: Frame-Based Annotation and Analysis of Moralizing Speech Acts across Diverse Text Genres
Maria Becker | Mirko Sommer | Lars Tapken | Yi Wan Teh | Bruno Brocai
Maria Becker | Mirko Sommer | Lars Tapken | Yi Wan Teh | Bruno Brocai
Moralizations – arguments that invoke moral values to justify demands or positions – are a yet underexplored form of persuasive communication. We present the Moralization Corpus, a novel multi-genre dataset designed to analyze how moral values are strategically used in argumentative discourse. Moralizations are pragmatically complex and often implicit, posing significant challenges for both human annotators and NLP systems. We develop a frame-based annotation scheme that captures the constitutive elements of moralizations – moral values, demands, and discourse protagonists – and apply it to a diverse set of German texts, including political debates, news articles, and online discussions. The corpus enables fine-grained analysis of moralizing language across communicative formats and domains. We further evaluate several large language models (LLMs) under varied prompting conditions for the task of moralization detection and moralization component extraction and compare it to human annotations in order to investigate the challenges of automatic and manual analysis of moralizations. Results show that detailed prompt instructions have a greater effect than few-shot or explanation-based prompting, and that moralization remains a highly subjective and context-sensitive task. We release all data, annotation guidelines, and code to foster future interdisciplinary research on moral discourse and moral reasoning in NLP.
Many European languages possess rich biblical translation histories, yet existing corpora — in prioritizing linguistic breadth — often fail to capture this depth. To address this gap, we introduce a multilingual corpus of 651 New Testament translations, of which 334 are unique, spanning five languages with 2.4–5.0× more translations per language than any prior corpus: English (194 unique versions from 390 total), French (41 from 78), Italian (17 from 33), Polish (29 from 48), and Spanish (53 from 102). Aggregated from 12 online biblical libraries and one preexisting corpus, each translation is annotated with metadata that maps the text to a standardized identifier for the work, its specific edition, and its year of revision. This canonicalization allows researchers to define “uniqueness” for their own needs: they can perform micro-level analyses on translation families, such as the KJV lineage, or conduct macro-level studies by deduplicating closely related texts. By providing the first multilingual resource with sufficient depth per language for flexible, multilevel analysis, the corpus fills a gap in the quantitative study of translation history.
Trigger Warnings Are Grounded in a Shared Vocabulary: A Corpus Analysis with User-Generated Labels
Sebastian Heineking | Matti Wiegmann | Magdalena Wolska | Benno Stein | Martin Potthast
Sebastian Heineking | Matti Wiegmann | Magdalena Wolska | Benno Stein | Martin Potthast
Trigger warnings advise of potentially disturbing content. On that note: This document discusses abuse. But can we trust trigger warnings? For a warning to be credible, independent authors must have a shared understanding of the type of content that advises caution. We investigate for the first time whether trigger warnings are aligned with the vocabulary of texts written by uncoordinated authors. To quantify the lexical alignment of trigger warnings, we conduct a series of statistical tests on the texts of fan fiction authors who used warnings relating to emotional, physical, or sexual abuse. We find that the vocabulary of texts with these warnings is aligned with a curated dictionary of terms related to abuse. However, a high frequency of a term in texts with a warning does not necessarily indicate a semantic relation.
ENEIDE: A High Quality Silver Standard Dataset for Named Entity Recognition and Linking in Historical Italian
Cristian Santini | Sebastian Barzaghi | Paolo Sernani | Emanuele Frontoni | Laura Melosi | Mehwish Alam
Cristian Santini | Sebastian Barzaghi | Paolo Sernani | Emanuele Frontoni | Laura Melosi | Mehwish Alam
This paper introduces ENEIDE (Extracting Named Entities from Italian Digital Editions), a silver standard dataset for Named Entity Recognition and Linking (NERL) in historical Italian texts. The corpus comprises 2,111 documents with over 8,000 entity annotations semi-automatically extracted from two scholarly digital editions: Digital Zibaldone, the philosophical diary of the Italian poet Giacomo Leopardi (1798–1837), and Aldo Moro Digitale, the complete works of the Italian politician Aldo Moro (1916–1978). Annotations cover multiple entity types (person, location, organization, literary work) linked to Wikidata identifiers, including NIL entities that cannot be mapped to the knowledge graph. To the best of our knowledge, ENEIDE represents the first multi-domain, publicly available NERL dataset for historical Italian with training, development, and test splits. We present a methodology for semi-automatic annotations extraction from manually curated scholarly digital editions, including quality control and annotation enhancement procedures. Baseline experiments using state-of-the-art models demonstrate the dataset’s challenge for NERL and the gap between zero-shot approaches and fine-tuned models. The dataset’s diachronic coverage spanning two centuries makes it particularly suitable for temporal entity disambiguation and cross-domain evaluation. ENEIDE is released under a CC BY-NC-SA 4.0 license.
YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset
Peace Busola Falola | Jesujoba Alabi | Solomon O. Akinola | Folashade T. Ogunajo | Emmanuel Oluwadunsin Alabi | David Ifeoluwa Adelani
Peace Busola Falola | Jesujoba Alabi | Solomon O. Akinola | Folashade T. Ogunajo | Emmanuel Oluwadunsin Alabi | David Ifeoluwa Adelani
Named Entity Recognition (NER) is a foundational NLP task, yet research in Yorùbá has been constrained by limited and domain-specific resources. Existing resources, such as MasakhaNER (a manually annotated news-domain corpus) and WikiAnn (automatically created from Wikipedia), are valuable but restricted in domain coverage. To address this gap, we present YoNER, a new multidomain Yorùbá NER dataset that extends entity coverage beyond news and Wikipedia. The dataset comprises about 5,000 sentences and 100,000 tokens collected from five domains including Bible, Blogs, Movies, Radio broadcast and Wikipedia, and annotated with three entity types: Person (PER), Organization (ORG) and Location (LOC), following CoNLL-style guidelines. Annotation was conducted manually by three native Yorùbá speakers, with an inter-annotator agreement of over 0.70, ensuring high quality and consistency. We benchmark several transformer encoder models using cross-domain experiments with MasakhaNER 2.0, and we also assess the effect of few-shot in-domain data using YoNER and cross-lingual setups with English datasets. Our results show that African-centric models outperform general multilingual models for Yorùbá, but cross-domain performance drops substantially, particularly for blogs and movie domains. Furthermore, we observed that closely related formal domains, such as news and Wikipedia, transfer more effectively. In addition, we introduce a new Yorùbá-specific language model (OyoBERT) that outperforms multilingual models in in-domain evaluation. We publicly release the YoNER dataset and pretrained OyoBERT models to support future research on Yorùbá natural language processing.
Linking Rationale to Decision on Internet Standards: A Retrieval-Based Approach Using Synthetic Data
Jie Bian | Michael Welzl
Jie Bian | Michael Welzl
The Internet Engineering Task Force (IETF) develops Internet-Drafts (I-Ds) and Requests for Comments (RFCs) as formal specifications for Internet Protocols. While these documents capture finalized technical standards, the rich design rationales and deliberations that shape them are often buried in informal discussions across mailing lists. These discussions are rarely linked explicitly to the specifications they inform, making it difficult to trace the origins of specific design decisions. We address this gap by generating synthetic data that explicitly links discussion threads to their corresponding RFC/I-D sections, producing roughly 350 000 such aligned instances. This data enables training a semantic embedding-based information retrieval (IR) system that, given an email discussion, retrieves the most relevant specification content. Our experiments show that this synthetic supervision helps models learn associations between informal discourse and formal documentation, though the task remains challenging due to the implicit and context-dependent nature of the links.
This paper introduces GELATO (Government, Executive, Legislative, and Treaty Ontology), a dataset of U.S. House and Senate bills from the 118th Congress annotated using a novel two-level named entity recognition ontology designed for U.S. legislative texts. We fine-tune transformer-based models (BERT, RoBERTa) of different architectures and sizes on this dataset for first-level prediction. We then use LLMs with optimized prompts to complete the second level prediction. The strong performance of RoBERTa and relatively weak performance of BERT models, as well as the application of LLMs as second-level predictors, support future research in legislative NER or downstream tasks using these model combinations as extraction tools.
Controllable Sentence Simplification in Italian: Fine-Tuning Large Language Models on Automatically Generated Resources
Michele Papucci | Giulia Venturi | Felice Dell’Orletta
Michele Papucci | Giulia Venturi | Felice Dell’Orletta
This paper presents a study on readability-controlled Sentence Simplification for Italian, addressing the scarcity of annotated resources for low-resource languages. We introduce IMPaCTS (Italian Multilevel Parallel Corpus for Text Simplification), the first fully automatically created corpus of 1,444,160 original–simple sentence pairs automatically annotated with readability levels and linguistic features. It was generated using an Italian LLM prompted in zero-shot to produce multiple simplifications per input sentence. Increasing portions of the resource are used to fine-tune mono- and multilingual open-weight LLMs, conditioning them to generate simplifications at a target readability level. Results from automatic and human evaluations show that fine-tuning on IMPaCTS improves performance both in terms of task completion and adherence to the targeted readability levels compared to few-shot baselines.
Evaluating LLM-based Text Simplification for German: Effects on Post-Editing Effort, Quality Ratings, and User Comprehension
Luisa Carrer | Andreas Säuberli | Martin Kappus | Lukas Fischer | Sarah Ebling
Luisa Carrer | Andreas Säuberli | Martin Kappus | Lukas Fischer | Sarah Ebling
Automatic text simplification (ATS) seeks to automate the process of rewording within the same language to enhance readability and comprehension. Current evaluation practices for ATS systems predominantly rely on automatic metrics or assessments by experts and crowdworkers, often excluding the intended end users and other stakeholders, and thus limiting insights into the actual effectiveness of ATS models. In this study, we address this gap by conducting a multi-faceted, mixed-method evaluation of two LLM-based ATS systems for German (capito.ai and GPT-4o) and by involving end users, post-editors, and Easy Language experts. The findings highlight the effectiveness of the LLM-based ATS systems examined across several dimensions, including post-editing efficiency, expert quality assessments, and, in the case of GPT-4o-generated simplifications, user comprehension. Post-editing effort metrics, in particular, show an increase in productivity of around 30% compared to full manual simplification. Moreover, the results reveal substantial differences in perception and understanding among participant groups. These outcomes clearly indicate that ATS for German has recently made considerable progress and, crucially, underscore the importance of incorporating multiple stakeholders into ATS evaluation to better align system performance with accessibility goals.
Reading Time in the Wild: An Assessment of Readability Predictors Based on Naturally-Observed Reading Times
Sijbren van Vaals | Rik van Noord | Malvina Nissim
Sijbren van Vaals | Rik van Noord | Malvina Nissim
Reading time has surfaced as a viable proxy for readability and comprehension. However, most studies used reading times obtained in controlled experimental settings with eye-tracking or self-paced reading tasks, which differs from uncontrolled, more naturalistic reading behaviour in the wild. Through a collaboration with a newspaper, we have access to a dataset of Dutch news articles with corresponding clickstream reading times averaged across thousands of readers. To address the issue, we evaluate how well common proxies for readability and comprehension hold on data from online readers. We first group the proxies in four dimensions and compute the correlation between the proxies and the average reading time per token for each dimension. Then we assess if the proxies can meaningfully predict reading time per token. The results are surprising: we find no meaningful correlation between any proxy and the average reading time per token, nor can any proxy be used for reliable prediction. Additionally, we rerun the prediction on corresponding, automatically simplified texts and surprisingly find increased predicted reading times per token. These results imply that clickstream reading time must be considered with caution as a proxy for readability or comprehension.
Document-Level Text Simplification in Estonian Using Large Language Models
Meeri-Ly Muru | Eduard Barbu
Meeri-Ly Muru | Eduard Barbu
Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.
Plain Language and Easy-to-Read formats in text simplification are essential for cognitive accessibility. Yet current automatic simplification and evaluation pipelines remain largely automated, metric-driven, and fail to reflect user comprehension or normative standards. This paper introduces a hybrid framework that explicitly integrates human participation into LLM-based accessible text generation. Human-in-the-Loop (HiTL) contributions guide adjustments during generation, while Human-on-the-Loop (HoTL) supervision ensures systematic post-generation review. Empirical evidence from user studies and annotated resources is operationalized into (i) checklists aligned with standards, (ii) Event-Condition-Action trigger rules for activating expert oversight, and (iii) accessibility Key Performance Indicators (KPIs). The framework shows how human-centered mechanisms can be encoded for evaluation and reused to provide structured feedback that improves model adaptation. By embedding the human role in both generation and supervision, it establishes a traceable, reproducible, and auditable process for creating and evaluating accessible texts. In doing so, it integrates explainability and ethical accountability as core design principles, contributing to more transparent and inclusive NLP systems.
Automatic Analysis of Collaboration through Human Conversational Data Resources: A Review
Yi Yu | Maria Boritchev | Chloé Clavel
Yi Yu | Maria Boritchev | Chloé Clavel
Collaboration is a task-oriented, high-level human behavior. In most cases, conversation serves as the primary medium for information exchange and coordination, making conversational data a valuable resource for the automatic analysis of collaborative processes. In this paper, we focus on verbal aspects of collaboration and conduct a review of collaboration analysis using task-oriented conversation resources, encompassing related theories, coding schemes, tasks, and modeling approaches. We aim to address the question of how to utilize task-oriented human-human conversational data for collaboration analysis. We hope our review will serve as a practical resource and illuminate unexplored areas for future collaboration analysis.
Benchmarking Arabic Authorship Attribution and Style Transfer with Large Language Models
Injy Hamed | Bashar Alhafni | Nizar Habash | Thamar Solorio
Injy Hamed | Bashar Alhafni | Nizar Habash | Thamar Solorio
Writing style is a fundamental component of natural language. However, significant research gaps remain in two key style-centric tasks: authorship attribution (AA) and authorship style transfer, particularly for Arabic. In this work, we revisit both tasks in that context. We introduce a new AA dataset comprising texts in Modern Standard and Dialectal Arabic. We train transformer-based AA models using dual cross-entropy and contrastive learning loss objectives, and validate model performance through human evaluation. We then utilize the trained AA model to benchmark a range of large language models (LLMs) on style recognition and generation tasks, providing new insights into their capabilities in modeling Arabic writing styles. Our work reveals limitations of current models and provides resources to advance research in this direction.
ADHD-Lang: A Large-Scale Social Media Dataset for Verbal Behavior and Digital Phenotyping in Adult ADHD
Daniel Wiechmann | Elma Kerz | Edward Kempa | Yu Qiao
Daniel Wiechmann | Elma Kerz | Edward Kempa | Yu Qiao
We introduce ADHD-Lang, a large-scale language resource derived from Reddit to advance computational phenotyping of adult ADHD. The corpus is constructed using a high-precision self-disclosure pattern to confirm ADHD diagnoses and a matched control cohort, comprising 12,070 ADHD users (317,073 posts; 2.83M sentences) and 12,070 controls (174,765 posts; 1.27M sentences). In releasing ADHD-Lang to the research community, we also provide the first comprehensive baseline results, systematically examining the accuracy–transparency trade-off across three model families: (1) interpretable shallow machine learning models trained on clinically meaningful, expert-engineered language biomarkers; (2) a deep BiLSTM network trained on the same feature representations to capture temporal dynamics across users’ posts; and (3) black-box transformer-based models (BERT, RoBERTa, MentalRoBERTa) leveraging contextual embeddings—non-interpretable, high-dimensional representations. ADHD-Lang is released as a standardized benchmark to promote reproducible research and accelerate progress toward digital verbal-behavior phenotyping for adult ADHD.
SynBullying: A Multi-LLM Synthetic Conversational Dataset for Cyberbullying Detection
Arefeh Kazemi | Hamza Qadeer | Joachim Wagner | Hossein Hosseini | Sri Balaaji Natarajan Kalaivendan | Brian Davis
Arefeh Kazemi | Hamza Qadeer | Joachim Wagner | Hossein Hosseini | Sri Balaaji Natarajan Kalaivendan | Brian Davis
We introduce SynBullying, a synthetic multi-LLM conversational dataset for studying and detecting cyberbullying (CB). SynBullying provides a scalable and ethically safe alternative to human data collection by leveraging large language models (LLMs) to simulate realistic bullying interactions. The dataset offers (i) conversational structure, capturing multi-turn exchanges rather than isolated posts; (ii) context-aware annotations, where harmfulness is assessed within the conversational flow considering context, intent, and discourse dynamics; and (iii) fine-grained labeling, covering various CB categories for detailed linguistic and behavioral analysis. We evaluate SynBullying across five dimensions, including conversational structure, lexical patterns, sentiment/toxicity, role dynamics, harm intensity, and CB-type distribution. We further examine its utility by testing its performance as standalone training data and as an augmentation source for CB classification.
The Multilingual Euphemism Benchmark: Datasets and Baselines for Pragmatic Language Understanding
Whitney Poh | Julia Sammartino | Jasper Andrew | Witold Kieraś | Natalia Zawadzka-Paluektau | Iryna Dilai | Libby Barak | JIng Peng | Anna Feldman
Whitney Poh | Julia Sammartino | Jasper Andrew | Witold Kieraś | Natalia Zawadzka-Paluektau | Iryna Dilai | Libby Barak | JIng Peng | Anna Feldman
Euphemisms are words or phrases used to soften or indirectly refer to taboo or sensitive topics. They pose interpretation challenges because the same expression may appear in different senses depending on context: literal, figurative but non-euphemistic, or euphemistic. For example, pull the plug may refer euphemistically to ending a patient’s life support, figuratively to canceling a project or funding, or literally to unplugging a device. Euphemisms also vary across languages and cultures in both their surface forms and the contexts in which they are conventionally used. Previous work introduced datasets for the computational study of euphemisms in five languages. We extend this line of work by introducing two new annotated datasets for euphemism detection in Polish and Ukrainian and by standardizing resources for all seven languages into a unified benchmark format that supports cross-lingual evaluation. Finally, we provide zero-shot and few-shot baselines using GPT-5-nano. We ran each configuration five times and report the average score, establishing reference scores for multilingual pragmatic understanding. In addition, we performed pilot tests using Qwen3-4B on the English and Chinese datasets.
Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization
Sara Bourbour Hosseinbeigi | Mohammad Hossein Shalchian | Sina Asghari | Mohammad Ali Seif Kashani | Mohammad Amin Abbasi
Sara Bourbour Hosseinbeigi | Mohammad Hossein Shalchian | Sina Asghari | Mohammad Ali Seif Kashani | Mohammad Amin Abbasi
This paper examines the specific obstacles of constructing Retrieval-Augmented Generation (RAG) systems in low resource languages, with a focus on Persian’s complicated morphology and versatile syntax. The research aims to improve retrieval and generation accuracy by introducing Persian-specific models, namely MatinaRoberta (a masked language model) and MatinaSRoberta (a fine-tuned Sentence-BERT), along with a comprehensive benchmarking framework. Three datasets—general knowledge (PQuad), scientifically specialized texts, and organizational reports—were used to assess these models after they were trained on a varied corpus of 73.11 billion Persian tokens. The methodology involved extensive pretraining, fine-tuning with tailored loss functions, and systematic evaluations using both traditional metrics and the Retrieval-Augmented Generation Assessment (RAGAS) framework. The results show that MatinaSRoberta outperformed previous embeddings, achieving superior contextual relevance and retrieval accuracy across datasets. Temperature tweaking, chunk size modifications, and document summary indexing were explored to enhance RAG setups. Larger models like Llama-3.1 (70B) consistently demonstrated the highest generation accuracy, while smaller models faced challenges with domain-specific and formal contexts. The findings underscore the potential for developing RAG systems in Persian through customized embeddings and retrieval-generation settings and highlight the enhancement of NLP applications such as search engines and legal document analysis in low-resource languages.
Corpus and Baselines for Distinguishing Authentic, AI-Generated, and AI-Enhanced Resumes
Andrea Loizidou | Anshu Kiran Sharma | Adrian Esquivel | Mark A. Finlayson | Mustafa Ocal
Andrea Loizidou | Anshu Kiran Sharma | Adrian Esquivel | Mark A. Finlayson | Mustafa Ocal
Job applicants are increasingly turning to generative AI to create or enhance their resumes, leading to challenges in fairness, integrity, and efficiency of modern recruitment processes. We present the first curated corpus of resumes annotated as to whether they are authentic, AI-enhanced, or fully AI-generated. The corpus is balanced across the three classes, comprising 420 resumes spanning five job descriptions in the Information Technology (IT) sector, with the authentic resumes anonymized. We establish strong baselines for this task using traditional and neural supervised machine learning approaches, including Logistic Regression, SVM, Random Forest, XGBoost, BERT, and Longformer. For the featurized approaches, we pair sparse TF-IDF (word/character n-grams) with style features capturing length, punctuation, casing, contractions, lexical diversity (type-token ratio [TTR], number of hapax legomena), n-gram uniqueness, readability indices, and sentiment. Our analysis reveals systematic differences between the classes: AI-generated text features shorter, more uniform sentences, and fewer contractions; AI-enhanced text has the highest uniqueness and TTR; and authentic text has the widest variance across all features. XGBoost is the best performing method, achieving 95.29% accuracy and an F1 of 0.953. We make the corpus available for other researchers to build upon our work. We also benchmark two leading off-the-shelf AI–text detectors on our 420-resume corpus. Despite strong reports in other domains, Originality attains only 55.7% accuracy overall (71/140 authentic, 81/140 AI-generated, 82/140 AI-enhanced correct), and Writer attains 25.0%, with the largest failures on AI-enhanced resumes, highlighting domain shift and cautioning against uncalibrated deployment.
Mute Cods: A Multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection
Katarina Laken | Erik Bran Marino | Paloma Piot | Davide Bassi | Søren Kirkegaard Fomsgaard | Michele Joshua Maggini | Renata Vieira | Marcos Garcia | Sara Tonelli
Katarina Laken | Erik Bran Marino | Paloma Piot | Davide Bassi | Søren Kirkegaard Fomsgaard | Michele Joshua Maggini | Renata Vieira | Marcos Garcia | Sara Tonelli
The proliferation of conspiracy theories and hateful messages on social media poses significant challenges for content moderation and public discourse. Despite their societal impact, existing datasets for automated conspiracy detection remain limited in scope and language coverage. We present a multilingual dataset of conspiracy content on Telegram comprising 5750 messages across English, Dutch, Italian, Spanish and Portuguese from 87 channels documented as disseminating conspiracist and extremist content. Domain experts annotated messages for conspiracist tone, population replacement conspiracy theories, vaccine conspiracies, and hate speech. We extensively report on difficulties and caveats when creating and annotating this type of dataset. We establish classification baselines by evaluating six models in zero-shot fashion and fine-tuning three encoder models, achieving F1 scores up to 0.800 for conspiracist tone, 0.846 for PRCT, 0.843 for vaccine-related conspiracy theories, and 0.734 for hate speech. Inter-annotator agreement was moderate, consistent with the complexity documented in similar annotation tasks.
Push and Pull: Training Sentence Encoders with Contrastive Losses for Distance-Based Multi-Label Text Classification
Jens Van Nooten | Andriy Kosar
Jens Van Nooten | Andriy Kosar
Despite the potential of Distance-Based Classification (DBC), a method that assigns labels to text by measuring semantic similarity between the text and the label representations, it has received very little attention for Multi-Label Text Classification (MLTC). Previous studies have focused on determining optimal thresholds, reaching promising results with contextual sentence encoders. We demonstrate that the performance of these models can be further improved by training them with contrastive losses, i.e., by bringing text representations closer to the corresponding true label representations in an embedding space. Using three supervised contrastive losses and three sentence encoders (Stella, GIST-Large, and BGE), we evaluated our approach on five English datasets (SemEval, BioTech, Reuters, AAPD, and LitCovid) and one Dutch dataset (EventDNA). The results show consistent substantial improvements over base sentence encoders, thereby narrowing the gap between DBC methods and fine-tuned or zero-shot approaches.
PRIVaThe: An Annotated Dataset of Multi-Objectives Web Search Sessions
Claire Ibarboure | Ludovic Tanguy | Franck Amadieu | Josiane Mothe
Claire Ibarboure | Ludovic Tanguy | Franck Amadieu | Josiane Mothe
This paper presents PRIVaThe, a new French-language dataset, consisting of 200 web search sessions from 100 participants performing two multi-objective, multi-hop tasks, designed to enable cross-user comparison of session-level search strategies. Unlike existing datasets that capture only query sequences or final answers, PRIVaThe provides explicit sub-objective decomposition traces for each session. We automatically annotate 3,162 queries with their addressed sub-objective(s) using validated open-weight LLMs (Mistral, LLama3, and Gemma) against human gold annotations. This annotation enables systematic analyses of how users distribute and sequence sub-objectives throughout their sessions, revealing distinct search strategies such as logical, global, and exploratory approaches.
Towards Safer Calls for Everyone: Designing a Benchmark Dataset for Evaluating Voice Phishing Detection Models
Joeun Kang | Gyuri Choi | Chanhyuk Yoon | Yongbin Jeong | Younggyun Hahm | Shea Husband | Hansaem Kim
Joeun Kang | Gyuri Choi | Chanhyuk Yoon | Yongbin Jeong | Younggyun Hahm | Shea Husband | Hansaem Kim
Voice phishing is an evolving form of social engineering crime and requires the continuous advancement of detection technologies. We introduce a benchmark dataset designed to evaluate the practical performance of AI-based voice phishing detection models. The dataset includes diverse voice conversation scenarios and supports four evaluation tasks to assess open-source language models. Experimental results show that while some large-scale models demonstrate stable performance across multiple tasks, accuracy remains low in topic classification and dialogue structure recognition, regardless of model size. These findings highlight the complexity of voice phishing detection, which demands contextual reasoning and dialogue structure understanding beyond simple sentence-level comprehension. The proposed benchmark dataset provides a foundation for more robust evaluation and development of AI systems capable of detecting deceptive voice interactions, contributing to safer and more trustworthy communication environments
Learning Long-Document Embeddings via Chunk–Context Entailment
Waheed Ahmed Abro | Naïm Es-Sebbani | Zied Bouraoui
Waheed Ahmed Abro | Naïm Es-Sebbani | Zied Bouraoui
Learning faithful embeddings for long documents remains challenging, especially in domains like law and medicine where inputs are long, structured, and semantically heterogeneous. We introduce the Chunk Prediction Encoder (CPE), a self-supervised framework that treats chunk–context compatibility as an unsupervised NLI problem. Given a document, CPE masks a chunk and learns (i) a contrastive objective that aligns the masked document with its held-out chunk against in-batch negatives, and (ii) a binary entailment head that predicts whether a candidate chunk belongs to the document. This joint objective encourages both geometric smoothness and directional semantic consistency, yielding robust document-level embeddings. We evaluate CPE with hierarchical and sparse-attention backbones on five benchmarks spanning legal and biomedical domains under frozen-embedding and end-to-end fine-tuning protocols. CPE consistently outperforms baselines, and is more compute-efficient than prompt-only LLM baselines under matched token budgets. Ablations demonstrate the effect of chunk length, the contrastive-vs-entailment balance, and skimming strategies.
Scientific Article Section Classification (SASC) Dataset
Nicolau Duran-Silva | Julian Moreno-Schneider | César Parra-Rojas | Georg Rehm
Nicolau Duran-Silva | Julian Moreno-Schneider | César Parra-Rojas | Georg Rehm
We introduce a novel, publicly available dataset of scientific publications specifically designed to focused on the structural and semantic analysis of their full texts. This collection comprises 4,896 scholarly articles processed using GROBID and self-defined parsers for its segmentation and section parsing. To ensure broad utility and diversity, the dataset includes (≈1,000) papers from 4 specialized research areas: Energy, Cancer, Neuroscience, and Transportation, supplemented by an additional ≈1,000 papers randomly selected from general scientific domains. This dataset is annotated using a newly-defined hierarchical taxonomy comprising 2 levels: the first level contains 9 semantic classes (coarse-grained), while the second level contains 47 semantic classes (fine-grained). All source documents were ethically and legally sourced via OpenAIRE, and the corpus is restricted exclusively to content available under open licenses. License verification was performed through cross-referencing publisher metadata, landing pages, and the Unpaywall database. This curated dataset provides a robust and domain-diverse resource, ideal for developing and evaluating NLP models that require training on hierarchical structure of scientific literature.
JMTEB and JMTEB-lite: Japanese Massive Text Embedding Benchmark and Its Lightweight Version
Shengzhe Li | Masaya Ohagi | Ryokan Ri | Akihiko Fukuchi | Tomohide Shibata | Daisuke Kawahara
Shengzhe Li | Masaya Ohagi | Ryokan Ri | Akihiko Fukuchi | Tomohide Shibata | Daisuke Kawahara
We present JMTEB, a large-scale evaluation suite for Japanese text embedding models, designed to provide comprehensive coverage across multiple task types. The benchmark integrates 28 datasets across 5 tasks, enabling broad and challenging evaluation of model performance in diverse scenarios. While the full benchmark delivers thorough assessment, its scale poses practical challenges in terms of computation time and resource requirements. To address this, we construct JMTEB-lite, a lightweight version of JMTEB, by substantially reducing corpus size in retrieval-related tasks. JMTEB-lite significantly accelerates evaluation while maintaining high fidelity to the full benchmark. Together, JMTEB and JMTEB-lite form a flexible evaluation framework: the full version serves as a comprehensive standard for exhaustive benchmarking, while the lightweight version enables rapid iteration and efficient model selection. This dual approach facilitates both rigorous evaluation and practical development workflows, supporting the advancement of Japanese text embedding research.
Construction of a Japanese RAG Benchmark Using Synthetic Documents on Non-existent Entities and Events
Shengzhe Li | Masaya Ohagi | Hayato Tsukagoshi | Akihiko Fukuchi | Tomohide Shibata | Daisuke Kawahara
Shengzhe Li | Masaya Ohagi | Hayato Tsukagoshi | Akihiko Fukuchi | Tomohide Shibata | Daisuke Kawahara
Retrieval-augmented generation (RAG) is a technique in which a large language model (LLM) generates answers based on relevant documents retrieved from an external document collection. Existing RAG evaluation benchmarks often use public data, such as Wikipedia and news articles, as the external document collection. However, these data are highly likely to be already included in the LLM’s pre-training corpus, which may prevent an accurate evaluation of the model’s ability to generate answers based on the retrieved documents. In this study, we construct a Japanese RAG benchmark by having an LLM synthesize documents about non-existent entities and events and use this collection of synthetic documents as the search target. Since these synthetic documents are not included in the LLM’s training data, the ability to generate answers based on retrieved documents can be evaluated more accurately. In addition to the synthetic documents, the benchmark is composed of questions and correct answers, which are created using a combination of LLMs and human effort. We then evaluated and analyzed the RAG performance of existing LLMs using the constructed benchmark.
C4: A Multilingual Benchmark for Retrieval-Augmented Generation Based on the Catechism of the Catholic Church and Its Compendium
Pius von Däniken | Mark Cieliebak | Jan Deriu
Pius von Däniken | Mark Cieliebak | Jan Deriu
We introduce a new multilingual case study for evaluating retrieval augmented generation (RAG) systems, based on the Catechism of the Catholic Church and its Compendium. The Catechism is a structured document with numbered paragraphs, officially translated into many languages under strict editorial alignment. The Compendium reformulates this material into a question-answer format with explicit citations to the corresponding paragraphs. Together, they form a set of parallel monolingual corpora that share identical semantic structure, enabling direct, controlled comparison of RAG performance across languages. Beyond its theological origin, this text pair closely mirrors real-world applications of RAG in institutional contexts, such as querying internal policy documents with associated FAQ-style summaries, making it a practical testbed for multilingual retrieval and grounded answer generation. We release our data collection scripts and baseline results for further research.
Contrastively Pre-trained Event Embeddings with Schema-free LLM Annotations
Frank Mtumbuka | Steven Schockaert
Frank Mtumbuka | Steven Schockaert
Event extraction is a notoriously challenging problem, among others due to the scarcity of suitable training data. Moreover, event-centric knowledge bases are not available for most domains, making traditional distant supervision strategies difficult to implement. In this paper, we evaluate the potential of using LLM-generated annotations as an alternative distant supervision signal. Specifically, we create a synthetically labelled event extraction corpus, using an LLM to identify event triggers and arguments, and to provide corresponding free-text descriptions. We then pre-train event embedding models on this corpus using a contrastive loss, before fine-tuning them in the usual way. We empirically show the effectiveness of this approach.
A Dataset of Psychiatric Hospital Notes with Temporal Information Annotations
Timothy A. Miller | Gaby Dinh | David Harris | WonJin Yoon | Spencer Thomas | Boyu Ren | Meihua Hall | Guergana Savova
Timothy A. Miller | Gaby Dinh | David Harris | WonJin Yoon | Spencer Thomas | Boyu Ren | Meihua Hall | Guergana Savova
Temporal information extraction is the task of identifying temporal entities in a text and relating them to each other. In medicine, electronic health records (EHRs) contain text that documents the sequence of events during an encounter with a patient, and sometimes the events prior to the encounter (e.g., social history). Temporality is especially important for the specialty of psychiatry. In this work, we describe the updates to the guidelines that allowed us to create a corpus of temporally-annotated psychiatric discharge summaries and progress notes. These updated guidelines were used to create a corpus of over 18000 events, 2200 time expressions, and 13,000 temporal relations. Temporal information extraction performance with a baseline system trained on non-psychiatric data obtains an F1 score of 0.152 on relation extraction, indicating the importance of this new dataset for making progress on temporal information extraction in the psychiatric domain.
Format Matters: A Critical Evaluation of Output Formats for Prompting LLMs in SLU and NER
Pierre Lepagnol | Sahar Ghannay | Thomas Gerald | Christophe Servan | Sophie Rosset
Pierre Lepagnol | Sahar Ghannay | Thomas Gerald | Christophe Servan | Sophie Rosset
Output format is often an unreported factor in LLM evaluations for structured NLP tasks such as Slot Filling or Named Entity Recognition. This work proposes to explore the impact of the output structured format generated by LLMs. We show that measured performance and reliability depend on the requested format (JSON, XML or inline Key-Values). A study is performed across four SLU and three NER benchmarks and considering 13 instruction-tuned open-weight LLMs, using standardized and open-source prompts and parsers. This format-specific evaluation reveals statistically significant swings of 2-46 F1 points depending on model and dataset. Additionally, we propose a lightweight selection procedure to determine the best format per model-dataset combination using only a small development slice; thus reducing trial-and-error in practice.
Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM Approaches
Namu Park | Giridhar Kaushik Ramachandran | Kevin Lybarger | Fei Xia | Özlem Uzuner | Martin Gunn | Meliha Yetisgen
Namu Park | Giridhar Kaushik Ramachandran | Kevin Lybarger | Fei Xia | Özlem Uzuner | Martin Gunn | Meliha Yetisgen
Large language models (LLMs) have shown considerable promise in clinical natural language processing, yet few domain-specific datasets exist to rigorously evaluate their performance on radiology tasks. In this work, we introduce an annotated corpus of 6,393 radiology reports from 586 patients, each labeled for follow-up imaging status, to support the development and benchmarking of follow-up adherence detection systems. Using this corpus, we systematically compared traditional machine-learning classifiers—logistic regression (LR), support vector machines (SVM), Longformer, and a fully fine-tuned Llama3-8B-Instruct—with recent generative LLMs. To evaluate generative LLMs, we tested GPT-4o and the open-source GPT-OSS-20B under two configurations: a baseline (Base) and a task-optimized (Advanced) setting that focused inputs on metadata, recommendation sentences, and their surrounding context. A refined prompt for GPT-OSS-20B further improved reasoning accuracy. Performance was assessed using precision, recall, and F1 scores with 95% confidence intervals estimated via non-parametric bootstrapping. Inter-annotator agreement was high (F1 = 0.846). GPT-4o (Advanced) achieved the best performance (F1 = 0.832), followed closely by GPT-OSS-20B (Advanced; F1 = 0.828). LR and SVM also performed strongly (F1 = 0.776 and 0.775), underscoring that while LLMs approach human-level agreement through prompt optimization, interpretable and resource-efficient models remain valuable baselines.
Efficient Topic Extraction via Graph-Based Labeling: A Lightweight Alternative to Deep Models
Salma Mekaoui | Hiba Sofyan | Imane Benchrif | Imane Amaaz | Ilham Chaker | Arsalane Zarghili | Nikola S. Nikolov
Salma Mekaoui | Hiba Sofyan | Imane Benchrif | Imane Amaaz | Ilham Chaker | Arsalane Zarghili | Nikola S. Nikolov
Extracting topics from text has become an essential task, especially with the rapid growth of unstructured textual data. Most existing works rely on highly computational methods to address this challenge. In this paper, we argue that probabilistic and statistical approaches, such as topic modeling (TM), can offer effective alternatives that require fewer computational resources. TM is a statistical method that automatically discovers topics in large collections of unlabeled text; however, it produces topics as distributions of representative words, which often lack clear interpretability. Our objective is to perform topic labeling by assigning meaningful labels to these sets of words. To achieve this without relying on computationally expensive models, we propose a graph-based approach that not only enriches topic words with semantically related terms but also explores the relationships among them. By analyzing these connections within the graph, we derive suitable labels that accurately capture each topic’s meaning. We present a comparative study between our proposed method and several benchmarks, including ChatGPT-3.5 (CITATION), across two different datasets. Our method achieved consistently better results than traditional benchmarks in terms of BERTScore and cosine similarity and produced results comparable to ChatGPT-3.5, while remaining computationally efficient. Finally, we discuss future directions for topic labeling and highlight potential research avenues for enhancing interpretability and automation.
From Noise to Signal: When Outliers Seed New Topics
Evangelia Zve | Gauvain Bourgne | Benjamin Icard | Jean-Gabriel Ganascia
Evangelia Zve | Gauvain Bourgne | Benjamin Icard | Jean-Gabriel Ganascia
Outliers in dynamic topic modeling are often discarded as noise, yet some act as early signals of emerging topics. We introduce a temporal taxonomy of news document trajectories that distinguishes anticipatory outliers, documents that appear before a topic forms but later integrate into it, from those that reinforce existing topics or remain isolated. This taxonomy bridges weak-signal detection and dynamic topic modeling, clarifying how individual articles anticipate, initiate, or drift within evolving clusters. We implement it within a cumulative clustering framework using document- embeddings from eleven state-of-the-art language models and apply it retrospectively to HydroNewsFr, a French news corpus on the hydrogen economy curated for this study. Inter-model agreement on anticipatory outliers indicates that a small high-agreement subset yields robust confidence estimates. Complementary qualitative case studies further demonstrate their potential value as early indicators of emerging narratives. All reproducibility materials and results are available at https://anonymous.4open.science/status/lrec_from_noise_to_signal-B721.
Explore Political Discourse with Transformers. Emergent Paradigmatic and Syntagmatic Representations.
Laurent Vanni | Damon Mayaffre
Laurent Vanni | Damon Mayaffre
Textual data analysis lies at the heart of inductive reasoning in corpus linguistics. Corpus-driven approaches place the corpus at the center of working hypotheses and use statistical processing as an exploratory tool. With deep neural networks, the training corpus is also crucial, but the objectives are less exploratory. Nevertheless, the performance of Transformers in automatic language processing suggests that self-attention is an effective means of extracting structural information from corpora. In this article, we present interdisciplinary work that uses Transformers descriptively to shed light on linguistic phenomena present in a learning corpus. We propose using two feature-based interpretation methods in a case study of political speeches applied to a text generation task. The first method is a global approach that uses attention scores to analyse the training corpus. The second is a local approach that uses gradient-based features to analyse predictions. These methods are compared to standard statistical techniques, providing empirical confirmation of the observed phenomena. We conclude on the potential of Transformers as a heuristic tool for corpus linguistics.
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
Taja Kuzman Pungeršek | Peter Rupnik | Vit Suchomel | Nikola Ljubešić
Taja Kuzman Pungeršek | Peter Rupnik | Vit Suchomel | Nikola Ljubešić
Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general corpora for this language group: the CLASSLA-web 1.0 corpora. Building on this success, we established a continuous crawling infrastructure for iterative national top-level domain crawling across South Slavic and related webs. We present the first outcome of this crawling infrastructure - the CLASSLA-web 2.0 corpus collection, with substantially larger web corpora containing 17.0 billion words in 38.1 million texts in seven languages: Bosnian, Bulgarian, Croatian, Macedonian, Montenegrin, Serbian, and Slovenian. In addition to genre categories, the new version is also automatically annotated with topic labels. Comparing CLASSLA-web 2.0 with its predecessor reveals that only one-fifth of the texts overlap, showing that re-crawling after just two years yields largely new content. However, while the new web crawls bring growing gains, we also notice growing pains - a manual inspection of top domains reveals a visible degradation of web content, as machine-generated sites now contribute a significant portion of texts.
MaritimEmails: A Synthetic Dataset for Maritime Chartering Correspondence
Kevin Bruendler | Simon Clematide
Kevin Bruendler | Simon Clematide
We introduce MaritimEmails, a large-scale synthetic corpus of 19,817 English-language email threads simulating maritime chartering negotiations between brokers and charterers. Email remains a dominant medium for business communication, yet no public corpora exist for this highly specialized domain due to confidentiality constraints. To address this gap, we generate domain-plausible negotiation exchanges using five contemporary language models under multiple prompting strategies, including Attribute Prompting and Base–Refine (BARE) approaches. Each thread includes structured annotations for vessels, ports, commodities, and Incoterms, enabling supervised training for information extraction and related tasks. Our comparative evaluation covering lexical and semantic diversity, sentiment balance, and verbosity shows that BARE generation increases linguistic variation while maintaining coherence. However, all models exhibit a systematic positivity bias, yielding less negative sentiment than is observed in the Enron reference corpus and likely also in many real negotiation settings. Baseline information extraction experiments with GLiNER and generative Qwen models yield up to 0.86 macro F1 on entity extraction, supporting the dataset’s usefulness. MaritimEmails, together with prompts, scripts, and documentation, is released for research use.
eSciBench: An Extensible Scientific PDF Extraction Benchmark
Noah Tremblay Taillon | Phillippe Langlais
Noah Tremblay Taillon | Phillippe Langlais
Automatically extracting information from PDF documents (such as authors, affiliations, references, tables, equations) may be transformative in Digital Humanities where meta-data accompanying a document is typically manually collected, a cumbersome process. In this work, we conduct a systematic benchmarking of PDF extractors on a set of 100 scientific articles (1949 pages) of the STEM domain that have been processed automatically, then carefully curated. Our benchmark, named eSciBench is openly accessible. Putting to the test 13 extractors on it reveals that although some extractors perform well overall, extracting information from scientific articles is far from a solved problem.
Vrittanta-AS: Dataset Development and Benchmarking for Event Trigger Detection and Classification in Assamese
Chaitanya Kirti | Dhrubajyoti Pathak | Ashish Anand | Prithwijit Guha
Chaitanya Kirti | Dhrubajyoti Pathak | Ashish Anand | Prithwijit Guha
Event trigger detection and classification aim to identify and categorize events within unstructured text. While prior research has primarily focused on news or biomedical corpora, the literary domain, especially short stories, remains largely underexplored. This gap is particularly pronounced for low-resource languages such as Assamese, where limited annotated data and complex narrative structures hinder progress. To address this challenge, we introduce Vrittanta-AS, a manually curated Assamese event trigger detection and classification dataset comprising 13,171 annotated events extracted from short stories. The dataset is designed to advance research in information extraction and narrative understanding for low-resource Indian languages. We conduct a comprehensive evaluation using classical machine learning methods, neural sequential architectures, pre-trained transformer models, and large language models (LLMs) on the proposed dataset. Experimental results demonstrate that IndicBERT v2 achieves the highest performance for both event trigger detection (85.86% micro-F1) and classification (65.21% macro-F1). Vrittanta-AS serves as an important step toward developing benchmark resources for event trigger detection and classification in Assamese literary text.
From Facts to Hypotheses: Joint Detection of Biomedical Relations and Epistemic Commitment Using LLMs
Aleksandra Gabryszak | Phuc Tran Truong | Arne Binder | Nikola Milosevic | Felix-Sebastian Keese | Astrid Rheinländer | Philippe Thomas
Aleksandra Gabryszak | Phuc Tran Truong | Arne Binder | Nikola Milosevic | Felix-Sebastian Keese | Astrid Rheinländer | Philippe Thomas
Determining the factual status of biomedical statements, whether affirmed, negated, or uncertain, is essential for accurate understanding. To support research in this area, we introduce BioRelFact, a publicly available, expert-annotated dataset of 1,767 English biomedical sentences labeled with nine relation types and five levels of epistemic commitment. Using this dataset, we evaluate eight large language models (LLMs) from the GPT, Qwen, and Gemma families for joint relation extraction and epistemic classification. Among the evaluated models, GPT-OSS-20B performs best in both tasks (F1 77.3 for relation, 65.3 for commitment), followed by GPT-4o (75.9 and 60.2), while Qwen3-8B (Thinking) shows strong performance despite its smaller size (74.6 and 57.2). Domain adaptation has mixed effects: relative to their general-purpose counterparts, MedGemma-27B improves (+3.6 F1 for relation, +4.4 for factuality), whereas Qwen2.5-Aloe-Beta-7B declines (–4.3 and –3.5, respectively). Moreover, definition-based few-shot prompts consistently yield the best results for most models, and an explorative analysis of prediction errors suggests which specific linguistic features may drive model confusions.
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
Luca Foppiano | Sotaro Takeshita | Pedro Ortiz Suarez | Ekaterina Borisova | Raia Abu Ahmad | Malte Ostendorff | Fabio Barth | Julian Moreno-Schneider | Georg Rehm
Luca Foppiano | Sotaro Takeshita | Pedro Ortiz Suarez | Ekaterina Borisova | Raia Abu Ahmad | Malte Ostendorff | Fabio Barth | Julian Moreno-Schneider | Georg Rehm
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. The dataset construction and processing workflow demonstrates how open-source tools can enable large-scale, scientific data curation while maintaining high data quality. Finally, we pre-train a RoBERTa model on our dataset and evaluate it across a comprehensive set of benchmarks, achieving performance comparable to other scientific language models of similar size, validating the quality and utility of SciLaD. We publish the dataset and evaluation pipeline to promote reproducibility, transparency, and further research in natural scientific language processing and understanding including scholarly document processing.
CausalSense: Leveraging Common Sense Knowledge and LLMs for Joint Event Extraction and Relation Classification
Youssra REBBOUD | Pasquale Lisena | Raphael Troncy
Youssra REBBOUD | Pasquale Lisena | Raphael Troncy
Event Relation Extraction (ERE) aims to identify and classify semantic relationships between events expressed in text. While existing work has mainly addressed temporal or simple causal links, fine-grained causal relations such as enable, prevent, and intend remain insufficiently explored, partly due to limited and imbalanced labeled datasets. We present a novel framework that leverages large language models (LLMs) and common-sense knowledge to jointly perform event extraction and relation classification. Our contribution includes (1) the creation of the CausalSense large-scale dataset containing more than 500k sentences from news data and commonsense knowledge extracted from ATOMIC, and enriched synthetically; and (2) the evaluation of multiple architectures, including transformer-based models and end-to-end multitask systems for extracting fine-grained causal relationships. Experimental results show that our best-performing model achieves a 32.3% improvement in average F1-score over the current state of the art. The integration of commonsense knowledge substantially enhances fine-grained causal relation detection. The CausalSense dataset, our code and models are released as open source to support future research on causal event relationship extraction.
This paper systematically evaluates modern large language models for automatic term extraction (ATE), examining GPT-5 and Mistral across four domains and three languages using the ACTER corpus. The study compares model sizes, evaluates reasoning-enhanced variants, and tests prompting strategies aligned with human annotation guidelines. Beyond extracting term lists, models provide term labels, confidence scores, and terminology management remarks. Current large language models achieve F1 scores of .36-.72; while seemingly low, this is competitive with supervised approaches and approaches the human inter-annotator agreement ceiling of ~0.59. Larger models outperform smaller variants, with reasoning-enhanced models showing modest improvements. Qualitative error analysis reveals that evaluation methodology partly misrepresents model capabilities: many extractions classified as errors represent defensible boundary judgements, and apparent hallucinations are predominantly (though not exclusively) valid normalisations. Limitations remain in fine-grained categorisation and handling overly general expressions. However, the convergence of model scores with each other and with human inter-annotator agreement suggests that, for high-resource languages, basic ATE may no longer be the bottleneck in terminology management pipelines, and research should shift toward downstream tasks such as definition generation and ontology construction.
A Large-Scale Dataset for Linking-Based Geocoding
Hibiki Nakatani | Yuichiro Yasui | Ryosuke Wakamoto | Masayuki Ishii | Tetsuhisa Suizu | Hiroki Ouchi | Taro Watanabe
Hibiki Nakatani | Yuichiro Yasui | Ryosuke Wakamoto | Masayuki Ishii | Tetsuhisa Suizu | Hiroki Ouchi | Taro Watanabe
Linking-based geocoding is the task of linking location mentions in text to their corresponding entries in a geographic database (Geo-DB) and assigning precise coordinates. Although the task and its technology are essential for spatial information extraction, existing datasets are manually curated and lack sufficient data for training accurate models. To address this limitation, we automatically construct a large-scale dataset for linking-based geocoding by leveraging publicly available resources to generate data efficiently at scale. Specifically, we align location mentions in the first paragraphs of Japanese Wikipedia articles with their associated Wikidata entries containing geographic attributes. Wikipedia provides natural textual contexts, while Wikidata offers structured data such as coordinates, place types, and administrative divisions, which can serve as rich metadata for future extensions. Our experiments show that models trained on our dataset achieve strong performance not only on in-domain data, i.e., Wikipedia, but also on out-of-domain newspaper articles, and further confirm that hard negative mining substantially improves disambiguation among confusable candidates. Although the dataset focuses on Japanese, the construction method is language-agnostic and can be extended to other languages with sufficient Wikipedia and Wikidata coverage.
FiNERVINER: Fine-grained Named Entity Recognition for Vulnerable Languages of India’s North Eastern Region
Prachuryya Kaushik | Ashish Anand
Prachuryya Kaushik | Ashish Anand
Named entity recognition (NER), particularly fine-grained NER (FgNER), extracts domain-specific entity information for Natural Language Processing (NLP) applications such as knowledge base construction and relation extraction. While manual annotation for creating relevant data is expensive, distant supervision often produces noisy data. Moreover, resources for coarse-grained and fine-grained NER in Indian languages, particularly in the vulnerable languages of India’s North Eastern Region, remain scarce. This work aims at creating such a resource for three vulnerable languages: <i>Bodo/Boro (brx)</i>, <i>Manipuri/Meitei (mni)</i>, and <i>Mizo/Lushai (lus)</i>, which are regarded as official languages in three Indian states and spoken by more than six million people across five countries in South and Southeast Asia. We use annotations projection from high-resource FgNER datasets using source-to-target parallel corpora and a projection tool built on a multilingual encoder. The dataset comprises over 198k sentences, 282k entities, and 2.8M tokens in each low-resource language. Our thorough analyses validate the dataset’s high quality. We further explore zero-shot and cross-lingual settings, examining the impact of script similarity and multilingualism in cross-lingual FgNER performance. The dataset, expert detector models, the agentic tool, and the interactive web application are available as open-source resources at: https://hf.co/collections/prachuryyaIITG/finerviner.
APTFiNER: Annotation Preserving Translation for Fine-grained Named Entity Recognition
Prachuryya Kaushik | Adittya Gupta | Ajanta Maurya | Gautam Sharma | V. V. Saradhi | Ashish Anand
Prachuryya Kaushik | Adittya Gupta | Ajanta Maurya | Gautam Sharma | V. V. Saradhi | Ashish Anand
RelEx-PT: A Portuguese Sentence-Level Relation Extraction Dataset
Tomás Pinto | Catarina Silva | Hugo Goncalo Oliveira
Tomás Pinto | Catarina Silva | Hugo Goncalo Oliveira
We introduce RelEx-PT, a new sentence-level Relation Extraction dataset for Portuguese. Addressing the scarcity of high-quality, controlled resources for the language, RelEx-PT provides a balanced benchmark comprising 18 Wikidata-derived relation types across diverse domains. The dataset is built through a distant supervision pipeline that links Wikidata triples with Portuguese Wikipedia sentences and enhanced by a Natural Language Inference (NLI)-based filtering process, combining scalability with quality assurance. Additionally, we conduct baseline experiments to evaluate the dataset’s applicability across diverse extraction settings, including Relation Classification (RC), Relation Triple Extraction, and Open Information Extraction. These experiments leverage both prompting and fine-tuning strategies using Large Language Models. The results show that RelEx-PT effectively supports a range of extraction paradigms, yielding high performance in RC and competitive results in structured triple generation, while also highlighting key challenges in open-ended extraction.
Benchmarking Portuguese Open Information Extraction
Gabriel Silva | Mário Rodrigues | António Teixeira | Marlene Amorim
Gabriel Silva | Mário Rodrigues | António Teixeira | Marlene Amorim
Open Information Extraction (OIE) has seen significant advancements for English, but progress in Portuguese has been hindered by a lack of resources such as Datasets and standardized evaluation benchmarks. This work addresses this critical gap by establishing the a systematic and reproducible benchmark for Portuguese OIE systems. We conduct a comprehensive evaluation of eight systems, spanning a decade of research and encompassing both rule-based and neural architectures. The performance of these systems is measured against three distinct Portuguese corpora (WIKI200, CETEN200, and Gamalho) using the established CaRB methodology. Our results reveal that no single system excels across all three datasets. Rule-based models perform strongly on general text (WIKI200, CETEN200) but falter on specialized corpora (Gamalho), while neural systems demonstrate more consistent but not superior performance. With overall F1 scores averaging around 40%, our findings confirm that Portuguese OIE remains a largely unsolved task. This benchmark provides a baseline for future research and highlights the need for a high-quality, manually annotated gold-standard dataset to drive meaningful progress in the field. The evaluation benchmark/framework is made publicly available at https://github.com/gabrielrsilva11/PT-OIE-Benchmark.
A Scalable Pipeline for Novelty Detection in Skill Extraction Using Large Language Models
Gian Seifert | Simon Clematide
Gian Seifert | Simon Clematide
The rapid evolution of the labor market requires skill ontologies to be continuously updated, but manually identifying emerging skills in job advertisements is highly labor-intensive. This paper presents a scalable, multi-stage pipeline for automated novelty detection in skill extraction. The system combines Large Language Models (LLMs) for candidate generation, a re-matching and threshold-based filtering module (“Turbo”), that compares candidates against the existing ontology, and a two-step aggregation process that merges string-based and embedding-based clustering. Experiments on Swiss job advertisement datasets using GPT-4o, Gemini-2.0-flash, and DeepSeek-V3 show that the pipeline effectively reduces noise and manual curation effort: Turbo filtering lowered false positives by 82%, and aggregation reduced the number of items requiring review by 97%. Among the tested models, Gemini-2.0-flash achieved the highest precision, reaching a novelty detection ratio of up to 73% in the qualitative evaluation. These findings demonstrate the pipeline’s potential as an efficient tool for maintaining dynamic skill ontologies.
Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
Alistair Plum | Laura Maria Bernardy | Tharindu Ranasinghe
Alistair Plum | Laura Maria Bernardy | Tharindu Ranasinghe
We present judgeWEL, a dataset for named entity recognition (NER) in Luxembourgish, automatically labelled and subsequently verified using large language models (LLM) in a novel pipeline. Building datasets for under-represented languages remains one of the major bottlenecks in natural language processing, where the scarcity of resources and linguistic particularities make large-scale annotation costly and potentially inconsistent. To address these challenges, we propose and evaluate a novel approach that leverages Wikipedia and Wikidata as structured sources of weak supervision. By exploiting internal links within Wikipedia articles, we infer entity types based on their corresponding Wikidata entries, thereby generating initial annotations with minimal human intervention. Because such links are not uniformly reliable, we mitigate noise by employing and comparing several LLMs to identify and retain only high-quality labelled sentences. The resulting corpus is approximately five times larger than the currently available Luxembourgish NER dataset and offers broader and more balanced coverage across entity categories, providing a substantial new resource for multilingual and low-resource NER research.
From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence
Premtim Sahitaj | Jawan Kolanowski | Ariana Sahitaj | Veronika Solopova | Max Upravitelev | Daniel Röder | Iffat Maab | Junichi Yamagishi | Sebastian Möller | Vera Schmitt
Premtim Sahitaj | Jawan Kolanowski | Ariana Sahitaj | Veronika Solopova | Max Upravitelev | Daniel Röder | Iffat Maab | Junichi Yamagishi | Sebastian Möller | Vera Schmitt
Fact-checking articles encode rich supporting evidence and reasoning, yet this evidence remains largely inaccessible to automated verification systems due to unstructured presentation. We introduce PrimeFacts, a methodology and resource for extracting fine-grained evidence from full fact-checking articles. We compile 13,106 PolitiFact articles with claims, verdicts, and all referenced sources, and we identify 49,718 in-article hyperlinks as natural anchors to pinpoint key evidence. Our framework leverages large language models (LLMs) to rewrite these anchor sentences into stand-alone, context-independent premises and investigates the extraction of additional implicit evidence. In evaluations on cross-article evidence retrieval and claim verification, the extracted premises substantially improve performance. Decontextualized evidence yields higher retrievability, achieving up to a 30% relative gain in Mean Reciprocal Rank over verbatim sentences, and using the evidence for verdict prediction raises Macro-F1 by 10-20 points over the baseline. These gains are consistent across different verdict granularities (2-class vs. 5-class) and model architectures. A qualitative analysis indicates that the decontextualized premises remain faithful to the original sources. Our work highlights the promise of reusing fact-checkers’ evidence for automation and provides a large-scale resource of structured evidence from real-world fact-checks.
EpiGator: An Event-based Surveillance System for Infectious Disease Outbreaks
Yiheng Wu | Jue Hou | Trangcasanchai Sathianpong | Lidia Pivovarova | Roman Yangarber
Yiheng Wu | Jue Hou | Trangcasanchai Sathianpong | Lidia Pivovarova | Roman Yangarber
We present EpiGator, a novel event-based system for global surveillance of outbreaks of infectious epidemics that automatically processes streams of news articles and generates reports about the outbreaks, which is crucial for medical authorities. The goal of our work is to combine our experience in outbreak surveillance with state-of-the-art large language models (LLM), which allows us to reduce the overall cost of system development and maintenance. The EpiGator pipeline combines keyword filtering, relevance classification, event-based clustering, and multi-document summarization. A key novelty lies in using a fine-tuned LLM to identify articles relevant to ongoing outbreaks, followed by a zero-shot information extraction pipeline that normalizes the event features and clusters the related articles. For each cluster, we generate an outbreak summary using instruction-tuned LLMs. We evaluate EpiGator output against disease outbreak reports written by medical specialists.
Relation Extraction across Entire Books to Reconstruct Community Networks: The AffilKG Datasets
Erica Cai | Sean Mcquade | Kevin Young | Brendan O’Connor
Erica Cai | Sean Mcquade | Kevin Young | Brendan O’Connor
When knowledge graphs (KGs) are automatically extracted from text, are they accurate enough for downstream analysis? Unfortunately, current annotated datasets cannot be used to evaluate this question, since the knowledge graphs they correspond to, constructed by mapping entities in the text to nodes and relations to edges, are typically highly disconnected, too small, or overly complex. To address this gap, we introduce AffilKG, which is a collection of six datasets that are the first to pair complete book scans with large, labeled knowledge graphs. Each dataset features affiliation graphs, which are simple KGs that capture Member relationships between Person and Organization entities—useful in studies of migration, community interactions, and other social phenomena. In addition, three datasets include expanded KGs with a wider variety of relation types. Our preliminary experiments demonstrate significant variability in model performance across datasets, underscoring AffilKG’s ability to enable two critical advances: (1) benchmarking how extraction errors propagate to graph-level analyses (e.g., community structure), and (2) validating KG extraction methods for real-world social science research.
Vrittanta-EN: A Benchmark Dataset for Event Trigger Detection and Classification Advancing Event Understanding in English Narrative Discourse
Chaitanya Kirti | Ashish Anand | Prithwijit Guha
Chaitanya Kirti | Ashish Anand | Prithwijit Guha
Event trigger detection and classification involve identifying meaningful occurrences and categorizing them into predefined event types within narrative text. Despite extensive research on English event extraction in factual domains like news and biomedical text, narrative prose, such as short stories, has received comparatively little attention. To bridge this gap, Vrittanta-EN introduces a manually annotated English corpus comprising 11,272 event instances extracted from diverse short stories. The dataset captures a wide range of communicative, cognitive, and physical actions typical of narrative discourse. A comprehensive evaluation is conducted across a wide range of models, including classical machine learning baselines (SVM, Naive Bayes), neural sequential models (LSTM, BiLSTM, BiLSTM-CRF), encoder-only transformers (BERT, RoBERTa, ALBERT, DistilBERT, DeBERTa, ELECTRA), and encoder-decoder models (T5, BART), along with large language models (GPT-4.1, DeepSeek-V3.2-Exp, Claude Sonnet 4) under both zero-shot and five-shot settings. Experimental results show that ELECTRA achieved the highest overall performance for event trigger detection with an F1-score of 90.61%, while RoBERTa demonstrated superior performance for event classification with a macro F1 of 74.71%. These findings highlight the robustness of contextual transformer-based architectures for modeling narrative event structures in English short stories. The dataset, code, and annotation guidelines will be publicly released upon paper acceptance.
MUC-4 Revisited: Document-level Event Analysis beyond Span-based Arguments
Helene Bøsei Olsen | Erik Velldal | Lilja Øvrelid
Helene Bøsei Olsen | Erik Velldal | Lilja Øvrelid
Automatically predicting structured representations of events has long been a central goal in information extraction, yet most contemporary work remains limited to identifying contiguous text spans as event arguments. This span-centric formulation fails to capture higher-level aspects of real-world events, such as actor identities, temporal scope, and aggregated outcomes, that many event-centred applications depend on. While commonly treated as a standard extractive benchmark, MUC-4 originally combined span-based arguments with normalised, inferred, and categorical fields, reflecting a richer, application-driven design. In this paper, we revisit MUC-4 in its full original formulation, casting it as an abstractive event analysis task that connect traditional event extraction goals with modern generative and document-level paradigms. We provide the first systematic evaluation of fine-tuned generative models in this extended formulation on MUC-4, examining how post-training stages and model size affect performance across both span-based and higher-level, semantically grounded event information. An extensive error analysis highlights practical challenges and directions for future work.
Historical Medical Knowledge Graphs and Ontologies from the Medical History of British India Corpus (1850-1950)
Mehrdad Almasi | Tugce Karatas
Mehrdad Almasi | Tugce Karatas
This research presents a reproducible framework for constructing biomedical knowledge graphs and ontologies from digitized historical archives. Focusing on the Medical History of British India corpus (468 reports; ∼22.5M words; 1850–1950), our pipeline combines BioBERT-based entity recognition, LLM-guided relation extraction with LLM-based filtering, and clustering-based ontology induction. Reliability is strengthened through canonicalization, schema mapping to standardized biomedical relation types, and multi-metric edge scoring with temporal decay; a manual evaluation of 500 validated triples yields 0.892 precision. The resulting resources comprise 282,882 extracted relations, consolidated into 22,360 unique surface forms and organized into 71 thematic clusters. Frequent categories include After Treatment (∼1,242 mentions), Date of Inoculation (∼540), and diverse causal relations, while the induced ontology highlights six epidemic diseases: plague, cholera, malaria, kala azar, leprosy, and smallpox together with their characteristic interventions (e.g., quinine therapy, vaccination campaigns, hospital disinfection). Temporal analyses capture historically plausible trajectories: plague interventions peaking in the 1890s, cholera’s long-run decline, and tuberculosis departments rising after 1910. All code, relation inventories, ontologies, and visualizations are released in a GitHub Repository, enabling reproducibility and supporting research in historical NLP, biomedical informatics, and digital humanities.
Graph-TempCZ: A Graph Representation of Software Mentions for Predicting Software Usage in Scientific Publications
Congfeng Cao | Pengyu Zhang | Jelke Bloem
Congfeng Cao | Pengyu Zhang | Jelke Bloem
Predicting how software is used, shared, and evolves across publications is essential to studying scientific progress. Existing methods for representing software usage in publications rely mainly on tabular or textual formats, which limit their structural expressiveness and consequently their ability to predict software usage. We address these gaps by representing software mentions and citations as a graph and formulating software usage prediction as a link prediction task. To support this study, we construct the first large-scale graph dataset of publication and software mentions, Graph-TempCZ, covering 1959-2022 with over six million mention relationships. Experiments using both traditional machine learning and Graph Neural Network (GNN) show that graph-based models substantially outperform feature-based baselines, achieving a 5.98% improvement in test accuracy. Temporal experiments further reveal that models trained on one year generalize effectively to nearby years but show gradual performance decay as the temporal gap increases. This work provides the first comprehensive foundation for analyzing software usage through a temporal graph representation.
Automatic Suggestions Help Extending Eventive Ontology: A Case Study on SynSemClass
Jana Straková | Eva Fučíková | Zdenka Uresova | Jan Hajič
Jana Straková | Eva Fučíková | Zdenka Uresova | Jan Hajič
Despite substantial recent progress in many areas of NLP, semantic tasks remain particularly challenging. One such task is the creation (extension, or annotation) of semantic ontologies. In this work, we present a case study on the eventive SynSemClass ontology, focusing on the challenges of semantic annotation – that is extending the ontology with new lexical units and/or new concepts – both with and without automatic support. We consider two strategies for generating annotation suggestions: (i) a knowledge-driven approach based on a small, carefully curated corpus of verbal valency frames, and (ii) a corpus-driven approach using lemma-based suggestions from a large raw text collection, disregarding semantic homonymy. Our findings show that ontology annotation is inherently difficult, and that automatic annotations statistically significantly reduce this difficulty both in terms of inter-annotator agreement and when compared with gold expert annotations. We discuss the implications for semantic resource creation and extension, as well as the limits of automation in ontology annotation.
JPPB: Automatic Construction of a Soft-Labeled Japanese Patient Phrase Bank for Symptom Normalization
Tomohiro Nishiyama | Mana Kuramoto | Shoko Wakamiya | Eiji ARAMAKI
Tomohiro Nishiyama | Mana Kuramoto | Shoko Wakamiya | Eiji ARAMAKI
Patient-generated symptom expressions are linguistically diverse, often deviating from standardized medical terminology. This paper introduces the Japanese Patient Phrase Bank (JPPB), the first automatically constructed phrase-level normalization resource for Japanese patient language. JPPB introduces an embedding-based soft labeling framework that transforms traditional one-to-one dictionary mappings into graded and ambiguity-aware associations. This framework represents a shift from word-level to phrase-level normalization in Japanese. The resource covers 7,035 phrase–term pairs across 412 symptoms. Evaluation on the KEEPHA and MedNLP-SC datasets shows that soft labels consistently improve Top-1 accuracy and better approximate gold label distributions compared with hard labels. While LLM-based normalization achieved the highest scores, JPPB provides a lightweight and transparent alternative suitable for local deployment. This work demonstrates that large-scale, automatically generated phrase banks can achieve competitive performance relative to manually curated resources and serve as practical, scalable resources for medical natural language processing in Japanese.
How I Met Your Snowclone: Unsupervised Discovery of Snowclone Patterns in Large Datasets
Julien Bezançon | Gaël Lejeune | Marceau Hernandez
Julien Bezançon | Gaël Lejeune | Marceau Hernandez
Snowclones are a type of Multiword Expression (MWE) pattern that includes open slots, i.e. positions that can be filled with various words. For example, in the phrase “May the X be with you,” the slot X can be replaced with virtually any noun. A key feature of snowclones is that the original MWE remains recognizable, carrying its meaning into the new form. However, previous work has not shown whether such substitutions are limited to fixed positions. In practice, variations such as “May the force bee with you” are also possible. In this paper, we propose to use Locality Sensitive Hashing (LSH) to automatically extract snowclone patterns from the non-commercial IMDb dataset. This process results in the creation of the FROST lexicon, comprising 29,011 pattern candidates and 991,626 snowclone candidates distributed in 29 languages. We then annotate 1,500 discovered patterns and 1,000 snowclones from the FROST lexicon to assess its quality. Our findings suggest that (i) most substitutions in snowclones occur at consistent positions and (ii) snowclones can be reliably discovered at scale using LSH and similarity-based metrics. This work provides the first large-scale lexicon of snowclone-based MWEs and a method that can support future research on MWEs and snowclones discovery.
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
Shusaku Egami | Aoi Ohta | Tomoki Tsujimura | Masaki Asada | Tatsuya Ishigaki | Ken Fukuda | Masahiro Hamasaki | Hiroya Takamura
Shusaku Egami | Aoi Ohta | Tomoki Tsujimura | Masaki Asada | Tatsuya Ishigaki | Ken Fukuda | Masahiro Hamasaki | Hiroya Takamura
Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured knowledge. Integrating these two in a complementary manner enables the development of reliable and verifiable AI systems. In particular, knowledge graph question answering (KGQA) has attracted attention as a means to reduce LLM hallucinations and to leverage knowledge beyond the training data. However, existing KGQA benchmark datasets are biased toward encyclopedic knowledge, limited to a single modality, and lack fine-grained spatiotemporal data, which limits their applicability to real-world scenarios targeted by Embodied AI. We introduce HOME-KGQA, a novel KGQA benchmark dataset built on a multimodal KG of daily household activities. HOME-KGQA consists of complex, multi-hop natural language questions paired with graph database query languages. Compared to existing benchmarks, it includes more challenging questions that involve multi-level spatiotemporal reasoning, multimodal grounding, and aggregate functions. Experimental results show that the LLM-based KGQA methods fail to achieve performance comparable to that on existing datasets when evaluated on HOME-KGQA. This highlights significant challenges that should be addressed for the real-world deployment of KGQA systems. Our dataset is available at https://github.com/aistairc/home-kgqa.
Extending the Semantic Layer of the CompL-it Italian Lexicon: Traits, Semantic Types, and Definitions
Emiliano Giovannetti | Andrea Bellandi | Simone Marchi | Mafalda Papini
Emiliano Giovannetti | Andrea Bellandi | Simone Marchi | Mafalda Papini
The growing impact of Large Language Models has highlighted the need for explicit, interpretable linguistic knowledge. Lexical resources respond to this need by offering structured representations that complement and constrain the implicit semantics of neural models. This paper presents an extension of CompL-it, currently the most comprehensive open computational lexicon of Italian. Building on the semantic layer inherited from LexicO—itself derived from the PAROLE-SIMPLE-CLIPS resource—the work enriches CompL-it with semantic traits and references to semantic types. Moreover, an experiment was conducted to generate missing definitions through an automatic process supported by LLMs. The resulting resource thus combines human-curated and machine-extended knowledge, ensuring both linguistic precision and scalability. This enriched semantic layer enhances CompL-it’s interoperability within the Linguistic Linked Data framework and strengthens its usability for NLP tasks such as word sense disambiguation, semantic role labelling, and knowledge grounding.
Integrating Knowledge Graph with Large Language Models for Multi-hop Question Generation
Yllias Chali | Al Hasib Mahamud
Yllias Chali | Al Hasib Mahamud
Question generation (QG) is a fundamental task in natural language processing that involves generating fluent and grammatically correct questions from a given input context, optionally conditioned on an answer. Multi-hop question generation (MHQG), a more complex variant, requires reasoning over multiple pieces of information across diverse contexts to formulate coherent questions. In this work, we propose Knowledge Graph for Question Generation (KG4QG), a novel framework that integrates knowledge graphs with large language models to address the challenges of MHQG. Our approach constructs knowledge graphs from input contexts, encodes them using Graph Attention Networks (GAT), and leverages Sentence Transformers for contextual text embeddings. These enriched representations are then fed into large language models—specifically BART and T5—for multi-hop question generation. We evaluate KG4QG on the HotpotQA dataset, demonstrating that our method achieves superior performance compared to existing state-of-the-art approaches, highlighting the effectiveness of combining structured knowledge and pre-trained language models for complex question generation tasks.
LocalGovPL: A Corpus of Speaker-Attributed Polish Local Government Transcripts
Dariusz Czerski | Maciej Ogrodniczuk
Dariusz Czerski | Maciej Ogrodniczuk
We present LocalGovPL, a large-scale, speaker-annotated corpus of Polish local government meeting transcripts processed using an automatic two-stage LLM pipeline. The corpus consists of 31,900 sessions from 749 councils recorded between 2018–2025 (approximately 391M words). It is released in TEI P5 format with explicit links between utterances and registered participants. We collect transcripts from official local government portals using a dedicated crawler, normalize the text, and apply: (1) LLM-assisted extraction of person names and administrative roles; and (2) attribution of utterances to identified speakers using discourse cues. To evaluate attribution quality, we manually annotate 30 sessions and evaluate five LLM configurations using three evaluation protocols with speaker-aware word error rate (sWER). The strongest system, Gemini-2.5-pro, achieves 3.9% sWER for abstract speaker identification, 4.6% for known participants, and 5.9% for end-to-end processing with relaxed name matching. LocalGovPL enables large-scale analysis of local deliberative discourse and supports research on dialogue modeling, summarization, and political text analysis.
Amharic DBpedia Chapter: A Knowledge Graph for a Low-Resource Language
HIzkiel Mitiku Alemayehu | Tilahun Abedissa Taffa | Meti Adane Bayissa | Andargachew Asfaw Zewge | Hamada Zahera | Ricardo Usbeck | Axel-Cyrille Ngonga Ngomo
HIzkiel Mitiku Alemayehu | Tilahun Abedissa Taffa | Meti Adane Bayissa | Andargachew Asfaw Zewge | Hamada Zahera | Ricardo Usbeck | Axel-Cyrille Ngonga Ngomo
DBpedia is a community-driven project that extracts structured knowledge from Wikipedia via language-specific chapters. We present the first steps toward the Amharic DBpedia chapter by extending the DBpedia Extraction Framework (DEF) to support Amharic Wikipedia, including language-specific components such as Ethiopian date parsers, an Ethiopian–Gregorian calendar converter, an Arabic–Ge’ez number converter, and Amharic template mappings, together with automated extraction pipelines and the publication of the resulting knowledge graph through a live website, DBpedia Databus collection, and query endpoints. For mapping, we evaluate the zero-shot NLLB-200 translation model on Amharic infobox property names, achieving a BLEU score of 45.31. For ontology alignment, we link mapped properties to DBpedia ontology properties across 58 DBpedia classes and benchmark multilingual encoders with Amharic support, including Afro-XLM-R Base, XLM-R Base, and Amharic fine-tuned mBERT. The fine-tuned Afro-XLM-R model achieves 92.1% Top-10 accuracy and strong ranking performance, as measured by Mean Reciprocal Rank (MRR). We release all resources developed for the Amharic DBpedia chapter, including the Ethiopian date parser, Ethiopian–Gregorian calendar converter, Arabic–Geʽez numeral converter, Amharic template mappings, automated extraction workflows, and the resulting Amharic DBpedia knowledge graph with public access via the DBpedia Databus collection, Tentris query endpoint, and the live website at am.dbpedia.org.
The wordnet file specification empowers the creators of different wordnets by allowing them to encode the same information in multiple different ways. A drawback of this approach is that redundancy is introduced. As a consequence, different wordnets often contain conflicting records, creating issues when one attempts to conduct multilingual research using multiple wordnets simultaneously. To address this, we present the OMW Cygnet, an experimental reformulation of wordnet that is designed to eliminate conflicting records and improve modularity. We convert data in 47 languages from the Open Multilingual Wordnet into this format, and release a web browser which makes it easy to navigate multilingual wordnets.
Masrad: Arabic Terminology Management Corpora with Semi-Automatic Construction
Mahdi Nasser | Laura Sayah | Fadi Zaraket
Mahdi Nasser | Laura Sayah | Fadi Zaraket
This paper presents Masrad (i.e. glossary in Arabic), a terminology dataset for Arabic terminology management, and a method with supporting tools for its semi-automatic construction. The entries in Masrad are (f,a) pairs of foreign (non-Arabic) terms f, appearing in specialized, academic and field-specific books next to their Arabic a counterparts. Masrad-Ex systematically extracts these pairs as a first step to construct Masrad. Masrad helps improving term consistency in academic translations and specialized Arabic documents, and automating cross-lingual text processing. Masrad-Ex leverages translated terms organically occurring in Arabic books, and considers several candidate pairs for each term phrase. The candidate Arabic terms occur next to the foreign terms, and vary in length. Masrad-Ex computes lexicographic, phonetic, morphological, and semantic similarity metrics for each candidate pair, and uses heuristic, machine learning, and machine learning with post-processing approaches to decide on the best candidate. This paper presents Masrad after thorough expert review and makes it available to the interested research community. The best performing Masrad-Ex approach achieved 90.5% precision and 92.4% recall.
SentiMalti: A Maltese Sentiment Analysis Dataset and Models
Ian Caruana | Matthew Vella | Fabio Zammit | Kurt Micallef | Claudia Borg
Ian Caruana | Matthew Vella | Fabio Zammit | Kurt Micallef | Claudia Borg
We present SentiMalti, a new Maltese social media sentiment resource and accompanying baselines. We scrape user-generated content from YouTube, Reddit, and Facebook, then apply a Maltese-aware preprocessing pipeline (cleaning, personally identifiable information anonymisation, sentence splitting, and sentence-level language filtering) to retain Maltese sentences while tolerating realistic code-switching. The resulting crowdsourced dataset contains 2,327 sentences annotated for positive (39%), negative (31%), and neutral (30%) sentiment. We integrate prior Maltese datasets to create a combined benchmark of 3,772 instances. We evaluate fine-tuned encoder models (BERTu, Glot500) and few-shot prompting with instruction-tuned multilingual LLMs (Aya-101, Gemma 2 Instruct 9B). On the full test set, five-shot Aya-101 attains 68.65 macro-F1, closely followed by a fine-tuned BERTu at 68.36 macro-F1. Error analysis reveals complementary strengths: BERTu better separates polarised classes, while Aya-101 tends to over-predict the neutral class. We release the dataset splits, code, and a fine-tuned BERTu model to facilitate further work in Maltese NLP and sentiment analysis.
Multilingual Structured Sentiment Analysis for Environmental Sustainability
Muhammad Okky Ibrohim | Tommaso Caselli | Cristina Bosco | Valerio Basile
Muhammad Okky Ibrohim | Tommaso Caselli | Cristina Bosco | Valerio Basile
To effectively address global environmental challenges, we must have tools that allow us to carefully monitor how citizens, policy makers and other stakeholders debate sustainability. However, there are currently very few NLP resources and tools specialized for this topic. This paper presents EnviS, a multilingual corpus (Italian, English, and Indonesian) for investigating the debate on environmental sustainability in social media using Structured Sentiment Analysis. We introduce a framework for the automatic aggregation of span-level annotations that preserves the annotators’ perspective and avoids manual intervention by safeguarding the quality of the annotations. We performed a series of experiments with four open-source instruction-based Large Language Models in zero-shot and few-shot settings, where we have measures the impact of the order and number of shots. The results further confirm the ineffectiveness of LLMs in extracting fine-grained sentiment information, being outperformed by a supervised state-of-the-art neural method trained on very few data. This questions the suitability of LLMs for rich knowledge/information extraction tasks requiring manipulation of text spans. In particular, our error analysis indicates that LLMs mostly struggle in identifying the sentiment term or its associated polarity, failing to extract full sentiment triples.
LLM-as-an-Annotator: Training Lightweight Models with LLM-Annotated Examples for Aspect Sentiment Tuple Prediction
Nils Constantin Hellwig | Jakob Fehle | Udo Kruschwitz | Christian Wolff
Nils Constantin Hellwig | Jakob Fehle | Udo Kruschwitz | Christian Wolff
Training models for Aspect-Based Sentiment Analysis (ABSA) tasks requires manually annotated data, which is expensive and time-consuming to obtain. This paper introduces LA-ABSA, a novel approach that leverages Large Language Model (LLM)-generated annotations to fine-tune lightweight models for complex ABSA tasks. We evaluate our approach on five datasets for Target Aspect Sentiment Detection (TASD) and Aspect Sentiment Quad Prediction (ASQP). Our approach outperformed previously reported augmentation strategies and achieved competitive performance with LLM-prompting in low-resource scenarios, while providing substantial energy efficiency benefits. For example, using 50 annotated examples for in-context learning (ICL) to guide the annotation of unlabeled data, LA-ABSA achieved an F1 score of 49.85 for ASQP on the SemEval Rest16 dataset, closely matching the performance of ICL prompting with Gemma-3-27B (51.10), while requiring significantly lower computational resources.
Extending Czech Aspect-Based Sentiment Analysis with Opinion Terms: Dataset and LLM Benchmarks
Jakub Šmíd | Pavel Priban | Pavel Kral
Jakub Šmíd | Pavel Priban | Pavel Kral
This paper introduces a novel Czech dataset in the restaurant domain for aspect-based sentiment analysis (ABSA), enriched with annotations of opinion terms. The dataset supports three distinct ABSA tasks involving opinion terms, accommodating varying levels of complexity. Leveraging this dataset, we conduct extensive experiments using modern Transformer-based models, including large language models (LLMs), in monolingual, cross-lingual, and multilingual settings. To address cross-lingual challenges, we propose a translation and label alignment methodology leveraging LLMs, which yields consistent improvements. Our results highlight the strengths and limitations of state-of-the-art models, especially when handling the linguistic intricacies of low-resource languages like Czech. A detailed error analysis reveals key challenges, including the detection of subtle opinion terms and nuanced sentiment expressions. The dataset establishes a new benchmark for Czech ABSA, and our proposed translation–alignment approach offers a scalable solution for adapting ABSA resources to other low-resource languages.
AnnoABSA: A Web-Based Annotation Tool for Aspect-Based Sentiment Analysis with Retrieval-Augmented Suggestions
Nils Constantin Hellwig | Jakob Fehle | Udo Kruschwitz | Christian Wolff
Nils Constantin Hellwig | Jakob Fehle | Udo Kruschwitz | Christian Wolff
We introduce AnnoABSA, the first web-based annotation tool to support the full spectrum of Aspect-Based Sentiment Analysis (ABSA) tasks. The tool is highly customizable, enabling flexible configuration of sentiment elements and task-specific requirements. Alongside manual annotation, AnnoABSA provides optional Large Language Model (LLM)-based retrieval-augmented generation (RAG) suggestions that offer context-aware assistance in a human-in-the-loop approach, keeping the human annotator in control. To improve prediction quality over time, the system retrieves the ten most similar examples that are already annotated and adds them as few-shot examples in the prompt, ensuring that suggestions become increasingly accurate as the annotation process progresses. Released as open-source software under the MIT License, AnnoABSA is freely accessible and easily extendable for research and practical applications.
Zero-Shot to Full-Resource: Cross-lingual Transfer Strategies for Aspect-Based Sentiment Analysis
Jakob Fehle | Nils Constantin Hellwig | Udo Kruschwitz | Christian Wolff
Jakob Fehle | Nils Constantin Hellwig | Udo Kruschwitz | Christian Wolff
Aspect-based Sentiment Analysis (ABSA) extracts fine-grained opinions toward specific aspects within text but remains largely English-focused despite major advances in transformer-based and instruction-tuned models. This work presents a multilingual evaluation of state-of-the-art ABSA approaches across seven languages and four subtasks (ACD, ACSA, TASD, ASQP). We systematically compare different transformer architectures under zero-resource, data-only, and full-resource settings, using cross-lingual transfer, code-switching and machine translation. Fine-tuned Large Language Models (LLMs) achieve the highest overall scores, particularly in complex generative tasks, while few-shot counterparts approach this performance in simpler setups, where smaller encoder models also remain competitive. Cross-lingual training on multiple non-target languages yields the strongest transfer for fine-tuned LLMs, while smaller encoder or seq-to-seq models benefit most from code-switching, highlighting architecture-specific strategies for multilingual ABSA. We further contribute two new German datasets, an adapted GERestaurant and the first German ASQP dataset (GERest), to encourage multilingual ABSA research beyond English.
LoveHate: Stance Detection and Generation for Multiple Topics in User-generated Comments in Russian and English
Natalia Evgrafova | Veronique Hoste | Els Lefever
Natalia Evgrafova | Veronique Hoste | Els Lefever
This paper introduces LoveHate, a new multi-topic corpus of user-generated arguments in Russian, collected from the historical data of the debate platform lovehate.ru. The dataset contains nearly 19,000 posts spanning 16 socially and politically relevant topics, each mapped to binary pro and con stances. We test multiple approaches to stance detection and stance generation across Russian and English data, including translated variants, using both classifier-based (Roberta, RuRoberta) and instruction-tuned generative (Llama, Qwen) models. Results demonstrate that language-specific pretraining yields the strongest performance for stance classification (F1 = 0.892 with RuRoberta), while multilingual generative models – when fine-tuned on sufficient data – can effectively generate stance in Russian without explicit Russian pretraining. Cross-domain experiments show that English datasets generalize better across corpora, whereas Russian data capture language- and culture-specific argumentation but are less effective for generalizable models. Generating topics remains a more challenging task for both Russian and English data. The dataset and accompanying results contribute to multilingual stance research and provide a valuable new resource for argument mining in Russian.
From Trial by Fire to Sleep like a Baby: A Lexicon of Anxiety Associations for 20K English Multi-Word Expressions
Saif M. Mohammad
Saif M. Mohammad
Anxiety is the unease about a possible future negative outcome. In recent years, there has been growing interest in understanding how anxiety relates to our health, well-being, body, mind, and behaviour. This includes work on lexical resources for word–anxiety association. However, there is very little anxiety-related work on larger units of text such as multiword expressions (MWE). Here, we introduce the first large-scale lexicon capturing descriptive norms of anxiety associations for more than 20k English MWEs. We show that the anxiety associations are highly reliable. We use the lexicon to study prevalence of different types of anxiety- and calmness-associated MWEs; and how that varies across two-, three-, and four-word sequences. We also study the extent to which the anxiety association of MWEs is compositional (due to its constituent words). The lexicon enables a wide variety of anxiety-related research in psychology, NLP, public health, and social sciences. The lexicon is freely available: https://saifmohammad.com/worrylex.html
Entity-Level Sentiment Analysis with Sentence Relevance Detection
Egil Rønningstad | Roman Klinger | Lilja Øvrelid | Erik Velldal
Egil Rønningstad | Roman Klinger | Lilja Øvrelid | Erik Velldal
The task of entity-level sentiment analysis (Elsa) is to extract sentiment scores for a given entity (such as person names or organization names) from a text. Elsa is a challenging task and involves processing of longer documents, where several entities may be mentioned with varying importance for the final score aggregation. Fine-tuning encoder-based Transformers (such as BERT) constitutes the state of the art for sentiment predictions, however, these models are still limited by their restricted input lengths. Decoder-only models so far still underperform on the task. We approach the context limitation by learning to extract segments that are relevant for the sentiment prediction for a given entity, without preprocessing by chunking and aggregation. For decoder models, we explore fine-tuning these through supervised fine-tuning and pairwise comparison, a method borrowed from reward modeling for preference optimization. Both methods perform well and set a new standard for the Elsa task. We further show that pairwise classification is faster, simpler, and shows less variance than the more common direct supervision for this task.
Enhancing Multi-Label Emotion Analysis and Corresponding Intensities for Ethiopian Languages
Tadesse Destaw Belay | Dawit Ketema Gete | Abinew Ali Ayele | Olga Kolesnikova | Iqra Ameer | Grigori Sidorov | Seid Muhie Yimam
Tadesse Destaw Belay | Dawit Ketema Gete | Abinew Ali Ayele | Olga Kolesnikova | Iqra Ameer | Grigori Sidorov | Seid Muhie Yimam
Developing and integrating emotion-understanding models are essential for a wide range of human- computer interaction tasks, including customer feedback analysis, marketing research, and social media monitoring. Given that users often express multiple emotions simultaneously within a single instance, annotating emotion datasets in a multi-label format is critical for capturing this complexity. The EthioEmo dataset, a multilingual and multi-label emotion dataset for Ethiopian languages, lacks emotion intensity annotations, which are crucial for distinguishing varying degrees of emotion, as not all emotions are expressed with the same intensity. We extend the EthioEmo dataset to address this gap by adding emotion intensity annotations. Furthermore, we benchmark state-of-the-art encoder-only Pretrained Language Models (PLMs) and Large Language Models (LLMs) on this enriched dataset. Our results demonstrate that African-centric encoder-only models consistently outperform open-source LLMs, highlighting the importance of culturally and linguistically tailored small models in emotion understanding. Incorporating an emotion-intensity feature for multi-label emotion classification yields better performance. The data is available at https://huggingface.co/datasets/Tadesse/EthioEmo-intensities.
A Japanese Dataset for Aspect-based Sentiment Polarity Classification and Emotion Intensity Estimation
Kentaro Hanafusa | Kota Manabe | Yuki Maeda | Daisuke Maekawa | Tomoyuki Kajiwara | Hideaki Hayashi | Yuta Nakashima | Hajime Nagahara
Kentaro Hanafusa | Kota Manabe | Yuki Maeda | Daisuke Maekawa | Tomoyuki Kajiwara | Hideaki Hayashi | Yuta Nakashima | Hajime Nagahara
We manually construct and publicly release a Japanese dataset for Aspect-based Sentiment Analysis (ABSA), annotated with both sentiment polarity and the emotional intensities for Plutchik’s eight emotions. Existing datasets for Japanese ABSA only handle sentiment polarity classification. Therefore, we manually annotated Plutchik’s eight emotions with a four-point scale and sentiment polarity with a five-point scale to words in the Japanese sentiment analysis corpus WRIME. Analysis of this corpus revealed that word-level emotions more strongly reflect the reader’s objective impression than the writer’s subjective perspective. Furthermore, the results of evaluation experiments on word-level emotion estimation quantitatively demonstrated that while Large Language Models achieve high performance, they struggle with the estimation of the “trust” emotion. Additionally, we demonstrated that multi-task learning, utilizing both word and sentence levels, can improve performance on difficult-to-estimate subjective emotions.
Assessing the Persuasive Effect of AI-Generated Image Support of Arguments
Mackwyn Quadras | Manfred Stede | Henning Wachsmuth
Mackwyn Quadras | Manfred Stede | Henning Wachsmuth
Argumentation is, at its core, an inherently verbal activity. Yet, other modalities may support arguments, one of which are images. In the argument mining community, this combination has not received much attention yet. While a few previous works studied whether images can make argumentative texts more effective in persuading people, the images that were considered matched the texts loosely only, or they were heavily text-based themselves. In this paper, we take the step to study to what extent the persuasive effect of textual arguments can be supported by images specifically created for this purpose. For a consistent experiment design, we combine NLP with image generation to synthesize both arguments and images with generative AI, for five controversial topics and for two rhetorical strategies. In two consecutive user studies, we first determine the best-matching image for each argument and then compare the perceived effect of bare textual arguments to those that are supported by an image. Our results suggest that the images may increase the persuasive effect of argumentative texts, but with variance across topics.
CIARAM: Class Imbalance Aware Generative Framework for Relational Argument Mining
Nilmadhab Das | Sayan Pal | V. V. Saradhi | Ashish Anand
Nilmadhab Das | Sayan Pal | V. V. Saradhi | Ashish Anand
Relational Argument Mining (RAM) is a key task of computational argumentation, which aims to classify the relationships such as Support or Attack between argument component (AC) pairs. Traditional approaches primarily rely on graph-based modelling with external knowledge sources, which are complex in nature. Also, these approaches struggle with RAM datasets when relation classes are imbalanced, as they are not designed for class-imbalanced scenarios. In this work, we propose CIARAM framework to reformulate RAM as a text-to-text generation problem to generate relational labels in a flattened text format. To address the class imbalance, we employ a data augmentation strategy using a decoder-only Large Language Model (LLM) to balance the underrepresented relation classes. Across five standard RAM benchmarks, CIARAM produces strong results, specifically with the billion-parameter model, with a substantial gain in performance compared to the latest baseline, demonstrating the strong potential of our approach.
Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
Muhammed Yahia Gaffar Saeed | Muhammad Abdul-Mageed | Shady Shehata
Muhammed Yahia Gaffar Saeed | Muhammad Abdul-Mageed | Shady Shehata
Large language models (LLMs) are widely deployed for open-ended communication, yet most bias evaluations still rely on English, classification-style tasks. We introduce , a new multilingual, debate-style benchmark designed to reveal how narrative bias appears in realistic generative settings. Our dataset includes 8,400 structured debate prompts spanning four sensitive domains – Women’s Rights, Backwardness, Terrorism, and Religion – across seven languages ranging from high-resource (English, Chinese) to low-resource (Swahili, Nigerian Pidgin). Using four flagship models (GPT-4o, Claude 3.5 Haiku, DeepSeek-Chat, and LLaMA-3-70B), we generate over 100,000 debate responses and automatically classify which demographic groups are assigned stereotyped versus modern roles. Results show that all models reproduce entrenched stereotypes despite safety alignment: Arabs are overwhelmingly linked to Terrorism and Religion (≥89%), Africans to socioeconomic “backwardness” (up to 77%), and Western groups are consistently framed as modern or progressive. Biases grow sharply in lower-resource languages, revealing that alignment trained primarily in English does not generalize globally. Our findings highlight a persistent divide in multilingual fairness: current alignment methods reduce explicit toxicity but fail to prevent biased outputs in open-ended contexts. We release our benchmark and analysis framework to support the next generation of multilingual bias evaluation and safer, culturally inclusive model alignment
Prompt-Based Stance Control in German: An Evaluation of LLMs for Experimental Research on Attitude Change
Florian Omiecienski | Cornelia Sindermann | Agnieszka Falenska
Florian Omiecienski | Cornelia Sindermann | Agnieszka Falenska
How much can Large Language Models (LLMs) influence the attitudes and opinions of their users? Answering this question requires controlled pre/post-treatment experiments, where participants interact with LLMs that consistently adopt a predefined political stance. Such experiments, however, are only possible if LLMs can be reliably steered to hold these stances throughout the interactions. In this work, we evaluate whether state-of-the-art LLMs can be effectively stance-controlled in German, thereby enabling experiments on human–LLM interactions. First, using a corpus of realistic user prompts, we find that LLMs are predominantly neutral, making them infeasible for said experiments. We then show that a prompt-based stance control method can reliably guide models to argue for or against a particular topic. Finally, we analyze confounding factors like topic and stance of the initial user prompts. We find that control is easiest when the target stance aligns with topical priors of the model or a user’s prompt. Further, the models maintain a comparable style across target stances — a key prerequisite for pre/post-treatment experiments. Taken together, our results demonstrate that stance-controlled LLMs are feasible and practically useful for experiments on user attitude change.
CoSt-BR: A Language Resource for Conversational Stance Detection
Felipe Penhorate Carvalho da Fonseca | Ivandre Paraboni | Luciano Antônio Digiampietri
Felipe Penhorate Carvalho da Fonseca | Ivandre Paraboni | Luciano Antônio Digiampietri
Stance detection is the computational task of determining the attitude (e.g., for, against, neutral) expressed in text toward a specific target topic. In its more conventional form, the task focuses on isolated, context-free input utterances. Conversational stance detection, by contrast, analyzes messages embedded within dialogue threads, enabling the interpretation of responses in relation to preceding discourse, and takes into account a greater variety of stance relations (e.g., support, deny, query, comment, etc.). Despite growing research attention, however, conversational stance detection remains relatively under-resourced and largely limited to the English language. To address these gaps, this study introduces CoSt-BR, a new corpus for conversational stance detection composed of a large set of annotated Reddit discussions in Brazilian Portuguese. In addition, the paper also reports benchmark results obtained using various computational methods, including supervised and prompt-based strategies, applied to the corpus data, providing baseline references for future research in this area.
Less Is More? The Role of Demographic Author Information in Emotion Classification of Ambiguous Text
Sabine Weber | Lynn Greschner | Roman Klinger
Sabine Weber | Lynn Greschner | Roman Klinger
Emotion annotation in text is a challenging task that often yields low inter-annotator agreement. Missing context, differences in world knowledge and extra-linguistic factors such as the author’s identity influence how emotions are perceived. When the text does not provide sufficient information, details about the author may help resolve ambiguity. We test the hypothesis that providing annotators with demographic information reduces disagreement in emotion annotation. We compare one group of annotators who sees each text alongside demographic information about its author, with a group who sees only the text. We find in our study with 500 annotators and 250 texts that displaying demographic information about the author of the text does not improve agreement between annotators, nor does it improve agreement with the gold label. The only exception are cases where the emotion polarity (positive or negative) is unclear. We also find that annotators perform overall better at identifying the correct emotion label when it aligns with gender stereotypes. Zero-shot prompting experiments with large language models do resemble the human annotation experimental results. Our findings suggest that providing demographic information is not a straightforward remedy for ambiguity in emotion annotation and careful consideration is needed when incorporating such data.
Big Five Personality Prediction through Emotion-Conditioned Representations and Learnable Psycholinguistic Mapping
Lorenzo Zangari | Antonin Schnyder | Davide Picca
Lorenzo Zangari | Antonin Schnyder | Davide Picca
Personality traits influence human behavior and social interactions, making their accurate prediction essential across multiple domains. The Big Five Model, a widely recognized framework in psychological science for assessing personality traits, has become the foundation for different computational approaches to personality prediction. In recent years, a growing body of research has highlighted the dynamic interplay between emotions and personality, as individuals navigate diverse emotional experiences that evoke distinct responses and ultimately shape their behavioral patterns. In this work, we present a novel framework that systematically integrates affective information into Pre-trained Language Models for Big Five Personality trait prediction. Our framework leverages text-based embeddings, emotion-conditioned features, and learnable psycholinguistic information that bridges affective dimensions with personality traits. This design preserves established psycholinguistic knowledge while enabling adaptive refinement through data-driven learning. Our experiments showed that our framework outperformed sentence embedding-based methods and Large Language Models across various datasets from different domains, achieving an average F1-score improvement of at least 15% in out-of-domain scenarios.
SENSEI-ASG: A Challenging Dataset for Argument Summary Graph Parsing
Jonathan Clayton | Marco Damonte | Robert Gaizauskas
Jonathan Clayton | Marco Damonte | Robert Gaizauskas
We create, and make publicly available, a novel dataset for the task of Argument Summary Graph Parsing (ASGP), which we call SENSEI-ASG, based on annotating a subset of the SENSEI corpus. Given an argumentative dialogue, such as might be found in a social media exchange, ASGP is the task of creating an Argument Summary Graph, a data structure which consists of nodes containing summaries of arguments in a dialogue, and edges showing argumentative relations between them. We find that the only existing ASG dataset, Debatabase-ASG, is not representative of online debates in language use, length of the dialogues, or graph complexity. In contrast to Debatabase-ASG, which was created based on a curated debate collection, SENSEI-ASG contains examples of spontaneous debates arising in the comments sections of an online newspaper (namely, The Guardian). We achieve moderate inter-annotator agreement on the dataset, with a Cohen’s kappa of k=0.57, reflecting the inherent challenges in distinguishing argumentative from non-argumentative text. We propose baselines for the new dataset by fine-tuning Llama-3 for the ASGP task, using the two ASGP datasets and an additional out-of-domain argument mining dataset, the AAEC.
Categorical Emotions or Appraisals - Which Emotion Model Explains Argument Convincingness Better?
Lynn Greschner | Meike Bauer | Sabine Weber | Roman Klinger
Lynn Greschner | Meike Bauer | Sabine Weber | Roman Klinger
The convincingness of an argument does not only depend on its structure (logos), the person who makes the argument (ethos), but also on the emotion that it causes in the recipient (pathos). While the overall intensity and categorical values of emotions in arguments have received considerable attention in the research community, we argue that the emotion an argument evokes in a recipient is subjective. It depends on the recipient’s goals, standards, prior knowledge, and stance. Appraisal theories lend themselves as a link between the subjective cognitive assessment of events and emotions. They have been used in event-centric emotion analysis, but their suitability for assessing argument convincingness remains unexplored. In this paper, we evaluate whether appraisal theories are suitable for emotion analysis in arguments by considering subjective cognitive evaluations of the importance and impact of an argument on its receiver. Based on the annotations in the recently published ContArgA corpus, we perform zero-shot prompting experiments to evaluate the importance of gold-annotated and predicted emotions and appraisals for the assessment of the subjective convincingness labels. We find that, while categorical emotion information does improve convincingness prediction, the improvement is more pronounced with appraisals. This work presents the first systematic comparison between emotion models for convincingness prediction, demonstrating the advantage of appraisals, providing insights for theoretical and practical applications in computational argumentation.
Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale
Karl Gustav Gailit | Kadri Muischnek | Kairit Sirts
Karl Gustav Gailit | Kadri Muischnek | Kairit Sirts
This article presents the creation of an Estonian-language dataset for document-level subjectivity, analyzes the resulting annotations, and reports an initial experiment of automatic subjectivity analysis using a large language model (LLM). The dataset comprises of 1,000 documents—300 journalistic articles and 700 randomly selected web texts—each rated for subjectivity on a continuous scale from 0 (fully objective) to 100 (fully subjective) by four annotators. As the inter-annotator correlations were moderate, with some texts receiving scores at the opposite ends of the scale, a subset of texts with the most divergent scores was re-annotated, with the inter-annotator correlation improving. In addition to human annotations, the dataset includes scores generated by GPT-5 as an experiment on annotation automation. These scores were similar to human annotators, however several differences emerged, suggesting that while LLM based automatic subjectivity scoring is feasible, it is not an interchangeable alternative to human annotation, and its suitability depends on the intended application.
Mitigating Misinterpretation in Policy Documents through Automated Language Understanding
Momojit Biswas | Anka Chandrahas Tummepalli | Preethu Rose Anish
Momojit Biswas | Anka Chandrahas Tummepalli | Preethu Rose Anish
Policy documents often employ intricate and technical language, posing comprehension challenges for policyholders and increasing the risk of misinterpretation, financial losses, and legal disputes. To address these issues, we propose an automated framework leveraging Retrieval-Augmented Generation to identify and clarify potentially mis-interpretable paragraphs within policy documents. The framework consists of two key modules: the Annotation module and the Rectification module. The Annotation module employs both paragraph-level and document-level contextual reasoning to classify paragraphs into categories indicative of potential misinterpretation. The Rectification module resolves these ambiguities by generating targeted interpretation queries, retrieving relevant document-level context, and incorporating external knowledge sources. Applied to a corpus of 240 real-world policy documents, the Annotation module produced a benchmark dataset comprising 11,000 annotated paragraphs, enabling systematic evaluation of interpretability issues. We assessed the dataset’s quality through expert-driven manual reviews and large-scale automated evaluations using fine-tuned Pretrained Language Model. For the Rectification module, we evaluated five open-source Large Language Models: Mistral-2-7B, Mistral-3-7B, LLaMA-2-7B, LLaMA-3-8B, andSaul-7B. Among these, Mistral-2-7B achieved the highest human evaluation scores: 0.912 for Clarity, 0.914 for Fidelity, and 0.934 for Usefulness. This work demonstrates the practical feasibility of utilizing automated frameworks to enhance the clarity and comprehensibility of complex policy documents, thereby mitigating risks associated with misinterpretation and its adverse consequences.
Sovereign AI-based Public Services Are Viable and Affordable
António Branco | Luis M. S. Gomes | Rodrigo Santos | Eduardo Santos | João Ricardo Silva | Nuno Marques | Madalena Rodrigues
António Branco | Luis M. S. Gomes | Rodrigo Santos | Eduardo Santos | João Ricardo Silva | Nuno Marques | Madalena Rodrigues
The rapid expansion of AI-based remote services has intensified debates about the long-term implications of growing structural concentration in infrastructure and expertise. As AI capabilities become increasingly intertwined with geopolitical interests, the availability and reliability of foundational AI services can no longer be taken for granted. This issue is particularly pressing for AI-enabled public services for citizens, as governments and public agencies are progressively adopting 24/7 AI-driven support systems typically operated through commercial offerings from a small oligopoly of global technology providers. This paper challenges the prevailing assumption that general-purpose architectures, offered by these providers, are the optimal choice for all application contexts. Through practical experimentation, we demonstrate that viable and cost-effective alternatives exist—alternatives that align with principles of digital and cultural sovereignty. Our findings provide an empirical illustration that sovereign AI-based public services are both technically feasible and economically sustainable, capable of operating effectively on premises with modest computational and financial resources while maintaining cultural and digital autonomy. The technical insights and deployment lessons reported here are intended to inform the adoption of similar sovereign AI public services by national agencies and governments worldwide.
A Typology of Synthetic Datasets for Dialogue Processing in Clinical Contexts
Steven Bedrick | A. Seza Doğruöz | Sergiu Nisioi
Steven Bedrick | A. Seza Doğruöz | Sergiu Nisioi
Synthetic datasets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of clinical (healthcare) contexts, where there exist significant and long-standing challenges (e.g., privacy, anonymization, and data governance) which have led to the development of an increasing number of synthetic datasets. One increasingly important category of clinical dataset is that of clinical dialogues which are especially sensitive and difficult to collect. Therefore, they are commonly synthesized. While such synthetic datasets have been shown to be sufficient in some situations, little theory exists to inform how they may be best used and generalized to new applications. In this paper, we provide an overview of how synthetic datasets are created, evaluated and used for dialogue related tasks in the medical domain. Additionally, we propose a novel typology for use in classifying types and degrees of data synthesis, to facilitate comparison and evaluation.
Text+: A National Hub Including Legacy Language Data
Florian Barth | Christoph Draxler | Jennifer Ecker | Stefan Fischer | Philippe Genêt | Alina Hemmer | Timm Lehmberg | Thorsten Trippel | Andreas Witt | Arden Zimmermann | Claus Zinn
Florian Barth | Christoph Draxler | Jennifer Ecker | Stefan Fischer | Philippe Genêt | Alina Hemmer | Timm Lehmberg | Thorsten Trippel | Andreas Witt | Arden Zimmermann | Claus Zinn
Text+ is the German distributed research data infrastructure for literary studies, linguistics, and spoken and written language. Its resources consist of contemporary and historical literary and media texts, deeply annotated material, transcripts of spoken and sign language, and original recordings. Text+ provides access to its resources according to the FAIR guidelines: Findable due to standard-conformant metadata, Accessible with single sign-on authentication, Interoperable via open data formats, and Reproducible through web services and extensive documentation. The 30+ partners of Text+ are archives, libraries, universities, and other research institutions. The partners are autonomous, and they differ in the amount of data and processing capabilities they provide. In this paper, we describe the hub architecture of Text+, which gives users a central and FAIR point of access to research data that continues to be distributed across the Text+ partner institutions. The architecture serves as a blueprint to evolving research infrastructures that aim at maintaining (and empowering) their research data contributors.
Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech
Tanvi Dinkar | Aiqi Jiang | Simona Frenda | Poppy Gerrard-Abbott | Nancie A. Gunson | Gavin Abercrombie | Ioannis Konstas
Tanvi Dinkar | Aiqi Jiang | Simona Frenda | Poppy Gerrard-Abbott | Nancie A. Gunson | Gavin Abercrombie | Ioannis Konstas
Counterspeech, i.e. the practice of responding to online hate speech, has gained traction in NLP as a promising intervention. While early work emphasised collaboration with non-governmental organisation stakeholders, recent research trends have shifted toward automated pipelines that reuse a small set of legacy datasets, often without input from affected communities. This paper presents a systematic review of 74 NLP studies on counterspeech, analysing the extent to which stakeholder participation influences dataset creation, model development, and evaluation. To complement this analysis, we conducted a participatory case study that spanned close to two years with five NGOs specialising in online Gender-Based Violence (oGBV), identifying stakeholder-informed practices for counterspeech generation. Our findings reveal a growing disconnect between current NLP research and the needs of communities most impacted by toxic online content. We conclude with concrete recommendations for re-centring stakeholder expertise in counterspeech research.
Towards Complex Debate Understanding: Predicting Claim Impact Scores through the Modelling of Claim Interactions
Maxime Brouat | Mihai Surdeanu | Srdjan Vesic | Eduardo Blanco
Maxime Brouat | Mihai Surdeanu | Srdjan Vesic | Eduardo Blanco
Structured debates can be naturally modeled as argument graphs, with claims connected by support and attack relations, a representation formalised in Computational Argumentation Theory. In this paper, we propose a novel neural architecture that jointly models both the textual content of claims and their relational structure. Claims are encoded using contextualised embeddings and compressed through a feedforward compression layer. Then, a graph attention network explicitly captures attack/support interactions. Trained on real-world debates from the Kialo platform, our model predicts the distribution of user-assigned impact votes for each claim. It achieves a mean absolute error (MAE) of 0.068, significantly outperforming both text-only and structure-only baselines. Further experiments show strong out-of-domain generalisation across thematic clusters, as well as suggestive correlations between the model’s attention patterns and human voting behaviour. An analysis of linguistic and graph-based features suggests that the model relies on latent argumentative patterns as well as the text. Our findings also shed light on language differences between strong and weak claims, as determined by humans as well as by our best model.
Is There Anything More Deceptive than an Obvious Fact? Investigating Implicitness in User-Generated Argumentative Text
Ekaterina Sviridova | Elena Cabrio | Serena Villata
Ekaterina Sviridova | Elena Cabrio | Serena Villata
While various attempts towards unveiling implicitness in argumentation have been made, particularly towards improving automatic detection and reconstruction of implicit components and background knowledge, the task remains overly challenging. In this paper, we present, to the best of our knowledge, the first fine-grained typology of implicitness in argumentation, distinguishing among implicature, ambiguity, and presupposition. Applying this typology, we annotate 78 full-length discussions from the Change My View forum, building the largest publicly available dataset of real-world enthymemes with implicitness types labeled. For comparison, we additionally annotate 112 short argumentative texts from the Microtext corpus to examine how text length and complexity influence the automatic analysis of natural arguments. Leveraging these datasets, we establish strong baselines for two tasks: (i) enthymeme detection and (ii) fine-grained implicitness classification, with both encoder-only and large language models, highlighting the challenge of modeling implicit reasoning in long, unstructured discourse.
Best-Worst Scaling of Hype in Biomedical Research: Building an Intensity Lexicon of Promotional Adjectives
Neil Millar | Dipesh Satav | Bojan Batalo | Erica K. Shimomoto | Ryosuke L. Ohniwa
Neil Millar | Dipesh Satav | Bojan Batalo | Erica K. Shimomoto | Ryosuke L. Ohniwa
Promotional language, or “hype”, is increasingly common in biomedical research reporting. Adjectives such as groundbreaking, robust, and impactful can engage readers but also risk imposing value judgements and undermining objectivity. Detecting and assessing such language requires distinguishing degrees of promotional intensity (e.g., new < novel < groundbreaking < revolutionary), yet no such graded resource exists. We present an intensity-scaled lexicon of 303 promotional adjectives attested in biomedical writing across eight evaluative domains (e.g. IMPORTANCE, NOVELTY, RIGOUR). Ratings were obtained through Best–Worst Scaling (BWS) with human participants evaluating adjectives for promotional strength in the context of scientific research reporting. We refer to this as the Hyplex resource (Hype Lexicon). The ratings show high internal consistency (r = 0.87; 95% CI [0.85, 0.89]) and correlate most strongly with arousal and dominance in the NRC VAD Lexicon, suggesting that promotional intensity aligns more with reader activation and perceptions of assertiveness than simple positivity. We also release an online BWS platform integrated with the R package bwsTools to support intensity-scaling research in other domains.
Trust Me, I Can Convince You: The Contextualized Argument Appraisal Framework and the ContArgA Corpus
Lynn Greschner | Sabine Weber | Roman Klinger
Lynn Greschner | Sabine Weber | Roman Klinger
Emotions that somebody develops based on an argument do not only depend on the argument itself - they are also influenced by a subjective evaluation of the argument’s potential impact on the self. For instance, an argument to ban plastic bottles might cause fear of losing a job for a bottle industry worker, which lowers the convincingness – presumably independent of its content. While binary emotionality of arguments has been studied, such cognitive appraisal models have only been proposed in other subtasks of emotion analysis, but not in the context of arguments and their convincingness. To fill this research gap, we propose the Contextualized Argument Appraisal Framework to model the interplay between the sender, receiver, and argument. We adapt established appraisal models from psychology to argument mining, including argument pleasantness, familiarity, response urgency, and expected effort, as well as convincingness variables. To evaluate the framework and pave the way for computational modeling, we develop a novel role-playing-based annotation setup, mimicking real-world exposure to arguments. Participants disclose their emotion, explain the main cause, the argument appraisal, and the perceived convincingness. To consider the subjective nature of such annotations, we also collect demographic data and personality traits of both the participants and ask them to disclose the same variables for their perception of the argument sender. The analysis of the resulting corpus of 4000 annotations reveals that convincingness is positively correlated with positive emotions (e.g., trust) and negatively correlated with negative emotions (e.g., anger). The appraisal variables particularly point to the importance of the annotator’s familiarity with the argument.
Towards Clinical Applications of NLP: Detecting Emotion Regulation via Emotional Categories and Expression Modes in French Transcriptions
Salome Klein | Amalia Todirascu | Hélène Vassiliadou
Salome Klein | Amalia Todirascu | Hélène Vassiliadou
We present an annotated corpus of patient interview transcriptions, labeled for emotionality, polarity, intensity, and emotional category (at the sentence level), and for expression mode (at the token level). Three modes of expression are distinguished: Designated (explicit), Suggested (implicit causes), and Manifested (implicit consequences). The corpus has been collected during the GREMO-LING project and is used to measure the linguistic expressions of emotions in patients’ narratives. The corpus, consisting of 7,471 sentences, was used to fine-tune and evaluate several transformer-based language models, including the French BERT family. Sentence classification was performed for emotionality, emotion categories and expression modes. The best-performing models achieved F1 scores of 0.87 (emotionality, fine-tuned DistilCamemBERT), 0.58 (emotion categories, CamemBERTaV2), and 0.70 (expression modes, CamemBERT). We obtain solid results despite the high complexity of non-standard, spoken-derived data. These findings confirm the feasibility and relevance of automatic emotion detection in clinical discourse. We provide publicly available guidelines, annotated corpora and models, thereby establishing a methodological foundation for future research on the linguistic assessment of emotional regulation and its clinical implications, such as the evaluation of the Dialectical Behavioral Theray (DBT) in enhancing patients’ emotion regulation skills.
R.U.Psycho? A Framework for Robust Unified Psychometric Testing of Language Models
Julian Schelb | Orr Borin | David Garcia | Andreas Spitz
Julian Schelb | Orr Borin | David Garcia | Andreas Spitz
Generative language models are increasingly being subjected to psychometric questionnaires intended for human testing, in efforts to establish their traits, as benchmarks for alignment, or to simulate participants in social science experiments. While this growing body of work sheds light on the likeness of model responses to those of humans, concerns are warranted regarding the rigour and reproducibility with which these experiments may be conducted. Instabilities in model outputs, sensitivity to prompt design, parameter settings, and a large number of available model versions increase documentation requirements. Consequently, generalization of findings is often complex and reproducibility is far from guaranteed. In this paper, we present R.U.Psycho, a framework for designing and running robust and reproducible psychometric experiments on generative language models that reduces the required coding expertise. We demonstrate the capability of our framework on a variety of psychometric questionnaires, which lend support to prior findings in the literature. R.U.Psycho is available as a Python package at https://github.com/julianschelb/rupsycho.
Code-switching as a Bias Indicator in LLMs: "the Consequences Are Not the Same Para Nosotros"
Fanny Ducel | Aurélie Névéol | Vidit Khazanchi | Loïc Leclere | Arthur Pedrini | Léa Bouchet | Benjamin Caissial | Karen Fort
Fanny Ducel | Aurélie Névéol | Vidit Khazanchi | Loïc Leclere | Arthur Pedrini | Léa Bouchet | Benjamin Caissial | Karen Fort
Code-switching is a widespread linguistic practice among bilingual speakers. While recent studies have addressed the impact of code-switching on downstream task performance, the potential biases and harms that language models may cause when prompted with code-switching have yet to be investigated. The objective of this study is to investigate whether code-switching constitutes an implicit indicator of ethnicity that can be leveraged to unveil covert racist or xenophobic bias in language models. The present paper introduces a methodology to compare generated texts that were prompted with code-switching vs. with monolingual inputs. It is applied on both Hinglish and Spanglish, two popular forms of code-switching that are omnipresent in Indian and Hispanic communities. With a decision tree approach, we tackle various types of semantic differences through the use of semantic resources, stereotypes lists, POS-tagging and sentiment classifiers. Over 84k text pairs are generated with 3 popular large language models. Overall, around 50% of generated text pairs are not semantically equivalent, and 25% of the time, there is a potential for harm against the Indian or Hispanic community. The different possible harms are further discussed, relying on sociological studies to argue that bias and harms against socially discriminated communities have greater consequences.
Understanding how hate is framed in multimodal social media content is crucial for developing interpretable and robust hate detection systems. We present the MM-HateFrames Dataset, a large-scale resource encoding 2,298 Hate Frames (HFs) and their corresponding rationales discovered from two benchmark datasets—Hateful Memes and MMHS150K—comprising over 11K+ social media multimodal posts. This allowed us to explore several generative and non-generative methods to automatically discover the way hate is framed when relying on MM-HateFrames, including clustering-based methods and large multimodal models (LMMs) under zero-shot and few-shot settings. Experimental evaluations show that few-shot LMMs prompting generates the most coherent and sound frame articulations. The MM-HateFrames Dataset provides a valuable foundation for future research in hate speech understanding, frame articulation, and explainable multimodal NLP, enabling models to interpret not only whether content is hateful but also how hate is conceptually framed.
Are Social Biases in LLMs Consistent across Generative Tasks? A Case Study for Basque
Muitze Zulaika | Xabier Saralegi | Julia Shershneva | Lia Gonzalez | Arkaitz Fullaondo
Muitze Zulaika | Xabier Saralegi | Julia Shershneva | Lia Gonzalez | Arkaitz Fullaondo
Most bias benchmarks for Large Language Models (LLMs) rely on multiple-choice formats, overlooking subtler biases that emerge in open-ended text generation. This gap is particularly relevant for low-resource languages like Basque, where culturally grounded evaluation resources are limited. We introduce BasqBBG (Basque Bias Benchmark for Generation), the first systematic benchmark for social bias in Basque Natural Language Generation (NLG), covering eight bias categories—including a newly added feminism dimension—adapted from the BasqBBQ dataset. We validate an LLM-as-a-Judge framework against expert human evaluations on two NLG tasks (story continuation and generative QA), achieving strong agreement (agreement of 0.78 in bias presence and 0.92 in bias directionality). We scale this approach to ten additional tasks and five models. Results show that bias levels vary markedly across tasks and depend more on model family than size: Llama-based models exhibit higher and less consistent bias (45–50%), whereas GPT-4o and the Gemma-based Kimu-9B remain substantially fairer (≤20%). Our findings highlight the need for task-aware, language-specific frameworks to assess social bias in generative LLMs. Keywords: Large Language Models, Social Bias, Basque, Natural Language Generation, Benchmarking, Manual Evaluation, LLM-as-a-judge.
Fine-grained Narrative Classification in Biased News Articles
Zeba Afroz | Harsh Vardhan | Pawan Bhakuni | Aanchal Punia | Rajdeep Kumar | Md. Shad Akhtar
Zeba Afroz | Harsh Vardhan | Pawan Bhakuni | Aanchal Punia | Rajdeep Kumar | Md. Shad Akhtar
Narratives are the cognitive and emotional scaffolds of propaganda. They organize isolated persuasive techniques into coherent stories that justify actions, attribute blame, and evoke identification with ideological camps. In this paper, we propose a novel fine-grained narrative classification in biased news article. We also explore article-bias classification as the pre-cursor task to narrative classification and fine-grained persuassive technique identification. We develop INDI-PROP, the first ideologically grounded fine-grain narrative dataset with multi-level annotation for analyzing propaganda in Indian news media. Our dataset INDI-PROP comprises 1,266 articles focusing on two polarizing socio-political events in recent times: CAA/NRC and the Farmers’ protest. Each article is annotated at three hierarchical levels: (i) ideological article-bias (pro-government, pro-opposition, neutral), (ii) event-specific fine-grained narrative frames anchored in ideological polarity and communicative intent, and (iii) persuasive techniques. We propose FANTA and TPTC, two GPT-4o guided multi-hop prompt-based reasoning frameworks for the bias, narrative, and persuasive technique classification. FANTA leverages multi-layered communicative phenomenon by integrating information extraction and contextual framing for hierarchical reasoning. On the other hand, TPTC adopts systematic decomposition of persuasive cues via a two-stage approach. Our evaluation suggest substantial improvement over underlying baselines in each case.
A Shoal of Voices: Parallel Read Speech from Professional Swedish Narrators
Christina Tånnander | Jim O’Regan | Jens Edlund
Christina Tånnander | Jim O’Regan | Jens Edlund
We present a shoal of voices in Storspigg–TBI, a legally cleared, professionally recorded Swedish speech corpus derived from talking-book production at the Swedish Agency for Accessible Media (MTM). The corpus contains 1 000 information messages read by 99 narrators under controlled studio conditions. The material has undergone full legal assessment and a three-sweep adoption process ensuring provenance, FAIR/FACT compliance, and reproducibility in collaboration with the national research infrastructure Språkbanken Tal. The paper describes the legal framework, data-selection and curation pipeline, as well as initial automatic transcription using Swedish Whisper and wav2vec 2.0 models. The resulting corpus provides a high-quality reference resource for speech science and technology, supporting research on inter-speaker variation, prosody, and evaluation under consistent acoustic and linguistic conditions.
Deep Learning-Based Multi-Aspect Pronunciation Assessment for Individuals with Down Syndrome
David Fernández-García | César González-Ferreras | Valentín Cardeñoso-Payo | Mario Corrales-Astorgano
David Fernández-García | César González-Ferreras | Valentín Cardeñoso-Payo | Mario Corrales-Astorgano
This paper explores the use of an annotated speech corpus to assess multiple dimensions of speech quality—particularly phonetic, fluency and prosody—in individuals with Down syndrome, with the aim of informing the development of automated assessment tools. We conducted a series of experiments using the GOPT model, together with representations extracted from fine-tuning Wav2Vec models focused on phoneme classification. Model predictions were compared against expert annotations from a speech-language pathologist using Pearson correlation. Results demonstrate significant improvements over prior work, with correlations up to 0.49 in certain aspects, particularly for phonetic and fluency dimensions, while prosody remained more challenging to model. The study highlights the potential of Transformer-based architectures for atypical speech assessment and underscores the challenges inherent in assessing atypical speech, particularly due to variability linked to specific disfluency types.
WikIPA: Integrating WikiPron and Lingua Libre for Multilingual IPA Transcription
Pierluigi Cassotti | Jacob Lee Suchardt | Domenico De Cristofaro
Pierluigi Cassotti | Jacob Lee Suchardt | Domenico De Cristofaro
We present WikIPA, a new multilingual benchmark designed for automatic speech-to-IPA (STIPA) transcription. By integrating human-curated IPA transcriptions from WikiPron with spoken recordings and metadata from Lingua Libre, WikIPA connects textual phonetic representations with real speech across 78 languages. This open resource supports both broad (phonemic) and narrow (phonetic) transcription tasks, enabling fine-grained evaluation of multilingual phonetic transcription systems. WikIPA provides over 289,000 paired entries and serves as a large-scale foundation for STIPA. We benchmark several state-of-the-art STIPA systems, including MultIPA, (Lo)WhIPA, and ZIPA. Results show that ZIPA achieves the lowest mean error rates across most languages, outperforming Whisper- and Wav2Vec-based baselines. Error analyses reveal that remaining discrepancies largely stem from minor phonetic confusions rather than complete transcription failures, emphasizing the challenge of modeling fine-grained articulatory variation. WikIPA thus establishes the first systematic, multilingual evaluation framework for speech-to-IPA transcription and highlights the potential of combining open, community-driven resources to advance STIPA evaluation.
How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse
Saki Imai | Lee Kezar | Laurel Aichler | Mert Inan | Erin Walker | Alicia Wooten | Lorna Cobban Quandt | Malihe Alikhani
Saki Imai | Lee Kezar | Laurel Aichler | Mert Inan | Erin Walker | Alicia Wooten | Lorna Cobban Quandt | Malihe Alikhani
Most state-of-the-art sign language models are trained on interpreter or isolated vocabulary data, which overlooks the variability that characterizes natural dialogue. However, human communication dynamically adapts to contexts and interlocutors through spatiotemporal changes and articulation style. This specifically manifests itself in educational settings, where novel vocabularies are used by teachers, and students. To address this gap, we collect a motion capture dataset of American Sign Language (ASL) STEM (Science, Technology, Engineering, and Mathematics) dialogue that enables quantitative comparison between dyadic interactive signing, solo signed lecture, and interpreted articles. Using continuous kinematic features, we disentangle dialogue-specific entrainment from individual effort reduction and show spatiotemporal changes across repeated mentions of STEM terms. On average, dialogue signs are 24.6%-44.6% shorter in duration than the isolated signs, and show significant reductions absent in monologue contexts. Finally, we evaluate sign embedding models on their ability to recognize STEM signs and approximate how entrained the participants become over time. Our study bridges linguistic analysis and computational modeling to understand how pragmatics shape sign articulation and its representation in sign language technologies.
Setting the Stage for Disfluency: Implications of Contextual Task Framing Effects for the Design of Listening Tasks
Ambika Kirkland | Jens Edlund
Ambika Kirkland | Jens Edlund
Speech disfluencies have been shown to impact both judgments about a speaker’s competence and decisions about which source of information to rely on. However, fluency effects more broadly are highly sensitive to context: they are strongest when there is little other information available to inform judgments and decisions, and can be attenuated or even reversed by metacognitive processes. Speech is generally experienced in the context of interactions, where listeners have access to a plethora of information about the speaker and other parameters relevant to decision-making. It is hence crucial to consider how the outcomes of studies on speech disfluencies might be impacted by the framing of experimental tasks and the information available to participants. We carried out a decision-making task where participants had to choose which of two speakers, one fluent and one disfluent, had answered a trivia question correctly. The task was presented in the context of three scenarios which provided different information about the speakers. We replicated previous findings that listeners preferred fluent answers in only one of these three contexts, demonstrating the importance of task framing.
ACAData: Parallel Dataset of Academic Data for Machine Translation
Iñaki Lacunza | Javier Garcia Gilabert | Francesca De Luca Fornaciari | Javier Aula-Blasco | Aitor Gonzalez-Agirre | Maite Melero | Marta Villegas
Iñaki Lacunza | Javier Garcia Gilabert | Francesca De Luca Fornaciari | Javier Aula-Blasco | Aitor Gonzalez-Agirre | Maite Melero | Marta Villegas
We present ACAData, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-Train, which contains approximately 1.5 million human-generated paragraph pairs across 12 languages, and ACAD-Bench, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its usefulness, we fine-tune two Large Language Models (LLMs) on ACAD-Train and benchmark them on ACAD-Bench against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine tuning on ACAD-Train leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best proprietary and open-weight models on the academic translation domain. By releasing ACAD-Train, ACAD-Bench and the fine-tuned models, we provide the community with a valuable resource to advance research in the academic domain and long-context translation.
A Single Model Ensemble Framework for Neural Machine Translation Using Pivot Translation
Seokjin Oh | Keonwoong Noh | Woohwan Jung
Seokjin Oh | Keonwoong Noh | Woohwan Jung
Despite the recent remarkable advances in neural machine translation, translation quality for low-resource language pairs remains subpar. Ensembling multiple systems is a widely adopted technique to enhance performance, often accomplished by combining probability distributions. However, previous approaches face the challenge of high computational costs for training multiple models. Furthermore, for black-box models, averaging token-level probabilities at each decoding step is not feasible. To address the problems of multi-model ensemble methods, we present a pivot-based single model ensemble. The proposed strategy consists of two steps: pivot-based candidate generation and post-hoc aggregation. In the first step, we generate candidates through pivot translation. This can be achieved with only a single model and facilitates knowledge transfer from high-resource pivot languages, resulting in candidates that are not only diverse but also more accurate. Next, in the aggregation step, we select k high-quality candidates from the generated candidates and merge them to generate a final translation that outperforms the existing candidates. Our experimental results show that our method produces translations of superior quality by leveraging candidates from pivot translation to capture the subtle nuances of the source sentence.
Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures
Chiara Manna | Hosein Mohebbi | Afra Alishahi | Frederic Blain | Eva Vanmassenhove
Chiara Manna | Hosein Mohebbi | Afra Alishahi | Frederic Blain | Eva Vanmassenhove
While Large Language Models achieve state-of-the-art results across a wide range of NLP tasks, they remain prone to systematic biases. Among these, gender bias is particularly salient in MT, due to systematic differences across languages in whether and how gender is marked. As a result, translation often requires disambiguating implicit source signals into explicit gender-marked forms. In this context, standard benchmarks may capture broad disparities but fail to reflect the full complexity of gender bias in modern MT. In this paper, we extend recent frameworks on bias evaluation by: (i) introducing a novel measure coined ’Prior Bias’, capturing a model’s default gender assumptions, and (ii) applying the framework to decoder-only MT models. Our results show that, despite their scale and state-of-the-art status, decoder-only models do not generally outperform encoder-decoder architectures on gender-specific metrics; however, post-training (e.g., instruction tuning) not only improves contextual awareness but also reduces the masculine Prior Bias.
Building a One-Million-Pair Bokmål–Nynorsk Translation Corpus: A Quality-First Harvesting and Cleaning Pipeline
Per E. Kummervold | Thea Tollersrud | Angelina Zanardi
Per E. Kummervold | Thea Tollersrud | Angelina Zanardi
We present a high-quality parallel corpus for translation between Norwegian Bokmål (nb) and Nynorsk (nn), two closely related written standards of Norwegian. The corpus was assembled from two complementary sources: Nasjonal digital læringsarena (NDLA), an educational platform, and Nynorsk pressekontor (NPK), a newswire service. Our methodology prioritizes precision over volume, employing a multi-stage filtering pipeline designed to address the specific challenges of aligning near-neighbor languages. This pipeline combines paragraph-level alignment, deduplication, multilingual semantic similarity scoring, language identification confidence checks, structural consistency tests, and strict bidirectional adjudication by a Large Language Model (LLM). To address the common problem of untranslated or placeholder “pending” copies, we apply a rule that flags pairs with zero semantic distance when the Nynorsk side shows weak evidence of being distinctively Nynorsk. After filtering, we retained 191,695 pairs from NDLA and 809,164 pairs from NPK, resulting in a merged corpus of 1,000,859 parallel paragraphs. This resource demonstrates that a precision-oriented pipeline can produce data better suited for training robust machine translation systems and instruction-tuned models than larger but noisier alternatives.
New Trends for Modern Machine Translation with Large Reasoning Models
Sinuo Liu | Chenyang Lyu | Minghao Wu | Zifu Shang | Longyue Wang | Weihua Luo | Kaifu Zhang
Sinuo Liu | Chenyang Lyu | Minghao Wu | Zifu Shang | Longyue Wang | Weihua Luo | Kaifu Zhang
Recent advances in Large Reasoning Models (LRMs), particularly those leveraging Chain-of-Thought reasoning (CoT), have opened brand new possibilities for Machine Translation (MT). This position paper argues that LRMs substantially transform traditional neural MT as well as LLMs-based MT paradigms by reframing translation as a dynamic reasoning task that requires contextual, cultural, and linguistic understanding and reasoning. We identify three foundational shifts: 1) contextual coherence, where LRMs resolve ambiguities and preserve discourse structure through explicit reasoning over cross-sentence and complex context or even lack of context; 2) cultural intentionality, enabling models to adapt outputs by inferring speaker intent, audience expectations, and socio-linguistic norms; 3) self-reflection, LRMs can perform self-reflection during inference to correct the potential translation errors, particularly in extremely noisy cases, showing better robustness compared to simply mapping X->Y translation. We explore various scenarios in translation including stylized translation, document-level translation and multimodal translation by showcasing empirical examples that demonstrate the superiority of LRMs in translation. We also identify several interesting phenomena for LRMs for MT including auto-pivot translation as well as the critical challenges such as over-localisation in translation and inference efficiency. In conclusion, we argue that LRMs redefine translation systems not merely as text converters but as multilingual cognitive agents capable of reasoning about meaning beyond the text. This paradigm shift reminds us to think of problems in translation beyond traditional translation scenarios in a much broader context with LRMs - what we can achieve on top of it.
MaitH 1.0: A Parallel Corpus and Baseline for Low-Resource Maithili-Hindi Translation
Kamanksha Prasad Dubey | Chandresh Maurya | Kumar Padmanabh
Kamanksha Prasad Dubey | Chandresh Maurya | Kumar Padmanabh
Maithili is one of the 22 official languages recognized in the Indian Constitution. The literature of Maithili is rich; however, due to current socio-political changes, the language is on the verge of extinction. Therefore, it is crucial to develop a corpus for low-resource Indic languages like Maithili to ensure that the dream of “No Language Left Behind" (NLLB) is realized. With this in mind, we contribute a corpus (1,05,600 sentences) containing both manually curated and synthetically generated. Additionally, we propose a strong baseline on the Maithali-Hindi pair using multilingual pretrained models such as IndicTrans2, mBART50, mT5, and NLLB-200 distilled. We evaluate the translation systems using standard performance metrics, including BLEU, CHRF2, TER, COMET, METEOR, and BERTScore. Comparative experiments conducted against the existing NLLB dataset (5,50,300 sentence pairs) demonstrate that our proposed dataset consistently yields superior translation quality. Finally, these results demonstrate that, even with a smaller corpus size, high-quality, task-specific data significantly enhance translation accuracy for low-resource Indian languages, such as Maithili.
NRD: A Hybrid Disentanglement Framework for Mitigating Interference in Multilingual Machine Translation
Jiarui Zhang | Yifan Deng
Jiarui Zhang | Yifan Deng
Negative interference from cross-lingual conflicting syntactic patterns is a primary obstacle in Multilingual Neural Machine Translation (MNMT). We trace this problem to the entanglement of transferable, universal semantics with non-transferable, language-specific syntactic structures. Existing methods, relying on disjoint training-only specialization or inference-only filtering, fail to fully resolve this fundamental entanglement. To address this, we propose NRD (Neuron Representation Disentanglement), a two-stage hybrid framework that couples training-time specialization with inference-time filtering. First, a Specialization Fine-tuning stage identifies functional neurons via a semantic-invariant activation-variance metric and reinforces intrinsic modularity through sparse updates. Second, a Dynamic Representation Filtering stage purifies semantic representations at inference by adaptively suppressing syntax-sensitive neurons, guided by each language’s pre-computed gradient consistency. On the OPUS-100 benchmark, NRD outperforms strong baselines, achieving an average gain of +1.9 BLEU on supervised directions. On the WMT-10 zero-shot benchmark, it obtains a substantial +7.1 BLEU, demonstrating robust cross-lingual generalization. These results provide strong evidence that our hybrid approach effectively purifies semantic representations by mitigating syntactic interference, paving the way for more robust cross-lingual generalization.
Linguistic and Demographic Factors in an Online Free Translation Task
Tyler Lee | Irina Stenger | Tania Avgustinova
Tyler Lee | Irina Stenger | Tania Avgustinova
Humans are remarkably adept of understanding unfamiliar languages, in part by utilizing resources from languages they do know. In this study, we investigated how various linguistic factors (word order, lexical distance) and demographic factors affected the speed and correctness of translations in a multilingual scenario. In free translation task conducted online, participants read Polish noun phrases and translated them into English text. The noun phrases were varied between noun-adjective and adjective-noun word order, and the number of international words varied among the stimuli. Both the accuracy and total response time were recorded, and additional demographic data was recorded for all participants. Participants were more successful at translating noun phrases composed of two international terms than those with one or no such words. Additionally, speakers of other Slavic languages were more accurate despite not knowing Polish than participants who knew no Slavic languages. Although word order had little or no effect on accuracy for participants overall, speakers of Slavic languages translated the noun-adjective stimuli more accurately overall.
Biases in Translation: Assessing Opinion Distortion in Machine Translated Texts
Nazanin Shafiabadi | François Yvon
Nazanin Shafiabadi | François Yvon
Current machine translation (MT) evaluation practices largely assume that high lexical and semantic fidelity implies preservation of meaning. We question this assumption by introducing a framework for detecting and quantifying translation-induced distortion—the systematic alteration of a text’s subjective properties during translation. Focusing on stance as a socially consequential property, we formalize stance preservation as an invariance problem and adapt two classical statistical tests, McNemar’s test and the two-proportion Z-test, to diagnose systematic opinion shifts between source texts and their translations. Unlike standard MT metrics such as BLEU or COMET, which prioritize surface similarity and adequacy, our approach explicitly targets preservation of subjective meaning. In controlled experiments with synthetically distorted translations, we demonstrate that the proposed tests are sensitive to graded levels of stance manipulation. We apply our framework to evaluate twelve multilingual models and find that none reliably preserve stance across all tested language directions. Our findings reveal a critical gap in current MT evaluation practices and highlight the need for explicit evaluation of subjective meaning preservation in socially and politically sensitive contexts.
When Translations Surprise: Human Awareness of Predictability in Translations
Cristian García-Romero | Miquel Esplà-Gomis | Felipe Sanchez-Martinez
Cristian García-Romero | Miquel Esplà-Gomis | Felipe Sanchez-Martinez
Machine translation (MT) has achieved near-human quality for some language pairs, yet its output remains distinct from human translation, primarily in its predictability. While MT systems generate low-perplexity text, humans produce less predictable outputs. This raises the question of whether humans can intuitively use this difference in predictability to distinguish between human- and machine-translated text. We report on a study with 30 native Spanish speakers tasked with identifying the origin of English-to-Spanish translations. We compared their performance against two perplexity-based baselines: a large language model capturing fluency, and a neural MT model, conditioned on the source text, capturing both fluency and adequacy. Our findings reveal that human judgments correlate with fluency-based perplexity, but show no correlation with the perplexity that also accounts for adequacy. This suggests that annotators’ decisions are driven by the target text’s fluency. Consequently, a simple computational baseline using source-aware perplexity significantly outperforms human annotators. This work contributes to a deeper understanding of human perception of MT, highlighting a potential bias in current evaluation protocols toward fluency over adequacy. This bias may lead to an overestimation of the capabilities of highly fluent systems and underscores the need for evaluation methods ensuring translation adequacy is not overlooked.
Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation
Xinyue Ma | Pol Pastells | Mireia Farrus | Mariona Taule
Xinyue Ma | Pol Pastells | Mireia Farrus | Mariona Taule
Machine Translation (MT) evaluation has gone beyond metrics, towards more specific linguistic phenomena. Regarding English-Chinese language pairs, passive sentences are constructed and distributed differently due to language variation, thus need special attention in MT. This paper proposes a bidirectional multi-domain dataset of passive sentences, extracted from five Chinese-English parallel corpora and annotated automatically with structure labels according to human translation, and a test set with manually verified annotation. The dataset consists of 73,965 parallel sentence pairs (2,358,731 English words, 3,498,229 Chinese characters). We evaluate two state-of-the-art open-source MT systems with our dataset, and four commercial models with the test set. The results show that, unlike humans, models are more influenced by the voice of the source text rather than the general voice usage of the source language, and therefore tend to maintain the passive voice when translating a passive in either direction. However, models demonstrate some knowledge of the low frequency and predominantly negative context of Chinese passives, leading to higher voice consistency with human translators in English-to-Chinese translation than in Chinese-to-English translation. Commercial NMT models scored higher in metric evaluations, but LLMs showed a better ability to use diverse alternative translations. Datasets and annotation script will be shared upon request.
Proper treatment of terms is an important and critical aspect in machine translation. It is therefore necessary to use appropriate metrics to evaluate MT system outputs from terminology perspective. However, despite the great improvements witnessed in the recent NMT and LLM models, MT system evaluation metrics that shed light on specific aspects of term translations are yet to be fully explored. In this paper, we propose CoTERM, a new metric for automatic evaluation of term translations based on the Herfindahl-Hirshman Index (HHI). CoTERM measures target term closeness to one or more reference translations, taking into account the fundamental criteria for translating terms, i.e. (i) accuracy; (ii) consistency at document or corpus levels; and (iii) appropriateness to the domain conventions with regard to term variations. The proposed metric correlates strongly with human raters, and empirical evaluations of a wide range of NMTs and LLMs show that the best MT systems in standard metrics are not necessarily the best at treating terms. CoTERM is thus shown to be highly useful for diagnosing MT systems’ term translation performance and conveniently seen as complementary to generic measures for MT system evaluations.
SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
Hannah Liu | Junghyun Min | Annie En-Shiun Lee | Ethan Yue Heng Cheung | Shou-Yi Hung | Elsie Chan | Shiyao Qian | Runtong Liang | Kimlan Huynh | Wing Yu Yip | York Hay Ng | Tsz Fung Yau | Ka Ieng Charlotte Lo | You-Wei Wu | Richard Tzong-Han Tsai
Hannah Liu | Junghyun Min | Annie En-Shiun Lee | Ethan Yue Heng Cheung | Shou-Yi Hung | Elsie Chan | Shiyao Qian | Runtong Liang | Kimlan Huynh | Wing Yu Yip | York Hay Ng | Tsz Fung Yau | Ka Ieng Charlotte Lo | You-Wei Wu | Richard Tzong-Han Tsai
Despite major advances in machine translation (MT) in recent years, progress remains limited for many low-resource languages that lack large-scale training data and linguistic resources. In this paper, we introduce SINITICMTERROR, a novel fine-grained dataset that builds on existing parallel corpora to provide error span, error type, and error severity annotations in machine-translated examples from English to Mandarin, Cantonese, and Wu Chinese, along with a Mandarin-Hokkien component derived from a non-parallel source. Our dataset serves as a resource for the MT community to fine-tune models with error detection capabilities, supporting research on translation quality estimation, error-aware generation, and low-resource language evaluation. We also establish baseline results using language models to benchmark translation error detection performance. Specifically, we evaluate multiple open source and closed source LLMs using span-level and correlation-based MQM metrics, revealing their limited precision, underscoring the need for our dataset. Finally, we report our rigorous annotation process by native speakers, with analyses on pilot studies, iterative feedback, insights, and patterns in error type and severity.
Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models
Spyridon Mavromatis | Sokratis Sofianopoulos | Prokopis Prokopidis | Maria Giagkou
Spyridon Mavromatis | Sokratis Sofianopoulos | Prokopis Prokopidis | Maria Giagkou
Machine Translation (MT) for Ancient Greek (AG) to Modern Greek (MG) is a low-resource task, constrained by the lack of large-scale, high-quality parallel data. We address this gap by introducing the AG-MG Parallel Corpus, a new resource containing 132,481 sentence-aligned pairs derived from literary, historical, and biblical texts. We present a novel corpus creation pipeline that combines web-scraped, excerpt-level data with a multi-stage sentence-level alignment, and refinement process. Our method uses VecAlign with LaBSE embeddings, which we first fine-tune on a manually-aligned AG-MG subset, followed by an LLM-based error/misalignment correction phase using Gemini 2.5 Flash to ensure high alignment quality. Furthermore, we provide the first comprehensive benchmark of modern MT models on this task, evaluating three fine-tuning strategies across NMT models (NLLB, M2M100) and a Greek LLM (Llama-Krikri-8B). Our experiments show that fine-tuning yields significant improvements over base models, increasing performance by up to +10.3 BLEU points. Specifically, full-parameter fine-tuning of Llama-Krikri-8B achieves the highest overall performance with a BLEU score of 13.16, while the QLoRA-adapted M2M100-1.2B model demonstrates the largest relative gains and highly competitive results. Our dataset and models represent a significant contribution to Greek NLP.
Linguistic Knowledge-Infused Fine-Tuning for Mitigating Gender Bias in Machine Translation
Luis Ernesto Garcia Estrada | Audrey Mash | Carlos Escolano | Maite Melero | Christine Basta
Luis Ernesto Garcia Estrada | Audrey Mash | Carlos Escolano | Maite Melero | Christine Basta
Large Language Models (LLMs) achieve strong performance in machine translation (MT) but often encode gender bias, particularly when translating from non-gendered into gendered languages. This paper introduces a fine-tuning strategy to mitigate such bias in English-Spanish and English-Catalan translation. Using parameter-efficient LoRA fine-tuning, we apply linguistic knowledge infusion—a reasoning-based method that trains models to identify gendered referents and syntactic cues before generating translations. Experiments with Mistral–7B and Salamandrata–7B on MT-GenEval show that linguistically infused models improve gender accuracy by 15 percentage points and reduce gender gaps by 27 points in English-Spanish translation, with comparable trends for Catalan. Gains are strongest for Mistral, suggesting that explicit linguistic reasoning particularly benefits general-purpose LLMs. Overall, these results demonstrate that structured linguistic priors can enhance fairness and referential consistency in multilingual machine translation.
What Triggers My Model? Contrastive Explanations Inform Gender Choices by Translation Models
Janiça Hackenbuchner
Janiça Hackenbuchner
Interpretability can be implemented to understand decisions taken by (black box) models, such as neural machine translation (NMT) or large language models (LLMs). Yet, research in this area has been limited in relation to a manifested problem in these models: gender bias. In this work, we aim to move away from simply measuring bias to exploring its origins. Working with gender-ambiguous natural source data, this exploratory study examines which context, in the form of input tokens in the source sentence (EN), influences (or triggers) the NMT model’s choice of a certain gender inflection in the target languages (DE/ES). To analyse this, we compute saliency attribution based on contrastive translations. We first address the challenge of the lack of a scoring threshold and specifically examine different attribution levels of source words on the model’s gender decisions in the translation. We compare salient source words with human perceptions of gender and demonstrate a noticeable overlap between human perceptions and model attribution. Additionally, we provide a linguistic analysis of salient words. Our work showcases the relevance of understanding model translation decisions in terms of gender, how this compares to human decisions and that this information should be leveraged to mitigate gender bias.
ViKhoMT: A Vietnamese–K’Ho Neural Machine Translation Dataset and Evaluation for Community Health Communication
Tram Truong | Vinh Nguyen | Dang Van Thin | Ngan Nguyen
Tram Truong | Vinh Nguyen | Dang Van Thin | Ngan Nguyen
The Vietnamese government is prioritizing the socio-economic development and societal integration of ethnic minorities, including the K’Ho people. However, the lack of digital resources creates significant communication barriers, particularly in the critical domain of community health. To address this gap, we introduce ViKhoMT, a new, professionally curated Vietnamese-K’Ho parallel dataset containing approximately 10,000 sentence pairs focused on community health communication. To demonstrate the dataset’s quality and establish performance benchmarks, we conducted comprehensive evaluations by fine-tuning several pre-trained Neural Machine Translation (NMT) models. Our experiments show that a system based on the M2M100 architecture achieves BLEU scores of 60.5 for K’Ho-to-Vietnamese and 56.4 for Vietnamese-to-K’Ho, respectively. We release our dataset to the research community for free research purposes to support future studies and the development of practical translation tools for the K’Ho community. The dataset is publicly available at https://github.com/NgocTram2711/ViKhoMT.
Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
Malik Marmonier | Benoît Sagot | Rachel Bawden
Malik Marmonier | Benoît Sagot | Rachel Bawden
This paper investigates two complementary paradigms for predicting machine translation quality: source-side difficulty prediction and candidate-side quality estimation (QE). The rapid adoption of Large Language Models (LLMs) into machine translation (MT) workflows is reshaping the research landscape, yet its impact on established quality prediction paradigms remains underexplored. We study this issue through a series of “hindsight” experiments on a unique, multi-candidate dataset resulting from a genuine machine translation post-editing (MTPE) project. The dataset consists of over 6,000 English source segments with nine translation hypotheses from a diverse set of traditional neural MT systems and advanced LLMs, all evaluated against a single, final human post-edited reference. Using Kendall’s rank correlation, we assess the predictive power of source-side difficulty metrics, candidate-side QE models and position heuristics against two gold-standard scores: TER (as a proxy for post-editing effort) and COMET (as a proxy for human judgment). Our analysis yields three primary findings: (1) On the source side, the predictive power of difficulty metrics is highly contingent on the reference metric used; features that strongly correlate with COMET (e.g., segment length, neural predictors) show much weaker correlation to TER. (2) On the candidate side, we find a significant mismatch between QE model rankings and final human-adjudicated quality, and further show that modern QE metrics are significantly more aligned with the quality of traditional neural MT outputs than with those from general-purpose LLMs. (3) While we confirm a statistically significant positional bias in document-level LLMs (i.e., the tendency for translation quality to degrade for segments occurring later in a document) its practical impact on translation quality appears to be negligible. These findings highlight that the architectural shift towards LLMs alters the reliability of established quality prediction methods while simultaneously mitigating previous challenges in document-level translation.
PETra: A Multilingual Corpus of Pragmatic Explicitation in Translation
Doreen Osmelak | Koel Dutta Chowdhury | Uliana Sentsova | Cristina España-Bonet | Josef van Genabith
Doreen Osmelak | Koel Dutta Chowdhury | Uliana Sentsova | Cristina España-Bonet | Josef van Genabith
Translators often enrich texts with background details that make implicit cultural meanings explicit for new audiences. This phenomenon, known as pragmatic explicitation, has been widely discussed in translation theory but rarely modeled computationally. We introduce PeTra, the first multilingual corpus and detection framework for pragmatic explicitation. The corpus consists of 2,900 sentence pairs from TED-Multi and Europarl, covers twelve language pairs, and includes additions such as entity descriptions, measurement conversions, and translator remarks. We identify candidates through null alignments and refine them using active learning with human annotation. Our results show that entity and system-level (e.g., metric conversions) explicitations are most frequent, and that active learning improves classifier accuracy by 7-8 percentage points, achieving up to 0.88 accuracy and 0.82 F1 for the best transfer languages. PeTra establishes pragmatic explicitation as a measurable, cross-linguistic phenomenon and takes a step towards building culturally aware machine translation.
A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
Jenny Kunz | Anja Jarochenko | Marcel Bollmann
Jenny Kunz | Anja Jarochenko | Marcel Bollmann
Translations often carry traces of the source language, a phenomenon known as translationese. We introduce the first freely available English-to-Swedish dataset contrasting translationese sentences with idiomatic alternatives, designed to probe intrinsic preferences of language models. It includes error tags and descriptions of the problems in the original translations. In experiments evaluating smaller Swedish and multilingual LLMs with our dataset, we find that they often favor the translationese phrasing. Human alternatives are chosen more often when the English source sentence is omitted, indicating that exposure to the source biases models toward literal translations, although even without context models often prefer the translationese variant. Our dataset and findings provide a resource and benchmark for developing models that produce more natural, idiomatic output in non-English languages.
STAR-IL: A Dataset for Style-Aware Machine Translation of Product Reviews in Indian Languages
Ketaki Shetye | Dipti Misra Sharma | Parameswari Krishnamurthy
Ketaki Shetye | Dipti Misra Sharma | Parameswari Krishnamurthy
Product reviews on e-commerce platforms are a critical form of user-generated content that influence consumer decisions. However, these reviews are predominantly in English, creating a significant accessibility barrier for users who are not fluent in English. When translating into major Indian languages using the current models, the outputs often fail to capture domain-specific features and colloquial style, resulting in stylistically unnatural texts. To address this gap, we introduce STAR-IL, a human-annotated, multilingual, parallel corpus for style-aware translation of product reviews. We evaluate the performance of several state-of-the-art models on our dataset for the task of product review translation. Our experiments show that models fine-tuned on STAR-IL achieve significant average performance gain of 5.77 points in BLEU and 3.78 points in COMET, when compared to their baselines, across all languages. Our dataset provides a valuable benchmark for future research in style-aware product review translation. The STAR-IL dataset is publicly available at https://github.com/ltrc/STAR-IL-Corpus.
Cultural and Knowledge Biases in LLMs through the Lens of Entity-Aware Machine Translation
Lu Xu | Luca Moroni | Roberto Navigli
Lu Xu | Luca Moroni | Roberto Navigli
Large Language Models (LLMs) demonstrate strong multilingual capabilities yet exhibit systematic cultural biases that affect entity-aware machine translation. While external knowledge integration improves translation accuracy, the extent of these benefits across varying degrees of cultural specificity remains unexplored. We propose a three-level cultural specificity framework: Culturally Agnostic, Culturally Sensitive, and Culturally Local, to systematically analyze how cultural context affects entity translation difficulty and the utility of external knowledge. Through experiments spanning 11 LLMs and 10 languages, we demonstrate that external knowledge provides substantially greater improvements for culturally local entities (up to 70% in m-ETA) compared to culturally agnostic ones. Our analysis reveals distinct behavioral patterns across model tiers: closed and open-weight models show synergistic improvements in both entity accuracy and overall translation quality, while open-data models struggle with instruction-following despite improved entity accuracy.
Referenceless Evaluation of Machine Translation Models by Ranking Performance in Romanian to English Translate-train Settings
Mihail Feraru | Alexandra Diaconu | Bogdan Dumitru Alexe
Mihail Feraru | Alexandra Diaconu | Bogdan Dumitru Alexe
We propose a referenceless evaluation method for machine translation (MT) models by assessing their performance in translate-train scenarios across a variety of natural language processing (NLP) tasks. The approach ranks MT systems based on the downstream impact of their translations on independent NLP models trained on translated data, thus eliminating the need for professional ground-truth references. We evaluate four prominent MT tools — ChatGPT 3.5 Turbo, DeepL, Google Translate, and Mistral 7B Instruct v0.2 — on the Romanian→English language pair and analyze their influence on text summarization, sentiment analysis, and authorship identification. To further test the generalization and robustness of our method, we extend the evaluation to a cross-modality setup using out-of-domain speech data. In this setting, speech segments are transcribed with Whisper-Large, translated into English, and used in a four-class domain classification task (children’s stories, audiobooks, film dialogues, podcasts). Our findings show that translation improves downstream performance for sentiment analysis and summarization, while stylistically rich texts such as poetry or noisy ASR transcriptions suffer degradation. The proposed ranking metric correlates strongly with human judgments and remains sensitive to translation quality even in multimodal pipelines, providing a scalable and practical alternative to reference-based MT evaluation.
Every Word Presented in Context: Syntactic Coverage as Objective for Low-Resource Machine Translation with Large Language Models
Samuel Frontull | Thomas Ströhle
Samuel Frontull | Thomas Ströhle
Large Language Models (LLMs) have demonstrated strong capabilities in multilingual machine translation. However, they underperform for low-resource languages, indicating the need for more explicit instructional guidance. In this work, we introduce Fragment-Shot Prompting, a novel few-shot prompting method that aims to retrieve examples for every word occurring in the sentence to be translated, illustrating their use and meaning in context. We evaluate our method on translation between Italian, Ladin (Val Badia) and Ladin (Gherdëina) and compare its performance with zero-shot prompting, random few-shot prompting, as well as established lexical and semantic retrieval strategies. We conduct these experiments using state-of-the-art LLMs, including GPT-3.5, GPT-4o, o1-mini, LlaMA-3.3, and DeepSeek-R1. Our results demonstrate that LLMs can extract substantial value from limited data when translating from a low- to the high-resource language. However, this does not apply to translations into the low-resource languages, where the prompting method plays a much more important role. In particular, our method consistently delivers the best results and enables significant gains. Even though translation performance into Ladin remains limited with the available resources, our results highlight the importance of syntactic coverage for improving translation accuracy and ariant-specific adaptation in low-resource scenarios.
Multilingual KokoroChat: A Multi-LLM Ensemble Translation Method for Creating a Multilingual Counseling Dialogue Dataset
Ryoma Suzuki | Zhiyang Qi | Michimasa Inaba
Ryoma Suzuki | Zhiyang Qi | Michimasa Inaba
To address the critical scarcity of high-quality, publicly available counseling dialogue datasets, we created Multilingual KokoroChat by translating KokoroChat, a large-scale manually authored Japanese counseling corpus, into both English and Chinese. A key challenge in this process is that the optimal model for translation varies by input, making it impossible for any single model to consistently guarantee the highest quality. In a sensitive domain like counseling, where the highest possible translation fidelity is essential, relying on a single LLM is therefore insufficient. To overcome this challenge, we developed and employed a novel multi-LLM ensemble method. Our approach first generates diverse hypotheses from multiple distinct LLMs. A single LLM then produces a high-quality translation based on an analysis of the respective strengths and weaknesses of all presented hypotheses. The quality of “Multilingual KokoroChat” was rigorously validated through human preference studies. These evaluations confirmed that the translations produced by our ensemble method were preferred from any individual state-of-the-art LLM. This strong preference confirms the superior quality of our method’s outputs. The Multilingual KokoroChat is available at https://github.com/UEC-InabaLab/MultilingualKokoroChat.
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
Rupak Raj Ghimire | Bipesh Subedi | Balaram Prasain | Prakash Poudyal | Praveen Acharya | Nischal Karki | Rupak Tiwari | Rishikesh Kumar Sharma | Jenny Poudel | Bal Krishna Bal
Rupak Raj Ghimire | Bipesh Subedi | Balaram Prasain | Prakash Poudyal | Praveen Acharya | Nischal Karki | Rupak Tiwari | Rishikesh Kumar Sharma | Jenny Poudel | Bal Krishna Bal
Modern Translation Systems heavily rely on high-quality, large parallel datasets for state-of-the-art performance. However, such resources are largely unavailable for most of the South Asian languages. Among them, Nepali and Tamang fall into such category, with Tamang being among the least digitally resourced languages in the region. This work addresses the gap by developing NepTam20K, a 20K gold standard parallel corpus, and NepTam80K, an 80K synthetic Nepali–Tamang parallel corpus, both sentence-aligned and designed to support machine translation. The datasets were created through a pipeline involving data scraping from Nepali news and online sources, pre-processing, semantic filtering, balancing for tense and polarity (in NepTam20K dataset), expert translation into Tamang by native speakers of the language, and verification by an expert Tamang linguist. The dataset covers five domains: Agriculture, Health, Education and Technology, Culture, and General Communication. To evaluate the dataset, baseline machine translation experiments were carried out using various multilingual pre-trained models:mBART, M2M-100, NLLB-200, and a vanilla Transformer model. The fine-tuning on the NLLB-200 achieved the highest sacreBLEU scores of 40.92 (Nepali → Tamang) and 45.26 (Tamang → Nepali).
Scoring the Translation: On Target Automatic Keyword-Based Evaluation of Machine Translation in the Sports Domain
Steinþór Steingrímsson | Einar Freyr Sigurðsson
Steinþór Steingrímsson | Einar Freyr Sigurðsson
We take a closer look at the results of a recent translation shared task at WMT 2025 (the Conference on Machine Translation) and analyse the errors in the output of the four highest-scoring systems. We revise the automatic evaluation method used in Sigurðsson et al. (2025) and compare it to manual evaluation of six machine translation systems. We find that our results are in line with the manual evaluation, indicating that the test suite can be well suited for evaluating machine translation in this domain. Finally, we publish a list of domain-specific sports terms, namely, in the domains of basketball, chess, football, golf and gymnastics.
Towards Improving Multimodal Machine Translation with LLMs: A Focus on Indic Languages
Amulya Ratna Dash | Chirag Wadhwa | Yashvardhan Sharma
Amulya Ratna Dash | Chirag Wadhwa | Yashvardhan Sharma
Recent advances in Multimodal Machine Translation (MMT) have attempted to address ambiguity and polysemy in text alone by enabling models to draw additional contextual cues from paired images, thereby improving disambiguation and translation accuracy. Datasets such as Multi30K and Visual Genome have significantly advanced this line of research. However, these datasets do not always compel models to rely on visual information. The CoMMuTE dataset takes a stronger step in this direction by serving as an evaluation benchmark specifically designed around ambiguous English sentences that can only be correctly interpreted with their accompanying images. In this work, we extend CoMMuTE to two Indic languages, introducing IndicCoMMuTE — an evaluation dataset for assessing MMT systems on low-resource Indic languages. We benchmark a range of open-source multimodal Large Language Models (< 15B parameters) and a strong text-only baseline across eight languages. We fine-tune one of these LLMs on two Indic languages. Our findings provide insights into the strengths and limitations of LLMs and establish IndicCoMMuTE as a valuable benchmark for future research on Multimodal Machine Translation in Indic languages.
Parallel Sentence Filtering for Low-Resource Language Pairs: A Case Study for Upper Sorbian, German, and Czech
Ruiyang Jiang | Shu Okabe | Alexander Fraser
Ruiyang Jiang | Shu Okabe | Alexander Fraser
As parallel corpora for low-resource languages are scarce, and automatic approaches to mine sentence pairs can lead to noisy datasets, parallel sentence filtering aims to detect only actual translations. We study here two language pairs: Upper Sorbian–German and Czech–German to represent both high and low availability of data resources. To evaluate filtering performance, we generate synthetic datasets by combining existing parallel corpora with synthetic non-parallel pairs, notably with five types of local semantic changes on the German side, such as negation or modality transformations. We represent sentences using three multilingual language models, XLM-R, Glot500m, and LaBSE, and train classifiers for the task. All three model representations led to worse filtering quality when pairs were altered more subtly, such as an antonym replacement. We still observed that a language model pre-trained on the considered language achieves more robust classification performance when sentence pairs are more ambiguous. We also evaluated a cross-lingual approach where the classifier is trained on the Czech–German pair and then applied to the Upper Sorbian–German pair. Such a language transfer paves the way for filtering other low-resource language pairs in the future.
OpenSubtitles2024: A Massively Parallel Dataset of Movie Subtitles for MT Development and Evaluation
Joerg Tiedemann | Hengyu Luo
Joerg Tiedemann | Hengyu Luo
This paper introduces OpenSubtitles2024, a massively parallel dataset compiled from translated subtitles. The collection includes an extensive collection of aligned training data based on user-contributed subtitles derived from OpenSubtitles.org and a dedicated held-out dataset for development and evaluation of machine translation and multilingual language models. The collection provides an increased language coverage and doubles the size of the previous edition. Furthermore, a careful procedure was applied to reserve a subset of the most recent subtitles for system development and evaluation. The collection covers 92 languages and language variants, aligned in over 3,000 bitexts containing 40 billion tokens in 7.7 million subtitle files. The test set comprises 2,022 language pairs. In addition, we also provide a multi-parallel test set that refers to a subset of the held-out data with synchronized alignments across 40 languages and 15 subtitles.
CREST: Universal Safety Guardrails through Cluster-Guided Cross-Lingual Transfer
Lavish Bansal | Naman Mishra
Lavish Bansal | Naman Mishra
Ensuring content safety in large language models (LLMs) is essential for their deployment in real-world applications. However, existing safety guardrails are predominantly tailored for high-resource languages, leaving a significant portion of the world’s population underrepresented who communicate in low-resource languages. To address this, we introduce CREST (CRoss-lingual Efficient Safety Transfer), a parameter-efficient multilingual safety classification model that supports 100 languages with only 0.5B parameters. By training on a strategically chosen subset of only 13 high-resource languages, our model utilizes cluster-based cross-lingual transfer from a few to 100 languages, enabling effective generalization to both unseen high-resource and low-resource languages. This approach addresses the challenge of limited training data in low-resource settings. We conduct comprehensive evaluations across six safety benchmarks to demonstrate that CREST outperforms existing state-of-the-art guardrails of comparable scale and achieves competitive results against models with significantly larger parameter counts (≥ 2.5B parameters). Our findings highlight the limitations of language-specific guardrails and underscore the importance of developing universal, language-agnostic safety systems that can scale effectively to serve global populations.
Semantic Alignment across Ancient Egyptian Language Stages via Normalization-Aware Multitask Learning
He Huang
He Huang
We study word-level semantic alignment across four historical stages of Ancient Egyptian. These stages differ in script and orthography, and parallel data are scarce. We jointly train a compact encoder-decoder model with a shared byte-level tokenizer on all four stages, combining masked language modeling (MLM), translation language modeling (TLM), sequence-to-sequence translation, and part-of-speech tagging under a task-aware loss with fixed weights and uncertainty-based scaling. To reduce surface divergence we add Latin transliteration and IPA reconstruction as auxiliary views. We integrate these views through KL-based consistency and through embedding-level fusion. We evaluate alignment quality using pairwise metrics, specifically ROC-AUC and triplet accuracy, on curated Egyptian–English and intra-Egyptian cognate datasets. Translation yields the strongest gains. IPA with KL consistency improves cross-branch alignment, while early fusion demonstrates limited efficacy. Although the overall alignment remains limited, the findings provide a reproducible baseline and practical guidance for modeling historical languages under real constraints. They also show how normalization and task design shape what counts as alignment in typologically distant settings.
Conditioning LLMs to Generate Code-Switched Text
Maite Heredia | Gorka Labaka | Jeremy Barnes | Aitor Soroa
Maite Heredia | Gorka Labaka | Jeremy Barnes | Aitor Soroa
Code-switching (CS) is still a critical challenge in Natural Language Processing (NLP), due to the limited availability of large-scale, diverse CS datasets for robust training and evaluation. Despite recent advances, the capabilities and limitations of LLMs in handling CS are still not fully understood. In this work, we investigate the extent to which LLMs can be used in a framework for CS text generation, focusing on the English-Spanish language pair. Our proposed methodology consists of back-translating natural CS sentences into monolingual English, and using the resulting parallel corpus to fine-tune LLMs to turn monolingual sentences into CS. We thoroughly analyse the models’ performance through a study on human preferences, a qualitative error analysis, an evaluation with popular reference-based metrics and LLM-based judgment. Results show that fine-tuning can be a key step to ensure that current LLMs consistently generate fluent code-switched text and that our methodology generates high-quality outputs, expanding research opportunities in CS communication. We find that traditional metrics do not correlate with human judgement when assessing the quality of the generated CS data, but LLM-based judgment aligns more closely with human preferences. We release our code and generated dataset under a CC-BY-NC-SA license.
Are the LLMs Capable of Maintaining at Least the Language Genus?
Sandra Mitrović | David Kletz | Ljiljana Dolamic | Fabio Rinaldi
Sandra Mitrović | David Kletz | Ljiljana Dolamic | Fabio Rinaldi
Large Language Models (LLMs) display notable variation in multilingual behavior, yet the role of genealogical language structure in shaping this variation remains underexplored. In this paper, we investigate whether LLMs exhibit sensitivity to linguistic genera by extending prior analyses on the MultiQ dataset. We first check if models prefer to switch to genealogically related languages when prompt language fidelity is not maintained. Next, we investigate whether knowledge consistency is better preserved within than across genera. We show that genus-level effects are present but strongly conditioned by training resource availability. We further observe distinct multilingual strategies across LLMs families. Our findings suggest that LLMs encode aspects of genus-level structure, but training data imbalances remain the primary factor shaping their multilingual performance.
Gender Bias in MT for a Genderless Language: New Benchmarks for Basque
Amaia Murillo | Olatz Perez-de-Viñaspre | Naiara Perez
Amaia Murillo | Olatz Perez-de-Viñaspre | Naiara Perez
Large language models (LLMs) and machine translation (MT) systems are increasingly used in our daily lives, but their outputs can reproduce gender bias present in the training data. Most resources for evaluating such biases are designed for English and reflect its sociocultural context, which limits their applicability to other languages. This work addresses this gap by introducing two new datasets to evaluate gender bias in translations involving Basque, a low-resource and genderless language. WinoMTeus adapts the WinoMT benchmark to examine how gender-neutral Basque occupations are translated into gendered languages such as Spanish and French. FLORES+Gender, in turn, extends the FLORES+ benchmark to assess whether translation quality varies when translating from gendered languages (Spanish and English) into Basque depending on the gender of the referent. We evaluate several general-purpose LLMs and open and proprietary MT systems. The results reveal a systematic preference for masculine forms and, in some models, a slightly higher quality for masculine referents. Overall, these findings show that gender bias is still deeply rooted in these models, and highlight the need to develop evaluation methods that consider both linguistic features and cultural context.
Optimizing Multilingual LLMs via Federated Learning: A Study of Client Language Composition
Aleix Sant | Jordi Luque | Carlos Escolano
Aleix Sant | Jordi Luque | Carlos Escolano
Federated Learning (FL) of Large Language Models (LLMs) in multilingual environments presents significant challenges stemming from heterogeneous language distributions across clients and disparities in language resource availability. To address these challenges, we extended the FederatedScope-LLM framework to support multilingual instruction-tuning experiments with LLMs. We also introduced a novel client-specific early stopping mechanism, Local Dynamic Early Stopping (LDES-FL), which allows clients to pause and resume local training based on client-side validation performance, enhancing training efficiency and sustainability. Through a series of experiments, we studied how client language composition — from fully monolingual to increasingly multilingual clients — affects multilingual quality, fairness and training cost. Monolingual local fine-tuning remains the most effective for single-language specialization, whereas federated training is better suited to learning a single balanced multilingual model. In FL, increasing within-client multilinguality leads to stronger and fairer global models, narrows the gap to centralized multilingual fine-tuning, and yields the largest gains for lower-resource languages, albeit at the cost of more optimization steps. Overall, our results identify client language composition as a key design variable in multilingual FL, shaping performance, fairness and efficiency.
Social media enables data-driven analysis of public opinion on contested issues. Target-Stance Extraction (TSE) is the task of identifying the target discussed in a document and the document’s stance towards that target. Many works classify stance towards a given target in a multilingual setting, but all prior work in TSE is English-only. This work introduces the first multilingual TSE benchmark, spanning Catalan, Estonian, French, Italian, Mandarin, and Spanish corpora. It manages to extend the original TSE pipeline to a multilingual setting without requiring separate models for each language. Our model pipeline achieves a modest F1 score of 12.78, underscoring the increased difficulty of the multilingual task relative to English-only setups and highlighting target prediction as the primary bottleneck. We are also the first to demonstrate the sensitivity of TSE’s F1 score to different target verbalizations. Together these serve as a much-needed baseline for resources, algorithms, and evaluation criteria in multilingual TSE.
MUNIChus: MUltilingual News Image Captioning Benchmark
Yuji Chen | Alistair Plum | Hansi Hettiarachchi | Diptesh Kanojia | Saroj Basnet | Marcos Zampieri | Tharindu Ranasinghe
Yuji Chen | Alistair Plum | Hansi Hettiarachchi | Diptesh Kanojia | Saroj Basnet | Marcos Zampieri | Tharindu Ranasinghe
The goal of news image captioning is to generate captions by integrating news article content with corresponding images, highlighting the relationship between textual context and visual elements. The majority of research on news image captioning focuses on English, primarily because datasets in other languages are scarce. To address this limitation, we release the first multilingual news image captioning benchmark, MUNIChus, comprising 9 languages, including several low-resource languages such as Sinhala and Urdu. We evaluate various state-of-the-art neural news image captioning models on MUNIChus and find that news image captioning remains challenging. We also make MUNIChus publicly available as a public leaderboard with over 20 models already benchmarked. We hope that MUNIChus will enable further advancements in developing and evaluating multilingual news image captioning models.
GlossMATE: Multi-Agent Translator Explanations for Glosses
Changbing Yang | Patrick Littell | Gabriel Bernier-Colborne | Yanfei Lu | Mengzhe Geng
Changbing Yang | Patrick Littell | Gabriel Bernier-Colborne | Yanfei Lu | Mengzhe Geng
This paper introduces GlossMATE, a multi-agent critique-and-judge system that translates the gloss line in Interlinear Glossed Text (IGT) into fluent English using Large Language Models (LLMs). GlossMATE integrates linguist-provided resources (e.g., gloss-tag explanations, lexicon entries, curated IGT) with in-context learning and a multi-agent critique-and-judge procedure that iteratively evaluates and refines candidate translations. Our experiments show that leveraging analogous examples, explicit linguistic explanations, and collaborative agent interactions can enhance translation quality across several low-resource and polysynthetic languages. We also incorporate human linguists into the critique loop for selected languages. Case studies on three Indigenous languages further demonstrate the complementary strengths of human-in-the-loop feedback and multi-agent reasoning for language documentation tasks.
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
Klaudia Thellmann | Bernhard Stadler | Michael Färber
Klaudia Thellmann | Bernhard Stadler | Michael Färber
Machine-translated benchmark datasets reduce costs and offer scale, but noise, loss of structure, and uneven quality weaken confidence. What matters is not merely whether we can translate, but also whether we can measure and verify translation reliability at scale. We study translation quality in the EU20 benchmark suite, which comprises five established benchmarks translated into 20 languages, via a three-step automated quality assurance approach: (i) a structural corpus audit with targeted fixes; (ii) quality profiling using a neural metric (COMET, reference-free and reference-based) with translation service comparisons (DeepL / ChatGPT / Google); and (iii) an LLM-based span-level translation error landscape. Trends are consistent: datasets with lower COMET scores exhibit a higher share of accuracy/mistranslation errors at span level (notably HellaSwag; ARC is comparatively clean). Reference-based COMET on MMLU against human-edited samples points in the same direction. We release cleaned/corrected versions of the EU20 datasets, and code for reproducibility. In sum, automated quality assurance offers practical, scalable indicators that help prioritize review – complementing, not replacing, human gold standards.
Resource-Lean Lexicon Induction for German Dialects
Robert Litschko | Barbara Plank | Diego Frassinelli
Robert Litschko | Barbara Plank | Diego Frassinelli
Automatic induction of high-quality dictionaries is essential for building lexical resources, yet low-resource languages and dialects pose several challenges: limited access to annotators, high degree of spelling variations, and poor performance of large language models (LLMs). We empirically show that statistical models (random forests) trained on string similarity features are surprisingly effective for inducing German dialect lexicons. They outperform LLMs, enable cross-dialect transfer, and offer a lightweight data-driven alternative. We evaluate our models intrinsically on bilingual lexicon induction (BLI) and extrinsically on dialect information retrieval (IR). On BLI, random forests outperform Mistral-123b while being more resource-lean. On dialect IR with BM25, using our dialect dictionaries for query expansion yields relative improvements of up to 28.9% in nDCG@10 and 50.7% in Recall@100. Motivated by the resource scarcity in dialects, we further investigate the extent to which models transfer across different German dialects, and their performance under varying amounts of training data.
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
Mirae Kim | Seonghun Jeong | Youngjun Kwak
Mirae Kim | Seonghun Jeong | Youngjun Kwak
Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, available resources for jailbreak detection are scarce, particularly in finance. To address this gap, we present FENCE, a bilingual (Korean–English) multimodal dataset for training and evaluating jailbreak detectors in financial applications. FENCE comprises 10k finance-domain text–image pairs across more than 15 finance categories, constructed via a three-step pipeline: transforming real-world financial FAQs into harmful queries using GPT-4o, collecting query-relevant images via keyword-based crawling, and fusing text and images with diverse layout strategies. Labels were assigned using GPT-4o as an evaluator, with human validation confirming 95% agreement. Experiments on 15 commercial and open-source VLMs reveal consistent vulnerabilities, with GPT-4o showing measurable attack success rates and open-source models displaying greater exposure. A baseline detector trained on FENCE achieves 99% in-distribution accuracy and maintains strong performance on external benchmarks. FENCE provides a focused resource for advancing multimodal jailbreak detection in finance and supporting safer AI deployment in sensitive domains. Content Warning: This paper includes example data that may be offensive.
Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
Keito Sasagawa | Shuhei Kurita | Daisuke Kawahara
Keito Sasagawa | Shuhei Kurita | Daisuke Kawahara
Multimodal Large Language Models (MLLMs) have seen rapid advances in recent years and are now being applied to visual document understanding tasks. They are expected to process a wide range of document images across languages, including Japanese. Understanding documents from images requires models to read what are written in them. Since some Japanese documents are written vertically, support for vertical writing is essential. However, research specifically focused on vertically written Japanese text remains limited. In this study, we evaluate the reading capability of existing MLLMs on vertically written Japanese text. First, we generate a synthetic Japanese OCR dataset by rendering Japanese texts into images, and use it for both model fine-tuning and evaluation. This dataset includes Japanese text in both horizontal and vertical writing. We also create an evaluation dataset sourced from the real-world document images containing vertically written Japanese text. Using these datasets, we demonstrate that the existing MLLMs perform worse on vertically written Japanese text than on horizontally written Japanese text. Furthermore, we show that training MLLMs on our synthesized Japanese OCR dataset results in improving the performance of models that previously could not handle vertical writing. The datasets and code are publicly available (https://github.com/llm-jp/eval_vertical_ja).
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
Kimihiro Hasegawa | Wiradee Imrattanatrai | Masaki Asada | Susan E. Holm | Yuran Wang | Xuanang Zhou | Ken Fukuda | Teruko Mitamura
Kimihiro Hasegawa | Wiradee Imrattanatrai | Masaki Asada | Susan E. Holm | Yuran Wang | Xuanang Zhou | Ken Fukuda | Teruko Mitamura
Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. To foster system development, we propose a new multimodal QA evaluation dataset on assembly activities. Our dataset, ProMQA-Assembly, consists of 646 QA pairs that require multimodal understanding of human activity videos and their instruction manuals in an online-style manner. For cost effectiveness in the data creation, we adopt a semi-automated QA annotation approach, where LLMs generate candidate QA pairs and humans verify them. We further improve QA generation by integrating fine-grained action labels to diversify question types. Additionally, we create 81 instruction task graphs for our target assembly tasks. These newly created task graphs are used in our benchmarking experiment, as well as in facilitating the human verification process. With our dataset, we benchmark models, including competitive proprietary multimodal models. We find that ProMQA-Assembly contains challenging multimodal questions, where reasoning models showcase promising results. We believe our new evaluation dataset contributes to the further development of procedural-activity assistants.
K-MIND: Korean Multimodal INteraction Data for Dyadic Conversation Analysis
Jae Hee Yang | Yuha Shin | Saim Shin | Je Woo Kim | Jin Yea Jang
Jae Hee Yang | Yuha Shin | Saim Shin | Je Woo Kim | Jin Yea Jang
We present the Korean Multimodal INteraction Data (K-MIND), a large-scale corpus of dyadic Korean dialogue that is designed to capture the multimodal richness of social interaction. The dataset includes 292 participants and 200 sets (935 clips) spanning 115 hours and 30 minutes, all aligned across verbal, paraverbal, and nonverbal modalities such as transcripts, acoustic features, and visual signals. For these modalities, we propose a comprehensive annotation scheme that enables nuanced yet consistent labeling of complex communicative behaviors, balancing theoretical soundness with practical feasibility. We further report analysis results of the corpus, including label distributions, within- and cross-layer analyses. These analyses illuminate the key properties of dyadic K-MIND and demonstrate its utility for advancing research in human–computer interaction as well as in interdisciplinary domains. To ensure continuous refinement, the corpus and framework are being validated in complementary studies and have been extended to triadic interactions (K-MIND Triadic) that model group dynamics, which will be included in upcoming releases.
Do Multimodal LLMs Understand Order? Measuring the Fragility of Multimodal Reasoning under Input Order Perturbations
Sheng-Lun Wei | Yu-Ling Liao | Hen-Hsen Huang | Hsin-Hsi Chen
Sheng-Lun Wei | Yu-Ling Liao | Hen-Hsen Huang | Hsin-Hsi Chen
Multimodal reasoning has progressed rapidly with large vision-language models (LVLMs), yet their robustness under input variations remains underexplored. This study investigates positional bias in LVLMs for multimodal multiple-choice questions. Our analysis shows that model predictions are sensitive to both choice and modality ordering. We conduct a large-scale evaluation on MMMU, CVQA, and MMBench using fourteen representative models. Further analysis examines how question properties, including difficulty, domain, and image type, affect robustness. We also assess whether text-based mitigation strategies transfer to the VQA setting and perform ablation studies on self-consistency and reasoning complexity. Overall, our findings provide the first comprehensive understanding of positional bias from a vision-language perspective, highlighting key challenges in achieving stable multimodal reasoning.
Early Fusion with Contrastive Learning: A Lightweight Alternative for Multi-modal Classification
Felix Wernlein | Abhik Jana | Sandipan Sikdar
Felix Wernlein | Abhik Jana | Sandipan Sikdar
With the emergence of numerous modalities, such as text, image, audio, etc., the use of effective multimodal systems has increased significantly. However, one of the significant challenges faced by such multimodal systems is effectively aligning and integrating diverse modalities. Several models have been proposed to address these issues; however, state-of-the-art performance is achieved by complex, heavyweight models (complexity measured in terms of trainable parameters) alone. Hence, we propose a simple yet effective lightweight framework explicitly designed for multimodal classification tasks, utilising the early fusion method combined with a contrastive learning approach. The early fusion method focuses on fusing different modalities at the input level, whereas contrastive learning allows a single modality to capture intra-modality relationships. Experiments on three different genres of multimodal classification datasets demonstrate that the proposed lightweight framework achieves performance comparable to the most competitive heavyweight state-of-the-art models and, in some cases, even outperforms them.
Multimodal Entrainment and Feedback in Online Group Meetings
Patrizia Paggio | Manex Agirrezabal | Giulia Di Cristina | Bart Jongejan | Costanza Navarretta
Patrizia Paggio | Manex Agirrezabal | Giulia Di Cristina | Bart Jongejan | Costanza Navarretta
This paper presents the results of a study on multimodal speaker behaviour in a corpus of online Zoom meetings. We investigate two questions: i) whether speakers display a higher degree of head movement when they exchange verbal feedback than when they don’t, as would be expected if verbal and gestural feedback reinforce one other, and ii) whether they move more or less similarly under the same conditions. Several linear mixed models were fitted to test the difference in head movement values in target and control intervals of two different durations. The results indicate that speakers indeed entrain by moving their heads more in target intervals where verbal feedback is present. This result confirms our expectations. However, speakers also appear to move in less similar ways in the same target intervals. This dissimilarity can be explained by the fact that not all speakers give the same type of gestural feedback, but also by noise created by non-communicative movements in which speakers adjust their positions or reach out for objects during the meeting.
MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling
Hyeyeon Kim | Sungwoo Han | Jingun Kwon | Hidetaka Kamigaito | Manabu Okumura
Hyeyeon Kim | Sungwoo Han | Jingun Kwon | Hidetaka Kamigaito | Manabu Okumura
In this study, we introduce a novel cover image generation task that produces both a concise summary and a visually corresponding image from a text-only document. Because no existing datasets are available for this task, we propose a multimodal pseudo-labeling method to construct high-quality datasets at low cost. We first collect documents with summaries, multiple images, and captions, and then exclude factually inconsistent instances. Our approach selects one image from multiple images accompanying each document. Using the gold summary, we independently rank both the images and their captions. Then, we annotate a pseudo-label for an image when both the image and its corresponding caption are ranked first in their respective rankings. Finally, we remove documents that contain direct image references within texts. Experimental results demonstrate that the proposed multimodal pseudo-labeling method constructs more precise datasets and generates higher quality images than text- and image-only pseudo-labeling methods, which consider captions and images separately.
Multimodal Reference by Means of the Pronoun We and Hand Gestures in a Novel Corpus of Parliamentary Opening Debates
Costanza Navarretta
Costanza Navarretta
Political discourse has persuasion as its main goal and the identification of the referents of pronouns in it is of great importance. This paper presents a novel multimodal corpus of Danish parliamentary opening debates. It also describes a study of multimodal reference by means of the first-person plural pronoun “vi” (we) and the co-occurring hand gestures in a subset of the corpus. The data in the study consists of 219 speeches of two prime ministers from the opening debates in 2013, 2014, and 2021-2024. In the speeches, the prime ministers answer questions of parliament members from the opposition. The uses of the first-person plural pronoun in political speeches are particularly interesting since the pronouns can refer to different groups, such as the government, the parliament, the country, or a specific party and can be used by politicians to achieve consensus or distinguish their politics from that of others. The main hypothesis we want to investigate in the study is whether the pointing gestures vary in their trajectory depending on the intended referents. The results of our study confirm this hypothesis for the most frequent referent types and show how pointing hand gestures are used by the two prime ministers to help their audience individuating the correct referents of “vi”, and emphasise them. Our data also indicates that co-speech hand gestures are in some cases used to show the attitude of the speakers toward what they are saying.
Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
Lukas Arana | Julen Etxaniz | Ander Salaberria | Gorka Azkune
Lukas Arana | Julen Etxaniz | Ander Salaberria | Gorka Azkune
Current Multimodal Large Language Models exhibit very strong performance for several demanding tasks. While commercial MLLMs deliver acceptable performance in low-resource languages, comparable results remain unattained within the open science community. In this paper, we aim to develop a strong MLLM for a low-resource language, namely Basque. For that purpose, we develop our own training and evaluation image-text datasets, leveraging state-of-the-art translation systems. Using two different Large Language Models as backbones, the Llama-3.1-Instruct model and a Basque-adapted variant called Latxa, we explore several data mixtures for training, encompassing Basque and English languages for both multimodal and text-only data. Evaluating our MLLMs for close-ended and open-ended generation tasks, we show that: i) low ratios of Basque multimodal data (around 20%) are already enough to obtain solid results on Basque benchmarks, and ii) contrary to expected, a Basque instructed backbone LLM is not required to obtain a strong MLLM in Basque. Additionally, we specify the optimal data mixture strategy, the effects of multimodal data in text-only tasks, and analyze evaluation approaches for open-ended generation tasks. Our results pave the way to develop MLLMs for other low-resource languages by openly releasing our resources.
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal | Yuki Saito | Hiroya Takamura | Katsuhito Sudoh | Shinnosuke Takamichi | Graham Neubig | Florian Matthes | Tatsuya Ishigaki
Anum Afzal | Yuki Saito | Hiroya Takamura | Katsuhito Sudoh | Shinnosuke Takamichi | Graham Neubig | Florian Matthes | Tatsuya Ishigaki
Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential decisions: what to say and when to say it. While recent prompting-based approaches using multimodal large language models (MLLMs) have shown strong performance in content generation, they largely ignore the timing aspect. We investigate whether in-context prompting alone can support real-time commentary generation that is both semantically relevant and well-timed. We propose two prompting-based decoding strategies: 1) a fixed-interval approach, and 2) a novel dynamic interval-based decoding approach that adjusts the next prediction timing based on the estimated duration of the previous utterance. Both methods enable pause-aware generation without any fine-tuning. Experiments on Japanese and English datasets of racing and fighting games show that the dynamic interval-based decoding can generate commentary more closely aligned with human utterance timing and content using prompting alone. We release a multilingual benchmark dataset, trained models, and implementations to support future research on real-time video commentary generation.
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
Sara Ghaboura | Shubham Patle | Ketan More | Wafa Hamad Mohamed Alghallabi | Omkar Thawakar | Jorma Laaksonen | Hisham Cholakkal | Salman Khan | Rao Anwer
Sara Ghaboura | Shubham Patle | Ketan More | Wafa Hamad Mohamed Alghallabi | Omkar Thawakar | Jorma Laaksonen | Hisham Cholakkal | Salman Khan | Rao Anwer
As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most existing benchmarks remain focused on English, overlooking languages with rich linguistic and cultural depth such as Arabic. To address this gap, we introduce the Comprehensive Arabic Multimodal Reasoning Benchmark (ARB), the first benchmark designed to evaluate step-by-step reasoning in Arabic across both textual and visual modalities. ARB covers 11 diverse domains and over 40 subfields, including visual reasoning, optical character recognition, scientific analysis, and cultural interpretation. It comprises 2,219 multimodal samples paired with over 8K human-curated reasoning steps and corresponding actions, verified through a human-in-the-loop process. We evaluated 15 state-of-the-art open- and closed-source LMMs and found persistent challenges in coherence, faithfulness, and cultural grounding. ARB provides a structured framework for diagnosing multimodal reasoning in underrepresented languages, marking a critical step toward inclusive, transparent, and culturally aware AI systems. The benchmark, rubric, and evaluation suite are publicly available
Event Chronography in Multi-modal Data: The BME Method for Quantitative Analyses
Anaïs Claire Murat | Maria Koutsombogera | Carl Vogel
Anaïs Claire Murat | Maria Koutsombogera | Carl Vogel
Methods for investigating multi-modality in human interactions remain open to refinement. Although the annotation process has been facilitated by tools like Elan, synchronising exported cross-tier data for further quantitative analyses remains challenging. We present the BME method: a new approach to data alignment. The idea is straightforward: instead of comparing exact times of onsets, durations, etc., the BME method focuses on their organisation. First, the method describes every annotation by at least two events: its beginning (B) and end (E). Then, it aligns them in chronological order. Middles (M) are precipitated to track events from other tiers which might occur between Bs and Es. We explore three cases in which such an arrangement of multi-modal data can benefit the scientific community: first, in getting insights about the dynamics and dependencies between tiers, second, in contemplating event-based duration rather than time-based ones, and, third, in contributing cross-annotator agreement assessment methods.
CANVAS: A Multimodal Dataset of Chinese Textbook Images for Bias and Representation Analysis
Haotian Zhu | Kefan Yu | Min Li
Haotian Zhu | Kefan Yu | Min Li
Social biases in educational materials can subtly shape students’ perceptions of social roles and participation. However, most existing bias benchmarks for Chinese language models focus on text or isolated images, overlooking the multimodal scenes commonly found in educational textbooks. To address this gap, we introduce CANVAS (Chinese ANnotated Visual And Social scenes), a multimodal dataset constructed from Chinese elementary science textbooks and annotated across multiple social dimensions. CANVAS provides fine-grained labels for each depicted character’s demographics, social roles, interactions, and power-related attributes within visual scenes. The dataset is created using a semi-automated pipeline in which a vision–language model generates preliminary structured annotations that are subsequently verified and refined by human annotators. The current release focuses on the Grade 6 science subset and serves as an initial annotated version of the dataset. Using this subset, we present an illustrative case study demonstrating how scene-level and interactional annotations in CANVAS can be used to analyze gender representation in textbook images. By extending bias analysis to full educational scenes, CANVAS provides a new resource for studying representation and fairness in multimodal educational materials and supports future research in NLP, computer vision, and education.
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
Anna Deichler | Jim O’Regan | Fethiye Irmak Dogan | Anna Klezovich | Lubos Marcinek | Iolanda Leite | Jonas Beskow
Anna Deichler | Jim O’Regan | Fethiye Irmak Dogan | Anna Klezovich | Lubos Marcinek | Iolanda Leite | Jonas Beskow
Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous expressions in spontaneous, multi-turn dialogue. We address this gap by introducing MM-Conv—speak, point, look—a benchmark for referential communication in dynamic 3D environments, built from 6.7 hours of egocentric VR interaction with synchronized speech, motion, gaze, and 3D scene geometry. The benchmark includes over 4,200 manually verified referring expressions spanning full, partitive, and pronominal types, enabling systematic evaluation of multimodal reference resolution.
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
June Hyoung Kwon | Jungmin Yun | Youngbin Kim
June Hyoung Kwon | Jungmin Yun | Youngbin Kim
Large Vision-Language Models (LVLMs), trained on web-scale data, risk memorizing and regenerating copyrighted visual content like characters and logos, creating significant challenges. Machine unlearning offers a path to mitigate these risks by removing specific content post-training, but evaluating its effectiveness, especially in the complex multimodal setting of LVLMs, remains an open problem. Current evaluation methods often lack robustness or fail to capture the nuances of cross-modal concept erasure. To address this critical gap, we introduce the CoVUBench benchmark, the first framework specifically designed for evaluating copyright content unlearning in LVLMs. CoVUBench utilizes procedurally generated, legally safe synthetic data coupled with systematic visual variations—spanning compositional changes and diverse domain manifestations—to ensure realistic and robust evaluation of unlearning generalization. Our comprehensive, multimodal evaluation protocol assesses both forgetting efficacy from the copyright holder’s perspective and the preservation of general model utility from the deployer’s viewpoint. By rigorously measuring this crucial trade-off, CoVUBench provides a standardized tool to advance the development of responsible and effective unlearning methods for LVLMs.
DREAM: A Multicultural Multimodal Dataset Linking Dialogues and Realistic Image Sequences
Juan Mallo | Marcos Estecha-Garitagoitia | Ricardo Cordoba | Luis Fernando D’Haro
Juan Mallo | Marcos Estecha-Garitagoitia | Ricardo Cordoba | Luis Fernando D’Haro
An ongoing challenge in multimodal language research is creating and interpreting dialogues that preserve visual and cultural consistency across turns. We introduce DREAM (Dialogue to REAlistic Multicultural Image Sequences), a multicultural multimodal resource that ties dialogues grounded in explicit persona profiles to photorealistic, storyboard-like image sequences. Each of the 1,000 dialogues includes two rich persona profiles (structured traits plus descriptive language), two matching photorealistic portraits, and a collection of scene-level images depicting key dialogue moments. The pipeline integrates profile augmentation, culturally-sensitive prompt engineering, and turn selection to craft cohesive visual narratives, promoting character consistency across images. This is accomplished through a controlled generation process employing large language and image models. Beyond dialogue grounding, DREAM supports appearance-based demographic perception and culture-aware rendering: models can be evaluated on their ability to (i) perceive age, gender presentation, and broad ethnicity appearance clusters from profile portraits, and (ii) maintain these characteristics in dialogue scenes. We provide a unified JSON format integrating profiles, dialogue text, and visual turns, facilitating research on visually anchored dialogue understanding, consistency, and generation. A dual evaluation protocol combines human judgments (realism, coherence, consistency, and demographic perception) with automated portrait analysis via GPT-5. Ethical considerations, limitations, and recommended applications are discussed.
Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
Masayuki Kawarada | Tatsuya Ishigaki | Hiroya Takamura
Masayuki Kawarada | Tatsuya Ishigaki | Hiroya Takamura
Task interference, the performance degradation caused by task switches within a single conversation, has been studied exclusively in text-only settings despite the growing prevalence of multimodal dialogue systems. We introduce a benchmark for evaluating this phenomenon in multimodal LLMs, covering six tasks across text and vision with systematic variation of history-target along three axes: modality mismatch, reasoning mismatch, and answer format mismatch. Experiments on both open-weights and proprietary models reveal that task interference is highly directional: switching from text-only to image-based targets causes severe performance drops, while the reverse transition yields minimal degradation. Interference is further amplified when mismatches co-occur across multiple dimensions, and is driven most strongly by modality differences, followed by answer format, while reasoning requirement shifts cause minimal degradation.
Can Video LLMs See Through Illusions? Video-Illusion QA Benchmark Dataset
Souto Ohira | Tosho Hirasawa | Mamoru Komachi
Souto Ohira | Tosho Hirasawa | Mamoru Komachi
Recent advances in multimodal learning have sparked growing interest in understanding how large vision-language models interpret optical illusions. While the behavior of image LLMs—which handle one image and text but not video input—on visual illusion images has been actively explored, research on their video counterparts remains limited. Video LLMs, which process sequential frames, are gaining prominence in areas such as robotics and autonomous driving. Understanding how they handle visual illusions over time is crucial for safety and may also reveal their potential as computational models of human cognition. To address this gap, we present the Video-Illusion QA Benchmark (VILQA), a novel video question answering (QA) benchmark mainly composed of carefully curated illusion videos that exhibit temporally driven perceptual phenomena. To the best of our knowledge, VILQA is the largest and most comprehensive benchmark for temporally-driven visual illusions. We evaluate several video LLMs on this benchmark from multiple perspectives. Some models were able to perceive visual illusions in a way similar to the general human experience and demonstrated an ability to resist illusions even more effectively than humans. The constructed dataset is available at https://github.com/SDS-NLP/VILQA.
To Skip, to Swap or to Not Swap? Identifying Step Transition Types in Instructional Manuals
Hsiu-Yu Yang | Michael Roth | Andreas Bulling | Carina Silberer
Hsiu-Yu Yang | Michael Roth | Andreas Bulling | Carina Silberer
Large language models (LLMs) are increasingly used as procedural planners that provide guidance across applications. However, in human-assistive scenarios where the environment and users’ knowledge constantly change, their ability to detect various step types for generating alternative plans is underexplored. To address this gap, we introduce a novel evaluation task and dataset to assess if models can identify steps that are sequential, interchangeable, and optional in textual instructions across five domains in a step-by-step manner. We compare seven LLM families from both open-source and proprietary spaces across varying sizes to a visually-informed baseline based on procedural knowledge graphs (PKG). Our results suggest that LLMs encode procedural knowledge, enabling them to identify step types with increasing effectiveness as training parameters and data size grow. However, all LLMs exhibit inconsistencies in reasoning on the mutual exclusivity of interchangeable and sequential step pairs. In contrast, the symbolic PKG baseline demonstrates stronger consistency in this aspect. Comprehensive analyses furthermore uncover limitations in LLMs’ procedural reasoning abilities.
Fruitcakes and Cupcakes Emerging from Noise: The ComposiGen Dataset of Compounds and Their Compositionality
Jule Godbersen | Sinan Cem Kurtyigit | Emma Raimundo Schulz | Tonmoy Rakshit | Diego Frassinelli | Sabine Schulte im Walde | Carina Silberer
Jule Godbersen | Sinan Cem Kurtyigit | Emma Raimundo Schulz | Tonmoy Rakshit | Diego Frassinelli | Sabine Schulte im Walde | Carina Silberer
Compounds are a complex linguistic phenomenon, as variation in their degree of compositionality often makes their interpretation non-straightforward. We consider the task of visual-linguistic compositionality prediction for English noun-noun compounds, i.e., predicting the degrees to which a compound’s meaning is predictable from its constituents. We introduce a new dataset, ComposiGen, which provides constituent-specific human-elicited compositionality ratings for compounds of different concreteness categories, and includes generated visual representations for both compounds and their constituents. To enable controlled comparisons, we structure ComposiGen such that head constituents are shared across multiple compounds (e.g., wedding cake, cup cake). We suggest a novel parameter-based approach leveraging constituent-to-compound image transformations to predict different degrees of visual constituent contributions to compound meaning. While our novel approach requires further exploration for validation, our overall results show that the generated images, in particular in combination with text, provide valuable information, and that simple late fusion outperforms multimodal transformers. Taken together, our findings highlight a promising avenue for future research on more efficient multimodal models for compositionality prediction. Our novel dataset offers a rich resource for future in-depth research, including the exploration of visual, constituent-based compound formation.
Large language models (LLMs) excel at modeling relationships between strings in natural language and have shown promise in extending to other symbolic domains like coding or mathematics. However, the extent to which they implicitly model symbolic music remains underexplored. This paper investigates how LLMs represent musical concepts by generating symbolic music data from textual prompts describing combinations of genres and styles, and evaluating their utility through recognition and generation tasks. We produce a dataset of LLM-generated MIDI files without relying on explicit musical training. We then train neural networks entirely on this LLM-generated MIDI dataset and perform genre and style classification as well as melody completion, benchmarking their performance against established models. Our results demonstrate that LLMs can infer rudimentary musical structures and temporal relationships from text, highlighting both their potential to implicitly encode musical patterns and their limitations due to a lack of explicit musical context, shedding light on their generative capabilities for symbolic music.
Entity Image and Mixed-Modal Image Retrieval Datasets
Cristian-Ioan Blaga | Paul Suganthan G C | Sahil Dua | Krishna Srinivasan | Enrique Alfonseca | Peter Dornbach | Tom Duerig | Imed Zitouni | Zhe Dong
Cristian-Ioan Blaga | Paul Suganthan G C | Sahil Dua | Krishna Srinivasan | Enrique Alfonseca | Peter Dornbach | Tom Duerig | Imed Zitouni | Zhe Dong
Despite advances in multimodal learning, challenging benchmarks for mixed-modal image retrieval that combines visual and textual information are lacking. This paper introduces a novel benchmark to rigorously evaluate image retrieval that demands deep cross-modal contextual understanding. We present two new datasets: the Entity Image Dataset (EI), providing canonical images for Wikipedia entities, and the Mixed-Modal Image Retrieval Dataset (MMIR), derived from the WIT dataset. The MMIR benchmark features two challenging query types requiring models to ground textual descriptions in the context of provided visual entities: single entity-image queries (one entity image with descriptive text) and multi-entity-image queries (multiple entity images with relational text). We empirically validate the benchmark’s utility as both a training corpus and an evaluation set for mixed-modal retrieval. The quality of both datasets is further affirmed through crowd-sourced human annotations.
Generating Sign Language Poses from HamNoSys and Natural Language Descriptions
Santiago Máximo | Luis Chiruzzo
Santiago Máximo | Luis Chiruzzo
One of the steps involved in the process of sign language generation is generating a sequence of poses that represent the signs. This paper presents a method for using textual information to improve the translation of signs in HamNoSys format into sequences of poses. The method comprises a description generator that translates HamNoSys into a textual description, an LLM fine-tuned to the task of predicting a pose sequence from a HamNoSys description, and a VQ-VAE network that encodes and decodes pose sequences as a list of discrete symbols. Our experiments found that even using simple dictionary descriptions of HamNoSys, it is possible to improve the predictions of pose sequences by leveraging the information from a pretrained LLM.
We study the discriminative ability of vision-language models (VLMs). This ability refers to processing information by distinguishing key details from unnecessary or redundant parts to achieve specific goals. It is vital for the practical use of VLMs in applications like visual chatbots. Whereas recent VLMs have shown decent performance on various multimodal capabilities, their discriminative ability has not been thoroughly explored to date. To this end, we construct DiscriBench to evaluate the discriminability of VLMs in various daily life activities. We carefully design the dataset to require distinguishing information in both vision and language modalities, and semi-manually craft questions in English and Japanese, making them solvable without relying on external knowledge or expertise. Experimental results demonstrate a large performance gap (14.0 to 69.3 points) between humans and existing VLMs in discriminability, where humans can solve the task with an accuracy of 90% or higher. By reducing the difficulty of discriminability, our ablation studies elucidate that vision encoders cannot distinguish visual details well, given generally similar but partially different images. Besides, we observe that VLMs show inconsistent inference between modalities. We will publish DiscriBench (1,200 samples) to foster research in this direction.
Seeing the Other Side: Diagnostic Tasks for Viewpoint Reasoning in Vision–Language Models
Makoto Takenaka | Hitomi Yanaka
Makoto Takenaka | Hitomi Yanaka
Humans can integrate multiple visual perspectives and infer how an object appears from unseen sides. This study investigates whether Large Vision Language Models (LVLMs) exhibit a comparable ability for reference-grounded spatial reasoning. We propose two diagnostic tasks: Opposite-Side Reasoning, which determines whether two images show the same object from opposite viewpoints, and Viewpoint Identification, which predicts the viewpoint of a target image using a reference image and its label. An additional condition, Viewpoint Identification (no-ref), removes reference information to reveal cases solvable without it, distinguishing genuine reasoning from bias-driven shortcuts. Our evaluation shows that both open and proprietary LVLMs fall far short of human performance. Even state-of-the-art proprietary LVLMs with relatively high accuracy retain many correct answers when reference information is removed, suggesting that their success often relies on linguistic or dataset-driven priors rather than genuine reference-based reasoning. These findings indicate that current LVLMs have not yet achieved consistent, reference-grounded spatial reasoning. Our datasets in this work will be released on the Hugging Face Hub to support future research on multimodal viewpoint reasoning and spatial understanding.
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
Masanari Oi | Masahiro Kaneko | Naoaki Okazaki | Nakamasa Inoue
Masanari Oi | Masahiro Kaneko | Naoaki Okazaki | Nakamasa Inoue
Vision-language models (VLMs) have shown impressive abilities across a range of multi-modal tasks. However, existing metrics for evaluating the quality of text generated by VLMs typically focus on an overall evaluation for a specific task, such as image captioning. While the overall evaluation is essential for any task, the criteria prioritized can differ depending on the task, making it challenging for current metrics to adapt to multi-task scenarios. To address this limitation, we propose HarmonicEval, a reference-free comprehensive evaluation metric that aggregates criterion-wise scores to produce the overall score in a bottom-up manner. Furthermore, to assess the generalizability of automatic evaluation metrics in multi-task scenarios, we construct the Multi-task Multi-criteria Human Evaluation (MMHE) benchmark, which comprises 18,000 expert human judgments across four multi-modal tasks. Our experiments demonstrate that HarmonicEval achieves higher correlations with human judgments than conventional metrics while providing numerical scores for each criterion. Our code and data will be available publicly.
Challenges in Image-Caption Association in Portuguese: Evaluating the CLIP Model on the FM30K Dataset
Vitória Colonetti Benedet | Gustavo Lopes Tamiosso | Rafael Oleques Nunes | Dennis Giovani Balreira
Vitória Colonetti Benedet | Gustavo Lopes Tamiosso | Rafael Oleques Nunes | Dennis Giovani Balreira
In recent decades, multimodal models such as CLIP have achieved significant advances in associating images and texts. However, most of these advances stem from models trained almost exclusively in English, which limits their effectiveness in other languages. This challenge is particularly relevant for Brazilian Portuguese, a language that still lacks dedicated multimodal resources and relies predominantly on automatic translations. This work investigates the performance of CLIP-based multimodal models in the task of associating images and descriptions written in Brazilian Portuguese. The analysis begins with a zero-shot scenario, in which different CLIP variants are directly evaluated on the FM30k dataset, composed of images and captions originally written in Portuguese. An additional experiment with automatic translations is also conducted to examine the impact of language on cross-modal retrieval tasks. Subsequently, fine-tuning is performed on the textual encoder of the ViT-B/32 model, keeping the visual encoder frozen, with the goal of adapting the model to the target language. The results show that models originally trained in English perform worse in Portuguese, while linguistically adapted variants, either multilingual or Portuguese-specific, achieve superior performance. The proposed fine-tuning approach was able to reduce this performance gap, leading to notable improvements. In the image-to-text scenario, the model achieved an absolute increase of 27.65 percentage points in the Accuracy@1 metric, representing a 209% relative gain over the original CLIP ViT-B/32. In the text-to-image scenario, the gain was 15.47 percentage points, amounting to an even higher 385% relative improvement, contributing to a more balanced association between images and captions.
A Large-Scale Instruction-Tuning Dataset and Models for Slovenian Vision-Language Tasks
Matej Martinc | Domen Vreš
Matej Martinc | Domen Vreš
Vision-language models (VLMs) represent a significant leap forward in artificial intelligence, yet their development has been predominantly focused on English, creating a digital divide for speakers of less-resourced languages. This paper addresses this gap by introducing the first large-scale, general instruction-tuning dataset for the less-resourced Slovenian language. Comprising over one million text-image pairs, the dataset was constructed through a multi-pronged approach: automatic curation from Slovenian news media and Wikipedia, and machine translation of the English LLaVA-665k dataset. To demonstrate the dataset’s efficacy, we fine-tuned two pre-trained, multilingual Gemma-3 models (4B and 12B parameters) on this new resource. Our evaluation, conducted on a new manually curated test set, reveals that the fine-tuned models named SVILA (Slovenian Vision Language Assistant) exhibit substantial performance gains on a variety of vision question answering, visual grounding, and optical character recognition tasks when compared to their baseline counterparts. This establishes our methodology as an effective blueprint for enhancing VLM capabilities in other less-resourced languages. The dataset is publicly available in the Slovenian language resource repository CLARIN.SI (http://hdl.handle.net/11356/2050) and both fine-tuned models are published on the Hugging Face platform (https://huggingface.co/GaMS-Beta/SVILA-1-12B and https://huggingface.co/GaMS-Beta/SVILA-1-4B).
A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
Dilara Torunoğlu-Selamet | Doğukan Arslan | Rodrigo Wilkens | Wei He | Doruk Eryiğit | Thomas Pickard | Adriana S. Pagano | Aline Villavicencio | Gülşen Eryiğit | Ágnes Abuczki | Aida Cardoso | Alesia Lazarenka | Dina Almassova | Amália Mendes | Anna Kanellopoulou | Antoni Brosa-Rodriguez | Baiba Valkovska | Beata Wojtowicz | Bolette Pedersen | Carlos Manuel Hidalgo-Ternero | Chaya Liebeskind | Danka Jokić | Diego Alves | Eleni Triantafyllidi | Erik Velldal | Fred Philippy | Giedre Valunaite Oleskeviciene | Ieva Rizgeliene | Inguna Skadina | Irina Lobzhanidze | Isabell Stinessen Haugen | Jauza Akbar Krito | Jelena M. Marković | Johanna Monti | Josue Alejandro Sauca | Kaja Dobrovoljc Zor | Kingsley O. Ugwuanyi | Laura Rituma | Lilja Øvrelid | Maha Tufail Agro | Manzura Abjalova | Maria Chatzigrigoriou | María del Mar Sánchez Ramos | Marija Pendevska | Masoumeh Seyyedrezaei | Mehrnoush Shamsfard | Momina Ahsan | Muhammad Ahsan Riaz Khan | Nathalie Carmen Hau Norman | Nilay Erdem Ayyıldız | Nina Hosseini-Kivanani | Noémi Ligeti-Nagy | Numaan Naeem | Olha Kanishcheva | Olha Yatsyshyna | Daniil Orel | Petra Giommarelli | Petya Osenova | Radovan Garabik | Regina E. Semou | Rozane Rebechi | Salsabila Zahirah Pranida | Samia Touileb | Sanni Nimb | Sarfraz Ahmad | Sarvinoz Sharipova | Shahar Golan | Shaoxiong Ji | Sopuruchi Christian Aboh | Srdjan Sucur | Stella Markantonatou | Sussi Olsen | Vahide Tajalli | Veronika Lipp | Voula Giouli | Yelda Yeşildal Eraydın | Zahra Saaberi | Zhuohan Xie
Dilara Torunoğlu-Selamet | Doğukan Arslan | Rodrigo Wilkens | Wei He | Doruk Eryiğit | Thomas Pickard | Adriana S. Pagano | Aline Villavicencio | Gülşen Eryiğit | Ágnes Abuczki | Aida Cardoso | Alesia Lazarenka | Dina Almassova | Amália Mendes | Anna Kanellopoulou | Antoni Brosa-Rodriguez | Baiba Valkovska | Beata Wojtowicz | Bolette Pedersen | Carlos Manuel Hidalgo-Ternero | Chaya Liebeskind | Danka Jokić | Diego Alves | Eleni Triantafyllidi | Erik Velldal | Fred Philippy | Giedre Valunaite Oleskeviciene | Ieva Rizgeliene | Inguna Skadina | Irina Lobzhanidze | Isabell Stinessen Haugen | Jauza Akbar Krito | Jelena M. Marković | Johanna Monti | Josue Alejandro Sauca | Kaja Dobrovoljc Zor | Kingsley O. Ugwuanyi | Laura Rituma | Lilja Øvrelid | Maha Tufail Agro | Manzura Abjalova | Maria Chatzigrigoriou | María del Mar Sánchez Ramos | Marija Pendevska | Masoumeh Seyyedrezaei | Mehrnoush Shamsfard | Momina Ahsan | Muhammad Ahsan Riaz Khan | Nathalie Carmen Hau Norman | Nilay Erdem Ayyıldız | Nina Hosseini-Kivanani | Noémi Ligeti-Nagy | Numaan Naeem | Olha Kanishcheva | Olha Yatsyshyna | Daniil Orel | Petra Giommarelli | Petya Osenova | Radovan Garabik | Regina E. Semou | Rozane Rebechi | Salsabila Zahirah Pranida | Samia Touileb | Sanni Nimb | Sarfraz Ahmad | Sarvinoz Sharipova | Shahar Golan | Shaoxiong Ji | Sopuruchi Christian Aboh | Srdjan Sucur | Stella Markantonatou | Sussi Olsen | Vahide Tajalli | Veronika Lipp | Voula Giouli | Yelda Yeşildal Eraydın | Zahra Saaberi | Zhuohan Xie
Potentially idiomatic expressions (PIEs) carry meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows evaluation of language model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision–Language Models
Shiho Matta | Lis Kanashiro Pereira | Peitao Han | Shigeru Kitazawa | Fei Cheng
Shiho Matta | Lis Kanashiro Pereira | Peitao Han | Shigeru Kitazawa | Fei Cheng
Modern vision–language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe this gap with a deceptively simple but revealing challenge: judging the arrow of time (AoT)—whether a short clip is played forward or backward. We introduce AoT-PsyPhyBENCH, a psychophysically validated benchmark that tests whether VLMs can infer temporal direction in natural videos using the same stimuli and behavioral baselines established for humans. Our comprehensive evaluation of open-weight and proprietary, reasoning and non-reasoning VLMs reveals that most models perform near chance, and even the best model lags far behind human accuracy on physically irreversible processes (e.g., free fall, diffusion/explosion) and causal manual actions (division/addition) that humans recognize almost instantly. These results highlight a fundamental gap in current multimodal systems: while they capture rich visual–semantic correlations, they lack the inductive biases required for temporal continuity and causal understanding. We release the code and data for AoT-PsyPhyBENCH to encourage further progress in the physical and temporal reasoning capabilities of VLMs.
I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
Shijia Zhou | Saif M. Mohammad | Barbara Plank | Diego Frassinelli
Shijia Zhou | Saif M. Mohammad | Barbara Plank | Diego Frassinelli
Internet memes represent a popular form of multimodal online communication and often use figurative elements to convey layered meaning through the combination of text and images. However, it remains largely unclear how multimodal large language models (MLLMs) combine and interpret visual and textual information to identify figurative meaning in memes. To address this gap, we evaluate eight state-of-the-art generative MLLMs across three datasets on their ability to detect and explain six types of figurative meaning. In addition, we conduct a human evaluation of the explanations generated by these MLLMs, assessing whether the provided reasoning supports the predicted label and whether it remains faithful to the original meme content. Our findings indicate that all models exhibit a strong bias to associate a meme with figurative meaning, even when no such meaning is present. Qualitative analysis further shows that correct predictions are not always accompanied by faithful explanations.
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
Toshiki Katsube | Fukuhara Taiga | Kenichiro Ando | Yusuke Mukuta | Kohei Uehara | Tatsuya Harada
Toshiki Katsube | Fukuhara Taiga | Kenichiro Ando | Yusuke Mukuta | Kohei Uehara | Tatsuya Harada
Vision-and-Language (V&L) models depend on large-scale, high-quality datasets, yet most resources are English-centric, and existing Japanese V&L datasets face a fundamental trade-off: manually annotated corpora offer quality but limited scale, translated datasets introduce unnatural phrasing and cultural bias, and web-crawled collections achieve scale but suffer from noise and poor grounding. To resolve this trade-off, we propose DEJIMA, a novel pipeline whose key idea is detection-guided LLM refinement: object detection first extracts visually verifiable evidence (labels and bounding boxes), then an LLM generates or refines Japanese text conditioned on this evidence, ensuring both factual grounding and linguistic naturalness without costly human annotation. Using this pipeline, we build two resources: an image–caption dataset (DEJIMA-Cap) and a VQA dataset (DEJIMA-VQA), each containing approximately 3.88M image–text pairs—over 20 times larger than existing Japanese V&L datasets. Human evaluations demonstrate that DEJIMA achieves substantially higher Japaneseness and linguistic naturalness than translation- or annotation-based baselines, while maintaining factual correctness comparable to human-annotated corpora. Models trained on DEJIMA show consistent improvements across multiple Japanese multimodal benchmarks, confirming that culturally grounded, large-scale resources play a key role in enhancing model performance. All pipeline components are commercially licensed, and we publicly release the dataset and metadata to support further research and applications. Our project page is available at https://mil-tokyo.github.io/DEJIMA-dataset/.
Vision-language models (VLMs) often struggle to interpret spatial referring expressions that require relational reasoning rather than reliance on surface-level cues. These models frequently identify referents through explicit visual attributes such as color or shape, rather than understanding spatial relationships (e.g., "to the left of the red cube”). To systematically analyze these limitations, we introduce CLEVR-3D-DeRef, a synthetic and extensible benchmark dataset modeled after CLEVR-Ref+, designed to evaluate spatial reasoning in multi-modal systems. CLEVR-3D-DeRef extends the original framework by incorporating depth information for 3D spatial reasoning, introducing de-identified context-dependent referring expressions that require relational inference to disambiguate referent objects, and expanding the range of spatial relations beyond the original four. We further extend our dataset by producing expressions with and without ordinal language and diversifying the language and structure of expressions while preserving meaning.
Bridging Text-to-Sign Translation via Codebook-Oriented Pretraining
Ninlawat Phuangchoke | Chantri Polprasert
Ninlawat Phuangchoke | Chantri Polprasert
Sign Language Production (SLP), the automatic translation from spoken to sign languages, faces several challenges due to the intricate mapping between linguistic semantics and the spatial–temporal motion domain. Existing SLP methods employing a transformer model with a Vector Quantization (VQ) method exhibit poor translation performance due to weak semantic alignment between the codebook and the text representation. In this work, we propose a novel text-to-sign translation based on model pretraining, which enhances semantic alignment by inheriting codebook-oriented prior knowledge from masked self-supervised models. Our approach involves two stages: (i) transforming sign language into discrete values by employing VQ with masked self-attention learning to create pre-tasks that bridge the semantic gap between text and codebook representations, (ii) constructing an end-to-end architecture with an encoder-decoder-like structure that inherits the parameters of the model from the first stage. The integration of these designs forms a robust sign language representation and significantly improves the translation model, which surpass prior baselines.
A Resource and Evaluation Method for Phonological Continuity in Japanese Sign Language
Jundai Inoue | Daisuke Hara | Makoto Miwa
Jundai Inoue | Daisuke Hara | Makoto Miwa
Computational models for sign language processing often represent phonological components as categories. This approach, however, does not adequately capture the continuous nature of sign articulation, obscuring nuanced phonetic variation. Furthermore, the field has lacked resources and standardized methods to evaluate a model’s ability to represent this continuity. In this work, we address these limitations. First, we introduce the JSL Ordered Triplet Dataset, a new manually-annotated resource designed to benchmark the modeling of gradual phonological progressions in Japanese Sign Language. Second, we propose a learning framework that reframes the task from classification to ranking, using Positive-Unlabeled (PU) learning to optimize the Area Under the ROC Curve (AUC). Our intrinsic evaluation on the new dataset shows that the learned continuous embeddings significantly outperform a cross-entropy baseline in ordering intermediate forms, improving the average accuracy on the continuity ranking task across phonological components from 81.52% to 91.71%. These embeddings also maintain strong discriminative power for standard component classification. This work provides the community with a valuable resource and a method for learning and evaluating more linguistically-grounded representations of sign language.
Sentiment Analysis of German Sign Language Fairy Tales
Fabrizio Nunnari | Siddhant Jain | Patrick Gebhard
Fabrizio Nunnari | Siddhant Jain | Patrick Gebhard
We present a dataset and a model for sentiment analysis of German sign language (DGS) fairy tales. First, we perform sentiment analysis for three levels of valence (negative, neutral, positive) on German fairy tales text segments using four large language models (LLMs) and majority voting, reaching an inter-annotator agreement of 0.781 Krippendorff’s alpha. Second, we extract face and body motion features from each corresponding DGS video segment using MediaPipe. Finally, we train an explainable model (based on XGBoost) to predict negative, neutral or positive sentiment from video features. Results show an average balanced accuracy of 0.631. A thorough analysis of the most important features reveal that, in addition to eyebrows and mouth motion on the face, also the motion of hips, elbows, and shoulders considerably contribute in the discrimination of the conveyed sentiment, indicating an equal importance of face and body for sentiment communication in sign language.
A Critical Study of Automatic Evaluation in Sign Language Translation
Shakib Yazdani | Yasser HAMIDULLAH | Cristina España-Bonet | Eleftherios Avramidis | Josef van Genabith
Shakib Yazdani | Yasser HAMIDULLAH | Cristina España-Bonet | Eleftherios Avramidis | Josef van Genabith
Automatic evaluation metrics are crucial for advancing sign language translation (SLT). Current SLT evaluation metrics, such as BLEU and ROUGE, are only text-based, and it remains unclear to what extent text-based metrics can reliably capture the quality of SLT outputs. To address this gap, we investigate the limitations of text-based SLT evaluation metrics by analyzing six metrics, including BLEU, chrF, and ROUGE, as well as BLEURT on the one hand, and large language model (LLM)-based evaluators such as G-Eval and GEMBA zero-shot direct assessment on the other hand. Specifically, we assess the consistency and robustness of these metrics under three controlled conditions: paraphrasing, hallucinations in model outputs, and variations in sentence length. Our analysis highlights the limitations of lexical overlap metrics and demonstrates that while LLM-based evaluators better capture semantic equivalence often missed by conventional metrics, they can also exhibit bias toward LLM-paraphrased translations. Moreover, although all metrics are able to detect hallucinations, BLEU tends to be overly sensitive, whereas BLEURT and LLM-based evaluators are comparatively lenient toward subtle cases. This motivates the need for multimodal evaluation frameworks that extend beyond text-based metrics to enable a more holistic assessment of SLT outputs.
How Much Data Is Enough Data? A New Motion Capture Corpus for Probabilistic Sign Language Generation
Anna Klezovich | Johanna Mesch | Gustav Eje Henter | Jonas Beskow
Anna Klezovich | Johanna Mesch | Gustav Eje Henter | Jonas Beskow
We present a new 4.1 hours long high-quality motion capture sign language dataset for Swedish Sign Language — STS Mocap v1. The dataset consists of high quality multimodal data: body tracked with markers, fingers tracked with Manus Quantum Metagloves, face tracked with iPhone LiveLink app in MetaHuman Animator mode, and corresponding textual sentence translation to spoken Swedish. With the help of this dataset, we show that four hours of motion capture data is enough for generative modeling of sign language conditioned on 2D pose. In comparison, training the same flow-matching model on only 30 minutes of this data, which is a common size for sign language motion capture datasets, shows a significant degradation in the quality of the synthesized data.
Decomposing Sign Language Movements: A Multi-Band Visualization Method for Articulatory Analysis
Antonio F. G. Sevilla | José María Lahoz-Bengoechea
Antonio F. G. Sevilla | José María Lahoz-Bengoechea
Understanding the structure of sign language movements requires methods that can isolate and analyze the hierarchical and simultaneous nature of sign articulation. We present a method for tracking and visualizing sign language movements that progressively isolates dependent movements within the articulatory chain: hand rotation from arm displacement and finger movement from hand movement. Using MediaPipe hand tracking on ordinary 2D video, we decompose motion into separate gestural components and compute velocity and direction for each articulator. We present these movement channels in a time-aligned multi-band visualization that reveals temporal structure, bimanual synchronization patterns, and the coordination of different articulatory components. An interactive web-based viewer synchronizes the visualization with video, enabling researchers to efficiently explore movement patterns and their relationship to signing. We demonstrate the method with examples from isolated signs and continuous signing, showing how it reveals patterns that are difficult to observe in raw video, including bimanual coordination, internal movements, and the distinction between linguistic and non-linguistic segments. This approach provides accessible tools for empirical investigation of rhythmic and prosodic patterns in sign languages.
Implicit Bias in Peer Review: Through the Lens of Language Abstraction
Xulang Zhang | Rui Mao | Erik Cambria
Xulang Zhang | Rui Mao | Erik Cambria
Peer review is essential for the scholarly publishing process. However, its credibility is increasingly brought to questions. Bias is one of the aspects worthy of investigation. Existing research mostly focuses on predefined, explicit bias types, which are insufficient for analyzing the myriad of implicit biases in peer review. Thus, we proposed to study the bias in peer review through the lens of language abstraction, informed by the cognitive theories which suggest that frequency of abstraction in descriptions plays a latent yet important role in bias transmission. Hence, we trained a model to assess the abstraction level of text, and applied it to a review dataset to examine the connection between abstraction and the implicit biases in peer reviews. Results show that there are indeed observable quantitative differences in the abstraction use of reviews recommending to reject versus recommending to accept. Furthermore, reviews for the rejected papers tend to be more abstract than ones for the accepted papers, indicating possible transmission of implicit bias. To the best of our knowledge, our study is the first to study generalized Linguistic Intergroup Bias in the academic text domain.
The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer’s Disease
Franziska Braun | Christopher Witzl | Florian Hönig | Elmar Nöth | Tobias Bocklet | Korbinian Riedhammer
Franziska Braun | Christopher Witzl | Florian Hönig | Elmar Nöth | Tobias Bocklet | Korbinian Riedhammer
Early and accessible detection of Alzheimer’s disease (AD) remains a major challenge, as current diagnostic methods often rely on costly and invasive biomarkers. Speech and language analysis has emerged as a promising non-invasive and scalable approach to detecting cognitive impairment, but research in this area is hindered by the lack of publicly available datasets, especially for languages other than English. This paper introduces the PARLO Dementia Corpus (PDC), a new multi-center, clinically validated German resource for AD collected across nine academic memory clinics in Germany. The dataset comprises speech recordings from individuals with AD-related mild cognitive impairment and mild to moderate dementia, as well as cognitively healthy controls. Speech was elicited using a standardized test battery of eight neuropsychological tasks, including confrontation naming, verbal fluency, word repetition, picture description, story reading, and recall tasks. In addition to audio recordings, the dataset includes manually verified transcriptions and detailed demographic, clinical, and biomarker metadata. Baseline experiments on ASR benchmarking, automated test evaluation, and LLM-based classification illustrate the feasibility of automatic, speech-based cognitive assessment and highlight the diagnostic value of recall-driven speech production. The PDC thus establishes the first publicly available German benchmark for multi-modal and cross-lingual research on neurodegenerative diseases.
Lexical and Discourse Semantics in a Reading-time Corpus of English
Jakub Dotlacil | Laia Colina Fortuny | Li Kloostra | Johan Bos
Jakub Dotlacil | Laia Colina Fortuny | Li Kloostra | Johan Bos
We present a novel language resource that combines a reading-time corpus, constructed in psycholinguistics, with rich lexical, compositional, and discourse meaning representation annotations. While existing psycholinguistic corpora typically provide morphological and syntactic annotations, no comparable corpora with comprehensive semantic information have been made available until now. We enriched the UCL corpus (361 sentences of self-paced reading, eye-tracking, and EEG data) with annotations in the style of the Parallel Meaning Bank (PMB) project, including WordNet synsets, VerbNet thematic roles, Combinatory Categorial Grammar (CCG) parses, and Discourse Representation Theory (DRT) structures. We demonstrate the utility of this resource through two case studies examining (1) encoding interference effects due to gender similarity and (2) integration costs in semantic role assignment. Both studies reveal processing patterns consistent with established psycholinguistic theories and/or previous findings. This resource fills a significant gap in psycholinguistic research, enabling the evaluation of semantic processing theories on naturalistic corpus data and extending the existing pool of annotated reading-time corpora. It should be useful to psycholinguists, as well as to cognitive scientists interested in language processing.
Semantic Capacity in Language Learners and LLMs: A Case Study of Quantifier Scope
Shaohua Fang | Yue Li | Yan Cong
Shaohua Fang | Yue Li | Yan Cong
This study investigates the semantic capacity of large language models (LLMs) through the lens of quantifier scope interpretation. Sentences containing multiple quantifiers often give rise to interpretive ambiguities, and the range of available readings can vary across languages. Adopting a cross-linguistic perspective, we examine how LLMs interpret quantifier scope in English and Chinese, using model-generated probabilities to assess the relative likelihood of competing interpretations. Human similarity (HS) scores were used to quantify the extent to which LLMs emulate human performance across language groups. Results reveal that most LLMs prefer the surface scope interpretations, aligning with human tendencies, while only some differentiate between English and Chinese in the inverse scope preferences, reflecting human-similar patterns. HS scores highlight variability in LLMs’ approximation of human behavior, but their overall potential to align with humans is notable. Linguistic identity, instantiated through monolingual and bilingual personas of English or Chinese, was found to influence LLM behavior. Differences in model architecture, scale, and particularly models’ pre-training data language background, significantly influence how closely LLMs approximate human quantifier scope interpretations.
How Long Does a Quick Kiss Take? Studying Event Duration of Light Verb Constructions Using Explicit Word Embeddings
Lin de Huybrecht | Geraint A. Wiggins
Lin de Huybrecht | Geraint A. Wiggins
Psycholinguistic research indicates that choosing one syntactic construction over another to describe an event can influence its perceived duration: Light Verb Constructions (LVCs) such as punctive events in count syntax (to give a kiss) and durative events in mass syntax (to do research) are perceived as taking less time than their Full Verb Constructions (FVCs; to kiss and to research). Similar computational results were achieved using BERT embeddings to semantically project events onto a one-dimensional Duration scale. We reproduce and further develop this experiment with explicit word embeddings from our own co-occurrence count-based vector space. By semantically projecting 158 LVC-FVC pairs onto our Duration scale, we find that LVCs are modelled as significantly shorter than FVCs. However, we do not find an overall statistically significant difference in duration between sentences containing the target LVCs and FVCs. We demonstrate that semantic properties observed in human experiments and in BERT embeddings can also be modelled using explicit word embeddings, which have the advantage of being fully transparent. However, using transcripts from spoken conversations can be challenging when studying a specific construction: optimising the extraction of sentences containing the target expressions and composition of their meanings are to be addressed in future work.
Disambiguation of Emotion Annotations by Contextualizing Events in Plausible Narratives
Johannes Schaefer | Roman Klinger
Johannes Schaefer | Roman Klinger
Ambiguity in emotion analysis stems both from potentially missing information and the subjectivity of interpreting a text. The latter did receive substantial attention, but can we fill missing information to resolve ambiguity? We address this question by developing a method to automatically generate reasonable contexts for an otherwise ambiguous classification instance. These generated contexts may act as illustrations of potential interpretations by different readers, as they can fill missing information with their individual world knowledge. This task to generate plausible narratives is a challenging one: We combine techniques from short story generation to achieve coherent narratives. The resulting dataset of Emotional BackStories, EBS, allows for the first comprehensive and systematic examination of contextualized emotion analysis. We conduct automatic and human annotation and find that the generated contextual narratives do indeed clarify the interpretation of specific emotions. Particularly relief and sadness benefit from our approach, while joy does not require the additional context we provide.
Identifying Contexts of Distress in College Students’ Reddit Posts: A Comparative Study of Classical NLP and Large Language Models
Carine Graff | Nikhil Krishnaswamy
Carine Graff | Nikhil Krishnaswamy
Mental health is a salient and growing societal concern among college students. Social media platforms such as Reddit offer a rich source of data regarding how students talk about their mental health, and NLP tools may potentially assist in identifying when a student is struggling. In this paper, we investigate how different NLP tools can be used to extract context surrounding college students expressions of distress. We construct a novel dataset from Reddit posts (College Distress on Reddit, or CDR), and examine the “classical NLP pipeline”, and modern generative LLMs on this data. Our dataset exploration is conducted in parallel with, and contrasted against the Dreaddit dataset to examine cross-domain variation. Results show that standard or “classical” NLP tools extract a limited number of concrete entities, whereas generative models can infer more nuanced causes. However, LLMs struggle with knowledge extraction in specific content areas. Our work shows how important it is to be wary of LLMs, especially in mental health contexts.
TiC-MuFormer: Time-Aware Caption-Integrated Multimodal Transformers for User-Level Mental Health Modeling
Georgios Tsoumplekas | Yannis Spyridis | Vasileios Argyriou
Georgios Tsoumplekas | Yannis Spyridis | Vasileios Argyriou
User-level affective modeling from social media requires integrating heterogeneous signals that unfold over time. While prior work has focused predominantly on textual analysis, visually expressed affect and temporal posting patterns also carry important psychological cues. However, these modalities are difficult to combine in practice due to sparse emotional evidence, asynchronous posting behavior, and frequent semantic misalignment between images and accompanying text. This paper introduces TiC-MuFormer, a time-enriched caption-integrated multimodal transformer that addresses these challenges by verbalizing visual content through image captioning before fusion and injecting temporal structure prior to cross-modal attention, enabling user trajectories to be modeled in a time-aware semantic space. We instantiate the method on a mental health detection task and demonstrate that it achieves state-of-the-art results across all user-level metrics, outperforming both unimodal and multimodal baselines. Ablation studies further show that temporal coverage, batch size and encoder choice jointly influence downstream accuracy, underscoring the importance of aligned temporal and semantic representations. Overall, this work highlights caption-guided temporal multimodality as a principled modeling strategy for general affective or psychiatric risk inference in social platforms.
Improving Neural Argumentative Stance Classification in Controversial Topics with Emotion-Lexicon Features
Mohammad Yeghaneh Abkenar | Weixing Wang | Manfred Stede | Mark A. Finlayson | Davide Picca | Panagiotis Ioannidis
Mohammad Yeghaneh Abkenar | Weixing Wang | Manfred Stede | Mark A. Finlayson | Davide Picca | Panagiotis Ioannidis
Argumentation mining comprises several subtasks, among which stance classification focuses on identifying the standpoint expressed in an argumentative text toward a specific target topic. While arguments—especially about controversial topics—often appeal to emotions, most prior work has not systematically incorporated explicit, fine-grained emotion analysis to improve performance on this task. In particular, prior research on stance classification has predominantly utilized non-argumentative texts and has been restricted to specific domains or topics, limiting generalizability. We work on five datasets from diverse domains encompassing a range of controversial topics and present an approach for expanding the Bias-Corrected NRC Emotion Lexicon using DistilBERT embeddings, which we feed into a Neural Argumentative Stance Classification model. Our method systematically expands the emotion lexicon through contextualized embeddings to identify emotionally charged terms not previously captured in the lexicon. Our expanded NRC lexicon (eNRC) improves over the baseline across all five datasets (up to +6.2 percentage points in F1 score), outperforms the original NRC on four datasets (up to +3.0), and surpasses the LLM-based approach on nearly all corpora. We provide all resources—including eNRC, the adapted corpora, and model architecture—to enable other researchers to build upon our work
Emotion Transcription in Conversation: A Benchmark for Capturing Subtle and Complex Emotional States through Natural Language
Yoshiki Tanaka | Ryuichi Uehara | Koji Inoue | Michimasa Inaba
Yoshiki Tanaka | Ryuichi Uehara | Koji Inoue | Michimasa Inaba
Emotion Recognition in Conversation (ERC) is critical for enabling natural human-machine interactions. However, existing methods predominantly employ categorical or dimensional emotion annotations, which often fail to adequately represent complex, subtle, or culturally specific emotional nuances. To overcome this limitation, we propose a novel task named Emotion Transcription in Conversation (ETC). This task focuses on generating natural language descriptions that accurately reflect speakers’ emotional states within conversational contexts. To address the ETC, we constructed a Japanese dataset comprising text-based dialogues annotated with participants’ self-reported emotional states, described in natural language. The dataset also includes emotion category labels for each transcription, enabling quantitative analysis and its application to ERC. We benchmarked baseline models, finding that while fine-tuning on our dataset enhances model performance, current models still struggle to infer implicit emotional states. The ETC task will encourage further research into more expressive emotion understanding in dialogue. The dataset is publicly available at https://github.com/UEC-InabaLab/ETCDataset.
SETUP: Sentence-level English-To-Uniform Meaning Representation Parser
Emma Markle | Javier Gutierrez Bach | Shira Wein
Emma Markle | Javier Gutierrez Bach | Shira Wein
Uniform Meaning Representation (UMR) is a novel graph-based semantic representation which captures the core meaning of a text, with flexibility incorporated into the annotation schema such that the breadth of the world’s languages can be annotated (including low-resource languages). While UMR shows promise in enabling language documentation, improving low-resource language technologies, and adding interpretability, the downstream applications of UMR can only be fully explored when text-to-UMR parsers enable the automatic large-scale production of accurate UMR graphs at test time. Prior work on text-to-UMR parsing is limited to date. In this paper, we introduce two methods for English text-to-UMR parsing, one of which fine-tunes existing parsers for Abstract Meaning Representation and the other, which leverages a converter from Universal Dependencies, using prior work as a baseline. Our best-performing model, which we call SETUP, achieves an AnCast score of 84 and a SMATCH++ score of 91, indicating substantial gains towards automatic UMR parsing.
This One or That One? A Study on Accessibility via Demonstratives with Multimodal Large Language Models
Yu Wang | Emmanuele Chersoni | Chu-Ren Huang
Yu Wang | Emmanuele Chersoni | Chu-Ren Huang
Accessibility refers to the ease with which a speaker can acquire an object, and it is often conveyed through demonstrative pronouns like “this” and “that”, indicating proximal or distal objects. Most importantly, accessibility also involves perspective shifts, which are essential for understanding differing viewpoints. In this case study, we adopt an evaluation dataset with a pair-to-pair question structure for referent identification based on demonstratives. Our experiments show that current Multimodal Large Language Models (MLLMs) exhibit markedly low performance in accessibility tasks requiring perspective shifts, with accuracies around 2.33% (Chinese) and 1.83% (English). Moreover, models struggle with qualitative characteristics and frame-based reasoning, often failing to apply implicit contextual rules unless explicitly encoded in training data. These limitations suggest that MLLMs rely heavily on surface co-occurrence instead of truly grounded, embodied experience. Our evaluation framework provides a robust lens revealing that MLLMs lack both self-other distinction—an essential aspect of self-awareness—and the embodied cognition necessary for reliable performance in practical embodied AI applications.
AMR Parsing beyond English: An Experiment on Bulgarian, French, Hungarian and Ukrainian
Ivaylo Mitov | Tadzhat Marharian | Zsofia F. Hauk | Samba FALL | Maxime Amblard | Bruno Guillaume
Ivaylo Mitov | Tadzhat Marharian | Zsofia F. Hauk | Samba FALL | Maxime Amblard | Bruno Guillaume
Under the assumption that the meaning of a sentence should be unchanged when it is translated into another language, recent work has developed on cross-lingual semantic parsing in an effort to extend the access to semantic resources beyond English. In this paper, we develop the automatic production of Abstract Meaning Representations (AMR), a graph-based semantic formalism, for four languages – Bulgarian, French, Hungarian and Ukrainian. We achieve high-performance on French and Hungarian, and execute, to our knowledge, the first semantic parsing of Bulgarian and Ukrainian on translations of the AMR3.0 corpus (Knight et al., 2020). Furthermore, we perform a complementary experiment on a novel parallel corpus of gold AMR annotations of the first chapter of “The Adventures of Pinocchio” in Bulgarian and Ukrainian. The experiment reveals that, despite their above-average performance, the models’ performance decreases when probed on texts outside of the domain of the training data.
Semantic Parsing for Evaluating Large Language Models: Separating Linguistic Abilities with YARN
Rémi DE VERGNETTE | Maxime Amblard
Rémi DE VERGNETTE | Maxime Amblard
We evaluate large language models (LLMs) through semantic parsing into Yarn, a structured meaning representation that distinguishes predicate–argument structure from higher-level linguistic features such as tense, aspect, and modality. For evaluation, we employ SmatchY, a fine-grained metric designed to assess different layers of meaning independently. Our experiments test multiple LLMs under varied conditions, including inference modes, linearization formats (JSON and logic-inspired CFG), and the presence or absence of auxiliary supervision via partial semantic parses. Results show that model performance is highly sensitive to both representational design and supervision, with no single configuration consistently outperforming the others. While some models gain from additional semantic information in prompts, others are negatively affected. A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification. Consistent with prior work, our findings highlight the limited capacity of current LLMs to generate fully formal meaning representations.
Two Ojibwe Constraint Grammars: Morphological Disambiguation and Dependency Parsing
Matthias Diederichsen | Christopher Hammerly
Matthias Diederichsen | Christopher Hammerly
This paper presents the first iteration of two connected Ojibwe constraint grammars, one for morphological disambiguation and one for syntactic parsing. Due to the polysynthetic nature of Ojibwe, along with its status as a low-resource language, the disambiguation grammar proves to be an effective and resource-efficient tool for morphological disambiguation, successfully eliminating 32% of redundant readings and fully resolving 41% of ambiguous tokens. The dependency grammar focuses on assigning dependency relations to model argument structure, where the constraint grammar once again proves to be an effective paradigm, with F1 scores of 0.97 for subject and 0.94 for object relations. The rule-based design of both grammars is linguistically informed, allowing for precise modeling of language-specific phenomena such as animacy, obviation, and verb-argument agreement. Applications of the two constraint grammars include building a disambiguated morphologically-tagged corpus of the Ojibwe language and creating a treebank for the Ojibwe language following the widely adopted CoNLL-U format.
Multimodal LLMs Do Not Compose Skills Optimally across Modalities
Paula Ontalvilla | Aitor Ormazabal | Gorka Azkune
Paula Ontalvilla | Aitor Ormazabal | Gorka Azkune
Skill composition is the ability to combine previously learned skills to solve new tasks. As neural networks acquire increasingly complex skills during their pretraining, it is not clear how successfully they can compose them. In this paper, we focus on Multimodal Large Language Models (MLLM), and study their ability to compose skills across modalities. To this end, we design three evaluation tasks which can be solved sequentially composing two modality-dependent skills, and evaluate several open MLLMs under two main settings: i) prompting the model to directly solve the task, and ii) using a two-step cascaded inference approach, which manually enforces the composition of the two skills for a given task. Even with these straightforward compositions, we find that all evaluated MLLMs exhibit a significant cross-modality skill composition gap. To mitigate the aforementioned gap, we explore two alternatives: i) use chain-of-thought prompting to explicitly instruct MLLMs for skill composition and ii) a specific fine-tuning recipe to promote skill composition. Although those strategies improve model performance, they still exhibit significant skill composition gaps, suggesting that more research is needed to improve cross-modal skill composition in MLLMs.
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
Maha Tufail Agro | Atharva A. Kulkarni | Karima Kadaoui | Zeerak Talat | Hanan Aldarmaki
Maha Tufail Agro | Atharva A. Kulkarni | Karima Kadaoui | Zeerak Talat | Hanan Aldarmaki
Motivated by a growing research interest into automatic speech recognition (ASR), and the growing body of work for languages in which code-switching (CS) often occurs, we present a systematic literature review of code-switching in end-to-end ASR models. We collect and manually annotate papers published in peer reviewed venues. We document the languages considered, datasets, metrics, model choices, and performance, and present a discussion of challenges in end-to-end ASR for code-switching. Our analysis thus provides insights on current research efforts and available resources as well as opportunities and gaps to guide future research.
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in VideoLMs for Multimodal Sarcasm Detection.
Anisha Saha | Varsha Suresh | Timothy Hospedales | Vera Demberg
Anisha Saha | Varsha Suresh | Timothy Hospedales | Vera Demberg
Sarcasm is a specific type of irony which involves discerning what is said from what is meant. Detecting sarcasm depends not only on the literal content of an utterance but also on non-verbal cues such as speaker’s tonality, facial expressions and conversational context. However, current multimodal models struggle with complex tasks like sarcasm detection, which require identifying relevant cues across modalities and pragmatically reasoning over them to infer the speaker’s intention. To explore these limitations in VideoLMs, we introduce MUStReason, a diagnostic benchmark enriched with annotations of modality-specific relevant cues and underlying reasoning steps to identify sarcastic intent. In addition to benchmarking sarcasm classification performance in VideoLMs, using MUStReason we quantitatively and qualitatively evaluate the generated reasoning by disentangling the problem into perception and reasoning and aim to pinpoint the current gaps in these VideoLMs. Furthermore, to facilitate structured pragmatic reasoning, we propose PragCoT, a framework that steers VideoLMs to focus on implied intentions over literal meaning, a property core to detecting sarcasm. Code and dataset are available at https://github.com/anisha0325/MUStReason
Human-Centered Multimodal Fusion for Sexism Detection in Memes with Eye-Tracking, Heart Rate, and EEG Signals
Iván Arcos Gabaldón | Paolo Rosso | Elena Gomis Vicent
Iván Arcos Gabaldón | Paolo Rosso | Elena Gomis Vicent
The automated detection of sexism in memes is a notoriously challenging task due to multimodal ambiguity, cultural nuance, and the use of humor to provide plausible deniability. As a result, content-only models often fail to capture the complexity of human perception. To address this fundamental limitation, we introduce and validate a human-centered paradigm that augments standard content features with rich physiological data. We created a novel resource by recording Eye-Tracking (ET), Heart Rate (HR), and Electroencephalography (EEG) from 16 subjects (8 per experiment) while they viewed 3,984 memes from the EXIST 2025 dataset. Our statistical analysis reveals significant physiological differences in how subjects process sexist versus non-sexist content. Sexist memes were associated with higher cognitive load (evidenced by increased fixation counts and longer reaction times), and with differences in EEG spectral power across the Alpha, Beta, and Gamma frequency bands. This pattern, commonly linked in previous research to increased attentional engagement and cognitive effort during visual processing, suggests that sexist memes may elicit more demanding neural activity compared to non-sexist ones. Building on these findings, we propose a novel multimodal fusion model that integrates these physiological signals with enriched textual-visual features derived from a Vision-Language Model (VLM). Our final model achieves an AUC of 0.794 in binary sexism detection, a statistically significant 3.4% improvement over a powerful VLM-based baseline. The fusion of physiological data proves particularly effective for nuanced and ambiguous cases, boosting the F1-score for the most challenging fine-grained category, Misogyny and Non-Sexual Violence, by an unprecedented 26.3%. Our work demonstrates that human physiological responses provide a robust, objective signal of perception that can significantly enhance the accuracy and human-awareness of automated systems for countering online sexism.
Nos_Brais-GL: A FAIR Galician TTS Corpus for Neural Speech Synthesis
Adina Ioana Vladu | Antonio Moscoso Sánchez | Carmen Magariños | María Perez Lago | Elisa Fernández Rei
Adina Ioana Vladu | Antonio Moscoso Sánchez | Carmen Magariños | María Perez Lago | Elisa Fernández Rei
This paper introduces Nos_Brais-GL, a new open-access high-quality Galician speech corpus designed for the development of neural Text-to-Speech (TTS) systems. Nos_Brais-GL contains approximately 18 hours of professionally recorded male speech and a carefully curated set of utterances selected to ensure linguistic variation and phonetic and prosodic richness. Beyond its immediate application in synthetic speech generation, Nos_Brais-GL exemplifies good practices in TTS corpus design for lesser-resourced languages, emphasizing methodological transparency, open licensing, and interoperability.
DR-CUP: A Dataset on Real-time Commentary in U.S. Presidential Debates
Yu-Yu Chang | Huan-Wen Ho | Chung-Chi Chen | Ming-Hung Wang
Yu-Yu Chang | Huan-Wen Ho | Chung-Chi Chen | Ming-Hung Wang
Presidential debates are critical platforms for political discourse, yet existing research lacks datasets tailored for analyzing real-time professional commentary. To address this, we introduce the Dataset on Real-time Commentary in U.S. Presidential debates (DR-CUP), which aligns U.S. presidential debate transcripts (2016–2024) with professional commentary and annotations. DR-CUP supports research on commentary understanding, planning, and generation, offering insights into expert analysis and its role in contextualizing complex political discourse. In pilot studies, we evaluated state-of-the-art large language models (LLMs), revealing notable performance differences in understanding expert commentary and planning for generating professional commentary. DR-CUP is the first dataset to incorporate real-time cross-document alignment for debate data, providing a comprehensive resource for advancing research in political communication and computational social science.
Russian Generative Spelling, Punctuation and Capitalization Correction
Nikita Martynov | Danil Astafurov | Ulyana Isaeva | Ivan Vasil’yevich Maksimov | Joqsan Azocar | Dmitrii Kosenko | Alena Fenogenova
Nikita Martynov | Danil Astafurov | Ulyana Isaeva | Ivan Vasil’yevich Maksimov | Joqsan Azocar | Dmitrii Kosenko | Alena Fenogenova
This paper presents SAGE, an open-access framework that encloses a set of models specifically designed for the generative correction of spelling, punctuation, and capitalization errors in Russian. The release includes four models, featuring a Russian-English version and a distilled version for easy use and cost-effectiveness. The models are pre-trained using a sequence-to-sequence approach on artificial errors that mimic human mistakes and fine-tuned on annotated multi-domain texts. A set of carefully engineered auxiliary learning objectives is employed during pre-training to enrich the models with additional semantic and syntactic information. Evaluations indicate that SAGE models, despite having a small number of parameters, outperform top-tier multilingual and Russian-specific large language models, including both closed- and open-source options, and are considered state-of-the-art. We release the online demo powered by a single Nvidia A100 80GB GPU as a Web service, which allows to simultaneously test the most advanced SAGE model of 1.7B parameters, its distilled version and the Russian-English SAGE model.
Using Multimodal and Language-Agnostic Sentence Embeddings for Abstractive Summarization
Chaimae Chellaf El Hammoud | Salima Mdhaffar | Yannick Estève | Stéphane Huet
Chaimae Chellaf El Hammoud | Salima Mdhaffar | Yannick Estève | Stéphane Huet
Abstractive summarization aims to generate concise summaries by creating new sentences, allowing for flexible rephrasing. However, this approach can be vulnerable to inaccuracies, particularly ‘hallucinations’ where the model introduces non-existent information. In this paper, we leverage the use of multimodal and multilingual sentence embeddings derived from pre-trained models such as LaBSE, SONAR, and BGE-M3, and feed them into a modified BART-based French model. A Named Entity Injection mechanism that appends tokenized named entities to the decoder input is introduced, in order to improve the factual consistency of the generated summary. Our novel framework, SBARThez, is applicable to both text and speech inputs and supports cross-lingual summarization; it shows competitive performance relative to token-level baselines, especially for low-resource languages, while generating more concise and abstract summaries.
Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering
Purva Chiniya | Kevin Joseph Scaria | Sagar Chaturvedi
Purva Chiniya | Kevin Joseph Scaria | Sagar Chaturvedi
Large language models (LLMs) remain susceptible to jailbreak and direct prompt-injection attacks, yet the strongest defensive filters frequently over- refuse benign queries and degrade user experience. Previous work on prompt injection detection such as, GradSafe, detects unsafe prompts with a single “accept all” anchor token, but its threshold is brittle and it offers no deterministic guarantee that harmful content will not be emitted once decoding begins. We introduce Gradient-Controlled Decoding (GCD), a training-free guardrail that combines with both an acceptance anchor (“Sure”) and refusal anchor (“Sorry”) tightening the decision boundary and lowering false positives. In the mitigation stage, if a prompt is flagged, GCD preset-injects one or two refusal tokens ("Sorry, I can’t . . . ") before autoregressive decoding resumes, guaranteeing first- token safety regardless of sampling strategy. On ToxicChat, XSTest-v2, and AdvBench, GCD reduces false positives by 52% vs. GradSafe at comparable recall, lowers attack success rate by up to 20% vs. the strongest decoding-only baseline, adds under 15-20 ms latency on an average on V100 instances, transfers to LLaMA-2-7B, Mixtral-8×7B, and Qwen-2-7B, and requires only 20 template prompts. GCD is a lightweight, scalable safety layer for real-time LLM deployment.
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
Pavel Braslavski | Dmitrii Iarosh | Nikita Sergeevich Sushko | Andrey Sakhovskiy | Vasily Konovalov | Elena Tutubalina | Alexander Panchenko
Pavel Braslavski | Dmitrii Iarosh | Nikita Sergeevich Sushko | Andrey Sakhovskiy | Vasily Konovalov | Elena Tutubalina | Alexander Panchenko
We present a configurable pipeline and the associated code that can be used to generate multilingual sets of entities with specified characteristics, such as domain, geographical location and popularity, using data from Wikipedia and Wikidata. These datasets are intended for evaluating the factuality of LLMs’ long-form generation, thereby complementing evaluation based on short-form QA datasets. We present the RiDiC dataset as an example of this approach. RiDiC contains 3,000 entities from three domains – rivers, natural disasters, and car models – spanning different popularity tiers. Each entity is accompanied by its geographical location, English and Chinese names (if available) and relevant English and Chinese Wikipedia content, which is used to evaluate LLMs’ responses. Generations about RiDiC entities were obtained from three LLMs in English and Chinese. These were then evaluated using a third-party factuality checker, which showed that entities from our dataset caused even frontier models to hallucinate. The code, data and generation/evaluation scripts have been released to enable the approach to be extended to new LLMs, languages and domains.
MeteoGalEus: An Iberian Multilingual Weather Dataset in Galician, Euskera, and Spanish
Ainhoa Vivel-Couso | Nella Zabrina Pramata | David Robredo | Aitor Soroa | Jose Maria Alonso-Moral
Ainhoa Vivel-Couso | Nella Zabrina Pramata | David Robredo | Aitor Soroa | Jose Maria Alonso-Moral
This paper introduces MeteoGalEus, a multilingual weather dataset that combines meteorological observations from two Spanish regional agencies, Euskalmet and MeteoGalicia. The dataset contains daily records spanning 4 years and 6 months, with aligned observations for both sources. MeteoGalEus captures key meteorological variables including temperature, wind and state of the sky. The dataset is provided in a structured format, facilitating data analysis and integration, with textual forecasts available in the official languages for each region (i.e., Galician and Spanish for MeteoGalicia; Euskera and Spanish for Euskalmet). By merging and harmonizing data from two regional agencies, MeteoGalEus is a unique resource for cross-regional weather analysis and multilingual climate studies. This dataset is suited for tasks requiring high-quality, aligned, and standardized weather data across multiple languages and regions. We conducted baseline experiments using LLaMA-based models in both zero-shot and fine-tuned settings to illustrate the use of MeteoGalEus for natural language generation (NLG). Fine-tuning led to consistent improvements across all metrics, with BERTScore increasing from 0.68 to 0.79, ROUGE from 0.20 to 0.35, and BLEU from 0.02 to 0.17 in the best-performing model. The experiments show how MeteoGalEus can be taken as a benchmark for multilingual and cross-regional NLG tasks.
RadTimeline: Timeline Summarization for Longitudinal Radiological Lung Findings
Sitong Zhou | Meliha Yetisgen | Mari Ostendorf
Sitong Zhou | Meliha Yetisgen | Mari Ostendorf
Tracking findings in longitudinal radiology reports is crucial for accurately identifying disease progression, and the time-consuming process would benefit from automatic summarization. This work introduces a structured summarization task, where we frame longitudinal report summarization as a timeline generation task, with dated findings organized in columns and temporally related findings grouped in rows. This structured summarization format enables straightforward comparison of findings across time and facilitates fact-checking against the associated reports. The timeline is generated using a 3-step LLM process of extracting findings, generating group names, and using the names to group the findings. To evaluate such systems, we create RadTimeline, a timeline dataset focused on tracking lung-related radiologic findings in chest-related imaging reports. Experiments on RadTimeline show tradeoffs of different-sized LLMs and prompting strategies. Our results highlight that group name generation as an intermediate step is critical for effective finding grouping. The best configuration has some irrelevant findings but very good recall, and grouping performance is comparable to human annotators.
InstructSum: A Benchmark to Evaluate Instruction-Following Capability of Large Language Models in Summarization
Kosuke Nishida | Kyosuke Nishida | Itsumi Saito
Kosuke Nishida | Kyosuke Nishida | Itsumi Saito
Pre-trained large language models (LLMs) align their outputs with user intent through natural language instructions. In the summarization task, conciseness of the output is inherently required, which makes the instruction-following capability of LLMs particularly important. That is, providing supplementary information beyond the instruction can be undesirable. In this study, we introduce a novel benchmark, InstructSum, consisting of 3,309 types of instructions to evaluate the instruction-following capability in the summarization task. InstructSum has multiple instructions per source text, and thus it enables the evaluation of how LLMs adjust the content of the summary according to the instructions. Our experiments with six LLM families revealed the challenges that LLMs face in this task. For example, LLMs provide polite and helpful responses with irrelevant information; they go beyond instructions and fail to respond with a concise summary.
NOVELSUM: Evaluating Long-Form Summary Generation for Historical Scandinavian Novels
Ali Al-Laith | Alexander Conroy | Kirstine Nielsen Degn | Jens Bjerring-Hansen | Daniel Hershcovich
Ali Al-Laith | Alexander Conroy | Kirstine Nielsen Degn | Jens Bjerring-Hansen | Daniel Hershcovich
We study long-form summarization of late-19th-century Danish and Norwegian novels and propose NOVELSUM, an evaluation resource and protocol tailored to literary narrative. We use a curated set of historical novels paired with professional reference summaries to establish baselines with long-document encoder–decoder models and prompt-based large-context LLMs. We evaluate with automatic metrics, expert human judgments, and LLM-as-judge scoring. Our human study identifies evaluation dimensions and literary facets that achieve substantial inter-annotator agreement and align with scholarly expectations. We further analyze reference-free evaluation, showing when it tracks expert trends and where it fails (notably for factual and setting-related criteria), thereby clarifying its utility when gold references or expert readers are unavailable. Our results benchmark long-context and prompted LLM approaches on historical literary prose and offer a practical path for human-grounded and reference-free assessment.
Evaluating Large Language Models for Text-to-Gloss Translation in Kazakh-Russian Sign Language: A Pilot Study
Zhanibek Kozhirbayev | Alfarabi Imashev
Zhanibek Kozhirbayev | Alfarabi Imashev
Conceptual glossing involves a systematic linguistic transformation in which the models must preserve meaning, grammatical integrity, and punctuation while turning the real language into a more structured structure. The purpose of this study is to assess the accuracy and dependability of glosses produced by these models by juxtaposing them with human-annotated standards, investigating whether the models maintain essential linguistic characteristics. By identifying the strengths and weaknesses of each model, we want to determine which architectures are most suitable for organized language tasks, such as glossing. This may reduce the manual labor required for linguistic annotation by experts while maintaining superior quality outcomes. And help deaf signers with weak reading skills interpret written paragraphs into glosses, making them more comprehensible and naturally looking to them. Text-to-gloss translation converts written or spoken language into sign language glosses, enhancing accessibility for the Deaf and Hard of Hearing (DHH) community. This pilot study evaluates four large language models (LLMs): GPT-4-turbo, Grok 3, Deepseek-V3, and Gemini 20 Flash to generate conceptual glosses in Kazakh-Russian Sign Language (K-RSL), still an under-resourced sign language. Using a dataset of 250 Russian sentences with expert-annotated K-RSL glosses, we assess performance across METEOR, BLEU, BERTScore, and WER. Results show Deepseek-V3 excels on complex texts (METEOR: 0.426 for K-RSL word order, 0.377 for fairytale paragraphs), while Gemini 20 Flash performs strongly on short sentences (METEOR: 0.602). These findings demonstrate LLMs’ potential to automate gloss production, reducing manual annotation and aiding DHH individuals with reading comprehension. Challenges include K-RSL’s unique grammar and limited datasets. This is the first study to apply LLMs to K-RSL glossing and examine the potential efficacy of autonomous gloss production.
HotelCheckSpan: A Benchmark Dataset for LLM Faithfulness
Patricia Schmidtova | Ondrej Dusek | Saad Mahamood
Patricia Schmidtova | Ondrej Dusek | Saad Mahamood
Hallucinations are among the most persistent and challenging issues in large language model (LLM) outputs. This particularly holds in domains that combine both objective and subjective content, such as hotel descriptions, that are intended to be enticing advertisements for the hotel. Distinguishing between factual errors and interpretative exaggeration is often subtle, complicating both human and automated evaluation. To address this, we present HotelCheckSpan, the first span-level faithfulness dataset for the hotel domain. Each example aggregates one or more hotel descriptions, and human-annotated summaries are labeled with three error types: Incorrect, Misleading, and Not Checkable. By marking the precise spans where errors occur, the dataset captures fine-grained information about the nature of hallucinations and factual inconsistencies. In addition to human annotations, we collect span-level judgments from multiple LLMs, enabling direct human–model comparisons. Our analysis shows that inter-annotator agreement varies substantially across aggregation levels: example-level agreement can mask subtle span-level disagreements, while soft and hard F1 variants highlight discrepancies in both span placement and error categorization. HotelCheckSpan provides a benchmark for studying ambiguity and disagreement, validating automatic faithfulness metrics, and evaluating LLMs as judges, offering a rich resource for research on faithfulness, subjectivity, and annotation practices in mixed-content domains
The availability of many fine-tuned neural language models for different tasks naturally leads to the question of whether it is worthwhile to combine them, particularly through parameter merging, which is the least resource-intensive option. Among the many existing methods, some focus on parameter alignment before actual merging. In this article, we propose a new method within this research area, based on Procrustes analysis. We evaluate this method for merging fine-tuned models for the same task, derived from the same encoder-based model. Considering nine tasks from the GLUE benchmark, three Named Entity Recognition tasks, and six reference merging methods, we show that our proposal can improve upon existing merging methods in most tested configurations.
MetaCORA: A Meta-Learned Curriculum for Adversarial and Contrastive Robustness in Speech Recognition
Yuqian Dai | Chun Fai Chan | Ying Ki Wong | Tsz Ho Pun
Yuqian Dai | Chun Fai Chan | Ying Ki Wong | Tsz Ho Pun
Pre-trained speech models like Whisper demonstrate impressive performance under ideal conditions but still face robustness challenges in low-resource language scenarios. We introduce Meta Curriculum Optimization for Robust ASR (MetaCORA), a novel meta-curriculum adaptive framework that improves speech recognition for low-resource Hong Kong Cantonese by integrating adversarial training with feature contrastive learning. Our approach dynamically adjusts three critical hyperparameters: adversarial perturbation magnitude, optimization step size, and contrastive learning temperature, allowing the model to adapt to varying training difficulties throughout the learning process. Unlike traditional meta-learning approaches, our framework does not rely on end-to-end differentiability but instead utilizes validation performance as a signal to guide hyperparameter adjustments. Experimental results demonstrate that our approach achieves lower WER than standard Whisper fine-tuning, commercial speech recognition systems, and LLM-based methods. Ablation studies confirm the necessity of each component, as removing any single element leads to a measurable drop in performance. The model also exhibits robustness under noisy conditions, achieving consistently lower WER than baseline systems. Further analysis shows that MetaCORA effectively compresses the distance between adversarial feature representations while maintaining well-separated class boundaries in the embedding space, providing a mechanistic explanation for its improvement.
Insights from Transfer Learning Experiments with Word-in-Context and Word Sense Disambiguation Models
Alp Mujko | Dominik Schlechtweg
Alp Mujko | Dominik Schlechtweg
We investigate the relationship between Word-in-Context (WiC) and Word Sense Disambiguation (WSD) by examining how training on one or both tasks affects performance on the other. Using established English datasets we train a unified sentence transformer (xlm-roberta-large) with target-word highlighting and contrastive loss. Models are evaluated on WiC and WSD benchmarks across single-task, joint, and combined dataset configurations. Results show that joint training consistently improves or maintains WiC performance, particularly in low-resource settings, while WSD benefits mainly when annotated data is limited. Cross-task experiments demonstrate strong transfer: WSD-trained models generalize effectively to WiC, and WiC-trained models outperform baselines on WSD, indicating shared context-sensitive lexical representations. Combining multiple WiC datasets further enhances accuracy and stability. These findings highlight the complementary nature of WiC and WSD and demonstrate that unified training strategies can yield more robust and generalizable sense disambiguation models. The results provide practical guidance for designing datasets and models in multilingual and low-resource contexts, emphasizing the value of leveraging shared semantic representations.
Joint Identification and Induction of Semantic Frames with Scalable Semi-Supervised Graph Clustering
Fabian Barteld | Steffen Remus | Saba Anwar | Julian Stawecki | Alexander Ziem | Chris Biemann
Fabian Barteld | Steffen Remus | Saba Anwar | Julian Stawecki | Alexander Ziem | Chris Biemann
Current methods for automatically assigning frames to their evoking words can be divided into frame identification and frame induction. In frame identification, frame names coming from a labeled dataset are assigned to unseen instances, a classical supervised labeling task. However, the training datasets are known to be incomplete in terms of real-world frames, resulting in an issue with potentially new frame labels. In frame induction, instances are clustered regarding the frames they evoke, a classical unsupervised clustering task. However, existing training data is not used to identify known frames. To overcome these shortcomings, we propose to use semi-supervised clustering for combined frame identification and frame induction. By using constrained clustering with hard constraints coming from labeled data, the resulting clusters contain only labeled instances with the same label. Thus, frame names can be easily assigned. We show for English and German datasets that using semi-supervised clustering improves the quality of frame induction compared to unsupervised clustering methods and results in notably good performance regarding frame identification.
Low-Rank Compression of Language Models via Differentiable Rank Selection
Sidhant Sundrani | Francesco Tudisco | Pasquale Minervini
Sidhant Sundrani | Francesco Tudisco | Pasquale Minervini
Approaches for compressing large-language models using low-rank decomposition have made strides, particularly with the introduction of activation and loss-aware SVD, which improves the trade-off between decomposition rank and downstream task performance. Despite these advancements, a persistent challenge remains–selecting the optimal ranks for each layer to jointly optimise compression rate and downstream task accuracy. Current methods either rely on heuristics that can yield sub-optimal results due to their limited discrete search space or are gradient-based but are not as performant as heuristic approaches without post-compression fine-tuning. To address these issues, we propose Learning to Low-Rank Compress (LLRC), a gradient-based approach that directly learns the weights of masks that select singular values in a fine-tuning-free setting. Using a calibration dataset, we train only the mask weights to select fewer and fewer singular values while minimising the divergence of intermediate activations from the original model. Our approach outperforms competing methods that similarly require no post-compression fine-tuning across various compression rates on common-sense reasoning and open-domain question-answering tasks. For instance, with a compression rate of 20% on Llama-2-13B, LLRC outperforms the competitive Sensitivity-based Truncation Rank Searching (STRS) on MMLU, BoolQ, and OpenbookQA by 12%, 3.5%, and 4.4%, respectively. Compared to other compression techniques, our approach consistently outperforms fine-tuning-free variants of SVD-LLM and LLM-Pruner across datasets and compression rates. Our approach also performs competitively with LLM-Pruner after fine-tuning on Llama-2-7B and Llama-2-13B.
Self-supervised Data Augmentation for Text Classification in Low-Data Settings
Deyu Ding | Mengying Wang | Andreas Spitz
Deyu Ding | Mengying Wang | Andreas Spitz
Due to data sparsity and high annotation cost, data augmentation has established itself as an effective tool for boosting model performance on supervised NLP tasks. Where task-agnostic augmentation methods tend to act as simple regularizers for the data, task-aware methods also leverage labels for the generation of data that are most suitable for downstream tasks. While prior work has investigated generation and sampling strategies individually, the potential of a self-supervised approach that leverages multiple pre-trained models in generation and sampling remains underexplored. To address this issue, we present an ensemble-based framework of language models that proposes augmentation candidates and internally reviews their suitability for low-resource text classification tasks. We evaluate our model on six classification benchmarks and find that it consistently outperforms state-of-the-art data augmentation baselines in classification accuracy by an average of 0.97 points in low-data scenarios.
Distribution-aware Low-bitwidth Quantization for Large Language Models
Bao Tan Duy Huynh | Takashi Tsunakawa | Masafumi Nishida
Bao Tan Duy Huynh | Takashi Tsunakawa | Masafumi Nishida
The increasing scale and complexity of large language models (LLMs) present significant computational and memory challenges, limiting their widespread deployment. Post-training quantization (PTQ) has emerged as a key technique for mitigating these challenges without costly retraining. However, compressing models to ultra-low bitwidths (e.g., 2-3 bits) while maintaining accuracy remains a major challenge. In this study, we present a comprehensive PTQ framework that addresses this problem by compressing LLM weights through three core innovations: (1) a calibration process guided by Kullback-Leibler divergence minimization to preserve the original weight distribution, (2) a learnable codebook optimization mechanism employing noise substitution for vector quantization to enable robust gradient estimation, and (3) a layer-grouping strategy based on statistical distribution similarity to improve parameter efficiency. Experimental evaluations on large-scale models show that the proposed framework achieves competitive performance compared with state-of-the-art quantization techniques. Importantly, these results are obtained without any post-quantization fine-tuning, highlighting the efficiency and practical applicability of our approach for deploying highly compressed LLMs.
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
ChengYeh Yang | Chien-Chun Wang | Li-Wei Chen | Hung-Shin Lee | Hsin-Min Wang | Berlin Chen
ChengYeh Yang | Chien-Chun Wang | Li-Wei Chen | Hung-Shin Lee | Hsin-Min Wang | Berlin Chen
Low-resource automatic speech recognition remains a critical challenge due to the scarcity of transcribed data for many languages.Taiwanese Hokkien exemplifies this problem as, although extensive speech content exists in television dramas and online videos, transcriptions are scarce and most available subtitles are in Mandarin.To address this gap, this paper presents TG-ASR for Taiwanese drama speech recognition, a translation-guided ASR framework that leverages multilingual translation embeddings to enhance recognition in low-resource conditions.The framework centers on the parallel gated cross-attention (PGCA) mechanism, which adaptively integrates embeddings from multiple auxiliary languages into the ASR decoder.This mechanism enables robust cross-linguistic semantic guidance while maintaining stable optimization and avoiding interference between languages.To support future research, we release YT-THDC, a 30-hour corpus of Taiwanese drama speech with aligned Mandarin subtitles and manually verified Taiwanese transcriptions.Extensive experiments and analysis identify which auxiliary languages most effectively improve Taiwanese ASR, achieving a 13.51% relative reduction in character error rate and demonstrating the potential of translation-guided learning for underrepresented languages in real-world scenarios.
Harnessing Synergy in Context and Emoji for Joint Detection of Harmful Online Content in Multi-turn Conversations
Feiyan Hu | Ciara Anne Byrne | Jiang Zhou | Rena Maycock | Mark Langan
Feiyan Hu | Ciara Anne Byrne | Jiang Zhou | Rena Maycock | Mark Langan
Detecting harmful content, such as cyberbullying, self-harm, and grooming, in self-generated content or conversations is an emerging research area with significant potential for positive social impact. However, challenges such as the scarcity of real-world conversational data, labor-intensive annotation processes, and inconsistent content policies hinder understanding and evaluating the performance of harmful content detection systems. In this study, we utilize openly available forum data to construct conversation proxies, facilitating the analysis and detection of harmful content. We undertook extensive efforts to label the conversational data using a consistent content policy developed by experts, with ten annotators contributing to the labeling process. Our experiments investigated the impact of context window size and found that performance in joint detection improved gradually up to a context window of 16 sentences, after which performance plateaued. Additionally, experiments with emojis demonstrated that using a tokenizer capable of decoding emojis yielded the best performance, while either removing emojis or converting them to text resulted in inferior outcomes.
Dynamic Layer Selection for Efficient Tone Recognition in Self-Supervised Speech Models
Saint Germes B. BENGONO OBIANG | Norbert TSOPZE | Paulin MELATAGIA YONTA
Saint Germes B. BENGONO OBIANG | Norbert TSOPZE | Paulin MELATAGIA YONTA
Low-resource tonal languages present significant challenges to speech processing technologies, due to limited training data and the critical role of pitch variation in expressing meaning. This paper applies established weighted layer combination methods to tone recognition in such languages, with a specific focus on Yoruba and Yemba. Building on our previous work with Wav2vec 2.0 representations and the weighted-sum methodology from Yang et al. (2024), we investigate layer specialisation in the SSA-HuBERT self-supervised speech model for tonal tasks. Our systematic analysis reveals significant performance differences between different layers, with middle layers generally outperforming both lower and upper layers for tonal recognition tasks. While typical approaches only use the output of the last layer, our experiments show that weighted layer combination outperforms the last layer by 20.4% and 15.8% relative improvement in tone error rate (TER) for Yoruba and Yemba, respectively. In addition to performance improvements, our approach provides dramatic computational efficiency gains, reducing the resources required by over 90% compared to evaluating each layer separately. Analysis of the learned layer weights reveals language-specific patterns, with Yoruba favouring middle layers and Yemba giving more weight to early layers. These results provide valuable insights into how tonal information is encoded in self-supervised speech models, and demonstrate a practical application of established layer combination methods in low-resource language contexts.
Intent Recognition in Speech-to-Text Processing in the Context of Natural Interaction with Cognitive Assistive Systems
Behnam Ensan | Magnus Jung | Matthias Busch | Adreas Wendemuth
Behnam Ensan | Magnus Jung | Matthias Busch | Adreas Wendemuth
This study investigates efficient speech-to-intent recognition for human–robot interaction in elderly-care environments in German, targeting deployment on resource-constrained platforms such as the Jetson AGX Orin. To benchmark performance, we created a domain-specific German dataset with two sub-datasets (PaSID and PaSynTex) that simulate specific nursing home communication scenarios. Two alternative speech-to-intent pipelines were developed and evaluated: a two-stage system combining automatic speech recognition (ASR) with a large language model (LLM), and an end-to-end large audio–language model (LALM) architecture. The performance of Whisper-based ASR systems was evaluated across a wide variety of LLMs and several LALMs, comparing intent-classification accuracy, latency, and resource efficiency. The results indicate that optimized ASR + LLM configurations, particularly Whisper Turbo coupled with Phi-3.5-mini or Qwen 2.5-7B, outperform unified LALM approaches while maintaining substantially lower memory and inference costs. Also, the analysis shows that, the unified LALM models outperform the two-step integration of ASR + LLM in the same configuration, but at the cost of higher resource utilization, likely due to limited optimization for edge deployment. Overall, the findings provide initial evidence that modular ASR + LLM pipelines provide a more practical solution for real-time, on-device intent recognition in assistive robotics in German, offering an effective trade-off between performance and deployability on resource-constrained platforms.
Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
Kentaro Ueda | François Portet | Hirohiko Suwa | Keiichi Yasumoto
Kentaro Ueda | François Portet | Hirohiko Suwa | Keiichi Yasumoto
While LLMs excel at general tasks, they struggle in specialized domains like finance, requiring diverse skills in domain knowledge, mathematical reasoning, and multilingual processing. Merging domain-specific Continual Pre-training (CPT) “experts” offers a practical alternative to costly and unstable multi-skill training. However, unlike established Supervised Fine-Tuning (SFT) model-based merging, CPT model merging remains largely unexplored. We address this gap by creating financial LLMs from experts in finance, math, and Japanese. We propose a three-stage evaluation focusing on knowledge recovery, complementarity, and emergence, and assess three merging methods (Task Arithmetic, TIES, and DARE-TIES) on a comprehensive financial benchmark curated from 18 tasks across 8 established datasets. Results show that merging an expert with its base model recovers general knowledge lost during CPT, while merging experts improves performance and can yield emergent cross-domain skills. Among the methods, Task Arithmetic performs strongly but is hyperparameter-sensitive, whereas TIES is more robust. Our findings also suggest that while model similarity correlates with merging success, emergent skills depend on more complex factors. This work presents the first foundational analysis of CPT model merging, establishing a principled framework and providing clear guidance for building multi-skill LLMs from existing assets.
Phonetic-based Ranking for Improved Pseudo-Labeling in Low-Resource ASR
Marco Matassoni | Roberto Gretter | Falavigna Daniele | Mohamed Nabih Ali Mohamed Nawar | Alessio Brutti | Matteo Negri | Mauro Cettolo | Marco Gaido | Sara Papi | Luisa Bentivogli
Marco Matassoni | Roberto Gretter | Falavigna Daniele | Mohamed Nabih Ali Mohamed Nawar | Alessio Brutti | Matteo Negri | Mauro Cettolo | Marco Gaido | Sara Papi | Luisa Bentivogli
The rise of large language models has boosted speech and language technologies; however, where transcripts of audio data are limited, the performance of current technology is not yet satisfactory. One common strategy to tackle data scarcity is leveraging pseudo-labels, for example automatically transcribing data with a pre-trained ASR. One critical issue of this approach is assessing the quality of the automatic transcriptions, that may be rather bad for low-resourced languages. While several filtering approaches exist in literature, they typically work with decent pre-trained ASR models but may fail otherwise. In this work we propose a phonetic-based ranking, enabling an effective selection with controllable computational resources; the resulting subset of pseudo-labels serves as additional material for fine-tuning the source ASR models. Experiments on common benchmarks in three low-resource languages demonstrate the effectiveness of the proposed approach, yielding up to a 3-point reduction in WER.
Privacy-Preserving Information Extraction with Local LLMs: A Comparative Study on Dutch Debt Collection Letters
Beyza Celep | Natalia Amat-Lefort | Joost Visser
Beyza Celep | Natalia Amat-Lefort | Joost Visser
For individuals in financial distress, understanding debt collection letters is critical. These documents are often unstructured, use complex legal language, and contain highly sensitive personal data. Automating information extraction is essential for assisting caseworkers, who currently perform this task manually; a slow and error-prone process. The sensitive nature of this data requires efficient, privacy-preserving, locally-deployed solutions. This paper compares the feasibility of various local NLP models for this task. We evaluated a feature-engineered Conditional Random Field (CRF), a fine-tuned spaCy NER model, and several Large Language Models (LLMs) (1.1B to 14B parameters) on a new synthetic dataset of 1,000 Dutch debt letters. Models were compared using accuracy (F1-score) and deployment metrics (CPU runtime, memory usage). Our results show a clear performance-resource trade-off. Lightweight CRF and spaCy models efficiently extracted structured data but failed in many critical unstructured fields. In contrast, LLM performance scaled directly with model size. The 14B DeepSeek model achieved the highest accuracy (95.2% average F1), successfully handling all field types. In conclusion, larger local LLMs are the most viable solution for accurate, private document processing. Alternatively, a hybrid approach using lightweight models for structured data and LLMs only for complex, unstructured fields, would also be adequate.
Forewarned Is Forearmed: When Non-Sequential Embedding Turns into an Anomaly Detector
Elys Allesiardo | Antoine Caubrière | Valentin Vielzeuf
Elys Allesiardo | Antoine Caubrière | Valentin Vielzeuf
This paper offers an in-depth analysis of non-sequential multimodal sentence-level embeddings, with a particular focus on the SONAR model. We demonstrate that certain embedding dimensions are sensitive to perturbations and can serve as indicators of decoding anomalies. By leveraging the consistency between successive encoding and decoding, we successfully build an accurate detector. Additionally, we explore modifying specific dimensions of interest to attempt to correct them. This work underscores the importance of understanding and analyzing the embeddings themselves to enhance the reliability of multimodal representations.
A Joint Detection Framework for Latvian Loanwords and Calques Using Monolingual Data
Yelingyun Zhang | Atis Kapenieks | Marina Platonova
Yelingyun Zhang | Atis Kapenieks | Marina Platonova
Lexical borrowing is pervasive across languages with extensive cultural contact, yet its automatic detection remains challenging for low-resource languages, especially regarding calques. Existing methods depend heavily on bilingual resources and focus almost exclusively on phonological loanwords, leaving structural borrowing phenomena like calques largely unaddressed by automated tools. This paper proposes a novel joint binary classification pipeline based solely on monolingual data and mBERT, introducing the first large-scale annotated Latvian borrowing dataset with over 3,000 manually labeled entries across three categories: loanwords, calques, and local words. The pipeline adopts a staged decision process grounded in language contact theory, separating surface-level loanwords before tackling the more ambiguous calque category. Experiments demonstrate that our semi-supervised strategy with pseudo-labeling achieves a macro-F1 of 0.854 on an external test set, outperforming both a direct three-way classifier and a GPT-4o zero-shot baseline. These results establish a performance benchmark for the previously unaddressed task of automatic borrowing detection in Latvian, providing empirical tools for borrowing detection in resource-scarce contexts.
Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
Phuong-Hang Le | Valentin Pelloin | Arnault Chatelain | Maryem Bouziane | Mohammed Ghennai | Qianwen Guan | Kirill Milintsevich | Salima Mdhaffar | Aidan Mannion | Nils Defauw | Shuyue Gu | Alexandre Daniel Audibert | Marco Dinarelli | Yannick Estève | Lorraine Goeuriot | Steffen Lalande | Nicolas Hervé | Maximin Coavoux | François Portet | Étienne Ollion | Marie Candito | Maxime Peyrard | Solange Rossato | Benjamin Lecouteux | Aurélie Nardy | Gilles Sérasset | Vincent Segonne | Solène Evain | Diandra Fabre | Didier Schwab
Phuong-Hang Le | Valentin Pelloin | Arnault Chatelain | Maryem Bouziane | Mohammed Ghennai | Qianwen Guan | Kirill Milintsevich | Salima Mdhaffar | Aidan Mannion | Nils Defauw | Shuyue Gu | Alexandre Daniel Audibert | Marco Dinarelli | Yannick Estève | Lorraine Goeuriot | Steffen Lalande | Nicolas Hervé | Maximin Coavoux | François Portet | Étienne Ollion | Marie Candito | Maxime Peyrard | Solange Rossato | Benjamin Lecouteux | Aurélie Nardy | Gilles Sérasset | Vincent Segonne | Solène Evain | Diandra Fabre | Didier Schwab
We release Pantagruel models, a new family of self-supervised encoder models for French text and speech. Instead of predicting modality-tailored targets such as textual tokens or speech units, Pantagruel learns contextualized target representations in the feature space, allowing modality-specific encoders to capture linguistic and acoustic regularities more effectively. Separate models are pre-trained on large-scale French corpora, including Wikipedia, OSCAR and CroissantLLM for text, together with MultilingualLibriSpeech, LeBenchmark, and INA-100k for speech. INA-100k is a newly introduced 100,000-hour corpus of French audio derived from the archives of the Institut National de l’Audiovisuel (INA), the national repository of French radio and television broadcasts, providing highly diverse audio data. We evaluate Pantagruel across a broad range of downstream tasks spanning both modalities, including those from the standard French benchmarks such as FLUE or LeBenchmark. Across these tasks, Pantagruel models show competitive or superior performance compared to strong French baselines such as CamemBERT, FlauBERT, and LeBenchmark2.0, while maintaining a shared architecture that can seamlessly handle either speech or text inputs. These results confirm the effectiveness of feature-space self-supervised objectives for French representation learning and highlight Pantagruel as a robust foundation for multimodal speech-text understanding.
Merge and Conquer: Instructing Multilingual Models by Adding Target Language Weights
Eneko Valero | Maria Ribalta i Albado | Oscar Sainz | Naiara Perez | German Rigau
Eneko Valero | Maria Ribalta i Albado | Oscar Sainz | Naiara Perez | German Rigau
Large Language Models (LLMs) remain heavily centered on English, with limited performance in low-resource languages. Existing adaptation approaches, such as continual pre-training, demand significant computational resources. In the case of instructed models, high-quality instruction data is also required, both of which are often inaccessible for low-resource language communities. Under these constraints, model merging offers a lightweight alternative, but its potential in low-resource contexts has not been systematically explored. In this work, we explore whether it is possible to transfer language knowledge to an instruction-tuned LLM by merging it with a language-specific base model, thereby eliminating the need of language-specific instructions and repeated fine-tuning processes whenever stronger instructed variants become available. Through experiments covering four Iberian languages (Basque, Catalan, Galician, and Spanish) and two model families, we show that merging enables effective instruction-following behavior in new languages and even supports multilingual capability through the combination of multiple language-specific models. Our results indicate that model merging is a viable and efficient alternative to traditional adaptation methods for low-resource languages, achieving competitive performance while greatly reducing computational cost.
SemiAdapt: Semi-Supervised and Efficient LoRA-Based Domain Adaptation for Low-Resource Irish Machine Translation with Transformers
Josh Mcgiff | Nikola S. Nikolov
Josh Mcgiff | Nikola S. Nikolov
Fine-tuning is widely used to adapt multilingual Transformer models for machine translation (MT) in specific domains. However, full-parameter fine-tuning of large multilingual models with billions of parameters is computationally expensive, thus creating a barrier to entry for researchers working on low-resource tasks such as Irish translation. Parameter-efficient fine-tuning (PEFT) addresses this by updating a fraction of the original model parameters, with the Low-Rank Adaptation approach (LoRA) introducing small, trainable adapter layers. We introduce SemiAdapt-Full and SemiAdapt-LoRA as semi-supervised approaches that leverage inferred domains to improve overall performance in MT. SemiAdapt-LoRA employs dynamic routing at inference time, eliminating the need to load multiple separately fine-tuned models. Instead, a single shared base model is maintained while lightweight domain-specific adapters, updating only 1.39% of the model parameters in our case, are activated dynamically. We demonstrate that SemiAdapt-Full can outperform full-model fine-tuning and SemiAdapt-LoRA can propel PEFT methods to compete with full-model fine-tuning. We further evaluate corpus-level domain fine-tuning and demonstrate that our embedding-based inference methods perform especially well on larger and noisier corpora. Code and training configurations are released to support reproducibility. Ultimately, our approach narrows the performance gap between PEFT and full-parameter fine-tuning, offering resource-constrained researchers a computationally efficient alternative.
Data Selection Effects on Self-Supervised Learning of Audio Representations for French Audiovisual Broadcasts
Valentin Pelloin | Lina Bekkali | Reda Dehak | David Doukhan
Valentin Pelloin | Lina Bekkali | Reda Dehak | David Doukhan
Audio and speech self-supervised encoder models are now widely used for a lot of different tasks. Many of these models are often trained on clean segmented speech content such as LibriSpeech. In this paper, we look into how the pretraining datasets of such SSL (Self-Supervised Learning) models impact their downstream results. We build a large pretraining corpus of highly diverse TV and Radio broadcast audio content, which we describe with automatic tools. We use these annotations to build smaller subsets, which we use to train audio SSL models. Then, we evaluate the models on multiple downstream tasks such as automatic speech recognition, voice activity and music detection, or speaker recognition. The results show the potential of pretraining SSL models on diverse audio content without restricting it to speech. We also perform a membership inference attack to evaluate the encoder ability to memorize their training datasets, which highlight the importance of data deduplication. This unified training could bridge speech and music machine learning communities.
SENS-ASR: Semantic Embedding Injection in Neural-transducer for Streaming Automatic Speech Recognition
Youness Dkhissi | Valentin Vielzeuf | Elys Allesiardo | Anthony Larcher
Youness Dkhissi | Valentin Vielzeuf | Elys Allesiardo | Anthony Larcher
Many Automatic Speech Recognition (ASR) applications require streaming processing of the audio data. In streaming mode, ASR systems need to start transcribing the input stream before it is complete, i.e., the systems have to process a stream of inputs with a limited (or no) future context. Compared to offline mode, this reduction of the future context degrades the performance of Streaming-ASR systems, especially while working with low-latency constraint. In this work, we present SENS-ASR, an approach to enhance the transcription quality of Streaming-ASR by reinforcing the acoustic information with semantic information. This semantic information is extracted from the available past frame-embeddings by a context module. This module is trained using knowledge distillation from a sentence embedding Language Model fine-tuned on the training dataset transcriptions. Experiments on standard datasets show that SENS-ASR significantly improves the Word Error Rate on small-chunk streaming scenarios.
Efficient Financial Language Understanding via Distillation with Synthetic Data
Wen-Fong (Xavier) Huang | Edwin Simpson
Wen-Fong (Xavier) Huang | Edwin Simpson
Large instruction-following models are powerful but costly to deploy, particularly in finance, where labelled data are limited by confidentiality and expert annotation cost. We present an efficient framework for financial sentiment analysis through distillation with synthetic data, transferring knowledge from a large instruction-tuned teacher to compact student models. The framework is designed for low-resource conditions, where a small set of real examples are collected and labelled by hand. The framework then clusters the examples and uses the clusters to select seeds for generating synthetic examples via structured few-shot prompting. Experiments show that clustering-based seed selection yields more representative synthetic data than random sampling, enabling compact models to achieve strong performance with minimal supervision. Notably, on a more complex and noisy text domain, the compact model trained on the complete synthetic–seed corpus even outperforms the teacher model, while remaining competitive on formal text. The framework provides a practical route toward resource-efficient domain adaptation in financial NLP with minimal human labelling effort.
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
Aditya Kamlesh Parikh | Cristian Tejedor-García | Catia Cucchiarini | Helmer Strik
Aditya Kamlesh Parikh | Cristian Tejedor-García | Catia Cucchiarini | Helmer Strik
Reliable and interpretable automated assessment of second-language (L2) speech remains a central challenge, as large speech-language models (SpeechLLMs) often struggle to align with the nuanced variability of human raters. To address this, we introduce a rubric-guided reasoning framework that explicitly encodes multi-aspect human assessment criteria: accuracy, fluency, and prosody, while calibrating model uncertainty to capture natural rating variability. We fine-tune the Qwen2-Audio-7B-Instruct model using multi-rater human judgments and develop an uncertainty-calibrated regression approach supported by conformal calibration for interpretable confidence intervals. Our Gaussian uncertainty modeling and conformal calibration approach achieves the strongest alignment with human ratings, outperforming regression and classification baselines. The model reliably assesses fluency and prosody while highlighting the inherent difficulty of assessing accuracy. Together, these results demonstrate that rubric-guided, uncertainty-calibrated reasoning offers a principled path toward trustworthy and explainable SpeechLLM-based speech assessment.
Leveraging Semi-Supervised Learning for Multimodal Hate Speech Data Annotation and Detection
Rathi Adarshi Rammohan | Zhao Ren | Dominik Puchała | Aleksandra Świderska | Dennis Küster | Tanja Schultz
Rathi Adarshi Rammohan | Zhao Ren | Dominik Puchała | Aleksandra Świderska | Dennis Küster | Tanja Schultz
While the Internet and social media have fundamentally transformed our lives, they can also rapidly spread hate speech, i.e., derogatory statements targeting individuals or groups based on their immutable characteristics. Automatic detection systems could help limit this harmful phenomenon. However, the lack of large-scale annotated datasets remains a major bottleneck for developing better algorithms. In this work, we employ semi-supervised learning (SSL) to leverage the advantages of limited labeled data alongside large amounts of unlabeled data. We apply three SSL approaches, Fix-match, Full-match, and All-match learning, to enhance the performance of end-to-end pre-trained speech and text models for hate speech detection. Our findings indicate that SSL methods enhance the performance, achieving F1 scores of 0.851 on speech, 0.957 on text, and 0.959 with multimodal fusion. Furthermore, we analyze the impact of different weak augmentation strategies on labeled data and assess the quality of generated pseudo-labels to evaluate their potential use in data annotation.
Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization
Yiming Liang | Fang Zhao
Yiming Liang | Fang Zhao
Recent years have seen growing interest in applying neural networks and contextualized word embeddings to the parsing of historical languages. However, most advances have focused on dependency parsing, while constituency parsing for low-resource historical languages like Middle Dutch has received little attention. In this paper, we adapt a transformer-based constituency parser to Middle Dutch, a highly heterogeneous and low-resource language, and investigate methods to improve both its in-domain and cross-domain performance. We show that joint training with higher-resource auxiliary languages increases F1 scores by up to 0.73, with the greatest gains achieved from languages that are geographically and temporally closer to Middle Dutch. We further evaluate strategies for leveraging newly annotated data from additional domains, finding that fine-tuning and data combination yield comparable improvements, and our neural parser consistently outperforms the currently used PCFG-based parser for Middle Dutch. We further explore feature-separation techniques for domain adaptation and demonstrate that a minimum threshold of approximately 200 examples per domain is needed to effectively enhance cross-domain performance.
Quantifying the Accuracy and Cost Impact of Design Decisions in Budget-Constrained Agentic LLM Search
Kyle A. McCleary | James M. Ghawaly
Kyle A. McCleary | James M. Ghawaly
Agentic Retrieval-Augmented Generation (RAG) systems combine iterative search, planning prompts, and retrieval backends, but deployed settings impose explicit budgets on tool calls and completion tokens. We present a controlled measurement study of how search depth, retrieval strategy, and completion budget affect accuracy and cost under fixed constraints. Using Budget-Constrained Agentic Search (BCAS), a model-agnostic evaluation harness that surfaces remaining budget and gates tool use, we run comparisons across six LLMs and three question-answering benchmarks. Across models and datasets, accuracy improves with additional searches up to a small cap, hybrid lexical and dense retrieval with lightweight re-ranking produces the largest average gains in our ablation grid, and larger completion budgets are most helpful on HotpotQA-style synthesis. These results provide practical guidance for configuring budgeted agentic retrieval pipelines and are accompanied by reproducible prompts and evaluation settings.
Reason-to-Learn (R2L): Multi-Agent Knowledge Distillation for Lightweight LLMs in Sentiment Analysis
Le-Huy Tu | Quan Nguyen | Vincent NGUYEN | Johanna Bjorklund | Xuan-Son Vu
Le-Huy Tu | Quan Nguyen | Vincent NGUYEN | Johanna Bjorklund | Xuan-Son Vu
Large Language Models (LLMs) boast remarkable capabilities but face deployment challenges due to computational demands. We introduce Reason-to-Learn (R2L), a novel multi-agent collaborative knowledge distillation framework enabling small LLMs to learn from a distributed system of specialized agent models. Our architecture employs multiple autonomous teacher agents, each with distinct expertise and reasoning capabilities, coordinated by a meta-agent that orchestrates knowledge synthesis and conflict resolution. Unlike prior methods, our flexible four-phase process (Detection, Processing, Rationale Generation, Aggregation) leverages agent-based communication protocols and consensus mechanisms for cross-architecture knowledge transfer, demonstrated primarily on Vietnamese sentiment analysis. Experimental results are definitive: our lightweight R2L-Students (1-1.5B) consistently outperform the individual specialized agents (Qwen32B, Llama70B) and the GPT-4o meta-agent coordinator, especially on complex ABSA tasks. Ablation studies confirm our multi-agent collaborative approach outperformed traditional fine-tuning and single-agent distillation. Furthermore, R2L enhance generalizability of lightweight LLMs: our Vietnamese-trained student achieves strong zero-shot cross-lingual performance on Swedish ABSA (Svensk ABSAbank-Imm), with Krippendorff’s Alpha scores competitive with the specialized agents. R2L offers an efficient path to compact, high-performing specialist models through coordinated multi-agent learning.
PRiSM: Partial Ranking via Inter-layer Semantic Measurement for Efficient Fine-tuning of Language Models
Aldrin Kabya Biswas | Md Fahim | Md. Ashraful Amin | Amin Ahsan Ali | AKM Mahbubur Rahman
Aldrin Kabya Biswas | Md Fahim | Md. Ashraful Amin | Amin Ahsan Ali | AKM Mahbubur Rahman
The growing scale of pre-trained language models poses a challenge in fine-tuning for downstream tasks, especially in resource-constrained settings. Recent studies highlight that not all layers in transformer-based language models contribute equally to downstream task performance, giving rise to various partial fine-tuning strategies. However, current methods often introduce significant training overhead or rely on simple heuristics that yield suboptimal performance and poor generalization. We propose PRiSM (Partial Ranking via inter-layer Semantic Measurement), a training-free approach for layer-wise partial fine-tuning that leverages the cosine similarity between pre-trained aggregate token representations across layers to identify inter-layer relationships. comprises two stages: (i) scoring layers based on their relevance to the task via a single forward pass, and (ii) fine-tuning a subset of block-wise highest-scoring layers, while keeping others frozen. We conduct experiments on 15 diverse NLP datasets, including single-sentence and sentence-pair classification tasks. Our method achieves competitive performance compared to full fine-tuning, with an average training speedup of 1.5× and a reduction of trainable parameters by 75%, and outperforms all the comparative baselines. Additionally, our approach does not cause any notable drop in performance when the domain is changed for the evaluation tasks, demonstrating robust cross-domain generalizability.
SEFL: A Framework for Generating Synthetic Educational Assignment Feedback with LLM Agents
Mike Zhang | Amalie Pernille Dilling | Léon Gondelman | Niels Erik Ruan Lyngdorf | Euan D. Lindsay | Johannes Bjerva
Mike Zhang | Amalie Pernille Dilling | Léon Gondelman | Niels Erik Ruan Lyngdorf | Euan D. Lindsay | Johannes Bjerva
Providing high-quality feedback on student assignments is crucial for student success, but it is heavily limited by time and budgetary constraints. In this work, we introduce Synthetic Educational Feedback Loops (SEFL), a synthetic data framework designed to generate data that resembles immediate, on-demand feedback at scale without relying on extensive, real-world student assignments and teacher feedback. To obtain this type of data, two large language models (LLMs) operate in a teacher-student role to simulate assignment completion and formative feedback, generating 19.8K synthetic pairs of student work and corresponding critiques and actionable improvements from a teacher. With this data, we fine-tune smaller, more computationally efficient LLMs on these synthetic pairs, enabling them to replicate key features of high-quality, goal-oriented feedback. Through comprehensive evaluations with three LLM judges and three human experts, across a subset of 900 outputs, we demonstrate that SEFL-tuned models outperform both their untuned counterparts and an existing baseline in terms of feedback quality. The potential for societal impact is reinforced by extensive qualitative comments and ratings from human stakeholders — both students and higher education instructors. SEFL has the potential to transform feedback processes for higher education and beyond.
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
Hailay Kidu Teklehaymanot | Dren Fazlija | Wolfgang Nejdl
Hailay Kidu Teklehaymanot | Dren Fazlija | Wolfgang Nejdl
Adapting pretrained language models to low-resource, morphologically rich languages remains a significant challenge. Existing vocabulary expansion methods typically rely on arbitrarily segmented subword units, resulting in fragmented lexical representations and loss of critical morphological information. To address this limitation, we propose the Lexically Grounded Subword Embedding Initialization (LGSE) framework, which introduces morphologically informed segmentation for initializing embeddings of novel tokens. Instead of using random vectors or arbitrary subwords, LGSE decomposes words into their constituent morphemes and constructs semantically coherent embeddings by averaging pretrained subword or FastText-based morpheme representations. When a token cannot be segmented into meaningful morphemes, its embedding is constructed using character n-gram representations to capture structural information. During Language-Adaptive Pretraining, we apply a regularization term that penalizes large deviations of newly introduced embeddings from their initialized values, preserving alignment with the original pretrained embedding space while enabling adaptation to the target language. To isolate the effect of initialization, we retain the original pre-trained model vocabulary and tokenizer and update only the new embeddings during adaptation. We evaluate LGSE on three NLP tasks: Question Answering, Named Entity Recognition, and Text Classification, in two morphologically rich, low-resource languages: Amharic and Tigrinya, where morphological segmentation resources are available. Experimental results show that LGSE consistently outperforms baseline methods across all tasks, demonstrating the effectiveness of morphologically grounded embedding initialization for improving representation quality in underrepresented languages. Project resources are available1.
A Cheap Lunch: Synthetic Annotation With Reduced Human Effort for Medical Text Mining
Shutao Chen | Piek T.J.M. Vossen
Shutao Chen | Piek T.J.M. Vossen
Electronic Health Records are rich resources of patient knowledge and information among which knowledge about the functioning of patients as defined in the International Classification of Functioning (ICF) by the WHO. However, the patient notes have yet to be explored as the knowledge is packaged in sometimes cryptic language exchanged between caretakers. Recent research started to use NLP techniques to extract this knowledge but often requires laborious annotation. In this paper, we report on how the annotation can (partly) be done by a generative LLM, both for ICF categories that were previously manually annotated and for new ICF categories for which there was no annotation. We show that a domain specific encoder finetuned with both manual and synthetic annotations outperforms finetuning with just the manual annotations on a dedicated test set that was adapted for the new categories with minimal manual effort. We also assessed the quality of the synthetic annotations of the training data. Our process shows how competitive text classifiers for medical text mining can be developed and extended to new categories with minimal manual effort by experts.
Active Few-Shot Learning (AFSL) is an effective paradigm for improving the performance of large language models under limited annotation budgets. To address the inefficiency of conventional fine-tuning objectives in AFSL, this paper proposes a supervised contrastive fine-tuning framework specifically designed for natural language processing (NLP) text classification tasks. By integrating Supervised Contrastive Learning (SCL) with Hard Negative Mining (HNM), the proposed framework optimizes the embedding space through an enhanced hybrid loss function, thereby improving the utilization efficiency of labeled samples. Extensive experiments on five benchmark datasets show that, under a fixed state-of-the-art (SOTA) query strategy, our method consistently outperforms baseline models in text classification performance, and exhibits strong generalizability across different backbone architectures and acquisition functions. These findings demonstrate that optimizing how to learn—through improved learning objectives—provides a complementary direction to existing query strategies in advancing AFSL.
Simulating Student Interactions for Virtual Pretesting with In-Context Learning
Arthur Thuy | Luca Benedetto | Ekaterina Loginova | Dries F. Benoit
Arthur Thuy | Luca Benedetto | Ekaterina Loginova | Dries F. Benoit
Recent research has experimented with using Large Language Models (LLMs) for simulating student responses to exam questions. This approach, known as virtual pretesting, potentially offers a scalable alternative to traditional pretesting, which is costly and time-intensive, by enabling the creation of datasets of virtual students’ responses. Prior studies focused on zero-shot role-playing, prompting one LLM to imitate students of different levels, but showed limited alignment with response patterns of real students. This work introduces a framework that improves the alignment of LLM-based student simulations through in-context learning (ICL), leveraging previous question-answer records to provide the model with richer information about students’ skills and misconceptions. Our experiments show that not all models can leverage the additional contextual information. However, a multi-model approach, which combines simulations from several models, significantly improves alignment of the simulated responses when provided with relevant context: we observe a reduction of up to 30% in difficulty estimation RMSE with respect to the non contextual and individual contextual models. Overall, our findings indicate that LLMs can be used with ICL to create synthetic datasets of student responses approximating some patterns of learner behavior, however their ability to align with authentic student performance remains limited.
An Exploration-Analysis-Disambiguation Reasoning Framework for Word Sense Disambiguation with Low-Parameter LLMs
Deshan Koshala Sumanathilaka | Nicholas Micallef | Julian Hough
Deshan Koshala Sumanathilaka | Nicholas Micallef | Julian Hough
Word Sense Disambiguation (WSD) remains a key challenge in Natural Language Processing (NLP), especially when dealing with rare or domain-specific senses that are often misinterpreted. While modern high-parameter Large Language Models (LLMs) such as GPT-4-Turbo have shown state-of-the-art WSD performance, their computational and energy demands limit scalability. This study investigates whether low-parameter LLMs (<4B parameters) can achieve comparable results through fine-tuning strategies that emphasize reasoning-driven sense identification. Using the FEWS dataset augmented with semi-automated, rationale-rich annotations, we fine-tune eight small-scale open-source LLMs (e.g. Gemma and Qwen). Our results reveal that Chain-of-Thought (CoT)-based reasoning combined with neighbour-word analysis achieves performance comparable to GPT-4-Turbo in zero-shot settings. Importantly, Gemma-3-4B and Qwen-3-4B models consistently outperform all medium-parameter baselines and state-of-the-art models on FEWS, with robust generalization to unseen senses. Furthermore, evaluation on the unseen "Fool Me If You Can” dataset confirms strong cross-domain adaptability without task-specific fine-tuning. This work demonstrates that with carefully crafted reasoning-centric fine-tuning, low-parameter LLMs can deliver accurate WSD while substantially reducing computational and energy demands.
Building Effective Japanese Medical LLMs with an Open Recipe for Domain Adaptation through Continued Pre-training
Akiko Aizawa | Yuki Arase | Fei Cheng | Jiahao Huang | Zhiyi Huang | Junfeng Jiang | Teruhito Kanazawa | Daisuke Kawahara | Kazuma Kobayashi | Takashi Kodama | Sadao Kurohashi | Yusuke Oda | Yuma Tsuta | Zhen Wan | Zhishen Yang | Rio Yokota
Akiko Aizawa | Yuki Arase | Fei Cheng | Jiahao Huang | Zhiyi Huang | Junfeng Jiang | Teruhito Kanazawa | Daisuke Kawahara | Kazuma Kobayashi | Takashi Kodama | Sadao Kurohashi | Yusuke Oda | Yuma Tsuta | Zhen Wan | Zhishen Yang | Rio Yokota
In high-stakes domains such as medicine, ensuring transparency of the training corpus is essential, with careful consideration of local healthcare landscapes; however, the majority of existing medical large language models (LLMs) have not disclosed the details of their training corpora. Here, we introduce an open recipe for domain adaptation of LLMs to the Japanese medical domain. We employed fully open-source Japanese general-domain LLMs as base models, whose pre-training datasets are also disclosed. To establish effective corpora for domain adaptation through continued pre-training, we started with small-scale medical datasets and ultimately constructed a medical corpus consisting of 79.6B tokens, incorporating local clinical guidelines, medical textbooks, and other domain-specific resources. The resulting LLM from continued pre-training, namely SIP-med-llm-8x13B, with an active parameter count of 22B, demonstrated favorable accuracy on benchmarks including the Japanese National Medical Examination. This performance was comparable to that of 70B-parameter open-weight models whose construction details remain non-transparent. This represents the first case in the Japanese medical field where complete corpus details have been disclosed for fully from-scratch development, providing important insights for future efforts to construct medical LLMs tailored to the specific characteristics of local contexts. The model is available publicly at this Hugging Face repository: https://huggingface.co/SIP-med-LLM/SIP-jmed-llm-2-8x13b-OP-instruct.
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
Julia Wunderle | Anton Ehrmanntraut | Jan Pfister | Fotis Jannidis | Andreas Hotho
Julia Wunderle | Anton Ehrmanntraut | Jan Pfister | Fotis Jannidis | Andreas Hotho
Encoders remain essential for efficient German NLP and NLU scenarios despite the rise of decoder-only LLMs. This work studies two routes to high-quality German encoders under identical data and training constraints: a) training from scratch and b) converting decoders via LLMVec. We introduce two resources: ModernGBERT (134M, 1B), fully transparent German encoders in the ModernBERT style, and LLäMmleinVec (120M, 1B, 7B), decoder-to-encoder conversions trained with masked next-token prediction, both undergoing a context extension to 8192 tokens. Across SuperGLEBer, ModernGBERT 1B sets a new state of the art (avg 0.808), surpassing GBERTlarge (+4%) and the seven-times larger converted 7B model (0.787). On German MTEB after supervised fine-tuning, ModernGBERT 1B (0.551) approaches the converted 7B model (0.557). We release all models, checkpoints, datasets, and full training records, and introduce an encoder-adapted QA-NIAH evaluation. All in all, our results provide actionable guidance: when parameter efficiency and latency matter, from-scratch encoders dominate. When a pre-trained decoder exists and compute is a limited, conversion offers an effective alternative.
Arabic ChartSumm: An English-to-Arabic Benchmark for Metadata-to-Text Summarization
Passant Elchafei | Amany Fashwan
Passant Elchafei | Amany Fashwan
Generating summaries from chart metadata in Arabic presents unique challenges at the intersection of cross-lingual transfer and data-to-text generation. Chart-to-text benchmarks have advanced English-language research, yet Arabic remains without a comparable resource, underscoring its continued underrepresentation in NLP. To cover this gap, we construct the first Arabic ChartSumm benchmark by translating chart metadata and reference summaries from English into Modern Standard Arabic (MSA). Two high-quality machine translation models with contrasting architectures are employed: NLLB-200-distilled-600M, designed for low-resource coverage, and Qwen2.5-1.5B, an open large language model with general multilingual capabilities. A central contribution of this work is a translation quality evaluation that systematically assesses both systems using BLEU, chrF, COMET_ref, and COMET_QE metrics against a Google-Translate Arabic pivot. Results demonstrate that NLLB achieves markedly higher lexical and semantic fidelity. Building on this foundation, we fine-tune two models, mT5 (multilingual) and CAMeL-Lab’s AraBART (Arabic-specific), to generate Arabic summaries from structured chart metadata. Experimental results show that AraBART trained on NLLB translations outperforms other configurations, achieving ROUGE-L = 63.8 and BLEU = 33.1, highlighting the strong dependency of downstream summarization quality on translation accuracy and demonstrating its superior capacity for Arabic generation.
Introducing a Bangla Sentence – Gloss Pair Dataset for Bangla Sign Language Translation and Research
Neelavro Saha | Rafi Shahriyar | Nafis Ashraf Roudra | Saadman Sakib | Annajiat Alim Rasel
Neelavro Saha | Rafi Shahriyar | Nafis Ashraf Roudra | Saadman Sakib | Annajiat Alim Rasel
Bangla Sign Language (BdSL) translation represents a low-resource NLP task due to the lack of large-scale datasets that address sentence-level translation. Correspondingly, existing research in this field has been limited to word and alphabet level detection. In this work, we introduce Bangla-SGP, a novel parallel dataset consisting of 1,000 human-annotated sentence–gloss pairs which was augmented with around 3,000 synthetically generated pairs using syntactic and morphological rules through a rule-based Retrieval-Augmented Generation (RAG) pipeline. The gloss sequences of the spoken Bangla sentences are made up of individual glosses which are Bangla sign supported words and serve as an intermediate representation for a continuous sign. Our dataset consists of 1000 high quality Bangla sentences that are manually annotated into a gloss sequence by a professional signer. The augmentation process incorporates rule-based linguistic strategies and prompt engineering techniques that we have adopted by critically analyzing our human annotated sentence-gloss pairs and by working closely with our professional signer. Furthermore, we fine-tune several transformer-based models such as mBart50, Google mT5, GPT4.1-nano and evaluate their sentence-to-gloss translation performance using BLEU scores, based on these evaluation metrics we compare the model’s gloss-translation consistency across our dataset and the RWTH-PHOENIX-2014T benchmark.
Language Models as Semantic Augmenters for Sequential Recommenders
Mahsa Valizadeh | Xiangjue Dong | Rui Tuo | James Caverlee
Mahsa Valizadeh | Xiangjue Dong | Rui Tuo | James Caverlee
Large Language Models (LLMs) excel at capturing latent semantics and contextual relationships across diverse modalities. However, in modeling user behavior from sequential interaction data, performance often suffers when such semantic context is limited or absent. We introduce LaMAR, a LLM-driven semantic enrichment framework designed to enrich such sequences automatically. LaMAR leverages LLMs in a few-shot setting to generate auxiliary contextual signals by inferring latent semantic aspects of a user’s intent and item relationships from existing metadata. These generated signals, such as inferred usage scenarios, item intents, or thematic summaries, augment the original sequences with greater contextual depth. We demonstrate the utility of this generated resource by integrating it into benchmark sequential modeling tasks, where it consistently improves performance. Further analysis shows that LLM-generated signals exhibit high semantic novelty and diversity, enhancing the representational capacity of the downstream models. This work represents a new data-centric paradigm where LLMs serve as intelligent context generators, contributing a new method for the semi-automatic creation of training data and language resources.
Efficient Adaptation of English Language Models for Morphologically Rich and Underrepresented Languages: The Case of Arabic
Ahmed Samy Eldamaty | Mohamed Maher Zenhom Abdelrahman | Mohamed Mostafa Ibrahim Elbehery | Mariam Ashraf | Radwa Elshawi
Ahmed Samy Eldamaty | Mohamed Maher Zenhom Abdelrahman | Mohamed Mostafa Ibrahim Elbehery | Mariam Ashraf | Radwa Elshawi
Transformer-based language models have revolutionized NLP, yet their adaptation to morphologically rich and dialectally diverse languages such as Arabic remains non-trivial. We introduce ModernAraBERT, a resource-efficient adaptation of the English-pretrained ModernBERT for Arabic, employing continued pretraining on large Arabic corpora followed by lightweight head-only fine-tuning with a frozen encoder. This strategy retains cross-lingual knowledge while capturing Arabic morphology and orthographic variation, offering a scalable alternative to training monolingual models from scratch. We evaluate ModernAraBERT on three representative Arabic NLP tasks, sentiment analysis, named entity recognition, and extractive question answering, against strong Arabic-specific and multilingual baselines (AraBERTv1, AraBERTv2, MARBERT, mBERT). Across all tasks, ModernAraBERT achieves consistent and often substantial improvements, particularly for sentence and token-level understanding, demonstrating that modern English encoder architectures can be efficiently transferred to Arabic through language-adaptive pretraining. Beyond Arabic, our findings highlight a generalizable paradigm for extending state-of-the-art models to morphologically complex and underrepresented languages with reduced computational overhead.
GhostWriter: Hidden AI-Generated Texts over Multiple Languages, Domains and Generators
Manuel Schaaf | Kevin Bönisch | Alexander Mehler
Manuel Schaaf | Kevin Bönisch | Alexander Mehler
The advent of Transformer-based Large Language Models (LLMs) has led to an unprecedented surge of AI-generated text (AIGT) across online platforms and academic domains. While these models exhibit near-human fluency and stylistic coherence, their widespread adoption has raised concerns about authorship integrity, research quality, and the recursive contamination of training corpora with synthetic data. These developments underscore the need for reliable AIGT detection methods and benchmark datasets, particularly for malicious or deceptive ghostwriting scenarios where AIGT is intentionally crafted to evade detection. To address this, we present GhostWriter, a large-scale, bilingual (German and English), multi-generator, and multi-domain dataset for AIGT detection. The dataset comprises human- and AI-authored texts produced under domain-specific ghostwriting conditions, including examples intentionally embedded within otherwise human-written texts to obscure their AI origin. With GhostWriter, we (i) aim to expand the resources available for German AIGT datasets, (ii) emphasize mixed or fused synthesizations—since most existing corpora are limited to the document level—and (iii) introduce specifically crafted malicious ghostwriting scenarios across multiple domains and generators.
Using LLMs to Extract Instances of Schematic Constructions from Unannotated L2 Learner Corpora
Jelena Kallas | Ahto Kiil | Heete Sahkai | Geda Paulsen | Kertu Saul
Jelena Kallas | Ahto Kiil | Heete Sahkai | Geda Paulsen | Kertu Saul
Our previous study found that generative LLMs can be successfully used to identify instances of schematic constructions (as defined in Construction Grammar) in unannotated L1 corpus data. This study tests the applicability of LLMs to also identify instances of constructions in unannotated L2 data. L2 learner corpora are notoriously difficult to annotate and query since they contain errors. Using LLMs can thus simplify the retrieval of construction data from L2 corpora. The identification of instances of constructions in L2 learner data has many possible uses in pedagogical applications of Construction Grammar and constructicography, like the identification of error-prone (properties of) constructions and the distribution of constructional instances across CEFR levels. Using the Estonian Nominal Quantifier Construction as the example construction and an Estonian CEFR-graded learner corpus as the source of L2 data, we tested several prompts and several models (OpenAI’s o3-mini, o3, gpt-5-mini and gpt-5, Google DeepMind’s Gemini Flash 2.5, Anthropic’s Claude Sonnet 4.5 and Opus 4.1). We found that the best model, gpt-5, achieved F1-scores from 0.90 to 0.96, depending on the level of detail of the prompt.
Corruption-Based Data Augmentation for Arabic Essay Scoring: A Preliminary Study on the Organization Trait
May Saed Bashendy | Tamer Elsayed
May Saed Bashendy | Tamer Elsayed
Despite significant advances in Automated Essay Scoring (AES), progress in Arabic AES remains limited by the scarcity and imbalance of publicly available datasets. Manual curation of such data is labor-intensive and lacks scalability. To address this, we introduce COrE, a corruption-based data augmentation method that targets the organization trait of Arabic essays. COrE generates synthetic essays by intentionally disrupting the organization of well-written essays through controlled, distance-aware sentence swapping. Our experiments are conducted on TAQAE, a dataset of 620 essays across 4 distinct writing prompts. We evaluate the effectiveness of COrE using two widely-adopted pre-trained models: AraBERTv2 and CAMeLBERT-mix. Both models show improved performance with COrE, achieving gains of 9-17% over the no-augmentation baseline. These results highlight the potential of trait-specific augmentation to address data scarcity and enhance AES performance for low-resource languages.
Structured Prompting for Arabic Essay Proficiency: A Trait-Centric Evaluation Approach
Salim Al Mandhari | Hieu Pham Dinh | Mo El-Haj | Paul Rayson
Salim Al Mandhari | Hieu Pham Dinh | Mo El-Haj | Paul Rayson
This paper presents a novel prompt engineering framework for trait specific Automatic Essay Scoring (AES) in Arabic, leveraging large language models (LLMs) under zero-shot and few-shot configurations. Addressing the scarcity of scalable, linguistically informed AES tools for Arabic, we introduce a three-tier prompting strategy (standard, hybrid, and rubric-guided) that guides LLMs in evaluating distinct language proficiency traits such as organization, vocabulary, development, and style. The hybrid approach simulates multi-agent evaluation with trait specialist raters, while the rubric-guided method incorporates scored exemplars to enhance model alignment. In zero and few-shot settings, we evaluate eight LLMs on the QAES dataset, the first publicly available Arabic AES resource with trait level annotations. Experimental results using Quadratic Weighted Kappa (QWK) and Confidence Intervals show that Fanar-1-9B-Instruct achieves the highest trait level agreement in both zero and few-shot prompting (QWK = 0.28 and CI = 0.41), with rubric-guided prompting yielding consistent gains across all traits and models. Discourse-level traits such as Development and Style showed the greatest improvements. These findings confirm that structured prompting, not model scale alone, enables effective AES in Arabic. Our study presents the first comprehensive framework for proficiency oriented Arabic AES and sets the foundation for scalable assessment in low resource educational contexts.
While large general-purpose Transformer-based encoders excel at general language understanding, their performance diminishes in specialized domains like manufacturing due to a lack of exposure to domain-specific terminology and semantics. In this paper, we address this gap by introducing ManufactuBERT, a RoBERTa model continually pretrained on a large-scale corpus curated for the manufacturing domain. We present a comprehensive data processing pipeline to create this corpus from web data, involving an initial domain-specific filtering step followed by a multi-stage deduplication process that removes redundancies. Our experiments show that ManufactuBERT establishes a new state-of-the-art on a range of manufacturing-related NLP tasks, outperforming strong specialized baselines. More importantly, we demonstrate that training on our carefully deduplicated corpus significantly accelerates convergence, leading to a 33% reduction in training time and computational cost compared to training on the non-deduplicated dataset. The proposed pipeline offers a reproducible example for developing high-performing encoders in other specialized domains. Our model, code and curated corpus will be publicly available.
Śmigiel Dataset: Laying Foundations for Investigating Machine-Generated Text Detection in Polish
Jakub Strebeyko | Alina Wróblewska | Piotr Przybyła
Jakub Strebeyko | Alina Wróblewska | Piotr Przybyła
We present Śmigiel, the first open dataset for training and evaluating machine-generated text (MGT) in Polish. The dataset includes a collection of human-written text fragments from six domains, which are used to prompt text generation by eight language models capable of producing credible Polish text. In addition to the raw corpus of over 462K generated texts, we also release a cleaned source- and domain-balanced dataset suitable for training and evaluating MGT detectors. Finally, we conduct preliminary experiments with text classifiers, showing that task difficulty depends on the text domain, the generating language model, and the availability of similar data in training. The results indicate that MGT detection in Polish can be approached with general-purpose classifiers that generalize well to new LLMs, but struggle to adapt to genres not represented in the training data.
Extracting Medical Image-Related Entities from Spanish Electronic Health Records Using NER Methods
Alexander Platas | Marcos Merino | Elena Zotova | Montse Cuadros | Karen López-Linares | Mikel Pérez de Mendiola | María Gálvez | Cristina Barba | Antón Asla
Alexander Platas | Marcos Merino | Elena Zotova | Montse Cuadros | Karen López-Linares | Mikel Pérez de Mendiola | María Gálvez | Cristina Barba | Antón Asla
This paper presents a novel corpus in Spanish tailored for the extraction of medical image-related entities from radiological reports using Named Entity Recognition (NER) methods. The dataset was created by aggregating and refining multiple existing corpora, focusing on entities that can be visually interpreted in associated medical images. This resource aims to bridge the gap between natural language processing and computer vision in the biomedical domain. The study evaluates various NER methods, including encoder-only, encoder-decoder, and decoder-only architectures. It explores fine-tuning, zero-shot, and few-shot In-Context Learning (ICL) strategies to determine the most effective approach for entity extraction. The resulting dataset is publicly available.
A Novel Synthetic Dataset for Few-Shot Legal Relation Extraction in German
Shiva Banasaz Nouri | Elena Leitner | Julian Moreno-Schneider | Georg Rehm
Shiva Banasaz Nouri | Elena Leitner | Julian Moreno-Schneider | Georg Rehm
The legal domain is particularly challenging for natural language processing due to the personal and confidential information it contains. Despite the significant advances of large language models (LLMs), applying them to relation extraction (RE) in legal texts remains challenging, not only because of the task’s linguistic and semantic complexity, but also due to privacy, compliance, and infrastructure constraints under regulations such as the EU AI Act. To address these challenges, we propose a novel synthetic dataset for German legal relation extraction, created using LLMs through a controlled, privacy-preserving, template-based pipeline. The dataset allows for reproducible and legally compliant experimentation. We benchmark it using two few-shot learning paradigms, a description-enhanced Model-Agnostic Meta-Learning (MAML) framework and Prototypical Networks with supervised contrastive loss and curriculum-aware prototype enrichment. Our results demonstrate that combining few-shot learning with structured semantic knowledge achieves robust and interpretable results, with the curriculum-aware Proto-Contrastive model reaching an F1-score of 99.83%.
LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs
Tian Huang | Tom Bourgeade | Irina Illina
Tian Huang | Tom Bourgeade | Irina Illina
Objective Structured Clinical Examinations (OSCEs) are the standard method for assessing medical students’ clinical and communication skills through structured patient interviews. In France, however, the organization of training sessions is limited by human and logistical constraints, restricting students’ access to repeated practice and structured feedback. Recent advances in Natural Language Processing (NLP) and Large Language Models (LLMs) now offer the opportunity to automatically evaluate such medical interviews, thereby alleviating the need for human examiners during training. Yet, real French OSCE annotated transcripts remain extremely scarce, limiting reproducible research and reliable benchmarking. To address these challenges, we investigate the use of LLMs for both generating and evaluating French OSCE dialogues in a low-resource context. We introduce a controlled pipeline that produces synthetic doctor–patient interview transcripts guided by scenario-specific evaluation criteria, combining ideal and perturbed performances to simulate varying student skill levels. The resulting dialogues are automatically silver-labeled through an LLM-assisted framework supporting adjustable evaluation strictness. Benchmarking multiple open-source and proprietary LLMs shows that mid-size models (≤32B parameters) achieve accuracies comparable to GPT-4o (~90%) on synthetic data, highlighting the feasibility of locally deployable, privacy-preserving evaluation systems for medical education.
Instruction-Tuned Urdu LLMs: Efficient Adaptation of Llama Models and Evaluation Resources for Urdu
Munief Hassan Tahir | Sana Shams | Sarmad Hussain | Miriam Butt
Munief Hassan Tahir | Sana Shams | Sarmad Hussain | Miriam Butt
This paper presents UrduLLaMA 1.1 and UrduLLaMA 1.1 Tiny, two instruction-tuned large language models (LLMs) designed to advance natural language processing for Urdu, a low-resource language with limited representation in multilingual corpora. These instruction-tuned models are derived from Llama-3.1-8B-Instruct and Llama-3.2-3B-Instruct architectures, respectively by conducting continual pretraining on 800 million diverse Urdu tokens curated from public and proprietary sources, followed by Supervised Fine-Tuning (SFT) using LoRA on 432K Urdu instructions spanning diverse NLP tasks. Rigorous evaluation across 14 culturally-specific domains using our novel Urdu LLM Evaluation Dataset demonstrates superior performance. UrduLLaMA 1.1 achieves 65.3 average accuracy (GPT-5 Nano evaluation), outperforming its Llama-3.1-8B-Instruct base (50.7) across all categories and surpassing Llama-3.3-70B-Instruct (62.7) in 8 out of 14 domains. UrduLLaMA 1.1 Tiny transforms Llama-3.2-3B-Instruct (38.8) into a (61.2) performer. Human evaluation by native Urdu linguists confirms these gains (3.51/5 vs. 2.61/5 base). Our results validate targeted adaptation strategies combining continual pretraining with instruction tuning as computationally efficient solutions for low-resource languages, enabling state-of-the-art Urdu LLM models with accessible hardware.
Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus
Aidan Mannion | Cécile Macaire | Armand Violle | Stéphane Ohayon | Xavier Tannier | Didier Schwab | Lorraine Goeuriot | François Portet
Aidan Mannion | Cécile Macaire | Armand Violle | Stéphane Ohayon | Xavier Tannier | Didier Schwab | Lorraine Goeuriot | François Portet
Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains, yet their adaptation to specialized fields remains challenging, particularly for non-English languages. This study investigates domain-adaptive pre-training (DAPT) as a strategy for specializing small to mid-sized LLMs in the French biomedical domain through continued pre-training. We address two key research questions: the viability of specialized continued pre-training for domain adaptation and the relationship between domain-specific performance gains and general capability degradation. Our contributions include the release of a fully open-licensed French biomedical corpus suitable for commercial and open-source applications, the training and release of specialized French biomedical LLMs, and novel insights for DAPT implementation. Our methodology encompasses the collection and refinement of high-quality French biomedical texts, the exploration of causal language modeling approaches using DAPT, and conducting extensive comparative evaluations. Our results cast doubt on the efficacy of DAPT, in contrast to previous works, but we highlight its viability in smaller-scale, resource-constrained scenarios under the right conditions. Our findings further suggest that model merging post-DAPT is essential to mitigate generalization trade-offs, and in some cases even improves performance on specialized tasks at which the DAPT was directed.
TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation
Toms Bergmanis | Ingus Jānis Pretkalniņš | Martins Kronis | Davis Nicmanis | Jeļizaveta Jelinska | Roberts Rozis | Rinalds Vīksna | Marcis Pinnis
Toms Bergmanis | Ingus Jānis Pretkalniņš | Martins Kronis | Davis Nicmanis | Jeļizaveta Jelinska | Roberts Rozis | Rinalds Vīksna | Marcis Pinnis
Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in training data. This paper presents TildeOpen LLM, a 30-billion-parameter open-weight foundational model trained on 34 European languages to promote linguistic equity and improve performance for low-resource languages. To address the data imbalance, we combine dataset upsampling with a curriculum-based training schedule that alternates between uniform and natural language distributions. The resulting model performs favorably compared to other multilingual LLMs despite being trained with significantly fewer computing resources. Evaluation across multiple multilingual benchmarks shows that TildeOpen surpasses existing open-weight models in text generation and comprehension, particularly for Baltic, Finno-Ugric, and Slavic languages. Human evaluations confirm an up to tenfold reduction in linguistic errors relative to leading baselines. The model and associated resources are fully open-weight and publicly available at huggingface.co/TildeAI/TildeOpen-30b. These outcomes demonstrate that careful data curation and balanced training strategies can substantially enhance multilingual model quality without increasing model size or training volume.
Common Sense vs. Morality: The Curious Case of Narrative Focus Bias in LLMs
Saugata Purkayastha | Pranav Kushare | Pragya Paramita Pal | Sukannya Purkayastha
Saugata Purkayastha | Pranav Kushare | Pragya Paramita Pal | Sukannya Purkayastha
Large Language Models (LLMs) are increasingly deployed across diverse real-world applications and user communities. As such, it is crucial that these models remain both morally grounded and knowledge-aware. In this work, we uncover a critical limitation of current LLMs—their tendency to prioritize moral reasoning over commonsense understanding. To investigate this phenomenon, we introduce CoMoral, a novel benchmark dataset containing commonsense contradictions embedded within moral dilemmas. Through extensive evaluation of ten LLMs across different model sizes, we find that existing models consistently struggle to identify such contradictions without prior signal. Furthermore, we observe a pervasive narrative focus bias, wherein LLMs more readily detect commonsense contradictions when they are attributed to a secondary character rather than the primary (narrator) character. Our comprehensive analysis underscores the need for enhanced reasoning-aware training to improve the commonsense robustness of large language models.
“Emphasizing the Commendable”: A Study of Homogenized Transitive Verb Constructions in Machine Generated Peer Reviews
Hing-Yuet Fung | Chi-kiu Lo | Samuel Larkin
Hing-Yuet Fung | Chi-kiu Lo | Samuel Larkin
We present a study of machine generated text (MGT) output homogenization with a focus on the relative usage of the prototypical object construction of verbs (the O construction), which takes a noun phrase as its accusative argument. Verbs of different semantics have different tendencies of selecting a direct object or clausal complement; and hence lead to natural variation away from the prototypical usage. However, our results in the study between scientific peer reviews written by human and machines show a shift to unusually high usage of the O construction in MGT and greatly suppressing the frequency of other construction types. This is considered a serious case of syntactic homogenization. A major finding is that frequent verbs, like “emphasize”, appear top on the list of such homogenized syntactic construction. This is more striking than identifying disproportionately more frequent usage of naturally rare words such as “commendable” in previous work. Our results will contribute to the prevention of further homogenization of MGT before they merge deeper into the ecosystem of human-written text.
CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data Augmentation
Shuzhou Yuan | Willliam LaCroix | Hardik Ghoshal | Ercong Nie | Michael Färber
Shuzhou Yuan | Willliam LaCroix | Hardik Ghoshal | Ercong Nie | Michael Färber
Large Language Models (LLMs) are increasingly employed as AI tutors in education due to their scalability and potential for personalized instruction. However, off-the-shelf LLMs often underperform in educational settings, exhibiting limitations such as providing answers too readily, failing to adapt their responses to students’ uncertainty, and remaining susceptible to emotionally manipulative prompts. To address these challenges, we introduce CoDAE, a framework that adapts LLMs for educational use through Chain-of-Thought (CoT) data augmentation. We collect real-world dialogues between students and a ChatGPT-based tutor and enrich them using CoT prompting to promote step-by-step reasoning and pedagogically aligned guidance. Furthermore, we design targeted dialogue cases to explicitly mitigate three key limitations: over-compliance, low response adaptivity, and threat vulnerability. We fine-tune four open-source LLMs on different variants of the augmented datasets and evaluate them in simulated educational scenarios using both automatic metrics and LLM-as-a-judge assessments. Our results show that models fine-tuned with CoDAE deliver more pedagogically appropriate guidance, promote student reflection and more effectively prevent premature answer disclosure.
Synthetic Instruction Generation for Low-Resource Nordic Languages: Viability and Limitations in LLM Instruction-Tuning
Mathias Stenlund | Annika Simonsen | Lars Bungum | Jan Ebert | Jiangtao Wang | Oleg Filatov | Hemanadhan Myneni | Morris Riedel | Hafsteinn Einarsson
Mathias Stenlund | Annika Simonsen | Lars Bungum | Jan Ebert | Jiangtao Wang | Oleg Filatov | Hemanadhan Myneni | Morris Riedel | Hafsteinn Einarsson
Pretrained large language models (LLMs) gain instruction-following abilities through instruction-tuning, a method which relies on datasets of instruction–response pairs. However, for low-resource languages, collecting human-authored instructions is costly, raising the question of whether synthetic instructions can substitute human-authored instructions for non-English languages. We compare instruction-tuning of a smaller pretrained LLM in four Nordic languages using (a) human-authored instructions paired with synthetic responses and (b) fully synthetic instruction–response pairs generated with a minimal-effort pipeline. Native-speaker evaluations show that models instruction-tuned on synthetic instructions perform on par with those trained on human-authored instructions for the largest Nordic languages, suggesting that minimal-effort synthetic instructions can serve as a practical alternative. In contrast, response quality deteriorates sharply for Icelandic, underscoring the limitations of current synthetic data generation pipelines when the LLM competence in the target language is weak. Overall, our results highlight that while synthetic instructions can enable cost-efficient instruction-tuning for the largest Nordic languages, they remain insufficient for Icelandic, clarifying when minimal-effort synthetic approaches suffice and when they fall short.
AYN: A Tiny Yet Competitive Indian Legal Language Model Pretrained from Scratch
Mitodru Niyogi | Eric Gaussier | Arnab Bhattacharya
Mitodru Niyogi | Eric Gaussier | Arnab Bhattacharya
Decoder-only Large Language Models (LLMs) are currently the model of choice for many Natural Language Processing (NLP) applications. Through instruction fine-tuning and prompting approaches, such LLMs have been efficiently used to solve both general and domain-specific tasks. However, they are costly to train and, to a certain extent, costly to use as well, and one can wonder whether LLMs can be replaced by domain-specific Tiny Language Models (TLMs), which typically contain less than 100M parameters. We address this question in this study by comparing the performance of an 88M TLM pretrained from scratch for 185 A100 hours on a specific domain with a domain-specific tokenizer (here, the Indian legal domain) with LLMs of various sizes between 1B and 8B for solving domain-specific tasks. We show in particular that our legal TLM, Ayn, can indeed outperform LLMs up to 80 times larger on the legal case judgment prediction task, rival LLMs up to 30 times larger on the summarization task, and still be competitive with these larger LLMs on general tasks.
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
Eeham Khan | Firas Saidani | Owen Van Esbroeck | Richard Khoury | Leila Kosseim
Eeham Khan | Firas Saidani | Owen Van Esbroeck | Richard Khoury | Leila Kosseim
Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain largely confined to a small number of high-resource languages for which there is abundant training data. Recently, continual pre-training (CPT) has emerged as a means to fine-tune these models to low-resource regional dialects. In this paper, we study the use of CPT for dialect learning under tight data and compute budgets. Using low-rank adaptation (LoRA) and compute-efficient continual pre-training, we adapt three LLMs to the Québec French dialect using a very small dataset and benchmark them on the COLE suite. Our experiments demonstrate an improvement on the minority dialect benchmarks with minimal regression on the prestige language benchmarks with around 1% of model parameters updated. Analysis of the results demonstrate that gains are highly contingent on corpus composition. These findings indicate that CPT with parameter-efficient fine-tuning (PEFT) can narrow the dialect gap by providing cost-effective and sustainable language resource creation, expanding high-quality LLM access to minority linguistic communities. To support reproducibility and broaden access, we release the first Québec French LLMs on Hugging Face.
Reformulate and Create, Don’t Translate: Creating Natural Prompts for Underserved Languages
Annika Simonsen | Mathias Stenlund | Lars Bungum | Marc Daníel Skipstað Volhardt | Hafsteinn Einarsson
Annika Simonsen | Mathias Stenlund | Lars Bungum | Marc Daníel Skipstað Volhardt | Hafsteinn Einarsson
We present a methodology for creating high-quality instruction prompts for low-resource Germanic languages that addresses a critical challenge: small annotator pools risk producing datasets reflecting narrow individual interests rather than diverse user needs. In this work, native speakers reformulate existing English prompts from OpenAssistant or create entirely original prompts, adapting them to reflect local contexts and natural language patterns while preserving broad task and topic diversity. This approach produced high-quality prompt datasets totaling 6,950 prompts across seven Germanic languages (German, Dutch, Swedish, Norwegian Bokmål/Nynorsk, Danish, Icelandic and Faroese) with validated coverage of diverse tasks and topics. Blind evaluation demonstrates that human-reformulated prompts significantly outperform synthetically generated prompts in naturalness and comprehensibility, particularly for low-resource languages like Icelandic and Faroese. For the bigger Scandinavian lan- guage, Danish, the difference was less pronounced. The prompt dataset is released under an open-source license at https://huggingface.co/datasets/AnnikaSimonsen/TrustLLM-reformulation-prompts.
Generating High Quality Synthetic Data for Dutch Medical Conversations
Cecilia Kuan | Aditya Kamlesh Parikh | Henk van den Heuvel
Cecilia Kuan | Aditya Kamlesh Parikh | Henk van den Heuvel
Medical conversations offer insights into clinical communication often absent from Electronic Health Records. However, developing reliable clinical Natural Language Processing (NLP) models is hampered by the scarcity of domain-specific datasets, as clinical data are typically inaccessible due to privacy and ethical constraints. To address these challenges, we present a pipeline for generating synthetic Dutch medical dialogues using a Dutch fine-tuned Large Language Model, with real medical conversations serving as linguistic and structural reference. The generated dialogues were evaluated through quantitative metrics and qualitative review by native speakers and medical practitioners. Quantitative analysis revealed strong lexical variety and overly regular turn-taking, suggesting scripted rather than natural conversation flow. Qualitative review produced slightly below-average scores, with raters noting issues in domain specificity and natural expression. The limited correlation between quantitative and qualitative results highlights that numerical metrics alone cannot fully capture linguistic quality. Our findings demonstrate that generating synthetic Dutch medical dialogues is feasible but requires domain knowledge and carefully structured prompting to balance naturalness and structure in conversation. This work provides a foundation for expanding Dutch clinical NLP resources through ethically generated synthetic data.
DeepICD-R1: Medical Reasoning through Hierarchical Rewards and Unsupervised Distillation
Tom Röhr | Thomas Maximilian Josef Steffek | Roman Teucher | Keno Bressem | Alexei Figueroa | Paul Grundmann | Peter Troeger | Felix Alexander Gers | Alexander Löser
Tom Röhr | Thomas Maximilian Josef Steffek | Roman Teucher | Keno Bressem | Alexei Figueroa | Paul Grundmann | Peter Troeger | Felix Alexander Gers | Alexander Löser
Large language models (LLMs) show strong reasoning abilities, but full retraining for the medical domain is often infeasible because of lacking data or compute resources. We present DeepICD-R1, a framework for efficient medical reasoning fine-tuning that unites hierarchical rewards with distilled supervision. We reformulate ICD-10-CM prediction as a reinforcement learning problem and design a hierarchical outcome-based reward that reflects the ICD code structure across chapter, category, and full-code levels. In parallel, we publish a large-scale distilled dataset of over 90k reasoning traces derived from MIMIC-IV admission notes, integrating clinical validation and official coding guidelines. Fine-tuning smaller instruction-tuned LLMs with this data and GRPO reinforcement yields consistent gains in diagnostic accuracy and reasoning coherence. Extensive ablations confirm that hierarchical supervision and verifiable outcome rewards enable competitive, domain-specialized reasoning models without additional pretraining, providing a reproducible foundation for clinical NLP research. Keywords: Clinical NLP, Large Reasoning Model, GRPO, Supervised Fine-Tuning
SynthLLM: An LLM-based Scalable Synthetic Data Generation Pipeline for Low-Resource Languages
Solmaz Panahi | Vasudevan Nedumpozhimana | John Kelleher
Solmaz Panahi | Vasudevan Nedumpozhimana | John Kelleher
Large Language Models (LLMs) have enabled scalable synthetic data generation, yet their effective adaptation to low-resource languages remains underexplored. We introduce an LLM-based generate and annotate paradigm to create synthetic datasets for low-resource NLP classification tasks. The framework employs a smaller model for text generation and a stronger model for automatic annotation. Using Farsi Natural Language Inference (NLI) as a case study, we construct a large-scale synthetic dataset of 100,000 labeled instances. We provide a systematic empirical analysis of annotation quality, label-distribution effects, and training regimes. We compare GPT-4o-mini, Aya-23-35B, and DeBERTa as annotators and examine how annotation variability propagates to downstream performance. Our results show that a warm-up phase with synthetic data consistently outperforms data mixing and reversed ordering. Notably, open-source annotation (Aya-23-35B) achieves comparable downstream performance to the proprietary model (GPT-4o-mini), with significant cost implications for deploying pipelines in low-resource settings. The dataset and code are publicly available at https://huggingface.co/datasets/Solmazp/text2entail.
Persona-Conditioned Generation of Patient Self-Reports from EHRs
Yuexin Wu | Jianming Wei | Vasile Rus
Yuexin Wu | Jianming Wei | Vasile Rus
Accurate diagnosis depends not only on clinical expertise but also on how patients describe their symptoms at first contact. Yet large English corpora of patient-authored self-reports are scarce, limiting advances in natural, context-aware narrative modeling. We address this gap by generating first-person self-reports from structured EHR content conditioned on persona attributes that capture social and clinical context. Reports are produced by two generators and scored by two independent graders using a rubric with four dimensions, complemented by a rubric-free preference test. Across 10k stratified cases, we compare two generators under a reliable evaluation protocol and select the higher-scoring one based primarily on Clinical Correctness and Faithfulness, yielding a dataset composed of narratives from the stronger system. Our contributions are threefold: (I) we developed and release a large, persona-conditioned dataset of patient-style self-reports grounded in patient-stated EHR facts, (II) we introduce a transparent evaluation framework that combines rubric-based scoring with rubric-free preference to mitigate grader bias and enable cross-validation, (III) we find that graders exhibit systematic stylistic preferences in rubric-free approach that influence scores independent of clinical content, and (IV) we study large language models for producing first-person self-reports from structured EHRs, highlighting where they succeed, where they fail, and how this affects use in telemedicine and triage.
SocialStep: Fast Prediction of Social Determinants of Health
Paul Landes | Adam Richard Cross | Jimeng Sun
Paul Landes | Adam Richard Cross | Jimeng Sun
Given thousands of medical documents, how can we automatically uncover patients’ social risk factors? Social Determinants of Health (SDoH) constitute a growing class of non-clinical risk factors that shape patient trajectories. While clinically significant, automatic detection of SDoH from free text remains understudied due to scarce and imbalanced training data. Current approaches often rely on monolithic large language models. We present SocialStep, a two-step hybrid pipeline that first uses a lightweight classifier to triage sentences and then applies a Large Language Model (LLM) for multilabel classification to the relevant subset. On the Medical Information Mart for Intensive Care III (MIMIC-III) dataset, SocialStep improves macro F1 by 5 points over the state-of-the-art baseline while running 12.2× faster. These findings demonstrate that integrating compact neural encoders with large language models provides a scalable and highly accurate framework for clinical NLP tasks, including SDoH extraction. Notably, we also observe some unexpected patterns in LLM performance. SocialStep offers a practical blueprint for hybrid model deployment that identifies critical social risk factors without prohibitive computational cost.
Dynamically Acquiring Text Content to Enable the Classification of Lesser-known Entities for Real-world Tasks
Fahmida Alam | Ellen Riloff
Fahmida Alam | Ellen Riloff
Existing Natural Language Processing (NLP) resources often lack the task-specific information required for real-world problems and provide limited coverage of lesser-known or newly introduced entities. For example, business organizations and health care providers may need to be classified into a variety of different taxonomic schemes for specific application tasks. Our goal is to enable domain experts to easily create a task-specific classifier for entities by providing only entity names and gold labels as training data. Our framework then dynamically acquires descriptive text about each entity, which is subsequently used as the basis for producing a text-based classifier. We propose a novel text acquisition method that leverages both web and large language models (LLMs). We evaluate our proposed framework on two classification problems in distinct domains: (i) classifying organizations into Standard Industrial Classification (SIC) Codes, which categorize organizations based on their business activities; and (ii) classifying healthcare providers into healthcare provider taxonomy codes, which represent a provider’s medical specialty and area of practice. Our best-performing model achieved macro-averaged F1-scores of 82.3% and 72.9% on the SIC code and healthcare taxonomy code classification tasks, respectively.
RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts
Darya Kharlamova | Irina Proskurina
Darya Kharlamova | Irina Proskurina
Many errors in student essays can be explained by influence from the native language (L1). L1 interference refers to errors influenced by a speaker’s first language, such as using stadion instead of stadium, reflecting lexical transliteration from Russian. In this work, we address the task of detecting such errors in English essays written by Russian-speaking learners. We introduce RILEC, a large-scale dataset of over 18,000 sentences, combining expert-annotated data from REALEC with synthetic examples generated through rule-based and neural augmentation. We propose a framework for generating L1-motivated errors using generative language models optimized with PPO, prompt-based control, and rule-based patterns. Models fine-tuned on RILEC achieve strong performance, particularly on word-level interference types such as transliteration and tense semantics. We find that the proposed augmentation pipeline leads to a significant performance improvement, making it a potentially valuable tool for learners and teachers to more effectively identify and address such errors.
Critical Foreign Policy Decision (CFPD) Benchmark: Measuring Diplomatic Preferences of Large Language Models
Benjamin Jensen | Ian J. Reynolds | Yasir Atalan | Michael Garcia | Austin Woo | Anthony Chen | Trevor Howarth
Benjamin Jensen | Ian J. Reynolds | Yasir Atalan | Michael Garcia | Austin Woo | Anthony Chen | Trevor Howarth
As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. We present a novel benchmark designed to evaluate biases and preferences of models in the context of international relations (IR), which we apply to eight prominent foundation models: Llama 3.1 8B Instruct, Llama 3.1 70B Instruct, GPT-4o, Gemini 1.5 Pro-002, Mixtral 8x22B, Claude 3.5 Sonnet, DeepSeek V3, and Qwen2 72B. We designed a bias discovery study around core topics in IR using 400 expert-crafted scenarios to analyze results from our selected models. These scenarios focused on four topical domains: military escalation, military and humanitarian intervention, cooperative behavior, and alliance dynamics. Analysis reveals noteworthy variation among model recommendations based on the four tested domains. Particularly, DeepSeek V3, Qwen2 72B, Gemini 1.5 Pro-002, and Llama 3.1 8B Instruct models offered significantly more escalatory recommendations than Claude 3.5 Sonnet and GPT-4o models. All models exhibit some degree of country-specific biases. These findings highlight the necessity for controlled deployment of LLMs in high-stakes environments, emphasizing the need for domain-specific evaluations and model fine-tuning to align with institutional objectives.
CrisisCL: A Domain Incremental Learning Benchmark for Crisis Management
Paul Le Van Kiem | Romain Meunier | Farah Benamara | Véronique MORICEAU
Paul Le Van Kiem | Romain Meunier | Farah Benamara | Véronique MORICEAU
This paper proposes CrisisCL, a domain incremental learning benchmark for crisis management. Based on previous crisis management protocols, it improves consistency by allowing continual learning (CL) of new crises. A set of experiments have been conducted on multilingual datasets relying on continual learning methods and transformers to improve performance and ensure model generalization. Results reveal that regularization methods are more effective on large, coherent domains, whereas replay strategies struggle under constrained memory. Additional experimental protocols further expose the limitations of current CL methods when generalizing to unforeseen crisis events.
Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation
Neha Sharma | Navneet Agarwal | Kairit Sirts
Neha Sharma | Navneet Agarwal | Kairit Sirts
Text-based automated Cognitive Distortion detection is a challenging task due to its subjective nature, with low agreement scores observed even among expert human annotators, leading to unreliable annotations. We explore the use of Large Language Models (LLMs) as consistent and reliable annotators, and propose that multiple independent LLM runs can reveal stable labeling patterns despite the inherent subjectivity of the task. Furthermore, to fairly compare models trained on datasets with different characteristics, we introduce a dataset-agnostic evaluation framework using Cohen’s kappa as an effect size measure. This methodology allows for fair cross-dataset and cross-study comparisons where traditional metrics like F1 score fall short. Our results show that GPT-4 can produce consistent annotations (Fleiss’s Kappa = 0.78), resulting in improved test set performance for models trained on these annotations compared to those trained on human-labeled data. While human expert verification was inconclusive on our target dataset, our findings suggest that LLMs can offer a scalable and internally consistent alternative for generating training data that supports strong downstream performance in subjective NLP tasks.
LLMs as Annotators: Evaluating Model–Human Alignment in Detecting Contentious Language in Historical Corpora
Yahui Zhao | Clemencia Siro | Laura Hollink
Yahui Zhao | Clemencia Siro | Laura Hollink
Historical texts often contain terminology that reflects outdated or harmful social values. Identifying such contentious terms is essential for the Galleries, Libraries, Archives, and Museums (GLAM) community, but manual annotation requires cultural expertise and is difficult to scale. This study evaluates whether large language models (LLMs) can support this process by aligning with human judgments of contentiousness in historical Dutch corpora. Using the Dutch Contentious Contexts Corpus (ConConCor), we formalize the task as context-dependent binary classification and compare two LLMs across multiple prompt configurations and evaluation scenarios. The models achieve near-human-level agreement on explicit cases but diverge when contextual or historical reasoning is required. Analysis of disagreement patterns shows that LLMs capture overtly harmful expressions yet tend to over-predict contentiousness for identity-related and colonial terms and under-predict for semantically shifted or figurative uses. These findings suggest that LLMs can act as auxiliary annotators for sensitive language detection in historical materials, provided that human oversight and contextual interpretation remain central to annotation workflows.
Widespread Gender and Pronoun Bias in Moral Judgments across LLMs
Gustavo Lucius Fernandes | Jeiverson Santos | Pedro O.S Vaz-de-Melo
Gustavo Lucius Fernandes | Jeiverson Santos | Pedro O.S Vaz-de-Melo
Large language models (LLMs) are increasingly used to assess moral or ethical statements, yet their judgments may reflect social and linguistic biases. This work presents a controlled, sentence-level study of how grammatical person, number, and gender markers influence LLM moral classifications of fairness. Starting from 550 balanced base sentences from the ETHICS dataset, we generated 26 counterfactual variants per item, systematically varying pronouns and demographic markers to yield 14,850 semantically equivalent sentences. We evaluated six model families (Grok, GPT, LLaMA, Gemma, DeepSeek, and Mistral), and measured fairness judgments and inter-group disparities using Statistical Parity Difference (SPD). Results show statistically significant biases: sentences written in the singular form and third person are more often judged as "fair”, while those in the second person are penalized. Gender markers produce the strongest effects, with non-binary subjects consistently favored and male subjects disfavored. We conjecture that these patterns reflect distributional and alignment biases learned during training, emphasizing the need for targeted fairness interventions in moral LLM applications.
Frame2KG: A Benchmark and Evaluation Toolkit for Interpretable Frame-to-Graph Generation
Lewis N. Watson | Carl Strathearn | Kenny Mitchell | Yanchao Yu
Lewis N. Watson | Carl Strathearn | Kenny Mitchell | Yanchao Yu
Interpretable frame-to-knowledge-graph (Frame2KG) generation enables structured visual scene representation while supporting on-device inference to enhance privacy, improve interpretability, and minimise compute. We introduce Frame2KG-YC2, a synthetic, reproducible dataset derived from YouCook2 that pairs keyframes with schema-valid JSON knowledge graphs containing typed, spatially grounded entities and semantic predicates, alongside faithful textual paraphrases. Using this corpus, we fine-tune Qwen2.5-VL models (3B and 7B) with parameter-efficient LoRA adapters on attention layers (QKVO), with and without GateProj/Up/Down MLP projections. For evaluation and benchmarking, we propose a deterministic toolkit featuring two-stage node matching, an IoU gate followed by Hungarian assignment on blended spatial-semantic similarity, and comprehensive metrics spanning node/edge precision-recall-F1, matched-pair IoU, and structural validity. On a held-out test set, our models achieve Node F1μ up to 0.621 and Edge F1μ up to 0.208, with mean matched IoU of ≈0.61 and >98% schema conformity. We show that MLP gating consistently improves predicate accuracy and spatial grounding, while post-training quantisation maintains accuracy and improves deployability on edge hardware. We release the dataset, code, adapters, and evaluation toolkit to establish an open, interpretable baseline for future temporal and multi-view extensions.
Injecting Structured Biomedical Knowledge into Language Models:Continual Pretraining vs. GraphRAG
Jaafer Klila | Sondes Bannour Souihi | Rahma Boujelbane | Nasredine Semmar | Lamia Hadrich-Belguith
Jaafer Klila | Sondes Bannour Souihi | Rahma Boujelbane | Nasredine Semmar | Lamia Hadrich-Belguith
The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current approaches rely on unstructured text corpora, this study explores two complementary strategies for leveraging structured knowledge from the UMLS Metathesaurus: (i) Continual pretraining that embeds knowledge into model parameters, and (ii) Graph Retrieval-Augmented Generation (GraphRAG) that consults a knowledge graph at inference time. We first construct a large-scale biomedical knowledge graph from UMLS (3.4 million concepts and 34.2 million relations), stored in Neo4j for efficient querying. We then derive a ~100-million-token textual corpus from this graph to continually pretrain two models: BERTUMLS (from BERT) and BioBERTUMLS (from BioBERT). We evaluate these models on six BLURB (Biomedical Language Understanding and Reasoning Benchmark) datasets spanning five task types and evaluate GraphRAG on the two QA (Question Answering) datasets (PubMedQA, BioASQ). On BLURB tasks, BERTUMLS improves over BERT, with the largest gains on knowledge-intensive QA. Effects on BioBERT are more nuanced, suggesting diminishing returns when the base model already encodes substantial biomedical text knowledge. Finally, augmenting LLaMA 3-8B with our GraphRAG pipeline yields over than 3 points accuracy on PubMedQA and 5 points on BioASQ without any retraining, delivering transparent, multi-hop, and easily updated knowledge access. We release the processed UMLS Neo4j graph to support reproducibility.
Linguistic Knowledge Graphs for Sense Prediction: A Case-study on Latin
Eleonora Ghizzota | Paola Marongiu | Pierpaolo Basile | Stefano Ferilli | Barbara McGillivray
Eleonora Ghizzota | Paola Marongiu | Pierpaolo Basile | Stefano Ferilli | Barbara McGillivray
This paper investigates the integration of the Linguistic Knowledge Graph (LKG) and Large Language Models (LLMs) for word sense prediction in Latin, a morphologically rich and low-resource historical language. Building on recent work in word sense disambiguation (WSD) and semantic change detection, we use a LKG that integrates information from a diachronic Latin corpus, a sense-annotated dataset of Latin, Latin WordNet, and Wikidata, as a structured representation of semantic and contextual relations. We present sense prediction as a binary classification task over the Latin dataset, using a Graph Retrieval-Augmented Generation approach that combines knowledge graph retrieval with LLM prompting. Two types of graph metadata are tested: author-related information (work, period, occupation) and linguistic metadata (synset and hypernyms derived from WordNet for each word sense). Experiments conducted on GPT-4o-mini, LLaMA-3.1-8B and LLaMA-3.3-70B show varying performance, with F1 scores ranging from 0.53 to 0.77. While GPT-4o-mini achieves the best overall accuracy, LLaMA-3.3-70B benefits the most from graph-based metadata, improving its F1 score by up to 3 points. Analysis by word type reveals that concrete and semantically shifting words are more easily disambiguated than abstract and semantically stable words. Results highlight both the promise and the challenges of combining graph-structured linguistic knowledge with LLMs for historical WSD.
ACID: On the Perception of Online Classism
Arianna Muti | Elisa Bassignana | Amanda Cercas Curry | Federica Durante | Dirk Hovy | Debora Nozza
Arianna Muti | Elisa Bassignana | Amanda Cercas Curry | Federica Durante | Dirk Hovy | Debora Nozza
Socioeconomic status (SES) structures social inequality and underlies class-based discrimination that is often rationalised through stereotypes expressed in public discourse. However, despite extensive research on hate speech detection in Natural Language Processing, classism detection remains an underexplored phenomenon. We introduce ACID, a cross-cultural corpus with over 1.15 million instances, to investigate classism across YouTube and Twitter from 14 English-speaking countries. We examine (i) which stereotypes are invoked towards lower-SES, (ii) whether blame for lower-SES is attributed to individuals or structural factors, and (iii) whether these people are portrayed offensively. Across platforms, explanations are predominantly framed in terms of individual responsibility. Across countries, class stereotypes consistently revolve around moralized notions of dependency, laziness, and ignorance, revealing a shared global structure of class-based stigma. Our dataset and analysis are a foundation to advance research on class-based discrimination and its representation in online discourse.
The Spectrum of Sentiment: Optimistic, Pessimistic, and Neutral Voices in Online Depression Discourse
Stefana Arina Tabusca | Ana-Maria Bucur | Liviu P. Dinu
Stefana Arina Tabusca | Ana-Maria Bucur | Liviu P. Dinu
The relationship between depression and the concepts of optimism and pessimism has been extensively researched by psychologists. In this paper, we use computational approaches to study how optimism and pessimism are expressed in the online discourse of people with a depression diagnosis. Publicly available datasets are used for the development of an optimism/pessimism detection model, as well as for the analyses performed on social media posts of individuals with depression, as measured by BDI-II, a validated depression questionnaire. To analyze the optimistic and pessimistic posts by individuals with depression, we use LIWC features and perform topic modeling. We also investigate specific words driving mislabeling using SHAP. Our results show that while there may not be significant differences in the number of optimistic versus pessimistic posts between individuals in the depression and control groups, the content of the posts differs meaningfully, both in terms of linguistic features and approached topics.
A Benchmark Dataset and Comparative Evaluation of Phonemized and Romanized Urdu for Text-to-Speech
M Kaab Bin Shahid | Muhammed Izharuddin
M Kaab Bin Shahid | Muhammed Izharuddin
Text-to-Speech (TTS) system for the Urdu language presents significant challenges, primarily due to the scarcity of high-quality datasets and an insufficient focus on modeling pronunciation. Urdu is spoken by 250 million people worldwide, but its research on computational linguistics remains underrepresented. In this paper, we introduce URDUTTS, a comprehensive and publicly available Urdu TTS dataset containing 89 hours of studio-quality speech, with accompanying transcriptions in three formats: Urdu Script, Phonemized Script, and Romanized Script. The dataset includes both mono-speaker and multi-speaker configurations. As Urdu relies heavily on phonetic features, accurate pronunciation is highly essential for the language. Therefore, we benchmark our dataset using VITS and GlowTTS models to compare the widely used Romanized script format with the Phonemized representation. To make the evaluation highly comprehensive, we combined both objective and subjective evaluation strategies. For objective evaluation, Mel-Cepstral Distortion (MCD with Plain, Dynamic Time-Warping, and Slope-Limitation variants), Signal-to-Noise Ratio (SNR), Word Error Rate (WER), and Character Error Rate (CER) were taken. Subjective evaluation was governed by Mean Opinion Score (MOS) ratings from 40 native speakers. Results show that using VITS and GlowTTS with Phonemized transcriptions performs significantly better than Romanized ones, with an improvement of 9.6% and 26.5% in MOS. The data and code are available at github.com/KAABSHAHID/URDUTTS.
S-VoCAL: A Dataset and Evaluation Framework for Inferring Speaking Voice Character Attributes in Literature
Abigail Berthe-Pardo | Gaspard Michel | Elena V. Epure | Christophe Cerisara
Abigail Berthe-Pardo | Gaspard Michel | Elena V. Epure | Christophe Cerisara
With recent advances in Text-to-Speech (TTS) systems, synthetic audiobook narration has seen increased interest, reaching unprecedented levels of naturalness. However, larger gaps remain in synthetic narration systems’ ability to impersonate fictional characters, and convey complex emotions or prosody. A promising direction to enhance character identification is the assignment of plausible voices to each fictional characters in a book. This step typically requires complex inference of attributes in book-length contexts, such as a character’s age, gender, origin or physical health, which in turns requires dedicated benchmark datasets to evaluate extraction systems’ performances. We present S-VoCAL (Speaking Voice Character Attributes in Literature), the first dataset and evaluation framework dedicated to evaluate the inference of voice-related fictional character attributes. S-VoCAL entails 8 attributes grounded in sociophonetic studies, and 952 character-book pairs derived from Project Gutenberg. Its evaluation framework addresses the particularities of each attribute, and includes a novel similarity metric based on recent Large Language Models embeddings. We demonstrate the applicability of S-VoCAL by applying a simple Retrieval-Augmented Generation (RAG) pipeline to the task of inferring character attributes. Our results suggest that the RAG pipeline reliably infers attributes such as Age or Gender, but struggles on others such as Origin or Physical Health. The dataset and evaluation code are available at https://github.com/AbigailBerthe/S-VoCAL.
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
Yunseung Lee | Subin Kim | Youngjun Kwak | Jaegul Choo
Yunseung Lee | Subin Kim | Youngjun Kwak | Jaegul Choo
Large language models (LLMs)-based chatbots are increasingly being adopted in the financial domain, particularly in digital banking, to handle customer inquiries about products such as deposits, savings, and loans. However, these models still exhibit low accuracy in core banking computations—including total payout estimation, comparison of products with varying interest rates, and interest calculation under early repayment conditions. Such tasks require multi-step numerical reasoning and contextual understanding of banking products, yet existing LLMs often make systematic errors—misinterpreting product types, applying conditions incorrectly, or failing basic calculations involving exponents and geometric progressions. However, such errors have rarely been captured by existing benchmarks. Mathematical datasets focus on fundamental math problems, whereas financial benchmarks primarily target financial documents, leaving everyday banking scenarios underexplored. To address this limitation, we propose BankMathBench, a domain-specific dataset that reflects realistic banking tasks. BankMathBench is organized in three levels of difficulty—basic, intermediate, and advanced—corresponding to single-product reasoning, multi-product comparison, and multi-condition scenarios, respectively. When trained on BankMathBench, open-source LLMs exhibited notable improvements in both formula generation and numerical reasoning accuracy, demonstrating the dataset’s effectiveness in enhancing domain-specific reasoning. With tool-augmented fine-tuning, the models achieved average accuracy increases of 57.6%p (basic), 75.1%p (intermediate), and 62.9%p (advanced), representing significant gains over zero-shot baselines. These findings highlight BankMathBench as a reliable benchmark for evaluating and advancing LLMs’ numerical reasoning in real-world banking scenarios.
TR-TEB: Turkish Text Embedding Benchmark
Omer Arslan | Atalay Celik | Yusuf Aslan | Hasan Fatih Durkaya | Mustafa Furkan Zenginoglu | Musa Alperen Yilmaz | Merve Gul Kantarci | Mehmet Haklidir
Omer Arslan | Atalay Celik | Yusuf Aslan | Hasan Fatih Durkaya | Mustafa Furkan Zenginoglu | Musa Alperen Yilmaz | Merve Gul Kantarci | Mehmet Haklidir
Text embeddings are central to modern natural language processing, enabling several downstream tasks. Despite their significance, existing evaluation frameworks primarily target English and other high-resource languages, leaving critical gaps for languages such as Turkish. To address this, we present TR-TEB (Turkish Text Embedding Benchmark), the first comprehensive, standardized, and reproducible benchmark for Turkish text embeddings. TR-TEB spans five core task categories: classification, pair classification, clustering, retrieval, and semantic textual similarity. It is supported by a diverse dataset portfolio that integrates 14 curated open-source resources, 26 high-quality translated datasets, and 7 newly constructed Turkish-specific datasets designed to capture the language’s unique characteristics. We test our framework by comparing 45 well-known open-source embedding models. As the first unified evaluation suite, TR-TEB serves as a core tool for the Turkish embedding research community, establishing a systematic basis for model comparison and improvement. Furthermore, its benchmarking methodology and dataset creation process provide a blueprint for extending robust embedding evaluation to other low-resource languages.
Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+
Mason Shipton | York Hay Ng | Aditya Khan | Phuong H. Hoang | Xiang Lu | A. Seza Doğruöz | Annie En-Shiun Lee
Mason Shipton | York Hay Ng | Aditya Khan | Phuong H. Hoang | Xiang Lu | A. Seza Doğruöz | Annie En-Shiun Lee
The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups.
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Xanh Ho | Yun-Ang Wu | Sunisth Kumar | Tian Cheng Xia | Florian Boudin | Andre Greiner-Petter | Akiko Aizawa
Xanh Ho | Yun-Ang Wu | Sunisth Kumar | Tian Cheng Xia | Florian Boudin | Andre Greiner-Petter | Akiko Aizawa
We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, directly extracted from published papers. To create refuted claims, we introduce a novel approach that modifies the supporting evidence (figures and tables), rather than altering the claims or relying on large language models (LLMs) to fabricate contradictions. The dataset provides cross-modal evidence with diverse representations: figures are available as images, while tables are provided in multiple formats, including images, LaTeX source, HTML, and JSON. SciClaimEval contains 1,664 annotated samples from 180 papers across three domains, machine learning, natural language processing, and medicine, validated through expert annotation. We benchmark 11 multimodal foundation models, both open-source and proprietary, across the dataset. Results show that figure-based verification remains particularly challenging for all models, as a substantial performance gap remains between the best system and human baseline.
Localizing Events in Space: Comparing Humans and AI Models
Derrick Eui Gyu Kim | Kenneth Lai | James Pustejovsky
Derrick Eui Gyu Kim | Kenneth Lai | James Pustejovsky
Understanding how Large Language Models (LLMs) and Text-to-Image models (T2Is) acquire and apply implicit spatial knowledge remains an open challenge. In this paper, we present a novel dataset and evaluation framework designed to probe event localization capabilities in both humans, LLMs and T2Is. Our dataset includes 134 sentence pairs derived from Flickr30k captions, where explicit location information is systematically removed via Abstract Meaning Representation (AMR) parsing and manual refinement. Using this dataset, we analyze the effects of location ablation on spatial reasoning across human annotators, LLMs, and T2Is. Results show that while humans maintain robust location inferences after ablation, LLMs exhibit degraded performance, particularly for semantically polysemous verbs. T2Is demonstrate similar limitations, often generating visually inconsistent spatial contexts when locative cues are missing. Our findings highlight the gap between human and LLMs and T2Is in recovering implicit situational knowledge and suggest future directions for improving spatial reasoning in multimodal AI systems. This dataset contribution work serves as a proof-of-concept for systematic evaluation of implicit spatial reasoning and paves the way for larger-scale studies.
STRUDEL: Unrolling a Benchmark for Evaluating Vision-Language Models on Structured Diagram Understanding across Domains
Daniel Steinigen | Lucie Flek | Sebastian Houben
Daniel Steinigen | Lucie Flek | Sebastian Houben
Vision-Language Models (VLMs) have achieved impressive progress across diverse multimodal tasks, yet their ability to interpret structured diagrams, such as circuit schematics, molecular structures, musical notation, business process flow charts or class diagrams, which are central to scientific and engineering communication, remains underexplored. We introduce STRUDEL (STRUctured Diagram EvaLuation), a benchmark for evaluating VLMs on structured diagram understanding across 8 domains and 20 image categories. STRUDEL leverages Large-Language Models (LLMs) to synthesize code in domain-specific formal representation languages (FRLs) (e.g. circuit netlists, SMILES, ABC-Notation, BPMN or PlantUML), which are rendered into valid diagrams and paired with generated tasks, functional descriptions, and captions. A multi-stage pipeline filters invalid, cluttered, or redundant samples and employs LLM-as-a-judge scoring to ensure correctness. Through targeted experiments, we evaluate the ability of LLMs to generate valid code in distinct FRLs, demonstrating their capability to successfully perform this task. The resulting benchmark comprises diverse task types covering identification, quantification, structural analysis, image-text association, and image-to-code translation. Evaluating 35 VLMs using STRUDEL reveals that models excel at association tasks, demonstrating strong visual-textual alignment, yet struggle with quantification and identification, where precise structural understanding is required. Performance varies markedly in image-to-code translation, reflecting significant differences in how models connect visual inputs to formal representations. Overall, STRUDEL establishes a scalable foundation for assessing and advancing VLMs torward deeper and more systematic understanding of structured visual information across domains.
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
Byeonggeuk Lim | Kyeonghyun Kim | Jungmin Yun | Youngbin Kim
Byeonggeuk Lim | Kyeonghyun Kim | Jungmin Yun | Youngbin Kim
The advancement of Large Vision-Language Models (LVLMs) requires precise local region-based reasoning that faithfully grounds the model’s logic in actual visual evidence. However, existing datasets face limitations in scalability due to extensive manual annotation and lack explicit alignment between multi-step reasoning and corresponding image regions, which constrains the evaluation of model trustworthiness. To address these challenges, we propose the Visual Grounding Chain-of-Thought (VG-CoT) dataset, which explicitly links each reasoning step to real visual evidence within the image through a fully automated three-stage pipeline. The pipeline first extracts object- and text-level visual evidence using state-of-the-art detection and OCR models, then generates step-by-step grounded reasoning with GPT-4o, and finally refines the grounding through a rationale-driven open-set detection process. In addition, we introduce a new benchmark that comprehensively evaluates LVLMs reasoning across three complementary dimensions: Rationale Quality, Answer Accuracy, and Reasoning–Answer Alignment. Experiments with representative LVLMs, including LLaVA-1.5 and Qwen2-VL, demonstrate consistent improvements across all evaluation metrics, confirming that VG-CoT effectively enhances trustworthy, evidence-based reasoning while maintaining scalable and cost-efficient dataset construction. The dataset and code will be released publicly upon acceptance to facilitate further research.
VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
Josef Kuchar | Marek Kadlcik | Michal Spiegel | Michal Stefanik
Josef Kuchar | Marek Kadlcik | Michal Spiegel | Michal Stefanik
We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language edit instructions. Our dataset enables training and evaluation of models that modify vector graphics based on textual commands. We describe the data collection process, including image pairing via CLIP similarity and instruction generation with vision-language models. Initial experiments with state-of-the-art large language models reveal that current methods struggle to produce accurate and valid edits, underscoring the challenge of this task. To foster research in natural language-driven vector graphic generation and editing, we make our resources created within this work publicly available.
ViWikiFC: Fact-Checking for Vietnamese Wikipedia-Based Textual Knowledge Source
Hung Tuan Le | Long Truong To | Manh Trong Nguyen | Kiet Van Nguyen
Hung Tuan Le | Long Truong To | Manh Trong Nguyen | Kiet Van Nguyen
Fact-checking is essential due to the explosion of misinformation in the media ecosystem. Although false information exists in every language and country, most research to solve the problem has mainly concentrated on huge communities like English and Chinese. Low-resource languages like Vietnamese are necessary to explore corpora and models for fact verification. To bridge this gap, we construct ViWikiFC, the first manually annotated open-domain corpus for Vietnamese Wikipedia Fact Checking more than 20K claims generated by converting evidence sentences extracted from Wikipedia articles. We analyze our corpus through many linguistic aspects, from the new dependency rate, the new n-gram rate, and the new word rate. We conducted various experiments for Vietnamese fact-checking, including evidence retrieval and verdict prediction. BM25 and InfoXLMLarge achieved the best results in two tasks, with BM25 achieving an accuracy of 88.30% for SUPPORTS, 86.93% for REFUTES, and only 56.67% for the NEI label in the evidence retrieval task. InfoXLMLarge achieved an F1 score of 86.51%. Furthermore, we also conducted a pipeline approach, which only achieved a strict accuracy of 67.00% when using InfoXLMLarge and BM25. These results demonstrate that our dataset is challenging for the Vietnamese language model in fact-checking tasks.
Automated Extraction of Answer Candidates for Question Generation
Claudia Preda | Mihai Dascalu | Stefan Ruseti | Danielle S. McNamara
Claudia Preda | Mihai Dascalu | Stefan Ruseti | Danielle S. McNamara
Answering questions based on a reference text is a frequently employed comprehension assessment method that enables teachers to effectively and efficiently evaluate students. Various tools and methods were developed to tackle automated question generation, however, selecting valid answer candidates as a first step is less addressed. Thus, we introduce a solution built on top of FairytaleQA and tailored for training a DeBERTa-based model to classify the quality of each candidate to be part of a strong answer-question pair. First, we extract answer candidates by syntactically parsing the context (i.e., selecting text spans from the reference text based on the nodes in the constituency tree); then, questions are generated for the extracted candidates using a pre-trained LLM model on this task. Next, we assess a candidate’s quality by relying on another fine-tuned model’s capability to answer the previously generated question for that candidate. This enables us to categorize answers using a four-class system: very good, good, average, and unusable. A significant advantage of our method is that the encoder classifier can score all potential answer candidates in a single inference step for the entire context. We compare our selection against both the answers from explicit questions in the original dataset and a fine-tuned LLM for answer selection using an Elo ranking system. In addition, we propose three strategies based on semantic similarity and text position to ensure coverage and diversity of candidates’ selection.
Green Bots versus Red Bots: Evaluating Large Language Models for Simulating Persuasion Dynamics in Online Influence Campaigns
Majd Eddin Al Ali | Filip Mihai Muntean | Lucia Donatelli | Jurriaan van Diggelen
Majd Eddin Al Ali | Filip Mihai Muntean | Lucia Donatelli | Jurriaan van Diggelen
Large language models (LLMs) are increasingly used to simulate social interaction and persuasion dynamics, yet their validity as proxies for human cognition and behavior remains unverified. We propose a dual-level evaluation framework to assess LLM-based agents at both the individual and collective levels. At the individual level, we examine agent fidelity by comparing LLM-generated political personas to human benchmark data. We find that while agents capture broad partisan orientations, they underestimate within-group variability and reproduce stereotypical ideological biases. At the collective level, we deploy Big Five personality-differentiated agents in 1080 structured dialogues to test the effect of rhetorical strategy on persuasive success. Our simulations reproduce theoretically expected interaction patterns; nevertheless, belief shifts are exaggerated relative to human baselines, supporting LLMs’ tendency toward over-responsiveness. These findings suggest a trade-off between engagement-optimized training objectives and psychological realism, confirming the need to use LLMs with caution to simulate human behavior. We contribute three resources: a persuasion dynamics dataset, a standardized agent taxonomy of “red” and “green” bots, and a framework for evaluating both individual-agent fidelity and emergent group-level behavior.
Towards Expectation Detection in Language: A Case Study on Treatment Expectations in Reddit
Aswathy Velutharambath | Amelie Wührl
Aswathy Velutharambath | Amelie Wührl
Patients’ expectations towards their treatment have a substantial effect on the treatments’ success. While primarily studied in clinical settings, online patient platforms like medical subreddits may hold complementary insights: treatment expectations that patients feel unnecessary or uncomfortable to share elsewhere. Despite this, no studies examine what type of expectations users discuss online and how they express them. Presumably this is because expectations have not been studied in natural language processing (NLP) before. Therefore, we introduce the task of Expectation Detection, arguing that expectations are relevant for many applications, including opinion mining and product design. Subsequently, we present a case study for the medical domain, where expectations are particularly crucial to extract. We contribute RedHOTExpect, a corpus of Reddit posts (4.5K posts) to study expectations in this context. We use a large language model (LLM) to silver-label the data and validate its quality manually (label accuracy ~78%). Based on this, we analyze which linguistic patterns characterize expectations and explore what patients expect and why. We find that optimism and proactive framing are more pronounced in posts about physical or treatment-related illnesses compared to mental-health contexts, and that in our dataset, patients mostly discuss benefits rather than negative outcomes. The RedHOTExpect corpus can be obtained from https://www.ims.uni-stuttgart.de/data/RedHOTExpect
Empathy Speaks in Metaphors: The Empathy-Metaphor Corpus of Figurative Language in Empathetic Text
Gyeongeun Lee | Natalie Parde
Gyeongeun Lee | Natalie Parde
Metaphorical language is a powerful vehicle for expressing empathy, yet it has received limited attention in computational studies of supportive communication. We introduce Empathy-Metaphor, the first corpus that explicitly annotates metaphorical spans in empathetic online peer-support. Building on 2,492 empathetic posts from an acne support forum, the dataset contains over 2,100 manually identified metaphorical spans with strong inter-annotator agreement (κ=0.85). Analyses show that metaphors are frequent, diverse, and strategically positioned, often framing acne as a battle, journey, or shared struggle. Lexical and semantic clustering highlight recurring themes of encouragement and emotional hardship, while psycholinguistic analysis emphasizes the prominence of conflict and negative emotion framings. Benchmark experiments demonstrate that transformer models, especially DeBERTa-v3, substantially outperform linear and recurrent baselines, achieving a token-level macro F1 of 0.634 and a span-level macro F1 of 0.440 under relaxed evaluation. These contributions establish a new resource for studying figurative language in empathetic text, providing insights into the creative role of metaphors in online support.
A Computational Diachronic Analysis of Gen Z Mental Health Discourse: A Large-scale Reddit Corpus Study from Pre- to Post-COVID
Felix Mao
Felix Mao
Generation Z’s mental health discourse has been uniquely shaped by digital saturation and the COVID-19 pandemic. This study introduces a large-scale corpus of Gen Z mental health discourse on Reddit, comprising over 3 million posts across 11 subreddits (2017–2025), identified through behavioral cross-posting between mental health and Gen Z-identified communities. Using a hybrid methodology that integrates statistical corpus linguistics with NLP techniques, we conduct diachronic keyness analysis, sentiment tracking, and topic modeling to examine lexical, syntactic, and semantic patterns across pre-, during-, and post-COVID periods. Our analysis reveals: (1) ritualized support exchanges more pronounced in Gen Z where highly negative self-disclosure functions as an authenticity signal; (2) a pandemic-induced reframing of existing mental health topics, particularly a rise in physical symptoms, followed by a sustained post-pandemic sentiment decline; and (3) a generational divergence where Gen Z favors abstract, existential concerns, unlike the pragmatic focus of non-Gen Z users. This study contributes a replicable approach for analyzing youth discourse and underscores the importance of culturally and linguistically informed digital mental health interventions, which can support Gen Z’s modes of expressing distress rather than pathologizing them.
"Oat Milk Vegan Chocolate Taste Great!": Monitoring the Food Transition Debate in Reddit
Greta Zella | Jan Willem Bolderdijk | Saskia Peels | Gerry Wakker | Tommaso Caselli
Greta Zella | Jan Willem Bolderdijk | Saskia Peels | Gerry Wakker | Tommaso Caselli
We present DRiFT (Debates on Reddit involving Food Transition), a new large-scale corpus and set of computational methods for using language as an early indicator of social change in the protein transition, i.e., the shift from a diet predominantly based on animal proteins to one based mainly on plant sources. DRiFT comprises 17.5M Reddit comments (2010–2022) from 29 subreddits grouped into two speaker communities: SUSTAINABLE (early adopters/innovators) and GENERIC (general public). Building on neologism analysis, lexical semantic change detection, and connotative profiling, we introduce three linguistic measures of innovation awareness, meaning shift, and attitudinal valence. We extract neonyms and retronyms to quantify awareness; apply static and contextual embedding-based Lexical Semantic Change methods (PPMI, SGNS, BERT substitutions) to probe semantic reconceptualization; and adapt an embedding-based connotation hyperplane to measure polarity changes for targeted terms. Results show marked diastratic differences, with SUSTAINABLE users both using innovation-specific lexicon more frequently and having reconceptualized core food terms in ethical/environmental frames, while the GENERIC community exhibits rapid proportional growth in neologism use and emerging positive connotations for some plant-based products. Diachronic denotational shifts over the 12-year window are weak, suggesting shortcoming of embedding-based methods to capture subtle meaning changes. DRiFT and our analyses demonstrate that language can function as a sensitive “thermometer” of subtle social change, revealing attitudinal dynamics before observable behavioral shifts.
ClimateChat-300K: A Multi-Modal Facebook Dataset for Understanding Diverse Perspectives in Climate Communication
Wajdi Zaghouani | Md. Rafiul Biswas | Mabrouka Bessghaier | Shimaa Amer Ibrahim | George Mikros
Wajdi Zaghouani | Md. Rafiul Biswas | Mabrouka Bessghaier | Shimaa Amer Ibrahim | George Mikros
We present ClimateChat-300K, a large-scale dataset of 299,329 public Facebook posts about climate change collected between May 2020 and May 2024 through the CrowdTangle platform. The dataset contains 41 metadata features including post content, engagement metrics, and page attributes, covering material from more than 26,000 global pages. Each post includes rich contextual information such as language, timestamp, page category, and interaction counts, enabling comprehensive analyses of public discourse around climate communication. Using topic modeling and sentiment analysis, we identify ten main themes grouped into five domains: policy, activism, cooperation, science, and conservation. The results reveal that emotional tone, post format, and page identity strongly influence audience engagement, with visually rich and emotionally charged content receiving the highest levels of interaction. The dataset also demonstrates how online discussions evolved in response to major events such as international climate summits and the COVID-19 pandemic period. ClimateChat-300K provides an open resource for reproducible and interdisciplinary research on polarization, misinformation, and the dynamics of digital climate discourse. By releasing this dataset, we aim to support transparent, data-driven research and contribute to a deeper understanding of how public engagement with climate issues develops across time, geography, and institutional contexts.
HateMirage: An Explainable Multi-Dimensional Dataset for Decoding Faux Hate and Subtle Online Abuse
Sai Kartheek Reddy Kasu | Shankar Biradar | Sunil Saumya | Md. Shad Akhtar
Sai Kartheek Reddy Kasu | Shankar Biradar | Sunil Saumya | Md. Shad Akhtar
Subtle and indirect hate speech remains an underexplored challenge in online safety research, particularly when harmful intent is embedded within misleading or manipulative narratives. Existing hate speech datasets primarily capture overt toxicity, underrepresenting the nuanced ways misinformation can incite or normalize hate. To address this gap, we present HateMirage, a novel dataset of Faux Hate comments designed to advance reasoning and explainability research on hate emerging from fake or distorted narratives. The dataset was constructed by identifying widely debunked misinformation claims from fact-checking sources and tracing related YouTube discussions, resulting in 4,530 user comments. Each comment is annotated along three interpretable dimensions: Target (who is affected), Intent (the underlying motivation or goal behind the comment), and Implication (its potential social impact). Unlike prior explainability datasets such as HateXplain and HARE, which offer token-level or single-dimensional reasoning, HateMirage introduces a multi-dimensional explanation framework that captures the interplay between misinformation, harm, and social consequence. We benchmark multiple open-source language models on HateMirage using ROUGE-L F1 and Sentence-BERT similarity to assess explanation coherence. Results suggest that explanation quality may depend more on pretraining diversity and reasoning-oriented data rather than on model scale alone. By coupling misinformation reasoning with harm attribution, HateMirage establishes a new benchmark for interpretable hate detection and responsible AI research.
MindSET: Advancing Mental Health Benchmarking through Large-Scale Social Media Data
Saad Mankarious | Edward Kempa | Daniel Wiechmann | Elma Kerz | Yu Qiao | Ayah Zirikly
Saad Mankarious | Edward Kempa | Daniel Wiechmann | Elma Kerz | Yu Qiao | Ayah Zirikly
Social media data has become a vital resource for studying mental health, offering real-time insights into thoughts, emotions, and behaviors that traditional methods often miss. Progress in this area has been facilitated by benchmark datasets for mental health analysis; however, most existing benchmarks have become outdated due to limited data availability, inadequate cleaning, and the inherently diverse nature of social media content (e.g., multilingual and harmful material). We present a new benchmark dataset, MindSET, curated from Reddit using self-reported diagnoses to address these limitations. The annotated dataset contains over 13M annotated posts across seven mental health conditions—more than twice the size of previous benchmarks. To ensure data quality, we applied rigorous preprocessing steps, including language filtering, and removal of Not Safe for Work (NSFW) and duplicate content. We further performed a linguistic analysis using LIWC to examine psychological term frequencies across the eight groups represented in the dataset. To demonstrate the dataset’s utility, we conducted binary classification experiments for diagnosis detection using both fine-tuned language models and Bag-of-Words (BoW) features. Models trained on MindSET consistently outperformed those trained on previous benchmarks, achieving up to an 18-point improvement in F1 for Autism detection. Overall, MindSET provides a robust foundation for researchers exploring the intersection of social media and mental health, supporting both early risk detection and deeper analysis of emerging psychological trends.
We present a new Turkish social media corpus annotated for verbal irony. The ironic post candidates are identified by a distant supervision method relying on reports of misunderstood irony in social media platforms. The data collected through this method, as well as irony-tagged posts and a random sample of posts are annotated by three annotators, resulting in a corpus of 3000 tweets with high quality annotations that may be useful for linguistic analysis as well as for training automatic irony detection systems or testing irony understanding of large language models. Since irony interpretation typically involves context, our dataset also includes the preceding conversational context of the potentially ironic expression. Besides the description of the corpus and the annotation process, this paper presents an analysis of the corpus. Our findings indicate that relying on distant supervision alone may result in suboptimal labels for irony/sarcasm corpora. We also investigate the usefulness of context for the annotators in identifying irony.
A Corpus of Joint EEG and Self-Paced Reading of Natural Dutch Texts
Sara Møller Østergaard | Lenneke Doris Lichtenberg | Laura Boon | Bruno Nicenboim
Sara Møller Østergaard | Lenneke Doris Lichtenberg | Laura Boon | Bruno Nicenboim
We present the Tilburg corpus of Natural Dutch Texts (TiNT): A corpus of joint electroencephalography (EEG) and self-paced reading (SPR) of natural, medium-length, Dutch texts. The corpus contains recordings from 71 native Dutch speakers reading eight naturally occurring texts of around 600 words each. The texts are of varying genres and were chosen based on overall fluency and comprehensibility. To assess the quality of the corpus, we examined participant responses to comprehension questions, self-reported familiarity with the texts, and whether well-established effects replicated for both reading times and event-related potentials (ERPs) (N400 and P600). The corpus contributes to a small collection of corpora with simultaneous recording of reading times and EEG. While this is often achieved using eye-tracking, the use of SPR offers methodological advantages, particularly in aligning neural signals with word-level processing. In addition, the use of natural texts with longer dependencies makes the corpus a unique resource for psycholinguistic research. The corpus enables research into the relationship between neural and behavioral responses in naturalistic reading contexts.
Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?
Prateek Kumar Rajput | Yewei Song | Iyiola Emmanuel Olatunji | Jacques Klein | Tegawendé Bissyande
Prateek Kumar Rajput | Yewei Song | Iyiola Emmanuel Olatunji | Jacques Klein | Tegawendé Bissyande
Can large language models reliably express a human-like personality, or are they merely mimicking surface cues without a stable underlying profile? We study this question on the long-form Essays Dataset, preferred over short, mood-driven text to target stable traits. Using a questionnaire-based (self-evaluation) test: IPIP-NEO, we ask: (i) does post-training (SFT, DPO, ORPO) stabilize questionnaire scores under prompt rephrasings, and (ii) can it induce target Big Five profiles from unguided essays? Across five models, fine-tuning consistently reduces variance in questionnaire responses, mitigating the fragility seen in pre-trained models. Yet accuracy on the full five-dimensional profile remains near chance even when single-trait scores improve, indicating that unguided essays lack the cues needed for faithful personality expression. We argue for scenario-grounded datasets or interactive elicitation that accumulates test-aligned evidence over time.
A Multi-Dialectal, Longitudinal Corpus of Human-AI Hybrid Language Production
Qiao Gan | Jonathan Dunn | Andrea Nini | Benjamin Adams
Qiao Gan | Jonathan Dunn | Andrea Nini | Benjamin Adams
This paper presents a multi-dialectal, longitudinal corpus of human-AI hybrid language production, comprising purely human-written texts, purely LLM-generated texts, and hybrid texts produced under different LLM-assistance modes (e.g., stylistic suggestions, short continuations, partial essay generation). The corpus includes 693 participants from five national English dialects, with natural and hybrid samples paired within individuals over a four-week period. This design enables investigation of both short- and longer-term effects of LLM assistance on language use across geographic and social contexts. To illustrate the corpus’s utility, we analyze linguistic features across three dimensions: lexical diversity, syntactic complexity, and stylistic variation. The results show that LLM assistance enhances lexical diversity without a corresponding increase in syntactic complexity, revealing distinct effects across linguistic dimensions. Overall, this corpus offers a valuable resource for studying human-AI interaction, dialectal variation, and the influence of AI assistance on written language.
Semantic Information: A Difference That Makes a Difference
J. Nathanael Philipp | Max Kölbl | Michael Richter
J. Nathanael Philipp | Max Kölbl | Michael Richter
In the framework of distributional semantics, we introduce a novel notion and operationalisation of semantic information for natural language. The key idea is as follows: a linguistic sign carries semantic information about a document if it reduces the amount of surprisal for a language processor. We consider two systems, an informed one and an uninformed one, and describe semantic information in their terms. Processing effort is quantified via surprisal where the informed system is ‘aware’ of the linguistic sign and the uninformed one is not. On an English fairy tale corpus and on two German news corpora, we tested successfully the prediction that if the linguistic sign in question carries pre-information through semantic surprisal, the current level of surprisal for the language processor is reduced. The conclusion is that the degree of semantic information results from the degree of semantic prior information.
Modeling the Memory-Surprisal Trade-Off over Time: Communicative Efficiency Decreases with Lexico-Grammatical Change in Scientific English
Julius Steuer | Marie-Pauline Krielke | Stefania Degaetano-Ortlieb | Elke Teich | Dietrich Klakow
Julius Steuer | Marie-Pauline Krielke | Stefania Degaetano-Ortlieb | Elke Teich | Dietrich Klakow
The memory-surprisal trade-off (MST) has been shown to hold cross-linguistically as a general principle of communicative efficiency: languages that exhibit information locality tend to have word orders that allow for efficient memory use, i.e., lower surprisal at a fixed memory budget. In this paper, we explore the influence of diachronic variation on the MST. We compare scientific English in the Royal Society Corpus (RSC, 18thc. – 20thc.) to “general language” in the Corpus of Historical American English (COHA) to assess the impact of intra-linguistic variation (register). We find that both time and register influence the shape of the tradeoff: Over time, vocabulary expansion raises minimal surprisal, while the shape of the MST curves changes. Decreasing distances between syntactic dependencies due to more local nominal encodings change how predictive information is distributed across memory scales. The effects are stronger for the RSC than for COHA.
Mechanistic Interpretability Meets Cognitive Linguistics: Modelling Locative Image Schemas in the Circuit Framework
Mattia Proietti | Afra Alishahi | Grzegorz Chrupała | Alessandro Lenci
Mattia Proietti | Afra Alishahi | Grzegorz Chrupała | Alessandro Lenci
Large Language Models are often considered the best computational testbeds for linguistic theorisation at our disposal. However, their inner workings remain largely opaque, and the mechanisms behind their behaviour cannot always be easily connected with theoretical linguistic assumptions. Mechanistic Interpretability (MI) is surging as a specialised field to reverse engineer models’ internals and shed light on the causal relationships happening under the hood. Nevertheless, MI is predominantly focused on AI-Safety problems, and the attempts to understand linguistically motivated behaviours with these tools are still limited. In this work, we investigate whether an LLM, namely LlaMA-3.2-1b, has developed specialised mechanisms governing the selection of the locative preposition in simple copular clauses. To frame the problem as a next-token prediction objective, we introduce the Stranded Locative Preposition Selection task along with a small dataset aptly curated to test it. We make use of several MI tools to scan the model’s internals and relate their mechanisms to classic theory in Cognitive Linguistics, which assumes that the two basic locative prepositions in and on are the respective linguistic encoding of two different Image Schemas: Containment and Surface
Variation Is the Norm: Embracing Sociolinguistics in NLP
Anne-Marie Lutgen | Alistair Plum | Verena Blaschke | Barbara Plank | Christoph Purschke
Anne-Marie Lutgen | Alistair Plum | Verena Blaschke | Barbara Plank | Christoph Purschke
In Natural Language Processing (NLP), variation is typically seen as noise and “normalised away” before processing, even though it is an integral part of language. Conversely, studying language variation in social contexts is central to sociolinguistics. We present a framework to combine the sociolinguistic dimension of language with the technical dimension of NLP. We argue that by embracing sociolinguistics, variation can actively be included in a research setup, in turn informing the NLP side. To illustrate this, we provide a case study on Luxembourgish, an evolving language featuring a large amount of orthographic variation, demonstrating how NLP performance is impacted. The results show large discrepancies in the performance of models tested and fine-tuned on data with a large amount of orthographic variation in comparison to data closer to the (orthographic) standard. Furthermore, we provide a possible solution to improve the performance by including variation in the fine-tuning process. This case study highlights the importance of including variation in the research setup, as models are currently not robust to occurring variation. Our framework facilitates the inclusion of variation in the thought-process while also being grounded in the theoretical framework of sociolinguistics.
Appraisal Theory-Informed Emotion Prediction
Xiaowei Wang | Jayant Teotia | Rui Mao | Wandeep Kaur Ratan Singh | Sabrina Binti Tiun | Erik Cambria
Xiaowei Wang | Jayant Teotia | Rui Mao | Wandeep Kaur Ratan Singh | Sabrina Binti Tiun | Erik Cambria
Emotion Recognition in Conversation (ERC) focuses on identifying static emotional states, overlooking the cognitive mechanisms that drive emotional transitions. This work introduces a novel emotion prediction task grounded in Appraisal Theory, which conceptualizes emotion as a cognitive evaluation of expectations and their violations. To address this task, we develop a prompt-based reasoning framework that breaks emotional dynamics into three interpretable stages, e.g., expectation inference, violation detection, and emotion-shift prediction, thereby explaining not only which emotion is expressed, but also why it emerges. To examine whether LLMs exhibit human-like affective reasoning, we design six appraisal-informed prompting tasks and evaluate eight representative LLMs across four conversational corpora. A unified two-level evaluation, which measures both emotion classification and transition dynamics, reveals that explicit expectation cues improve accuracy by up to +2.4%, whereas violation-only cues often degrade performance. Our analysis uncovers a robust appraisal pattern across models and datasets: expectation construction is the primary contributor to accurate emotion prediction, while isolated violation cues tend to induce misattribution rather than improve causal reasoning. Beyond label accuracy, transition-level evaluation shows that LLMs capture emotion-shift direction above chance but exhibit a marked stability bias, over-predicting no-change trajectories and under-detecting fine-grained shifts. These findings demonstrate both the promise and the current limits of LLMs in appraisal-driven affective reasoning, and motivate a new cognitively-grounded research direction.
The Evolution of Philosophy: A Metaphorical Cognition Perspective
Rui Mao | Dapeng Chen | Zihao Huang | Xulang Zhang | Erik Cambria
Rui Mao | Dapeng Chen | Zihao Huang | Xulang Zhang | Erik Cambria
We present a large-scale study of philosophical cognition through the lens of Conceptual Metaphor Theory. Using a computational metaphor processing system that extracts target concepts, source concepts, and concept mappings from a curated corpus of 50+ canonical texts (300k sentences) spanning ten schools from antiquity to the late twentieth century, we quantify how metaphor organizes philosophical argument. We model temporal dynamics with year-level cosine series, authorial neighborhoods with PCA projections, and school signatures with heatmaps of normalized frequencies. The study demonstrates that the history of philosophy is structured by stable cross-domain schemas that are selectively recombined to address new problems.
Predicting States of Understanding in Explanatory Interactions Using Cognitive Load-Related Linguistic Cues
Yu Wang | Olcay Türk | Angela Grimminger | Hendrik Buschmeier
Yu Wang | Olcay Türk | Angela Grimminger | Hendrik Buschmeier
We investigate how verbal and nonverbal linguistic features, exhibited by speakers and listeners in dialogue, can contribute to predicting the listener’s state of understanding in explanatory interactions on a moment-by-moment basis. Specifically, we examine three linguistic cues related to cognitive load and hypothesised to correlate with listener understanding: the information value (operationalised with surprisal) and syntactic complexity of the speaker’s utterances, and the variation in the listener’s interactive gaze behaviour. Based on statistical analyses of the MUNDEX corpus of face-to-face dialogic board game explanations, we find that individual cues vary with the listener’s level of understanding. Listener states (’Understanding’, ’Partial Understanding’, ’Non-Understanding’ and ’Misunderstanding’) were self-annotated by the listeners using a retrospective video-recall method. The results of a subsequent classification experiment, involving two off-the-shelf classifiers and a fine-tuned German BERT-based multimodal classifier, demonstrate that prediction of these four states of understanding is generally possible and improves when the three linguistic cues are considered alongside textual features.
Figurative Language in Alzheimer’s Discourse: Linguistic and Neural Alignment in Clinical Narratives
Diana Kylymnyk | Vitória Hilgert Tomasel | Helena Caseli | Edward Watkins | Aline Villavicencio | Rodrigo Wilkens
Diana Kylymnyk | Vitória Hilgert Tomasel | Helena Caseli | Edward Watkins | Aline Villavicencio | Rodrigo Wilkens
Figurative language, including multiword expressions and metaphors, provides a sensitive lens on cognitive functioning but remains largely overlooked in computational studies of Alzheimer’s Disease (AD). This work investigates figurative-language patterns in AD and whether they can help in distinguishing AD from non-clinical discourse and whether a neural model encodes comparable linguistic tendencies. We propose a two-step framework that combines relevant linguistic features with neural representations. Figurative expressions are automatically identified using Large Language Models focusing on idiomaticity and metaphor detection. These figurative language indicators are integrated with lexical, syntactic, and readability features and used to train classifiers on the ADReSS dataset. Correlation and proxy-model analyses reveal significant alignment between linguistic indicators and model predictions: participants with AD produce fewer figurative constructions, lower lexical diversity, and more concrete language. The results obtained demonstrate that contextual embeddings implicitly encode linguistic cues associated with cognitive decline and highlight the value of figurative-language metrics for transparent and linguistically grounded clinical NLP.
Prompting Instruction-tuned LLMs for Semantic Similarity Values
Xander Akiko Snelder | Yunchong Huang | Jelke Bloem
Xander Akiko Snelder | Yunchong Huang | Jelke Bloem
The impressive few-shot performance of generative decoder transformer language models at novel tasks has raised interest in using them to estimate lexical-semantic properties of words, word pairs or multi-word expressions. We explore the task of eliciting semantic similarity scores between word pairs through prompting, comparing these scores to human benchmarks. We investigate different prompting approaches, different model architectures and different languages using the Dutch, English and Mandarin Chinese SimLex-999 benchmarks. The results show that prompting each word pair individually yields better correlations, and that models struggle with the distinction between similarity and relatedness, just as static and contextual word embedding models did. The new, open-weight gpt-oss-20b model yields the highest correlation with human ratings out of the models we evaluated.
Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective
Tianyi Zhang | David Traum
Tianyi Zhang | David Traum
In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by coherence, consistency, and shared understanding. However, many systems for open-domain and personalized dialogue use surface-level similarity metrics (e.g., BLEU, ROUGE, F1) as one of their main reporting measures, which fail to capture these deeper aspects of conversational quality. We re-examine a notable retrieval-augmented framework for personalized dialogue, LAPDOG, as a case study for evaluation methodology. Using both human and LLM-based judges, we identify limitations in current evaluation practices, including corrupted dialogue histories, contradictions between retrieved stories and persona, and incoherent response generation. Our results show that human and LLM judgments align closely but diverge from lexical similarity metrics, underscoring the need for cognitively grounded evaluation methods. Broadly, this work charts a path toward more reliable assessment frameworks for retrieval-augmented dialogue systems that better reflect the principles of natural human communication.
Evaluating Multimodal Large Language Model Narrative Interpretation through the Lens of Appraisal Theory
Jayant Teotia | Xiaowei Wang | Xulang Zhang | Rui Mao | Erik Cambria
Jayant Teotia | Xiaowei Wang | Xulang Zhang | Rui Mao | Erik Cambria
Narrative interpretation is an essential aspect of human cognition, enabling individuals to comprehend complex sequences of events, form emotional connections, and engage in nuanced social reasoning. At the heart of this interpretive ability lies emotional understanding, which cognitive scientists often frame through Appraisal Theory, a model that views emotions as the outcome of subjective evaluations of events in relation to goals, values, and beliefs. In this study, we explore whether multimodal large language models (MLLMs) are able to replicate aspects of this human-like narrative and emotional reasoning. Specifically, we examine how well MLLMs interpret visual narratives, with a focus on their ability to identify and appraise emotional content within scenes. We also investigate whether these models can utilize additional narrative descriptions generated by them to enhance their emotional recognition capabilities, as humans often do. To probe these questions, we conducted a series of experiments using two publicly available datasets, EMOTIC and HECO. Contrary to our expectations, our results reveal a consistent and noteworthy pattern: rather than improving the models’ performance, the inclusion of supplementary narrative or contextual information frequently diminishes their ability to accurately recognize emotions. This counterintuitive finding suggests that current MLLMs face significant challenges in integrating multimodal information in a coherent, context-sensitive way. These findings underscore key limitations in the emotional and narrative reasoning capabilities of existing MLLMs and highlight a critical gap between human cognitive processes and current AI approaches.
Mapping Liberty Metaphors across Cultures and Time
Sidney Suen | Rui Mao | Kenneth Kwok | Erik Cambria
Sidney Suen | Rui Mao | Kenneth Kwok | Erik Cambria
Cognitive metaphors provide a lens for understanding how societies construct and negotiate ideas, including liberty discourse. This study explores conceptual metaphors in liberty discourse by applying a scalable, corpus-driven approach for cognitive analysis. A curated list of thematic keywords related to liberty topics is used to extract relevant sentences from the Corpus of Historical American English (COHA) and the News on the Web (NOW) corpus. MetaPro, a framework grounded in Conceptual Metaphor Theory, processes these sentences to identify metaphorical mappings at scale. Embedding visualizations and frequency counts were applied to both corpora; in COHA, line graphs captured temporal shifts in metaphor usage across time, while in NOW, two-dimensional heatmaps highlighted spatial variation across countries. Selected example phrases illustrate how metaphorical mappings extend across diverse issues and domains. Thus, metaphor distributions and shifts provide a useful empirical lens for identifying changing thematic concerns in liberty discourse, offering a scalable, cognitively grounded method for cultural analysis across time and space. This demonstrates the value of computational methods for large-scale culture research.
Sensorimotor information plays a crucial role in the conceptual representation of linguistic knowledge. While previous studies have established sensorimotor norms for nouns and adjectives, little is known about how Chinese numeral classifiers encode perceptual and action-based experiences. The present study constructs the first large-scale sensorimotor norms for Chinese classifiers, collecting perceptual and action ratings for 357 classifiers from 288 native Chinese speakers. Participants evaluated each classifier along six perceptual modalities (vision, hearing, taste, smell, touch, and interoception) and five action effectors (foot/leg, hand/arm, mouth/throat, head, and torso). The resulting dataset provides detailed sensorimotor profiles for each classifier and reveals systematic mappings between classifier semantics and embodied dimensions. The findings demonstrate that Chinese classifiers are not purely syntactic markers but encode distinct sensorimotor features grounded in perceptual and motor systems, highlighting the embodied foundation of the classifier system and offering valuable resources for future psycholinguistic and computational modelling studies of Chinese semantics.
DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance
Ali Khoramfar | Ali Ramezani | Mohammad Mahdi Mohajeri | Mohammad Javad Dousti | Majid Nili Ahmadabadi | Heshaam Faili
Ali Khoramfar | Ali Ramezani | Mohammad Mahdi Mohajeri | Mohammad Javad Dousti | Majid Nili Ahmadabadi | Heshaam Faili
While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge this gap, we introduce DeepQuestion, a scalable, automated framework that systematically elevates the cognitive complexity of existing datasets through controlled task transformations grounded in explicit cognitive hierarchies. Based on Bloom’s taxonomy, DeepQuestion generates (1) scenario-based problems to test the application of knowledge in noisy, realistic contexts, and (2) instruction-based prompts that require models to create new questions from a given solution path, assessing synthesis and evaluation skills. Our extensive evaluation across ten leading open-source and proprietary models, covering both general-purpose and reasoning LLMs, reveals a stark performance decline—with accuracy dropping by up to 70%—as tasks ascend the cognitive hierarchy across evaluation settings. These findings underscore that current benchmarks overestimate true reasoning abilities and highlight the critical need for cognitively diverse evaluations to guide future LLM development.
Pragmatic Modelling in Language Learning: Caregiver Question-Answer Feedback in Child-Directed Dialogue
Maryam Bala | Johannes Heim | Elspeth Edelstein | Arabella Sinclair
Maryam Bala | Johannes Heim | Elspeth Edelstein | Arabella Sinclair
In language development, children learn to form Question–Answer (QA) sequences through caregiver feedback that adapts dynamically to their evolving linguistic abilities. Using expert annotated child-caregiver interaction, we examine four feedback types that guide children’s acquisition of adult-like QA behaviour: caregiver instructions through reformulating and affirming a child’s output as well as caregiver demonstrations through exemplifying and modelling adult-like behaviour. Our analysis reveals that feedback incidence, frequency and complexity progress and adapt over the course of development, akin to a tailored curriculum for pragmatic development. We release our annotated dataset which offers a rich resource for studying pragmatic feedback and provides the first large-scale empirical evidence of adaptive, tailored caregiver feedback on QA behaviour.
Modular Approach to Automating Morphological Components in Grammar Engineering
Ekaterina Voloshina | Krasimir Angelov
Ekaterina Voloshina | Krasimir Angelov
Creating formal grammars is a time-consuming and complex task. We present a method to automatically create the morphological components of a formal grammar in Grammatical Framework. Our method is linguistically interpretable and modular, consisting of three stages: paradigm construction, extraction of inflectional classes, and prediction of inflectional classes. The modular structure allows human interventions after each stage. Moreover, our method supports encoding pre-existing language knowledge in form of Python APIs. Experiments show that automatically extracted morphological rules yield results comparable with manual grammars and that incorporating prior linguistic knowledge leads to improvement in low-resourced scenarios. Our findings show that our method simplifies the process of grammar development while preserving quality and interpretability.
MorfFlex: Handling Rich Morphology
Jaroslava Hlaváčová | Marie Mikulová | Barbora Štěpánková | Milan Straka | Jan Hajič
Jaroslava Hlaváčová | Marie Mikulová | Barbora Štěpánková | Milan Straka | Jan Hajič
We present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the primary example of MorfFlex in use we introduce MorfFlex CZ, a morphological dictionary of Czech. It is distributed as a simple, unstructured list of <wordform, lemma, tag> triplets, however, its manually maintained, unpublished source files and conversion scripts encode a sophisticated system of inflectional and derivational patterns. These patterns dramatically reduce the otherwise enormous size of the dictionary, which currently contains over 100 million wordforms and more than 1 million lemmas. The MorfFlex CZ dictionary serves as an essential resource for ensuring the consistency of manual morphological annotation in the Prague Dependency Treebanks and underpins state-of-the-art automatic tools such as MorphoDiTa. In this paper, we focus on: (i) presenting an effective method for managing the rich morphological system within the dictionary, and (ii) demonstrating the utility of such a language resource for maintaining annotation consistency in corpora and supporting the development of advanced NLP applications.
Using Valency Inheritance in Building a Valency Lexicon
Václava Kettnerová | Veronika Kolářová | Jiří Mírovský | Michal Olbrich
Václava Kettnerová | Veronika Kolářová | Jiří Mírovský | Michal Olbrich
Derived words often share certain characteristics with their base words, which leads to the idea that identical properties are inherited from the base words. These properties also cover valency. Valency inheritance has not been used to automatically build lexical resources providing information on valency, the manual annotation of which requires significant human effort. In this paper, we propose a procedure for generating valency frames of selected semantic categories of Czech nouns and adjectives exhibiting a significant level of valency inheritance, thus covering the productive and systemic core of the lexicon. Based on a semiautomatic comparison of the noun and adjectival valency frames from NomVallex and the verbal valency frames from VALLEX, rules describing valency changes in the valency frames of noun and adjectival derivatives are formulated. The conditions imposed by the rules on valency frames identify individual base lemmas in these lexicons for which direct noun and adjectival derivatives are searched in DeriNet. Based on the changes in valency determined in the rules, more than 23,000 valency frames assigned to more than 10,000 noun and adjectival derivatives were derived, achieving high accuracy. These valency frames were included in DeriVallex, a database providing a solid basis for extending current lexical resources.
From CHAT to Coded CoNLL-U: A Reproducible Pipeline for the Syntactic Annotation and Querying of Child Language Data
Achim Stein
Achim Stein
The CHILDES database is a core resource for language acquisition research, yet its CHAT format poses significant challenges for modern computational analysis. To address this, we present a reproducible, open-source pipeline that transforms CHAT transcripts into annotated tabular (CSV) and CoNLL-U formats. Its core script, childes.py, automates the conversion and integrates part-of-speech tagging and dependency parsing. A key innovation is dql.py, a tool that uses a Grew dependency query language to systematically add user-defined linguistic codings to the parsed data. While the script is parametrised for various languages, the pipeline’s utility is demonstrated by applying it to the French CHILDES corpus to conduct a large-scale analysis of object clitic production. The resulting structured data reveals clear developmental trajectories, such as the gradual convergence of children’s dative clitic usage towards the adult input. The workflow and the resources it generates facilitate reproducible, data-driven research in language acquisition.
Language use is different across different language communities. Social media provides a rich source for studying how language varies, as it contains large data for a wide variety of sub-communities. In this paper, we study language usage on Danish TikTok. TikTok is a video-based platform, but most users are mainly active in the text-based comment sections. With the goal of analyzing language usage on this language variety, we contribute: 1) the first Danish social media treebank annotated for Universal Dependencies 2) evaluation of a variety of parsers using the new treebank, showing that cross-lingual in-domain data provides a valuable signal 3) a comparison of syntactic trends on standard Danish languages and TikTok language.
Adaptive Chunking: Optimizing Chunking-Method Selection for RAG
Paulo Roberto de Moura Júnior | Jean Lelong | Annabelle Blangero
Paulo Roberto de Moura Júnior | Jean Lelong | Annabelle Blangero
The effectiveness of Retrieval-Augmented Generation (RAG) is highly dependent on how documents are chunked, that is, segmented into smaller units for indexing and retrieval. Yet, commonly used “one-size-fits-all” approaches often fail to capture the nuanced structure and semantics of diverse texts. Despite its central role, chunking lacks a dedicated evaluation framework, making it difficult to assess and compare strategies independently of downstream performance. We challenge this paradigm by introducing Adaptive Chunking, a framework that selects the most suitable chunking strategy for each document based on a set of five novel intrinsic, document-based metrics: References Completeness (RC), Intrachunk Cohesion (ICC), Document Contextual Coherence (DCC), Block Integrity (BI), and Size Compliance (SC), which directly assess chunking quality across key dimensions. To support this framework, we also introduce two new chunkers, an LLM-regex splitter and a split-then-merge recursive splitter, alongside targeted post-processing techniques. On a diverse corpus spanning legal, technical, and social science domains, our metric-guided adaptive method significantly improves downstream RAG performance. Without changing models or prompts, our framework increases RAG outcomes, raising answers correctness to 72% (from 62-64%) and increasing the number of successfully answered questions by over 30% (65 vs. 49). These results demonstrate that adaptive, document-aware chunking, guided by a complementary suite of intrinsic metrics, offers a practical and effective path to more robust RAG systems. Code available at https://github.com/ekimetrics/adaptive-chunking.
Do Large Language Models Grasp the Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgish
Lujun LI | Yewei Song | Lama Sleem | Yiqun Wang | Yangjie Xu | Cedric LOTHRITZ | Niccolo’ Gentile | Radu State | Tegawendé F. Bissyandé | Jacques Klein
Lujun LI | Yewei Song | Lama Sleem | Yiqun Wang | Yangjie Xu | Cedric LOTHRITZ | Niccolo’ Gentile | Radu State | Tegawendé F. Bissyandé | Jacques Klein
Grammar refers to the system of rules that governs the structural organization and the semantic relations among linguistic units such as sentences, phrases, and words within a given language. In natural language processing, there remains a notable scarcity of grammar-focused evaluation protocols, a gap that is even more pronounced for low-resource languages. Moreover, the extent to which large language models genuinely comprehend grammatical structure, especially the mapping between syntactic structures and meanings remains under debate. To investigate this issue, we propose a Grammar-Book–Guided evaluation pipeline intended to provide a systematic and generalizable framework for grammar evaluation consisting of four key stages, and in this work we take Luxembourgish as a case study. The results show a weak positive correlation between translation performance and grammatical understanding, indicating that strong translations do not necessarily imply deep grammatical competence. Larger models perform well overall due to their semantic strength but remain weak in morphology and syntax, struggling particularly with Minimal Pair tasks, while strong reasoning ability offers a promising way to enhance their grammatical understanding.
Survey of Tools for Manual Linguistic Annotation: Supporting Diversity through Interactive Exploration
Ludovica Pannitto | Kaja Dobrovoljc Zor | Bruno Guillaume
Ludovica Pannitto | Kaja Dobrovoljc Zor | Bruno Guillaume
Manual annotation tools are core infrastructure for corpus creation, enabling the development of linguistically informed language resources relevant for both linguistic discovery and computational applications. We present a comprehensive survey of 21 tools supporting morphosyntactic and multi-word expression annotation, systematically documenting more than 50 features relevant for annotation workflows—from software architecture and usability to linguistic coverage and annotation scope. The survey results are published as an open dataset and made accessible through an interactive online platform that allows users to filter and compare tools according to their specific needs. Our initial analysis highlights a robust and open ecosystem of annotation tools, but advanced needs for complex and language-independent annotation are inconsistently addressed.
TextLens & LeTTuce: Automated Corpus Annotation and Multilingual Tagging as a Service
Cynthia Van Hee | Jonas Doumen | Vincent Prins | Pranaydeep Singh | Vincent Vandeghinste | Els Lefever
Cynthia Van Hee | Jonas Doumen | Vincent Prins | Pranaydeep Singh | Vincent Vandeghinste | Els Lefever
We present TextLens, a web-based platform for automated linguistic annotation designed to lower technical barriers for researchers in digital humanities, linguistics and translation studies. Hosted by the Dutch Language Institute (INT), TextLens allows users to upload and annotate corpora in a variety of formats (.txt, .tsv, CoNLL-U, FoLiA, TEI, and NAF) using state-of-the-art NLP tools, without the need for local installation or computational resources. The platform supports multilingual data processing and provides a persistent dashboard for managing, monitoring and sharing annotation projects. Alongside this service, we introduce the LeTTuce-PoS Dataset, a new multilingual, manually annotated dataset for part-of-speech tagging in English, French, Dutch and German, covering multiple genres and offering a valuable resource to the research community. This paper also reports benchmark results for different PoS taggers (LeTs Preprocess, LeTTuce, spaCy and Stanza) on the dataset. Together, TextLens and the LeTTuce-PoS Dataset provide an accessible, scalable platform for high-quality annotation and a robust multilingual dataset that support comparable and reproducible research in multilingual contexts.
The Corpus of Contemporary Polish — a New Reference Corpus with Rich Syntactic Annotations
Witold Kieraś | Małgorzata Marciniak | Marcin Woliński | Katarzyna Krasnowska-Kieraś | Marek Łaziński
Witold Kieraś | Małgorzata Marciniak | Marcin Woliński | Katarzyna Krasnowska-Kieraś | Marek Łaziński
In the paper, we describe the Corpus of Contemporary Polish (KWJP) and its rich syntactic annotation. The corpus covers a wide range of text originally published between 2011 and 2020. Although it carries on the idea of providing up-to-date reference corpora of Polish initiated by the National Corpus of Polish (NKJP) project, the principles underlying its development are not the same. In this article, we outline the different choices that affect corpora content and give an explanation for them. The article focuses mainly on the description of annotation layers in KWJP which are generated with a neural network based tool specially developed for this purpose. We describe in details syntactic structure annotation, which is represented by hybrid trees combining information typical to constituency and dependency trees. Finally, we provide several examples showing how annotation with hybrid trees facilitates querying and effective searching for information in the corpus.
Prague Dependency Treebank - Consolidated 2.0: Enriching a Complex Annotation Scheme
Marie Mikulová | Jiří Mírovský | Milan Straka | Pavlína Synková | Jan Štěpánek | Barbora Štěpánková | Jan Hajič
Marie Mikulová | Jiří Mírovský | Milan Straka | Pavlína Synková | Jan Štěpánek | Barbora Štěpánková | Jan Hajič
The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relation. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence.
Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies
Marie Mikulová | Barbora Štěpánková | Daniel Zeman | Jan Štěpánek | Milan Straka | Jan Hajič
Marie Mikulová | Barbora Štěpánková | Daniel Zeman | Jan Štěpánek | Milan Straka | Jan Hajič
Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague Dependency Treebank being order of magnitude larger than most other UD treebanks. More recently, three other datasets from the Prague family were added and the annotations thoroughly revisited, forming the “Prague Dependency Treebank-Consolidated” (PDT-C). In comparison to the original PDT, PDT-C is more than twice as large, but it is also much more diverse in terms of genres and domains. In this paper, we describe the conversion of the new resource to Universal Dependencies. While the two annotation schemes are relatively similar at the first sight, there are numerous small differences in topology of the dependency structures and in granularity of the POS and relation type inventories. We demonstrate a selection of such differences on examples, discuss the diverging motivations, as well as ways to overcome the differences during conversion. We argue that while PDT is less “universal” and more tightly bound to one language, its multi-layer annotation is rich and provides all information needed for basic UD trees, and much more.
Encoding Logical Relations of Chinese Complex Sentences within the Universal Dependencies Framework
Hongpu Zhu | Hongzhi Xu
Hongpu Zhu | Hongzhi Xu
Clauses in complex sentences always entail certain logical relations such as conjunctive, causative, and concessive. Such logical relations, however, are not properly represented in the universal dependencies (UD) framework, being collapsed into a adverbial clause (advcl) or clausal complement (ccomp) relation between clausal heads. This study extends the UD framework by encoding 13 logical relations. With the new framework, which is structurally identical to UD, we construct a training corpus containing about 1,769 sentences extracted from Chinese newswire and annotated an existing Chinese corpus (GSD-simp test) in UD as a test set. We trained a BERT-based biaffine parser and fine-tuned the Qwen-3 model with the training corpus and evaluated the models on the UD test data. They are compared against four general purpose LLMs including GPT-4o, GPT-5, Claude 4 and DeepSeek V3.2. We find that the fine-tuned Qwen-3-8B model achieves a UAS/LAS of 0.840/0.757, higher than the BERT-based parser and the general purpose LLMs. The results confirm the feasibility of our framework and highlight the inherent challenges of parsing hierarchical and implicit inter-clause relations.
Unsupervised Labelling of Mutation Triggers in Welsh
Nicolás Gutiérrez-Rolón | Fernando Alva-Manchego
Nicolás Gutiérrez-Rolón | Fernando Alva-Manchego
Initial consonant mutation is a key feature of Welsh, but its complexity poses significant challenges for both language learners and natural language processing (NLP) systems. While existing tools can reliably detect mutated forms, they provide no information about why a mutation occurs, i.e. what grammatical or lexical factors trigger the change. This paper introduces the novel task of mutation trigger labelling, representing the first computational attempt to analyse and explain the reasons behind Welsh mutations. Two preliminary approaches are explored: (i) a linguistically-informed rule-based system integrating Constraint Grammar rules, and (ii) large language models (LLMs), prompted in few-shot settings. Our experiments test the feasibility of automatically identifying and labelling linguistic triggers behind Welsh mutations using a dataset constructed from grammar reference books and public corpora, and establish baseline insights into how context-aware mutation analysis can be achieved. By framing mutation trigger labelling as a linguistic computational problem, this work lays important groundwork within Welsh NLP and contributes to the broader development of explainable grammatical analysis for low-resource languages.
In this paper, we present a new Universal Dependencies treebank for Uzbek language(UzUDT) developed as a gold-standard resource with full manual annotation. The treebank includes 684 sentences (7,582 tokens) from Uzbek literary texts, and is larger and more domain-diverse than the existing Uzbek UD treebank. The corpus was developed through rigorous multi-annotator adjudication, achieving very high inter-annotator agreement (multi-rater agreement coefficients >0.90) across lemmatization, PoS tagging, and morphological features. Alongside comprehensive corpus profiling, we establish robust computational baselines by evaluating graph-based (Stanza) and transition-based (spaCy) parsing architectures using both static and monolingual contextual embeddings. Our evaluations reveal a critical architectural trade-off for low-resource agglutinative parsing: joint transition-based models excel at morphosyntactic tagging, whereas graph-based models remain strictly superior for resolving complex structural dependencies. Furthermore, we demonstrate that cross-treebank data augmentation yields substantial, synergistic accuracy gains. The resource provides a much-needed high-quality treebank for Uzbek to assist in developing better NLP tools and to enable linguistic research in the low-resource language
BRAGD: Constrained Multi-Label POS Tagging for Faroese
Annika Simonsen | Barbara Scalvini | Uni Johannesen | Iben Nyholm Debess | Hafsteinn Einarsson | Vésteinn Snæbjarnarson
Annika Simonsen | Barbara Scalvini | Uni Johannesen | Iben Nyholm Debess | Hafsteinn Einarsson | Vésteinn Snæbjarnarson
We present the first multi-label part-of-speech (POS) tagger for Faroese using linguistically-informed constraints, addressing the data sparsity problem inherent in compound tag approaches. We propose the BRAGD tagset, which decomposes compound morphological tags into independent features (word class, gender, number, case, etc.). The BRAGD tagset is the third iteration of a tagset previously released for Faroese, with substantial modifications that are better aligned with Faroese grammar. We annotate the previously released Sosialurin corpus with the tagset, as well as a new annotated out-of-domain test corpus of 500 sentences from more varied and contemporary texts. To train the tagger, we use a constrained loss function that dynamically masks morphologically invalid features based on the word class (noun, verb, adjective, etc.). We fine-tune a Scandinavian transformer language model using the constrained multi-label loss, achieving an overall accuracy of 97.5%. We find that models trained with multi-label loss perform better, converge faster, and show significantly lower error rates on out-of-domain data than single-label approaches or previously reported methods for Faroese POS tagging. This confirms that the multi-label approach learns robust morphological patterns rather than memorizing domain-specific tag distributions. We release models, code, and the systematically revised Sosialurin-BRAGD corpus, featuring the new BRAGD tagset and a new out-of-domain evaluation corpus from diverse and contemporary text types.
Syntactic Sugar for Syntactic Queries: Sequential Representations for Dependency Queries
Niklas Deworetzki | Arianna Masciolini
Niklas Deworetzki | Arianna Masciolini
Syntactic query languages such as Grew and dep_search allow looking for grammatical patterns in linguistically annotated corpora. However, these languages are often unsupported by large-scale corpus management tools, where queries are of an essentially sequential nature. In this paper, we present CQP/Tree, a tool to convert syntactic queries into CQL, the Corpus Query Language used in Corpus Workbench, SketchEngine, Korp and several other such systems. In this framework, syntactic queries act as syntactic sugar: they allow expressing complex CQL queries in a more readable and concise fashion, thus bridging the gap between expressive linguistic search and large-scale corpora. CQP/Tree is available as a web and command-line tool, as well as an open source Python library.
Context Is (Almost) Everything: Llama-3 on Structured Output and AMR Parsing
Maja Buljan | Stephan Oepen | Lilja Øvrelid
Maja Buljan | Stephan Oepen | Lilja Øvrelid
This paper evaluates the ability of an open-source LLM (Llama-3.1) to compute sentence-level semantics and encode it in formal language. We here compare two versions of the model on the task of generating a meaning representation graph for a given English sentence in the form of Abstract Meaning Representation. We explore the model’s in-context learning capability, comparing zero-shot prompting to few-shot demonstrations of varying levels of specificity. We find that Llama-3.1 frequently makes errors when reproducing the syntactic structure of both seen and unseen structured output, and that it only achieves near-SotA parsing performance when shown highly specific demonstrations similar in structure to the target sentence graph. We include an in-depth analysis of the model output, considering performance through the lens of fine-grained semantic phenomena, graph properties (e.g. top node accuracy), and graph complexity.
Low German (Low Saxon, ISO 639-2 nds) is an underresourced West Germanic language spoken in Northern Germany (Plattdütsch), in the Netherlands (Nedersaksisch) and in an international diaspora (Plautdietsch, Pomerano, etc.). As a minority language, it is under pressure from the respective national languages, and considered threatened. Although NLP and digital language resources might play a role in facilitating the use of the language on the web and to support intergenerational transmission, no NLP tools are known to exist, and no adequate corpora that such tools could be trained on. This paper describes the construction of a novel corpus of North Markian, a dialect of East Low German, its morphosyntactic annotation and morphological analysis, and in particular explores methods to bootstrap and develop such resources in the face of a complete lack of training data.
Cross-Dataset Inconsistencies in Morphological Annotation: Evidence from Universal Dependencies
Vlasta Ohlídalová
Vlasta Ohlídalová
Ensuring annotation consistency is a challenging task in language dataset development. While difficulty is typically increasing at higher levels of linguistic complexity, we show that it is a critical issue even for fundamental linguistic tasks such as morphological annotation. Contrary to previous research that targeted intra-dataset inconsistencies, this study investigates inconsistencies across various pre-existing datasets for the same language. On the example of Universal Dependencies datasets, we examined what morphological categories exhibit the most disagreement. The analysis revealed that there are specific categories with low inconsistency score that indicates good agreement on these features (namely Case, Gender, Number and to a lesser extent Animacy). On the other hand, the Part-of-Speech (UPOS) tag stands out as a “red flag” due to high inconsistency score. Analysis of the most frequent inconsistencies suggest that they are dataset-specific artifacts rather than inherently language-specific phenomena.
Improving Latvian Morphosyntactic Parsing with Pretrained Encoders and Analyzer-Constrained Decoding
Arturs Znotins
Arturs Znotins
We present a systematic evaluation of Latvian morphosyntactic parsing with pretrained transformer encoders in a unified joint architecture for tagging, lemmatization, and dependency parsing. We benchmark multilingual and Latvian-specific models and show that language-specific adaptation, even with modest in-language data, substantially improves performance. We further demonstrate that factored morphological modeling improves robustness and that integrating a Latvian morphological analyzer through constrained decoding yields consistent gains in XPOS tagging and lemmatization. The best system achieves new state-of-the-art results, reaching 95.22% XPOS accuracy, 98.72% lemma accuracy, and 93.19% LAS.
CommonMorph: Participatory Morphological Documentation Platform
Aso Mahmudi | Sina Ahmadi | Kemal Maulana Kurniawan | Rico Sennrich | Eduard H. Hovy | Ekaterina Vylomova
Aso Mahmudi | Sina Ahmadi | Kemal Maulana Kurniawan | Rico Sennrich | Eduard H. Hovy | Ekaterina Vylomova
Collecting and annotating morphological data present significant challenges, requiring linguistic expertise, methodological rigour, and substantial resources. These barriers are particularly acute for low-resource languages and varieties. To accelerate this process, we introduce CommonMorph, a comprehensive platform that streamlines morphological data collection development through a three-tiered approach: expert linguistic definition, contributor elicitation, and community validation. The platform minimises manual work by incorporating active learning, annotation suggestions, and tools to import and adapt materials from related languages. It accommodates diverse morphological systems, including fusional, agglutinative, and root-and-pattern morphologies. Its open-source design and UniMorph-compatible outputs ensure accessibility and interoperability with NLP tools. Our platform is accessible at https://common-morph.com, offering a replicable model for preserving linguistic diversity through collaborative technology.
Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies
Giuseppe Samo | Paola Merlo
Giuseppe Samo | Paola Merlo
Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this work, we present curated paradigm-based datasets for four languages, designed to probe systematic cross-sentence knowledge of verb alternations (change-of-state and object-drop constructions in English, German and Italian, and Hebrew binyanim). The datasets comprise thousands of the Blackbird Language Matrices (BLMs) problems. The BLM task – an RPM/ARC-like task devised specifically for language – is a controlled linguistic puzzle where models must select the sentence that completes a pattern according to syntactic and semantic rules. We introduce three types of templates varying in complexity and apply linguistically-informed data augmentation strategies across synthetic and natural data. We provide simple baseline performance results across English, Italian, German, and Hebrew, that demonstrate the diagnostic usefulness of the datasets.
A Large and Balanced Multi-Domain Arabic Corpus Annotated for Morphology, Syntax, and Readability
Khalid N. Elmadani | Adel Mahmoud Wizani | Hanada Taha Thomure | Nizar Habash
Khalid N. Elmadani | Adel Mahmoud Wizani | Hanada Taha Thomure | Nizar Habash
We present BAREC-10M, an expanded version of the Balanced Arabic Readability Evaluation Corpus (BAREC). This new release extends the original 1M-word corpus to 10 million words and broadens its scope to include balanced multi-domain coverage annotated for morphology, syntax, and readability. The corpus integrates 45 sub-corpora drawn from diverse sources, including news, educational materials, literature, children’s texts, and religious discourse. Each text is labeled for domain, readership level, and genre, and automatically analyzed using state-of-the-art morphological and syntactic tools. To enhance coverage of underrepresented varieties, we manually digitized and included children’s materials, magazines, and curriculum-based content. The resulting dataset provides a balanced resource for studying Arabic linguistic variation across styles, audiences, and levels of complexity.
Precision computational grammars encode detailed linguistic analyses and compositional semantics that support rigorous investigation of grammatical phenomena, but their development requires substantial expertise and maintenance. To ensure long-term sustainability and accessibility of these resources, we present the DELPH-IN Grammary, a curated collection of twenty three HPSG grammars spanning seventeen languages and eight language families. The repository includes mature broad-coverage grammars (English, German, Japanese, Norwegian, Spanish) with associated treebanks, as well as grammars for typologically diverse and less-resourced languages. Each grammar is standardized with metadata, compiled using the ACE parser/generator, and loaded into the Linguistic Type Database for detailed inspection. Following FAIR principles, all resources are version-controlled and archived on Zenodo with annual releases synchronized to community development cycles. The Grammary enables reproducible grammar research, cross-linguistic typological studies, semantic parsing development, and grammar engineering pedagogy, providing the depth and theoretical grounding that complements data-driven approaches in computational linguistics. Our goal is to establish a sustainable model for preserving these valuable resources which bridge the critical gap between theoretical linguistics and empirical corpus based research. The Grammary is available at https://github.com/delph-in/grammary (Zenodo doi: zenodo.18945956).
Morphemes without Borders: Evaluating Root–Pattern Morphology in Arabic Tokenizers and LLMs
Yara Yousif Alakeel | Chatrine Qwaider | Hanan Aldarmaki | Sawsan Alqahtani
Yara Yousif Alakeel | Chatrine Qwaider | Hanan Aldarmaki | Sawsan Alqahtani
This work investigates how effectively large language models (LLMs) and their tokenization schemes represent and generate Arabic root–pattern morphology, probing whether they capture genuine morphological structure or rely on surface memorization. Arabic morphological system provides a rich testbed for analyzing how LLMs handle complex, non-concatenative forms and how tokenization choices influence this process. Our study begins with an evaluation of morphological fidelity across Arabic and multilingual tokenizers against gold-standard segmentation, followed by an analysis of LLM performance in productive root–pattern generation using a newly developed benchmark. Our findings across seven Arabic-centric and multilingual LLMs and their respective tokenizers reveal that tokenizer morphological alignment is not necessary nor sufficient for morphological generation, which questions the role of morphological tokenization in downstream performance.
APODICTUS: Automatic Processing of DICTionary Update candidateS
Felix Blessing | Johannes S. Sax | Julian Kaufmann | Wei Zhao | Nikolay Arefyev | Dominik Schlechtweg
Felix Blessing | Johannes S. Sax | Julian Kaufmann | Wei Zhao | Nikolay Arefyev | Dominik Schlechtweg
Dictionaries have to be regularly updated. Some dictionary-makers gather proposals for updates of sense entries in internal databases. We automate the process of verifying and prioritizing such sense proposals, and facilitate their addition to a dictionary, by building a sophisticated processing pipeline relying on state-of-the-art language models. Our pipeline presents the first systematic, large-scale, and comprehensive solution for processing candidates for inclusion in a dictionary, which is tested in an industry-relevant context. We conduct several experiments to evaluate the pipeline and provide an annotated dataset for future work. Model performance is acceptable for words which are not yet in the dictionary, but low for in-dictionary words. Through an error analysis and model component ablation, we gain further insight on directions of future model improvements.
We evaluate a focused test collection at the intersection of part-of-speech tagging and word-sense disambiguation. The collection targets words such as train, novel, and lean, where part-of-speech contrasts align with clear meaning differences. We use it to detect regressions across tagger versions, track quantitative and qualitative progress over time, and test robustness to orthographic variation. Experiments with the Stanford and TnT taggers show 68% accuracy, compared with 92% for a recent spaCy transformer model. Earlier taggers erred mainly on noun–verb distinctions; spaCy’s errors more often involve noun–adjective distinctions. Uppercase text roughly doubles error rates for all taggers. We discuss common problems and propose directions for future testing.
Creating a Hybrid Rule and Neural Network Based Semantic Tagger Using Silver Standard Data: The PyMUSAS Framework for Multilingual Semantic Annotation
Andrew Moore | Paul Rayson | Dawn Archer | Tim Czerniak | Dawn Knight | Daisy Monika Lal | Gearóid Ó Donnchadha | Mícheál J. Ó Meachair | Scott Piao | Elaine Uí Dhonnchadha | Johanna Vuorinen | Yan Yabo | Xiaobin Yang
Andrew Moore | Paul Rayson | Dawn Archer | Tim Czerniak | Dawn Knight | Daisy Monika Lal | Gearóid Ó Donnchadha | Mícheál J. Ó Meachair | Scott Piao | Elaine Uí Dhonnchadha | Johanna Vuorinen | Yan Yabo | Xiaobin Yang
Word Sense Disambiguation (WSD) has been widely evaluated using the semantic frameworks of WordNet, BabelNet, and the Oxford Dictionary of English. However, for the UCREL Semantic Analysis System (USAS) framework, no open extensive evaluation has been performed beyond lexical coverage or single language evaluation. In this work, we perform the largest semantic tagging evaluation of the rule based system that uses the lexical resources in the USAS framework covering five different languages using four existing datasets and one novel Chinese dataset. We create a new silver labelled English dataset, to overcome the lack of manually tagged training data, that we train and evaluate various mono and multilingual neural models in both mono and cross-lingual evaluation setups with comparisons to their rule based counterparts, and show how a rule based system can be enhanced with a neural network model. The resulting neural network models, including the data they were trained on, the Chinese evaluation dataset, and all of the code will be released as open resources.
Scare Quotes as Markers of "Questionable" Word Usages and Misalignment in Conversation: An Annotation Study
Aina Garí Soler | Juan Carlos Zevallos Huaco | Matthieu Labeau | Chloé Clavel
Aina Garí Soler | Juan Carlos Zevallos Huaco | Matthieu Labeau | Chloé Clavel
Scare quotes are a subtle yet powerful device: they can mark irony, distance, or disagreement about word meaning or lexical choices. We present a large-scale manual annotation of quoted word usages focused on the scare versus non-scare quote distinction as well as on their role in managing (mis)alignment in conversation. Our analysis reveals that scare quotes can mark problematic word usages, and they are often used to contest or criticize other speakers’ word choices. However, non-scare, meta-linguistic usages of quotes are also often involved in explicit efforts toward lexico-semantic alignment.
Modeling Clinical Uncertainty in Radiology Reports: From Explicit Uncertainty Markers to Implicit Reasoning Pathways
Paloma Rabaey | Jong Hak Moon | Jung-Oh Lee | Min Gwan Kim | Hangyul Yoon | Thomas Demeester | Edward Choi
Paloma Rabaey | Jong Hak Moon | Jung-Oh Lee | Min Gwan Kim | Hangyul Yoon | Thomas Demeester | Edward Choi
Radiology reports are invaluable for clinical decision-making and hold great potential for automated analysis when structured into machine-readable formats. These reports often contain uncertainty, which we categorize into two distinct types: (i) Explicit uncertainty reflects doubt about the presence or absence of findings, conveyed through hedging phrases. These vary in meaning depending on the context, making rule-based systems insufficient to quantify the level of uncertainty for specific findings; (ii) Implicit uncertainty arises when radiologists omit parts of their reasoning, recording only key findings or diagnoses. Here, it is often unclear whether omitted findings are truly absent or simply unmentioned for brevity. We address these challenges with a two-part framework. We quantify explicit uncertainty by creating an expert-validated, LLM-based reference ranking of common hedging phrases, and mapping each finding to a probability value based on this reference. In addition, we model implicit uncertainty through an expansion framework that systematically adds characteristic sub-findings derived from expert-defined diagnostic pathways for 14 common diagnoses. Using these methods, we release Lunguage++, an expanded, uncertainty-aware version of the Lunguage benchmark of fine-grained structured radiology reports. This enriched resource enables uncertainty-aware image classification, faithful diagnostic reasoning, and new investigations into the clinical impact of diagnostic uncertainty.
ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination
Wajdi Zaghouani | Shimaa Amer Ibrahim | Mabrouka Bessghaier | Houda Bouamor
Wajdi Zaghouani | Shimaa Amer Ibrahim | Mabrouka Bessghaier | Houda Bouamor
We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014–2024) discussing racism and discrimination. Unlike existing Twitter-centric datasets, ArabDiscrim integrates platform-native engagement signals, including reactions, shares, comments, and page metadata, enabling joint analysis of language and audience response. The resource includes 200 curated terms (100 racism, 100 discrimination) with morphological regex families (13+ inflections per lemma), and 20 discrimination axes capturing identity-based grounds for unequal treatment. It also provides explicit attribution patterns. Released under a restricted research-use license for ethical compliance with platform terms, ArabDiscrim supports weak supervision, axis-aware sampling, and platform ecology research. By bridging lexical depth and ecological validity, it establishes a foundation for fairness-oriented, platform-aware Arabic NLP.
DAMETA: An LLM Benchmark for Danish Metaphor Interpretation with Systematically Varied Distractors
Nina Skovgaard Schneidermann | Sanni Nimb | Nathalie Carmen Hau Norman | Sussi Olsen | Bolette Pedersen
Nina Skovgaard Schneidermann | Sanni Nimb | Nathalie Carmen Hau Norman | Sussi Olsen | Bolette Pedersen
We present DAMETA, the first evaluation benchmark for Danish metaphor interpretation in language models, derived from the following sources: an annotated corpus (the Dafig Corpus), the Danish dictionary (DDO) and culture reviews in Danish newspapers. Each of the 900 data instances contains a sentence with a metaphorical target word and four human-created paraphrase options; one correct interpretation and three systematic errors or distractors: i) a false literal paraphrase (typically concrete), ii) a false figurative paraphrase (typically abstract), and iii) a false contradictory paraphrase. The benchmark is tested on seven language models, and 5% of the data is further tested on humans for comparison. Results show, among others, that when informed in the prompt that the target word is a metaphor, the models tend to be most distracted by the false figurative paraphrase; in contrast, when uninformed about the metaphorical setting, the models are more distracted by the false literal paraphrase. The dataset goes beyond standard by incorporating descriptive metadata regarding metaphor conventionality on a 3-graded scale (lexicalised, implicit, and ad-hoc), alongside a range of dictionary-derived source domains (military, gastronomy, health, meteorology, etc.). These metadata enable deeper analysis and potentially innovative insights of model performance regarding creativity, language change, and culture-sensitivity.
A New Semantic Artifact Based Framework for Studying and Documenting Algospeak and Related Phenomena
Fahad Khan | Elisa Gugliotta | Elisa Squadrito | Maura Tarquini | Francesca Frontini
Fahad Khan | Elisa Gugliotta | Elisa Squadrito | Maura Tarquini | Francesca Frontini
In this paper we present a new framework for analysis, documenting and publishing resources about the recent linguistic phenomenon of algospeak. This proposed framework features the use of two semantic artifacts (both of which we make available as SKOS semantic artifacts in RDF), and a cross-lingual lexicon of algospeak terms which follows a schema intended to facilitate the comparison of algospeak across languages and cultural contexts. Our article also features a discussion of the use of algospeak in two non-anglophone contexts (Italian and Arabic) which resulted from a period of data collection which the authors undertook as preparation for the creation of our framework and the categories which underlie it.
Creating a High Quality Abstract Meaning Representation Dataset Automatically
Johannes Heinecke | Munshi Asadullah | Frédéric Herledan | Geraldine Damnati
Johannes Heinecke | Munshi Asadullah | Frédéric Herledan | Geraldine Damnati
As only a few gold training datasets are available today, Abstract Meaning Representation (AMR) parsers are mainly trained on AMR 3.0, the largest dataset (Knight et al., 2020) which contains 55k sentences for training. Even if great progress has been made, leading to parsers that can reach Smatch scores higher than 83% on the AMR 3.0 test dataset, this is not accurate enough to be used in real world application pipelines. More data could help improve performance, but manually annotating sentences is costly. So, we have investigated an approach to automatically create synthetic data using different existing tools and models trained on AMR 3.0. This leads to better parsing performance with Smatch scores increased by 1 to 2 points (depending on the 3 gold test datasets used) with models trained on the augmented data.
Towards a Comprehensive English Wordnet-Wikidata Mapping
John P. McCrae | Johann Bergh | Krasimir Angelov
John P. McCrae | Johann Bergh | Krasimir Angelov
In this study, we present a comprehensive investigation into the mapping of English Wordnet to Wikidata, focusing on the existing mappings created by different projects. We systematically analyze the current mapping methodologies and their effectiveness, highlighting the strengths and limitations of each approach. Through a comparative analysis, we identified overlaps and discrepancies among the mappings, revealing insights into the relationships between the data sets. Our findings underscore the need for a more unified dataset that consolidates disparate mappings into a comprehensive unified Wordnet-Wikidata mapping. We propose a novel construction methodology for this unified data set, taking advantage of existing mappings while addressing their shortcomings. In addition, we discuss future perspectives and advanced techniques for mapping the remaining unmapped records, such as machine learning algorithms. This work not only contributes to the enhancement of data interoperability between Wordnet and Wikidata but also sets the stage for future research aimed at refining mapping techniques and expanding coverage.
Two fundamental tasks in computational linguistics are Lexical Semantic Change Detection and Word Sense Disambiguation. Both commonly rely on large annotated datasets. Most available datasets cover only one of two areas: diachronic corpora used for Semantic Change Detection, or synchronic datasets for Word Sense Disambiguation. To address this gap, the AmDi dataset is introduced as a German-language resource that supports a more fine-grained diachronic analysis of word meanings, while also enabling the investigation of embeddings generated with corresponding models, as well as providing a foundation for Word Sense Disambiguation tasks.
GerVLPro: A CEFR-Graded Vocabulary List of L2 Learners’ Productive Vocabulary in German
Noah-Manuel Michael | Anna Huelsing | Andrea Horbach
Noah-Manuel Michael | Anna Huelsing | Andrea Horbach
CEFR-graded vocabulary lists are a valuable tool for second-language (L2) learners as they provide guidance on the order in which to acquire vocabulary items. Thus, they are essential for informing computer-assisted language learning solutions that target vocabulary development in learners. However, the vast majority of GVLs are prescriptive in that they determine which items learners should learn at each level, and they provide little information about which items learners actually know. Moreover, in the case of German, almost all established GVLs focus exclusively on learners’ receptive vocabulary. To remedy this, we introduce GerVLPro: A CEFR-Graded Vocabulary List of L2 learners’ Productive vocabulary in German. We derived GerVLPro from a comprehensive aggregation of available CEFR-annotated German L2 learner corpora to represent a wide range of learners and contexts. The resulting list comprises 4,015 lemma-POS entries (A1: 611; A2: 1,134; B1: 903; B2: 1,103; C1: 249; C2: 15), assigned via a normalized share-based method. We then conducted a large-scale cross-evaluation against seven established GVLs and six prominent frequency lists. Despite sizable lexical overlap among resources, we found only weak to moderate alignment with GerVLPro. Finally, we investigated whether Gpt-4o and Gpt-5 can reliably grade the productive vocabulary items in GerVLPro. Although both models exhibit roughly similar predictive capacity, they underperform most of the established GVLs on alignment and do not accurately capture productive difficulty. Overall, our findings suggest that established GVLs, frequency lists, and LLM grading insufficiently reflect the trajectory of learners’ productive vocabulary, underscoring the need for descriptive, learner-based resources such as GerVLPro.
Building Bridges between Student and Curricular Language: Creating a Corpus of Abstract Meaning Representations for the Classroom
Kristin Wright-Bettner | Zheng Cai | Zekun Zhao | James H. Martin | Jeffrey Flanigan | Martha Palmer
Kristin Wright-Bettner | Zheng Cai | Zekun Zhao | James H. Martin | Jeffrey Flanigan | Martha Palmer
The potential of AI conversational agents to foster student learning and reduce teacher strain in classroom settings has made the development of pedagogical agents a prime research target. An effective AI agent in particular must be able to understand both student language and the content they are learning and, furthermore, map between them. Curricular terminology and student speech, though topically and semantically related, differ significantly in surface-form expression. We present the JIA-AMRs Collection, a new resource for exploring whether Abstract Meaning Representations (AMRs) can optimize interventions by a conversational AI agent in a middle-school classroom by providing structured semantic representations of classroom language. This resource also provides an avenue by which we can verify interventions by the agent. We discuss the challenges of creating a corpus of meaning representations that map across highly-dissimilar classroom data (multimedia curriculum, student spoken language, and student written language) and our promising results of a nearly 30-point gain in trained-parser performance over the off-the-shelf model.
Mu’jam Arriyadh: A Comprehensive Lexicon for Contemporary Arabic Language
Afrah A. Altamimi | Abdulrahman Alosaimy | Halah Munif Alharbi | Hawra Aljasim | Muneera Alhoshan | Amal Almazrua | Hanan Alharbi | Abdulrahman Saeed Alshehri | Bayan M. Almuqhim | Maryam H. Algarny | Yahya A. Asiri | Abdullah I. Alharbi | Saleh Zaidan Albalawi | Fawziah Mohammed Asiri | Sara Ali Alhifthi | Abdullah Alfaifi
Afrah A. Altamimi | Abdulrahman Alosaimy | Halah Munif Alharbi | Hawra Aljasim | Muneera Alhoshan | Amal Almazrua | Hanan Alharbi | Abdulrahman Saeed Alshehri | Bayan M. Almuqhim | Maryam H. Algarny | Yahya A. Asiri | Abdullah I. Alharbi | Saleh Zaidan Albalawi | Fawziah Mohammed Asiri | Sara Ali Alhifthi | Abdullah Alfaifi
This paper provides an overview of Contemporary Arabic Lexicon (Mu’jam Arriyadh). It is a contemporary and inclusive Arabic dictionary that has been specifically developed to cater to the needs of both native and non-native Arabic speakers. The corpus utilized in this study is derived from the Arabic Contemporary Corpus for Analysis (ACCA), which encompasses a vast collection of 450 million words of Modern Standard Arabic spanning the previous century. Significantly, the lexicon in question prioritizes lemma-based entries over root forms, hence enhancing its user-friendliness and adaptability across different contexts. The resource offers comprehensive linguistic data pertaining to a wide array of Arabic vocabulary, encompassing morphological, morph-syntactic, and semantic aspects. The Lexicon has been developed in accordance with the ISO 24613 standard, which improves its ability to be processed by machines and facilitates the utilization of natural language processing systems. The database encompasses a range of linguistic aspects, such as synonyms, antonyms, and root forms, offering a comprehensive compilation. Mu’jam Arriyadh is a contemporary Arabic lexicon that is designed to be accessible to users, compatible with machine processing, and highly beneficial for anyone studying the language, conducting research, and utilizing natural language processing technologies.
The Romanian Corpus Annotated with Multiword Expressions. PARSEME-Ro Version 2.0
Verginica Barbu Mititelu | Mihaela Cristescu | Elena Irimia | Carmen Mîrzea Vasile
Verginica Barbu Mititelu | Mihaela Cristescu | Elena Irimia | Carmen Mîrzea Vasile
The Romanian journalistic corpus previously annotated with verbal multiword expressions (PARSEME-Ro) has been extended recently with other journalistic texts and annotated with multiword expressions of all parts of speech closely observing version 2.0 of the PARSEME guidelines. The corpus size has been increased by about 40%, it underwent automatic morpho-syntactic annotation following the Universal Dependencies principles, as well as extensive semi-automatic annotation of multiword expressions of all morphological types (nominal, adjectival, adverbial, determiner, pronominal, prepositional, conjunction, interjection, and verbal for the newly added texts). We present here our work methodology, which involves an automatic annotation phase, but the manual work prevails in checking the annotation and its consistency. We also offer quantitative data about the new version of the corpus, the types of multiword expressions existing in Romanian and occurring therein, and characteristics thereof. The new version of the PARSEME-Ro corpus contributes to the field of developing multiword expressions resources per se, i.e. describing this language phenomenon, as well as resources for training, tuning and testing the performance of tools and large language models when dealing with this linguistic phenomenon.The paper also discusses some remarks on the MWE paraphrasing subtask in which a part of the corpus was used. The corpus is released with a permissive license.
Missing Links: LLM-Augmentation of Event Triggers of State Changes in the OpenPI Dataset
Kyeongmin Rim | James Pustejovsky
Kyeongmin Rim | James Pustejovsky
Effective computational understanding of procedural text requires modeling not just the state changes that occur (entity transformations), but also the specific actions that cause them (event triggers). A lack of datasets that explicitly link these two primary information sources has hindered progress in theory-oriented research and applications of NLP. This paper presents two primary contributions: (i) a new silver-standard dataset where event trigger annotations are added to existing state-change data on task-oriented procedural text, enabling both theoretical investigation and practical benchmarking; and (ii) inverse annotation, a framework for recovering missing linguistic annotations from existing semantic annotations—which we apply to recover event triggers from OpenPI’s state-change outcomes. We provide detailed pipeline analysis including error modes and quality filtering, and validate the dataset through comprehensive baseline evaluation of diverse trigger detection systems. Our work delivers both a reusable methodological framework applicable to other annotation recovery tasks and a new benchmark resource for modeling the relationship between linguistic actions and their semantic outcomes in procedural domains.
This article proposes the Conventional and Novel Metaphor Identification Procedure (CNMIP) for Mandarin Chinese and applies this replicable protocol to annotate the VUPMC dataset, a new Political Metaphor Corpus developed at VU University Amsterdam. The VUPMC corpus contains three Chinese political genres (Policy Documents, Remarks, News Reports) and includes over 220,000 tokens of concordance sentences for the node word 贸易 ‘trade’. The corpus analysis shows that 6.64% of lexical units in the VUPMC dataset are used as metaphor-related words (MRWs) to frame trade (e.g., using ‘war’ to frame trade as a war). Further tests show that distributions of MRWs differ significantly across genres and Parts of Speech. Similarities in MRW distributions between the VUPMC and other datasets confirm the reliability of the CNMIP procedure. The differences, however, highlight the methodological advances in manual annotation of conventional and novel MRWs as well as the distinctive features of Chinese political genres. The VUPMC dataset serves as a valuable language resource for computational detection of Chinese conventional and novel metaphors.
Not All Disneys Are the Same: Making Coreference Metonymy-Aware
Bingyang Ye | Jingxuan Tu | James Pustejovsky
Bingyang Ye | Jingxuan Tu | James Pustejovsky
Metonymy, a type of referential transfer in which a name evokes a conceptually related entity (e.g., “Disney” for the theme park), is a pervasive and systematic feature of natural language. Yet, despite its impact on entity interpretation, coreference research has rarely treated metonymy explicitly. Computational models of metonymy, in turn, typically analyze local, sentence-level cases, leaving unexplored how metonymic reference interacts with discourse-level coreference phenomena. We bridge this gap by introducing CoNLL-Coref-Met, a metonymy-aware annotation layer on top of CoNLL-2012 that flags metonymic mentions in context. Using this lens, we show that state-of-the-art neural resolvers and LLMs systematically underperform on metonymic clusters relative to literal counterparts. We then (i) correct clusters affected by metonymy to reflect semantic reference rather than surface form and (ii) introduce a metonymy-aware LLM procedure to resolve semantic ambiguities introduced by metonymic shifts. Our pipeline introduces a novel way to see, measure, and mitigate metonymy effects on coreference.
JSTS-Neg: Japanese Semantic Textual Similarity Dataset for Evaluating Negation Understanding Ability
Reiko Yuasa | Yoshihide Kato | Shigeki Matsubara
Reiko Yuasa | Yoshihide Kato | Shigeki Matsubara
Negation is a common linguistic phenomenon in natural language. Thus, datasets and benchmarks focused on negation are being constructed to evaluate the negation understanding abilities of language models. Negation is especially crucial when estimating the semantic similarity between sentences because it inverses their meaning. Although semantic textual similarity (STS) is one of the useful tasks to evaluate the abilities of large language models (LLMs), few STS datasets focus on negation. In this research, we introduce JSTS-Neg, a new Japanese STS dataset focusing on negation. Most instances in JSTS-Neg include negations and they are composed of both clausal and sub-clausal negations to reflect a variety of negation types. Moreover, JSTS-Neg consists of negation minimal pairs that only differ in the presence or absence of a negation cue. We evaluate the performance of existing LLMs on JSTS-Neg using negation minimal pairs to explore their abilities and limitations in understanding negation. LLMs tend to predict the similarity of two sentences ignoring negation cues in specific settings.
Few-shot Prompting or Supervised Tuning? A Comparative Study of LLMs for Linguistically Distant Language Pairs in BDI
Deepen Naorem | Sanasam Ranbir Singh | Telem Joyson Singh | Priyankoo Sarmah
Deepen Naorem | Sanasam Ranbir Singh | Telem Joyson Singh | Priyankoo Sarmah
Bilingual Dictionary Induction (BDI) presents significant challenges in distant language pairs, particularly in light of the non-isomorphic nature and complexity of linguistic structures. This paper systematically evaluates the performance of unsupervised, supervised fine-tuning, and few-shot prompting approaches on BDI using Large Language Models (LLMs) on a diverse set of distant language pairs. The unsupervised approach explores the inherent multilingual capabilities of LLMs without fine-tuning, while the supervised fine-tuning method utilizes extensive labeled datasets to train models explicitly for BDI tasks. On the other hand, few-shot prompting leverages minimal examples to elicit accurate responses from the LLMs in a zero-shot or few-shot learning paradigm. Our experimental results reveal that the 5-shot prompting approach outperforms unsupervised and zero-shot settings in all cases and surpasses supervised settings in 82.86% of the cases. Few-shot prompting demonstrates robustness against overfitting, leveraging LLMs’ in-context learning and multilingual capabilities, making it particularly effective in target-to-source translation, even for morphologically complex language pairs. At the same time, few-shot prompting in LLM models, such as Llama, remains ineffective for morphologically rich language pairs like En-Mn and En-Ta in source-to-target BDI tasks. These findings suggest that few-shot prompting is a cost-effective and powerful alternative for BDI tasks, with future work enhancing BDI tasks in morphologically rich pairs.
When Structure Matters: Cross-Lingual Hyperbolic Embeddings for Chinese and English Wordnets
Mao-Chang Ku | Da-Chen Lian | Pin-Er Chen | Po-Ya Angela Wang | Wei-Ling Chen | Shu-Kai HSIEH
Mao-Chang Ku | Da-Chen Lian | Pin-Er Chen | Po-Ya Angela Wang | Wei-Ling Chen | Shu-Kai HSIEH
Hyperbolic embeddings such as the Poincaré model effectively represent lexical hierarchies with low distortion, yet their cross-lingual generalizability remains largely unexplored. This study investigates cross-lingual transfer by training 20-dimensional Poincaré embeddings exclusively on Open English WordNet (OEWN) hypernymy relations and evaluating on aligned Chinese Wordnet (CWN) synsets under a vocabulary-constrained transfer setting, where CWN-relevant synsets appear in OEWN training data but no Chinese-language supervision is used. We report robust statistical evidence based on the final 10 training checkpoints: Poincaré embeddings achieve 2.57× higher Mean Reciprocal Rank (MRR) than Euclidean embeddings on CWN (0.030 ± 0.001 vs 0.012 ± 0.000, p < 0.001, Cohen’s d = 34.48) and 5.61× higher on OEWN (0.016 ± 0.000 vs 0.003 ± 0.000, p < 0.001, d = 42.48). Furthermore, hierarchical filtering leveraging the radial dimension of hyperbolic space provides substantial additional gains: +74.6% MRR improvement on CWN and +25.8% on OEWN (both p < 0.001). The model achieves higher absolute performance on the zero-shot CWN test set (MRR = 0.052 ± 0.002) than on the in-domain OEWN test set (MRR = 0.020 ± 0.001). We attribute this to structural alignment: CWN’s broader branching factor (4.32 vs 1.10) and moderate depth naturally suit hyperbolic geometry’s capacity to compactly represent hierarchies. Our findings demonstrate that geometric properties learned from English hypernymy transfer robustly across languages when semantic structures align. We release the aligned CWN–OEWN hypernymy evaluation dataset and complete evaluation framework to facilitate future research on geometry-based cross-lingual semantic modeling.
up
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Reinhard Rapp | Ayla Rigouts Terryn | Serge Sharoff | Pierre Zweigenbaum
Reinhard Rapp | Ayla Rigouts Terryn | Serge Sharoff | Pierre Zweigenbaum
Keynote: The Cross-Lingual Transfer Myth: Why Modern LLMs Still Fail Without Comparable Corpora and Representations
Els Lefever
Els Lefever
Comparable corpora have long served as a foundation for multilingual NLP, supporting transfer across languages in tasks such as classification, retrieval, translation, and argument mining. Yet in the era of multilingual transformers and generative models, a central question is no longer simply whether texts are comparable, but what kinds of internal representations and downstream behaviors that comparability actually enables. In this keynote, I argue that cross-lingual transfer is best understood as a continuum oscillating between shared semantic structures and language-specific realizations. Drawing on two complementary studies, I demonstrate how this tension manifests both in the data models learn from and in the representations they develop. The first case study investigates multilingual stance and argument mining using the new Russian LoveHate corpus alongside English debate data. The results indicate that translated or multilingual resources are useful but insufficient proxies for language-specific corpora: local topics, culturally situated argumentation patterns, and stance expression still shape model performance and generalization. The second case study presents a neuron-level analysis of multilingual emotion detection, showing that multilingual encoders such as XLM-R develop both polyglot neurons, which respond consistently across languages, and monolingual neurons, which remain tied to particular linguistic systems. This reveals that even successful cross-lingual emotion transfer depends on only partial internal alignment. Together, these findings suggest that multilingual NLP needs corpora that preserve culturally specific meaning while supporting robust transfer, as well as interpretability frameworks that can diagnose where multilingual systems genuinely share representations and where they merely approximate them. Comparable corpora are not just training material; they are essential to understand how cross-lingual generalization succeeds, where it breaks down, and how truly multilingual NLP can move beyond English-centric assumptions and conclusions.
A Comparative Study of Parkinsonian Speech Corpora for Deep Learning-Based Detection of Dysarthria
Clara Ponchard | Pierre Serrano
Clara Ponchard | Pierre Serrano
Idiopathic Parkinson’s disease is associated with motor speech impairments collectively referred to as hypokinetic dysarthria, which can appear at early disease stages and remain challenging to assess objectively in clinical practice. Most automatic assessment studies rely on individual speech corpora analyzed in isolation, leaving open questions regarding their comparability and their suitability for joint use within unified classification frameworks. This study explicitly investigates the cross-corpus comparability of existing Parkinsonian speech datasets designed for hypokinetic dysarthria assessment. Rather than assuming their compatibility, we evaluate it empirically through the generalization performance of classification systems trained on single or multiple corpora. We examine which datasets can be effectively combined and whether multi-corpus training improves robustness across heterogeneous recording conditions and speech tasks. Four corpora are evaluated under intra-corpus, cross-corpus, and out-of-domain settings. Results demonstrate that multi-corpus training enhances robustness and generalization performance, while also revealing substantial differences in cross-dataset compatibility. These findings provide a clearer understanding of the degree of comparability between existing resources and offer practical guidelines for the design of future corpora and more generalizable tools for the automatic clinical assessment of Parkinsonian speech.
Computing Semantic Similarity for Aligning Bilingual Semi-parallel Texts: A Case Study
Steffen Frenzel | Maximilian Krupop | Manfred Stede
Steffen Frenzel | Maximilian Krupop | Manfred Stede
Semi-parallel text refers to versions of the same text that have to some extent been edited by authors, translators, or others. They are of relevance especially in the social sciences and in literary genres. In this paper, we consider the bilingual (English/German) variant of the problem. The philosopher Hannah Arendt, for example, wrote political essays that often exist in multiple versions and in both languages. She repeatedly modified her texts, added or deleted parts, and framed topics differently for target audiences. For researchers to explore the history of such material in detail, and at the same time at scale, automatic alignment (i.e., finding the best match of semantically similar sentences) is a very valuable preprocessing step. In this paper, we compare the performances of a range of methods for this task, based on computing semantic similarity. We present the results and conduct a qualitative error analysis to identify recurring sources of error.
A Comparative Study in Corpus Linguistics Applied to Automatic Terminology Extraction
Mercè Vàzquez | Sergi Alvarez-Vidal | Antoni Oliver
Mercè Vàzquez | Sergi Alvarez-Vidal | Antoni Oliver
Parallel and comparable corpora are the main linguistic resources to identify multilingual terminology using automatic term extraction tools. However, parallel corpora are available only for certain languages, domains and genres, and comparable corpora have some limitations when identifying corresponding terms. To implement a more efficient selection of multilingual terminology, we compared the performance of using specialised parallel and comparable corpora applied to languages with various forms of capital in linguistic resources. This paper presents a comparative study in corpus linguistics in which we automatically identify terms in Catalan, Spanish and English in legislation and administrative law using parallel corpora, comparable corpora and a combined methodology based on both typologies of corpora together with word embeddings. We observe that the combined methodology implemented obtains a higher number of term candidates than when working exclusively with parallel or comparable corpora. The evaluation of the results is performed using a terminological thesaurus as a gold standard. The new methodology presented in our study permits us to identify multilingual terminology in an efficient way, especially in Catalan-Spanish languages.
Comparable Corpora in Cross-linguistic Research: Nominal Number in English, Czech, and Greek
Konstantinos Diamantopoulos | Magda Ševčíková
Konstantinos Diamantopoulos | Magda Ševčíková
The paper examines the use of comparable corpora for contrastive research on the category of nominal number across three languages—English, Czech, and Greek. Two objectives are pursued: a cross-linguistic analysis of number and an assessment of the impact of automatic annotation on linguistic findings. For this study, corpora of comparable size and composition were compiled for the three languages from the Leipzig Corpora Collection. The data were automatically annotated using two open-access tools, Stanza and UDPipe, producing six datasets (two per language), each containing about 5 million sentences and 100 million tokens. Although derived from the same source, the paired datasets for each language differ in sentence and word segmentation, in the number of nouns identified, and in the number values assigned. These differences, nevertheless, do not appear to substantially affect the overall picture of number in the languages examined. The distribution of lemmas by the ratio of singular and plural forms challenges the view commonly presented in grammars that most nouns occur in both numbers and that singular-only and plural-only nouns are rare. However, a closer analysis of nouns assumed to have defective number indicates that answers to more nuanced questions vary depending on the annotation tool used.
We present a multi-register (web, news, and government texts), diachronic (2015-2024), comparable corpus annotated for lexical gender-inclusive language (gil) features in German and Spanish. Apart from rule-based annotations, we train a transformer-based classifier to resolve semantically ambiguous neutral expressions like epicenes to reliably annotate true human referents. In a sample study, we analyze register variation in the three registers in terms of gil features both contrastively and diachronically. We show that gil usage increases and varies diachronically in terms of register in both languages. German texts show a higher overall frequency and diversity of gil features than Spanish texts. However, across languages, registers behave similarly, with government text showing the strongest usage of gil followed by news and web texts, and web texts showing the strongest innovation in terms of features. The results of our study are valuable to linguistic areas such as human and machine translation, SLA, and contribute to register-conform gender inclusive NLP downstream tasks such as machine translation, summarization or textgeneration. From a diachronic point of view, our corpus and analyses are a valuable contribution to observing language change in the making.
A Diachronic Comparable Corpus of Spanish Digital News (2017–2026) for the Study of Stylistic Convergence in the GenAI Era
Hugo Sanjurjo-González
Hugo Sanjurjo-González
This study introduces a comparable corpus of Spanish digital news (2017–2026) designed to analyze potential linguistic shifts coinciding with the widespread adoption of Generative AI. We propose an analytical framework structured across three levels: lexical statistics, semantic topology, and neural classification. By implementing a protocol of NER-masking, we isolate structural discourse markers from topical content to identify the stylistic patterns of the contemporary period. Our results suggest a measurable structural shift within the analyzed corpus, indicating a trend toward a more standardized professional register. While macro-statistical metrics like Shannon entropy remain stable —indicating statistical consistency— Zipf-Mandelbrot distributions and SVD mapping reveal a concentration of unique vocabulary into more predictable clusters. In this scenario, the 2023–2026 subcorpus exhibits a discernible topological displacement compared to the 2017–2021 baseline. The study identifies a ‘Gray Zone’ where highly structured technical reporting and hybridized production become indistinguishable, suggesting a structural stylistic convergence within this digital environment. These findings provide a methodological baseline for analyzing discursive stabilization in professional domains without assuming definitive authorship.
Align and Shine: Building High-quality Sentence-aligned Corpora for Multilingual Text Simplification
Luis Kenji Hilasaca Sanchez | Nouran Khallaf | Serge Sharoff
Luis Kenji Hilasaca Sanchez | Nouran Khallaf | Serge Sharoff
Text simplification plays a crucial role in improving the accessibility and comprehensibility of written information for diverse audiences, including language learners and readers with limited literacy. Despite its importance, large-scale, high-quality datasets for training and evaluating text simplification models remain scarce for languages other than English. This paper reports an experimental study on the collection and processing of crowd-sourced simplification data to construct a corpus suitable for both training and testing text simplification systems across multiple languages (Catalan, English, French, Italian and Spanish). We report mechanisms for sentence-level alignment from document-level data. The resulting dataset of the aligned sentence pairs is publicly available.
Bi-Text Mining across German Dialects: On the Role of Synthetic Training Data for Dialect Adaptation
Jing Wang | Barbara Plank | Robert Litschko
Jing Wang | Barbara Plank | Robert Litschko
Cross-dialect bi-text mining relies on robust multilingual sentence representations to identify semantically equivalent sentence pairs across languages. While recent multilingual bi-encoder models achieve strong performance on standardized written languages, their behavior on dialectal varieties is largely unknown. In this study, we use Tatoeba to evaluate the performance of four widely-used bi-encoders on dialect-to-standard German translation retrieval, covering German documents and queries written in three dialects: Low German, Bavarian, and Alemannic. Motivated by the lack of resources, we examine the extent to which synthetic translations (from dictionaries and large language models; LLMs) can serve as weak supervision for dialect adaptation. Our results reveal that bi-encoders, when applied in a zero-shot setting, exhibit deficiencies in capturing semantic similarity between German and dialects, while fine-tuning on synthetic data substantially improves their retrieval effectiveness, with larger gains obtained from LLM-translated training data. We further analyze retrieval performance on Bavarian across varying dialect word proportions and observe a drop when dialect words make up more than 60% of the text.
Parallel Corpora of Scholarly Documents for English-French Machine Translation
Ziqian Peng | Lichao Zhu | Rachel Bawden | Maud Bénard | Éric de la Clergerie | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | François Yvon
Ziqian Peng | Lichao Zhu | Rachel Bawden | Maud Bénard | Éric de la Clergerie | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | François Yvon
The growing ability of large language models (LLMs) to process long-range context opens new perspectives for document-level machine translation (MT), especially in scholarly communication. In fact, translating scholarly texts requires to integrate both local and long-range contextual information to ensure the consistency and coherence across the full document. However, document-level parallel corpora for such text types remain scarce, limiting both evaluation and domain adaptation of MT systems for this task. To address this gap, we introduce ParaEPS (Earth and Planetary Sciences Bilingual Corpus) and ParaNLP (Natural Language Processing Bilingual Corpus), two new parallel corpora covering 14k abstracts and 103 full-length articles in two scientific domains to be used for fine-tuning and evaluation purposes. We compare the performance of eight MT systems on these test sets and find that fine-tuning on document-level data closes the gap between open systems based on Large Language Models (LLMs) and commercial systems. We also find that the performance of recent LLMs can worsen when translating full articles instead of translating them on a per paragraph basisfine-tuning. These experiments underscore the need for corpora such as ParaEPS and ParaNLP.
Validating a Pipeline to Create a Comparable Corpus of Government-Issued Travel Advisories from the Internet Archives
Laura Braun | Christian Oswald
Laura Braun | Christian Oswald
Government-issued travel advisories are used by citizens to get information about destination countries for tourism and other purposes such as temporary work stays or permanent relocation plans. However, qualitative evidence suggests that travel advisories may be influenced by considerations beyond current security situations. Systematic and rigorous quantitative analyses of advisories are scarce because relevant corpus data are not readily available and official government websites often provide practical obstacles. We validate a pipeline to generate a time-series cross-sectional dataset of government-issued travel advisories for three English-speaking issuing countries based on the Internet Archive’s Wayback Machine. Using official government data sources that are prohibited to be scraped and used for research, we illustrate that our approach provides (near-)complete coverage. The resulting corpus and code are intended to support downstream research on comparative risk communication, international relations, and text analysis using natural language processing methods.
Leveraging Comparable Toxicity Lexicons in Prompt Instructions for Multilingual Text Detoxification
Yassir El Attar | Esra Dönmez | Nina K. Ohlendorf | Agnieszka Falenska
Yassir El Attar | Esra Dönmez | Nina K. Ohlendorf | Agnieszka Falenska
To mitigate the prevalence of toxic language on digital social media, various NLP approaches have been proposed for automatic text detoxification. However, the potential of toxic expression lexicons as a comparable cross-lingual resource to guide this process remains largely unexplored. In this work, we investigate how such resources can be effectively used to inform multilingual language models about what should and should not be considered toxic. We evaluate four models under two settings—zero-shot prompting and fine-tuning—to assess the impact of incorporating toxic expressions in prompt instruction, including in cross-lingual transfer scenarios. Our results show that both zero-shot prompting and fine-tuning approaches benefit considerably from adding toxic expressions in prompt instructions during training and/or inference. Our findings demonstrate that comparable, lightweight, language-specific toxic expression lexicons constitute an effective mechanism for injecting explicit information about lexical toxicity into multilingual language models.
up
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Ingo Siegert | Maria Irena Szawerna | Khalid Choukri | Simon Dobnik | Paweł Kamocki | Therese Lindström Tiedemann | Pierre Lison | Ricardo Muñoz Sánchez | Ildikó Pilán | Lisa Södergård | Kossay Talmoudi | Elena Volodina | Xuan-Son Vu
Ingo Siegert | Maria Irena Szawerna | Khalid Choukri | Simon Dobnik | Paweł Kamocki | Therese Lindström Tiedemann | Pierre Lison | Ricardo Muñoz Sánchez | Ildikó Pilán | Lisa Södergård | Kossay Talmoudi | Elena Volodina | Xuan-Son Vu
Transparency as Architecture: Structural Compliance Gaps in EU AI Act Article 50 II
Vera Schmitt | Niklas Kruse | Premtim Sahitaj | Julius Schöning
Vera Schmitt | Niklas Kruse | Premtim Sahitaj | Julius Schöning
Art. 50 II of the EU Artificial Intelligence Act mandates dual transparency for AI-generated content: outputs must be labeled in both human-understandable and machine-readable form for automated verification. This requirement, entering into force in August 2026, collides with fundamental constraints of current generative AI systems. Using synthetic data generation and automated fact-checking as diagnostic use cases, we show that compliance cannot be reduced to post-hoc labeling. In fact-checking pipelines, provenance tracking is not feasible under iterative editorial workflows and non-deterministic LLM outputs; moreover, the assistive-function exemption does not apply, as such systems actively assign truth values rather than supporting editorial presentation. In synthetic data generation, persistent dual-mode marking is paradoxical: watermarks surviving human inspection risk being learned as spurious features during training, while marks suited for machine verification are fragile under standard data processing. Across both domains, three structural gaps obstruct compliance: (a) absent cross-platform marking formats for interleaved human-AI outputs; (b) misalignment between the regulation’s ’reliability’ criterion and probabilistic model behavior; and (c) missing guidance for adapting disclosures to heterogeneous user expertise. Closing these gaps requires transparency to be treated as an architectural design requirement, demanding interdisciplinary research across legal semantics, AI engineering, and human-centered design.
Towards Robust Evaluation for Privacy QA Systems
Anna Leschanowsky | Zahra Kolagar | Erion Çano | Ivan Habernal | Dara Hallinan | Emanuël Habets | Birgit Popp
Anna Leschanowsky | Zahra Kolagar | Erion Çano | Ivan Habernal | Dara Hallinan | Emanuël Habets | Birgit Popp
The transparency principle of the General Data Protection Regulation requires data-processing information to be clear, precise, and accessible. While Large Language Models (LLMs) show promise in this context, their probabilistic nature raises challenges for ensuring truthfulness and comprehensibility. This paper presents an exploratory evaluation of eight Privacy Question Answering (QA) systems – including LLMs, retrieval-augmented generation, and alignment-based approaches – on two datasets. We propose an evaluation framework that maps both traditional NLP and LLM-as-a-judge metrics to the legal requirements of comprehensibility and precision. Results show that no single system consistently excels across all metrics, and that system rankings can vary depending on the choice of metric and thresholding. We highlight open questions and emphasize the need to translate legal requirements into technical evaluation criteria. Our work provides a foundation for a more robust evaluation of Privacy QA systems.
LDS Contractual Framework: Principles, Status and Implementation
Penny Labropoulou | Kossay Talmoudi | Dimitrios Gkoumas | Katerina Gkirtzou | Miltos Deligiannis | Leon Voukoutis | Athanasia Kolovou | Khalid Choukri | Stelios Piperidis | Dimitrios Galanis
Penny Labropoulou | Kossay Talmoudi | Dimitrios Gkoumas | Katerina Gkirtzou | Miltos Deligiannis | Leon Voukoutis | Athanasia Kolovou | Khalid Choukri | Stelios Piperidis | Dimitrios Galanis
To strengthen competitiveness and digital sovereignty, the European Union has promoted the development of Common European Data Spaces to enable secure and interoperable data sharing between participants for various sectors. Data spaces combine technical infrastructure with governance mechanisms to ensure trust, transparency, data sovereignty and interoperability. Their operation must comply with the evolving European regulatory framework as well as contractual law. This paper presents the strategy adopted in the Language Data Space (LDS) to operationalise these requirements, focusing on its contractual framework and supporting instruments. It outlines the governing principles designed to ensure lawful, transparent, and fair data transactions while safeguarding the rights and obligations of data providers and consumers alike. It further describes the actual framework, and the recommended data sharing licences, with a particular emphasis on the LDS standard licence. Finally, it presents the automation tools designed and developed to support the relevant workflows while serving a wide range of users that have little or no knowledge of technical and legal complexities.
Authorship Attribution in the Times of LLMs within the Framework of the CRediT Taxonomy
Pawel Kamocki | Andreas Witt
Pawel Kamocki | Andreas Witt
This article examines the concept of authorship in the context of generative language models and other uses of Artificial Intelligence, and how this new ‘authorshipness’ can be represented in metadata. It analyses authorship under copyright law and proposes a metadata-based approach to disclosing the use of AI in publications, drawing on the widely adopted CRediT taxonomy developed by the National Information Standards Organization (NISO), and informed by guidance from the United States Copyright Office (USCO) and the International Association of Scientific, Technical and Medical Publishers (STM).
DeID-Clinic: A Risk-Aware Pseudonymization Framework for Clinical Text De-identification and Re-identification Risk Assessment
Angel Paul | Dhivin Shaji | Lifeng Han | Warren Del-Pinto | Goran Nenadic | Suzan Verberne
Angel Paul | Dhivin Shaji | Lifeng Han | Warren Del-Pinto | Goran Nenadic | Suzan Verberne
The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sharing while preserving downstream utility. This paper presents DeID-Clinic, a multi-layered framework for automated pseudonymization and re-identification risk assessment of clinical free-text data. Our approach integrates domain-adapted transformer models, including BioBERT and ClinicalBERT, into the MASK de-identification framework to improve the detection and masking of protected health information (PHI). Beyond entity recognition, we introduce a novel document-level risk assessment module that quantifies residual re-identification risk using a combination of k-anonymity, l-diversity, t-closeness, contextual similarity, and entity co-occurrence analysis. Experiments conducted on the i2b2 2014 de-identification dataset demonstrate strong performance, achieving macro-level F1 scores above 0.96 for several entity categories, while enabling quantitative prioritization of high-risk documents for further review. Our results highlight the effectiveness of combining neural de-identification with explicit risk modeling, supporting privacy-preserving data sharing in sensitive domains. Although evaluated on clinical text, the proposed framework is generalizable to other privacy-critical domains such as legal and administrative documents, where reliable pseudonymization and risk-aware anonymization are essential.
Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
Gabriel Loiseau | Damien Sileo | Damien Riquet | Maxime Meyer | Marc Tommasi
Gabriel Loiseau | Damien Sileo | Damien Riquet | Maxime Meyer | Marc Tommasi
Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving NLP. Recent work has shown that LLMs can serve as reliable privacy evaluators, achieving strong agreement with human judgments; however, their computational cost and impracticality for processing sensitive data at scale limit real-world deployment. We address this gap by distilling the privacy assessment capabilities of Mistral Large 3 (675B) into lightweight encoder models with as few as 150M parameters. Leveraging a large-scale dataset of privacy-annotated texts spanning 10 diverse domains, we train efficient classifiers that preserve strong agreement with human annotations while dramatically reducing computational requirements. We validate our approach on human-annotated test data and demonstrate its practical utility as an evaluation metric for de-identification systems.
Birds of a Feather: Do Embedding Representations of Personal Information Flock Together?
Maria Irena Szawerna | Simon Dobnik
Maria Irena Szawerna | Simon Dobnik
Personally identifiable information (PII or PI) can appear in a wide variety of linguistic data, posing both ethical and legal challenges for conducting research and developing applications involving such texts. In this paper, we investigate the alignment between automatic clustering of FastText and Transformer embedding representations of personal information spans sourced from essays written by adult learners of Swedish as a second language and the general and detailed personal information labels assigned to these spans by expert annotators. Our goals are to assess the extent of overlap between the semantic categories and evaluate the semantic coherence of the human-assigned classes, which may have implications for de-identification procedures. We observe that while contextual embeddings, especially ones from a specialized word-in-context model, produce relatively good clustering results, they only partly map to the human understanding of how to classify personal information.
Modelling Legal Compliance in a Consent Wizard Application as Part of a Research-Centered and User-Oriented Data Infrastructure
Aliena Strathmann | Marc-Levin Joppek | Maryam Mohammadi | Katja Politt | Paul T. Schrader | Annett B. Jorschick | Hendrik Buschmeier
Aliena Strathmann | Marc-Levin Joppek | Maryam Mohammadi | Katja Politt | Paul T. Schrader | Annett B. Jorschick | Hendrik Buschmeier
Recent research calls for data management infrastructures that explicitly operate within the bounds of ethical and legal constraints, and facilitate adherence to Open Science principles by integrating automated support for planning, collection, storage, use, reuse, and sharing of data within. Legal and ethical requirements of data processing have become increasingly complex, introducing administrative barriers to scientific research investigating data generated by human participants, which encompasses a vast majority of humanities research. In response to this, we present RUDI (“Research-centered User-oriented Data Infrastructure”), a modular framework grounded in an interdisciplinary approach informed by legal, computational and linguistic expertise. This paper introduces its first component; a configurable and dynamically adaptive consent form generator in the form of a “wizard” web application. We outline how legal aspects are modeled within, and highlight its concrete benefits for administrative aspects of research. Further, we discuss the contextualization of data within the research domain by leveraging the use of standardized ontology within the framework.
Balancing FAIR and GDPR: A Governance Framework for Oral Archives
Elvira Mercatanti | Monica Monachini | Giovanni Abete | Silvia Calamai | Sergio Canazza | Alessandro Casellato | Virginia Niri | Cesarina Vecchia | Giulia Zitelli Conti | Giada Zuccolo
Elvira Mercatanti | Monica Monachini | Giovanni Abete | Silvia Calamai | Sergio Canazza | Alessandro Casellato | Virginia Niri | Cesarina Vecchia | Giulia Zitelli Conti | Giada Zuccolo
This paper presents a governance framework developed within the research project ROADS to support thesustainable management of oral archives, which constitute essential linguistic resources for interdisciplinary research and cultural heritage preservation. Oral archives raise complex ethical and legal challenges due to the hybrid nature of voice data, which function simultaneously as historical documents, scientific sources and biometric identifiers, thereby creating tensions between open science principles and data protection regulations. The proposed framework integrates FAIR principles (Findable, Accessible, Interoperable, Reusable) with Privacy by Design and the GDPR accountability principle through a multilayered approach. It introduces an access model that distinguishes between publicly available metadata and controlled access to identifiable audio materials, following trusted repository standards. The framework also incorporates consent management procedures and safeguards for legacy collections, enabling responsible data sharing while preserving scientific usability. More broadly, ROADS provides a transferable model to guide the transition from project-based archives to FAIR, sustainable and reusable research resources, ensuring compliance with data protection requirements and respect for the sensitivity of the documented contexts.
Legal Considerations in the Use of Synthetic Data for AI Development and Finetuning: The Case of LLMs4EU
Kossay Talmoudi | Khalid Choukri | Amélie Gourgeot | Florine Astruc
Kossay Talmoudi | Khalid Choukri | Amélie Gourgeot | Florine Astruc
This paper examines the legal implications of using synthetic data to develop and fine-tune general-purpose AI models in the European Union, using the LLMs4EU project as a case study. It situates synthetic data within the Union’s broader data policy and highlights it as a candidate tool for reconciling data availability with regulatory constraints. From a data-protection perspective, it analyses whether and when synthetic data should be classified as “personal data” under the GDPR. From a copyright and contractual standpoint, the paper assesses the risks that synthetic datasets may embed infringing content or derive from unlawfully trained models, in light of the GEMA v. OpenAI ruling on memorised works and emerging analyses of liability for AI-generated outputs, and considers the constraints imposed by model licensing and acceptable-use policies on using models to generate training data for other models. The paper concludes that synthetic data can play a valuable role in mitigating legal risks and enabling compliant AI development in LLMs4EU, but only if its generation and use are embedded in robust governance frameworks that address data protection, copyright and contractual obligations across the entire data value chain.
Evaluating Encoder- and LLM-Based Approaches for Robust Indirect Personal Identifier Detection
Christoph Otto | Ibrahim Baroud | Akiko Aizawa | Sebastian Möller | Roland Roller | Lisa Raithel
Christoph Otto | Ibrahim Baroud | Akiko Aizawa | Sebastian Möller | Roland Roller | Lisa Raithel
Removing explicit protected health information does not fully eliminate re-identification risk in clinical text. Contextual attributes such as socio-economic status, institutional affiliations or detailed life circumstances may still enable linkage attacks. These heterogeneous and sparsely distributed elements, termed Indirect Personal Identifiers, extend de-identification beyond fixed identifier lists and pose new modeling challenges. Therefore, we present the first systematic comparison of encoder-only models, prompt-based LLMs and hybrid pipelines for span-level IPI detection in English discharge summaries. A fine-tuned RoBERTa-large model improves on an existing baseline and substantially outperforms ChatGPT-5.2, achieving 0.906 micro-F1 and 0.724 macro-F1, compared to 0.509 micro-F1 and 0.487 macro-F1. Our findings indicate that IPI detection constitutes a distinct modeling regime characterized by class imbalance and high intra-class variability, where scaling model capacity alone does not guarantee macro-level robustness. We show that supervised encoder models currently provide the most reliable foundation for extending anonymization guarantees and future research.
VEIL: A Benchmark for Value-Preserving Entity Identification Limitation
Darina Gold | Shadi Rastegar | Alina Liebel | Alessandra Zarcone
Darina Gold | Shadi Rastegar | Alina Liebel | Alessandra Zarcone
Large Language Models (LLMs) are linked to several issues regarding Personally Identifiable Information (PII). PII can occur in the training data and can thus be accidentally leaked or extracted with malicious intent, or it can be inputted in LLM-based technologies by users through their prompts. A viable strategy to limit the LLMs’ exposure to PII is to filter input and output data by de-identifying PII, including personal names. This however poses a challenge: a name could refer to a private person in a context containing sensitive information (e.g., Michelangelo is an atheist), or it could refer to a famous artist in another context (e.g., Michelangelo’s Sistine Chapel), and masking the latter may hinder the LLMs’ capabilities in general-knowledge tasks. We tackle the problem of personal name de-identification and focus on the decision of which personal names need to be removed (and which should be kept), based on context. We present VEIL, a challenging benchmark for Value-preserving Entity Identification Limitation, for context-aware de-identification decisions on LLM training data, and compare the performance of different state-of-the-art systems on the task.
up
Proceedings of Computational Affective Science (CAS) @ LREC 2026
Proceedings of Computational Affective Science (CAS) @ LREC 2026
Christopher Bagdon | Krishnapriya Vishnubhotla | Kristen A. Lindquist | Lyle Ungar | Roman Klinger | Saif M. Mohammad
Christopher Bagdon | Krishnapriya Vishnubhotla | Kristen A. Lindquist | Lyle Ungar | Roman Klinger | Saif M. Mohammad
Quality and Agreement in Multilabel Emotion Annotation: A Case Study and Evaluation Framework
Emily Sofi Ohman | Anna Koufakou
Emily Sofi Ohman | Anna Koufakou
Emotion annotation is inherently subjective, yet most NLP pipelines still assume “gold” labels, typically produced by majority voting, and treat annotator variation as noise. In this paper, we present a multilabel emotion annotation case study and use it to examine how annotator behavior and aggregation choices affect both agreement estimates and downstream emotion classifiers. Rather than collapsing disagreement into a single label, we represent targets as soft vote-share labels (including an intensity-weighted variant) and evaluate models using both thresholded metrics (macro-/micro-F1) and probabilistic alignment (Bernoulli cross-entropy SoftBCE), alongside data-derived disagreement diagnostics. Across annotation regimes, we show that disagreement is structured and leaves measurable traces in model behavior: hard labels may maximize F1 metrics, while soft supervision yields predictions that better reflect empirical annotator variance and uncertainty. Our results provide practical guidance for designing, aggregating, and evaluating multilabel emotion datasets when multiple interpretations are plausible.
Understanding Irony through Explanations and Background Knowledge
Aaron Maladry | Els Lefever | Cynthia Van Hee | Veronique Hoste
Aaron Maladry | Els Lefever | Cynthia Van Hee | Veronique Hoste
This article investigates the automatic explanation of irony in English tweets. The work covers the development and validation of a conceptual framework for annotating knowledge-informed explanations for figurative language as well as the training and evaluation of specialized generative models. Human judgements confirm that both fine-tuned open-source models (Llama 3) and proprietary models (GPT-4) can produce high-quality explanations, effectively incorporating relevant world knowledge. While metrics like BLUE and ROUGE do not seem to align with human judgement, we find that semantic similarity measures align well with human quality estimations. The resulting models and datasets for irony explanations, published as the iRONNIE collection, actively bridge the gap between theoretical understanding of irony and the technical innovations of the NLP domain. The models are be released to the public to facilitate a deeper linguistic analysis of world knowledge involved in understanding irony on social media in future work.
Speech-Based Emotion Recognition and Classification Integrating a CNN and BiLSTM Network
Fatima Uroosa | Asim Abbas | Muhammad Tayyab Zamir | Grigori Sidorov
Fatima Uroosa | Asim Abbas | Muhammad Tayyab Zamir | Grigori Sidorov
Speech emotion recognition (SER) has gained significant interest in recent times, which utilizes speech signals to identify the emotional state of speakers. Accurate recognition of subtle emotional variations in speech, such as distinguishing closely related emotional states, remains a challenging problem due to the variability of speech signals and the acoustic similarity among emotion classes across different speakers and linguistic contexts. This paper proposes a hybrid deep learning model that integrates a Convolutional Neural Network (CNN) with a Bidirectional Long Short-Term Memory (BiLSTM) network to effectively identify both spectral and temporal features of speech. The log Mel-frequency spectral coefficients (MFSC) are used as input features to represent discriminative spectral representations, while the BiLSTM layer model represents long-range temporal dependencies in speech signals. The proposed framework is evaluated on the Toronto Emotional Speech Set (TESS), a publicly available dataset of acted emotional speech containing seven emotion classes. The experimental findings show that the hybrid CNN-BiLSTM achieved an overall classification accuracy of 96.36%, significantly outperforming baseline models including GRU (91.84%), BiLSTM (93.12%), and CNN–GRU (94.67%). These findings highlight the effectiveness of combining spectral and temporal modeling for improved speech emotion recognition performance. Furthermore, our CNN+BiLSTM approach offers a computationally efficient and data-efficient alternative to transformer-based models, while still effectively capturing both spatial and temporal emotional cues in speech, making it suitable for real-time and resource-constrained applications.
A Corpus-Based Comparison of two Approaches for Emotion Annotation in French Texts
Valentina Dragos | Delphine Battistelli
Valentina Dragos | Delphine Battistelli
Emotion annotation in texts remains a challenging task in the field of Natural Language Processing (NLP), as, unlike voice or images, texts might not only contain peculiar cues to express emotions. Methods for emotion annotation are based on lexicons or on machine learning techniques which are based on the use of manually annotated corpora. This paper aims to explore if and how the combination of these two types of methods might be useful for the annotation of emotions in texts. Four data sets are used for comparison of the two approaches, and then to investigate to what extent the results are distinct or complementary on three aspects: (i) identification of emotional sentences; (ii) identification of emotion categories; (iii) identification of one specific mode of expression of emotions called “behavioral emotions” (e.g. shout, cry). Findings show that not all emotions are equally easy to annotate, and, most specifically, the learning-based approach tends to over detect Admiration.
Clarifying the Role of Psychological Factors in Language Acquisition: A Psycholinguistic Lexical Ratings Dataset
Wanwan Zheng
Wanwan Zheng
Lexical acquisition extends beyond the learning of surface-level word forms and encompasses underlying cognitive characteristics as well as broader processes that involve the learner. Nevertheless, in Japanese, as in many other languages, word difficulty has been characterized primarily by frequency and surface-level properties. Far less is known about the cognitive and affective dimensions that influence whether words are easier or more difficult to process. To address this gap, this study introduces a novel dataset that incorporates six psycholinguistic dimensions—familiarity, affective valence, arousal, imageability, abstractness, and understandability—collected through large-scale surveys of Japanese second language learners. Preliminary analyses of responses from 536 participants across 15 countries demonstrated that the dataset is both theoretically coherent and empirically reliable, consistent with established theories and findings while also yielding new insights into lexical processing. In addition to supporting more accurate estimations of word difficulty, the dataset provides a resource for future research by enabling systematic exploration of how lexical processing is shaped through the interaction of visual, affective, cognitive, and contextual factors.
Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research
Taryn Wong | Zeerak Talat | Hanan Aldarmaki | Anjalie Field
Taryn Wong | Zeerak Talat | Hanan Aldarmaki | Anjalie Field
Critical analyses of emotion recognition technology have raised ethical concerns around task validity and potential downstream impacts, urging researchers to ensure alignment between their stated motivations and practice. However, these discussions have not adequately influenced or drawn from research on speech emotion recognition (SER). We address this gap by conducting a systematic survey of SER research to uncover what stated motivations drive this work and if they align with the datasets and emotions studied. We find that while SER research identifies appealing goals—such as well-situated voice-activated systems or healthcare applications—commonly-used datasets do not reflect these proposed deployment contexts, thus presenting a gap between motivations and research practices. We argue that such gaps engender ethical concerns, and that SER research should reassert itself with concrete use-cases to prevent misinterpretations, misuse, and downstream harms.
Feeling First, Speaking Second: A Dual-Process Cognitive-Affective Architecture for LLM Agents
Nicolò Buscaroli | Fabio Tamburini
Nicolò Buscaroli | Fabio Tamburini
Current Large Language Models (LLMs) demonstrate exceptional generative capabilities but lack a coherent “inner life”, failing to model the dynamics of emotion regulation essential for believable affective behavior. Constrained by statelessness and a lack of theoretical grounding, standard models struggle to maintain psychological depth. To address this, we propose a computational cognitive-affective architecture grounded in Dual-Process Theory. Our system computationally distinguishes between visceral emotional reaction (Appraisal) and strategic verbal expression (Formulation), effectively operationalizing the gap between “feeling” and “saying”. This modular design allows agents to embody specific personas by integrating long-term memory, dynamic emotional states, personality and goals. We evaluated the system in simulated narrative scenarios using an LLM-as-a-judge protocol. We frame this system as a computational experiment to investigate the mechanics of artificial affect. Results confirm the feasibility of simulating a coherent and believable inner emotional monologue. However, analysis in high-pressure scenarios reveals a rational bias where strategic planning can override this emotional authenticity. These findings contribute to Computational Affective Science by demonstrating that while cognitive sequentiality successfully generates inner lives, enforcing affective primacy in the decision cycle is critical to prevent excessive rational regulation.
Emotion Recogniton in Conversations - empirical study
Rufaida Kashif | Benjamin Piwowarski | Helena Gomez Adorno
Rufaida Kashif | Benjamin Piwowarski | Helena Gomez Adorno
Emotion Recognition in Conversations (ERC) requires modeling complex contextual dependencies across dialog turns. While transformer-based models achieve strong performance on ERC benchmarks, several key design choices including context construction, optimization strategies, and imbalance handling remain insufficiently examined. In this work, we conduct a systematic empirical study of transformer-based ERC models across three benchmark datasets. We analyze the impact of context length and directionality, layer freezing, learning rate scheduling, parameter-efficient fine-tuning, and class imbalance mitigation strategies. Our results show that short-to-medium conversational context and moderate layer freezing provide stable and strong performance, while very long context windows, aggressive freezing, and parameter-efficient adaptation offer limited gains. Furthermore, imbalance-aware losses and data augmentation do not consistently outperform standard cross-entropy training. Overall, our findings provide practical insights into effective and stable design choices for transformer-based conversational emotion recognition.
Exploring Cross-Modal Interactions in Unimodal and Multimodal Emotion Recognition: An Empirical Study
Quanqi Du | Loic De Langhe | Els Lefever | Veronique Hoste
Quanqi Du | Loic De Langhe | Els Lefever | Veronique Hoste
Understanding how cross-modal interactions influence unimodal and multimodal emotion recognition remains an open question in multimodal affective computing. This study presents a systematic empirical investigation of how multimodal inputs affect both unimodal and multimodal emotion recognition performance. Using the UniC dataset, which provides modality-specific and global multimodal annotations across text, audio, and visual modalities, we conduct experiments based on the Tensor Fusion Network (TFN) under unimodal, bi-modal, and tri-modal configurations. Results show that cross-modal interactions exert complex and asymmetric effects. While additional modalities can provide complementary emotional cues, they may also introduce interference when signals diverge. Models continue to struggle with less frequent or extreme emotions such as disgust. Notably, multimodal embeddings combined with unimodal annotations outperform fully multimodal supervision in the same setup, highlighting the role of annotation consistency and cue reliability. These findings provide a systematic empirical validation of the long-assumed notions, demonstrating that cross-modal effects are not simply additive and highlighting the need for more interpretable multimodal fusion strategies.
Annotation Matters: Resolving Cross-Corpus Performance Drops in Hebrew Offensive Language Detection
Gili Berger Hefetz | Yossef Haim Shrem | Natalia Vanetik | Chaya Liebeskind
Gili Berger Hefetz | Yossef Haim Shrem | Natalia Vanetik | Chaya Liebeskind
Cross-dataset generalization remains a major challenge in offensive language detection, especially for culturally sensitive languages such as Hebrew. A large Hebrew dataset introduced in prior work (citation omitted for double-blind review) was annotated via a taxonomy-grounded, prompt-guided LLM protocol and achieved strong in-domain results. However, performance degraded sharply on two external Hebrew corpora. We investigate whether this degradation reflects domain shift or annotation shift, i.e., differences in how offensiveness is operationalized across datasets. Using the same prompt framework and a dual-LLM agreement procedure, we re-annotate both external corpora and quantify label divergence. We observe substantial mismatch between the original and new annotations, consistent with the view that offensiveness is not objective but depends on cultural context, discourse conventions, political framing, and the interpretation of irony. Evaluating models against the new labels yields markedly improved performance, and fine-tuning with the new external labels further improves results. Overall, our findings suggest that cross-dataset failure in affective NLP tasks may often be driven by annotation mismatch rather than domain adaptation limitations, highlighting the importance of annotation validity and culturally grounded labeling protocols.
Multi-Source Emotion Annotation in Children’s Language: When LLM Consensus Diverges from Human Judgment
Farida Said | Jeanne Villaneau
Farida Said | Jeanne Villaneau
Automated emotion annotation increasingly relies on inter-LLM agreement as a proxy for label quality. We test this assumption on 2,106 clause-level segments from interviews with French-speaking children (ages 6-11) about parental roles, a setting where affect is often implicit rather than lexically explicit. Using a 500-segment expert gold standard, we show that internal consensus can be seriously misleading: Dawid-Skene, a probabilistic label aggregation method, estimates GPT-5.2 valence accuracy at 90.7%, whereas evaluation against human gold yields 71.0%, revealing substantial overestimation driven by shared neutralization bias. Conversely, Dawid-Skene underestimates Claude Sonnet 4, reversing model ranking. Majority Vote, Dawid-Skene, and MACE produce near-identical consensus labels, suggesting that the main source of error lies in shared annotator bias rather than in the aggregation rule itself. We release the expert gold subset and the probabilistic corpus to support future work. Our results show that high inter-LLM agreement cannot replace external human validation for affect annotation.
This paper introduces EVOKE (Emotion Vocabulary of Korean and English), a Korean-English parallel dataset of emotion words. The dataset offers comprehensive coverage of emotion words in each language, in addition to many-to-many translations between words in the two languages and identification of language-specific emotion words. The dataset contains 1,426 Korean words and 1,397 English words, and we systematically annotate 819 Korean and 924 English adjectives and verbs. We also annotate multiple meanings of each word and their relationships, identifying polysemous emotion words and emotion-related metaphors. The dataset is, to our knowledge, the most systematic and theory-agnostic dataset of emotion words in both Korean and English to date. It can serve as a practical tool for emotion science, psycholinguistics, computational linguistics, and natural language processing, allowing researchers to adopt different views on the resource reflecting their needs and theoretical perspectives. The dataset is publicly available at https://github.com/yoonwonj/EVOKE.
Linguistic Distancing on Social Media: Indicators of Emotion Regulation Across Age Groups
Daniela Teodorescu | Saif M. Mohammad | Alona Fyshe
Daniela Teodorescu | Saif M. Mohammad | Alona Fyshe
Managing our emotional responses to events is key to emotional well-being, a process referred to as emotion regulation in psychology. Previous work has established that the degree to which we distance events is a type of emotion regulation. When we psychologically distance from events there can be markers in our language. These markers have been referred to as linguistic distancing. We build upon a previous metric to operationalize linguistic distancing, and explore how it changes across the lifespan. We explore this systematically by analyzing large amounts of social media text, a venue where people express their emotions. By investigating how distancing varies across age groups we can better understand how emotion regulation varies with age and provide initial benchmarks on social media data. We provide additional evidence further strengthening the hypothesis that linguistic distancing occurs in proportionally more instances with age. These findings align with past work in psychology which indicate improved well-being with older age. Better understanding how linguistic distancing changes with age is important because it functions as a marker of well-being and can inform effective health interventions. We provide a foundation for further exploring emotion regulation through linguistic distancing in text data.
MOSAIC : a Corpus of Small-Group Interactions During a Collaborative Task
Amine Benamara | Celine Clavel | Brian Ravenet | Nicolas Sabouret | Mathilde Sassier–Roublin | Julien Saunier
Amine Benamara | Celine Clavel | Brian Ravenet | Nicolas Sabouret | Mathilde Sassier–Roublin | Julien Saunier
This paper presents MOSAIC (Multimodal Observations of Social Affect, Intimacy, and Cohesion), a multimodal interaction corpus of video and audio recordings of 17 groups of 4 participants (68 participants in total) playing a collaborative board game. The aim of this corpus collection is to support the design of socially interactive agents. We describe the experimental protocol of this corpus collection, the perceptive questionnaires completed by participants, the automatic annotation process of game specific elements and non-verbal behaviors, and the manual verbal annotations collected. We provide a preliminary characterization of this corpus with descriptive results for some of the perceptive scales used in this study. We illustrate the possibilities offered by this corpus on the question of interpersonal social relations and group dynamics.
Age and Affect in Language: How Emotion Expression on Social Media Varies Across Adulthood
Daniela Teodorescu | Jan Philip Wahle | Saif M. Mohammad
Daniela Teodorescu | Jan Philip Wahle | Saif M. Mohammad
As we age, the way we experience and express emotions changes. This is because of a number of factors, including: changes in our body, differing types of experiences at different ages, improved emotion regulation strategies, and increasing experience of dealing with affective situations. However, work in psychology points to differing findings on how emotions, typically valence and happiness, changes with age. Psychologists measure happiness and well-being through questionnaires, which can have biases and result in limited data. Thus corpus analyses can provide useful complementary insights. We compile and release a large dataset of social media posts annotated with the age of the author at the time of posting. We refer to it as AgeCorpus. Using this dataset, we apply simple and interpretable methods to explore research questions pertaining to how social media posts, especially emotion expression through these posts, varies by age groups. Analyzing the emotions expressed in the posts, we find that the average valence increases until the middle ages, and then decreases; arousal decreases (Reddit)/plateaus with age (Twitter); and dominance follows the inverted U-shape (Reddit)/increases with age (Twitter). For categorical emotions, we find they follow the inverted U-shape on Reddit and increase in intensity with age on Twitter. We hope our dataset enables further research into age related phenomenon, such as well-being and language use.
Multimodal Affective Modeling in an LLM-based Intelligent Tutoring System for Foreign Language Learning
Dionysios Koulouris | Athasios Kallipolitis | Melina Tziomaka | Argyrios Zafeiriou | Stamatios Orfanos | Andreas Menychtas | Ilias Maglogiannis | George Tsoulouhas | Stamatia Michalopoulou | Athina Sioupi | Voula Giouli
Dionysios Koulouris | Athasios Kallipolitis | Melina Tziomaka | Argyrios Zafeiriou | Stamatios Orfanos | Andreas Menychtas | Ilias Maglogiannis | George Tsoulouhas | Stamatia Michalopoulou | Athina Sioupi | Voula Giouli
Foreign language learning is a cognitively and affectively demanding process, in which fluctuations in attention and motivation can negatively impact learner engagement. Emotions play a central role in this process, yet they are rarely modelled in a systematic, data-driven manner in authentic learning environments. At the same time, positive emotions. This paper presents a prototype affective computing architecture that incorporates various modalities (audio, video, biosignals) to facilitate real-time or near-real-time emotion recognition in an educational scenario; the architecture is integrated within an emotion-aware and adaptive Language Learning application that harnesses Large Language Models in view of providing appropriate educational scenarios to learners. The system comprises modules for acquiring data for each modality and a processing pipeline for synchronizing and analyzing heterogeneous affective signals. We demonstrate both the feasibility and applicability of the approach through a proof-of-concept implementation and discuss its relevance for studying learner affect and supporting affect-aware educational scenarios. The results highlight both the applicability of multimodal affective data in educational settings and the need for further research on their pedagogical interpretation and use.
Emotion classification has been extensively studied, with numerous datasets enabling progress in both textual and multimodal settings. However, most existing text-based resources treat emotion as an utterance-level property, assuming that the emotional content is fully encoded in the sentence itself. This assumption is problematic: in the absence of paralinguistic cues such as prosody, facial expressions, or emojis, textual emotions are often highly context-dependent. Many utterances lack explicit emotion markers, and even when present, such cues may be overridden by broader situational context. Sentence-level emotion annotation, thus, is driven by the annotator’s ability to imagine the context in which the given utterance would elicit a given emotion. An utterance may be able to express an emotion completely (Emotion Obvious), or it can express an emotion when imagined in a certain context (Emotion Plausible). Also, for an utterance, certain emotions might be implausible to express given the specific wording of a sentence (Emotion-Implausible). To address these issues, we create a new paradigm for emotion classification by categorizing utterance and emotion pairs into context-dependency classes. We present the PoETIC benchmark dataset, where sentences in the GoEmotions dataset are human-annotated for the three aforementioned classes across seven emotions (Fear, Anger, Sadness, Joy, Disgust, Surprise, and Neutral). We observe that gold-tagged emotions in GoEmotions do not have a clear correlation with human judgment with respect to the ability to express other emotions, given different contexts. Human annotators identify significantly more plausible emotions for a given utterance if asked to imagine a plausible context per utterance-emotion pair. We also present baselines using three popular large language models and two “small” language models in zero-shot and few-shot settings on the benchmark dataset.
From Sentiment to Valence in Metaphor: a Comparison of BERT-based Sentiment and Prompted Large Language Models
Rebecca Guolo | Ginevra Martinelli | Chiara Barattieri di San Pietro | Valentina Bambini
Rebecca Guolo | Ginevra Martinelli | Chiara Barattieri di San Pietro | Valentina Bambini
Although the affective dimension is a key aspect of metaphor, computational studies of figurative language have largely overlooked psycholinguistic variables such as valence. This study investigates whether computational models can reliably estimate the affective aspects of Italian and German metaphors and whether metaphor valence is compositionally derived. Outputs of BERT-based sentiment analysis and a valence-prompted LLM were compared with human ratings. Results show that the former exhibit limited alignment with human judgments, whereas higher agreement is achieved when the explicit concept of valence is prompted in a LLM. Both humans and models rely on the combined valence of the individual lemmas, suggesting a compositional contribution to metaphor valence.
Affect, Body, Cognition, Demographics, and Emotion: The ABCDE of Text Features for Computational Affective Science
Jan Philip Wahle | Krishnapriya Vishnubhotla | Bela Gipp | Saif M. Mohammad
Jan Philip Wahle | Krishnapriya Vishnubhotla | Bela Gipp | Saif M. Mohammad
Work in Computational Affective Science and Computational Social Science explores a wide variety of research questions about people, emotions, behavior, and health. Often they make use of language data that is first labeled with relevant information such as the use of emotion words and age of the speaker. Even though many resources and algorithms exist to enable such labeling, finding and using them is still a substantial impediment, especially to practitioners in fields outside of computer science. Here, we present the ABCDE dataset (“Affect, Body, Cognition, Demographics, and Emotion”), a large-scale collection of over 400 million released text instances from social media, blogs, books, and AI-generated sources, annotated for a number of features relevant to computational affective and social science. ABCDE facilitates inter-disciplinary research in wide range of fields, including affective science, cognitive science, the digital humanities, sociology, political science, and computational linguistics.
Affective computing is the development of systems that recognize, interpret, and simulate human emotions and has advanced rapidly through deep learning and multimodal fusion techniques. Yet this technological progress has significantly outpaced foundational understanding of what emotions are, how they function socially, and what it means for machines to simulate them. This position paper argues that current affective systems are critically flawed because they rely on correlational patterns rather than established psychological theories of affect, are trained on biased data that systematically fails to generalize across demographic groups, and operate within inadequate ethical and regulatory frameworks that cannot protect emotional privacy or prevent harm. Drawing on empirical work in affective science, computational fairness research, and philosophical accounts of emotional expression, we argue for a “theory-first” approach that integrates psychological models, mandates rigorous fairness auditing across intersectional demographics, treats affective data as a protected category requiring heightened safeguards, and recognizes fundamental limits to emotion recognition systems. Without such grounding, affective computing risks systematically encoding bias, enabling emotional manipulation, and eroding authentic human connections that define meaningful social experience.
Beyond Toxic Positivity: Interpersonal Affect Regulation in LLM-Based Dialogue Agents using Discourse Politeness Theory
Rina Sakagami | Emmanuel Ayedoun | Masataka Tokumaru
Rina Sakagami | Emmanuel Ayedoun | Masataka Tokumaru
In affective science, effective interpersonal emotion regulation requires behavioral inhibition when responding to severe emotional disclosures, temporarily suppressing intimacy to validate distress. However, while current Large Language Models (LLMs) excel at immediate sentiment recognition, long-term companion agents built upon them often adjust their conversational style based primarily on accumulated interaction time (psychological distance). This architectural overreliance on chronological intimacy causes systems to ignore the fluctuating emotional weight of specific topics. This results in “toxic positivity”: exaggerated optimism that invalidates negative affect and damages psychological safety. We propose a computational framework grounded in Discourse Politeness Theory that dynamically regulates interpersonal affect by calculating conversational strategy using two variables: Psychological Distance and Affective Weight of the Topic. When users disclose heavy emotional burdens, the system executes behavioral inhibition by suppressing Positive Politeness Strategies (intimacy, cheerfulness) and engaging Negative Politeness Strategies (hedging, validation). Through an 8-week longitudinal simulation evaluated by 18 third-party observers, our affective-regulation framework showed consistent advantages over a distance-only baseline. The framework was rated as more natural, empathetic, and fostering psychological safety. Among empathy-seeking participants, preference for the proposed model was consistent across all respondents. These exploratory findings suggest that computational interpersonal emotion regulation requires context-aware behavioral inhibition, not uniform friendliness.
up
Proceedings of the Third Workshop on Computation and Written Language (CAWL 2026) @ LREC 2026
Proceedings of the Third Workshop on Computation and Written Language (CAWL 2026) @ LREC 2026
Kyle Gorman
Kyle Gorman
Diacritics are orthographic marks that clarify pronunciation, distinguish similar words, or alter meaning. They play a central role in many writing systems, yet their impact on language technology has not been systematically quantified across scripts. While prior work has examined diacritics in individual languages, there’s no cross-linguistic, data-driven framework for measuring the degree to which writing systems rely on them and how this affects downstream tasks. We propose a data-driven framework for quantifying diacritic complexity using corpus-level, information-theoretic metrics that capture the frequency, ambiguity, and structural diversity of character-diacritic combinations. We compute these metrics over 24 corpora in 15 languages, spanning both single- and multi-diacritic scripts. We then examine how diacritic complexity correlates with performance on the task of diacritics restoration, evaluating BERT- and RNN-based models. We find that across languages, higher diacritic complexity is strongly associated with lower restoration accuracy. In single-diacritic scripts, where character-diacritic combinations are more predictable, frequency-based and structural measures largely align. In multi-diacritic scripts, however, structural complexity exhibits the strongest association with performance, surpassing frequency-based measures. These findings show that measurable properties of diacritic usage influence the performance of diacritic restoration models, demonstrating that orthographic complexity is not only descriptive but functionally relevant for modeling.
Private-Use Area Characters in the Wild: Signal or Noise?
Alexander Gutkin | Adrian Benton | Christo Kirov | Brian Roark | Lawrence Wolf-Sonkin
Alexander Gutkin | Adrian Benton | Christo Kirov | Brian Roark | Lawrence Wolf-Sonkin
The Private-Use Area (PUA) designation plays an important role in the Unicode standard. It covers several ranges of Unicode code points with no official character assignments. A PUA range is primarily used as a temporary representation mechanism for characters falling outside the official standard, to facilitate text entry and display of orthographies that are not otherwise adequately represented. The primary downside of PUA use is that characters lose their semantics if the pairing with the corresponding display font is broken, in which case they cannot be faithfully displayed in a general setting. Large-scale multilingual web corpora invariably contain PUA code points of unclear provenance, which may commonly be treated as noise and discarded. We investigate the distribution of PUA characters within large-scale web corpora, and analyze the resulting distributions across both scripts and writing systems. We show that, while the proportion of PUA-bearing paragraphs in the original corpora are small, PUA-bearing tokens can signal texts from under-represented languages. We additionally explore whether an off-the-shelf large language model (LLM) can classify PUA characters as constituting relevant orthographic signals versus punctuation or other noise. Our methods identify millions of paragraphs making use of such characters, and we argue that such data is important for the long tail of data-scarce orthographies. Moreover, as a primary Unicode mechanism for poorly represented writing systems, PUA characters are here to stay.
HAnnoI: A Handwriting Annotation Interface to Extract Data for Linguistic Analyses of Graphetic Detail
Joshua Wieler | Simon Petitjean | Kristian Berg | Henriette Huber | Stefan Hartmann
Joshua Wieler | Simon Petitjean | Kristian Berg | Henriette Huber | Stefan Hartmann
In this paper, we present HAnnoI – short for Handwriting Annotation Interface –, an open-source GUI application developed in Python that allows its users to identify and annotate so-called Regions of Interest (ROIs) within digital images. Several meta data such as their coordinates are retained for each ROI and they can be annotated on user-defined annotation layers. HAnnoI comes with a function to export all annotations to a CSV file, enabling further processing as well as quantitative analyses. HAnnoI also has a function to extract single PNG image files of all ROIs. It was developed to mark and annotate single letters in scans of handwritten (alphabetic) texts for linguistic analyses, yet it is not limited to this particular use case. In this paper, we first provide information on HAnnoI’s conception and technical details as well as an overview of alternative applications. We then showcase HAnnoI’s capabilities in a letter annotation task, where five annotators marked instances of lower case <s> in handwritten texts. Finally, we report on an exploratory analysis of this data, showing what kinds of investigations are enabled by using HAnnoI. The tool is available for free use at https://github.com/pywielR/HAnnoI.
SoriGraph: A New Database of Visual Feature-Level Descriptions of Written Korean
Wednesday Bushong | Hala Habahbeh | Ryan Jiang | Yoolim Kim
Wednesday Bushong | Hala Habahbeh | Ryan Jiang | Yoolim Kim
Phoneticians and phonologists have developed featural systems that enable systematic description of human speech sounds. However, no such systems exist for describing the visual features of writing systems. It is critical to understand the features of writing systems given their central role in many language users’ everyday experience. Just as phonetic and phonological features provide insight into speech perception, visual features can play a similar role for studying reading. In this paper, we introduce SoriGraph, a database of visual feature descriptions and IPA transcriptions for the full lexicon of Korean, drawing on a recent large-scale study of the visual features of writing systems. This database enables analysis of the visual and phonological properties of Korean and will be a critical resource for researchers. We describe the construction of the database and provide an overview of several potential uses of the database, and demonstrate one potential usage (information-theoretic analysis of lexicon structure).
Confusable Characters as Endangered Language Markers: The Case of North Caucasus Writing Systems
Alexander Gutkin | Adrian Benton | Christo Kirov | Brian Roark
Alexander Gutkin | Adrian Benton | Christo Kirov | Brian Roark
The Abkhaz-Adyghe and Nakh-Daghestanian language families encompass 35 living languages that possess arguably the most complex modern Cyrillic orthographies due to their very sophisticated phonology. The relevant online data displays idiosyncratic patterns among which the use of confusable characters in input methods is the most prevalent. This work studies one such character—letter palochka—that is shared by most of the writing systems in question. We investigate whether patterns including variants of this character can act as markers of these languages in large-scale web-crawled data. We use GlotLID, a wide-coverage off-the-shelf language identification (LID) model, to label paragraph-level web text that contains a palochka confusable, and estimate the effect of confusable character normalization on the quality of GlotLID’s predictions in 14 supported North Caucasian languages. According to GlotLID, the normalization significantly increases the recall (discovery of new language data) for some languages, while degrading it for others. However, manual evaluation reveals that overall, only 41% of ensuing wins and 46% of losses are accurate due to GlotLID prediction errors. We argue that, despite finding useful signals, higher precision LID approaches tailored to these long-tail languages are needed to improve the quality of mined data.
Abbreviations are an entrenched feature of the majority of the world’s writing systems, and the ability to expand abbreviations in context is important for speech technologies and language understanding tasks. This study presents several experiments applying large language models, prompt engineering, and fine-tuning to the expansion of ad-hoc abbreviations in English, showing substantial improvements against noisy channel models previously used for this task.
Evaluating Data Augmentation Strategies for Training Spanish Misspelling Detection Models
Manuel Castillo-Sancho | Jordi Porta | Asunción Gómez-Pérez
Manuel Castillo-Sancho | Jordi Porta | Asunción Gómez-Pérez
This paper evaluates three data augmentation strategies for training misspelling detection models in Spanish. Using the Spanish CORRSIC corpus of naturally occurring misspellings, we compare three misspelling generation methods: random perturbations, keyboard-based errors, and a statistical model derived from empirical edit patterns encoded as weighted finite-state transducers. We also analyze two word selection strategies (random and length-based) and two augmentation configurations designed to balance data diversity and reduce spurious correlations. This study shows that the statistical model produces misspellings most similar to real data, showing the lowest Jensen–Shannon divergence (0.148 nats) with the empirical distribution. In downstream detection experiments, performance improves with training size, and differences between word selection strategies remain minimal. Overall, the results highlight the value of statistically grounded misspelling generation for realistic and effective data augmentation in spell-checking tasks in Spanish.
Kazakh is written in the Arabic, Cyrillic, and Latin script which present unique challenges for OCR and post-OCR correction research. Despite this complexity, NLP research on Kazakh and its low-resource scripts remains extremely scarce. We analyze common OCR error patterns in all three Kazakh scripts using Tesseract and evaluate four large language models (LLMs) for post-OCR correction using minimal, confusion-aware, and few-shot prompting strategies. Our results reveal three systematic, writing-system-driven failure modes in LLM-based post-OCR correction: script switching, hallucination, and instruction-following breakdown. Arabic script post-OCR correction remains unsuccessful across all setups. In the Cyrillic script, post-OCR correction improvements are minimal due to the high baseline OCR performance on Cyrillic. For the Latin script, few-shot prompting with Gemini 2.5 Flash yields substantial improvements, reducing CER by 8.58 points and WER by 32.49 points to levels better than high-resource Kazakh Cyrillic script OCR. These findings demonstrate that LLM post-OCR correction failure modes are predictable from writing system properties such as script resource asymmetry and co-existing script dominance and demonstrate the need for typology-aware evaluation frameworks for multi-script and under-resourced languages.
Grapheme-to-phoneme (G2P) conversion plays a central role in speech technologies. This paper introduces G&P2P, a multi-source framework that integrates multiple pronunciation dictionaries to enhance G2P modeling. We evaluate both expert-curated and crowd-sourced resources using attentive LSTM, pointer-generator LSTM, and transformer architectures. Results indicate that combining high-quality expert dictionaries yields substantial improvements, achieving an 11.26-point absolute (22% relative) reduction in word error rate. In contrast, incorporating noisy crowd-sourced resources may degrade performance. Statistical analyses further suggest that dataset quality exerts a greater influence on outcomes than the choice of fusion strategy, offering practical guidance for the design of multi-source G2P systems.
We present a lightweight, corpus-based approach to abbreviation expansion that relies solely on contextual N-gram statistics. The method models local context using two-sided and one-sided bigram and trigram counts extracted from a large domain-specific corpus. Candidate expansions are selected through linear interpolation of context-specific evidence, enhanced with reliability-based scaling to mitigate sparse data effects. The approach does not require external linguistic resources, pretrained language models, or explicit morphosyntactic analysis, making it suitable for domain-specific and resource-constrained settings. Experiments conducted on a large Slovene medical corpus demonstrate that interpolation generally outperforms strict backoff strategies, with notable improvements for medium- and low-frequency abbreviations. Despite its simplicity, the proposed framework achieves robust performance while remaining computationally efficient and scalable.
Inverse Text Normalization for Arabic Numbers in Streaming ASR
Enas Albasiri | Myungjong Kim | Nourchene Ferchichi | Oluwatobi Olabiyi
Enas Albasiri | Myungjong Kim | Nourchene Ferchichi | Oluwatobi Olabiyi
Streaming multilingual speech recognition benefits from unified systems that produce numbers in their written form ‘34’ rather than their spoken form ‘thirty-four’. By generating digits directly, these systems eliminate the post-processing latency inherent in cascaded architectures that require a separate inverse text normalization (ITN) step. Arabic presents a formidable challenge for ITN; the system must not only determine the correct numerical value but also navigate complex rules for gender, number, and case marking that are determined by the counted noun. For instance, the digit ‘7’ (as in ‘47’) exhibits gender polarity: it must take a masculine form if modifying a feminine noun (e.g., Halala) and a feminine form if modifying a masculine noun (e.g., Riyal). While Arabic dialects typically exhibit simplified numeral systems by omitting case and gender markers, they vary significantly in verbalization patterns. This study explores the efficacy of a unified streaming Automatic Speech Recognition (ASR) system with integrated ITN features, comparing it against a traditional cascaded approach utilizing a post-processing rule-based ITN module. We utilize a FastConformer cache-aware streaming model trained on English and a diverse Arabic corpus spanning Modern Standard (MSA), dialectal, and Classical Arabic, while maintaining diacritics where contextually appropriate. We evaluate the system using Word Error Rate (WER) for ASR accuracy and exact match for ITN capability. Our results demonstrate that integrating ITN does not degrade core ASR performance and that the unified model achieves accuracy competitive with cascaded systems across Arabic variants. However, error analysis reveals that the primary failures in ITN are rooted in diacritization, gender polarity, and orthographic variation, highlighting the challenges of Arabic’s unique linguistic features in end-to-end modeling.
up
Proceedings of the Second workshop on Challenges in Processing South Asian Languages (CHiPSAL2026)
Proceedings of the Second workshop on Challenges in Processing South Asian Languages (CHiPSAL2026)
Kengatharaiyer Sarveswaran | Ashwini Vaidya
Kengatharaiyer Sarveswaran | Ashwini Vaidya
Findings of the Second Workshop on Challenges in Processing South Asian Languages (CHiPSAL 2026)
Kengatharaiyer Sarveswaran | Surendrabikram Thapa | Ashwini Vaidya | Tafseer Ahmed | Bal Krishna Bal
Kengatharaiyer Sarveswaran | Surendrabikram Thapa | Ashwini Vaidya | Tafseer Ahmed | Bal Krishna Bal
This paper presents the findings of the second workshop on Challenges in Processing South Asian Languages (CHiPSAL 2026), held as part of LREC 2026. South Asia is one of the most linguistically diverse regions in the world, yet its languages remain severely underrepresented in language resources and technologies, particularly in the era of large language models (LLMs). The workshop brings together research addressing key challenges in this space, including data scarcity, morphological complexity, code-mixing, script diversity, and the lack of culturally grounded evaluation benchmarks. The workshop received 57 submissions, covering a wide range of languages, tasks, and modalities, including both widely spoken languages (e.g., Bengali, Hindi, Tamil, and Urdu) and extremely low-resource and endangered languages such as Burushaski, Limbu, and Nepal Bhasha (Newari). Several contributions introduce arguably first-of-their-kind resources and benchmarks for these languages, spanning both text and speech domains, and focusing on linguistically informed and culturally grounded data creation. In addition to the main track, the workshop hosted a shared task on Multimodal Hate and Sentiment Understanding in Low-Resource Memes for Nepali, attracting strong community participation. The results highlight the effectiveness of multimodal approaches while also revealing persistent challenges in modelling culturally nuanced and low-resource data. Across the accepted papers and shared task, key insights include the central role of high-quality data, the limitations of current multilingual models in low-resource settings, and the need for culturally aware and data-centric approaches. Overall, CHiPSAL 2026 demonstrates the growing momentum in South Asian language processing and highlights the importance of sustained, community-driven efforts to build inclusive and representative language technologies.
Development of Burushaski Speech - English Text Translation Dataset
Tauqeer Saleem | Abdul Samad | Azkaa Nasir | Adina Adnan Mansoor | Fatima Faisal | Mahrukh Yousuf
Tauqeer Saleem | Abdul Samad | Azkaa Nasir | Adina Adnan Mansoor | Fatima Faisal | Mahrukh Yousuf
Burushaski is a language isolate spoken in northern Pakistan with a predominantly oral tradition, limited standardized orthography, and virtually no existing speech technology infrastructure. These characteristics make conventional text-centric NLP pipelines unsuitable and position speech data collection as the primary scientific challenge. This paper introduces an audio-first, linguistically informed methodology for developing a Burushaski–English speech translation resource. Rather than prioritizing model architecture, we focus on principled corpus design tailored to the language’s morphological complexity, ergative-absolutive alignment, and four-gender agreement system. The dataset combines structured elicitation targeting high-frequency and morphologically diverse constructions, functional and formulaic speech, and oral narratives that capture discourse-level phenomena. We describe the design of a custom data collection application, community-embedded crowdsourcing strategy, and translation-aligned workflow for generating parallel speech–English data. The resulting pilot corpus comprises approximately 10 hours of curated audio from 42 speakers across controlled and naturalistic settings. While we present the results of preliminary translation experiments using Whisper, the primary contribution of this work is methodological: a scalable framework for speech-first corpus development in morphologically rich, under-resourced, and predominantly oral languages. We argue that for languages lacking stable orthography and large textual corpora, data design, not model selection constitutes the central research problem
Despite negation being one of the core element in any language, it remains a challenging phenomenon for modern Large Language Models(LLMs). Recently, there have been growing efforts to evaluate how models handle negation. However, the existing probing datasets are mostly English-centric. To facilitate evaluation for Indian Languages especially Telugu, which has complex morphological features, we present NEGTEG benchmark. This benchmark is a test suite that contains 5 tasks: Negation Detection, Negation Translation, Paraphrase Detection, Sentiment Analysis and Polarity Flipping. The test suite is designed based on strong linguistic analysis and includes annotations of different negation types. This helps us evaluate how models perform across various forms of negation. We use the benchmark to probe the negation handling capabilities of multilingual language models at different levels and our evaluation reveals that most of the models struggle significantly with Telugu negation across all tasks.
This paper presents the first ever morphological transducer for the Limbu language, also known by its endonym Yakthung Pan, an endangered Sino-Tibetan language primarily spoken in the area known as Limbuwan/Koshi Province in Eastern Nepal, with a minority population in the Sikkim state of India, where Limbu enjoys official status. Using a corpus of Limbu text produced by field interviews (Michailovsky, 1977), and a translation of the Holy Bible into the Limbu language, the paper presents various elements of the morphology of the language, how they were implemented into the transducer, and an evaluation of the transducer against identified corpora and a gold standard. With a relatively small lexicon, the transducer was found to have reasonable coverage, with high precision but low recall. The paper discusses future expansion through further involvement with the community, which can help in the maintenance and revitalisation of the endangered language.
Evaluating Large Language Models for Medical Named Entity Recognition in Urdu: A Benchmark Study
Bushra Nasim | Kinza Latif | Muhammad Zohair | Muhammad Hassan Asif | Zarmeen Nasim
Bushra Nasim | Kinza Latif | Muhammad Zohair | Muhammad Hassan Asif | Zarmeen Nasim
Medical named entity recognition (NER) is a crucial task in natural language processing (NLP) for extracting meaningful entities such as diseases, symptoms, medications, body parts, and treatments from clinical text. However, NER in low-resource languages like Urdu remains underexplored due to limited annotated datasets. In this study, we evaluated the performance of two state-of-the-art large language models (LLMs), ChatGPT-4o and LLAMA 3.2, on Urdu medical NER using a dataset of 2,057 health-related Urdu news headlines manually annotated across five entity categories. Both models were evaluated using precision, recall, and F1-score. It was found that both models exhibited low precision and moderate recall. ChatGPT-4o achieved the highest F1 for Disease (0.35) while LLAMA 3.2 reached slightly lower F1 scores for Disease (0.33). Both models performed poorly on treatment-related terms, with F1 scores of 0.036 (LLAMA 3.2) and 0.011 (ChatGPT-4o). Micro-average F1-scores were 0.187 for ChatGPT-4o and 0.183 for LLAMA 3.2, indicating comparable overall performance. These findings highlight the challenges of medical NER in low-resource languages and underscore the need for domain-specific fine-tuning, transfer learning, few-shot learning, and prompt engineering to improve performance.
DR-RAG: Addressing Retrieval Misalignment in Low-Resource Urdu Question Answering
Saad Ahmad | Muhammad Hammad | Muhammad Zeeshan | Faizad Ullah | Asim Karim
Saad Ahmad | Muhammad Hammad | Muhammad Zeeshan | Faizad Ullah | Asim Karim
Retrieval-Augmented Generation performs well on English QA benchmarks, but degrades considerably in morphologically rich, low-resource languages. Urdu presents a particularly challenging case: heavy inflectional morphology, Nastaliq script inconsistencies, and limited training data produce a systematic mismatch between query representations and indexed document content that standard retrieval architectures cannot bridge. We propose DR-RAG (Dual-Representation Retrieval-Augmented Generation), which addresses this through dual indexing. Each document is represented as overlapping text chunks and as automatically generated question-answer pairs. Queries are first matched against the QA index, which aligns more reliably with natural query phrasing than declarative document chunks. When retrieval confidence falls below τ = 0.80, the system falls back to chunk-based retrieval, maintaining coverage without sacrificing precision. Evaluated on Urdu UQA and English SQuAD 2.0, DR-RAG improves Urdu METEOR by 38×, ROUGE-1 by 140%, and reduces generation latency by 43%. LLM-as judge scores show higher faithfulness (3.03 vs 1.93) and overall quality (2.99 vs 2.21) over MultiVector. English performance remains competitive throughout. These results indicate that representation-level alignment between queries and indexed content, rather than increased model complexity, is the critical factor for reliable retrieval in underserved South Asian languages.
Cross-Domain Evaluation of Transformer-Based Models for Punjabi Speech Emotion Recognition
Fatima Tu Zahra | Kulsoom Asim | Sandesh Kumar | Abdul Samad
Fatima Tu Zahra | Kulsoom Asim | Sandesh Kumar | Abdul Samad
Speech Emotion Recognition (SER) is an important part of human–computer interaction, but most existing research focuses on high-resource languages, with very limited work on regional languages such as Punjabi. This paper focuses on detecting emotions from Punjabi speech using machine learning and deep learning techniques. We curated our own Punjabi speech emotion dataset using volunteer recordings and real-world sources, covering four emotion classes: angry, happy, sad, and neutral. The data was preprocessed for consistency and evaluated using a multi-strategy framework (E1–E4) to test domain generalization. Three models were evaluated: CNN, ResNet-34, and the transformer-based Wav2Vec 2.0. Among these, the ResNet-34 model performed the best in the combined-domain strategy (E4), achieving a test accuracy of 96%. While cross-corpus evaluations (E2, E3) highlighted challenges in generalizing to neutral emotions, the model achieved perfect scores for happy and sad classes in E4. These results demonstrate the effectiveness of residual networks and combined-domain training for emotion recognition in low-resource languages and highlight the potential for further work on Punjabi SER.
SiPaKosa: A Comprehensive Corpus of Canonical and Classical Buddhist Texts in Sinhala and Pali
Ranidu Hansaka Gurusinghe | Nevidu Jayatilleke
Ranidu Hansaka Gurusinghe | Nevidu Jayatilleke
SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhist documents alongside the complete web-scraped Tripi t. aka canonical texts. The corpus was created through high-quality OCR using Google Document AI on historical manuscripts, combined with systematic web scraping of canonical repositories, followed by rigorous quality control and metadata annotation. The corpus is organised into language-specific subcorpora: Sinhala and Mixed Sinhala-Pali. We evaluate the performance of language models using ten pretrained models, with perplexity scores ranging from 1.09 to 189.67 on our corpus. This analysis shows that proprietary models significantly outperform open-source alternatives by factors of three to six times. This corpus supports the pretraining of domain-adapted language models, facilitates historical language analysis, and aids in the development of information retrieval systems for Buddhist scholarship while preserving Sinhala cultural heritage.
BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference
Farah Binta Haque | Md Yasin | Shishir Saha | Md Shoaib Akhter Rafi | Farig Sadeque
Farah Binta Haque | Md Yasin | Shishir Saha | Md Shoaib Akhter Rafi | Farig Sadeque
Despite the growing progress in Natural Language Inference (NLI) research, resources for the Bengali language remain extremely limited. Existing Bengali NLI datasets exhibit several inconsistencies, including annotation errors, ambiguous sentence pairs, and inadequate linguistic diversity, which hinder effective model training and evaluation. To address these limitations, we introduce BNLI, a refined and linguistically curated Bengali NLI dataset designed to support robust language understanding and inference modeling. The dataset was constructed through a rigorous annotation pipeline emphasizing semantic clarity and balance across entailment, contradiction, and neutrality classes. We benchmarked BNLI using a suite of state-of-the-art transformer-based architectures, including multilingual and Bengali-specific models, to assess their ability to capture complex semantic relations in Bengali text. The experimental findings highlight the improved reliability and interpretability achieved with BNLI, establishing it as a strong foundation for advancing research in Bengali and other low-resource language inference tasks. The link to the BNLI dataset: https://github.com/FarahHaque/BNLI-Dataset.git
Exploring Large Language Models for Multitask Learning in Bengali Text Classification
Md. Sajjad Hossain | Kawsar Ahmed | Suny Md Ashraf Khan | Mohammed Moshiul Hoque
Md. Sajjad Hossain | Kawsar Ahmed | Suny Md Ashraf Khan | Mohammed Moshiul Hoque
Text classification in low-resource languages has become increasingly important due to the rapid growth of user-generated digital content. While multitask learning has long been studied in NLP, the use of LLMs for multitask text classification in low-resource languages such as Bengali remains underexplored. Although LLMs are inherently multilingual and multitasking, their effectiveness in structured multitask classification settings for Bengali has not been systematically evaluated. In this work, we investigate how LLMs can be leveraged for multitask Bengali text classification across five domains: sentiment analysis, aggressive text detection, fake news detection, news categorization, and emotion analysis. We compare in-context learning strategies—including zero-shot, one-shot, and chain-of-thought prompting—with parameter-efficient fine-tuning approaches. Our findings show that CoT prompting does not consistently improve performance and often degrades performance, highlighting the instability of prompt-based adaptation in low-resource settings with limited pretraining exposure. Moreover, reasoning-optimized models such as DeepSeek-R1 exhibit substantial performance drops, indicating that enhanced reasoning capabilities alone cannot overcome the challenges posed by low-resource settings. Among the evaluated mLLMs, Gemma-3-4B demonstrates the most stable and balanced cross-task performance under both in-context learning and parameter-efficient fine-tuning, making it a strong backbone candidate for multitask Bengali text classification. These results provide empirical evidence on the limitations of prompting and the advantages of lightweight fine-tuning for low-resource multilingual NLP.
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
Rishikesh Kumar Sharma | Safal Narshing Shrestha | Jenny Poudel | Rupak Tiwari | Arju Shrestha | Rupak Raj Ghimire | Bal Krishna Bal
Rishikesh Kumar Sharma | Safal Narshing Shrestha | Jenny Poudel | Rupak Tiwari | Arju Shrestha | Rupak Raj Ghimire | Bal Krishna Bal
Nepal Bhasha (Newari), an endangered language of the Kathmandu Valley, remains digitally marginalized due to the severe scarcity of annotated speech resources. In this work, we introduce NwÄchÄ MunÄ, a newly curated 5.39-hour manually transcribed Devanagari speech corpus for Nepal Bhasha, and establish the first benchmark using script-preserving acoustic modeling. We investigate whether proximal cross-lingual transfer from a geographically and linguistically adjacent language (Nepali) can rival large-scale multilingual pretraining in an ultra-low-resource Automatic Speech Recognition (ASR) setting. Fine-tuning a Nepali Conformer model reduces the Character Error Rate (CER) from a 52.54% zero-shot baseline to 17.59% with data augmentation, effectively matching the performance of the multilingual Whisper-Small model despite utilizing significantly fewer parameters. Our findings demonstrate that proximal transfer from Nepali language serves as a computationally efficient alternative to massive multilingual models. We openly release the dataset and benchmarks to digitally enable the Newari community and foster further research in Nepal Bhasha.
From Romanized to Devanagari: Enhancing Nepali Sentiment Analysis with NepaliXlit
Suraj Patel | Kashish Kumari Dhami | Norden Sherpa | Supriya Khadka
Suraj Patel | Kashish Kumari Dhami | Norden Sherpa | Supriya Khadka
Romanized Nepali is the dominant medium of social media communication in Nepal, yet most multilingual NLP models are trained on Devanagari, creating a noticeable drop in performance in informal settings. To address this script mismatch, we develop NepaliXlit, a transliteration model fine-tuned from IndicXlit to better handle the phonetic variability of Romanized Nepali. Trained on 2,943 informal word pairs and evaluated on 736 held-out pairs, NepaliXlit improves transliteration accuracy by 8% and reduces character error rate by 11%. We use sentiment analysis as a testbed to understand whether transliteration actually helps downstream NLP tasks. We curate over 6,500 Romanized social media comments and construct a balanced subset of 1,518 manually annotated instances. Baseline experiments show that multilingual encoder models struggle with Romanized input; however, transliterating text into Devanagari using NepaliXlit consistently improves sentiment classification accuracy with mBERT and MuRIL. Comparative evaluation against large language models (LLMs) further reveals that generative models such as Gemini and GPT variants exhibit strong cross-script generalization and outperform encoder-based baselines. Our results indicate that adaptive transliteration enhances conventional multilingual models, while modern LLMs offer a better alternative for multi-script, low-resource settings.
Why Does Low-Rank Adaptation Work for Hindi-English Code-Mixing? A Geometric Analysis
Shashank Vishwakarma | Rakesh Kumar
Shashank Vishwakarma | Rakesh Kumar
Low-Rank Adaptation (LoRA) enables efficient fine-tuning of large language models, yet why it works particularly well for code-mixed text remains unexplained. We propose that LoRA’s efficiency stems from geometric structure in multilingual pre-trained models: code-mixed embeddings concentrate in low-dimensional cross-lingual subspaces. Through spectral analysis of mBERT and MuRIL on Hindi-English (Hinglish) data, we establish that pre-trained attention weights have effective ranks of 437–441, while LoRA updates (r = 4,8,16) exhibit ranks of 2.1–5.9—a 136× average compression. Cross-lingual geometry measured via Centered Kernel Alignment shows Hinglish embeddings align strongly with Hindi (CKA=0.279) but weakly with English (0.093), compared to a monolingual baseline of 0.074. Statistical tests (Wilcoxon p < 10−19) and permutation ablations confirm these differences are robust. We interpret the convergence of geometric overlap (3.77× baseline) and empirical compression (136×) as evidence that low-rank adaptation exploits pre-existing multilingual structure. Findings are demonstrated on token-level language identification; extensions to other language pairs and tasks remain open questions.
Hi-SEMFLOW: Lie Algebra–Based Semantic Flow for Span-Level Informal Language Identification in Hindi
Manikandan Ravikiran | Tanmay Tiwari | Vibhu Gupta | Rohit Saluja
Manikandan Ravikiran | Tanmay Tiwari | Vibhu Gupta | Rohit Saluja
Informal Hindi text frequently contains multi-token slang and idiomatic expressions whose correct identification requires consistent span boundaries. Transformer-based token classifiers, despite strong contextual representations, often produce fragmented or structurally invalid BIO sequences due to largely local predictions. We propose Hi-SEMFLOW, a Lie algebra–based semantic flow framework that models span consistency as a continuous refinement process over label logits. Instead of discrete structured decoding (e.g., CRFs), Hi-SEMFLOW learns context-dependent transition operators derived from antisymmetric generators and propagates structural information through smooth, fully differentiable transformations. This formulation integrates structural bias directly into end-to-end training without requiring dynamic programming or hard decoding constraints. Experiments on the HiSlang-4.9k benchmark show that Hi-SEMFLOW improves span-level F1 by up to 2–3 absolute points and yields consistent macro-F1 gains across Hindi-pretrained encoders. Extensive ablations demonstrate that continuous geometric refinement provides a flexible and effective alternative to discrete structured decoding for span-centric sequence labeling.
NeCCo: Nepali Cultural Commonsense Benchmark for Large Language Model Evaluation
Sanket Shrestha | Raunak Regmi | Sadikshya Ghimire | Satyam Rana | Supriya Khadka
Sanket Shrestha | Raunak Regmi | Sadikshya Ghimire | Satyam Rana | Supriya Khadka
Large language models perform strongly on standard evaluations, yet these benchmarks prioritize high-resource languages and culturally dominant knowledge, leaving culture-specific commonsense underexamined. In low-resource languages such as Nepali, everyday communication depends on culturally embedded cues, including kinship hierarchies, ritual practices, food systems, idioms, and honorific distinctions that literal translation often fails to capture. As a result, models that appear competent on global metrics can perform poorly in local contexts. To address this gap, we introduce NeCCo, a curated multiple-choice benchmark for culturally situated reasoning across five domains: kinship and social hierarchy; festivals, rituals, and geography; idioms, proverbs, and metaphors; commonsense and daily life; and gastronomy, agriculture, and nature. The dataset was created through structured authoring, cross-review, and normalization, and is released in Devanagari, English, and Romanized formats. We evaluate multiple state-of-the-art LLMs using standardized prompting and controlled decoding. Results show substantial variation: models perform better on globally documented knowledge such as geography, but struggle with relational and linguistically implicit tasks, including extended kinship reasoning and proverb interpretation. The most culturally dense categories expose brittleness and increased hallucination. These findings suggest that multilingual competence requires more than translation coverage and highlight the need for culturally grounded benchmarks and training signals.
Reward-Guided Fine-Tuning of Whisper for Low-Resource Nepali Speech Recognition
Aadarsh Pandit | Yudhin Khanal | Ishan Pandey | Kushal Kunwar | Sunil Regmi
Aadarsh Pandit | Yudhin Khanal | Ishan Pandey | Kushal Kunwar | Sunil Regmi
Fine tuning speech recognition models on noisy real world data is tricky. The model has no way of knowing which training samples are reliable and which are not, so it ends up learning from bad examples just as readily as good ones. This is a real problem for Nepali, where most available training data comes from YouTube videos with automatically generated subtitles that are often inaccurate. In this work, we tried a simple fix. Instead of feeding everything to the model, we first asked humans to rate the quality of a sample of transcriptions, trained a small Random Forest classifier on those 2,000 ratings, and used it to filter out the bad samples before each retraining round. The classifier uses four automatically computable features, Word Error Rate (WER), Character Error Rate (CER), length ratio, and length difference, and achieves 81% accuracy on a held out set. Running two filtering and retraining cycles on a 40,000 clip training subset drawn from a 68.4 hour corpus improves substantially over our own standard fine tuning baseline of 5.60% WER and 5.10% CER, reaching 4.89% WER and 4.52% CER, which corresponds to an 11 to 13% relative gain. The approach is much lighter than full Reinforcement Learning from Human Feedback but still uses real human judgment to guide training.
Evaluating Linguistic Knowledge of LLMs in Tamil: The ILAKKANAM Benchmark
Jeyarajalingam Varsha | Menan Velayuthan | Sumirtha Karunakaran | Rasan Nivethiga | Kengatharaiyer Sarveswaran
Jeyarajalingam Varsha | Menan Velayuthan | Sumirtha Karunakaran | Rasan Nivethiga | Kengatharaiyer Sarveswaran
Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich languages such as Tamil remains largely unexplored. Existing multilingual benchmarks often rely on translated English datasets, failing to capture the language specific linguistic and cultural nuances of the target language. To address this gap, we introduce ILAKKANAM, the first Tamil-specific linguistic evaluation benchmark manually curated using 820 questions from Sri Lankan school-level Tamil subject examination papers spanning Grades 1–13. Each question is annotated by trained linguists under five linguistic categories and a factual knowledge category. We evaluate both closed-source and open-source LLMs using a standardized evaluation pipeline. Our results show that Gemini 2.5 achieves the highest overall performance, while open-source models lag behind, highlighting the gap in linguistic grounding. Category- and grade-wise analyses reveal that all models perform well on lower-grade questions but show a clear decline as the grade level and the linguistic complexity of the questions increase. Further, no strong correlation is observed between a model’s overall performance and its ability to identify linguistic categories, suggesting that performance may be driven by exposure rather than genuine understanding. The code and dataset used in this study are publicly available in our repository, where the dataset consists only of extracted examination questions to mitigate potential data leakage. Keywords: Tamil, Linguistic Benchmark, Linguistic diagnostics, Low-resource language
A Feature-Fusion Ensemble Approach for Tamil Hate Speech Detection
Sathasivam Nerujan | Kengatharaiyer Sarveswaran
Sathasivam Nerujan | Kengatharaiyer Sarveswaran
Detecting online toxicity in morphologically rich, low-resource languages like Tamil remains a major computational challenge. Standard transformer models often struggle with sub-word fragmentation, which can dilute the semantic intensity of regional insults and out-of-vocabulary slang. To mitigate this limitation, we train a multi-layer hybrid framework that fuses the deep contextual representations of L3Cube-TamilBERT with the character-level robustness of FastText embeddings. Our architecture leverages Last-4 Layers averaging and a dual pooling strategy (Mean + Max) to capture both global sentence intent and extract high-activation spikes of offensive cues typically lost in single layer representations. Experiments show that this hybrid model achieves a Macro-F1 of 0.7883, notably enhancing Hate Recall (0.7503) for detection of offensive content. Additionally, as reported by other studies, stacking ensemble achieves peak hate precision (0.9296), providing a high accuracy alternative for moderation scenarios requiring minimal false positives. By combining deep contextual hidden states with FastText embeddings, the proposed feature-fusion ensemble approach with multi-layer hybrid framework approach establishes a new benchmark for hate speech detection for Tamil.
Comparative Analysis of Tokenizers in Tamil Text Classification in Low Resource Settings
Gokulan Sivakumaran | Randil Pushpananda | ERAD Bandara
Gokulan Sivakumaran | Randil Pushpananda | ERAD Bandara
Tokenization is crucial in NLP, influencing performance for morphologically rich, low resource languages like Tamil. This study comprehensively analyzes WordPiece, SentencePiece, and Byte-Level Byte Pair Encoding (BBPE) for Tamil text classification. We assess tokenization efficiency using metrics including token count, fragmentation, OOV rate, and compression ratio. Additionally, we analyze downstream impact through Tamil news title classification using a custom lightweight BERT based Transformer architecture. Tokenizers were pretrained on a 5.45 GB Tamil Corpus and evaluated on a Kaggle Tamil News Dataset. Results indicate WordPiece and SentencePiece outperform BBPE in efficiency and accuracy. While BBPE eliminates OOV words, excessive fragmentation hinders model learning. Increasing vocabulary size improves WordPiece and SentencePiece but not BBPE. Misclassification analysis highlights overfragmentation challenges. This study contributes to Tamil NLP by comparing tokenizers, aiding researchers in selecting appropriate strategies for agglutinative languages.
Improving Public Health Safety in Low-Resource Languages Using a Human-Verified Health Misinformation Corpus and Large Language Models
Sujal Maharjan | Astha Shrestha | Laxmi Thapa | Sweta Poudel | Shuvam Shiwakoti | Rabin Thapa | Kritesh Rauniyar | Surendrabikram Thapa
Sujal Maharjan | Astha Shrestha | Laxmi Thapa | Sweta Poudel | Shuvam Shiwakoti | Rabin Thapa | Kritesh Rauniyar | Surendrabikram Thapa
The proliferation of health misinformation in Low-Resource Languages (LRLs) poses a severe threat to public health, yet automated detection remains critically under-studied due to the scarcity of high-quality benchmarks. We address this gap by introducing Nep-Health-Misinfo, a novel human-verified corpus for health misinformation identification in Nepali. The dataset was developed by adapting four foundational benchmarks (Monkeypox-V1, Monkeypox-V2, COVID-19, and CoAID) through a systematic Machine Translation Post-Editing (MTPE) protocol involving native experts. Our evaluation of Neural Machine Translation (NMT) systems reveals a significant translation asymmetry: while state-of-the-art (SOTA) systems achieve a BLEU score of 43.21 on factual health data, performance degrades sharply on deceptive narratives, with BLEU and TER scores dropping to 19.11 and 62.42, respectively. To establish robust baselines, we benchmark seven recent open-weight Large Language Models (LLMs), including Qwen2.5-7B-Instruct, Gemma-3-4B-IT, and Ministral-8B-Instruct, across zero-shot and few-shot settings. For the few-shot evaluation, we compare stochastic sampling against a K-means centroid-based approach for semantically representative exemplar selection. Experimental results indicate that Qwen2.5-7B-Instruct achieves a peak Macro F1-score of 0.8488, improving over its zero-shot performance (0.7188) on the same dataset. Our findings demonstrate that while few-shot prompting effectively mitigates distribution shifts in low-resource medical contexts, performance remains highly sensitive to the semantic density of exemplars. This work provides the first human-verified Nepali health misinformation corpus. All code and resources are available at https://github.com/SUJAL390/Nep-Health-Misinfo-CHIPSAL-LREC.
Multimodal Hate and Sentiment Understanding in Low-Resource Text-Embedded Images for Online Safety and Digital Well-being
Surendrabikram Thapa | Shuvam Shiwakoti | Siddhant Bikram Shah | Kritesh Rauniyar | Laxmi Thapa | Surabhi Adhikari | Kristina T Johnson | Kengatharaiyer Sarveswaran | Bal Krishna Bal | Usman Naseem
Surendrabikram Thapa | Shuvam Shiwakoti | Siddhant Bikram Shah | Kritesh Rauniyar | Laxmi Thapa | Surabhi Adhikari | Kristina T Johnson | Kengatharaiyer Sarveswaran | Bal Krishna Bal | Usman Naseem
This paper presents an overview of the Shared Task on Multimodal Hate and Sentiment Understanding in Low-Resource Memes, organized as part of the Second Workshop on Challenges in Processing South Asian Languages (CHiPSAL 2026) at LREC 2026. The task addresses automated content understanding in low-resource settings by focusing on monolingual Nepali memes written in Devanagari script. Built upon the NeMeme dataset, the task comprises two subtasks: (1) binary hate speech detection and (2) three-class sentiment analysis. The competition attracted 23 teams for hate detection and 13 teams for sentiment analysis. Participating teams employed diverse strategies, including late-fusion multimodal architectures combining multilingual text encoders with vision models, caption-based approaches using large vision-language models, and ensemble techniques. The top-performing system achieved macro-F1 scores of 80.52% on hate detection and 68.81% on sentiment analysis using a late-fusion hybrid architecture with discriminative learning rates. Our analysis reveals that multimodal fusion consistently outperforms unimodal baselines, sentiment analysis poses greater challenges than hate detection due to increased semantic nuance, and the scarcity of Devanagari-centric pretrained models remains a significant bottleneck. This shared task establishes a benchmark for multimodal understanding in low-resource South Asian languages and provides insights for developing inclusive content moderation systems.
Unigoa@CHiPSAL 2026: Early vs Late Fusion for Multimodal Hate and Sentiment Detection in Nepali Memes
Ashweta Fondekar | Milind Shivolkar | Jyoti Pawar
Ashweta Fondekar | Milind Shivolkar | Jyoti Pawar
Internet memes pose significant challenges for automatic content moderation due to the interaction of visual and textual cues, sarcasm, and cultural context. In this work, we participate in the CHiPSAL 2026 shared task on multimodal hate and sentiment understanding in Nepali memes. The task consists of two subtasks: binary hate speech detection and three-class sentiment classification. We investigate both early-fusion and late-fusion multimodal architectures. Our primary system employs a late-fusion dual-encoder architecture combining XLM-RoBERTa for multilingual text representation and CLIP for visual encoding. We further evaluate an early-fusion ViLT-based joint vision–language transformer using NepBERTa tokenization as a baseline. Experimental results show that late-fusion models consistently outperform early-fusion architectures, particularly for code-mixed memes containing Devanagari Nepali and Roman-script English text. Our best system achieves a Macro-F1 of 0.6564 for hate speech detection and 0.4859 for sentiment classification. We provide analysis highlighting the challenges of multilingual code-mixing, sarcasm, and implicit sentiment in low-resource multimodal settings.
HasNat@CHiPSAL 2026: Multimodal Hate Speech Detection in Low-Resource Nepali Memes Using Aligned Vision–Language Models
Alvee Hasan Chowdhury | Md. Abul Hasnat | Adnan Faisal
Alvee Hasan Chowdhury | Md. Abul Hasnat | Adnan Faisal
Memes are widely used for communication on social media but are increasingly exploited to spread hate and harmful stereotypes. Detecting hate speech in memes is particularly challenging because meaning is conveyed jointly through images and embedded text, and the problem becomes more complex in low-resource languages such as Nepali. In this work, we participate in Subtask A of the CHiPSAL 2026 Shared Task, focusing on hate speech detection in Nepali-only memes. We benchmark three multimodal vision language backbones, ViT-B-32 (OpenCLIP), AltCLIP, and BLIP2+mT5, under controlled preprocessing and augmentation settings. Our best-performing system uses AltCLIP to extract aligned text and image representations, followed by a late-fusion classifier trained with stratified 5-fold cross-validation to address class imbalance. The proposed model achieves a macro F1-score of 0.66 on the validation set. Experimental results highlight the effectiveness of aligned vision language representations and demonstrate that preprocessing and augmentation strategies have model-dependent effects in low-resource multimodal hate speech detection.
TeamHerald@CHIPSAL 2026: Hate Speech Detection and Sentiment Analysis of Nepali Memes Using Transformer-based Architectures and Ensemble Learning
Ashish Acharya | Anish Khatiwada | Rohit Khadka | Pragya Aryal
Ashish Acharya | Anish Khatiwada | Rohit Khadka | Pragya Aryal
The analysis of internet memes in the Nepali language is complicated by frequent code-mixing and a lack of established baseline resources. While memes inherently combine visual and textual elements, this study focuses on a text-centric approach by extracting embedded text using an OCR layer and modeling it with Transformer-based architectures. We evaluate six distinct models and investigate the comparative effectiveness of Hard and Soft Voting ensemble strategies across two tasks: binary hate speech detection and three-class sentiment analysis. Experimental results show that a standalone decoder-only model achieved the highest performance for binary classification, whereas the Soft Voting ensemble performed best for the multi-class sentiment task, yielding a 15.8% relative improvement in Macro F1-score over the strongest standalone baseline. These findings suggest that ensemble strategies behave differently across binary and multi-class tasks, highlighting the importance of selecting aggregation methods suited to the classification objective.
NeuralNoodles@CHiPSAL 2026: Late-Fusion Multimodal Stacking for Nepali Meme Sentiment Classification
Sidratul Muntaha | Sabila Anzum | Arpita Mallik | Hasan Murad
Sidratul Muntaha | Sabila Anzum | Arpita Mallik | Hasan Murad
Memes have emerged to be an essential medium of online expression, where the sentiment is determined by the interaction of text and image. Sentiment analysis of memes is particularly challenging when the language is low-resource, such as Nepali, due to the lack of resources and the complex relationships between text and image modalities. In this paper, we report our submission to Subtask B of CHiPSAL 2026, where the task was sentiment analysis of Nepali text-embedded memes for three sentiment classes: Negative, Neutral, and Positive. Through this submission, we present a late fusion multimodal framework that encompasses lexical, semantic, and visual models through a cross-validated stacking approach. Our submission to the shared task competition received a Macro F1 of 0.5045 on the official test set, achieving 6th place in the leaderboard. This demonstrates the strength of well-structured late fusion approaches to multimodal sentiment analysis of text-embedded memes.
Multi-Modal-Minds@CHiPSAL 2026: A Comparative Study of Textual, Visual and Multimodal Architecture for Nepali Meme Moderation
Sandesh Shrestha | Bikram K.C. | Akshyat Shah | Ashish Acharya | Rabin Thapa
Sandesh Shrestha | Bikram K.C. | Akshyat Shah | Ashish Acharya | Rabin Thapa
Memes have become ubiquitous on social media platforms blending text and imagery to express complex and culturally nuanced messages. While a high degree of automation in meme moderation has been achieved for high-resource languages, low-resource languages, such as Nepali, still remain largely neglected. In this paper, we describe our system submission to the CHiPSAL 2026 Shared Task on Multi-modal Hate and Sentiment Understanding in Low-Resource Nepali Memes, which features two main sub-tasks: (1) Detection of HateSpeech as binary classification and (2) Sentiment Analysis as multi-class classification in Nepali memes. We perform a comprehensive analysis of the following models: uni-modal textual models (mBERT, XLM-RoBERTa,MuRIL), uni-modal visual models (ResNet, ConvNeXt, ViT), nine different late-fusion multimodal models, and the vision-language foundation model, SigLIP. Among all models, the ViT model achieved the best macro F1-score(0.6278) for the hate speech detection task, while SigLIP achieved the best score (0.5481) for the sentiment analysis task. We hypothesize that the under-performance of fusion models may be attributed to OCR noise and inadequate low-resource textual representations that act as a bottleneck when paired with more advanced visual encoders. These results highlight the unique challenges of multimodal meme comprehension in low-resource contexts and underscores the requirement for culturally grounded, noise-robust approaches to content moderation in Nepali.
linus@CHiPSAL 2026: Multimodal Hate Speech and Sentiment Detection in Low-Resource Memes Using Late-Fusion Hybrid Architecture
Sunil Regmi | Bipesh Subedi | Saugat Singh | Suman Shrestha
Sunil Regmi | Bipesh Subedi | Saugat Singh | Suman Shrestha
The increased sharing of memes on social media creates serious challenges for automated moderation, especially in low-resource and code-mixed languages such as Nepali. In this paper, we present our system for the CHiPSAL 2026 Shared Task on Multimodal Hate and Sentiment Understanding in Low-Resource Memes. We propose a late-fusion hybrid architecture that combines OpenAI’s Vision Transformer (CLIP ViT-B/32) with a domain-specific Nepali language model (NepBERTa) to capture both visual features and linguistic information. To address data scarcity, we introduce a cross-task label mapping and data augmentation strategy between the hate speech and sentiment datasets. By applying controlled hyperparameter settings and balanced loss optimization, our framework achieved a Macro F1 score of 0.8052 on Subtask A (Hate Speech Detection) and 0.6881 on Subtask B (Sentiment Analysis) in the official CodaBench evaluation, demonstrating the effectiveness of the proposed multimodal approach.
ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification
Nitiz Khanal
Nitiz Khanal
This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subtasks: binary hate speech classification and three-class sentiment analysis. Our approach adapts the Robust Adaptation of Hateful Meme Detection (RA-HMD) framework using Qwen3-VL-8B-Instruct, a state-of-the-art vision-language model with native Devanagari support. We employ a two-stage training pipeline: (1) LoRA fine-tuning with an MLP projection head for generative classification, and (2) contrastive backbone fine-tuning with supervised InfoNCE loss. We handle class imbalance through minority oversampling, image augmentation, and focal loss. At inference, we ensemble Stage 1 token probabilities with Stage 2 classifier scores using validation-tuned weights. Our end-to-end approach eliminates error propagation from separate OCR and translation pipelines by leveraging the model’s native Devanagari understanding. Our system achieved 2nd place on hate speech detection (F1: 0.797) and 4th place on sentiment analysis (F1: 0.518). We provide detailed ablations, error analysis, and insights into adapting large vision-language models for low-resource South Asian languages.
Team Oryu@CHiPSAL 2026: Integrating Text and Vision Transformers for Multimodal Hate Speech Detection in Memes
Noore Tamanna Orny | Joyeta Barua Moni | Md. Abtahee Kabir | Hasan Murad
Noore Tamanna Orny | Joyeta Barua Moni | Md. Abtahee Kabir | Hasan Murad
With the proliferation of multimodal content on various social media platforms, automated hate speech detection has emerged as a challenge, especially in meme-based communication, where meaning arises from interactions between text and images. In these situations, unimodal techniques are inadequate in capturing semantics. In order to address such issues, a late-fusion-based multimodal hate speech detection framework has been proposed and implemented for the CHiPSAL shared task. In the proposed framework, multimodal content is processed by utilizing XLM-RoBERTa for multilingual text representation and a Vision Transformer (ViT) for visual representation. Both modal representations are fused using a fully connected classification head and are used for binary hate speech detection. The findings suggest that multimodal content effectively captures features from individual modalities and helps improve hate speech detection accuracy by obtaining a Macro F1-score of 0.66 and ranking 5th on the leaderboard. Also, transformer-based multimodal fusion performs effectively and acts as a reliable baseline for hate speech detection in low-resource multilingual meme-based communication scenarios.
Cuet Yet Another Baseline@CHiPSAL LREC 2026: Shared Task on Multimodal Sentiment Understanding in Low-Resource Memes
Rotna Dipika Debnath | Shahrin Afroz Hoque Ruhi | Ayesha Labiba | Arpita Mallik | Hasan Murad
Rotna Dipika Debnath | Shahrin Afroz Hoque Ruhi | Ayesha Labiba | Arpita Mallik | Hasan Murad
Memes serve as a method to express feelings such as humor, sarcasm, and diverse viewpoints. The task of identifying sentiment in memes is becoming increasingly complex, particularly in low-resource languages like Nepali where memes often combine images, texts, and code-mixed language. However, multimodal methods for sentiment analysis in Nepali memes seem to be insufficient. In this paper, we present our system for the Subtask B(Sentiment Analysis) for Shared Task on Multimodal Hate and Sentiment Understanding in Low-Resource Memes@CHiPSAL LREC 2026. We implement various unimodal models, such as XLM-RoBERTa-large,MuRIL-base, Twitter-XLM-R for text. Moreover, we incorporate BLIP-2 captions to enhance visual-text understanding and adopted a multimodal approach that fuses textual embeddings, image embeddings, caption embeddings, and similarity scores. The fused features process through cross-attention and a dense neural network for classification, with focal loss and class weighting used to improve performance. Our approach achieved a macro F1 score of 0.50 securing 7th place and highlighting the importance of cross-modal interaction and large-scale pretrained vision-language models for robust meme understanding in sentiment analysis.
MEME-Fusion@CHiPSAL 2026: Multimodal Ablation Study of Hate Detection and Sentiment Analysis on Nepali Memes
Samir Wagle | Reewaj Khanal | Abiral Adhikari
Samir Wagle | Reewaj Khanal | Abiral Adhikari
Hate speech detection in Devanagari-scripted social media memes presents compounded challenges: multimodal content structure, script-specific linguistic complexity, and extreme data scarcity in low-resource settings. This paper presents our system for the CHiPSAL 2026 shared task, addressing both Subtask A (binary hate speech detection) and Subtask B (three-class sentiment classification: positive, neutral, negative). We propose a hybrid cross-modal attention fusion architecture that combines CLIP (ViT-B/32) for visual encoding with BGE-M3 for multilingual text representation, connected through 4-head self-attention and a learnable gating network that dynamically weights modality contributions on a per-sample basis. Systematic evaluation across eight model configurations demonstrates that explicit cross-modal reasoning achieves a 5.9% F1-macro improvement over text-only baselines on Subtask A, while uncovering two unexpected but critical findings: English-centric vision models exhibit near-random performance on Devanagari script, and standard ensemble methods catastrophically degrade under data scarcity (N ≈ 850 per fold) due to correlated overfitting. Code and implementation details are available at a repository that has been anonymized for the review process and will be fully disclosed in the final version
eGrantha.ai@CHiPSAL 2026: Stochastic Image Captioning for Robust Hate Speech Detection in Low-Resource Nepali Memes
Anish Thapaliya
Anish Thapaliya
This paper presents a system for hate speech detection in low-resource Nepali memes, submitted as part of Subtask A of the Shared Task on Multimodal Understanding at CHiPSAL 2026. Detecting hateful memes is particularly challenging due to the combination of images, text, and emojis used to portray humor, satire, or sociopolitical commentary, as well as the low-resource nature of the Nepali language. We investigate a range of unimodal and multimodal modeling strategies, including text-only, vision-text, and caption-based approaches. For caption generation, the Gemini family of models (Gemini 2.X and Gemini 3.X) was used to produce contextually rich captions, which are publicly released as NeMeme-CAP on Hugging Face. Caption-based modeling leverages stochastic caption augmentation to address class imbalance and Test-Time Augmentation (TTA) to reduce prediction variance and improve model robustness. The best-performing system fine-tunes an encoder-only transformer model, RoBERTa-base, on the generated captions, achieving third place on the official leaderboard with a macro-averaged F1-score of 0.7397. The code is publicly available at https://github.com/thapaliya123/LREC-CHiPSAL-2026.
EthosAI@CHiPSAL2026: Hate and Sentiment Understanding in Low-Resource Memes Using a Multimodal Approach
Vinayak Bansal | Deepawali Sharma | Aakash Singh | Vivek Kumar Singh
Vinayak Bansal | Deepawali Sharma | Aakash Singh | Vivek Kumar Singh
Memes have become a popular way for people to share opinions and emotions on social media, but they are also often used to spread hate and negative sentiments. In this paper, we present our multimodal approach to the CHiPSAL 2026 shared task on multimodal hate and sentiment detection in Nepali memes, which includes two subtasks: hate detection and sentiment analysis. Since memes usually combine both text and images, we first experimented with different unimodal models for text and images separately. After identifying the top two best-performing text and image models, combined them using different fusion techniques. The results show that multimodal models outperform unimodal ones, highlighting that both textual and visual information are important for understanding the context of memes. The multi- modal model, which combines sentence-transformers/LaBSE for text and ResNet-18 for image using weighted Fusion technique, achieved a macro F1 score of 0.6614 for Subtask A and sentence-Transformers/LaBSE for text and deit- Base for image using simple Fusion technique, achieved a macro F1 score of 0.4839 for SubTask B, on the test dataset.
up
Proceedings of the Third Workshop on Patient-Oriented Language Processing (CL4Health) @ LREC 2026
Proceedings of the Third Workshop on Patient-Oriented Language Processing (CL4Health) @ LREC 2026
Deepak Gupta | Paul Thompson | Sophia Ananiadou | Dina Demner-Fushman
Deepak Gupta | Paul Thompson | Sophia Ananiadou | Dina Demner-Fushman
FHIRPath-QA: Executable Question Answering over FHIR Electronic Health Records
Michael Frew | Nishit Bheda | Bryan Tripp
Michael Frew | Nishit Bheda | Bryan Tripp
Though patients are increasingly granted digital access to their electronic health records (EHRs), existing interfaces may not support precise, trustworthy answers to patient-specific questions. Large language models (LLM) show promise in clinical question answering (QA), but retrieval-based approaches are computationally inefficient, prone to hallucination, and difficult to deploy over real-life EHRs. This work introduces FHIRPath-QA, the first open dataset and benchmark for patient-specific QA that includes open-standard FHIRPath queries over real-world clinical data. A text-to-FHIRPath QA paradigm is proposed that shifts reasoning from free-text generation to FHIRPath query synthesis. For o4-mini, this reduced average token usage by 391× relative to retrieval-first prompting (629,829 vs 1,609 tokens per question) and lowered failure rates from 0.36 to 0.09 on clinician-phrased questions. Built on MIMIC-IV on FHIR Demo, the dataset pairs over 14k natural language questions in patient and clinician phrasing with validated FHIRPath queries and answers. Empirically, the evaluated LLMs achieve at most 42% accuracy, highlighting the challenge of the task, but benefit strongly from supervised fine-tuning, with query synthesis accuracy improving from 27% to 79% for 4o-mini. These results highlight that text-to-FHIRPath synthesis has the potential to serve as a practical foundation for safe, efficient, and interoperable consumer health applications, and the FHIRPath-QA dataset and benchmark serve as a starting point for future research on the topic. The full dataset and generation code can be accessed on GitHub.
COACH Meets QUORUM: A Framework and Pipeline for Aligning User, Expert, and Developer Perspectives in LLM-Generated Health Counselling
Yee Man Ng | Bram van Dijk | Pieter Beynen | Otto Boekesteijn | Joris Jansen | Gerard van Oortmerssen | Max J. van Duijn | Marco Spruit
Yee Man Ng | Bram van Dijk | Pieter Beynen | Otto Boekesteijn | Joris Jansen | Gerard van Oortmerssen | Max J. van Duijn | Marco Spruit
Systems that collect data on sleep, mood, and activities can provide valuable lifestyle counselling to populations affected by chronic disease and its consequences. Such systems are, however, challenging to develop; in addition to reliably extracting patterns from user-specific data, systems should contextualise these patterns with validated medical knowledge to ensure the quality of counselling and generate counselling that is relevant to a real user. We present QUORUM, an evaluation framework that unifies these developer-, expert-, and user-centric perspectives, and show with a real case study that it meaningfully tracks convergence and divergence in stakeholder perspectives. We also present COACH, a Large Language Model-driven pipeline to generate personalised lifestyle counselling for our Healthy Chronos use case, a diary app for cancer patients and survivors. Applying our framework indicates that, overall, users, medical experts, and developers converge on the view that the generated counselling is relevant, of good quality, and reliable. However, stakeholders also diverge on the tone of the counselling, sensitivity to errors in pattern-extraction, and potential hallucinations. These findings highlight the importance of multi-stakeholder evaluation for consumer health language technologies and illustrate how a unified evaluation framework can support trustworthy, patient-centered NLP systems in real-world settings.
Addressing Domain Shift in Health Coaching Note Analysis through Factorized Synthetic Data Generation
Michael Tänzer | Iva Bojic | Ashwini Yuvraj Lawate | Andy Hau Yan Ho | Andy Khong
Michael Tänzer | Iva Bojic | Ashwini Yuvraj Lawate | Andy Hau Yan Ho | Andy Khong
Automatic extraction of behavioral goals from health coaching notes is essential for scalable monitoring of coaching programs, yet training data is scarce and exhibits substantial domain shift across programs. We collect and annotate 157 notes from a coaching program and show that models trained on the only existing public corpus, SMARTSpan (173 notes), suffer a drop of up to 30 points in exact-match F1 when transferred to our data. To address this, we propose a factorized synthetic data generation pipeline that decomposes note variation into three largely independent axes, health coach documentation structure, patient goal content, and patient persona, extracts empirical priors from a small in-domain seed set, and samples from them to produce diverse synthetic notes with embedded goal-span labels validated via cycle-consistency filtering. In low-resource experiments with only 57 in-domain training notes, our approach outperforms rephrasing and backtranslation baselines on both exact-match and partial-match F1. Ablation analysis demonstrates that augmentation must target the in-domain distribution to be effective, and a human evaluation confirms that synthetic notes are structurally faithful, with detection driven by surface artifacts rather than content or organizational flaws.All code and generated data will be published at GitHub repository: https://github.com/Michael-Tanzer/cl4health-factorized-augmentation.
Scalable Generation of Adult-Oriented Therapeutic Reading Texts for Russian Aphasia Rehabilitation
Anastasia Kolmogorova | Anastasia Margolina | Alina Telnova | Igor Ilchenko
Anastasia Kolmogorova | Anastasia Margolina | Alina Telnova | Igor Ilchenko
Texts are widely used in aphasia rehabilitation to support the recovery of comprehension and narrative planning. In routine practice, clinical impact depends strongly on patient motivation and on the availability of age-appropriate reading materials: adults are often offered child-oriented texts, which can be perceived as demeaning and may reduce engagement. We present a controllable generation pipeline for building a repository of Russian therapeutic reading texts for adult aphasia therapy. An anonymized repository with code and data is available at https://github.com/z00logist/aphasia-exercises-generation. The pipeline conditions each story on an explicit semantic triplet (12 topics-10 archetypes-11 objects) and enforces three clinically motivated complexity regimes (Basic/Intermediate/Advanced). Using batched prompting, we generate 1,296 unique stories. We evaluate the corpus with classical linguistic metrics and a LLM-as-a-judge protocol (18 binary criteria); on a stratified sample of 198 stories, overall rubric compliance is 80.0%. Surface metrics show a monotonic increase in lexical and syntactic complexity across regimes, and Basic texts closely match a small clinical anchor set of 10 therapist-authored texts. Judge-based analysis indicates near-perfect adherence to high-level narrative constraints but persistent limitations in fine-grained phonotactic control, motivating hybrid neuro-symbolic enforcement.
An Open-Resource Knowledge Augmentation for Biomedical Lay Summarization
João Pedro Veloso | Evelin Amorim
João Pedro Veloso | Evelin Amorim
Automatic summarization aims to generate concise versions of texts while retaining relevant information. Summaries can be either extractive, using direct excerpts, or abstractive, rephrasing content to convey the same meaning. Lay summarization applies abstractive techniques to simplify complex texts, such as scientific literature, for broader audiences, thereby promoting public understanding of specialized knowledge. Prior work shows that knowledge augmentation improves lay summarization. Still, biomedical applications often rely on closed resources like the Unified Medical Language System (UMLS), which require expert curation and are costly to scale. We propose a four-step approach that leverages keyword extraction and DBpedia, an open general domain knowledge base, ideal to bridge the gap between expert and lay knowledge. First, we extract keywords from biomedical texts using YAKE!, a well-established unsupervised method. Second, we query DBpedia using these keywords to retrieve relevant concept entries. Third, we construct a graph of concepts for each document based on cosine similarity between DBpedia entries. Finally, we combine each graph with the original abstract to train a summarization model. Our method achieves competitive performance compared to UMLS-based systems in the eLife dataset (ROUGE-1: 58.44 vs. 60.26, ROUGE-L: 43.45 vs. 45.45), demonstrating that open-resource approaches can provide viable alternatives to licensed knowledge bases while maintaining accessibility for resource-constrained organizations.
TabMedQA: From Structured Data to Question-Answer Datasets in Early Clinical Decision-Making
Gabriel Iturra Bocaz | Petra Galuščáková | Sol Gedde Vedde | Alvaro Fernandez-Quilez
Gabriel Iturra Bocaz | Petra Galuščáková | Sol Gedde Vedde | Alvaro Fernandez-Quilez
The rising adoption of Large Language Models (LLMs) and Retrieval Augmented Generation (RAG) in clinical general practice demands datasets that capture realistic early-stage clinical decision-making, where experts must decide on follow-up actions based on sparse, structured patient data. Existing medical Question–Answering (QA) resources primarily address post-diagnostic or specialist settings and rarely reflect how General Practitioners (GPs) document and justify early decisions based on clinical observations from Electronic Health Records (EHRs) and grounded on clinical guidelines. We present TabMedQA, a framework for synthesizing QA collections that emulate how GPs formulate and document decisions in encounter notes during early patient assessments. TabMedQA leverages instruction-tuned LLMs, guided by disease-specific clinical guidelines, to generate full encounter notes composed of a guideline-grounded justification and a corresponding follow-up recommendation directly from structured EHR inputs. The framework further supports RAG-based evaluation, simulating how GPs might consult previous patient encounters to inform new consultations. We demonstrate the application and resulting resource use of TabMedQA on prostate cancer using the publicly available PI-CAI collection and release the resulting PI-CAI QA collection, resource generation templates, and TabMedQA code. To the best of our knowledge, TabMedQA provides the first open framework for creating guideline-grounded, EHR-based QA collections that enable the generation and holistic evaluation of LLM-produced clinical encounter notes, bridging decision-making accuracy with clinical encounter quality in general practice
SimpliMED: Automatic Simplification of Cardiology Discharge Reports Using Large Language Models
Lucas Molino-Piñar | Manuel Carlos Díaz Galiano | María-Teresa Martín-Valdivia | Jose Angel Urbano-Moral | Elena Sola-Garcia
Lucas Molino-Piñar | Manuel Carlos Díaz Galiano | María-Teresa Martín-Valdivia | Jose Angel Urbano-Moral | Elena Sola-Garcia
Medical discharge reports frequently contain highly technical language that creates significant communication barriers between healthcare professionals and patients, potentially compromising treatment adherence and post-discharge care quality. In this paper, we present SimpliMED, a modular system designed to automatically simplify cardiology discharge reports using Large Language Models (LLMs) and advanced Natural Language Processing techniques (NLP). Our architecture integrates section-based preprocessing with specialized prompts, explicit handling of medical abbreviations, and therapeutic explanations of medications to enhance accessibility. We evaluate our system using a corpus of 307 anonymized cardiology discharge reports from a Spanish medical center. For abbreviation detection, our fine-tuned Small Language Model (SLM) achieves an F1-score of 0.90, significantly outperforming regex-based approaches (F1: 0.67). For medication recognition, we achieve F1-scores of 0.91 for commercial names and 0.70 for active principles. We also contribute a therapeutic dictionary containing 14,611 medications with patient-friendly explanations extracted from the Spanish Agency of Medicines. Expert evaluation by two cardiologists yields an overall quality score of 75%, with highest performance for admission reason (91%) and current illness (75%) sections. While results demonstrate the potential of LLM-based medical text simplification for Spanish clinical language, we identify areas requiring further development before clinical deployment.
Dutch Metaphor Extraction from Cancer Patients’ Interviews and Forum Data Using LLMs and Human in the Loop
Lifeng Han | David Lindevelt | Sander Puts | Erik van Mulligen | Suzan Verberne
Lifeng Han | David Lindevelt | Sander Puts | Erik van Mulligen | Suzan Verberne
Metaphors and Metaphorical Languages (MLs) play an important role in healthcare for the information communication between clinicians, patients, and patients’ family members. In this work, we focus on the Dutch language and cancer patients’ data. We extract the metaphors used by patients using two data resources: 1) cancer patient storytelling interview data, 2) online forum, data including cancer patients’ posts, comments, and questions to professionals. We investigate how current state of the art LLMs and perform on this task by exploring different prompting strategies such as Chain of Thought, few-shot learning, and self-prompting. With human in the loop, we verify the extracted metaphors and collect the output as a corpus, named “HealthQuote.NL”. We believe the extracted metaphors can be useful for supporting better patient care, e.g. shared decision making, helping communication between patients and clinicians, patient health literacy, etc. It can also be integrated into the design of a care path. We share our prompts and resources at https://github.com/4dpicture/HealthQuote.NL
Compressed Representations of Patient Records: A Comparative Study of Template-Based and LLM-Based Methods for Clinical Data Summarization and Visualization
Andreas Stöckl | Oliver Krauss | Sophie Bauernfeind
Andreas Stöckl | Oliver Krauss | Sophie Bauernfeind
Electronic Health Records (EHRs) contain comprehensive patient information that is often voluminous and challenging to review efficiently. This paper presents a systematic evaluation of multiple methods for compressing patient records into standardized, comparable formats. Four compression approaches are implemented and compared: two template-based methods (structured extraction, extractive key-phrase) and two LLM-based methods (LLM, and hybrid LLM with 8 different models). Using a synthetic cohort of 75 patient records generated with realistic clinical patterns, each method is evaluated on information preservation (diagnosis, medication, allergy, lab value recall, and vital accuracy), compression efficiency, and output quality. Across methods, diagnosis recall ranged from 0.637 to 1.000, with medication and allergy recall consistently exceeding 0.880. In the test setup, the template-based approach yielded the highest compression ratio (7.6×), while the hybrid methods provided the most balanced trade-off between compression and clinical utility. These results suggest that combining structured extraction with LLM-generated summaries can be an effective strategy for scenarios requiring both compact representations and contextual clinical information.
Medical Text Rewriting for Non-Experts: A Guideline-Driven LLM Approach
Mana Kuramoto | Hiroyuki Nagai | Keiko Yamada | Hiroo Ide | Masayo Hayakawa | Tomohiro Nishiyama | Shoko Wakamiya | Eiji Aramaki
Mana Kuramoto | Hiroyuki Nagai | Keiko Yamada | Hiroo Ide | Masayo Hayakawa | Tomohiro Nishiyama | Shoko Wakamiya | Eiji Aramaki
Medical research is highly specialized, making it difficult for patients and general readers to understand recent findings.Traditionally, text simplification, replacing technical terms with more accessible expressions, has been employed. However, this approach alone is limited in addressing a lack of background knowledge and often results in the loss of important information.Therefore, this study defines “rewriting for non-experts” as a rewriting process that, in addition to simplification, supplements essential background knowledge such as the significance of the research and reasons it is needed and proposes a method for implementing this process using large language models (LLMs).To verify the effectiveness of the proposed approach, a quantitative evaluation using automatic metrics was conducted. The results showed that the method combining the guidelines for human text creation with few-shot examples of reference texts achieved the highest scores.The expansion of the guidelines is planned as part of future work to enable the rewriting of scientific and technological information in a form that is accessible to a broader audience.
Patient-Specific Care Pathway Visualisation for Medical Nursing Staff
Sophie Bauernfeind | Selina Adlberger | Oliver Krauss | Andreas Stöckl
Sophie Bauernfeind | Selina Adlberger | Oliver Krauss | Andreas Stöckl
Nursing staff are increasingly confronted with extensive and detailed patient documentation, requiring much time to read through numerous possible care measures. Combined with rising patient loads, this underscores the need for a clearer and more immediately accessible overview of each patient’s situation. Patient-specific care pathway visualisations offer a promising approach to reduce cognitive load, support faster decision-making, and improve situational awareness. This work investigates two Artificial intelligence (AI)-assisted methods for generating such visualisations: (1) simple image generation based on structured textual prompts, and (2) automated code generation that produces graph-based representations of clinical pathways. Using a dataset of synthetic patient profiles and seven defined care pathways, evaluating multiple state-of-the-art foundation models. The results highlight clear differences between models and approaches, particularly in language sensitivity, structural consistency, and the level of detail achievable. Image-based outputs provided visually rich overviews but frequently introduced subtle logical inconsistencies, while code-based methods produced verifiable and structurally coherent pathways yet varied in their ability to preserve contextual and psychosocial information. Together, these findings indicate that AI-assisted visualisation can effectively support—but not yet fully automate—patient-specific pathway generation, and they point toward hybrid solutions that combine visual accessibility with logical robustness.
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
Adrian Cosma | Cosmin Dumitrache | Emilian Radoi
Adrian Cosma | Cosmin Dumitrache | Emilian Radoi
Text-based telemedicine has become an increasingly used mode of care, requiring clinicians to deliver medical advice clearly and effectively in writing. As platforms increasingly rely on patient ratings and feedback, clinicians face growing pressure to maintain satisfaction scores, even though these evaluations often reflect communication quality more than clinical accuracy. We analyse patient satisfaction signals in Romanian text-based telemedicine. Using a sample of anonymised text-based telemedicine consultations, we model feedback as a binary outcome, treating thumbs-up responses as positive and grouping negative or absent feedback into the other class. We extract from doctor responses interpretable, predominantly language-agnostic features (e.g., length, structural characteristics, readability proxies), along with Romanian LIWC psycholinguistic features and politeness/hedging markers where available. We train a classifier with a time-based split and perform SHAP-based analyses, which indicate that metadata dominates prediction, functioning as a strong prior, while characteristics of the response text provide a smaller but actionable signal. In subgroup correlation analyses, politeness and hedging are consistently associated with positive patient feedback, whereas lexical diversity shows a negative association.
Matching patients to clinical trials is a critical bottleneck hindered by complex eligibility criteria. While conversational AI offers a promising solution, its safe deployment depends on high-quality, domain specific data. This paper introduces three benchmark datasets designed to support the development and evaluation of conversational agents for clinical trial pre-screening. First, a manually-annotated paired-criterion dataset provides a gold standard for structuring raw criteria, which we used to objectively group 12,596 criteria. Second, we curated a human-authored question benchmark to validate the clinical fidelity and patient-centric clarity of questions generated by a medical LLM, ensuring the AI’s dialogue is accurate and understandable. Third, we constructed a human-validated assessment corpus of criterion-question-answer tuples with human-labeled outcomes to evaluate criterion classification based on a patient’s answer to a generated question. The primary contribution of this work is a foundational set of benchmark datasets, designed to support and evaluate key components for a chatbot for clinical trial search.
HealthTrajectory: Patient Journey Summaries and Visualizations for Patient-Clinician Communication Support
Rohmah Hidayah | Tomohiro Nishiyama | Shoko Wakamiya | Eiji Aramaki
Rohmah Hidayah | Tomohiro Nishiyama | Shoko Wakamiya | Eiji Aramaki
In recent years, patient narratives have been used to understand subjective experiences that are not recorded in clinical notes. However, narratives tend to be long and unstructured, requiring summarization. However, text-based summaries often require a lot of clarification from patients and make it difficult for clinicians to review events and changes in symptoms over time. In this study, we expanded the summary output by presenting a visualization of the patient’s journey to facilitate communication between patients and medical staff. Referring to the widespread use of LLM for summarization, we compared GPT-4.1 and Gemini-2.5-pro, and used Gemini-3-pro-image-preview for visualization. Data was collected from DIPEx-Japan, then the quality of the summaries was evaluated quantitatively and the visualizations qualitatively. Quantitative evaluation using BLEU and ROUGE metrics showed that Gemini-2.5-pro achieved higher summary scores than GPT-4.1, and Japanese summaries scored higher than English ones. Conversely, English performed better than Japanese in temporal expression extraction using precision, recall, and F1 metrics, and the Gemini-2.5-pro model consistently outperformed GPT-4.1. In qualitative evaluation using the pairwise method, the timetable-based model was far superior with an overall win rate of 0.865 in Japanese and 0.969 in English compared to the baseline.
Medical-FLAVORS-AECC: Spanish Oncological Metaphors Dataset
Lucia Pitarch | Jordi Bernad | Sergio LUIS Ojeda Trueba | Alec Sánchez-Montero | Maxim Ionov | Emma Anglés-Herrero | Ángel Óscar Corona Beomont | Gemma Bel-Enguix
Lucia Pitarch | Jordi Bernad | Sergio LUIS Ojeda Trueba | Alec Sánchez-Montero | Maxim Ionov | Emma Anglés-Herrero | Ángel Óscar Corona Beomont | Gemma Bel-Enguix
Metaphors play a central role in cancer narratives, helping patients and practitioners articulate complex experiences and technical concepts. While cancer metaphors in English have been extensively studied, Spanish remains underexplored in this regard, despite its global importance and rich cultural variation. This paper presents a new dataset of Spanish cancer metaphors designed to address these gaps. The resource comprises over 80K annotated words drawn from diverse forum posts, with detailed documentation of lexical units, contextual versus basic meanings, and inter-annotator agreements. To construct the dataset, we adapted the Metaphor Identification Procedure (MIP) for Spanish medical discourse, proposing methodological refinements to challenges such as defining lexical units or domain-specific Basic Meaning labels.
A Synthetic Conversational Dataset for Type 2 Diabetes Management
Stergios Ntanavaras | Maaike de Boer | Piek T.J.M. Vossen
Stergios Ntanavaras | Maaike de Boer | Piek T.J.M. Vossen
Access to real patient-doctor conversations in the medical domain is often restricted due to privacy concerns, making it difficult to build robust conversational AI systems. To address this, we present a novel methodology for generating a high-quality synthetic dataset designed for conversational triple extraction in Type 2 Diabetes management. Using structured prompting with GPT-4, we generated 16 demographically and medically diverse diabetic personas, and 256 multi-turn conversations between these personas and a caretaker agent, simulating realistic and context-rich interactions. The conversations incorporate critical properties such as personalization, empathy, contextual awareness, and medically grounded advice, as validated through both LLM-based and human expert evaluations. These synthetic conversations are further annotated with Subject-Predicate-Object (SPO) labels at the token level, integrating both manual and LLM-automated methods, forming the foundation for downstream tasks like triple extraction. Our work demonstrates the feasibility of using generative AI to simulate healthcare conversations at scale, offering a solution for data-scarce domains.
Italian Medical Term Simplification: From Patient Information Leaflets to Simplified Language Resources
Maria Pia di Buono
Maria Pia di Buono
Terminological simplification in patient information leaflets (PILs) is implemented through a variety of linguistic strategies. Although these strategies help improve text comprehensibility, their overall impact remains limited. The complexity of PILs is still influenced by multiple factors, including frequent cross-referential terminology, the presence of subordinate clauses, lengthy sentences, and the use of domain-specific terms. This paper introduces I-MTS (Italian Medical Term Simplification), the first resource specifically developed for medical term simplification in Italian. I-MTS is designed to support research on lexical simplification and to facilitate the automatic adaptation of medical texts for non-expert audiences, thereby enhancing the readability and accessibility of health information in Italian.
Evaluating Professional Acceptability of LLM-Generated Systematic Review Summaries in Healthcare: Psychiatrists’ Perspectives
Paul Thompson | Artemis Boulogeorgou | Fotini Kaponi | Efstathia Soufleri | Sophia Ananiadou
Paul Thompson | Artemis Boulogeorgou | Fotini Kaponi | Efstathia Soufleri | Sophia Ananiadou
Cochrane systematic reviews evaluate the effectiveness and safety of medical interventions. Patients can benefit from clinicians’ integration of outcomes of these reviews into their daily practices. However, systematic reviews are usually long documents; even their abstracts can extend to 1000 words, making rapid appraisal challenging for busy health professionals. Large language models (LLMs) offer potential to further distil these abstracts. Nevertheless, generating high-quality, clinician-oriented summaries in this context is non-trivial. They must comprehensively cover the original abstract, while remaining accurate and professionally acceptable, i.e., retaining all clinically important details. To address this challenge, we have developed a novel dataset, PsycSumEval, comprising summaries generated by four different LLMs for 115 Cochrane abstracts concerning mental health. Psychiatrists evaluated each summary across nine content dimensions, assigning scores and providing free-text justifications that highlight inaccuracies and missing details. The corpus provides fine-grained insight into how psychiatrists assess professional acceptability of compressed medical evidence. Rather than treating agreement as a merely statistical endpoint, we capture structured expert judgments alongside their rationales, enabling transparent analysis of where professional norms are stable and where interpretive latitude persists. We contribute both a rigorous evaluation dataset and an explicit model of expert acceptability criteria for medical evidence summarisation.
Reasoning, Contrastive, and In-Context Strategies for Opioid Use Stage Detection on Social Media
Vinu Ekanayake | Ramakanth Kavuluru
Vinu Ekanayake | Ramakanth Kavuluru
The opioid epidemic has ravaged the US for the past two decades and is still a persistent threat. During the same time, the increasing use of social media has created a new avenue for people to share their journeys regarding opioid use. In this context, research in automatically determining opioid use stages (e.g., misuse, addiction, recovery) based on self disclosures in social media posts is gaining traction. In this paper, using a recent benchmark, we assess different supervised strategies for identifying self-disclosed opioid use stages from Reddit posts. We consider distilled reasoning traces from DeepSeek R1 (an open weights reasoning model), supervised contrastive learning (SCL), and few-shot in-context learning (ICL) with GPT-5 to conduct a variety of experiments with encoder and encoder-decoder models. We also conduct direct zero-shot (ZS) experiments with GPT 5 and GPT 5.2. Across different models and datasets, our strategies provide improvements in performance with some nuances that are too subtle to elaborate in the abstract. A surprising finding is that ZS results with GPT-5 are better than all supervised results, which ushers a new frontier for LLM-based classification of opioid use in social media posts. Our code is available for reuse and replication: https://github.com/bionlproc/Opioid-Stage.
FoodBench-QA: Overview of the Shared Task on Grounded Food and Nutrition Question Answering
Tome Eftimov | Ana Gjorgjevikj | Matej Martinc | Gjorgjina Cenikj | Sašo Džeroski | Barbara Koroušič Seljak
Tome Eftimov | Ana Gjorgjevikj | Matej Martinc | Gjorgjina Cenikj | Sašo Džeroski | Barbara Koroušič Seljak
We present the results of the FoodBench-QA 2026 shared task at the CL4Health workshop, collocated with LREC 2026. FoodBench-QA challenges systems to answer food and nutrition questions using evidence from food composition databases and food-related ontologies. The shared task comprises three main tasks: nutrient estimation from recipe ingredients, evaluated using EU Regulation 1169/2011 tolerance thresholds; FSA traffic-light classification for fat, salt, saturates, and sugars; and food named entity recognition and linking to three ontologies, namely Hansard Taxonomy, FoodOn, and SNOMED CT. We received submissions from five participating teams across all tasks. For nutrient estimation, the best system achieved accuracy rates of 93.57% for protein, 86.50% for sugars, 84.65% for fat, and 86.26% for saturates. For FSA traffic-light prediction, the best macro F1 scores ranged from 0.65 to 0.90 across different nutrient-color combinations. For named entity linking, the best systems achieved macro F1 scores between 60.71% and 80.89% for natural text and 87.75% and 95.75% for artificial NEL datasets, depending on the ontology.
Overview of the CT-DEB’26 Shared Task on Predicting Dosing Errors in Interventional Clinical Trials
Sohrab Ferdowsi | Félicien Hêche | Anthony Yazdani | Edward Choi | Sara Sansaloni-Pastor | Douglas Teodoro
Sohrab Ferdowsi | Félicien Hêche | Anthony Yazdani | Edward Choi | Sara Sansaloni-Pastor | Douglas Teodoro
Dosing errors represent an important source of medication-related risk in interventional clinical trials, potentially affecting both participant safety and the validity of study outcomes. Despite their importance, systematic methods for predicting dosing error risk from trial design information remain largely unexplored. To address this gap, we organized the Clinical Trial Dosing Error Benchmark 2026 (CT-DEB’26) shared task, hosted at the CL4Health workshop at LREC 2026. The task focuses on predicting the risk of dosing errors in interventional clinical trials using heterogeneous information extracted from ClinicalTrials.gov, including structured protocol metadata and long-form textual descriptions. The released benchmark dataset contains over 42,000 clinical trial records spanning multiple study phases and therapeutic areas, annotated with binary labels indicating a significant high rate of dosing errors. Participants were asked to develop ML models capable of estimating trial-level dosing error risk, evaluated primarily using the ROC-AUC metric under strong class imbalance. The shared task was conducted in two phases and attracted 15 submissions in the development stage and 4 submissions in the final evaluation phase. This paper provides an overview of the shared task, describing the dataset construction, evaluation protocol, and participating systems. In addition, we present a schema-aware CatBoost baseline that leverages structured trial metadata and simple textual statistics, achieving ROC-AUC scores of 0.8606 and 0.8624 on the Phase 1 and Phase 2 leaderboards, respectively. We further summarize the approaches proposed by participating teams, which explore both feature-engineering pipelines and transformer-based text representations. The results highlight the importance of structured trial design variables and hybrid modeling strategies combining tabular and textual information. Finally, we discuss limitations of the benchmark and outline future directions for applying natural language processing and ML to improve medication safety in clinical trial design.
Overview of the CRF 2026 Shared Task on Clinical Case Report Forms Filling
Pietro Ferrazzi | Soumitra Ghosh | Alberto Lavelli | Bernardo Magnini
Pietro Ferrazzi | Soumitra Ghosh | Alberto Lavelli | Bernardo Magnini
Case Report Forms (CRFs) are structured instruments widely used in clinical research to systematically collect patient information according to predefined protocols. In practice, CRFs are often manually completed by clinicians based on patients’ clinical reports, a process that is time-consuming and prone to inconsistencies. Despite their central role in medical studies, automatic population of CRFs from clinical narratives remains largely underexplored in the Natural Language Processing community, partly due to the scarcity of publicly available datasets. In this paper, we present the CRF Filling Shared Task, organized at the CL4Health Workshop at LREC 2026, which aims to advance research on automatic extraction of structured clinical information from unstructured patient notes. The task consists of assigning the correct value to a set of predefined CRF items given a clinical note. The target dataset is derived from a real-world CRF for dyspnea assessment, comprising 134 medical items with predefined value sets. The task is provided in two languages, Italian and English. We describe the dataset, the task formulation, and the evaluation framework, and discuss the participating systems and their results. By introducing this shared task, we aim to stimulate research on clinically applicable NLP systems for structured data extraction in healthcare.
Overview of the ArchEHR-QA 2026 Shared Task on Grounded Question Answering from Electronic Health Records
Sarvesh Soni | Dina Demner-Fushman
Sarvesh Soni | Dina Demner-Fushman
We present an overview of the ArchEHR-QA 2026 Shared Task on grounded question answering from electronic health records (EHRs), organized at the CL4Health Workshop at LREC 2026. The 2026 task decomposes grounded EHR question answering (QA) into four complementary subtasks: question interpretation, evidence identification, answer generation, and evidence alignment. We evaluated submitted systems for the text-generation subtasks (question interpretation and answer generation) using lexical, semantic, and grounding-sensitive automatic metrics, and for the evidence-centric subtasks (evidence identification and evidence alignment) using precision, recall, and F1. The shared task received 198 submitted runs from 43 teams, and 17 teams additionally provided system descriptions for this overview. The highest-ranked systems differed across subtasks, and gains over the organizer baseline were largest on the evidence-centric subtasks. Across submitted system descriptions, prompt-based large language model (LLM) pipelines were dominant, whereas task-specific fine-tuning was rare; retrieval, self-consistency, and ensembling were especially common in the strongest evidence-centric systems. In this paper, we describe the task design, data, evaluation protocol, baselines, participation, official results, and common system characteristics, and discuss implications for developing clinically faithful and transparent QA systems.
Structured Radiology Intelligence: Extracting Structured Data from MRI Reports Using LLMs
Sushvin Marimuthu | Parameswari Krishnamurthy | Dipti Misra Sharma | Goldwin H | Anu Eapen | Betty Simon | Anuradha Chandramohan
Sushvin Marimuthu | Parameswari Krishnamurthy | Dipti Misra Sharma | Goldwin H | Anu Eapen | Betty Simon | Anuradha Chandramohan
This study presents efforts focused on extracting and structuring doctor notes, specifically Magnetic Resonance Imaging (MRI) reports, into a standardized format using large language models (LLMs). We introduce a novel benchmark dataset comprising of 55 clinically relevant variables given by doctors, making it the first of its kind in the automated processing of unstructured medical texts. The annotations to the dataset were generated using a systematic prompt-tuning approach that was manually validated. It was then evaluated across three experimental stages: baseline, intermediate, and fine-tuned. Each stage assessed the impact of different prompt strategies on the performance of various LLMs (LLaMA, Qwen, and DeepSeek). Among the models tested, LLaMA 3.1 8B Instruct consistently achieved the highest composite Score in both the intermediate and final phases, resulting in an 18.42% improvement in performance.
Useful to Whom? A Persona-Driven Evaluation of Knowledge-Adapted Health Question Reformulation via LLM Simulation
Jooyeon Lee | Luan Huy Pham | Özlem Uzuner
Jooyeon Lee | Luan Huy Pham | Özlem Uzuner
Automatic metrics such as F1 and BERTScore are often insufficient for evaluating user-centric generative tasks like Consumer Health Question (CHQ) reformulation. A high F1-score may not correlate with user satisfaction, especially when the user’s knowledge level (UKL) dictates their needs. We propose a robust, Persona-Driven Evaluation Framework (PDEF), grounded in cognitive science and health literacy literature, to measure persona-specific utility. This framework assesses reformulations from the perspectives of a ‘Layperson’ (requiring foundational context) and an ‘Expert’ (requiring efficient, precise answers). We apply this framework to a set of reformulated questions generated by LLMs, and test the robustness of our evaluation by using three state-of-the-art LLMs (GPT-4o, Llama 3.3, and Mistral Large) as the evaluators. Our results reveal a significant disconnect between automatic metrics and user-perceived quality: the model with the highest F1-score (0.6134) was consistently outperformed in user preference by a Pipelined model, with experts preferring the latter by a statistically significant margin (p < 0.001). Furthermore, our persona-driven ablation analysis provides robust evidence that specific architectural components, specifically UKL inference and Entailment logic, are linked to significant gains in persona-driven utility for Layperson cohorts. This work demonstrates the critical need for user-centric evaluation and shows that its findings are generalizable across different LLM architectures.
MedGore: An Approach and a Dataset for Identification of Sensitive Medical Images
Soumya Gayen | Rory Mulcahey | Russell Loane | Dina Demner-Fushman | Deepak Gupta
Soumya Gayen | Rory Mulcahey | Russell Loane | Dina Demner-Fushman | Deepak Gupta
Medical images are invaluable in illustrating health issues for the patients. While biomedical publications are a good source of such images, some of the images are not appropriate for the patient viewing without a warning. To enable development of automated tools for selection of patient-safe images and generation of warnings, we created a dataset MedGore of over 78,000 sensitive medical images and 183,000 non-sensitive images published in the biomedical literature. The sensitive content includes gore, severe disease, nudity, surgical openings, internal organs, and other medical images of this nature. The set of the manually identified seed 300 images was expanded using a combination of human curation and a nearest neighbor clustering algorithm. The quality of the automatically labeled images was evaluated manually, yielding a total of more than 4,000 doubly-manually annotated images. The automatically labeled images proved to approach the utility of the manually labeled images for training the models in our experiments that validated the dataset in the task of labeling unseen images using the image features, the figure captions or both.
Diagnostic Reasoning with Large Language Models for a Rare Disease: Case Study of Primary Ciliary Dyskinesia
Swati Rajwal | Mary Ellen M. Fain | Lokesh Guglani | Abeed Sarker
Swati Rajwal | Mary Ellen M. Fain | Lokesh Guglani | Abeed Sarker
Primary ciliary dyskinesia (PCD) is a rare pediatric lung disease that is frequently underdiagnosed due to nonspecific early symptoms and limited clinical exposure. We investigate whether large language models (LLMs) can support early diagnostic reasoning using real-world pediatric pulmonology notes written before the final diagnosis. We curated 58 de-identified first-visit notes (28 confirmed PCD, 30 controls) and evaluated five open-source LLMs using a standardized zero-shot prompt to produce structured outputs, including PCD evaluation recommendations, justifications, and suggested tests. Quantitative performance was assessed against expert-validated labels using sensitivity, specificity, and accuracy, and a clinician qualitatively reviewed all explanations and testing recommendations for clinical soundness. Sensitivity ranged from 0.48 to 1.00 and specificity from 0.10 to 0.48 (excluding uncertain outputs), with a best accuracy of 0.75. A majority-vote ensemble of five open-source LLMs achieved perfect sensitivity (1.00) with accuracy of 0.73. While models often identified clinically relevant signals in unstructured notes, explanations and testing recommendations were frequently only partially sound. These findings suggest LLMs may serve as cautious early screening aids for rare disease suspicion, but not as standalone diagnostic tools. This work further highlights the need for larger, multi-site evaluation on longitudinal clinical text.
Beyond One-Size-Fits-All: Multi-Agent Refinement Framework for Persona-Based Biomedical Summarization
Rohan Charudatt Salvi | Chirag Chawla | Md. Shad Akhtar | Shweta Yadav
Rohan Charudatt Salvi | Chirag Chawla | Md. Shad Akhtar | Shweta Yadav
Lay summarization aims to make biomedical research accessible to non-experts, but most approaches assume a uniform audience, overlooking variation in medical literacy and information needs. We present MAPS (Multi-Agent Persona-based Summarization), a framework that generates persona-specific summaries through iterative cross-agent feedback. Human evaluation shows MAPS improves quality over single-agent baselines, while automatic metrics fail to capture these gains. LLM-based judges also exhibit limited sensitivity, assigning inflated scores and misdetecting errors. These findings highlight the need for improved evaluation methods for persona-based summarization.
Automated Detection of Dosing Errors in Clinical Trial Narratives: A Multi-Modal Feature Engineering Approach with LightGBM
Mohammad AL-Smadi
Mohammad AL-Smadi
Clinical trials require strict adherence to medication protocols, yet dosing errors remain a persistent challenge affecting patient safety and trial integrity. We present an automated system for detecting dosing errors in unstructured clinical trial narratives using gradient boosting with comprehensive multi-modal feature engineering. Our approach combines 3,451 features spanning traditional NLP (TF-IDF, character n-grams), dense semantic embeddings (all-MiniLM-L6-v2), domain-specific medical patterns, and transformer-based scores (BiomedBERT, DeBERTa-v3), used to train a LightGBM model. Features are extracted from nine complementary text fields (median 5,400 characters per sample) ensuring complete coverage across all 42,112 clinical trial narratives. On the CT-DEB benchmark dataset with severe class imbalance (4.9% positive rate), we achieve 0.8725 test ROC-AUC through 5-fold ensemble averaging (cross-validation: 0.8833 ± 0.0091 AUC). Systematic ablation studies reveal that removing sentence embeddings causes the largest performance degradation (2.39%), demonstrating their critical role despite contributing only 37.07% of total feature importance. Feature efficiency analysis demonstrates that selecting the top 500-1000 features yields optimal performance (0.886-0.887 AUC), outperforming the full 3,451-feature set (0.879 AUC) through effective noise reduction. Our findings highlight the importance of feature selection as a regularization technique and demonstrate that sparse lexical features remain complementary to dense representations for specialized clinical text classification under severe class imbalance.
CaresAI at CT-DEB’26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models
Leon Hamnett | Favour Igwezeke | Joseph Itopa Abubakar | Mary Adetutu Adewunmi
Leon Hamnett | Favour Igwezeke | Joseph Itopa Abubakar | Mary Adetutu Adewunmi
Medication errors, particularly dosing errors in clinical trials (CT), can lead to patient harm, adverse drug events and worse patient outcomes. Dosing errors are preventable, and early identification can improve trial integrity and mitigate subsequent clinical and financial burden. This study aims to detect dosing errors within CT protocols by evaluating text representations of trial information using transformer-based language models trained on biomedical corpora. CT textual data was encoded using several models, including ClinicalBERT, PubMedBERT, BioBERT, and MedCPT, and integrated with categorical features. These text embeddings were used as input to classical machine learning models and neural network architectures within an experimental framework. Performance was primarily assessed using ROC-AUC with respect to predicting dosage error. Under a logistic regression baseline, BioBERT consistently outperformed alternative encoders, achieving an ROC-AUC of 0.794, a 3.95 % improvement over the ClinicalBERT baseline. Combining multiple embeddings did not yield improvements, indicating that domain alignment outweighs representational stacking. Gradient boosting models, support vector classifiers, logistic regression, and residual neural networks achieved the strongest performance for predicting dosage error, achieving ROC-AUCs: 0.821 to 0.853. Overall, the integration of domain-specific transformer embeddings with structured metadata enables discrimination of trials meeting a predefined elevated dosing error risk criterion, advancing safety monitoring and supporting informed regulatory decision-making.
CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation
Wei-Chun Chen | Yu-Xuan Chen | I-Fang Chung | Ying-Jia Lin
Wei-Chun Chen | Yu-Xuan Chen | I-Fang Chung | Ying-Jia Lin
Accurate nutrient estimation from unstructured recipe text is an important yet challenging problem in dietary monitoring, due to ambiguous ingredient terminology and highly variable quantity expressions. We systematically evaluate models spanning a wide range of representational capacity, from lexical matching methods (TF-IDF with Ridge Regression), to deep semantic encoders (DeBERTa-v3), to generative reasoning with large language models (LLMs). Under the strict tolerance criteria defined by EU Regulation 1169/2011, our empirical results reveal a clear trade-off between predictive accuracy and computational efficiency. The TF-IDF baseline achieves moderate nutrient estimation performance with near-instantaneous inference, whereas the DeBERTa-v3 encoder performs poorly under task-specific data scarcity. In contrast, few-shot LLM inference (e.g., Gemma-3-27B) and a hybrid LLM refinement pipeline (TF-IDF combined with Gemini 2.5 Flash) deliver higher accuracy across all nutrient categories. These improvements likely arise from the ability of LLMs to leverage pre-trained world knowledge to resolve ambiguous terminology and normalize non-standard units, which remain difficult for purely lexical approaches. However, these gains come at the cost of substantially higher inference latency, highlighting a practical deployment trade-off between real-time efficiency and nutritional precision in dietary monitoring systems.
Food and nutrition question answering involves resolving ambiguous ingredient terminology and diverse household measurement expressions, and converting them into representations compatible with nutrient databases. In FoodBench-QA, recipe-level nutrient estimation requires consistent handling of heterogeneous and imprecise measurement descriptions. We propose FoodComponentProfiler (FCProfiler), a deterministic pipeline that treats nutrient estimation as a structured measurement resolution problem. The pipeline is composed of multiple stages, including parsing, normalization, unit canonicalization, gram conversion, and nutrient estimation, with each step designed to remain transparent and traceable. Unit canonicalization combines rule-based standards with data-driven unit expansion from large-scale recipe corpora, enabling broader coverage of real-world measurement variations. Gram conversion grounds quantities in ingredient-specific portion information, enabling accurate and traceable mass computation. Experimental results show that accurate nutrient estimation mainly depends on reliable unit normalization and ingredient-specific measurement conversion. Additionally, FCProfiler achieves performance comparable to FoodyLLM, demonstrating that explicit measurement grounding serves as an effective alternative to implicit reasoning. The proposed methodology preserves interpretability while maintaining strong performance in food and nutrition question answering.
DocUA at CRF Filling 2026: LLM StructCore — Schema-Guided Reasoning Condensation and Deterministic Compilation
Serhii Zabolotnii
Serhii Zabolotnii
Automatically filling Case Report Forms (CRFs) from clinical notes is challenging due to noisy language, strict output contracts, and the high cost of false positives. We describe our CL4Health 2026 submission for Dyspnea CRF filling (134 items) using a contract-driven two-stage design grounded in Schema-Guided Reasoning (SGR) (Abdullin, 2025). The key task property is extreme sparsity: the majority of fields are unknown, and official scoring penalizes both empty values and unsupported predictions. We shift from a single-step “LLM predicts 134 fields” approach to a decomposition where (i) Stage 1 produces a stable SGR-style JSON summary with exactly 9 domain keys, and (ii) Stage 2 is a fully deterministic, 0-LLM compiler that parses the Stage 1 summary, canonicalizes item names (optionally using a UMLS alias map with 134/134 coverage), normalizes predictions to the official controlled vocabulary (13 categories), applies evidence-gated false-positive filters, and expands the output into the required 134-item format. On the dev80 split, the best teacher configuration (Mistral Large 3 Stage 1 → Stage 2 deterministic) achieves macro-F1 0.6543 (EN) and 0.6905 (IT); on the hidden test200, the submitted English variant scores 0.63 on Codabench. The pipeline is language-agnostic: Italian results match or exceed English with no language-specific engineering.
GREYC at CRF Filling 2026: Rewrite Before You Extract - Rewriting Clinical Notes for Automated CRF
Jesus Lovon-Melgarejo | Jérémie Pantin | Gaël Dias
Jesus Lovon-Melgarejo | Jérémie Pantin | Gaël Dias
This paper describes the system we submitted to the CRF:filling 2026 shared task. We propose a modular, LLM-based framework including an LLM as rewriter, which enhances the original clinical note from the perspective of each target CRF item; an LLM extractor, which retrieves the relevant value using a k-shot prompting strategy; and an LLM as a judge, which determines whether the clinical note contains evidence to support a given answer, defaulting to ’unknown’ otherwise. We evaluated our system on the English portion of the dataset; our complete framework achieves a macro-F1 of 0.64 on the development set. Our analysis reveals that while the rewriting step effectively generates correct factual information, it also increases false positives. The judge component mitigates this by adopting a conservative prediction strategy that substantially reduces false positives at the cost of a moderate reduction in true positives, yielding higher precision and better alignment with the shared task metric. On the test set, a light version of our system ranked 21 out of 32 public submissions, achieving a macro-F1 of 0.45.
Innov8rs at CRF Filling 2026: An Iterative Multi-LLM Ensemble Pipeline with Dynamic Few-Shot Retrieval and Data-Driven Precision Filtering
Samminga Sainath Rao | Sumit Mishra | Chanchal Suman
Samminga Sainath Rao | Sumit Mishra | Chanchal Suman
In this paper, we present the technical report on the CL4Health 2026 Shared Task on Case Report Form (CRF) filling for our team Innov8rs. The paper explains the complete development of our system for the CL4Health 2026 Shared Task. We describe every phase of our system – from initial catastrophic failures with small models producing over 4,800 false positives, through prompt engineering breakthroughs, to our final multi-LLM ensemble combining Gemini 2.5 Flash and Llama 3.3 70B with dynamic TF-IDF-based few-shot retrieval. The main contribution of this work is a data-driven precision filter that suppresses predictions for CRF items with historically high false-positive rates. This single intervention reduced false positives from 816 to 171 on the English development set, boosting macro-F1 from 0.541 to 0.703. We document the engineering challenges of multi-API-key rotation across 11 Google API keys and 2 Groq keys, the design of four distinct ensemble strategies, and the critical analysis of why development-calibrated filters suffered from distribution shift on test data (final test F1: 0.47).
Aurum at CRF Filling 2026: Modular DSPy Extractors with Qwen3-Max for Multilingual CRF Filling
Vinay Babu Ulli | Jyoti Kumari | Anindita Mondal
Vinay Babu Ulli | Jyoti Kumari | Anindita Mondal
This paper describes the submission by Team Aurum to the CL4Health @ LREC 2026 Shared Task on Case Report Form (CRF) Filling from dyspnea patient clinical notes. Extracting 134 structured clinical fields using a single Large Language Model (LLM) call often leads to schema-following errors, hallucination, and poor attention over complex instructions. To address this, we propose a modular extraction pipeline built with DSPy, which decomposes the 134 CRF fields into 14 specialized, domain-specific extractors (e.g., Medical History, Lab Values, Acute Diagnoses). We conducted extensive experiments across multiple multilingual LLMs, including Llama4 Maverik, GPT-4o, GPT-4o Mini, DeepSeek-V3, Gemma-3-12B-Instruct, and Qwen-series models. Among these, Qwen3-Max (Thinking) with our optimized v2 prompts achieved the best performance on the development set with a Macro-F1 of 0.70, outperforming other evaluated models such as GPT-4o (0.68) and DeepSeek-V3 (0.66). Prompt optimization resulted in measurable gains, improving Qwen3-Max performance from 0.67 to 0.70. Using this configuration, our pipeline achieved an official Codabench Test Macro-F1 score of 0.68 in English and 0.67 in Italian, securing the 1st place ranking overall in the shared task.
sebis at CRF Filling 2026: A Two-Stage Local LLM Pipeline for Medical CRF Filling
Katharina Sommer | Tristan Till | Florian Matthes
Katharina Sommer | Tristan Till | Florian Matthes
The extraction of structured clinical information from unstructured EHR notes is a persistent bottleneck in healthcare informatics. While large language models (LLMs) offer high performance, their deployment in clinical settings is hindered by privacy risks, inference costs, and the tendency to hallucinate beyond textual evidence. We address these challenges for the CL4Health 2026 Case Report Form (CRF) filling task by proposing a fully local, domain-adapted pipeline using the MedGemma-27B model. Our two-stage architecture, which separates binary presence classification from value extraction, enforces strict adherence to textual evidence and ensures deterministic outputs for negated, uncertain, or unknown states. By leveraging item-specific, few-shot in-context learning without external API calls or fine-tuning, our approach achieves a macro-F1 score of 0.55 on the official English test track. This result secures second place among all locally-hosted, open-source submissions. Our work demonstrates that privacy-preserving, on-premise LLM pipelines can achieve near-competitive performance with proprietary frontier models, providing a practical, data-sovereign framework for clinical NLP.
Polimi at CRF Filling 2026: Prompt-Based Information Extraction from Italian Clinical Notes
Vittorio Torri | Francesca Ieva
Vittorio Torri | Francesca Ieva
In this paper we describe the system developed by the Polimi team for the CRF Filling Shared Task 2026, which focuses on extracting structured variables from clinical notes. The task is challenging due to scarce annotations, heterogeneous clinical language, and the sparsity of the 134 items to be extracted. Our approach relies on prompt-based information extraction using locally deployed open-weight Large Language Models (LLMs). We focused on the Italian subset of the dataset. The pipeline performs zero-shot extraction using task-specific prompts augmented with a glossary of abbreviations derived from unlabeled notes. To improve reliability and reduce hallucinations, the extraction schema is decomposed into multiple prompts targeting groups of variables, whose outputs are merged and refined through deterministic post-processing rules to normalize values and recover missing labels. During development we explored verification stages based on LLM-based prediction validation and synthetic example generation, but these strategies did not improve performance and were not included in the final system. On the development set, the best configuration based on Mistral Small 3.2 24B Instruct achieved an F1-score of 67.51%. On the official test set, our system ranked third overall and second among systems evaluated on the Italian subset, achieving an F1-score of 63%.
Cohere Labs Community at FoodBench-QA 2026: The Cake Makes the Ingredients
Ravi Ranjan | Roshan Santhosh | Lucien Carroll
Ravi Ranjan | Roshan Santhosh | Lucien Carroll
People intuitively ask natural language dialogue systems for advice on nutrition and dietary guidelines, but systems based on prompted text generation are susceptible to fabricating details, which could be hazardous to non-specialist users. The FoodBench-QA shared task grounds answers in knowledge bases with linked ontologies, in order to evaluate and mitigate fabrication of nutrition information. Our system treats nutrient estimation and entity linking not as a generative problem (predicting numbers from scratch), but as a retrieval problem. We operate on the hypothesis that for structured data like food composition, finding a “real” recipe that is 95% similar is more likely to approximate the correct values than letting the language model fabricate values from sparse context. Our system performed well on food safety labeling from recipe ingredients alone, and it did not benefit from the additional information of recipe titles. In the NER and NEL tasks, our system handled the recipe-focused FCD corpus well, but suffered from poor recall on scientific abstracts and the artificial dataset. These results show the importance of basing information retrieval and question answering in data that is well-matched to the target data.
HiTZ-IXA at ArchEHR-QA 2026: Evidence Alignment Through Self-Consistency and Prompt Curation in Memory-Constrained Environments
Xabier Irastortza-Urbieta | Maite Oronoz | Alicia Pérez
Xabier Irastortza-Urbieta | Maite Oronoz | Alicia Pérez
The development of question-answering systems capable of grounding their answers in Electronic Health Records could provide patients with faithful assistance while reducing the clinical workload. The ArchEHR-QA 2026 Shared Task was organized to advance progress in this context. In this paper, we present our strategies for addressing this shared task, which are focused primarily on evidence alignment and, to a lesser extent, on evidence identification. Our approaches rely exclusively on open-source models with up to 8 billion parameters, aiming to produce systems suitable for environments with memory constraints. We experimented with methods based on embedding models, prompt curation, self-consistency, and combination of LLMs. We concluded that prompt curation together with an effective post-processing step was crucial for creating stable systems, while self-consistency yielded considerable gains in performance. The results of our approaches suggest that small LLMs can substantially improve their accuracy in the evidence alignment task via simple and affordable techniques.
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
Richard A. A. Jonker | Alexander Christiansen | Alexandros Maniatis | Rúben Garrido | Rogério Braunschweiger de Freitas Lima | Roman Jurowetzki | Sérgio Matos
Richard A. A. Jonker | Alexander Christiansen | Alexandros Maniatis | Rúben Garrido | Rogério Braunschweiger de Freitas Lima | Roman Jurowetzki | Sérgio Matos
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without weight updates. We evaluate several state-of-the-art proprietary models and locally deployable open-source alternatives using various prompt engineering strategies, including task decomposition, Chain-of-Thought, and in-context learning. Furthermore, we explore majority voting and LLM-as-a-judge ensembling techniques to maximize predictive robustness. Our results demonstrate that while proprietary models exhibit strong resilience to prompt variations, domain-adapted open-source models (such as MedGemma 3 27B) achieve highly competitive performance when paired with the right prompt. Overall, our prompt-based approach proved highly effective, securing 1st place in Subtask 4 (evidence citation alignment) and 3rd place in Subtask 3 (patient-friendly answer generation). All code, results, and prompts are available on our GitHub repository: https://github.com/bioinformatics-ua/ArchEHR-QA-2026.
WisPerMed at ArchEHR-QA 2026: Retrieval-Augmented Prompting for Grounded EHR Question Answering
Jan-Henning Büns | Tabea Margareta Grace Pakull | Hendrik Damm | Bohao Chu | Christoph M. Friedrich | Felix Nensa | Elisabeth Livingstone | Peter A. Horn | Norbert Fuhr
Jan-Henning Büns | Tabea Margareta Grace Pakull | Hendrik Damm | Bohao Chu | Christoph M. Friedrich | Felix Nensa | Elisabeth Livingstone | Peter A. Horn | Norbert Fuhr
ArchEHR-QA is a grounded question-answering (QA) task for electronic health records (EHRs) comprising four subtasks: (1) question rewriting, (2) evidence identification, (3) grounded answer generation, and (4) answer-evidence alignment. In this work, we present a modular pipeline centered on retrieval-augmented generation (RAG). For Subtask 1, RAG few-shot prompting outperformed both PEFT and prompt-only baselines on the development set; however, Claude few-shot proved substantially more robust on the test set, ranking 6th out of 13 participating teams (score: 26.94). For Subtask 2, a union ensemble of open-weight LLMs (GPT-OSS-120B and Qwen3-30B-A3B) achieved a 56.7 micro-F1, rivaling the proprietary Claude Opus 4.6 while demonstrating higher recall (53.6). For Subtask 3, our RAG few-shot approach using Claude Opus 4.5 achieved the 1st place out of 13 participating teams (score: 36.33). Finally, for Subtask 4, a zero-shot Claude Opus 4.6 configuration ranked 2nd out of 16 participating teams (score: 81.3).
sebis at ArchEHR-QA 2026: How Much Can You Do Locally? Evaluating Grounded EHR QA on a Single Notebook
Ibrahim Ebrar Yurt | Fabian Tobias Karl | Tejaswi Choppa | Florian Matthes
Ibrahim Ebrar Yurt | Fabian Tobias Karl | Tejaswi Choppa | Florian Matthes
Clinical question answering over electronic health records (EHRs) can help clinicians and patients access relevant medical information more efficiently. However, many recent approaches rely on large cloud-based models, which are difficult to deploy in clinical environments due to privacy constraints and computational requirements. In this work, we investigate how far grounded EHR question answering can be pushed when restricted to a single notebook. We participate in all four subtasks of the ArchEHR-QA 2026 shared task and evaluate several approaches designed to run on commodity hardware. All experiments are conducted locally without external APIs or cloud infrastructure. Our results show that such systems can achieve competitive performance on the shared task leaderboards. In particular, our submissions perform above average in two subtasks, and we observe that smaller models can approach the performance of much larger systems when properly configured. These findings suggest that privacy-preserving EHR QA systems running fully locally are feasible with current models and commodity hardware. The source code is available at https://github.com/ibrahimey/ArchEHR-QA-2026.
MedEvi-NS at ArchEHR-QA 2026: Using Clinical Reasoning Principles to Improve Zero-shot Capabilities of Large Language Models in Evidence Alignment
Mengxuan Sun | Nicolay Rusnachenko
Mengxuan Sun | Nicolay Rusnachenko
The ArchEHR-QA shared task focuses on grounded question answering using patient EHR data. For the given clinical interpretation of the patient question, note excerpt (E) and answer text (A), subtask 4 (evidence alignment) aims to cite supporting sentences from E for each sentence in A. In this paper, we propose a prompt-engineering methodology that features clinical-reasoning principles in related alignment. We adopt this methodology for GPT-5.2 in zero-shot learning mode. According to our experiments on ArchEHR-QA, incorporating clinical reasoning principles into the prompt improves F 1overall by +2.0%. Our final submission resulted in 77.4% by F 1overall, which positions us at 10th out of 16 teams. Our code is publicly available: https://github.com/nicolay-r/ArchEHR-QA-2026-Task-4-MedEvi-NS
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
Mohammad Arvan | Hossein Haeri | Natalie Parde | Rebecca Feinstein
Mohammad Arvan | Hossein Haeri | Natalie Parde | Rebecca Feinstein
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer. For Subtask 4, we apply self-consistency voting over five independent model calls, retaining links by vote threshold. Our pipeline ranked third on evidence identification (Strict Micro F1 62.90), ninth on answer generation (Overall 31.90), and fifth on answer-evidence alignment (F1 79.81). A post-hoc linguistic analysis of 45 stylistic features reveals that model outputs remain 3.2 Flesch-Kincaid grade levels harder to read than clinician-authored references despite matching their word and sentence counts, suggesting readability warrants explicit optimization in clinical NLP systems. Code and prompts are available at https://github.com/mo-arvan/archehr-qa-2026-uic-aihealth4all.
tt501 at ArchEHR-QA 2026: Few-Shot Prompting with Retrieval-Augmented Generation for Grounded Clinical EHR Question Answering
Tai Tan Tran
Tai Tan Tran
We present the ArchEHR-QA 2026 shared task system of team tt501, which addresses evidence identification (Subtask 2), answer generation (Subtask 3), and evidence alignment (Subtask 4) from electronic health record notes. Our approach relies entirely on prompt engineering with xAI’s Grok models, without any task-specific fine-tuning or external knowledge. For evidence identification we compare a hybrid BM25 plus large language model (LLM) reranker with a full-context chain-of-thought ensemble and refinement step, finding that full-note reasoning yields higher recall and F1. For answer generation we implement a retrieval-augmented generation pipeline that conditions on predicted evidence sentences and few-shot examples, improving lexical and semantic faithfulness over a zero-shot baseline. For evidence alignment we design a recall-oriented few-shot prompt enriched with explicit rationales that teach the model how to map each answer sentence back to its supporting note sentences. We report official shared task results and analyse the impact of these design choices across the three subtasks.
HealthNLP_Retrievers at ArchEHR-QA 2026: Cascaded LLM Pipeline for Grounded Clinical Question Answering
Md Biplob Hosen | Md Alomgeer Hussein | Md Akmol Masud | Omar Faruque | Tera L Reynolds | Lujie Karen Chen
Md Biplob Hosen | Md Alomgeer Hussein | Md Akmol Masud | Omar Faruque | Tera L Reynolds | Lujie Karen Chen
Patient portals now give individuals direct access to their electronic health records (EHRs), yet access alone does not ensure patients understand or act on the complex clinical information contained in these records. The ArchEHR-QA 2026 shared task addresses this challenge by focusing on grounded question answering over EHRs, and this paper presents the system developed by the HealthNLP_Retrievers team for this task. The proposed approach uses a multi stage cascaded pipeline powered by the Gemini 2.5 pro large language model to interpret patient authored questions and retrieve relevant evidence from lengthy clinical notes. Our architecture comprises four integrated modules. (1) A few shot query reformulation unit which summarizes verbose patient queries; (2) A heuristic based evidence scorer which ranks clinical sentences to prioritize recall; (3) A grounded response generator which synthesizes professional caliber answers restricted strictly to identified evidence; (4) A high precision many to many alignment framework which links generated answers to supporting clinical sentences. This cascaded approach achieved highly competitive results. Across the individual tracks, the system ranked 1st in question interpretation (Subtask 1), 5th in answer generation, 7th in evidence identification, and 9th in answer evidence alignment. These results show that integrating large language models within a structured multi stage pipeline improves grounding, precision, and the professional quality of patient oriented health communication. To support reproducibility, our source code is publicly available in our GitHub repository.
Yale-DM-Lab at ArchEHR-QA 2026: Deterministic Grounding and Multi-Pass Evidence Alignment for EHR Question Answering
Elyas Irankhah | Samah Fodeh
Elyas Irankhah | Samah Fodeh
We describe the Yale-DM-Lab system for the ArchEHR-QA 2026 shared task. The task studies patient-authored questions about hospitalization records and contains four subtasks (ST): clinician-interpreted question reformulation, evidence sentence identification, answer generation, and evidence–answer alignment. ST1 uses a dual-model pipeline with Claude Sonnet 4 and GPT-4o to reformulate patient questions into clinician-interpreted questions. ST2–ST4 rely on Azure-hosted model ensembles (o3, GPT-5.2, GPT-5.1, and DeepSeek-R1) combined with few-shot prompting and voting strategies. Our experiments show three main findings. First, model diversity and ensemble voting consistently improve performance compared to single-model baselines. Second, the full clinician answer paragraph is provided as additional prompt context for evidence alignment. Third, results on the development set show that alignment accuracy is mainly limited by reasoning. The best scores on the development set reach 88.81 micro F1 on ST4, 65.72 macro F1 on ST2, 34.01 on ST3, and 33.05 on ST1.
Razreshili at ArchEHR-QA 2026: Evidence Alignment via LLM Prompting and Cross-Encoder Fine-tuning
Arina Zemchyk
Arina Zemchyk
We describe our system for Subtask 4 (Evidence Alignment) of the ArchEHR-QA 2026 shared task, which requires aligning each sentence of a clinician-authored answer to the supporting sentence(s) in a clinical note excerpt derived from MIMIC. The task is challenging due to many-to-many alignment structure, answer sentences with no note support, and the semantic gap between clinical note language and answer paraphrases. We explore two approaches: few-shot chain-of-thought prompting with Qwen2.5-7B-Instruct and LoRA fine-tuning of a cross-encoder with combined InfoNCE and BCE loss. Our best system achieves a micro F1 of 67.93 on the test set.
Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs
Abrar Majeedi | Viswanatha Reddy Gajjala | Sai Prasanna Teja Reddy Bogireddy | Siddhant Rai
Abrar Majeedi | Viswanatha Reddy Gajjala | Sai Prasanna Teja Reddy Bogireddy | Siddhant Rai
Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit grounding of answers in clinical notes. In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, which comprises of four subtasks: question interpretation, evidence identification, answer generation, and evidence alignment. Our approach decouples the task into independent, modular stages and employs DSPy’s MIPROv2 optimizer to automatically discover high-performing prompts, jointly tuning instructions and few-shot demonstrations for each stage. Within every stage, self-consistency voting over multiple stochastic inference runs suppresses spurious errors and improves reliability, while stage-specific verification mechanisms (e.g., self-reflection and chain-of-verification for alignment) further refine output quality. Among all teams that participated in all four subtasks, our method ranks second overall (mean rank 4.00), placing 4th, 1st, 4th, and 7th on Subtasks 1–4, respectively. These results demonstrate that systematic, per-stage prompt optimization combined with self-consistency mechanisms is a cost-effective alternative to model fine-tuning for multi-faceted clinical QA.
OptiMed at ArchEHR-QA 2026: GEPA Prompt Optimization and Multi-Agent Majority Voting for EHR-Grounded Question Answering
Feras AlMannaa | Talia Tseriotou | Maria Liakata
Feras AlMannaa | Talia Tseriotou | Maria Liakata
Despite the demonstrated promise of Large Language Models in medical question answering, existing work largely addresses closed-form, exam-style tasks and overlooks complex open-ended questions requiring reasoning over noisy, long clinical documents. In this work, we present our system, OptiMed, submitted to the ArchEHR-QA 2026 shared task on grounded clinical question answering over EHR notes. We combine GEPA, an evolutionary prompt optimization framework, with multi-agent majority voting across five diverse LLMs and a structured clinical abstraction strategy for question interpretation. OptiMed ranked 1st overall among teams completing all four subtasks with an average score of 52.0, achieving top AlignScore in both Question Interpretation and Answer Generation, reflecting strong factual grounding. GEPA optimization proved effective for structured tasks with sufficient development data, but failed to generalize on complex generative tasks under very limited number of supervisions. Multi-agent majority voting consistently lifted performance in evidence-oriented subtasks. Prompt analysis attributes GEPA’s gains to role prompting and procedural decomposition and failures to over-specification under limited supervision.
TAMU-NLP at ArchEHR-QA 2026: Grounded Clinical QA with Evidence Identification and Intent-Aware Answer Generation
Xinqi Su | Rongrong Wang | Sunyang Fu | Hongfang Liu | Ruihong Huang
Xinqi Su | Rongrong Wang | Sunyang Fu | Hongfang Liu | Ruihong Huang
Electronic Health Records (EHRs) contain rich clinical information and provide an important data source for medical question answering. However, generating reliable answers grounded in patient-specific clinical evidence remains challenging. In this work, we participate in the ArchEHR-QA 2026 shared task and focus on Subtask 2 (Evidence Identification) and Subtask 3 (Answer Generation). For evidence identification, we explore both traditional learning-to-rank methods and large language models (LLMs), and propose a two-stage LLM framework that improves prediction stability through few-shot prompting and self-reflection reasoning. For answer generation, we design an intent-aware few-shot prompting framework to generate concise answers grounded in clinical evidence. Experimental results show that our approach achieves strong performance despite limited training data. On the official leaderboard, our system ranks 5th in Subtask 2 and 2nd in Subtask 3. These results demonstrate that combining evidence-driven reasoning with the generative capabilities of LLMs is an effective approach for EHR-based clinical question answering.
GigitAI at ArchEHR-QA 2026: Prompting Strategies and Constitutional AI for Clinical Question Answering
Saran Krishnasamy | Inez Wihardjo
Saran Krishnasamy | Inez Wihardjo
Answering patient questions from electronic health records requires identifying relevant evidence in lengthy clinical notes and generating faithful, patient-friendly answers. We present a systematic study of LLM prompting strategies for both tasks, evaluating 21 evidence identification methods and 13 answer generation methods across 7 language models. For evidence identification, we find that LLM prompting outperforms traditional retrieval (BM25, SBERT, BioLinkBERT) by 19 F1 points, and that prompt framing alone controls precision–recall trade-offs: inclusive framing achieves 90% recall on dev while balanced framing reaches 67% precision. For answer generation, we introduce a Constitutional AI pipeline that critiques and revises answers against five clinical faithfulness principles, improving BLEU and ROUGE over the constrained baseline. Our analysis reveals that chain-of-thought effectiveness is strongly model-dependent, and that simple well-designed prompts outperform complex multi-step pipelines. We evaluate our approaches on the ArchEHR-QA 2026 shared task at CL4Health, achieving 58.0 F1 for evidence identification and 31.8 overall for answer generation.
up
Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026
Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026
Asma Ben Abacha | Steven Bethard | Danielle Bitterman | Tristan Naumann | Kirk Roberts
Asma Ben Abacha | Steven Bethard | Danielle Bitterman | Tristan Naumann | Kirk Roberts
Overview of the MEDIQA-EVAL 2026 Shared Task on Evaluation Metrics in Medical Multimodal Question Answering
Asma Ben Abacha | Wen-wai Yim
Asma Ben Abacha | Wen-wai Yim
Evaluating clinical text generation remains challenging, as automatic metrics often correlate weakly with clinician judgments. This issue is particularly pronounced in medical multimodal question answering (MMQA), where systems must integrate visual and textual information and evaluation must capture factual accuracy, visual grounding, completeness, and overall coherence. Despite rapid progress in MMQA, there is limited consensus on clinically meaningful evaluation, and existing metrics, largely adapted from general NLG or VQA, often fail to capture domain-specific criteria. We introduce MEDIQA-EVAL 2026, a shared task on evaluation metrics for medical multimodal QA. To our knowledge, this is the first shared task focused on evaluating automatic metrics in this setting. We release a dataset of medical visual question-answer pairs annotated with multidimensional clinician judgments. Systems are evaluated by the correlation of their metric scores with expert ratings on a held-out test set. Participants explored diverse approaches, including vision-language models, retrieval-augmented judging, metric-specific classifiers, reinforcement learning, and LLM-as-a-judge frameworks. Results show that model-based evaluators achieve stronger alignment with human judgments than traditional NLG metrics, particularly on English data, while performance remains lower on Chinese, highlighting challenges in multilingual evaluation. Notably, our MEDIQA LLM-as-a-judge approach achieves strong performance across both languages.
SUAT-BMI at MEDIQA-EVAL 2026: An Ensemble Approach to Language Models as Judges for Automatic Rating of Medical Responses
Xinzhe Peng | Liyuan E | Kun Feng | Jielin Li | Yuxuan Tang | Zhao Li
Xinzhe Peng | Liyuan E | Kun Feng | Jielin Li | Yuxuan Tang | Zhao Li
The MEDIQA-EVAL 2026 shared task focuses on developing automatic evaluation metrics for LLM-generated responses in dermatology and wound care. While LLMs have shown promise as judge models, the reliability of these metrics remains underexplored. In this work, we study how well judge models can approximate human expert ratings across clinical evaluation criteria. We evaluate multiple approaches, including few-shot prompting, BERT fine-tuning, and retrieval-augmented generation (RAG), and combine them in an ensemble framework. Our method achieves a correlation score of 0.481, ranking first among 41 participating teams. Our results provide insight into the reliability of LLM-based evaluation metrics and highlight their potential for scalable clinical assessment.
Overview of the MEDIQA-SYNUR 2026 Shared Task on Observation Extraction from Nurse Dictations
George Michalopoulos | Jean-Philippe Corbeil | Cari Bader | Nathan Bodenstab | Asma Ben Abacha
George Michalopoulos | Jean-Philippe Corbeil | Cari Bader | Nathan Bodenstab | Asma Ben Abacha
Hospital nurses spend a significant portion of their shifts performing manual data entry tasks. An automatic solution for extracting medical information from nurse dictations into large spreadsheet ontology (flowsheet) could reduce the documentation burden of nurses and alleviate nurse burnout. We introduce the MEDIQA-SYNUR shared task, the first challenge on extracting and normalizing clinical observations from nurse dictations and mapping them to a large ontology of clinical concepts. 13 teams participated in the challenge and experimented with a broad range of approaches. In this paper, we describe the MEDIQA-SYNUR task, the datasets, and the participant’s results and solutions.
SemAnTICA Lab at MediQA-SYNUR 2026: Route, Extract and Verify – An LLM-gated Ensemble for Parsing Nurse Dictations
Sy Hwang | Katherine S. Pitcher | Sue Hyon Kim | Yoonjae Lee | Hayoung K. Donelly | Harsh Bandhey | Andrew J. King | Karen O’Connor | Ryan J. Urbanowicz | Danielle L. Mowery
Sy Hwang | Katherine S. Pitcher | Sue Hyon Kim | Yoonjae Lee | Hayoung K. Donelly | Harsh Bandhey | Andrew J. King | Karen O’Connor | Ryan J. Urbanowicz | Danielle L. Mowery
We describe the Semantic Analysis of Text to Inform Clinical Action (SemAnTICA) Lab’s system for the MediQA-SYNUR 2026 shared task on extracting structured clinical observations from nurse dictation transcripts. The task requires mapping observations from disfluent conversational text to a large, fixed ontology and producing strictly normalized outputs, where small amounts of concept over-selection severely degrade micro-F1 score. Our approach evolved from a full-schema in-context baseline to a pipeline that explicitly separates concept selection from value extraction. We first preprocess transcripts, then generate transcript-specific concept candidates using hybrid sparse–dense retrieval. The candidates are then pruned with an evidence-based filter. For extraction, we adopt a system-level mixture-of-experts design with an online LLM router that selects a subset of domain-specialized experts per transcript. Each expert operates over a constrained schema partition to reduce spurious predictions. We enhance robustness with agreement-gated ensembling and targeted adjudication for ambiguous cases. Finally, we intersect complementary high-recall and high-precision runs to produce the best submission. Our system ranked first on the official test leaderboard with F1 = 0.814, P = 0.826, R = 0.801.
L2D-Clinical: Learning to Defer for Adaptive Model Selection in Clinical Text Classification
Rishik Kondadadi | John E. Ortega
Rishik Kondadadi | John E. Ortega
Clinical text classification requires choosing between specialized fine-tuned models (BERT variants) and general-purpose large language models (LLMs), yet neither dominates across all instances. We introduce Learning to Defer for clinical text (L2D-Clinical), a framework that learns when a BERT classifier should defer to an LLM based on uncertainty signals and text characteristics. Unlike prior L2D work that defers to human experts assumed universally superior, our approach enables adaptive deferral-improving accuracy when the LLM complements BERT. We evaluate on two clinical tasks: (1) ADE detection (ADE Corpus V2), where BioBERT (F1=0.911) outperforms the LLM (F1=0.765), and (2) treatment outcome classification (MIMIC-IV with multi-LLM consensus ground truth), where GPT-5-nano (F1=0.967) outperforms ClinicalBERT (F1=0.887). On ADE, L2D-Clinical achieves F1=0.928 (+1.7 points over BERT) by selectively deferring 7% of instances where the LLM’s high recall compensates for BERT’s misses. On MIMIC, L2D-Clinical achieves F1=0.980 (+9.3 points over BERT) by deferring only 16.8% of cases to the LLM. The key insight is that L2D-Clinical learns to selectively leverage LLM strengths while minimizing API costs.
TRUMEDIQA: A Modular Trustworthy RAG Pipeline for Multilingual Medical Question Answering
Meryem El Fatimi | Ayoub Nainia | Jihad Zahir
Meryem El Fatimi | Ayoub Nainia | Jihad Zahir
Medical question answering systems must balance usefulness with safety, particularly in low-resource linguistic settings where robustness is limited and hallucinations can cause harm. We present TRUMEDIQA, a reproducible multilingual medical QA pipeline for Moroccan Darija, Arabic, French, and English, deployed on WhatsApp with text and voice interactions. TRUMEDIQA uses layered decision-making: (i) language identification, (ii) a pre-retrieval intent router that maps queries to one of 38 clinical FAQ categories to constrain retrieval, and (iii) post-retrieval LLM-based re-ranking that selects the best candidate answer or returns a null decision to trigger a safe fallback (abstention). Answers are retrieved from a curated FAQ knowledge base validated by medical professionals. We evaluate TRUMEDIQA with 21 participants submitting 290 questions across four languages. An expert annotator labels each interaction as relevant, acceptable, or irrelevant, and we also measure correct abstentions when no suitable answer exists in the knowledge base. An ablation study shows that routing and re-ranking improve the weighted relevance score from 0.25 to 0.94 and precision from 0.53 to 0.98 versus a naïve retrieval baseline, while increasing correct abstention on unanswerable queries from 4.38% to 69.77%.
Evaluating the Retrieval Component in a Retrieval-Augmented Summarization System for Patient Records in French
Marco Naguib | Christel Gérardin | Victor Beaucoté | Cyril Charron | Adrien Joseph | Aurélie Névéol | Xavier Tannier
Marco Naguib | Christel Gérardin | Victor Beaucoté | Cyril Charron | Adrien Joseph | Aurélie Névéol | Xavier Tannier
In emergency and intensive care settings, clinicians must process large volumes of patient data to make time-sensitive decisions. Summarizing patient records can help reduce cognitive load and improve decision-making, but the complexity and variability of clinical documentation create challenges. This study explores a Retrieval-Augmented Generation (RAG) approach, consisting of two phases: (1) retrieval of relevant clinical information, and (2) generation of a summary. This paper evaluates the retrieval component of RAG systems, focusing on its performance in clinical contexts. Using French clinical text, we assess retrieval models and propose an annotation-based querying method to improve accuracy and consistency in retrieving core clinical information. We use an annotated dataset from anonymized-hospital to benchmark retrieval models tailored for French clinical records. The proposed annotation-based querying method is compared to traditional prompt-based approaches, demonstrating improved retrieval performance. The findings indicate that specialized retrieval techniques enhance the effectiveness of RAG systems in clinical settings, providing more accurate and relevant information for summarization. The study contributes to the development of clinical decision support tools by improving the retrieval process in RAG systems. The proposed methods offer a structured approach to summarizing patient records, which may help clinicians manage information more efficiently.
Recent advancements in Large Language Models (LLMs) have played a significant role in reducing human workload across various domains, a trend that is increasingly extending into the medical field. In this paper, we propose an automated pipeline designed to alleviate the burden on nurses by automatically extracting clinical observations from nurse dictations. To ensure accurate extraction, we introduce a method based on Retrieval-Augmented Generation (RAG). Our approach demonstrates effective performance, achieving an F1-score of 0.796 on the MEDIQA-SYNUR test dataset.
Automatic Generation of Discharge Summaries Using Large Language Models: A Systematic Literature Review
Lucas Molino-Piñar | Manuel Carlos Diaz Galiano | María-Teresa Martín-Valdivia
Lucas Molino-Piñar | Manuel Carlos Diaz Galiano | María-Teresa Martín-Valdivia
Discharge summaries are critical documents for continuity of care, yet their manual creation imposes significant burdens on clinical staff. This systematic literature review examines current approaches to automatic generation of discharge summaries using Natural Language Processing (NLP) and Large Language Models (LLMs). Following the Kitchenham guidelines for systematic reviews in software engineering, we searched Scopus and PubMed databases for studies published between 2023 and 2026, identifying 9 primary studies from an initial pool of 102 papers. Our analysis reveals that GPT-4 and its variants dominate current research (appearing in 6 of 9 studies), while open-source alternatives like LLaMA show promise for privacy-preserving deployments. Evaluation primarily relies on automatic metrics (ROUGE, BLEU) combined with human expert assessment. Key challenges include hallucination rates ranging from 33% to 64%, information omission, integration with Electronic Health Record (EHR) systems, and context window limitations. Studies addressing factuality employ human-in-the-loop validation, prompt engineering techniques, and knowledge graph-based correction mechanisms. Despite these challenges, recent implementations demonstrate clinical feasibility, with one study achieving a 94.35% System Usability Score. This review provides a comprehensive synthesis of the state-of-the-art and identifies opportunities for future research in this rapidly evolving field.
Smart_solutions at MEDIQA-SYNUR 2026: A Multi-Stage LLM Pipeline for Nursing Observation Extraction
Prateek Munjal
Prateek Munjal
Extracting clinical observations from nursing dictations addresses an important problem of addressing burden in clinical documentation. In this work, we describe our approach submitted to MEDIQA-SYNUR 2026, which achieved third place among participating teams with a balanced precision and recall of 0.80 on the unseen test set. Our approach, instead of finetuning LLMs, is to adopt a multi-stage pipeline of agents: Observation Agent, Ontology Matching Agent, Relevance Scoring Agent, Evidence Assignment Agent, and Formatting Agent. First, the Observation Agent extracts clinical observations and corresponding evidence from the nurse transcript. These observations are then processed by the Ontology Matching Agent, which maps them to a restricted set of candidate ontology fields via TF-IDF–based retrieval, and subsequently evaluated by the Relevance Scoring Agent, which assigns continuous support scores (1–5) to each candidate field. Finally, field value assignments are performed by the Evidence-Based Agent, which extracts values strictly from nurse transcripts and clinical observations (Observation Agent outputs) to populate each ontology field. These outputs are then formatted by the Formatting Agent to ensure correct submission structure with the necessary metadata. Our agentic system results suggest that combination of agents with prompt engineering can narrow the gap between general and specialized clinical NLP models, making it an immediately deployable alternative to traditional fine-tuning.
GS-BrainText: A Multi-Site Brain Imaging Report Dataset from Generation Scotland for Clinical Natural Language Processing Development and Validation
Beatrice Alex | Claire Grover | Arlene Casey | Richard Tobin | Heather Whalley | William Whiteley
Beatrice Alex | Claire Grover | Arlene Casey | Richard Tobin | Heather Whalley | William Whiteley
We present GS-BrainText, a curated dataset of 8,511 brain radiology reports from the Generation Scotland cohort, of which 2,431 are annotated for 24 brain disease phenotypes. This multi-site dataset spans five Scottish NHS health boards and includes broad age representation (mean age 58, median age 53), making it uniquely valuable for developing and evaluating generalisable clinical natural language processing (NLP) algorithms and tools. Expert annotations were performed by a multidisciplinary clinical team using an annotation schema, with 10–100% double annotation per NHS health board and rigorous quality assurance. Benchmark evaluation using EdIE-R, an existing rule-based NLP system developed in conjunction with the annotation schema, revealed some performance variation across health boards (F1: 86.13-98.13), phenotypes (F1: 22.22-100) and age groups (F1: 87.01-98.13), highlighting critical challenges in generalisation of NLP tools. The GS-BrainText dataset addresses a significant gap in available UK clinical text resources and provides a valuable resource for the study of linguistic variation, diagnostic uncertainty expression and the impact of data characteristics on NLP system performance.
SASTA Self Assessment: An efficient human-in-the-loop strategy for developmental and pathological language analysis
Jan Odijk | Jelte van Boheemen | Xander Vertegaal | Tessel Boerma | Marijn Schraagen
Jan Odijk | Jelte van Boheemen | Xander Vertegaal | Tessel Boerma | Marijn Schraagen
This paper introduces SASTA self-assessment, i.e., a self-assessment procedure for the SASTA application for semiautomatic analysis of spontaneous language transcripts for Dutch. We introduce SASTA and the methods that it supports. These methods are used to assess the language development of young children and to assess the language skills of patients with aphasia. We illustrate this with typical example utterances. The performance of SASTA is good but not good enough for fully automatic use. The self-assessment procedure attempts to automatically identify utterances that require revision by a human expert. The self-assessment procedure gives promising results for datasets for the ASTA and STAP methods. Significant improvements are still needed for the TARSP method, but there is still potential for such improvements, and we sketch some directions to achieve such improvements.
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
Michele Miranda | Xinlan Yan | Nishant Mishra | Rachel Murphy | Ameen Abu Hanna | Sébastien Bratières | Iacer Calixto
Michele Miranda | Xinlan Yan | Nishant Mishra | Rachel Murphy | Ameen Abu Hanna | Sébastien Bratières | Iacer Calixto
Protecting patient privacy in clinical narratives is essential for enabling secondary use of healthcare data under regulations such as GDPR and HIPAA. While manual de-identification remains the gold standard, it is costly and slow, motivating the need for automated methods that combine privacy guarantees with high utility. Historically, most automated text de-identification pipelines employed named entity recognition (NER) to identify protected entities for redaction. Although methods based on differential privacy (DP) provide formal privacy guarantees, more recently also large language models (LLMs) are increasingly used for text de-identification in the clinical domain. In this work, we present the first comparative study of DP, NER, and LLMs for Dutch clinical text de-identification. We investigate these methods separately as well as hybrid strategies that apply NER or LLM preprocessing prior to DP, and assess performance in terms of privacy leakage and extrinsic evaluation (entity and relation classification). We show that DP mechanisms alone degrade utility substantially, but combining them with linguistic preprocessing, especially LLM-based redaction, significantly improves the privacy–utility trade-off.
SQUCS at MEDIQA-SYNUR 2026: A Multi-Agent Open Source LLM System for Nursing Observation Extraction
Riham JeebAllah | Adhari AlZaabi | Abdulrahman Khalifa AAlAbdulsalam
Riham JeebAllah | Adhari AlZaabi | Abdulrahman Khalifa AAlAbdulsalam
Clinical nursing documentation contains detailed observational information that is essential for patient monitoring and clinical decision-making, yet this information is predominantly recorded in free-text form. The MEDIQA-SYNUR shared task addresses this challenge by requiring systems to extract structured nursing observations from clinical transcripts under strict constraints on evidence grounding and value normalization. In this work, we present a multi-agent large language model (LLM)–based system for the MEDIQA-SYNUR task. We utilize the Llama3 open source LLM for this purpose for ease of local deployment within hospital digital infrastructure. Our system decomposes the extraction process into specialized agents responsible for schema-guided extraction, rule-based validation, and precision-oriented filtering. Starting from a baseline multi-agent pipeline, we conduct a systematic error analysis over the entire development set, examining all false positive and false negative predictions. Our final configuration, selected after extensive exploration and error analysis, combined transcript segmentation, the precision agent, and a suppression table derived from development-set analysis. On the development set, this setup achieved an F1 score of 0.6930 (precision = 0.6427, recall = 0.7518). Applying the same configuration directly to the test set, without any additional tuning, yielded an F1 score of 0.5923 (precision = 0.5292, recall = 0.6725). These results represent the most effective balance of precision and recall achieved through our iterative refinements and reflect the final state of the system as submitted for the competition
LTRC-IIIT at MEDIQA-SYNUR 2026: Benchmarking a Fully Local, Training-Free RAG Pipeline
Aashwin Vaish | Dipti Misra Sharma
Aashwin Vaish | Dipti Misra Sharma
In this paper we present our solution to MEDIQA-SYNUR 2026 shared task organized at LREC-ClinicalNLP workshop. The goal of the task is to populate Electronic Health Record (EHR) flowsheets using the transcriptions of nurse dictations, to alleviate the extensive manual labor associated with sifting through large flowsheets of clinical concepts. We propose a modular architecture combining heuristic-driven Retrieval-Augmented Generation (RAG) with grammar-constrained decoding on an open-weight, quantized, 8B-parameter model (Llama 3.1 Instruct). Our system achieves an F1 score of 0.57, significantly trailing the initial zero-shot experiments with GPT-4o and placing it towards the lower end of the current leaderboard. We conduct a failure analysis of this approach while establishing a baseline for privacy-preserving, zero-shot documentation assistants.
BDI at MEDIQA-EVAL 2026: A ReAct-Style Multimodal Agent for Fine-Grained Medical Response Assessment
Justin Xu | Zizheng Zhang | Augustine Luk | Benjamin Khong | Haochen Cui | Samuel Hwang | Alyssa Pradhan | Kevin Yuan | David W. Eyre
Justin Xu | Zizheng Zhang | Augustine Luk | Benjamin Khong | Haochen Cui | Samuel Hwang | Alyssa Pradhan | Kevin Yuan | David W. Eyre
Free-text evaluation of multimodal clinical question answering (QA) systems remains a central challenge in medical NLP due to the complexity of medical knowledge, the necessity of integrating visual and textual information, and the limitations of existing automatic evaluation metrics for open-ended outputs. In this work, we present a training-free, agentic evaluation framework that formulates response scoring as evidence-guided orchestration of components rather than a task requiring conventional end-to-end fine-tuning of underlying LLMs/VLMs. Our ReAct-style evaluator combines (i) structured reasoning, (ii) multimodal retrieval of similar encounters, (iii) auxiliary explainable feature-based regression models that provide numeric priors and human-interpretable signals, (iv) VLM-generated visual QA references for comparison, and (v) optional image augmentation tools. Unlike standard LLM-as-a-judge approaches that rely on direct generative scoring, our agent decomposes evaluation into modular stages of evidence acquisition, structured feature modeling, and integrative reasoning. We apply this architecture to the MEDIQA-EVAL shared task - a multimodal, multilingual clinical evaluation challenge that assesses system-generated answers for patient queries paired with images along multiple clinical quality dimensions. We report results across both English and Chinese tracks, comparing against baseline prompting methods, and discuss the feasibility and limitations of lightweight agentic systems for clinical QA evaluation.
MasonNLP at MEDIQA-SYNUR 2026: Retrieval-Augmented Large Language Models for Schema-Constrained Clinical Information Extraction
A H M Rezaul Karim | Özlem Uzuner
A H M Rezaul Karim | Özlem Uzuner
Conversational nurse-patient transcripts contain actionable observations, but converting these transcripts into structured representations at scale remains challenging. Documentation burden is substantial, with prior studies showing clinicians spend large portions of their workday on documentation and related desk work rather than direct patient care. MEDIQA-SYNUR focuses on observation extraction from conversational nurse-patient transcripts, requiring systems to normalize these narratives into a predefined schema with value-type constraints. We propose a modular retrieval-augmented generation (RAG) pipeline that uses the training set as an exemplar corpus, combines schema-constrained prompting (full schema vs. pruned candidate schema), deterministic schema-based postprocessing, and a second-pass audit, with two LLM backbones: Llama-4-Scout-17B-16E-Instruct and GPT-5.2 with corresponding embedding models for RAG. Our best configuration uses GPT-5.2 with full schema, RAG, and a second-pass auditing, achieving 80.36% F1 score. Overall, our results show that RAG consistently improves performance, while the optimal degree of schema constraint depends on the model, and second-pass auditing yields modest additional gains by correcting residual schema-adherence errors.
Night Shift Nerds at MEDIQA-SYNUR 2026: Pushing Small Large Language Model Capability for Clinical Observation Extraction and Normalization from Nurse Dictation using RLVR
Bayu Aryoyudanta | Maria Yuliana | Mikie Rachman | I Made Agus Setiawan
Bayu Aryoyudanta | Maria Yuliana | Mikie Rachman | I Made Agus Setiawan
We presented a small decoder-only language model for clinical observation extraction and normalization from nurse dictation developed using Reinforcement Learning with Verifiable Rewards (RLVR). We fine-tune Qwen3-1.7B model using a two-stage pipeline: (1) supervised fine-tuning (SFT) with an augmented chain-of-thought (CoT) dataset generated by a teacher model to mitigate RL cold-start, followed by (2) GRPO-based RLVR with multi-component reward functions that verify output format, concept presence, value type, and value correctness using the shared-task ontology (193 concepts) as a verifier. On the development set, SFT+GRPO substantially outperforms GRPO-only (F1 0.803 vs. 0.620). After the test holdout was released, our final system achieved 0.700 precision, 0.785 recall, and 0.740 F1. Error analysis shows remaining challenges in concept over-detection and missed concepts, as well as boundary errors in categorical and multi-select value type extraction. Our results demonstrates that small language models can enable accurate, cost-effective, and privacy preserving automated clinical documentation for nurse dictation, supporting scalable deployment in low-resource healthcare settings to reduce nurses’ documentation burden.
Extracting Medication Instructions from Dutch General Practice Electronic Health Records with Local Natural Language Processing
Marya Dukmak | Constanza L. Andaur Navarro | Artuur Leeuwenberg
Marya Dukmak | Constanza L. Andaur Navarro | Artuur Leeuwenberg
The extraction of structured medication prescription data from unstructured clinical text remains a critical challenge for clinical research and data standardization. This study investigates the application of Natural Language Processing (NLP) techniques to Dutch electronic health records (EHRs) from the Julius General Practitioners Network. The goal is to automatically extract key prescription attributes including dosage, duration, and medication unit and prepare them for integration into the ConcePTION Common Data Model, to support scalable pharmacoepidemiological research. We compare a lightweight rule-based system with transformer-based models (RobBERT and MedRoBERTa) under the technical constraints of a Trusted Research Environment, where external resources and cloud-based solutions are restricted. Using a dataset of 1,819 manually annotated records, the approaches are evaluated on predictive performance and computational costs. Results show that the rule-based system achieves strong accuracy and computational costs for structured patterns, while transformer-based models demonstrate greater robustness to linguistic variability. However, both approaches encounter difficulties with ambiguous dosage formats and long treatment durations. Our findings indicate that NLP methods can substantially improve the structuring of Dutch prescription data and support scalable pharmacoepidemiological research. Future work should focus on improving generalization and expanding annotated datasets to enhance model reliability.
Gladiator at MEDIQA-SYNUR 2026: Contextual Clinical Extraction: Integrating Foundation Models with Domain-Specific Validation Rules
Siva Satyanarayana Raju Pusapati | Ankit Singh
Siva Satyanarayana Raju Pusapati | Ankit Singh
We present a hybrid extraction system that combines large language model capabilities with rule-based precision for extracting structured clinical observations from nursing dictation transcripts. Our approach leverages Claude Opus 4.5 as the primary extractor, enhanced with comprehensive prompt engineering that includes the complete 193-concept schema, few-shot examples, and detailed validation rules covering respiratory, cardiac, diagnosis, and mental status fields. The LLM output undergoes extensive post-processing with six specialized filters that remove speculative diagnoses, validate physiological ranges, ensure unit-field dependencies, and verify contextual appropriateness. Five correction mechanisms normalize breathing patterns, map dyspnea severity, standardize assistance levels, clean STRING fields, and handle multi-select conjunctions. A supplementary rule-based component employs 400+ regex patterns with contextual validation to capture high-confidence observations, particularly for vital signs and categorical fields. The system requires cardiac keywords for heart rate extraction and respiratory context for respiration rates, preventing false positives from unrelated numeric values. Results are merged through an intelligent strategy that prioritizes LLM comprehensiveness while supplementing with rule-based findings. A strict schema validation layer ensures all four value types (NUMERIC, STRING, SINGLE_SELECT, MULTI_SELECT) conform to enumerated options and physiological ranges. This multi-layered approach balances recall through LLM reasoning with precision through rule-based validation, effectively structuring natural nursing narratives into standardized EHR-ready observations.
MedAware at MEDIQA-EVAL 2026: Vision-Language Model Fine-Tuning with Logprob-Based Score Calibration for Medical Response Evaluation
Ziqi Hao | Pengbo Liu
Ziqi Hao | Pengbo Liu
We present MedAware, our MEDIQA-EVAL 2026 system for predicting human ratings of medical QA responses from text and images. We fine-tune Qwen3-VL models (4B/8B/32B) with supervised fine-tuning (SFT), and study GRPO as an optional second stage under both LoRA and full-parameter settings. To handle severe label skew and unstable correlation metrics, we use logprob-based continuous scoring with quantile calibration, converting token probabilities into calibrated metric scores without retraining. This reduces prediction collapse on skewed dimensions and improves metric stability in both English and Chinese. The approach follows the official reference-based shared-task setup and is designed to produce meaningful metric estimates even under extreme class imbalance. In the official shared-task submission setting (8B-LoRA SFT with discrete scoring), our system ranked 3rd on English and 1st among participants on Chinese. Separately, in post-competition offline re-evaluations with logprob scoring, the best tested configuration reaches 0.449 EN-ALL and 0.308 ZH-ALL, while SFT initialization remains critical for effective GRPO.
HSE NLP TEAM at MEDIQA-SYNUR 2026: Consensus Adjudication Ensemble (ACE): Balancing Precision and Recall for Schema-Bystander Clinical Extraction
Airat A. Valiev
Airat A. Valiev
Clinical documentation from nurse dictations is labor-intensive and error-prone, yet it contains high-value observations that must be transferred into structured flowsheets. The MEDIQA-SYNUR 2026 shared task evaluates systems that extract and ontology-align 193 clinical concepts (with heterogeneous value types) from synthetic speech transcripts derived from intensive care notes. We describe the Consensus Adjudication Ensemble (ACE), a three-stage pipeline that (i) maximizes candidate coverage via complementary generators, (ii) enforces high precision through a dedicated adjudicator that operates as a verifier rather than a generator, and (iii) restores strict schema compliance using a targeted, token-efficient repair step. On the official test set we achieve an exact-match micro-F1 of 0.7996 (P=0.7812, R=0.8188), ranking 4th on the leaderboard. Beyond the competitive result, we analyze clinically relevant failure modes - hallucinated interventions, over-confident categorical labels, and unit/normalization errors - and quantify adjudication trade-offs: 2,219 candidates removed, 91.3% of which are true false positives, at the cost of 8.7% mistakenly removed true positives. Finally, targeted schema repair reduces validation context from approx. 230k tokens to <2k per document while preserving most extraction gains. Keywords: clinical information extraction, nurse dictations, ontology alignment, ensemble methods, adjudication, error analysis
Lakefront AI Ramblers at MEDIQA-SYNUR 2026: Hybrid Retrieval and LLM Verification for Open-Source Schema-Guided Clinical Information Extraction
Michael T. Saban | Arsalan Yaghoubi | Behnaz Eslami | Samie Tootooni | Dmitriy Dligach
Michael T. Saban | Arsalan Yaghoubi | Behnaz Eslami | Samie Tootooni | Dmitriy Dligach
Schema-constrained clinical information extraction requires identifying text-supported observations and outputting exact schema identifiers and values. In the MEDIQA-SYNUR 2026 shared task, synthetic nursing dictations were mapped to structured JSON outputs aligned with a 193-concept clinical schema under strict exact-match evaluation. We extended the baseline pipeline, which consists of transcript segmentation, schema retrieval, and LLM-based extraction, with hybrid schema retrieval, supervised fine-tuning (SFT) of open-source LLMs, and LLM-based verification. Our hybrid retrieval approach combined dense embeddings with sparse BM25 representations using a convex combination strategy, improving schema coverage to 0.994 recall@60 on the development set. We evaluated GPT-4o, GPT-4o-mini, Llama-3-8B-Instruct, and Llama-3.3-70B-Instruct, applying LoRA-based SFT to open-source models. On the official test set, our best submitted configuration (Llama-3.3-70B-Instruct-SFT with union voting and GPT-4o-mini verification) achieved 0.711 F1. Post-competition experiments showed that Llama-3-8B-Instruct-SFT (train + dev) reached 0.723 F1 under the same post-processing pipeline. For reference, GPT-4o achieved 0.791 F1 and did not benefit from post-processing. Performance differences across development and test splits further highlight the sensitivity of post-processing strategies to variation across split distribution. Overall, integrating high-recall retrieval, SFT, and LLM verification substantially narrows the performance gap between open- and closed-source models for schema guided clinical extraction.
JMedWiC: A Japanese Word-in-Context Dataset in the Medical Domain
Koki Horiguchi | Seiji Sugiyama | Tomoyuki Kajiwara | Shoko Wakamiya | Eiji Aramaki
Koki Horiguchi | Seiji Sugiyama | Tomoyuki Kajiwara | Shoko Wakamiya | Eiji Aramaki
We release JMedWiC, a Japanese dataset for Word-in-Context (WiC) tasks specifically tailored to the medical domain. To address the challenge of word sense disambiguation, where the meaning of a word varies depending on its context, previous research has developed WiC datasets to evaluate word sense identity by determining whether a target word shares the same sense across two given contexts. In the medical domain, the misinterpretation of word senses can hinder the accurate comprehension of medical information; however, there is currently no Japanese WiC dataset specialized for this domain. Moreover, existing WiC datasets have been constructed using lexical resources with sense inventories, such as WordNet and UMLS, but such resources are not sufficiently developed for Japanese. Therefore, we construct a Japanese WiC dataset in the medical domain by manually annotating sense-identity labels for target words in context pairs automatically extracted from a large-scale corpus, without relying on lexical resources.
LTRC-Medicom at MEDIQA-SYNUR 2026: Schema-Guided Clinical Information Extraction with Hybrid Clustering-SFT-Verification
Pasumarthy Deepak | Sushvin Marimuthu | Parameswari Krishnamurthy
Pasumarthy Deepak | Sushvin Marimuthu | Parameswari Krishnamurthy
Extracting structured clinical data from unstructured patient transcripts is challenging due to large target schemas and inherent linguistic ambiguity. We address the extraction of 193 heterogeneous clinical attributes from nursing notes and clinician–patient dialogues, and demonstrate that zero-shot large language models (LLMs) are ineffective in this setting, achieving an F1 score below 0.15 due to context window saturation and hallucination. We propose a four-stage framework that combines semantic schema clustering, role-based chain-of-thought prompting, supervised fine-tuning of Llama-3.1-8B, and transcript-verified post-processing. Our approach achieves an F1 score of 0.66, representing a 4.4x improvement over the baseline, by balancing high recall from generative models with high precision from verification. These results highlight the effectiveness of hybrid pipelines for high-stakes clinical information extraction.
SloCal-Net at MEDIQA-Eval 2026: Investigating the Impact of Reasoning and External Context on Medical Answer Grading
Primoz Kocbek | Valentina Carbonari | Pierangelo Veltri | Pietro Hiram Guzzi | Gregor Stiglic
Primoz Kocbek | Valentina Carbonari | Pierangelo Veltri | Pietro Hiram Guzzi | Gregor Stiglic
Automated evaluation of multimodal medical answers is essential for scalable safety assessment, yet it remains difficult to align automatic scores with expert judgment across languages and image modalities. We describe SloCal-Net’s systems for the MEDIQA-EVAL 2026 shared task, framing evaluation as rubric-conditioned multimodal judging: the judge receives the question, image(s), candidate answer, and task-specific criteria, and outputs criterion-level scores and an overall rating. Evidence retrieval was initialized using ChatGPT Deep Research, producing a 25-document clinical corpus used for lightweight retrieval-augmented grounding. On the official leaderboard, our best submission (GPT-5-mini with web search and RAG) achieved Pearson correlations of 0.466 on English and 0.260 on Chinese expert ratings. In post-competition experiments with open-source judges, the best English Pearson reached 0.272 with GLM-4.6V and 0.212 with Qwen3-VL-30B-Thinking, while Chinese correlations were lower, highlighting remaining gaps in multilingual calibration and image–text grounding.
MIDAS_SYNUR at MEDIQA-SYNUR 2026: A Prompting Study for Clinical Observation Extraction from Nurse Dictation Transcriptions
Swetha Krishna Sriram | Akshitaa Sahoo
Swetha Krishna Sriram | Akshitaa Sahoo
This paper describes MIDAS_SYNUR, a system developed for the MEDIQA-SYNUR task at ClinicalNLP 2026 on observation extraction from nurse dictations. The primary system adopts a single-prompt, field-rich few-shot strategy using GPT-5.2, jointly generating all schema fields in one structured output. Few-shot demonstrations are curated and grouped by value type, with five examples per type, promoting consistency across heterogeneous value distributions while leveraging global context to resolve cross-field dependencies. To analyze design trade-offs, this holistic strategy is compared against a field-wise decomposed prompting baseline, where each schema field is extracted independently using explicit positive and NULL demonstrations to improve absence detection and reduce cross-field interference. Zero-shot variants of both approaches are also evaluated to isolate the contribution of in-context examples. The results highlight inference-time prompting as a simple, reproducible, and competitive baseline for large-scale clinical observation extraction from conversational nurse dictations. Keywords: Prompt Engineering, Few-Shot Prompting, In-Context Learning, Structured Output
AnotherOne at MEDIQA-SYNUR 2026: Detect, Extract, Normalize - Knowledge-Grounded LLM Pipeline for Clinical Observation Extraction
Jerrin John Thomas | Parameswari Krishnamurthy
Jerrin John Thomas | Parameswari Krishnamurthy
We present a system for the MEDIQA-SYNUR 2026 shared task on extracting structured clinical observations from nurse dictation transcripts. The transcripts contain spoken-style clinical language with disfluencies, filler words, and hesitations. Our approach is a four-stage LLM inference pipeline preceded by an offline knowledge enhancement step: (1) knowledge-enhanced concept detection using medical domain clustering, (2) evidence-grounded value extraction, (3) schema-constrained value normalization, and (4) deterministic post-processing with fuzzy matching and unit pairing. In the offline step, we use the task ontology and training examples to generate per-concept clinical definitions and extraction rules, and group the 193 concepts into 19 non-exclusive medical domain clusters. These are injected into all downstream prompts as domain priors. All LLM stages use gpt-oss-120b with structured JSON output and chain-of-thought reasoning. The task requires exact matching on concept ID and value pairs across a 193-concept ontology, making precision particularly challenging. We iteratively refine concept definitions and prompt guidelines based on error analysis of the training data. Our system achieves an F1 score of 0.806 on the test set.
hgkai26 at MEDIQA-EVAL 2026: Automated Evaluation of Visual Medical Question Answering Using LLM-as-a-Judge
Haritha Gangavarapu
Haritha Gangavarapu
As there is a rise in the use of multimodal large language models (LLMs) for medical response generation, it is necessary to have reliable automated evaluation mechanisms that can assess the quality of model-generated outputs. The MediQA-Eval 2026 shared task focuses on grading AI-generated dermatology and wound care responses using structured human-aligned rubrics. In this work, we explore a zero-shot multimodal LLM-as-a-Judge framework to assess candidate responses across multiple quality dimensions. System performance is evaluated using the official task metrics designed to reflect alignment with human judgments. Our findings provide preliminary insights into the feasibility and limitations of LLM-based evaluators for rubric-guided medical response assessment.
Role-Adapted Clinical Report Generation for Ultrasound Measurements in Low-Resource Settings
Ayoub Nainia | Tanya Akumu | Noussair Lazrak | Karim Lekadir
Ayoub Nainia | Tanya Akumu | Noussair Lazrak | Karim Lekadir
Obstetric ultrasound is critical for monitoring fetal growth, yet in many low-resource settings, healthcare workers who perform or receive ultrasound measurements lack the training to interpret them clinically. We present a system that automatically generates role-adapted clinical reports from fetal biometry measurements, targeting six healthcare worker roles across three expertise levels. The system combines Retrieval-Augmented Generation (RAG) from a knowledge base extracted from the World Health Organization (WHO) Manual of Diagnostic Ultrasound with deterministic fetal growth percentile computation based on INTERGROWTH-21st international standards. The knowledge base is designed for multilingual extensibility: since the source material is from an official WHO document, entries can be translated into any target language by domain experts or machine translation services. A key design principle is that clinical decision support (red, yellow, and green alerts) is derived deterministically from percentile thresholds, not from the language model, ensuring safety regardless of LLM output quality. Evaluation demonstrates sub-millimeter accuracy in percentile computation, 100% correctness in decision support classification, measurable readability differentiation across roles (Flesch-Kincaid grade 8.8 for community health workers vs. 11-13 for clinical roles), and 98% factual consistency across 42 generated reports spanning seven clinical scenarios. The system is designed for local deployment without internet connectivity.
A Comparative Study of Approaches to Anonymization of Clinical Free Text in Spanish
Florencia Luciana Brunello | Laura Alonso Alemany | Serena Villata | Milagro Teruel
Florencia Luciana Brunello | Laura Alonso Alemany | Serena Villata | Milagro Teruel
The anonymization of clinical free-text records is a prerequisite for enabling the secondary use of healthcare data while preserving patient privacy. This challenge is particularly acute for Spanish clinical text, where annotated resources are scarce and practitioners lack clear empirical guidance on which technological approaches are more adequate to their particular restrictions and capabilities. In this work, we present a controlled comparative study of representative anonymization paradigms for Spanish clinical narratives, including a baseline rule-based approach, a general-purpose large language model under prompt-based inference, an off-the-shelf industrial NLP toolkit (spaCy) and comparable neural sequence labeling architectures. To ensure a fair and contamination-aware evaluation, particularly given the opacity of pretrained model training data, we introduce a synthetic clinical dataset. Recurrent neural network architectures, particularly the off-the-shelf spaCy toolkit, consistently achieve the best balance between effectiveness, computational efficiency, and deployment feasibility. We further observe that training task-specific embeddings end-to-end yields stronger generalization than incorporating pretrained representations. Although limited to Spanish and to representative instances of each paradigm, the study identifies stable performance tendencies across datasets. These results provide actionable guidance for institutions seeking to implement anonymization pipelines. This work contributes reproducible evaluation procedures and empirical evidence for privacy-preserving clinical NLP in Spanish.
Disagreement-Driven Joint Refinement of Retrieval and Decision Rules for Imbalanced Counseling Risk Classification
Zhihao Shao | Ryo Sekizaki | Shengzhou Yi | Toshihiko Yamasaki
Zhihao Shao | Ryo Sekizaki | Shengzhou Yi | Toshihiko Yamasaki
With the rapid growth of online counseling services, timely and reliable risk classification of counseling records is essential for supporting early screening and prioritizing limited intervention resources. High-risk samples refer to high-acuity suicide risk and require expedited human review. However, this task is challenging due to severe class imbalance (93% low-risk and 7% high-risk samples) and complex decision boundaries. Large language models (LLMs) exhibit unstable predictions and systematic errors in such imbalanced clinical-text settings. To address this issue, we propose Disagreement-Driven Joint Refinement (DDJR), an iterative, parameter-free refinement framework. It uses prediction disagreement between two inference settings, zero-shot and retrieval-augmented in-context learning, as the primary signal for identifying high-value instances. These disagreement-identified instances are transformed into adaptive refinement signals and used to jointly update both the exemplar pool and an executable rule set, thereby sharpening decision boundaries and improving prediction stability. Experiments on 6,481 real-world counseling records demonstrate that the proposed DDJR outperforms existing methods, achieving an accuracy of 0.915 and a Matthews Correlation Coefficient (MCC) of 0.583. These results demonstrate that DDJR achieves more stable and reliable predictions for high-stakes counseling risk classification in real-world settings.
Context-Aware SNOMED CT Entity Linking for Clinical Text
Provia Kadusabe | Demian Gholipour Ghalandari | Lauren Cassidy | Jack Boylan | Chris Hokamp | Abhishek Kaushik | Fiona Lawless
Provia Kadusabe | Demian Gholipour Ghalandari | Lauren Cassidy | Jack Boylan | Chris Hokamp | Abhishek Kaushik | Fiona Lawless
Mapping free-text mentions in clinical notes to standardized terminologies such as SNOMED CT is essential for large-scale secondary use of electronic health records, but remains challenging due to linguistic variability, under-specified annotation guidelines, term ambiguity, and ontology scale. This work presents a two-stage entity linking pipeline that combines span detection with context-aware concept linking and evaluates it on the SNOMED CT Entity Linking Challenge dataset. Our work builds upon the SNOMED CT entity linking challenge (CITATION), resulting in a fully open-source system. To our knowledge, this is the first end-to-end open-source system for this task. For span detection, we compare multiple neural architectures together with dictionary-based matching. For concept linking, we adopt a context-aware bi-encoder, and construct a multi-source knowledge base enriched with context derived from the SNOMED CT ontology. Finally, we implement an agentic re-ranker and test the effectiveness of LLM-backed re-ranking with access to annotation guidelines. In contrast to findings from the original shared task submissions, we show that context is important for optimal performance, and that agentic re-ranking with a state-of-the-art LLM only marginally improves overall performance, suggesting that the current benchmark may be approaching its practical ceiling. This work provides the first fully open-source, reproducible system for SNOMED CT entity linking, offering a foundation for future research and practical deployment.
Temporal Structure in Clinical Narratives in Portuguese: Insights from Cross-Document Annotation
Ana Luisa Fernandes | Purificação Silvano | Luís Filipe Cunha
Ana Luisa Fernandes | Purificação Silvano | Luís Filipe Cunha
Medical reports, by documenting disease progression and patient responses to treatment, form continuous narratives in which each new document adds a chapter to the patient’s clinical story. Constructing coherent patient timelines requires identifying temporal relations across multiple medical reports that compose a patient’s clinical journey. However, cross-document temporal annotation remains an underexplored area, largely due to the methodological and conceptual challenges it entails. This study addresses these challenges by investigating the identification and characterization of cross-document temporal relations in Portuguese medical records. For this purpose, cross-document annotation was performed on different types of reports (Group Consultation Reports, Discharge Reports, and General Reports) from patients diagnosed with Acute Myeloid Leukemia and followed at IPO-Porto, Portugal. Annotation was carried out using the Med2Story scheme, specifically designed to capture both temporal and medical information. Our results indicate that, although cross-document annotation of temporal information is more demanding in terms of both the annotation scheme and the annotation process, it enables the construction of coherent chronological representations of patients’ clinical journeys. Furthermore, the analysis reveals key characteristics of these clinical narratives, including the predominance of nominal events and the prevalence of simultaneity as the most frequent temporal relation type.
MOSAIC: A Multilingual, Taxonomy-Agnostic, and Computationally Efficient Approach for Radiological Report Classification in Low-Resource Settings
Alice Schiavone | Marco Fraccaro | Lea Marie Pehrson | Silvia Ingala | Rasmus Bonnevie | Michael Bachmann Nielsen | Vincent Beliveau | Melanie Ganz | Desmond Elliott
Alice Schiavone | Marco Fraccaro | Lea Marie Pehrson | Silvia Ingala | Rasmus Bonnevie | Michael Bachmann Nielsen | Vincent Beliveau | Melanie Ganz | Desmond Elliott
Radiology reports contain rich clinical information that can be used to train imaging models without relying on costly manual annotation. However, existing approaches face critical limitations: rule-based methods struggle with linguistic variability, supervised models require large annotated datasets, and recent LLM-based systems depend on closed-source or resource-intensive models that are unsuitable for clinical use. Moreover, current solutions are largely restricted to English and single-modality, single-taxonomy datasets. We introduce MOSAIC, a multilingual, taxonomy-agnostic, and computationally efficient approach for radiological report classification. Built on a compact open-access language model (MedGemma-4B), MOSAIC supports both zero-/few-shot prompting and lightweight fine-tuning, enabling deployment on consumer-grade GPUs. We evaluate MOSAIC across seven datasets in English, Spanish, French, and Danish, spanning multiple imaging modalities and label taxonomies. The model achieves a mean macro F1 score of 88 across five chest X-ray datasets, approaching or exceeding expert-level performance, while requiring only 24 GB of GPU memory. With data augmentation, as few as 80 annotated samples are sufficient to reach a weighted F1 score of 82 on Danish reports, enabling large-scale cohort classification with minimal human effort. Code and models are open-source, offering a practical alternative to large or proprietary LLMs in clinical settings.
MedNormJ: A Benchmark Dataset for Medical Concept Normalization in Japanese Clinical Documents
Yuki Tashiro | Seiji Shimizu | Tomohiro Nishiyama | Shoko Wakamiya | Eiji Aramaki
Yuki Tashiro | Seiji Shimizu | Tomohiro Nishiyama | Shoko Wakamiya | Eiji Aramaki
Medical concept normalization in clinical text is a fundamental technology for the secondary use of clinical data. However, constructing annotated resources for this task is challenging because annotation is both expertise-intensive and methodologically complex. As a result, a standard evaluation dataset for Japanese has yet to be established. In this study, we introduce a Japanese dataset for medical concept normalization, MedNormJ, which will be publicly available. The dataset consists of 397 pairs of medical expressions and their corresponding normalized disease names, manually curated from 96 medical documents, including case reports and radiology reports. Furthermore, we conduct comparative experiments using existing normalization approaches to benchmark their performance on this dataset in terms of both accuracy and computational efficiency. Through these experiments, we clarify the present performance level and identify remaining challenges specific to Japanese medical concept normalization.
Pediatric Sepsis Cohort Detection Using In-Context Pointwise V-Usable Information
Yingya Li | Alon Geva | Steven Bethard | Timothy A. Miller | Kate Madden | Matthew A. Eisenberg | Daniel P. Kelly | Guergana Savova
Yingya Li | Alon Geva | Steven Bethard | Timothy A. Miller | Kate Madden | Matthew A. Eisenberg | Daniel P. Kelly | Guergana Savova
Pediatric sepsis diagnosis remains a major clinical challenge due to non-specific symptoms and a lack of reliable diagnostic criteria. Large language models (LLMs) provide a scalable solution for processing and understanding unstructured text in medical records. However, identifying the most suitable model is non-trivial given the rapid growth of available LLMs. In this work, we proposed using in-context pointwise V-usable information (pvi) to estimate task difficulty and guide model selection for pediatric sepsis cohort detection. We applied in-context pvi to estimate task difficulty and inform model selection across 12 state-of-the-art open LLMs on the task, using electronic medical record data from 507 patient encounters at a U.S. children’s hospital. We compared the performance of the best-fitting LLM to feature-rich baseline models and a fine-tuned transformer. Our results show that the pvi-selected LLM outperforms the baselines, although the feature-rich bag-of-words model with a support vector machine also achieves competitive performance. We believe our approach demonstrates a promising application of current LLM techniques to high-stakes clinical tasks.
RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems
Adarsh Srinivasan | Jacob Dineen | Muhammad Uzair Sarfraz | Muhammad Umar Afzal | Irbaz Riaz | Ben Zhou
Adarsh Srinivasan | Jacob Dineen | Muhammad Uzair Sarfraz | Muhammad Umar Afzal | Irbaz Riaz | Ben Zhou
Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for clinical trust. We present RECAP (Reflect–Extract–Calibrate–Align–Produce), an inference-time framework grounded in cognitive appraisal theory that decomposes patient input into auditable, appraisal-theoretic stages without retraining. Across multiple benchmarks and models from 8B to 120B parameters, RECAP improves alignment with human judgments, with gains inversely proportional to model scale. Intermediate outputs further reveal that models systematically underweight relational factors such as social support. In blinded evaluations, oncology fellows rated RECAP responses significantly higher than baselines with 76–88% win rates, demonstrating that principled prompting can enhance medical AI’s emotional intelligence while maintaining the transparency required for clinical deployment.
Disentangling Ambiguity from Instability in Large Language Models: A Clinical Text-to-SQL Case Study
Angelo Ziletti | Leonardo D’Ambrosi
Angelo Ziletti | Leonardo D’Ambrosi
Deploying large language models for clinical Text-to-SQL requires distinguishing two qualitatively different causes of output diversity: (i) input ambiguity that should trigger clarification, and (ii) model instability that should trigger human review. We propose CLUES, a framework that models Text-to-SQL as a two-stage process (interpretations –> answers) and decomposes semantic uncertainty into an ambiguity score and an instability score. The instability score is computed via the Schur complement of a bipartite semantic graph matrix. Across AmbigQA/SituatedQA (gold interpretations) and a clinical Text-to-SQL benchmark (known interpretations), CLUES improves failure prediction over state-of-the-art Kernel Language Entropy. In deployment settings, it remains competitive while providing a diagnostic decomposition unavailable from a single score. The resulting uncertainty regimes map to targeted interventions - query refinement for ambiguity, model improvement for instability. The high-ambiguity/high-instability regime contains 51% of errors while covering 25% of queries, enabling efficient triage.
An OMOP-Based Open-Source Text-to-SQL Benchmark Dataset
Paul Legrand | Kawsar Noor | Satyam Bhagwanani | Richard J. Dobson
Paul Legrand | Kawsar Noor | Satyam Bhagwanani | Richard J. Dobson
Access to electronic health record (EHR) warehouses is limited by SQL expertise and complex clinical schemas. We present an open-source OMOP Common Data Model text-to-SQL benchmark (CDM v5.4) with a safety contract: output one executable SQL statement or the abstention token (<NO_SQL>) for unanswerable requests. Inputs are concept-normalized (entities as OMOP concept IDs) to decouple SQL generation from entity linking. We evaluate by executing predicted and reference queries on a synthetic OMOP PostgreSQL database, reporting Execution Accuracy (result equivalence) and a reliability score that rewards correct abstention and penalizes unsafe attempts. The dataset includes 6,690 paraphrases from 75 OMOP-adapted templates with leakage-resistant template/SQL-variation splits. LoRA-tuned Llama-3-8B-Instruct achieves 93.55% execution accuracy with improved abstention reliability, while schema-injected baselines fail the contract. We release the dataset, splits, database dump, and a reproducible evaluation pipeline to support reliable clinical analytics assistants.
Profiling Hallucinations in Frontier LLMs for Entity Linking to Medical Ontologies
Logan Born | Nishant Kambhatla | Uliyana Kubasova | Maryam Siahbani | Andrei Vacariu | Timothy W. O’Connell | Anoop Sarkar
Logan Born | Nishant Kambhatla | Uliyana Kubasova | Maryam Siahbani | Andrei Vacariu | Timothy W. O’Connell | Anoop Sarkar
The integration of Large Language Models (LLMs) into healthcare promises to revolutionize clinical documentation and interoperability, yet reliability remains a concern. This study presents a comprehensive analysis of hallucinations by frontier LLMs tasked with mapping clinical text to SNOMED CT. Through rigorous experimentation, we identify a critical reliability gap: LLMs hallucinate medical codes at a rate that currently renders them unsuitable for autonomous clinical coding applications. Paradoxically, constraining models to use ground-truth mention spans exacerbates, rather than mitigates, these hallucinations. We further contribute a taxonomy of hallucination types – including deprecated codes and cross-ontology errors – and demonstrate that general-purpose LLMs significantly underperform compared to specialized zero-shot entity linking approaches. These findings underscore the need for robust verification mechanisms before clinical deployment.
up
Proceedings of the 15th Workshop on Cognitive Modeling and Computational Linguistics
Proceedings of the 15th Workshop on Cognitive Modeling and Computational Linguistics
Byung-Doh Oh | Tatsuki Kuribayashi | Giulia Rambelli | Ece Takmaz | Philipp Wicke | Jixing Li | Ryo Yoshida
Byung-Doh Oh | Tatsuki Kuribayashi | Giulia Rambelli | Ece Takmaz | Philipp Wicke | Jixing Li | Ryo Yoshida
Transformer Attention as a Unified Model of Encoding and Retrieval in Human Sentence Processing
Dan Parker
Dan Parker
ChineseDevBench: A Chinese Developmental Benchmark for Language Development
Shaonan Wang | Yiwen Wu | Na Li | Zesheng Chen | Gan Wang | Shuchen Zhang | Xin Sun | Luan Li | Yaran Chen
Shaonan Wang | Yiwen Wu | Na Li | Zesheng Chen | Gan Wang | Shuchen Zhang | Xin Sun | Luan Li | Yaran Chen
Predicting Sentence Acceptability Judgments in Multimodal Contexts
Hyewon Jang | Nikolai Ilinykh | Sharid Loaiciga | Jey Han Lau | Shalom Lappin
Hyewon Jang | Nikolai Ilinykh | Sharid Loaiciga | Jey Han Lau | Shalom Lappin
What Kind of Language is Easy to Language-Model Under Curriculum Learning?
Nadine El-Naggar | Tatsuki Kuribayashi | Ted Briscoe
Nadine El-Naggar | Tatsuki Kuribayashi | Ted Briscoe
Modeling semantic association in self-paced reading with language model embeddings
Sara Møller Østergaard | Kenneth Enevoldsen | Afra Alishahi | Bruno Nicenboim
Sara Møller Østergaard | Kenneth Enevoldsen | Afra Alishahi | Bruno Nicenboim
Character-aware Transformers Learn an Irregular Morphological Pattern Yet None Generalize Like Humans
Akhilesh Kakolu Ramarao | Kevin Tang | Dinah Baer-Henney
Akhilesh Kakolu Ramarao | Kevin Tang | Dinah Baer-Henney
Comparing Transformer Model Interpretability with Human Cognition: A Dual Analysis of Attention and Attribution
Lingchen Kong | Jinnie Shin | Pavlo Antonenko
Lingchen Kong | Jinnie Shin | Pavlo Antonenko
Correlating Language Model Surprisal With Cloze and Plausibility: Getting the Best of Both Measures
Kate Rebecca Belcher | Matthew Crocker
Kate Rebecca Belcher | Matthew Crocker
Is Cross-Lingual Transfer in Bilingual Models Human-Like? A Study with Overlapping Word Forms in Dutch and English
Iza Škrjanec | Irene Elisabeth Winther | Vera Demberg | Stefan L. Frank
Iza Škrjanec | Irene Elisabeth Winther | Vera Demberg | Stefan L. Frank
Brain-to-Text Decoding with Brain Atlases and Brain Foundation Models
Haruka Akama | Ryo Yoshida | Max Müller-Eberstein | Yohei Oseki
Haruka Akama | Ryo Yoshida | Max Müller-Eberstein | Yohei Oseki
Floating or Suggesting Ideas? A Large-Scale Contrastive Analysis of Metaphorical and Literal Verb–Object Constructions
Prisca Piccirilli | Alexander Fraser | Sabine Schulte im Walde
Prisca Piccirilli | Alexander Fraser | Sabine Schulte im Walde
Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting
Roland Mühlenbernd
Roland Mühlenbernd
Headlines You Won’t Forget: Can Pronoun Insertion Increase Memorability?
Selina Meyer | Magdalena Abel | Michael Roth
Selina Meyer | Magdalena Abel | Michael Roth
Aggregated Transformer Attention Measures Predict Reading Times Beyond Surprisal
Lukas Mielczarek | Laura Kallmeyer
Lukas Mielczarek | Laura Kallmeyer
’Layer su Layer’: Identifying and Disambiguating the Italian NPN Construction in BERT’s family
Greta Gorzoni | Ludovica Pannitto | Francesca Masini
Greta Gorzoni | Ludovica Pannitto | Francesca Masini
Teaching LLMs to unveil tendentious implicit contents of Italian political communication
Walter Paci | Lorenzo Gregori | Alessandro Panunzi
Walter Paci | Lorenzo Gregori | Alessandro Panunzi
Examining Algebraic Recombination for Compositional Generalisation
Joaquin Cardona Ruiz | Antske Fokkens | Lucia Donatelli
Joaquin Cardona Ruiz | Antske Fokkens | Lucia Donatelli
Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task
Mengyang Qiu | Zoe Brisebois | Siena Sun
Mengyang Qiu | Zoe Brisebois | Siena Sun
Translation from the Information Bottleneck Perspective: an Efficiency Analysis of Spatial Prepositions in Bitexts
Antoine Taroni | Ludovic Moncla | Frederique Laforest
Antoine Taroni | Ludovic Moncla | Frederique Laforest
Causal Inferences Are Driven by Noun Concept Specificity Evidence from Self-Paced Reading and Large Language Model Surprisal
Fabian Schlotterbeck | Raphael Barth
Fabian Schlotterbeck | Raphael Barth
Topic-context Dependency on Continuous Semantic Reconstruction of Language from fMRI Signals
Fermin Travi | Agustin Delmagro | Diego Fernandez Slezak | Bruno Bianchi | Juan E. Kamienkowski
Fermin Travi | Agustin Delmagro | Diego Fernandez Slezak | Bruno Bianchi | Juan E. Kamienkowski
Disentanglement and Compositionality of Letter Identity and Letter Position in Variational Auto-Encoder Vision Models
Bruno Bianchi | Aakash Agrawal | Stanislas Dehaene | Emmanuel Chemla | Yair Lakretz
Bruno Bianchi | Aakash Agrawal | Stanislas Dehaene | Emmanuel Chemla | Yair Lakretz
up
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Piotr Bański | Dawn Knight | Marc Kupietz | Andreas Witt | Alina Wróblewska
Piotr Bański | Dawn Knight | Marc Kupietz | Andreas Witt | Alina Wróblewska
TestiMole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996–2024) for Language Modeling and Sociolinguistic Research
Matteo Rinaldi | Rossella Varvara | Viviana Patti
Matteo Rinaldi | Rossella Varvara | Viviana Patti
We present TestiMole-Conversational a massive collection of discussion boards messages in the Italian language. The large size of the corpus, almost 30B word-tokens (1996–2024), brings challenges in the processing and curation of the resource, but it renders it an ideal dataset for native Italian Large Language Models’ pre-training. Furthermore, discussion boards’ messages are a relevant resource for linguistic as well as sociological analysis. The corpus captures a rich variety of computer-mediated communication, offering insights into informal written Italian, discourse dynamics, and online social interaction in a wide time span. Beyond its relevance for NLP applications such as language modelling, domain adaptation, and conversational analysis, it also support investigations of language variation and social phenomena in digital communication.
A Large Dataset Representing Bulgarian, with the Bulgarian National Corpus as Its Core
Svetla Peneva Koeva | Ivelina Stoyanova
Svetla Peneva Koeva | Ivelina Stoyanova
The paper introduces the IfGPT dataset, which integrates several Bulgarian text collections, including the Bulgarian National Corpus, and applies cleaning, deduplication, and LLM-oriented metadata such as personally identifiable information and bias scores. The composition of the IfGPT dataset is presented, along with the unified metadata schema and metadata management in a graph database, enabling efficient querying and document selection for specific tasks. The main contributions are the integration of multiple Bulgarian text collections into a unified dataset, the development of a standardised metadata schema with graph-based organisation, and the provision of efficient metadata querying mechanisms to support LLM development.
Merimënga: A Manifest-First Pipeline for Reproducible Albanian Web Corpus Construction
Besim Kabashi | Michael Ruppert
Besim Kabashi | Michael Ruppert
We present Merimënga, a pipeline for reproducible Albanian web-corpus construction from Common Crawl. Rather than distributing a static text dump, we publish versioned manifests and append-only JSONL ledgers that make every retrieval and filtering decision replayable at record level. Records are addressed by (WARC filename, byte offset, byte length) and retrieved via HTTP range requests with checksum validation, enabling selective download, resumability, and exact re-materialization. On top of deterministic cleaning and deduplication, Merimënga supports teacher–student filtering: a large LLM labels a stratified sample; the resulting policy is distilled into a faster student model applied at corpus scale. The paper contributes (i) a reproducibility specification for web-corpus construction based on coordinate-addressed retrieval and decision ledgers, (ii) a concrete instantiation for Albanian with language-specific filtering, and (iii) an evaluation protocol for rerun equivalence and filter-stack ablation. Large-scale download and full-corpus filtering are ongoing; this submission focuses on methodology and auditable artifacts rather than final corpus statistics.
Pop Lyrics through Time: Challenges in Corpus-Based Modeling of Linguistic and Emotional Dynamics in German Pop Lyrics
Roman Schneider
Roman Schneider
This paper presents a large-scale diachronic analysis of German pop lyrics based on a linguistically rich, TEI-encoded monitoring corpus. We describe multi-layer annotation and reproducible workflows for deriving higher-level features at scale, including lexical diversity indices, a pronoun-based subjectivity measure, modal particle density, and a length-normalized sentiment intensity score. Particular attention is paid to the development and evaluation of pipelines for two notoriously challenging phenomena: modal particles and sentiment. For modal particles, we build a manually curated gold standard and train sequence models whose performance we relate to inter-annotator agreement. For sentiment, we integrate a lexicon-based resource with a dedicated human annotation experiment to assess reliability and alignment with expert judgments. On this basis, we investigate how structural and affective features co-vary in the corpus and how they change over time, showing, among other trends, declining lexical diversity and sentiment intensity alongside a slight increase in first- and second-person pronouns. Beyond the empirical findings, the paper highlights practical challenges in managing culturally specific corpora, and makes evaluation materials available to support transparent, reusable corpus-based research on popular music and related domains.
The rapid advancement of digital humanities and Natural Language Processing (NLP) necessitates centralized access to high-quality, large-scale language resources. This paper presents the technical infrastructure and evolving ecosystem of Korpuss.lv, the central access platform for the Latvian National Corpora Collection (LNCC). The LNCC consolidates 42 corpora developed by 14 institutions, comprising 2.8 billion tokens of written and spoken Latvian across diverse genres and annotation layers. Korpuss.lv has evolved from a simple metadata index into a comprehensive digital infrastructure that enhances corpus discoverability, accessibility, and usability for researchers in linguistics, digital humanities, and natural language processing. The platform integrates noSketchEngine as its primary corpus analysis tool and extends its functionality with custom modules, including a metadata-driven Corpora Explorer, a client-side Federated Content Search system, and precomputed UD-based Word Sketches. The ecosystem is further supported by CLARIN DSpace repositories for persistent storage and citation management, as well as a federated academic authentication architecture built on SATOSA and Keycloak via the CLARIN Service Provider Federation. The paper outlines architectural decisions, integration strategies, and future development plans.
Optimized for AI: Curating the Icelandic Gigaword Corpus for Stable LLM Training
Jón Friðrik Daðason | Steinþór Steingrímsson
Jón Friðrik Daðason | Steinþór Steingrímsson
The Icelandic Gigaword Corpus (IGC) is a primary resource for Icelandic NLP, with its current version containing 2.7 billion words of curated text. The IGC is traditionally distributed in a TEI-XML format, a hierarchical structure that allows for rich linguistic annotation and metadata. However, this format introduces significant friction for modern machine learning workflows. Even high-quality curated corpora have been found to contain “unwanted” text sequences – such as fragmented lists or repetitive boilerplate that may trigger instabilities during training of large language models. In this paper, we present a new processing pipeline designed to optimize the IGC for AI development. We describe a filtering approach focusing on training stability, including fuzzy deduplication to reduce the risk of data leakage, with the aim to provide high-quality data for stable model convergence. Furthermore, we introduce a new JSONL distribution format that bridges the gap between TEI-XML and machine-actionable data, facilitating easier access and safer training for models aiming to work with Icelandic.
The Hellenic National Corpus (HNC) is an integrated online environment offering access to standard Modern Greek language material and to related analysis tools. The HNC corpus has been developed in two main phases, and currently comprises over 97 million words exclusively of written language, sourced from printed resources or scraped from the internet. The material has been automatically lemmatized and morphologically annotated, while a subset of 100,000 words has been further manually corrected, in order to produce a freely downloadable error-free corpus. Through the dedicated platform, the users have access to concordances, morphological analysis of words and statistical information (frequency) at word, lemma, part of speech and n-gram levels. Future steps include the expansion of the material in both historical and coverage dimensions: the inclusion of material from older phases of the language is foreseen, as well as the addition of dialectal material besides standards language.
Corpas Náisiúnta Na Gaeilge 2022-2029: A Project Overview
Mícheál J. Ó Meachair | Úna Bhreathnach | Kevin Scannell | Michal Mechura | Brian Ó Raghallaigh | Gearóid Ó Cleircín
Mícheál J. Ó Meachair | Úna Bhreathnach | Kevin Scannell | Michal Mechura | Brian Ó Raghallaigh | Gearóid Ó Cleircín
This paper reports the latest developments, planned works, and issues of the Corpas Náisiúnta na Gaeilge (henceforth: CNG, translation: the National Corpus of Irish) project, detailing the work that has been completed to date, current work, and planned future work. This report details the compilation of corpora, development of a project website and part-speech tagger, the challenges of expanding existing corpora, and the addition of historical and legal corpora. We also present the training and outreach activities of the project.
General Regionally Annotated Corpus of Ukrainian: Recent Developments and Future Plans
Maria Shvedova
Maria Shvedova
The General Regionally Annotated Corpus of Ukrainian (GRAC) effectively serves as a national corpus. GRAC v.19 (2025) contains 2 billion tokens from over 800,000 texts (1816–2025). The corpus has multi-level annotations: rich metadata including regional tags, morphological annotation based on the VESUM dictionary, and partial semantic annotation. GRAC is the source of several derivative projects, including UD_Ukrainian_ParlaMint, ParaRook parallel corpora, Rada_Trees, and others.
We present recent developments in the Bulgarian National Corpus, including data collection from various sources, cleaning of diverse datasets, enrichment with multimodal data, and extensive metadata, which resulted in the development of IfGPT, a large BulNC-based dataset. Typical methods for distributing the BulNC-based dataset are briefly described, with emphasis on effective searching within the metadata stored in a graph database.
The British National Corpus (BNC) is a 100 million word collection of samples of written and spoken language from a wide range of sources, designed to represent a wide cross-section of British English from the later part of the 20th century, both spoken and written. It is one of the first generation of monolingual, synchronic, general, representative corpora of its size, and led the way for other national corpora. It was created by a consortium of academic partners and publishers, with funding from the Department of Trade and Industry in the UK. This posters reflects on a number of lessons learning in more than thirty years, in terms of corpus representativeness, modes of access to the corpus, licensing, and managing the transition from a contemporary synchronic corpus to a historical corpus.
The Corpus of Contemporary Polish: 2011-2020 Decade and Beyond
Witold Kieraś | Małgorzata Marciniak | Katarzyna Krasnowska-Kieraś | Marcin Woliński
Witold Kieraś | Małgorzata Marciniak | Katarzyna Krasnowska-Kieraś | Marcin Woliński
It has been thirteen years since the release of the current version (v3) of the Croatian National Corpus (HNK). In terms of synchronicity in corpus linguistics, that many years may be considered quite some time. The preparatory phase for the composition of the new version of HNK (v4) has been going already for several years and in this paper we touch on several issues of concern. Apart of regular corpus parameters, e.g. text sources, text genres, coverage of language varieties, time span, we also discuss about metadata and linguistic annotation schemata. One of important technical prerequisites was the development of CorpRepo, a custom corpus data management system and file system, which enable us to do sustainable long-term maintenance of the data, and to produce newer versions of corpus more easily and more often. The selection of IPR-cleared data entails some restrictions and we give several examples of that kind of textual sources, but also discuss possible weaknesses of such approach to data selection. Regarding the linguistic annotation, the important shift is the decision to abandon the MulText East morphosyntacting descriptions and use solutions recommended by UD-initiative.
Managing Growth in a National Corpus: The Hungarian National Corpus 3.0 (MNSZ3)
Noémi Ligeti-Nagy | Enikő Héja | Ágnes Bánfi | Flóra Földesi | Bence Sárossy | Boglárka Skrabák | Tamás Váradi | Gábor Prószéky
Noémi Ligeti-Nagy | Enikő Héja | Ágnes Bánfi | Flóra Földesi | Bence Sárossy | Boglárka Skrabák | Tamás Váradi | Gábor Prószéky
The third generation of the Hungarian National Corpus (MNSZ3) aims to provide a large-scale, curated, and well-described corpus resource needed for the sustainable digital presence of Hungarian. Building on the domain structure and proportions of MNSZ2 (v2.0.5; 1.04 billion running words), the project targets a substantial increase in scale while also strengthening the coverage and metadata description of Hungarian language use outside Hungary. MNSZ3 retains the six traditional domains of the earlier corpus—press, fiction, scientific, official, personal, and transcribed spoken language—and is planned to reach approximately 10 billion tokens. This paper presents the motivation and design principles of the project, outlines the practical decisions and procedures used in data collection and cleaning, and discusses the annotation strategy developed for large-scale processing. In planning the linguistic analysis, we build on the complementary strengths of HuSpaCy and e-magyar: HuSpaCy provides the unified and efficient UD-oriented processing backbone, while e-magyar (emMorph) is preserved as an explicit additional layer for morphology and lemmatisation.
CoRoLa Version 2.0: Corpus Enrichment and a New Annotation Level
Elena Irimia | Verginica Barbu Mititelu | Radu Ion | Vasile Pais | Maria Mitrofan | Dan Ioan Tufis
Elena Irimia | Verginica Barbu Mititelu | Radu Ion | Vasile Pais | Maria Mitrofan | Dan Ioan Tufis
The paper gives an overview of the recent developments in the enrichment of the reference Corpus of Contemporary Romanian (CoRoLa), within on-going international projects. Statistics of the newly acquired data, work methodology and work towards inclusion of a new annotation layer, the syntactic one, are detailed. We briefly present RODNA, an updated Romanian text processor with state-of-the-art performance on POS tagging, lemmatization and dependency parsing that will be used to populate the syntactic layer of CoRoLa.
The German Medical Text Corpus: Early 2026 Update
Justin Hofenbitzer | Christina Lohr | Frank Meineke | Markus Löffler | Martin Boeker
Justin Hofenbitzer | Christina Lohr | Frank Meineke | Markus Löffler | Martin Boeker
Clinical text resources are a central component for the study of medical language, as well as the training and evaluation of large language models, chatbots, and artificial intelligence systems supporting clinical routines. With the German Medical Text Corpus (GeMTeX), we are currently working on the largest shareable clinical document dataset in German. The multi-centric project ensures diversity across different university hospitals, clinical domains, and text sorts. After a thorough de-identification process, the clinical texts are semantically annotated using Snomed CT, a language-independent, standardized medical ontology. While the corpus is still under active development, it is accessible upon request under controlled access conditions. As of February 2026, GeMTeX comprises more than 15k documents and 20M tokens. We refer researchers interested in the resource to visit https://kiinformatik.mri.tum.de/en/gemtex or reach out to us via gemtex.mi@mh.tum.de.
From Corpus to Community: New NLP Tools for Welsh Language Research and Learning
Dawn Knight | Fernando Alva-Manchego
Dawn Knight | Fernando Alva-Manchego
Launched in 2020, CorCenCC (Corpws Cenedlaethol Cymraeg Cyfoes – National Corpus of Contemporary Welsh) is the first large-scale corpus of the Welsh language to integrate spoken, written, and electronically mediated data, offering a comprehensive snapshot of contemporary Welsh use. Including contributions from over 2,000 speakers, the 11.2-million-word corpus represents the diversity of Wales’s linguistic landscape. As a national resource, CorCenCC enables users to explore real world Welsh. Several tools and resources were developed through the CorCenCC project, including the CyTag POS tagger and CySemTag (adapted from Lancaster University’s USAS semantic system), to enable the grammatical and semantic categorisation of the dataset. The team also built the pedagogic toolkit Y Tiwtiadur, to allow learners and teachers to access corpus-based examples and tasks. Additionally, Yr Amliadur provides curated frequency-based wordlists across modes and parts of speech, supporting linguistic analysis and vocabulary development. Since completing the corpus, the team has focused on extending its impact and reach, to ensure that the resources are maintained and sustained for future use; a challenge often faced when large-scale projects end. This poster profiles the tools and resources created from and inspired by CorCenCC and its associated tools and resources, as a means of supporting the democratisation of linguistic resources for minoritised language contexts.
Swiss-AL: Language Data Platform for Applied Sciences
Julia Krasselt | Philipp Dreesen | Dolores Lemmenmeier-Batinić | Sooyeon Geckeler | Klaus Rothenhäusler | Matthias Fluor
Julia Krasselt | Philipp Dreesen | Dolores Lemmenmeier-Batinić | Sooyeon Geckeler | Klaus Rothenhäusler | Matthias Fluor
This paper introduces Swiss-AL, a language data platform designed for the multilingual, comparative analysis of public discourse in Switzerland. Swiss-AL is an open research data resource providing browser-based access to a variety of corpora in all four of Switzerland’s official languages. Corpora contain journalistic, organisational, and parliamentary discourse. The platform supports research in applied linguistics as well as neighbouring disciplines (e.g., social sciences, communication and media studies).
EuReCo, KorAP and DeReKo: Updates on Ingestion and Annotation Pipelines, Backend, Interfaces, Operation, and Corpora
Marc Kupietz | Nils Diewald | Harald Lüngen | Eliza Margaretha Illig | Helge Stallkamp | Uyen-Nhu Tran | Rameela Yaddehige
Marc Kupietz | Nils Diewald | Harald Lüngen | Eliza Margaretha Illig | Helge Stallkamp | Uyen-Nhu Tran | Rameela Yaddehige
This paper reports on recent technical developments in the European Reference Corpus EuReCo and its current technical implementation based on the corpus search and analysis platform KorAP. We describe updates to the ingestion pipeline, including extensions to the TEI-to-KorAP-XML converter tei2korapxml and the KorAP tokenizer, as well as the newly introduced korapxmltool for annotation and index conversion. We further present Koral-Mapper, a service that enables cross-schema comparability of annotations and metadata at query time, and report on developments in the backend access control system Kustvakt, the web user interface Kalamar, API client libraries for R and Python that promote reproducibility and methodologically sound AI-assisted analysis, and containerized deployment. The corpora and languages currently represented in EuReCo are outlined, and the role of the German Reference Corpus DeReKo, including its metadata-driven virtual corpus design, predefined useful subcorpora, and TEI encoding, is discussed in detail. We further present the National Libraries as Corpus approach and DeLiKo-2025@DNB as its first full-scale proof of concept, and discuss the potential of this approach for extending EuReCo with comparable contemporary fiction corpora across European countries.
up
Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages
Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages
Atul Kr. Ojha | Sakriani Sakti | Claudia Soria | Maite Melero | John P. McCrae | Constantine Lignos | Chao-Hong Liu | German Rigau Claramunt | Georg Rehm
Atul Kr. Ojha | Sakriani Sakti | Claudia Soria | Maite Melero | John P. McCrae | Constantine Lignos | Chao-Hong Liu | German Rigau Claramunt | Georg Rehm
How Well Do Large Language Models Reason in Under-Resourced Languages? Evidence from Vietnamese
Tuan Anh Do | Jelke Bloem
Tuan Anh Do | Jelke Bloem
Despite advancements in Large Language Models, reasoning benchmarks remain centered on high-resource languages, leaving languages like Vietnamese under-evaluated. In this study, we aim to address this gap by evaluating four models: PhoGPT (native), Vistral and VBD-Llama (adapted), and Llama-2 (English-centric), on commonsense reasoning and arithmetic reasoning. As Vietnamese benchmarks for these tasks are lacking, we adapt two analogy datasets from English to Vietnamese and construct two sequence datasets, ensuring a range of structural complexity and difficulty levels. We evaluate diverse prompting strategies, including Chain-of-Thought, role-playing guidance, cross-lingual prompting, and few-shot learning. Our results reveal a baseline proficiency in analogical and arithmetic reasoning among the models, with Vistral and Llama-2 outperforming other models in multiple tasks. The effects of Chain-of-Thought and contextual guidance are limited in Vietnamese, while cross-lingual prompting and few-shot learning show promising performance improvements. The findings underscore the feasibility of adapting benchmarks to less-resourced languages and provide insights into strengths and weaknesses in the performance of Vietnamese LLMs, suggesting directions for model improvements.
Register Sensitivity in Scalar MT Evaluation: Evidence from Spanish–Basque Informal Discourse
Nora Aranberri
Nora Aranberri
Automatic scalar metrics are widely used for machine translation (MT) evaluation, yet their behavior under sociolinguistic variation remains underexplored, particularly in under-resourced and minority-language contexts. We present a small, controlled empirical analysis of reference-based evaluation in Spanish–Basque informal discourse. Register is operationalized as indexical density, capturing dialectal forms, informal lexicon, code-switching and orthographic stylization. Across two MT systems and prompting conditions, sentence-level scores from chrF++, COMET-DA, and XCOMET-XL show a consistent negative association with indexical density under the original informal reference. In a reference-perturbation design that holds MT outputs constant while replacing the informal reference with a standardized Batua version, scores increase systematically, particularly for high-density items, and the density–score association weakens. These results provide controlled evidence that evaluation outcomes in this setting depend in part on reference register configuration. In minority-language and informal domains, reference design choices may influence how translation quality is measured and interpreted.
Corpus-Linguists’ Little Helpers? Evaluating LLMs for Linguistic Annotation: The Case of Sensationalist Headlines Corpus
Petra Bago | Virna Karlić
Petra Bago | Virna Karlić
Manual annotation of pragmastylistic features in sensationalist media is a resource-intensive bottleneck for corpus- based research, particularly for lower-resource languages. This paper evaluates whether Large Language Models (LLMs) can reliably automate this process. We benchmark two proprietary models, OpenAI’s GPT-5 and Google’s Gemini 2.5 Pro, on annotating eight sensationalist linguistic and orthographic features within a corpus of 508 Serbian celebrity magazine headlines. Our methodology involves a systematic comparison of five prompting strategies: zero-shot, few-shot (1, 3, and 5 examples), and chain-of-thought. Results demonstrate that LLMs can achieve high alignment with a manually curated gold standard, reaching a peak macro-F1 score of 98.76%. Notably, the most effective and cost-efficient configuration was GPT-5 using a simple zero-shot prompt. Qualitative error analysis reveals that remaining inaccuracies are systematic, primarily involving pragmatic conventions, discourse scope, and quoted speech. We conclude that LLMs are viable for first-pass annotation of well-defined features in Serbian, though implicit and genre-dependent cues require further study. To support reproducibility and future research on underrepresented languages, we provide our full prompting setup, evaluation procedures, and a detailed cost comparison.
LLM as a Morphological Disambiguator for Belarusian: A Preliminary Study
Vladislav Poritski | Oksana Volchek | Ilia Afanasev
Vladislav Poritski | Oksana Volchek | Ilia Afanasev
We explore the use of large language models (LLMs) for morphological disambiguation in Belarusian, a low-resource language. The pipeline has two stages: a rule-based analyzer generates candidate lemmas and grammatical tags, which an LLM then disambiguates in context. Initial evaluation of ChatGPT, Claude, and Gemini on a gold-standard sample shows high accuracy. We scale this approach to a 375K-word corpus using Gemini and compare the results against a neural baseline (Stanza). Manual review of discrepancies suggests that the LLM-based approach outperforms the baseline, offering a solution for corpus annotation in Belarusian.
We present mobile and desktop keyboards for Idu Mishmi, an endangered Trans-Himalayan language spoken by approximately 11,000 people in Arunachal Pradesh, India. A Latin-based orthography, the Idu Azobra, was developed in 2018, but no digital input tools existed to use it. The orthography requires characters absent from standard keyboards, including schwa (ə), retracted vowels (ə̱, o̱, u̱), nasalized vowels, and accented forms, several of which involve multi-codepoint Unicode sequences that default keyboards do not support. Developed with the Idu Mishmi community, the keyboards comprise: (1) an Android mobile keyboard, published on the Google Play Store, and (2) a Windows desktop keyboard distributed as a single portable executable. Both tools support the complete character inventory, and operate fully offline with zero network permissions. The Android keyboard has been adopted by community leaders and teachers who currently know and actively use the Idu Azobra orthography. The Windows keyboard is currently undergoing testing with community leaders. We describe the design, implementation, and deployment as a replicable model for other endangered language communities.
SAINT: Multilingual Span-Level Interpretability for Sentiment Analysis
Seid Muhie Yimam | Tadesse Destaw Belay | Robert Geislinger | Shamsuddeen Hassan Muhammad | Adaeze Ngozi Ohuoba | Sukairaj Hafiz Imam | Abinew Ali Ayele | Martin Semmann | Chris Biemann | Serge Sharoff
Seid Muhie Yimam | Tadesse Destaw Belay | Robert Geislinger | Shamsuddeen Hassan Muhammad | Adaeze Ngozi Ohuoba | Sukairaj Hafiz Imam | Abinew Ali Ayele | Martin Semmann | Chris Biemann | Serge Sharoff
We investigate multilingual sentiment analysis and interpretability across high- and low-resource languages, focusing on Amharic, English, German, and Hausa. Our study evaluates encoder-only transformer models for both sequence-level sentiment classification and token-level attribution using Captum. Additionally, we assess zero- and few-shot decoder-only models for sequence-level sentiment prediction. Our results show that few-shot decoder-only models outperform encoder-only models on token-level sentiment classification in most languages, with the exception of Hausa, where a multilingual encoder-based model leads. For sequence-level sentiment classification, encoder-only models generally achieve strong performance across most languages, but decoder-only models are highly competitive, and may even surpass encoders, in the high-resource settings (German, English) and low-resource scenarios, depending on the prompting strategy. These findings highlight the utility of combining fine-tuned transformer models with prompt- based large language models to build interpretable sentiment analysis systems across both low- and high-resource languages. The SAINT dataset, annotation guideline, and evaluation scripts can be found at https://github.com/uhh-hcds/SAINT.
AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian
Wajdi Zaghouani | Kholoud Khalil Aldous | Isra Fejzullaj
Wajdi Zaghouani | Kholoud Khalil Aldous | Isra Fejzullaj
Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved. We present AlbanianLLMSafety, the first publicly available safety evaluation dataset for LLMs in Albanian, a linguistically distinct low-resource language with approximately 7.5 million speakers across Albania, Kosovo, North Macedonia, and the diaspora. The dataset contains 2,951 prompts spanning 11 safety categories, including self-harm, violence, racist content, child exploitation, and radicalization, with an average of 268 prompts per category. Each prompt is provided in Albanian with an English reference translation and a detailed category label. This resource addresses a significant gap in safety evaluation infrastructure for low-resource languages and provides an essential benchmark for developing safer, more inclusive LLMs. The dataset will be provided upon request to support safety evaluation, fine-tuning, red-teaming, and guardrail development for Albanian-speaking communities.
Urdu-CLEVR: A Novel Benchmark for Visual Reasoning in an Under-Resourced Linguistic Context
Sohail Ashraf | Adeel Zafar | Slawomir Nowaczyk | Ahthasham Sajid
Sohail Ashraf | Adeel Zafar | Slawomir Nowaczyk | Ahthasham Sajid
Visual Question Answering (VQA) bridges the gap between computer vision and natural language processing, yet progress remains largely confined to high-resource languages. For low-resource languages like Urdu, research is severely hindered by the total absence of large-scale reasoning-based datasets. To address this critical gap, we introduce the first synthetic Urdu VQA dataset modeled after the CLEVR framework, specifically designed to evaluate complex, multi-step visual reasoning. We conduct a rigorous comparative analysis using both transformer-based architectures (VisualBERT, LXMERT, ViLT) and neuro-symbolic models. Our results demonstrate that the neuro-symbolic approach achieves a superior accuracy of 85.3%, outperforming the strongest transformer baseline by 7.1% while maintaining competitive processing efficiency. This work establishes a primary benchmark for Urdu VQA, demonstrating that hybrid reasoning architectures provide a robust and scalable solution for advancing multimodal AI in under-resourced linguistic contexts.
A Database of Romance Clitics With Speech Samples
Abdelrahim Qaddoumi | Owen Rambow | Lori Repetti | Francisco Ordóñez
Abdelrahim Qaddoumi | Owen Rambow | Lori Repetti | Francisco Ordóñez
We present a new database of Romance clitics across nine varieties. The database includes speech, transcriptions, and linguistic annotations. The database concentrates on clitics, and includes varieties of Romance with stressed clitics. Specifically, the database includes data from the regions of Corsica, Pyrénées-Atlantiques, Sardinia, Liguria, Basilicata, Campania, Mallorca, Menorca, and Formentera. A publicly accessible interface allows easy searching. The database and, separately, the interface code will be made publicly available.
GreekCommonGen: A Benchmark for Evaluating Generative Commonsense Reasoning in Greek
Aristotelis Stamopoulos | Dimitrios Galanis
Aristotelis Stamopoulos | Dimitrios Galanis
This paper introduces GreekCommonGen, the first benchmark designed for generative commonsense reasoning in Greek. The dataset is created by automatically translating the original English CommonGen corpus and subsequently refining the outputs through manual post-editing to ensure linguistic and semantic quality. We conduct a comprehensive evaluation of a range of approaches/models on this benchmark, exploring the impact of different prompting strategies, decoding methods, and model sizes/architectures. Our findings provide valuable insights into the challenges of commonsense generation in Greek, paving the way for future research in the field.
Transfer Learning for Creole TTS: A Pilot Study on Whether Substrate Phonologies or Lexifier Vocabularies Matter More
Emmett Strickland | Marc Evrard | Valentina Fedchenko
Emmett Strickland | Marc Evrard | Valentina Fedchenko
In this early-stage study, we investigate whether transfer learning from lexifier or substrate languages can improve text-to-speech (TTS) performance for low-resource creoles. We conducted a controlled experiment using two creoles of distinct lexical origins: Nigerian Pidgin (English-based) and Guadeloupean Creole (French-based). Single-speaker TTS datasets of approximately 30 minutes each were recorded and used to fine-tune pretrained models for English, French, and Yoruba. Objective metrics and informal subjective evaluations were employed to assess synthesis quality. Though partially inconclusive, our results suggest that the French-based models outperform others for both creoles, while Yoruba-based models yield weaker performance. These findings may suggest that lexical similarity or historic influences alone do not fully predict transfer learning effectiveness, and that phonotactic compatibility and orthographic depth may also be relevant factors. Our work provides insight into TTS model development for creoles and other low-resource languages, and highlights avenues for further research on leveraging relevant linguistic and orthographic features for model development.
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models
Wajdi Zaghouani | Shimaa Amer Ibrahim | Aruzhan Muratbek | Olzhasbek Zhakenov | Adiya Akhmetzhanova
Wajdi Zaghouani | Shimaa Amer Ibrahim | Aruzhan Muratbek | Olzhasbek Zhakenov | Adiya Akhmetzhanova
Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation across eleven categories covering common risk areas such as self-harm, violence, child exploitation, sexual content, racist content, radicalization, and regulated goods or illegal activities. The dataset contains 5,717 prompts written natively in Kazakh (Cyrillic), organized by category, with English translations for cross-lingual analysis. Prompts resemble realistic user queries, often in a teen or child style, and are phrased as intent prompts without procedural instructions. We document the writing protocol, labeling procedures (including borderline-case decision rules), and quality-control steps (schema standardization, completeness checks, and deduplication). We also align the categories with widely used safety taxonomies to support integration with existing evaluation pipelines. Baseline results with GPT-4o show an overall refusal rate of 28.2%, varying from 5.5% to 53.8% across categories, indicating that Kazakh prompts expose category-specific safety gaps not captured by English-only evaluation.
Human judgements of word similarity have been a core benchmark for the intrinsic evaluation of word embedding models, and continue to be used for assessing the capabilities of large language models. While word similarity benchmarks have been collected for a range of languages, none existed for Greek. We develop a Modern Greek variant of the SimLex-999 word similarity dataset by gathering similarity judgements from 90 native speakers of Greek. We then use this as a benchmark for intrinsically evaluating several Greek language models.
Quality and Appropriateness of Large Text Datasets for Irish NLP
Abigail Walsh | Mark Andrade | Jane Lauren Adkins | Ornait O’Connell | Éanna O’Connor | Ellen Rushe | Brian Davis
Abigail Walsh | Mark Andrade | Jane Lauren Adkins | Ornait O’Connell | Éanna O’Connor | Ellen Rushe | Brian Davis
The value of high-quality datasets for training essential language tools has long been recognised for NLP research. Despite the importance of such datasets, most language data available for training consists of large, automatically curated corpora, often scraped from web content. The quality of such datasets is often an unknown factor. This presents a problem for already low-resourced languages (such as Irish), as existing datasets may not provide adequate, representative language data for training effective models. This paper examines existing monolingual and parallel Irish text corpora to evaluate the quality of the language data, through manual review, automatic metrics, and LLMs as judges.
BiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement
Abdullah Al Shafi | Swapnil Kundu Argha | M. A. Moyeen | Abdul Muntakim | Shoumik Barman Polok
Abdullah Al Shafi | Swapnil Kundu Argha | M. A. Moyeen | Abdul Muntakim | Shoumik Barman Polok
High-quality bilingual resources remain a critical bottleneck for advancing multilingual NLP in low-resource settings, particularly for Bangla. To mitigate this gap, we introduce BiST, a rigorously curated Bangla–English corpus for sentence-level grammatical classification, annotated across two fundamental dimensions: syntactic structure (Simple, Complex, Compound, Complex-Compound) and tense (Present, Past, Future). The corpus is compiled from open-licensed encyclopedic sources and naturally composed conversational text, followed by systematic preprocessing and automated language identification, resulting in 30,534 sentences, including 17,465 English and 13,069 Bangla instances. Annotation quality is ensured through a multi-stage framework with three independent annotators and dimension-wise Fleiss’ Kappa (κ) agreement, yielding reliable and reproducible labels with κ values of 0.82 and 0.88 for structural and temporal annotation, respectively. Statistical analyses demonstrate realistic structural and temporal distributions, while baseline evaluations show that dual-encoder architectures leveraging complementary language-specific representations consistently outperform strong multilingual encoders. Beyond benchmarking, BiST provides explicit linguistic supervision that supports grammatical modeling tasks, including controlled text generation, automated feedback generation, and cross-lingual representation learning. The corpus establishes a unified resource for bilingual grammatical modeling and facilitates linguistically grounded multilingual research.
LLM-Assisted Spanish Dialect Corpus Construction
Jessica Claribel RAMIREZ VIDAL | Hiroki Ouchi | Sakriani Sakti
Jessica Claribel RAMIREZ VIDAL | Hiroki Ouchi | Sakriani Sakti
This study presents a multi-dialect, pragmatically annotated Spanish corpus designed to address persistent gaps in the representation of regional varieties and communicative functions in existing linguistic and NLP resources. The corpus focuses exclusively on Spanish dialects spoken in the Americas, selecting one representative dialect per country and incorporating a single neutral Castilian variety for comparative purposes. Dialects are organized into five regional groups: Mexican, Central American, Caribbean, South American, and Rioplatense Spanish. Corpus development follows a multi-stage workflow in which a seed lexicon composed of openly licensed material from sources such as Wikipedia, Project Gutenberg, and curated random and synthetic data is used to initiate the LLM-based text generation. Each base sentence is expanded into dialect-specific variants and annotated with pragmatic and domain labels, producing a fully parallel dataset that supports cross dialect comparison. A multi-stage correction pipeline combining automated scripts, controlled LLM-based editing, and manual review ensures syntactic well-formedness and dialectal authenticity while eliminating language-switching and hallucination errors. The final version of the corpus covers 20 dialects and contains, 40,000 annotated sentences, released in both JSON and plain-text formats for use in a wide range of NLP tasks.
Structured Entity Extraction from Hawaiian Television Chyrons Using Vision-Language Models
Kelley Lynch | Owen King | Kyeongmin Rim | Gabrielle Keen | Yangyang Chen | James Pustejovsky
Kelley Lynch | Owen King | Kyeongmin Rim | Gabrielle Keen | Yangyang Chen | James Pustejovsky
Hawaiian (ʻŌlelo Hawaiʻi) is an endangered Polynesian language whose broadcast archives represent a critical yet underutilized resource for language documentation. We present the first evaluation of vision-language models (VLMs) for structured entity extraction from television chyrons, investigating the performance gap between Hawaiian-language content and mainland U.S. comparisons. Using our new HiChy dataset of 3,925 manually annotated images, we demonstrate that Hawaiian content remains significantly more challenging for current VLMs: for the best-performing model (Qwen2.5-VL-7B), character error rates roughly double from 0.064 on mainland data to 0.130 on Hawaiian content. We extend the task to key information extraction (KIE), finding that while models can perform structured parsing, they struggle specifically with names of Hawaiian linguistic origin, a difficulty that persists even when controlling for geographic source. Across five evaluated models spanning local quantized inference and commercial APIs, we find that OCR accuracy and structured extraction capability do not necessarily correlate: the best OCR model (Gemini 3 Flash) underperforms locally-deployed alternatives on KIE, while even a 2.2B-parameter model (SmolVLM2) achieves functional extraction. Our results provide a baseline for AI-assisted archival processing of underrepresented language media and highlight the need for models that better account for the orthographic and cultural specificities of Hawaiian.
The world’s languages are commonly categorised in terms of available language technologies such as speech recognition and machine translation. On this view, “under-resourced languages” suffer from language barriers which cut people off from markets, healthcare, human rights, and AI. The “solution” is more funding for language technologies, opening the way to a utopia of digital language equality and AI-enabled mobility. Yet the world’s linguistic diversity is not a set of language objects to be pushed up a cline from emerging to thriving. It consists of polyglossic communities who have long used vernaculars for local functions and dominant languages for external functions. I present a new theory of linguistic diversity which places the world of vernaculars alongside the world of institutional languages, and articulates diverse language technology agendas that lie within and between these worlds.
Interlinear Glosses as a Multilingual Pivot for Machine Translation: An Updated Study on Turkish with Restricted Resources
Volkan Ozer | Shu Okabe | Alexander Fraser
Volkan Ozer | Shu Okabe | Alexander Fraser
Translating very low-resource languages is a challenge that has been approached using available linguistic cues. Among them, interlinear glosses are linguistic annotations that can essentially bridge the gap between two languages thanks to both grammatical and lexical information. We perform a case study on a simulated low-resource condition for Turkish, a morphologically rich language, with a pipeline approach, following (Zhou et al., 2020). A source sentence is passed through a morphological analyzer and a bilingual dictionary to obtain a gloss-like representation. We then evaluate the current capacity of Neural Machine Translation systems and Large Language Models in performing the translation task from interlinear glosses into fluent English translations. We notably evaluate how performance scales with multilingual glossed data and how translation is affected by pseudo-glosses. Pivoting with glosses remains a better approach than a direct translation for languages with limited parallel data for training. Although glosses remain helpful resources, translations are sensitive to their quality, especially for lexical information.
Multilingual large language models are known to perform very well on high-resource languages, while their ability to process severely under-resourced languages remains underexplored. We investigate multilingual LLM translation performance on Fuzhounese, an under-resourced Sinitic language without a standardized orthography and almost no digital presence. Having adopted some methodological insights from the HKCanto-Eval benchmark, this paper presents a bidirectional translation framework based on a dataset of 305 sentences (300 constructed English sentences and 5 additional reference translations), that assesses the comprehension and generation of Fuzhounese, evaluated using automatic metrics and human Likert-scale judgments. The results reveal poor performance on Fuzhounese in both translation directions: BERTScore and chrF++ values consistently stay low when models are faced with comprehension tasks, while for generation tasks, scores are generally more than twofold lower than those for Mandarin or Cantonese. These findings highlight structural biases in multilingual LLMs toward high-resource languages and stress the need for resource-aware modeling and evaluation approaches in multilingual NLP systems.
Evaluating Nepali NER and POS Tagging Models on the Achhami Dialect
Samikshya Dhamala | Rishav Beejukchhen | Subresh Thakulla | Bikash Kadayat | Supriya Khadka
Samikshya Dhamala | Rishav Beejukchhen | Subresh Thakulla | Bikash Kadayat | Supriya Khadka
Nepali Natural Language Processing (NLP) models are typically trained and evaluated on Standard Nepali, which can introduce bias against regional dialects. This study investigates the performance of Named Entity Recognition (NER) and Part-of-Speech (POS) tagging models on Achhami, a Far-Western dialect of Nepal. A parallel corpus of 300 sentence pairs was created, covering news, cultural topics, and everyday conversations. Achhami translations were produced by native speakers to preserve linguistic authenticity. The evaluation compared fine-tuned Transformer models with large language models using zero-shot prompting. Across both tasks, all models showed consistent performance degradation on the Achhami dialect. For NER, F1 scores decreased by 2.12 to 3.97 percent. Claude 3.5 Haiku achieved the best NER performance, while the monolingual NepBertA model unexpectedly outperformed multilingual alternatives, challenging assumptions about multilingual advantages. POS tagging results showed a similar pattern, with accuracy dropping notably on Achhami data. Large language models also showed comparable weaknesses, with accuracy reductions ranging from 2.9 to 7.0 percent. These findings quantify Kathmandu-centered bias in Nepali NLP and highlight the importance of dialectally diverse training data for building more inclusive and equitable language technologies.
From LLM Prompts to Acoustic Baselines: A Scalable Pipeline for Under-Resourced Disfluent Code-Mixed Speech
Anuran Mitra | Anirvan Chakravarty | Tapabrata Mondal | Sivaji Bandyopadhyay
Anuran Mitra | Anirvan Chakravarty | Tapabrata Mondal | Sivaji Bandyopadhyay
Spontaneous speech in multilingual communities often involves rapid code-mixing (CM) and natural disfluencies, yet such patterns are rarely reflected in available training data for under-resourced languages. This gap limits the development of robust automatic speech recognition (ASR) systems. To address this, we introduce BEHE-CMDisfl, a fully synthetic Bengali–English and Hindi–English disfluent code-mixed speech corpus generated through a controlled Large Language Model (LLM) and Text-to-Speech (TTS) pipeline. The dataset explicitly incorporates conversational phenomena such as filled pauses, repetitions, and restarts. We evaluate its usefulness under two ASR settings. In a micro-resource scenario (∼1.3 hours), a GMM-HMM Kaldi baseline achieved a 37.74% Word Error Rate (WER) after phonetic normalization to reduce transliteration inconsistencies, and successfully retained disfluency markers in decoding. We also examined adaptation of a modern foundation model. In zero-shot testing, openai/whisper-small failed on the code-mixed speech due to severe hallucinations and looping behavior. After applying parameter-efficient fine-tuning (LoRA) for 1,000 steps, the model stabilized, reduced insertion errors, captured rapid language switching more reliably, and achieved a WER of 21.37%. These findings show that synthetic data combined with efficient fine-tuning offers a practical path for ASR development in complex low-resource disfluent CM settings.
Beyond Fine-Tuning: Procrustes Alignment of Multilingual Embeddings for Low-Resource Cross-Lingual Retrieval
Ali Faheem | Muhammad Hammad | Faizad Ullah | Ahmed Hassan | Fezan Rasool | Asim Karim
Ali Faheem | Muhammad Hammad | Faizad Ullah | Ahmed Hassan | Fezan Rasool | Asim Karim
Multilingual sentence-embedding models are widely used for cross-lingual retrieval; however, their performance drops significantly in low-resource languages. The Urdu language, which is considered a low-resource language by the NL community, poses this challenge, despite being spoken by over 246 million people worldwide. Its distribution in training corpora results in poor alignment with English within shared embedding spaces. To resolve this misalignment without model fine-tuning, we apply Procrustes transformation, which is an orthogonal post-hoc alignment method with a closed-form solution. We utilize SQuAD and UQA datasets to learn a rotation matrix from a small set of sentence pairs and evaluate its effect across five multilingual embedding models (MiniLM, DistilUSE, E5-Base, LaBSE, and E5-Large) and perform geometric alignment, cross-lingual retrieval, and question-answering tasks on these models. We find that cosine distances between parallel pairs decrease by up to 38.67%, and retrieval accuracy improves by 12.49% points in Recall@1. We also analyze that models with better pre-trained cross-lingual representations exhibit a saturation effect, showing minimal retrieval change even as geometric tightening increases. Our error analysis reveals that morphologically complex queries and colloquial expressions remain challenging, indicating representational limitations beyond the scope of a linear transformation. These findings demonstrate that a computationally inexpensive alignment step can meaningfully improve cross-lingual retrieval for low-resource languages, with implications for retrieval-augmented generation (RAG) in resource-constrained settings.
Formosan languages are a critically endangered group of Austronesian languages spoken in Taiwan, with severely limited representation in natural language processing (NLP) research and no support in existing language identification (LID) tools. We present the first systematic evaluation of machine learning models for the language identification of Kavalan, a Formosan language with fewer than 300 known speakers. We construct two benchmarks: a deployment-oriented benchmark with languages commonly confused with Kavalan by existing tools, and a linguistically motivated benchmark of typologically related Formosan languages. We evaluate random forest, support vector machine (SVM), and three pre-trained multilingual models using repeated stratified cross-validation. SVM models with character n-gram features achieve the strongest performance on both benchmarks, with a macro F1 of 0.993 on the deployment benchmark and a macro F1 of 0.906 on the Formosan benchmark, while remaining computationally inexpensive and effective with a low amount of data. Pre-trained multilingual models degrade significantly on the Formosan benchmark, with XLM-RoBERTa falling to a macro F1 of 0.505. These results demonstrate that traditional n-gram-based approaches are effective with low-resource Formosan LID and establish a foundation for downstream NLP tasks supporting the documentation and revitalization of low-resource Formosan languages.
Rebelòt: Datasets and Token-Level Language Identification for Lombard-Italian-English Code-Mixing
Edoardo Signoroni | Emma Bednaříková | Pavel Rychly
Edoardo Signoroni | Emma Bednaříková | Pavel Rychly
Lombard is an endangered and under-resourced Gallo-Italic language variety that exists with Standard Italian. As with other language varieties of Italy, code-switching and code-mixing is common between Lombard and Italian in everyday conversation and with English, online. This linguistic complexity, and the lack of a unified written standard, poses challenges for Natural Language Processing tools. We introduce Rebelòt, a novel multi-domain, token-level annotated dataset for Lombard-Italian-English code-mixing. Furthermore, we develop and evaluate three variants of a token-level Language Identification (LID) tool based on a pre-trained encoder architecture, fine-tuned using both authentic data from our corpus and synthetically generated code-mixed text. Our evaluation demonstrates that the optimal model variant achieves an accuracy of over 0.99 on token-level prediction, and substantially outperforms widely used off-the-shelf LID baselines at sentence-level.
This paper introduces the Spontaneous Persian Speech (SPS) dataset designed for automatic speech recognition (ASR) tasks and a methodology laying the groundwork for addressing the shortage of spontaneous speech data. The corpus aims to support research on natural and conversational Persian, which remains under-represented in current ASR resources. The dataset consists of 694 minutes of audio from a total of 65 speakers, including 34 male and 31 female speakers. It contains 526,585 tokens. The audio segmentation step produces intervals of 1.24 to 3.25 seconds, each containing 3 to 9 words. The recordings cover a variety of environments, from inside cars to homes and shopping areas, including both busy and quiet settings. We use the SPS dataset to fine-tune Whisper and the performance increases significantly for both the small and medium models based on Word Error Rate (WER). This could be an initiative toward building domain-oriented datasets for specific ASR tasks.
Small Language Models for Less-Resourced Languages in a Real-World Scenario: The Case for Catalan
Roser Saurí | Josep Sànchez-Ferreres | Lluis Padro | Josep Carmona
Roser Saurí | Josep Sànchez-Ferreres | Lluis Padro | Josep Carmona
Small Language Models (SLMs), typically ranging from a few million to 10–15 billion parameters, offer a promising solution towards constraints imposed by platform size–particularly mobile and IoT devices–and by the requirements of many organizations such as SMEs, which need solutions that ensure data privacy while remaining cost-effective. Their compactness and efficiency provide digital sovereignty and flexibility, though with more limited general-purpose capabilities. This makes them especially sensitive when working with underrepresented languages, such as Catalan, due to interference from majority languages that can increase bias risk. This paper evaluates state-of-the-art SLMs in a real-world Catalan use case: an AI assistant for older adults, assessing both user interactions and structured function call generation. Our work, which contributes to the Anonymized-Project initiative for deploying a connected SLM-based infrastructure under the Model Context Protocol, demonstrates that some SLMs are able to deliver high-quality performance even in resource-constrained, linguistically minority environments.
AmazoniaNLP: A Survey of Extreme Low-Resource Languages in the Peruvian-Brazilian Amazon
Rodolfo Joel Zevallos | Fabrício Carraro | John E. Ortega
Rodolfo Joel Zevallos | Fabrício Carraro | John E. Ortega
The Amazon basin along the Peru–Brazil border hosts extraordinary linguistic diversity, including many Indigenous languages whose speaker communities span national frontiers. Despite sustained documentation work, most remain extremely low-resource languages (ELRLs) for Natural Language Processing (NLP): reusable corpora are scarce, orthographies vary across countries and institutions, and basic tools such as tokenizers, taggers, and morphological analyzers are largely unavailable. We present a resource-oriented survey of five Indigenous languages of the Western Amazon—Matsés, Amahuaca, Kashinawa, Ticuna, and Kukama-Kukamiria—aimed at supporting more realistic NLP and speech work in extreme low-resource settings. Using a systematic search across academic venues, language archives, and public code/model repositories, we identify and cross-check available materials spanning lexical resources, text corpora, linguistic annotation, and speech collections. For each item we record practical reuse information, including the relevant task or modality, source location, and any stated access, licensing, or usage conditions. Our findings show strong cross-language asymmetries and fragmentation: most materials concentrate in documentation artifacts and lexicons, while standardized datasets with clear access and reuse conditions suitable for training and evaluation remain rare. We conclude with concrete recommendations to improve discoverability, normalize orthographic variation, and prioritize resource creation that maximizes interoperability across tools and benchmarks.
AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages
Lilian Wanzare | Cynthia Jayne Amol | Ezekiel Maina | Nelson Odhiambo | Hope Kerubo | Leila Misula | Vivian Oloo | Rennish Mboya | Edwin Onkoba | Edward Ombui | Joseph Muguro | Ciira wa Maina | Andrew Kipkebut | Alfred Omondi Otom | Ian Ndung’u Kang’ethe | Angela Wambui Kanyi | Brian Gichana Omwenga
Lilian Wanzare | Cynthia Jayne Amol | Ezekiel Maina | Nelson Odhiambo | Hope Kerubo | Leila Misula | Vivian Oloo | Rennish Mboya | Edwin Onkoba | Edward Ombui | Joseph Muguro | Ciira wa Maina | Andrew Kipkebut | Alfred Omondi Otom | Ian Ndung’u Kang’ethe | Angela Wambui Kanyi | Brian Gichana Omwenga
AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, Kalenjin, Maasai, and Somali. The dataset includes 750 hours of scripted speech and 2,250 hours of spontaneous speech, collected from 4,777 native speakers across diverse regions and demographics. This work addresses the critical underrepresentation of African languages in speech technology by providing a high-quality, linguistically diverse resource. Data collection followed a dual methodology: scripted recordings drew from compiled text corpora, translations, and domain-specific generated sentences spanning eleven domains relevant to the Kenyan context, while unscripted speech was elicited through textual and image prompts to capture natural linguistic variation and dialectal nuances. A customized mobile application enabled contributors to record using smartphones. Quality assurance operated at multiple layers, encompassing automated signal-to-noise ratio validation prior to recording and human review for content accuracy. Though the project encountered challenges common to low-resource settings, including unreliable infrastructure, device compatibility issues, and community trust barriers, these were mitigated through local mobilizers, stakeholder partnerships, and adaptive training protocols. AfriVoices-KE provides a foundational resource for developing inclusive automatic speech recognition and text-to-speech systems, while advancing the digital preservation of Kenya’s linguistic heritage.
HuNeBR: A Multitask Benchmark to Evaluate LLMs’ Understanding of Northeastern Brazilian Portuguese Humor
José Matheus do Nascimento Gama | David Candeia Maia | Leandro Balby Marinho | Fabio Morais | João Brunet
José Matheus do Nascimento Gama | David Candeia Maia | Leandro Balby Marinho | Fabio Morais | João Brunet
Humor recognition is a major challenge in Natural Language Processing (NLP) due to its subtle and context-dependent nature. Despite advances, Large Language Models (LLMs) still struggle with this task, especially in Brazilian Portuguese, where no dedicated benchmarks exist. This paper presents HuNeBR, a new benchmark of 475 annotated humorous texts from Northeastern Brazilian comedians. The benchmark evaluates LLMs on three tasks: identifying punchlines, classifying texts into eight comic styles, and explaining humor. This is the first benchmark to evaluate LLMs on the in-depth interpretation of humorous texts in Brazilian Portuguese, going beyond the binary tasks of traditional humor benchmarks. Both general-purpose and Portuguese-specialized LLMs were evaluated under zero-shot and few-shot settings. The findings indicate that LLMs perform very well at identifying punchlines, show inconsistent results in classifying comic styles, and produce humor interpretations that mostly align with human judgments. Among the models assessed, general-purpose multilingual systems like GPT-4 and Gemini 2.5 Flash achieved the top overall performance, whereas Sabiá 3.1, a model specialized in Brazilian Portuguese, demonstrated competitive results across all three tasks, highlighting the value of locally trained models in capturing linguistic and cultural subtleties.
Esperanto is a widespread constructed language, known for its regular grammar and productive word formation. Besides having substantial resources available thanks to its online community, it remains relatively underexplored in the context of modern machine translation (MT) approaches. In this work, we present the first comprehensive evaluation of open-source MT systems for Esperanto, comparing rule-based systems, encoder–decoder models, and LLMs across model sizes. We evaluate translation quality across six language directions involving English, Spanish, Catalan, and Esperanto using multiple automatic metrics as well as human evaluation. Our results show that the NLLB family achieves the best performance in all language pairs, followed closely by our trained compact models and a fine-tuned general-purpose LLM. Human evaluation confirms this trend, with NLLB translations preferred in approximately half of the comparisons, although noticeable errors remain. In line with Esperanto’s tradition of openness and international collaboration, we release our code and best-performing models publicly.
up
Proceedings of The 2nd Workshop on Language-driven Deliberation Technology
Proceedings of The 2nd Workshop on Language-driven Deliberation Technology
Lucas Anastasiou | Katarina Boland | Anna De Liddo | Neele Falk | Annette Hautli-Janisz | Gabriella Lapesa | Julia Romberg
Lucas Anastasiou | Katarina Boland | Anna De Liddo | Neele Falk | Annette Hautli-Janisz | Gabriella Lapesa | Julia Romberg
Using AI to Support Discursive Integration in Online Discussions
Maike Behrendt | Viviana Warnken | Dennis Friess | Marc Ziegele | Tobias Escher
Maike Behrendt | Viviana Warnken | Dennis Friess | Marc Ziegele | Tobias Escher
Online discussions can be rough, especially when it comes to political issues. They are often characterized by a harsh tone which discourages many people from participating in them at all. At the same time, these discussions are very important for democracy as they promote exchange and help individuals form their own opinions. While Artificial Intelligence (AI) may be detrimental to the quality of discussions (e.g. when used in spam bots), it also offers a promising opportunity to support constructive and inclusive discussions, for example by making them more civil. To strengthen such discursive integration we have engaged in a co-creation process with non-academic stakeholders to develop a discussion assistant prototype that i) identifies likely problematic comments for a possible rephrasing and ii) offers authors help with reformulation by letting generative AI suggest improvements like more civil wording. In this paper, we describe the process of co-creative research and the current status of the discussion assistant, which is still being developed and improved.
Accountable Human-AI Deliberation with LLMs: Scaling Collective Intelligence through Symbiotic Scaffolding
Wajdi Zaghouani
Wajdi Zaghouani
Large language models (LLMs) can support democratic deliberation at scales previously constrained by turn-taking and facilitation bandwidth. Recent work shows that LLM-generated group statements are often preferred over human-mediated outputs, while theoretical analyses argue that LLMs relax the simultaneity constraints limiting collective intelligence. Yet pure LLM mediation risks collapsing pluralism, over-optimizing for agreement, and undermining legitimacy when participants cannot contest how they are represented. We propose a symbiotic human-AI framework organized into three layers: observation and diversity amplification, facilitation with clause-level provenance, and human primacy for ratification. Our contributions include graded coverage, diversity, and erasure metrics with salience-aware weighting; a provenance pipeline combining cross-encoder similarity with causal knockout diagnostics; preference-conditioned trade-off control; equity-aware contestability workflows; adversarial robustness tests; and an evaluation protocol with ablation designs informed by evidence of LLM-as-judge limitations. The result is a testable blueprint for deliberation technology that scales collective intelligence while preserving agency and legitimacy.
InFACT: Benchmarking LLM Explanations Against Institutional Reasoning for Deliberation-Aware Fact-Checking
Diana Constantina Hoefels
Diana Constantina Hoefels
Explainability in deliberation-support NLP is usually evaluated through post-hoc rationales or model-internal attribution methods, and only rarely against explicit institutional reasoning procedures. We introduce , a Romanian corpus of professional fact-checking reports that preserves the workflow of editorial epistemic arbitration, namely claim articulation, contextualisation, verification scope, evidence-based verification narrative, and calibrated conclusion. contains 789 raw reports from factual.ro and a processed benchmark release of 788 instances after removal of a singleton non-standard verdict label. Beyond six-way verdict prediction, we position as a benchmark for LLM explanation alignment, where models must generate short explanations that can be compared directly to gold institutional reasoning. We evaluate primarily with instruction-tuned LLMs, reporting full-corpus experiments for open-weight models and a matched pilot comparison with GPT-4 Turbo. The resulting evidence shows that verdict prediction and institutional explanation alignment are not the same capability: models that improve verdict accuracy do not necessarily preserve institutional calibration or produce explanations that align with professional verification narratives. These results support the central claim of the paper, namely that measures not only whether a model reaches a verdict, but also whether it does so in a manner that resembles documented public reasoning.
Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
Panatchakorn Anantaprayoon | Nataliia Babina | Nima Asgharbeygi | Jad Tarifi
Panatchakorn Anantaprayoon | Nataliia Babina | Nima Asgharbeygi | Jad Tarifi
LLM alignment has progressed in single-agent settings through paradigms such as RL with human feedback (RLHF), while recent work explores scalable alternatives such as RL with AI feedback (RLAIF) and dynamic alignment objectives. However, these approaches remain limited in multi-stakeholder settings, where conflicting values arise and deliberative negotiation is required. This work proposes a multi-agent negotiation-based alignment framework that aligns LLMs to Collective Agency (CA)—an existing alignment objective introduced to promote the continual expansion of agency—while simultaneously improving conflict-resolution capability. To enable scalable training, two self-play LLM instances are assigned opposing personas and engage in turn-based dialogue to synthesize mutually beneficial solutions. We generate synthetic moral-dilemma prompts and conflicting persona pairs, and optimize the policy via RLAIF using Group Relative Policy Optimization (GRPO) with an external LLM reward model. While rewards are computed from CA scores assigned to the final completion, gradients are applied to dialogue tokens to directly improve deliberative interaction dynamics. Experiments show that the model achieves CA alignment comparable to a single-agent baseline while substantially improving conflict-resolution performance without degrading general language capabilities. These results suggest that negotiation-driven deliberation training provides a practical path toward LLMs that better support collective decision-making in value-conflict scenarios.
up
Proceedings of the 2nd Workshop on Evaluating Text Difficulty in a Multilingual Context (DeTermIt! 2026)
Proceedings of the 2nd Workshop on Evaluating Text Difficulty in a Multilingual Context (DeTermIt! 2026)
Giorgio Maria Di Nunzio | Federica Vezzani | Liana Ermakova | Hosein Azarbonyad | Jaap Kamps
Giorgio Maria Di Nunzio | Federica Vezzani | Liana Ermakova | Hosein Azarbonyad | Jaap Kamps
Cross-linguistic Readability and Controllable Difficulty: A Corpus-Based Comparison of Human and LLM Translations of Children’s Literature in Romanian
Karla Csuros | Madalina Chitez | Roxana Rogobete
Karla Csuros | Madalina Chitez | Roxana Rogobete
Translation can systematically alter text difficulty, particularly when moving into morphologically rich languages. This study examines whether readability-constrained Large Language Models (LLMs) can mitigate difficulty shifts observed in English–Romanian translation of children’s literature. We construct a paired four-condition corpus comprising English originals, published Romanian translations, readability-constrained LLM translations, and human readability adaptations (12 aligned passages; approx. 23,000 words). Readability is assessed using a Romanian grade-level index (LEMI) designed to be educationally comparable to Flesch–Kincaid Grade Level (FKGL), the cross-linguistic LIX metric, and morphologically informed measures derived from spaCy. Published Romanian translations are significantly more difficult than their English originals, showing higher LIX scores, grade-level estimates, and increased morphological variation. Readability-constrained LLM translation substantially reduces difficulty relative to the published versions (median delta approx. −1.46 grade levels), with significant decreases in LIX, morphological feature density, and lexical diversity (MTLD). Human adaptation yields a smaller reduction (median delta approx. −0.26). Although the direct comparison between LLM and human adaptation is marginal (p = .055, r = 0.64), LLM outputs generally produce larger reductions. These findings demonstrate that translation-induced difficulty shifts are measurable and that controllable LLM translation can modulate readability across structural, lexical, and morphological dimensions in multilingual educational contexts.
A Benchmark for Overgeneration Detection in Biomedical Text Simplification
Berkay Chakar | Liana Ermakova | Jaap Kamps
Berkay Chakar | Liana Ermakova | Jaap Kamps
Large Language Models deployed for biomedical text simplification frequently produce overgeneration: extraneous content appended beyond the faithful simplification, including leaked model instructions, ungrounded medical claims, and repetitive text. Despite its prevalence, this failure mode remains largely unaddressed. We present a benchmark for document-level overgeneration detection, releasing two resources: SimpleOG-manual, 500 abstract-level examples with human-validated positive labels, and SimpleOG-auto, over 46,000 automatically labeled abstract-level examples derived from submissions to the CLEF 2025 SimpleText Track. Our method exploits the positional regularity of overgeneration in simplification output through sequence alignment, identifying trailing content that lacks a corresponding segment in the source. Human validation of 117 automatically flagged positives confirms ∼95% precision, with leaked model instructions accounting for 75.7% of confirmed cases. Analysis across teams and models reveals that overgeneration is primarily driven by system-level choices, such as prompting and post-processing, rather than by model architecture. We evaluate three detection paradigms and find that sentence similarity (F1 = 0.731, ROC-AUC = 0.915) surprisingly outperforms both NLI-based and LLM-based approaches, suggesting that overgenerated content occupies distinct semantic regions from source material.
Conplext 1.0: A Multilingual Lexical Complexity Prediction Dataset for L2 Learning
David Alfter | Jasper Degraeuwe
David Alfter | Jasper Degraeuwe
This paper presents Conplext 1.0, a multilingual dataset designed for lexical complexity prediction in the context of second language (L2) learning. The resource covers 3,901 sentence contexts for 1,000 vocabulary items across five languages (English, French, Spanish, Swedish, and Dutch), each aligned with Common European Framework of Reference (CEFR) proficiency levels. Contexts were generated using a generative large language model and subsequently filtered for pedagogical suitability. A large-scale best–worst scaling (BWS) annotation experiment is being conducted with L2 learners to derive continuous, learner-informed lexical complexity values. The resulting dataset enables the development of context-aware word difficulty models that account for variation across both languages and learning stages. In addition to its primary use in lexical complexity prediction, Conplext provides valuable opportunities for research in word sense disambiguation, generative model evaluation, and adaptive language learning applications. By integrating computational and educational perspectives, this work advances the study of lexical difficulty in multilingual language learning environments.
From Complexity to Inclusivity: A Methodology for Drafting Patient-Centered Explanations of Gut-Brain Axis Concepts
Vanessa Bonato | Federica Vezzani | Giorgio Maria Di Nunzio
Vanessa Bonato | Federica Vezzani | Giorgio Maria Di Nunzio
Understanding specialized biomedical knowledge can be particularly challenging, posing significant barriers to the acquisition and use of medical information especially by patients. In this study, a methodology for drafting patient-centered explanations of concepts related to the gut-brain axis and related medical conditions is proposed. The explanations are specifically intended for patients affected by neurodegenerative diseases, who are experiencing cognitive decline. The methodology consists of the following steps: 1) the drafting of specialized definitions in the form of intensional definitions, which enable the structured representation of domain-specific knowledge, and 2) the simplification of specialized definitions into patient-centered explanations. In particular, explanations intended for patients are formulated using popular terms and plain language, considered as two complementary strategies aimed at enhancing the comprehension of specialized biomedical knowledge. This work lays the foundation for the future development of a terminology resource specifically designed to collect and systematically represent knowledge related to the gut-brain axis and associated health conditions.
Translation as Augmentation: Effect of Translated Data on Assessment of Difficulty
Yiheng Wu | Jue Hou | Roman Yangarber
Yiheng Wu | Jue Hou | Roman Yangarber
Reliable Text Difficulty Assessment is a prerequisite for valid text simplification workflows and personalized learning applications. However, the development of robust assessment models is severely hindered by a critical bottleneck: the scarcity of expert-annotated corpora containing fine-grained difficulty levels (e.g., CEFR), particularly for lower-resource languages. This paper addresses this data scarcity problem in the context of a low-resource European language. We propose a cross-lingual data augmentation strategy that leverages machine translation to transfer labeled resources from high-resource languages to the target low-resource language. We train BERT-based regression models to predict difficulty scores and investigate whether synthetic, translated data can effectively supplement native training sets. Our experiments demonstrate that augmenting scarce native data with machine-translated corpora significantly improves the accuracy of difficulty estimation, offering a viable solution for languages lacking extensive expert annotations.
Automatic Generation of Graded Texts in Old Church Slavonic
Iglika Nikolova-Stoupak | Gaël Lejeune | Aliona Shestakova-Stukun | Eva Schaeffer-Lacroix
Iglika Nikolova-Stoupak | Gaël Lejeune | Aliona Shestakova-Stukun | Eva Schaeffer-Lacroix
In the past few decades, graded readers have been valued within language education and have so much as extended onto the so-called classical (or ‘dead’) languages, such as Latin and Greek. The immersive reading and listening of adapted texts in these languages has been shown to increase students’ proficiency, independence and motivation. However, as of now there is only a small number of related resources as well as of classical languages represented. The present study will investigate the current potential for (semi-)automatic generation of adapted classical-language readers while focusing on the Old Church Slavonic language. From a Natural Language Processing (NLP) point of view, work with the language is challenging due to the variety of dialects and diachronic variations it encompasses. The following steps are taken within our study: 1) Representative measurable characteristics of professional classical-language readers, such as the Latin Lingua latina per se illustrata and the Greek Athenaze, are analysed. 2) Automatic generation of adapted Old Church Slavonic text is attempted through the use of a sequence-to-sequence model (mT5) as well as a Large Language Model (GPT-5) in a one-shot setting. 3) The derived texts’ quality is assessed through both human evaluation and a comparison of their textual characteristics with those of professional texts as defined in point 1). The edited versions of the GPT-based texts are shared for future reference and use.
A Calibrated and Interpretable Framework for Multilingual Text Difficulty Prediction
Voula Giouli | George Tsoulouhas | Athina Sioupi | Stamatia Michalopoulou
Voula Giouli | George Tsoulouhas | Athina Sioupi | Stamatia Michalopoulou
Text difficulty prediction in educational contexts requires models that balance predictive performance, interpretability, calibration, and pedagogical alignment. While transformer-based approaches increasingly dominate text difficulty classification, educational applications demand transparent and linguistically grounded modeling. This paper presents work aimed at developing a workbench for CEFR-based text difficulty prediction. The proposed platform comprises three main components: (i) a tool for CEFR-aligned dataset preparation incorporating a pipeline for documenting, processing, and enriching textual data, (ii) CEFR-aligned datasets, and (iii) three alternative modeling approaches, namely a rule-based baseline, a feature-based Machine Learning (ML) classifier, and a fine-tuned BERT model. Our approach integrates linguistically informed feature engineering with data-driven modeling techniques, thereby balancing transparency and predictive performance. The proposed workbench has been designed as a language-agnostic infrastructure that can be extended to any language. In its current implementation, it has been applied to the creation of a German CEFR dataset, while its Greek counterpart is currently under development.
Terminology-Augmented Generation for Intangible Cultural Heritage: A Controlled LLM-Based Translation Framework
Wanda Punzi Zarino | Pilar Sánchez Gijón
Wanda Punzi Zarino | Pilar Sánchez Gijón
This study examines the integration of a bilingual Italian–Spanish concept-oriented terminological resource into a controlled large language model (LLM) translation workflow within the domain of Campanian gastronomy. The termbase encodes structured conceptual, linguistic, and translational metadata, including grammatical information, translation strategies, and genre-sensitive usage recommendations. Through a local Model Context Protocol (MCP) architecture, the resource is dynamically connected to locally deployed LLMs, enabling the automatic identification and retrieval of relevant terminological units prior to generation. The system combines in-context terminological injection with deterministic post-processing enforcement: genre-specific policies are injected into the model prompt prior to generation and verified through a rule-based post-processing layer that enforces surface-level terminological consistency in the output. Two open-weight models — Mistral 7B Instruct and Gemma3 4B — are evaluated across three conditions and three discursive genres on a dataset of authentic texts. The findings suggest that the combination of terminological injection and deterministic enforcement can improve terminological compliance in controlled, domain-specific settings, while also highlighting differences in instruction-following behavior across models and genres.
Assessing Small Language Models as Text Simplification Evaluators
David Carranza Navarrete | Jan Bakker | Jaap Kamps
David Carranza Navarrete | Jan Bakker | Jaap Kamps
Text simplification requires reliable automatic evaluation, yet existing learnable metrics such as LENS and LENS-SALSA are specialized and costly to develop. Moreover, it remains unclear how these metrics compare to using large language models (LLMs) as evaluators. Exploring this question is important because LLM-based evaluation could make simplification research and deployment more flexible and easier to adapt than training new task-specific metrics for each setting. In this work, we empirically compare several small, open-weight instruction-tuned LLMs with LENS and LENS-SALSA in both reference-based and reference-free evaluation settings. We measure their alignment with human judgments across multiple datasets. Our results provide insight into when small LLMs can serve as effective evaluators and when specialized metrics remain preferable, informing the design of future evaluation pipelines for text simplification and related text generation tasks.
up
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Antonis Anastasopoulos | Stella Markantonatou | Angela Ralli | Marcos Zampieri | Stavros Bompolas | Vivian Stamou
Antonis Anastasopoulos | Stella Markantonatou | Angela Ralli | Marcos Zampieri | Stavros Bompolas | Vivian Stamou
A Bolu: A Structured Dataset for the Computational Analysis of Sardinian Improvisational Poetry
Silvio Calderaro | Johanna Monti
Silvio Calderaro | Johanna Monti
The growing interest of Natural Language Processing (NLP) in minority languages has not yet bridged the gap in the preservation of oral linguistic heritage. In particular, extemporaneous poetry — a performative genre based on real-time improvisation, metrical-rhetorical competence — remains a largely unexplored area of computational linguistics. This methodological gap necessitates the creation of specific resources to document and analyse the structures of improvised poetry. This is the context in which A Bolu was created, the first structured corpus of extemporaneous poetry dedicated to cantada logudorese, a variant of the Sardinian language. The dataset comprises 2,835 stanzas for a total of 141,321 tokens. The study presents the architecture of the corpus and applies a multidimensional analysis combining descriptive statistical indices and computational linguistics techniques to map the characteristics of the poetic text. The results indicate that the production of Sardinian extemporaneous poets is characterised by recurring patterns that support Parry and Lord’s theory of formulaicity. This evidence not only provides a new key to understanding oral creativity, but also offers a significant contribution to the development of NLP tools that are more inclusive and sensitive to the specificities of less widely spoken languages.
Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
Lena Sophie Oberkircher | Jesujoba Alabi | Dietrich Klakow | Jürgen Trouvain
Lena Sophie Oberkircher | Jesujoba Alabi | Dietrich Klakow | Jürgen Trouvain
Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread use, are underrepresented in linguistic resources and computational models, resulting in performance disparities. To address this gap, we introduce Saar-Voice, a six-hour speech corpus for the Saarbrücken dialect of German. The dataset was created by first collecting text through digitized books and locally sourced materials. A subset of this text was recorded by nine speakers, and we conducted analyses on both the textual and speech components to assess the dataset’s characteristics and quality. We discuss methodological challenges related to orthographic and speaker variation, and explore grapheme-to-phoneme (G2P) conversion. The resulting corpus provides aligned textual and audio representations. This serves as a foundation for future research on dialect-aware TTS, which is proposed in the form of zero-shot and fine-tuned model adaptation in low-resource scenarios.
MD_NLP: Reconstructing an Australian English Heritage Dialect Corpus from the Mitchell-Delbridge Recordings through LLM-Assisted Speaker Attribution
Steven Coats
Steven Coats
We present MD_NLP, a discourse-annotated and georeferenced corpus derived from the Mitchell–Delbridge (MD) recordings, a foundational archive of mid-20th-century Australian English. The corpus comprises word-aligned narrative recordings from 7,735 secondary school pupils across 327 locations, enriched with structured sociodemographic metadata (e.g., sex, birthplace, parental background) and geocoded institutional coordinates. The narratives were reconstructed from archival audio using an integrated pipeline combining WhisperX-based automatic speech recognition, neural speaker diarization, LLM-assisted discourse-role correction, and Montreal Forced Aligner (MFA) boundary refinement. Evaluation on manually annotated data shows that incorporating an LLM-based reasoning step improves turn-level speaker-role attribution from 62.70% (acoustic diarization alone) to 95.68%. Unlike prior uses of the MD archive, which focused on controlled sentence materials, MD_NLP makes the spontaneous narrative component accessible for large-scale analysis. The resulting resource supports research on regional and socially conditioned variation, discourse structure, and corpus phonetics in Australian English. The proposed architecture is directly transferable to other legacy dialect archives, providing a practical pathway for transforming interview-based recordings into temporally aligned, speaker-consistent corpora.
Challenges in the Detection of Dialect for Historical Languages; the Case of Old Irish Text Resources
Adrian Doyle
Adrian Doyle
Old Irish presents particular challenges for the study of automatic dialect detection. It is generally accepted that Old Irish presents little trace of dialect. Extant Old Irish text resources introduce a considerable amount of extra variation, which could impact dialect identification applications. While some scholarship has suggested that certain features may be indicative of dialect, such hypotheses are difficult to substantiate where authorship is anonymous, or where the text itself is not associated with a particular geographical region. This paper describes the application of stylometric dialect detection techniques to Old Irish texts, and discusses the features which emerge from this process as potential markers of dialect. The aim is not necessarily to identify Old Irish dialectal features, but rather to investigate the impact that Old Irish text resources could have on such applications. This paper does, however, add to the extant body of research by highlighting some features which might be identified as stylistically distinct by stylometric dialect identification techniques.
Phonologically-aware Automatic Speech Recognition Evaluation of Low-Resource Languages: The Case of Basque Dialects
Christoforos Souganidis | Asier Herranz | Ibon Saratxaga | Eva Navas | Inma Hernaez
Christoforos Souganidis | Asier Herranz | Ibon Saratxaga | Eva Navas | Inma Hernaez
Automatic speech recognition models are typically trained with data of standard languages. However, their performance degrades when dealing with non-standard dialectal speech. In this paper, we present the first evaluation of an automatic speech recognition system for Basque, a low-resource language, based on spontaneous broadcast speech with high representation of dialectal speech. It relies on a 140-h manually annotated propietary corpus of television programs broadcast by Basque Radio Television, including dialect-level labels, as well as standardized and pseudo-phonetic transcriptions. We find that recognition performance significantly degrades for dialectal compared to standard speech, for all dialects present in our corpus. Subsequently, we provide a quantitative analysis of phonological phenomena based on single-word substitution errors, and identify 52 recurrent phenomena, grouped into sound deletions, epentheses, and substitutions. We further show a modest but statistically significant correlation between the number of phonological phenomena in an utterance and its recognition error rate. Our findings highlight the limitations of dialect-agnostic evaluation and motivate linguistically informed, dialect-aware strategies for automatic speech recognition in low-resource and typologically diverse languages.
Literary transcriptions of spoken language often deviate from standard, written language. These variations can lead to higher than desirable error rates in NLP processing. This is particularly the case for spoken data of low resource varieties, including dialects and contact varieties of higher resource languages. This paper outlines a proposal for the systematic dialect-to-standard normalization of spoken language from language contact and dialect contact situations. This system is then tested on the Texas German Sample Corpus (~13 hours), a set of audio and transcripts of Texas German conversations. Texas German is an umbrella term for a set of a heritage varieties of German spoken in Texas, USA that descend from multiple German dialects and that have been in contact with English for 150+ years. The proposed normalization system, along with the accompanying language-tagging system, can act as a starting point for other projects interested in normalizing their mixed variety data.
Handling Cross-Dialect Syntactic Variation: a Theory-Driven Web Resource
Emanuela Li Destri | Marco Longhin | Gaia Sorge | Sofia Ferroni | Giovanni Battista Matteazzi | Andrea Artioli | Lorenzo Carletti | Federico Motta | Giuseppe Longobardi | Cristina Guardiano
Emanuela Li Destri | Marco Longhin | Gaia Sorge | Sofia Ferroni | Giovanni Battista Matteazzi | Andrea Artioli | Lorenzo Carletti | Federico Motta | Giuseppe Longobardi | Cristina Guardiano
Cross-dialect syntactic variation tests the limits of comparative analysis, owing to the entanglement of inheritance and contact in dialect systems. Addressing this challenge requires analytical tools combining the theoretical depth of formal models of grammatical competence with quantitative taxonomic techniques. The Parametric Comparison Method (PCM) embodies this integration by quantifying structural similarity across grammars through the comparison of abstract syntactic rules. The method has been shown to achieve a good degree of resolution in dialectal domains, capturing subtle contrasts and yielding configurations aligning with phylogenetic expectations while remaining sensitive to contact-induced convergence. Fully assessing its effectiveness as a resource for the quantitative study of syntactic dialectology, however, requires an infrastructure that ensures systematic data collection, consistent parameter setting, and robust statistical evaluation across diverse datasets. The PCM Hub is a web-based resource designed for this purpose. It integrates guided elicitation, automated parameter-setting procedures, data management, and the computation of distances and automatic classifications within a unified environment. By standardizing the transition from raw linguistic observations to a structured, replicable empirical apparatus, the PCM Hub provides the practical and quantitative support necessary to test the power of the PCM across expanded comparative domains.
Can LLM Agents Identify Spoken Dialects like a Linguist?
Tobias Bystrich | Lukas Hamm | Maria Hassan Akhter | Lea Fischbach | Lucie Flek | Akbar Karimi
Tobias Bystrich | Lukas Hamm | Maria Hassan Akhter | Lea Fischbach | Lucie Flek | Akbar Karimi
Due to the scarcity of labeled dialectal speech, audio dialect classification is a challenging task for most languages, including Swiss German. In this work, we explore the ability of large language models (LLMs) as agents in understanding the dialects and whether they can show comparable performance to models such as HuBERT in dialect classification. In addition, we provide an LLM baseline and a human linguist one. Our approach uses phonetic transcriptions produced by ASR systems and combines them with linguistic resources such as dialect feature maps, vowel history, and rules. Our findings indicate that, when linguistic information is provided, the LLM predictions improve. The human baseline shows that automatically generated transcriptions can be beneficial for such classifications, but also present opportunities for improvement.
Beyond Accuracy: Analyzing Dialect Confusion in Automatic Speech-Based Dialect Classification
Lea Fischbach | Alfred Lameli | Lucie Flek
Lea Fischbach | Alfred Lameli | Lucie Flek
Automatic dialect classification is commonly treated as a supervised task with a primary focus on overall accuracy. In this paper, we argue that classification errors and model uncertainty provide valuable insights into dialectal structure and variation. We analyze a speech-based dialect classification model trained on German dialect data from three generations and evaluated across 250 speaker-disjoint splits (median weighted F1=0.42). A systematic confusion analysis shows that misclassifications are largely explained by speaker diversity, dialectal similarity, geographical proximity, and speaker self-assessment. Among these factors, the number of speakers per dialect has the strongest impact on performance, while frequent confusions between closely related dialects reflect inherent linguistic similarity rather than model limitations. Generational analyses further indicate that younger speakers exhibit reduced dialectal distinctiveness, although core dialectal features remain shared across generations. By explicitly modeling classification uncertainty, the proposed approach enables the analysis of dialect transition areas and gradient dialect boundaries. Overall, this work demonstrates that automatic dialect classification can serve not only as a predictive task but also as a tool for dialectological analysis.
We present FLEURS-Kobani, a Northern Kurdish (ISO 639-3 KMR) spoken extension of FLEURS benchmark. Although FLEURS offers n-way parallel speech for 100+ languages, Northern Kurdish is absent, limits benchmarking automatic speech recognition and speech translation tasks for this language. The FLEURS-Kobani dataset consists of 5,162 validated utterances, totaling 18 hours and 24 minutes. As baselines, we fine-tuned Whisper v3-large for ASR and E2E S2TT. A two-stage fine-tuning strategy (Common Voice→FLEURS-Kobani) yields the best ASR performance (WER 28.11, CER 9.84 on test). For end-to-end S2TT (KMR→EN), Whisper achieves 8.68 BLEU on test; we additionally report pivot-derived targets and a cascaded S2TT setup. FLEURS-Kobani provides the first Northern Kurdish public benchmark for evaluation of ASR, S2TT and S2ST tasks. The data can be accessed (LINK IS BLANK DUE TO REVISION RULES) under CC BY 4.0 license.
Exploring the reusability of Northern Kurdish resources for Badini speech recognition
Mohammad Mohammadamini | Aveen Jalal Mohammed | Barzan Hussein Mohammed | Dezheen H. Abdulazeez | Imad Saeed Sadeeq | Dilgash Mohammed Salih | Amera Ismail Melhum | Abuobaida Abdullah Dheyab
Mohammad Mohammadamini | Aveen Jalal Mohammed | Barzan Hussein Mohammed | Dezheen H. Abdulazeez | Imad Saeed Sadeeq | Dilgash Mohammed Salih | Amera Ismail Melhum | Abuobaida Abdullah Dheyab
Badini is a variant of the Kurdish language spoken in the Duhok province of the Kurdistan Region of Iraq. It is written mainly in a modified version of the Arabic script. Although it shares the same script as Central Kurdish (CKB), it is linguistically classified under the Northern Kurdish (KMR) branch. In this paper, we explore the potential and limitations of Northern Kurdish ASR resources for the Badini variant. Firstly, we transliterate the Common Voice 18 dataset from the Latin script into the modified Arabic script and revised it to align with the orthographic conventions of Badini variant. Additionally, we introduce the first text collection for the Badini variant, containing 14,22 million tokens, which serves as a source for speech synthesis. A third resource developed in this research is a standard speech recognition benchmark recorded by 5 speakers which includes 2 hours and 46 minutes of multi-domain read speech. Results show that combining transliterated and synthetic data significantly improves recognition accuracy, achieving a 6.8% CER and 34% WER. All three resources curated during this research will be made available under the CC BY-NC-ND 4.0 license.
Wancho Dialectometry: Community-created data and the Living Dictionaries project
Kellen Parker van Dam
Kellen Parker van Dam
Community-created lexical resources for under-documented languages represent an underexplored data source for computational dialectology. This study evaluates the viability of such data for dialectometric analysis, using the Wancho (Glottocode: wanc1238) LivingDictionaries project as a case study. Wancho is a Tibeto-Burman language of the Southwestern Patkaian branch, spoken primarily in Longding District, Arunachal Pradesh, India. The dictionary is notable for being entirely community-built and speaker-facing, and uniquely among resources for Northeast India, it incorporates dialect-specific forms spanning village-level geolects and clanlects. We extract dialectal data via automated web scraping and apply a series of preprocessing steps to address inconsistencies in transcription, language labelling, and concept assignment. Pairwise linguistic distances are then computed using Sound Class Alignment (SCA, List 2010), which captures phonological similarity more accurately than raw edit distance by incorporating articulatory feature structure. The resulting distance matrix is analysed through UPGMA hierarchical clustering and NeighborNet split network inference. Despite the dataset’s uneven dialect coverage and absence of systematic cognate coding, SCA-based distances recover the traditional Upper/Lower/Middle Wancho distinction and correctly situate transitional varieties. These results hold even for dialects with as few as a dozen attested forms. We show that unlike Bayesian phylogenetic inference which is poorly suited to data of this density and distribution, SCA proves to be a reliable metric. Our findings suggest that SCA distance is robust to the kinds of noise and sparsity characteristic of community-generated lexical data, and that such resources constitute a viable, if imperfect, input for automated dialectometric workflows — particularly in contexts where fieldwork-based data collection is not currently feasible.
Dialectometry and Evaluation of the ePark Corpus for Low-Resource Formosan Language Dialects
Henry Gagnier
Henry Gagnier
Formosan languages are a critically endangered branch of the Austronesian family spoken in Taiwan, and many of their dialects remain poorly understood and computationally understudied. Subgrouping relationships in these languages are often contested and unresolved. We provide the first evaluation of the ePark corpus as a dialectal NLP resource, identifying its strengths and gaps for future NLP work, and present the first large-scale corpus-based computational analysis of dialect similarity across all officially recognized Formosan languages. We use the ePark corpus to analyze 42 dialects in 16 Formosan languages, and through word-level TF-IDF cosine similarity, Jaccard similarity over shared vocabulary, and Levenshtein distance, we quantify pairwise dialectal relationships within the Amis, Atayal, Seediq, Bunun, Paiwan, Rukai, and Puyuma languages. We find that simple lexical similarity methods can recover and confirm linguistically established dialectal subgroupings. We find that in multiple cases the two metrics diverge, offering insights on contested subgroupings such as Mantauran Rukai. This work establishes a scalable methodological framework for dialectometry in low-resource languages, demonstrates the value of the ePark corpus for Formosan NLP research, and encourages future work in NLP on Formosan dialects.
A Dialectal Corpus for Ukrainian: Collection, Classification, and Standardization
Yuliia Frund | Sina Ahmadi
Yuliia Frund | Sina Ahmadi
Ukrainian dialects remain largely excluded from the digital linguistic landscape despite their active everyday use. We present a regional dialect corpus covering 18 administrative regions of Ukraine, compiled from digitized fieldwork collections and an online dialect atlas. The corpus comprises over 284,000 tokens of dialect text, annotated by region and partially accompanied by manually standardized translations. Using these resources, we investigate language identification and dialect-to-standard standardization. Baseline language identification yields an F-score of 0.75, rising to 0.99 with dialect-inclusive training. Dialect classification reaches 0.58, with confusion patterns reflecting known regional boundaries. For standardization, the best-performing LLM achieves a COMET score of 0.80, though BLEU scores remain low (0.21–0.23) across all models. We release the corpus, labelled datasets, model outputs, and reference translations to support future work on inclusive language technologies for non-standard varieties.
German Dialects Across Situations, Generations, and Regions: The REDE corpus as an Oral Resource for NLP
Hanna Fischer | Alfred Lameli
Hanna Fischer | Alfred Lameli
Recent advances in speech and language technologies increasingly rely on large and diverse corpora that represent linguistic variation across dialect regions, communicative situations, and social speaker characteristics. While substantial resources are available for Standard German, comparable spoken corpora for German dialects have so far been largely lacking, limiting the development and evaluation of dialect-sensitive NLP systems. The REDE corpus addresses this gap by providing a methodologically uniform collection of spoken German for 148 locations that systematically covers all major dialect areas in Germany. It comprises contemporary recordings collected in multiple elicitation and interaction settings, capturing variation across speaking styles, situational contexts, and speaker generations. With more than 1,500 hours of speech and rich metadata on regional and social dimensions, the REDE corpus constitutes a large-scale oral resource suitable for both linguistic research and NLP applications. This paper presents the design, structure, and methodological foundations of the corpus and discusses its relevance for current speech technology requirements.
A Catalog of Basque Dialectal Resources: Online Collections and Standard-to-Dialectal Adaptations
Jaione Bengoetxea | Itziar Gonzalez-Dios | Rodrigo Agerri
Jaione Bengoetxea | Itziar Gonzalez-Dios | Rodrigo Agerri
Recent research on dialectal NLP has identified data scarcity as a primary limitation. To address this limitation, this paper presents a catalog of contemporary Basque dialectal data and resources, offering a systematic and comprehensive compilation of the dialectal data currently available in Basque. Two types of data sources have been distinguished: online data originally written in some dialect, and standard-to-dialect adapted data. The former includes all dialectal data that can be found online, such as news and radio sites, informal tweets, as well as online resources such as dictionaries, atlases, grammar rules, or videos. The latter consists of data that has been adapted from the standard variety to dialectal varieties, either manually or automatically. Regarding the manual adaptation, the test split of the XNLI Natural Language Inference dataset was manually adapted into three Basque dialects: Western, Central, and Navarrese-Lapurdian, yielding a high-quality parallel gold standard evaluation dataset. With respect to the automatic dialectal adaptation, the automatically adapted physical commonsense dataset (BasPhyCowest) underwent additional manual evaluation by native speakers to assess its quality and determine whether it could serve as a viable substitute for full manual adaptation (i.e., silver data creation).
WoVis: Interactive Visualization of Word Embeddings for Semantic Change in Historical and Dialectal Language Resources
Filip Miletić | Maximilian Henkel | Rene Cutura | Sophie Sadler | Quynh Quang Ngo | Michael Sedlmair | Sabine Schulte im Walde
Filip Miletić | Maximilian Henkel | Rene Cutura | Sophie Sadler | Quynh Quang Ngo | Michael Sedlmair | Sabine Schulte im Walde
Computational modeling of language variation and change often relies on comparisons of word embeddings induced from existing historical and dialectal language resources. However, their use in the wider linguistics research community and in application domains such as lexicography is challenged by their limited manipulability for non-technical users, which in turn exacerbates the underuse of such resources. Aiming to foster a broader uptake of embedding-based analyses, we introduce WoVis, an interactive visualization tool designed to compare word embedding models in analyses of semantic change. Our system supports simultaneous model comparisons along two dimensions (e.g., language varieties and time periods) and provides analyses at different levels of granularity: an overview of the full vocabulary across all word embedding models, distributional behavior of individual words, targeted comparisons of word pairs, and model-external lexical features such as frequency and affective norms. We illustrate the utility of our system on two languages, German and English, with analyses of word usage across language varieties as well as time: West vs. East Germany, 1950–1989; and general-domain US vs. scientific UK English, ca. 1800–2000.
Speaker Normalization via Voice Conversion Reveals a Human-Machine Dissociation in Dialect Classification
Caroline Kleen | Lea Fischbach | Akbar Karimi | Lucie Flek | Alfred Lameli
Caroline Kleen | Lea Fischbach | Akbar Karimi | Lucie Flek | Alfred Lameli
This study evaluates whether Retrieval-based Voice Conversion (RVC) can be used to normalize speaker-specific variability while preserving dialect-relevant acoustic cues, and what the response of human and machine systems to this manipulation reveals about the architecture of dialect recognition. In two perception experiments, speech samples from nine German dialect regions were presented either in their original form or after conversion to a single target speaker. We compared overall accuracy, confusion structures, item-level response distributions, and the interaction between listener origin and target dialect across conditions. Human classification remained stable under voice conversion. Accuracy did not differ between conditions, confusion matrices were highly correlated, and item-level divergences were minimal. The interaction between listener origin and target dialect—reflecting systematic regional bias—remained invariant. These findings indicate that RVC does not distort perceptually relevant dialectal cues and that human dialect recognition is robust to speaker normalization. In contrast, we evaluated a deep learning model under matched conditions: model accuracy improved significantly under RVC, while human performance remained unchanged. This dissociation reframes RVC as an experimental probe for investigating the divergence between human and machine speech processing, suggesting that this divergence is rooted in fundamentally different representational architectures.
This paper presents a developing oral resource for South Tyrolean, a German dialect spoken in Northern Italy. The dialect is ubiquitous in spoken communication but lacks a standardised orthography. In this context, strict transcription into dialect is of limited to no utility to the local community. Instead, there is a distinct and strong demand for technology capable of directly translating spoken dialect into Standard German. To address this specific need, we introduce a dynamic, incrementally growing dataset designed to fine-tune ASR models for this translation task. Our corpus aggregates diverse sources, including media and research interviews, totalling over 13 hours of aligned audio. We describe a collaborative workflow where community partners contribute audio archives in exchange for automated transcriptions, creating a virtuous cycle of data improvement. Additionally, we detail our iterative model fine-tuning strategy, data collection challenges and the resulting improvements in model performance.
TransVar – the Corpus for Variation and Change Study of the Historical Transcarpathian lects
Ilia Afanasev
Ilia Afanasev
The paper introduces TransVar – the corpus of the historical Transcarpathian lects (the first half of the XX century, the territories of modern Ukraine, Poland, Slovakia, and Romania). The corpus contains data from Lemko, Bojko and Hutsul small territorial lect groups. It is crucial for studies of the people of these territories, who witnessed forceful deportation from their homeland in the 1940s – 1950s, soon after the recordings were made (1920s – 1930s). The article also provides a brief overview of their linguistic properties, as evident in the material. The corpus is morphosyntactically tagged. It contains data on part-of-speech, morphological features, lemmata and syntactical dependencies. The study stresses the crux of manual analysis of the errata made in an automatic tagging phase for further improvement. The supplementary information includes named entities encountered in the text and the basic vocabulary. All the texts are accompanied by metalinguistic information, required for the sociolinguistic study. After the analysis of the current stage of the corpus creation, the article outlines further research prospects. Apart from more thorough manual annotation, one of the prospects is to add English translation with the purpose of making the material more accessible to scholars without a background in Slavic studies.
The Generator-Eraser Paradox: Community Guidelines for Responsible LLM-Assisted Dialect Resource Creation
Wajdi Zaghouani
Wajdi Zaghouani
Dialect resources occupy a unique position at the intersection of scientific description, cultural preservation, and computational infrastructure. Large language models offer powerful capabilities for accelerating dialect resource development through retrieval-grounded drafting, corpus navigation, metadata enrichment, and annotation workflow support. However, the same systems pose substantial risks: they can contribute to dialect erasure by privileging prestige varieties, homogenizing orthography, and enabling synthetic feedback loops that reduce linguistic diversity over time. These risks are particularly acute for language varieties characterized by diglossia, limited written standardization, or marginalized speaker communities. This paper makes three contributions. First, we integrate insights from variationist sociolinguistics and corpus linguistics to formalize the generator-eraser paradox as a theoretical framework for understanding the dual nature of LLM-assisted dialect work. Second, we derive 12 community guidelines that operationalize this framework into implementable design requirements for dialect resource creation and documentation. Third, we provide an in-depth case study of Arabic dialects, including a structured comparison of widely used resources, to demonstrate how these guidelines address language-specific challenges including diglossia, orthographic variability, and community governance. The contribution is conceptual and operational rather than experimental, with the goal of enabling dialect communities and resource builders across languages to adopt LLMs without sacrificing authenticity, variation, or sovereignty.
The Texas German Dialect Project Corpus as a Diachronic Resource for Investigating Language Contact
Thomas Schmidt | Margaret M. Blevins | Hans C. Boas | Glenn Gilbert
Thomas Schmidt | Margaret M. Blevins | Hans C. Boas | Glenn Gilbert
The Texas German Dialect Project (TGDP) is a long-standing effort to document the unique variety of German spoken in Texas since the 1840s. For 25 years, the TGDP has built up the freely accessible Texas German Dialect Archive Online (TGDA Online) with recordings and annotations of interviews and language tasks conducted between 2001 and today with some of the last speakers of the variety, which is expected to go extinct within the next 5-10 years. The present paper reports on the most recent addition of to the TGDP’s online corpus platform— a collection of Texas German data recorded in the 1960s—as well as historical translation elicitations that will be released later in 2026. Both the contemporary and the historical materials follow very similar elicitation methods and are processed using the same pipelines, increasing their comparability. This provides a comparable historical dimension to the resource, enabling diachronic and multidimensional analyses of this endangered variety. These data can help shed light on the dynamics of language contact, dialect contact, and language death.
This paper presents a multi-media corpus of Pontic Greek as spoken by Pontic Greek speakers in the Caucasus (Georgia). The corpus covers three major stages reflecting different sociolinguistic settings: (a) Ponitc Greek in small rural communities in Georgia (original settlements); (b) internal migration to urban centers (within Georgia), (c) external migration (to Greece). The dataset comprises 373 audio recordings (total duration 7h 26m; total word count: 43.073). The open-access resource includes audio files (wav) and annotations (xml). Annotations provide orthographical transcription, morphemic transcription and morpheme-by-morpheme and sentence-by-sentence translations in English (Toolbox); transcriptions are time-aligned with the audio files (ELAN). This collection is intended to linguists working on dialectology and language contact, as well as people with broader interests about the history and practices of this community. Pontic Greek in the Caucasus offers a unique opportunity to investigate contact between Greek and another Indo-European language (Russian) as well as two Non-Indo-European languages (Georgian, Turkish).
Meaning Over Morphology: A Multi-Metric Benchmark of LLMs for Bangla Dialect Translation
Soumik Deb Niloy | Subhey Sadi Rahman | Mahbub E Sobhani | Md. Golam Rabiul Alam | Farig Yousuf Sadeque | Md. Rezuwan Hassan
Soumik Deb Niloy | Subhey Sadi Rahman | Mahbub E Sobhani | Md. Golam Rabiul Alam | Farig Yousuf Sadeque | Md. Rezuwan Hassan
Regional dialects of Bangla, such as Sylheti and Chittagonian, pose significant challenges for natural language processing due to their low-resource nature and substantial linguistic variation from standard Bangla. In this work, we present a systematic evaluation of eight open-source LLMs for translating fifteen distinct Bangla dialects into standard Bangla. To achieve this comprehensive coverage, we utilize a combination of established benchmarks and a novel dataset curated from an ongoing regional linguistic project. We assess model performance using a multi-metric framework that combines exact-match and error-rate evaluations such as, Averaged BLEU, WER, and CER with embedding-based semantic metrics including BERTScore, METEOR, and COMET. Additionally, we perform a detailed dialect-level linguistic analysis to identify the deep-seated structural, orthographic, and semantic barriers inherent to dialectal translation. Our study highlights the strengths and limitations of current open-source models, provides empirical insights for future dialect-aware fine-tuning, and contributes a reproducible benchmark for the research community.
Sociolinguistic aspects of crowdsourcing for a vocal corpus of Alsatian
Pascale Erhart | Lucile Hamm | Sam Bigeard | Carole Werner | Malek Yaich | Slim Ouni
Pascale Erhart | Lucile Hamm | Sam Bigeard | Carole Werner | Malek Yaich | Slim Ouni
Alsatian is a regional low-resource language spoken in a majority-language context. In order to create a voice dataset suited for training automatic speech recognition and speech-to-text models, we launched a crowdsourcing campaign on the platform Mozilla Common Voice. We describe sociolinguistic issues we ran into, such as participants’ perception of their own language and its role in the AI landscape, which are vital to address to raise the participation in the crowdsourcing effort. We found that the participants are often confused about NLP and AI tools, and have a strong interested in preserving their language.
HeptaTAX: A Neuro-Symbolic Pipeline and Benchmark for Classifying 16th-Century Heptanesian Notarial Acts
Stergios Chatzikyriakidis | Eleni Karantzola | Vasiliki Makri
Stergios Chatzikyriakidis | Eleni Karantzola | Vasiliki Makri
This study originates in the investigation of lexical bundles and formulaic language within sixteenth-century Corfiot notarial documents. The observed functional variation across identical formulaic sequences motivated the development of a document classification framework designed to support the structural interpretation of such language. Given that 16th-century Corfiot notarial acts represent a rich, albeit understudied, dialectal resource, their systematic categorization into subgenres is essential for their full exploration. However, this task requires substantial manual work, while NLP tools for this task and dialect do not exist. In this paper, we attempt to take an initial step in this direction. First, we present a corpus of 1,088 notarial acts from 5 notaries spanning 1500-1567, a 3-tier annotation schema (17 core genres, extension subcategories, hybrid cross-cutting tags), and a 40-act benchmark with gold annotations at all three tiers. Then, we evaluate 12 LLMs across 4 architectures, zero-shot, few-shot, full-context and Neuro-Symbolic. For the latter, we introduce a symbolic engine comprising a set of deterministic rules for identifying discriminative legal formulae, whose output is then injected into the neural (LLM) engine. The results show that the NeSy architecture compresses the accuracy gap between stronger and weaker models from 47.5 pp to 12.5 pp, with the smallest model (Llama 3.1 8B) gaining 47.5% and matching frontier models that operate without symbolic support. Three models reach a ceiling of 72.5% on the core tier. However, consistent errors in procedurally dense material reveal the limits of lexical and formulaic cues for identifying legal effect, motivating the use of symbolic signals in the NeSy pipeline. Extension and hybrid classification remain open challenges, with best scores of ∼63% and ∼35% respectively.
Towards Semantic Access and Interoperability in Digital Dialectal Atlases. A Case Study
Paola Marongiu | Simonetta Montemagni
Paola Marongiu | Simonetta Montemagni
The increasing digital availability of dialectal atlases has significantly enhanced access to dialectal data and their potential for linguistic and cultural studies. However, despite their richness, such resources often remain difficult to integrate into contemporary data-driven research workflows, due to complex data structures and limitated interoperability. Most digital dialectal atlases still rely on traditional access models centered on maps, offering only implicit and coarse-grained semantic structures, which limits concept-based exploration and potential for integration with other linguistic resources in the Linguistic Linked Open Data (LLOD) ecosystem. This paper presents a case study carried out on the Atlante Lessicale Toscano (ALT), aimed at addressing these limitations through the introduction of an explicit semantic layer designed to support both user-oriented exploration and machine-actionable interoperability. While ALT already provides a conceptual organization of dialectal materials, this structure was originally conceived for human navigation and not for integration with other computational lexical-semantic resources. To bridge this gap, we align ALT concepts with ItalWordNet, leveraging its synset-based model as a widely adopted semantic backbone in NLP and LLOD infrastructures. The case study focuses on the domain of agriculture, whose historically grounded conceptual distinctions are often underrepresented in general-purpose lexical resources. The paper proposes a mapping strategy, analyzes coverage and mismatch patterns, and releases a new aligned resource mapping ALT agricultural concepts to ItalWordNet, thereby creating the prerequisites for interoperability and reusability of dialectal atlas data.
A CLDF-Compliant Lexical Database for Modern Greek Dialects: Resource Design and Dialectometric Analysis
Stavros Bompolas | Natalia Chousou-Polydouri | Manuela Genitsaridi | Danae Karatzanou | Georgios Kostopoulos | Elena Anagnostopoulou | Dimitra Melissaropoulou
Stavros Bompolas | Natalia Chousou-Polydouri | Manuela Genitsaridi | Danae Karatzanou | Georgios Kostopoulos | Elena Anagnostopoulou | Dimitra Melissaropoulou
This paper presents the first CLDF database that systematically documents lexical variation across 36 Modern Greek varieties (including Standard Modern Greek). The dataset aligns 14,378 lexical items over 345 concepts, links varieties to stable identifiers (Glottocodes), introduces a Greek-specific concept list, and maps meanings to standardized Concepticon concept sets, enabling interoperability and reproducible workflows. To assess whether the database preserves a meaningful dialectological signal, we conduct a dialectometric analysis by computing feature-sensitive string distances over IPA transcriptions and applying hierarchical clustering. The resulting similarity structure recovers major macro-divisions—most notably a broad Northern vs. Southern partition among Koine-descended mainland varieties—and isolates peripheral groups with distinct historical trajectories (e.g., Asia Minor, Italiot, Tsakonian). The database provides scalable infrastructure for quantitative dialectology, comparative Greek linguistics, and dialect-aware language technology.
A Speech Resource for the Pontic Greek Dialect: Transcription Choices and Baseline ASR Evaluation
Rodanna Konstantinidou | Chara Tsoukala | Vivian Stamou | Voula Giouli | Stella Markantonatou
Rodanna Konstantinidou | Chara Tsoukala | Vivian Stamou | Voula Giouli | Stella Markantonatou
Pontic Greek is a living but endangered Modern Greek dialect that lacks publicly available AI-oriented speech resources and ASR benchmarks. This work reports on the first systematic inference-only (zero-shot) ASR evaluation on authentic Pontic speech. Progress on Pontic ASR is hindered by two coupled challenges: the scarcity of transcribed speech data and the absence of a standardized orthography, which makes it difficult to create consistent reference transcriptions for evaluation. We address these challenges by releasing a new speech corpus of contemporary Pontic as spoken in Northern Greece, derived from natural conversations and provided with manual, utterance-level, time-aligned transcriptions. To reduce annotator bias and increase practical usability, we collect community evidence on written-form preferences via a small questionnaire and use the observed patterns to guide a consistent Greek-script transcription scheme. We use this corpus to perform inference-only (zero-shot) ASR evaluation, benchmarking four state-of-the-art pretrained speech recognition models under a unified evaluation protocol. Results show that zero-shot recognition remains challenging, establishing baseline figures and underscoring the need for dialect-specific data and adaptation.
First Steps in ASR for Cypriot Greek: Challenges and Insights
Vivian Stamou | Spyros Armostis | Antigoni Klimi | Georgios Paraskevopoulos | Vassilis Katsouros | Antonios Anastasopoulos
Vivian Stamou | Spyros Armostis | Antigoni Klimi | Georgios Paraskevopoulos | Vassilis Katsouros | Antonios Anastasopoulos
This paper presents the first automatic speech recognition (ASR) system for Cypriot Greek, a non-standardized variety of Modern Greek with distinctive phonological, lexical, and orthographic characteristics. We adapt Whisper, a state-of-the-art multilingual ASR model, to Cypriot Greek through fine-tuning on the mozilla common voice spontaneous speech dataset for Cypriot Greek. The phonological and lexical divergence of Cypriot Greek from Standard Modern Greek poses significant challenges for mainstream ASR, particularly under conditions of limited training data and dialectal variation. Results demonstrate that whisper-medium achieved a best word error rate (WER) of 37.85%, while whisper-large-v3 consistently outperformed it, reaching a minimum WER of 33.93%. In the light of these findings, increased model size, combined with targeted fine-tuning on normalized dialectical data, significantly improves recognition accuracy, indicating that careful handling of orthographic and dialectical variation provides an effective path for ASR adaptation to low-resource varieties.
Structural Divergence under Shared Language-Level Specification: Griko in Universal Dependencies
Stavros Bompolas | Emanuela Pinna | Josep Quer | Marika Lekakou | Stella Markantonatou
Stavros Bompolas | Emanuela Pinna | Josep Quer | Marika Lekakou | Stella Markantonatou
Dialectal varieties pose major challenges for NLP resource development, especially when annotation frameworks are organized around standardized language specifications. In Universal Dependencies (UD), dialects without independent ISO codes are subsumed under the corresponding standard language and inherit its language-level documentation, validator settings, and grammatical inventories. This paper examines Griko, a Greek variety spoken in southern Italy that developed in relative isolation from the Modern Greek dialect continuum while remaining in long-term contact with local Italo-Romance varieties. We assess the consequences of this organizational structure through controlled parsing experiments comparing intra-dialectal training, cross-dialectal transfer from Standard Modern Greek (SMG), script-controlled transfer using romanized SMG, and contact-related cross-lingual transfer from Italian. Our results show that, before romanization, the Italian model even surpasses SMG on several UD metrics and that, although romanization substantially improves SMG-based transfer, performance still remains far below the intra-dialectal baseline. We argue that this persistent gap reflects the interaction between structural divergence and language-level validation constraints, a phenomenon we term ISO-based validation coupling. Through analyses of auxiliary systems, voice marking, and progressive constructions, we show how standard-centric validation architectures can constrain the representation of dialect-specific grammar. More broadly, the Griko case highlights the limitations of language-centric organization in UD and underscores the need for variety-sensitive mechanisms when extending universal annotation frameworks to structurally divergent dialects.
Digital Preservation of Aromanian Through Knowledge Management and Automatic Speech Recognition Evaluation
Marija Pendevska | Hristina Nastevska
Marija Pendevska | Hristina Nastevska
This paper presents a knowledge management framework for the digital preservation of Aromanian, an endangered Eastern Romance language spoken across the Balkans, combined with the first systematic evaluation of automatic speech recognition (ASR) models on Aromanian dialectal speech. The proposed three-module framework encompasses localization of existing resources, distribution through digital platforms, and creation of new linguistic content through technological innovation. Empirical analysis of knowledge management practices among 176 respondents validates the framework design, revealing that information quality (28.3%) and personal utility (24.1%) are the most valued criteria for knowledge sharing, while digital information seeking (27.1%) is the dominant behaviour. Within this framework, we evaluate models from the OpenAI Whisper family across three sizes (medium, large-v2, large-v3) and multiple language settings on two Aromanian varieties: Gramosteanj and Crushova. All configurations yield word error rates (WER) above 88%, with character error rates (CER) as low as 34% under optimal conditions, indicating partial phonotactic capture despite word-level failure. The Latin language setting with large-v3 consistently achieves the best results. Romanization of non-Latin script output substantially reduces CER, confirming script mismatch as a major error source. These findings underscore the limitations of current pretrained ASR models for endangered languages and the need for dedicated resources and adaptation strategies within a broader language preservation framework
A Novel Typology of Mutually Intelligible Words: The Case of Slavic Languages
Edward Klyshinsky | Yulia Badryzlova
Edward Klyshinsky | Yulia Badryzlova
In this paper, we demonstrate that using the notion of cognate in the task of evaluating mutual intelligibility (MI) in closely related languages can be confusing. We suggest a new term – percipiants – which handles MI of words irrespective of whether they have a common origin. We propose four classes of percipiants,which are differentiated by the degree and type of the closeness of meanings within the word pair. Furthermore, we claim that MI of individual words across a set of languages may be established computationally via normalized Levenshtein Distance (LD). We verify our hypotheses by analyzing data from a psycholinguistic experiment where the respondents were to predict words missing in a text in their native language. In the experimental condition, the respondents had access to the text in their native language and in a language from the same language group (Slavic); in the control group, the respondents were performing the same test in the presence of their native language only. The analysis demonstrates that (a) psycholinguistically, MI may be defined as the difference between the average correctness of answers in the experimental and the control groups; (b) normalized LD may serve as an adequate predictor of the experimentally measured MI; (c) this MI corroborates the four classes of percipiants; (d) contextual factors weaken the predictive force of LD, requiring further investigation.
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
Sercan Karakas
Sercan Karakas
This paper presents new resources and baselines for Dependency Parsing in Pomak, an endangered Eastern South Slavic language with substantial dialectal variation and no widely adopted standard. We focus on the variety spoken in Turkey (Uzunköprü) and ask how well a dependency parser trained on the existing Pomak Universal Dependencies treebank, which was built primarily from the variety that is spoken in Greece, transfers across dialects. We run two experimental phases. First, we train a parser on the Greek-variety UD data and evaluate zero-shot transfer to Turkish-variety Pomak, quantifying the impact of phonological and morphosyntactic differences. Second, we introduce a new manually annotated Turkish-variety Pomak corpus of 650 sentences and show that, despite its small size, targeted fine-tuning substantially improves accuracy; performance is further boosted by cross-variety transfer learning that combines the two dialects.
up
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Jin Zhao | Claire Benet Post | Elizabeth Hoefer
Jin Zhao | Claire Benet Post | Elizabeth Hoefer
CxGr-AMR: Extending Abstract Meaning Representation Beyond Lexically Anchored Relations with Constructional Rolesets
Claire Bonial | Claire Benet Post | Paul Van Eecke | Katrien Beuls | Harish Tayyar Madabushi
Claire Bonial | Claire Benet Post | Paul Van Eecke | Katrien Beuls | Harish Tayyar Madabushi
Current Abstract Meaning Representation (AMR) annotation guidelines, which largely tie argument structure to lexical rolesets, systematically misrepresent cases in which key semantic roles stem from clause-level structure rather than the verb, leaving these meanings either unnaturally attached, incorrect, or unexpressed. To address this limitation, we present CxGr-AMR, a novel extension of AMR that captures the semantics of various types of phrasal constructions, including argument structure constructions. We first examine how such cases are handled under current Standard-AMR guidelines and show that these analyses are often inadequate when constructionally contributed roles clash with those assigned by the verb. We then provide a theoretical grounding for our CxGr-AMR rolesets that lay out the relationship between the syntactic signatures of constructional slots and particular semantic roles associated with them. Finally, we develop an annotation-expert-in-the-loop pipeline for the semi-automatic annotation of sentences, and release a dataset containing 355 instances of phrasal constructions annotated with both Standard and CxGr-AMR.
Adding Aspectual Information to Structured Meaning Representations
Claire Benet Post | Paul Bontempo | August Ulfelder Milliken | Alvin Po-Chun Chen | Nicholas Derby | Saksham Khatwani | Sumeyye Nabieva | Karthik Sairam | Alexis Palmer
Claire Benet Post | Paul Bontempo | August Ulfelder Milliken | Alvin Po-Chun Chen | Nicholas Derby | Saksham Khatwani | Sumeyye Nabieva | Karthik Sairam | Alexis Palmer
To fully capture the meaning of a sentence, semantic representations should encode aspect, which describes the internal temporal structure of events. In graph-based meaning representation frameworks such as Uniform Meaning Representations (UMR), aspect lets one know how events unfold over time, including distinctions such as states, activities, and completed events. Despite its importance, aspect remains sparsely annotated across semantic meaning representation frameworks. This has, in turn, hindered not only current manual annotation, but also the development of automatic systems capable of predicting aspectual information. In this paper, we introduce a new dataset of English sentences annotated with UMR aspect labels over Abstract Meaning Representation (AMR) graphs that lack the feature. We describe the annotation scheme and guidelines used to label eventive predicates according to the UMR aspect lattice, as well as the annotation pipeline used to ensure consistency and quality across annotators through a multi-step adjudication process. To demonstrate the utility of our dataset for future automation, we perform simple baseline experiments using three modeling approaches. Our results establish initial benchmarks for automatic UMR aspect prediction and provide a foundation for integrating aspect into semantic meaning representations more broadly.
Modelling Idiomatic Expressions in Abstract Meaning Representation
Venera Gareeva | Johannes Heinecke
Venera Gareeva | Johannes Heinecke
Idiomatic expressions, a subclass of multiword expressions (MWE), pose persistent challenges for semantic parsing, as their meanings often diverge from the compositional semantics of their constituent words and depend strongly on contextual cues. While Abstract Meaning Representation (AMR) parsers aim to capture sentence-level semantics in a structured graph form, existing datasets provide limited coverage of idiomatic language, constraining their ability to model such expressions accurately. To address this gap, we extended a subset of the MAGPIE dataset by constructing a corpus of potentially idiomatic expressions (PIE) annotated with their corresponding AMR graphs. The dataset includes both naturally occurring and synthetically generated sentences, covering idioms in literal and idiomatic contexts. We fine-tune a state-of-the-art AMR parser on this dataset and evaluate its capacity to generate context-sensitive graphs that correctly reflect idiomatic versus literal interpretations. Our results show that standard parsers often capture only literal meanings of such expressions, while fine-tuning on our dataset improves alignment with the intended interpretations.
Abstract Meaning Representation (AMR) is a graph-based semantic representation which captures the core elements of meaning of a text. AMR has been incorporated into a variety of downstream tasks, which rely heavily on the availability of gold-annotated AMR corpora. While the annotation process is fairly lightweight, annotator training is still required even for linguists due to the extensive nature of the annotation guidelines and comprehensive set of roles. Therefore, all corpus development projects for AMR (and extensions of AMR) require the dataset curators to first train annotators. In this paper, we develop an online AMR annotation training system called TrAinMR in order to ease this training process and thus motivate the development of additional AMR corpora. The two main components of TrAinMR are (1) a written tutorial covering the basics of AMR annotation, and (2) an interactive practice module with corrective feedback. To measure the effectiveness of this tool, we conduct two pilot studies with five human annotators each. We find that the majority of annotators state their understanding of AMR improved as a result of TrAinMR, and some annotators show a positive trend in SMATCH scores after completing the practice module.
Named Entity Recognition for Persian Literary Text: A Case Study on The Little Prince
Minoo Nassajian | Joakim Nivre | Daniel Zeman
Minoo Nassajian | Joakim Nivre | Daniel Zeman
Existing Persian named entity recognition (NER) research has focused predominantly on news and social media domains, leaving literary texts—with their distinct linguistic characteristics—virtually unexplored. This paper addresses this gap by developing a new literary NER corpus using the Persian translation of The Little Prince story and evaluating existing state-of-the-art Persian NER tools on this corpus, trained exclusively on news and social media corpora. Our analysis reveals significant performance degradation on literary text, identifying systematic errors related to narrative-specific entities, metaphorical language, and discourse structures that challenge conventional NER approaches.
Towards Consistent UMR Annotation of Deverbal Nouns: Evidence from Czech and Latin
Hana Hledíková | Federica Gamba | Marketa Lopatkova | Jan Štěpánek
Hana Hledíková | Federica Gamba | Marketa Lopatkova | Jan Štěpánek
Deverbal nouns pose challenges for semantic annotation frameworks that aim to represent event structures consistently across lexical categories. This paper examines problematic phenomena in the annotation of deverbal nouns in Czech and Latin within the Universal Meaning Representation (UMR) framework, addressing both manual graph construction and rule-based automatic conversion from existing resources. Current UMR guidelines lack operational criteria for deciding when a noun should be treated as an eventive concept, particularly in the absence of a PropBank-like lexicon with sufficient nominal coverage. We therefore propose practical annotation principles: deverbal nouns denoting events (such as učení ‘teaching’), results of events (řešení ‘solution’), or event participants (učitel ‘teacher’) should be related to underlying event concepts (represented as verbs in their particular senses, i.e., učit-001 ‘to teach’, vyřešit-001 ‘to solve’, and učit-001 ‘to teach’, respectively), while other deverbal nouns should remain unrelated to respective events (such as učebna ‘teaching room’). To reduce inter-annotator variation, we further suggest systematic strategies for selecting verbal labels, including the use of light-verb constructions, synonymous verbs, and a preference for imperfective verbs in Czech aspectual pairs. For automatic conversion, we outline a rule-based approach that combines multiple lexical resources and frequency-based heuristics to identify corresponding verb senses. Our findings provide guidelines for more consistent UMR annotation across languages.
SAVI: Web-based Multilayered Semantic Annotation Validation Interface
Sashank Tatavolu | Soma Paul | Pratibha Rani | Sukhada Sukhada
Sashank Tatavolu | Soma Paul | Pratibha Rani | Sukhada Sukhada
This paper presents SAVI, a web-based interface for multilayer semantic annotation validation of Universal Semantic Representation (USR). USR encodes meaning across interdependent lexical, constructional, relational, discourse, and co-reference layers, making validation challenging using conventional annotation tools. SAVI addresses this limitation through structured tab-based layer separation, constraint-aware editing mechanisms, and role-based review workflows. The system integrates a multilingual concept dictionary to ensure sense-level consistency, along with a Hindi text-generation module and dependency-based visualization to support interpretation and correction. SAVI is implemented using a Flask backend, Flutter frontend, and PostgreSQL for structured data management. Evaluation results demonstrate effective governance of concept proposals and improved efficiency in multilayer USR correction, positioning SAVI as a structured validation framework for scalable semantic corpus development.
Finding Meaning in Embeddings: Concept Separation Curves
Paul Keuren | Marc Ponsen | Robert Ayoub Bagheri
Paul Keuren | Marc Ponsen | Robert Ayoub Bagheri
Sentence embedding techniques aim to encode key concepts of a sentence’s meaning in a vector space. However, the majority of evaluation approaches for sentence embedding quality rely on the use of additional classifiers or downstream tasks. These additional components make it unclear whether good results stem from the embedding itself or from the classifier’s behaviour. In this paper, we propose a novel method for evaluating the effectiveness of sentence embedding methods in capturing sentence-level concepts. Our approach is classifier-independent, allowing for an objective assessment of the model’s performance. The approach adopted in this study involves the systematic introduction of syntactic noise and semantic negations into sentences, with the subsequent quantification of their relative effects on the resulting embeddings. The visualisation of these effects is facilitated by Concept Separation Curves, which show the model’s capacity to differentiate between conceptual and surface-level variations. By leveraging data from multiple domains, employing both Dutch and English languages, and examining sentence lengths, this study offers a compelling demonstration that Concept Separation Curves provide an interpretable, reproducible, and cross-model approach for evaluating the conceptual stability of sentence embeddings. The open-source code and a live interactive demo are available upon acceptance.
Extracting First Order Logic formulas from graphical semantic representations
Rémi de Vergnette | Vincent Tourneur | Maxime Amblard
Rémi de Vergnette | Vincent Tourneur | Maxime Amblard
In this paper, we present a method for interpreting Yarn structures as logical formulas in a modal first order logic with temporality. Yarn is a recent semantic formalism that aims to bridge the gap between graph-based and logic-based semantic representations, providing a flexible and expressive framework for capturing the meaning of natural language utterances. Our approach translates the elements of Yarn structures such as predicates, features, into corresponding logical constructs, allowing for an interpretation of the represented meaning. Given that Yarn allows ambiguous representations, we associate to each Yarn structure a set of possible interpretations. We account for a range of semantic phenomena, extending beyond ambiguity to capture aspects of dynamic quantification as well. This work contributes to the understanding of the expressive power of graphical semantic representations and their relationship to formal logic.
Meaning Representations as Variational Quantum Circuits
Tilen Gaetano Limbäck-Stokin | Tanishka A. Birdavade | Kin Ian Lo | Mehrnoosh Sadrzadeh
Tilen Gaetano Limbäck-Stokin | Tanishka A. Birdavade | Kin Ian Lo | Mehrnoosh Sadrzadeh
Large language and vision-language models (VLMs) struggle with a ‘compositionality gap’. They treat language as a sequence of tokens lacking any structure and thus rely on a large number of parameters making them computationally expensive. To address these issues, we propose CCG-VQC, a quantum framework that unifies statistical distributions with linguistic structure. Guided by Combinatory Categorial Grammar, our model maps syntactic rules into parametrised quantum circuits and models sentences as quantum states. We evaluate CCG-VQC on structural VLM benchmarks such as ARO and SVO-Swap. Our experiments show that CCG-VQC consistently outperforms a quantum bag-of-words model, as well as classical VLMs such as CLIP and OpenCLIP. CCG-VQC achieved 71.19% accuracy on ARO-Attribution, significantly outperforming the parameter-matched MicroCLIP, which struggled to surpass random chance with a maximum performance of 50.85%.
Semantic frame, role, and relation labels are important parts of symbolic meaning representations. The mainstream approach is to define large language-specific lexicons that map predicate senses to their frame and argument labels, and to use separate label inventories for modifiers. Maintaining lexicons is very labor-intensive and scales poorly to the multilingual case. There are schemas that aim to simplify the annotation task using a more coarse-grained inventory of frames and roles, but suffer from a lack of systematicity and clear definitions. We present a schema that 1) uses a small inventory of frames and roles for lexicon-free annotation, 2) systematizes the frame inventory by factoring out aspect and mode, 3) has a unified vocabulary for arguments and modifiers, and 4) is designed to be annotated atop Universal Dependencies syntactic annotation. We argue for the adoption of such a schema for multilingual annotation and demonstrate promising results in annotation experiments on German.
First Shared Task on UMR Parsing
Jan Štěpánek | Daniel Zeman | Marketa Lopatkova | Federica Gamba | Hana Hledíková | Nianwen Xue
Jan Štěpánek | Daniel Zeman | Marketa Lopatkova | Federica Gamba | Hana Hledíková | Nianwen Xue
The paper presents the first shared task on parsing Uniform Meaning Representation (UMR), a graph-based framework for cross-linguistic semantic annotation of typologically diverse languages. The task requires systems to enrich plain text with sentence-level structure, node–token alignment, and document-level relations. It involves processing data for seven languages from four language families (Indo-European, Sino-Tibetan, Na-Dene, and Algic). Six languages have at least some training data; for one language, data is not available, leading to a zero-shot scenario. The training dataset as well as the gold-standard test set for all seven languages is released and made available for follow-up research. We present the task setup and evaluation methodology, using two graph matching approaches – a traditional, and an alignment-sensitive one, tailored specifically for UMR. Two participating systems are compared, each representing different modeling approaches. Results highlight the challenges of UMR parsing, particularly for alignment prediction and document-level semantics, and reveal substantial variation across languages and annotation conditions.
Uniform Meaning Representation (UMR) is a novel meaning representation formalism emanating from Abstract Meaning Representation (AMR). Since it is more complex than AMR, including document level annotation it is more difficult to create a parsing pipeline which can predict an UMR document from a set of consecutive sentences. The UMR Parsing Shared Task was created to compare different approaches. We decided to use a 2-step approach to predict sentence level and document level annotation. Since the available data was limited, we opted for a multilingual model, even though unlike AMR, in UMR the concepts of the meaning graph are not drawn from a single source, but from language dependend resources. Our final score was 19.35%, 0.08 points behind the best participant (19.43%).
Sema System for the DMR 2026 Shared Task: Multistage UMR Parsing with Qwen3-4B
Rémi de Vergnette | Maxime Amblard
Rémi de Vergnette | Maxime Amblard
We present the Sema system for the DMR 2026 shared task on parsing from natural language to . Our approach relies on parameter-efficient fine-tuning of Qwen3-4B with a multistage training procedure. We first train on a capped subset of the noisy training data, then continue training on the clean split, and finally fine-tune a dedicated stage for word-to-node alignment prediction. The system generates sentence-level graphs, selected document-level information, and alignments in separate steps, followed by rule-based post-processing to satisfy the official evaluation format. Results show that the approach is viable across several languages and exhibits promising transfer to Italian despite the absence of Italian data for fine-tuning , while very low-resource languages remain challenging.
Meaning Annotation Experience. A Tribute to Petr Sgall
Marie Mikulová | Jan Štěpánek | Barbora Štěpánková | Jarmila Panevova | Eva Hajicova
Marie Mikulová | Jan Štěpánek | Barbora Štěpánková | Jarmila Panevova | Eva Hajicova
We present ongoing work on annotating fine-grained semantic distinctions for circumstantial meanings, focusing on spatial expressions. We describe our theoretical background, and annotation process, as well as how we evaluate the results obtained. Using multiple independent annotations across the 3-million-token, genre-diverse Prague Dependency Treebank – Consolidated corpus of Czech data, we analyse inter-annotator agreement, recurrent disagreement patterns, and the limits of semantic categorization. Our results highlight the inherent vagueness of linguistic meaning. We also propose strategies for handling disagreement, such as weighted annotations, intermediate labels, and fuzzy labels that preserve annotation nuance. This work builds on the legacy of Petr Sgall and the Functional Generative Description theory that underpins the multi-layer form–meaning framework.
Extending Uniform Meaning Representation to Persian: The First Corpus Resource
Minoo Nassajian | Daniel Zeman
Minoo Nassajian | Daniel Zeman
Uniform Meaning Representation (UMR) has cross-linguistic design principles that make it particularly well-suited as a semantic representation framework for capturing all language-specific phenomena. Despite its growing adoption, no UMR corpus currently exists for Persian. In this paper, we present the first version of a Persian UMR dataset created through a rule-based conversion of existing Persian AMR annotations from The Little Prince corpus, followed by manual mapping of split semantic roles from AMR to their finer-grained UMR counterparts. We report detailed statistics on the conversion, analyze the challenges of mapping Persian AMR structures into UMR, and provide illustrative examples. The resource is freely available and it lays the groundwork for subsequent enrichment of Persian UMR with additional semantic layers, including co-reference, named entities, and discourse relations.
Regression-Tested Compositional Semantics: A Graphical Development Environment for Glue and Description-by-Analysis
Mark-Matthias Zymla | Kascha Kruschwitz
Mark-Matthias Zymla | Kascha Kruschwitz
We present a platform for developing compositional semantic annotations for formal syntactic representations that allows users to interact with and explore annotations and to track their progress and quality. For this, we provide several forms of visualizations and take inspiration from research in linguistic treebanking. Thus, we contribute to the development of formal semantic parsers and corresponding meaning banks. The system is designed with a regression testing paradigm in mind and provides support for NLI so that the created semantic resources can be developed and validated in a task-driven environment. We defend this paradigm in comparison to modern approaches to semantic parsing that are mainly evaluated on the basis of gold standard annotations.
up
Proceedings of Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science (DTF) @ LREC 2026
Proceedings of Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science (DTF) @ LREC 2026
Florian Barth | Keli Du | José Calvo Tello | Philippe Genêt | Piroska Lendvai | Christof Schöch | Thorsten Trippel
Florian Barth | Keli Du | José Calvo Tello | Philippe Genêt | Piroska Lendvai | Christof Schöch | Thorsten Trippel
Derived Text Formats as Strategic Transformations of In-Copyright Materials to Support Open Science: A Survey
Christof Schöch
Christof Schöch
Derived Text Formats (DTFs) are the result of a strategic transformation of textual materials that are protected by copyright in their original form, such that the resulting data is useful for computational analyses and can be openly shared following best practices of Open Science without infringing copyright law. This paper aims to provide insights into several key aspects of this concept that is closely related to concepts such as corpus masking, non-consumptive research and extracted features. The paper establishes the motivation for using DTFs, discusses several foundational aspects of the concept and practice, describes ongoing research on issues including copyright, reconstructibility, evaluation and standardization of DTFs, and concludes with a roadmap for future work on DTFs. In this way, this paper provides a broad but concise overview of work on DTFs as a contribution to Open Science practices, with a focus on work in the Digital Humanities.
Derived Text Formats (DTFs) have been proposed as a solution to enable text and data mining while avoiding copyright infringement. Building on a review of recent empirical studies of DTFs on topic modeling, authorship classification, and sentiment analysis, this paper argues that DTFs should not be treated as static formats, but as variable and task-dependent representations shaped by multiple interacting factors. In response, we propose a multi-dimensional framework that conceptualizes DTFs as configurations within a structured space defined by both internal representation parameters and external constraints. The framework includes four internal representation dimensions—feature level, degree of reduction, transformation strategy, and aggregation level—as well as two external constraining forces: legal requirements and task-specific information needs. By emphasizing the interdependence of these dimensions, the proposed framework provides a systematic way to describe, compare, and design DTFs across different analytical contexts. Therefore, this paper contributes to a more theoretically grounded understanding of DTFs and offers guidance for their responsible and effective use in text and data mining in Digital Humanities.
Legal implications of Derived Text Formats - a copyright perspective
Gianna Iacino | Pawel Kamocki | Keli Du
Gianna Iacino | Pawel Kamocki | Keli Du
Text and Data Mining (TDM) methods are often used in order to analyse large amounts of text for scientific research. If the analysed text is protected by copyright, the use of such TDM methods has copyright implications. The existing copyright exceptions facilitate TDM within a narrow framework which limits the storage, publication and re-use of datasets. This paper examines the legal framework of converting the source text into a derived text format (DTF) which is no longer protected by copyright in order to allow the use of TDM without legal restrictions. First, the creation itself of a DTF is being examined: it entails copyright relevant acts which are covered by the TDM exception. In a second step the copyright status of the created DTF has to be evaluated based on three criteria: the DTF may not contain elements which are an expression of the intellectual creation of the author of the source material, the source material may not be easily reconstructable based on the DTF and the source material may not be recognizable.
Revisiting Masking After Fifteen Years: Early Approaches to Non-Reconstructable Linguistic Data in the current context
Georg Rehm | Thorsten Trippel | Andreas Witt
Georg Rehm | Thorsten Trippel | Andreas Witt
This paper revisits the masking approaches introduced in 2007 for enabling the distribution of linguistically annotated corpora without exposing copyrighted or sensitive source texts and situates them within the contemporary framework of Derived Text Formats (DTF). While the original work demonstrated how syntactic and morphological information could be preserved through parameterised masking, today’s landscape, which is shaped by large language models, FAIR requirements, and emerging standardisation efforts, demands more formalised, robust and reproducible methods. We outline how DTF extend early masking concepts by introducing explicit abstraction levels, reversibility classes, and machine- actionable provenance, supported by standards such as TEI, ISO linguistic annotation models, CMDI metadata, and the draft DIN DTF specification. Building on these foundations, we present a modern workflow for DTF generation, including enrichment pipelines, structural abstractions, statistical and embedding-based representations, and non-reversible transformation layers, illustrated through the MONA-pipe framework. We conclude that DTF constitute a sustainable and infrastructure - ready solution for open, reproducible and legally secure text-based research in the decades to come.
Multi-Label Text Classification of Derived Text Formats with DistilBERT
Jennifer Ecker | Roman Schneider
Jennifer Ecker | Roman Schneider
Derived Text Formats enable the distribution of copyrighted texts by systematically perturbing linguistic information to reduce reconstructability. However, the extent to which such information loss affects downstream text classification remains unclear. We investigate how controlled perturbations affect learning dynamics in transformer-based classification using two datasets and two strategies: POS-consistent replacement of 30%, 40%, and 50% of tokens, and random word-order shuffling. On Wikipedia data, POS replacement increases loss by 4-9% and reduces micro-F1 by 3-8%, depending on the replacement rate, while shuffling raises loss by 5% and lowers micro-F1 by 4%. Performance degrades monotonically with higher replacement rates, and shuffling yields results between the 30% and 40% conditions, indicating that DistilBERT relies more on lexical semantics than on word order. Experiments on specialist-domain data show the same pattern, demonstrating robustness across domains. To test cross-representation generalization, we train classifiers on both clean and perturbed texts and evaluate them on the respective alternate representation. Models trained on DTF data generalize better to clean text than vice versa, suggesting that perturbation-based training promotes more robust representations. Our findings position DTF as a promising strategy for reproducible, legally compliant, and robust NLP research.
Training data generation for context-dependent rubric-based short answer grading
Pavel Šindelář | Filip Prášil | Dávid Slivka | Christopher Bouma | Ondrej Bojar
Pavel Šindelář | Filip Prášil | Dávid Slivka | Christopher Bouma | Ondrej Bojar
Every four years, the PISA test is administered by the OECD to test the knowledge of teenage students worldwide and allow for comparisons of educational systems. However, having to avoid language differences and annotator bias makes the grading of student answers challenging. For these reasons, it would be interesting to consider methods of automatic student answer grading. To train some of these methods, which require machine learning, or to compute parameters or select hyperparameters for those that do not, a large amount of domain-specific data is needed. In this work, we explore a small number of methods for creating a large-scale training dataset using only a relatively small confidential dataset as a reference, leveraging a set of very simple derived text formats to preserve confidentiality. Using the proposed methods, we successfully created three surrogate datasets that are, at the very least, superficially more similar to the reference dataset than a straightforward result of prompt-based generation. Early experiments suggest one of these approaches might also lead to improved training of automatic answer grading models.
DUO_DE A1: An Annotated Corpus of Online Learning Material for Beginning Learners of German as a Foreign Language
Jammila Laâguidi | Vitaliia Ruban | Ronja Laarmann-Quante | Anastasia Drackert
Jammila Laâguidi | Vitaliia Ruban | Ronja Laarmann-Quante | Anastasia Drackert
This paper describes the creation of DUO_DE A1, a corpus based on A1-level learning material from the Deutsch-Uni Online (DUO) language courses for German as a foreign language. We split the material into small segments and manually annotated each with fine-grained information such as the type of segment (e.g. task description, description of grammar), the medium (e.g. text, table, audio), the text units it contains (e.g. words, phrases, sentences) and other special features (e.g. marking cloze texts). Furthermore, we automatically tokenized, POS tagged and lemmatized the corpus and compared the performance of three models on these steps for different kinds of segments. We publish the created corpus in a manner that respects copyright, releasing all structural features, metadata and POS tags.
This paper explores the limitations of reconstructing scrambled text within the context of Derived Text Formats (DTFs). While previous research has treated reconstruction as a technical challenge, this study shifts the focus to investigating the causes of reconstruction failure. Through a detailed analysis of outputs generated by language models on non-literary (IMDb reviews) and literary (Gutenberg texts) datasets, several systematic patterns were identified. First, reconstructed texts are generally shorter than the originals, indicating that the generated results are often incomplete. Second, models simplify expressions by omitting specific modifiers, thereby producing more general outputs. Third, high similarity at the string level does not guarantee semantic equivalence, revealing fidelity-related issues in text reconstruction. In literary texts, chunk-based segmentation poses additional challenges; this approach disrupts syntactic and contextual coherence, leading to sentences that are structurally correct but semantically distorted. These findings suggest that reconstruction difficulty is not merely a matter of model performance but also reflects the importance of higher-level textual organization. This study highlights the fundamental limitations of current language models and reframes reconstruction failure as an analytical perspective for understanding how meaning is constructed in text.
DIN 19461: A National Standard for Derived Text Formats
Thorsten Trippel | Florian Barth | Jose Calvo Tello | Keli Du | Philippe Genêt | Daniel Kurzawe | Peter Leinen | Piroska Lendvai | Christof Schöch | Andreas Witt | Arden Zimmermann
Thorsten Trippel | Florian Barth | Jose Calvo Tello | Keli Du | Philippe Genêt | Daniel Kurzawe | Peter Leinen | Piroska Lendvai | Christof Schöch | Andreas Witt | Arden Zimmermann
We present DIN 19461:2026-06 (E), a German draft national standard that defines categories, terminology, and process requirements for Derived Text Formats (DTFs) created from text documents in natural language. The standard specifies enrichment and information reduction operations, requirements for combining multiple DTFs, and documentation obligations for publication, archiving, and reuse. Its aim is to enable legally compliant sharing and analysis of texts–especially where copyright or data protection prevents distributing originals–while maintaining scientific utility and reproducibility through explicit process and parameter recording. We outline the scope, the key concepts, the four core reduction operations (retain, delete, replace, randomise), together with examples across token-, structure-, and vector-based DTFs, and implications for infrastructures (e.g., ISO 24622-based metadata). Finally, we discuss limitations, open questions (e.g., reconstruction risks with modern ML models), and next steps for adoption and maintenance.
up
The 7th Financial Narrative Processing Workshop
The 7th Financial Narrative Processing Workshop
Mo El-Haj | Antonio Moreno Sandoval | Ana Garcia-Serrano | Chung-Chi Chen | Paul Rayson | Yanco Amor Torterolo Orta | Paloma Martinez | Jordi Porta
Mo El-Haj | Antonio Moreno Sandoval | Ana Garcia-Serrano | Chung-Chi Chen | Paul Rayson | Yanco Amor Torterolo Orta | Paloma Martinez | Jordi Porta
LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank
Serhii Hamotskyi | Akash Kumar Gautam | Christian Hänig
Serhii Hamotskyi | Akash Kumar Gautam | Christian Hänig
Verifying the eligibility of securities as collateral is a key responsibility of the Deutsche Bundesbank. However, manually verifying these assets against legal and financial criteria within lengthy, semi-structured, and often bilingual prospectuses is a resource-intensive task. While previous efforts utilized traditional Named Entity Recognition (NER) for information extraction, these methods often struggle with OCR noise, linguistic variance, and rigid span-based constraints, as well as requiring manual annotation of documents to generate adequate training data for all the required annotation types. In this paper, we present the first case study applying Large Language Models (LLMs) to the eligibility examination process, shifting the paradigm toward a generative Information Extraction pipeline. Our approach decomposes the task into extraction, normalization, and interpretation, allowing for greater flexibility in handling noisy text and interleaved German-English content. We further introduce a value-based evaluation methodology using LLM-as-a-judge, which offers a more semantic assessment than offset-based metrics. Our results demonstrate that LLM-based systems achieve high precision (up to 91%) in document-level eligibility, exhibiting a conservative operating profile that minimizes false acceptance.
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
Virginie Mouilleron | Théo Lasnier | Anna Mosolova | Djamé Seddah
Virginie Mouilleron | Théo Lasnier | Anna Mosolova | Djamé Seddah
Vision-language models (VLMs) perform well on many document understanding tasks, yet their reliability in specialized, non-English domains remains underexplored. This gap is especially critical in finance, where documents mix dense regulatory text, numerical tables, and visual charts, and where extraction errors can have real-world consequences. We introduce SCRIBE FINANCE, the first multimodal benchmark for evaluating French financial document understanding. The dataset contains 1,204 expert-validated questions spanning text extraction, table comprehension, chart interpretation, and multi-turn conversational reasoning, drawn from real investment prospectuses, KIDs, and PRIIPs. We evaluate six open-weight VLMs (8B–124B parameters) using an LLM-as-judge protocol. While models achieve strong performance on text and table tasks (85–90% accuracy), they struggle with chart interpretation (34–62%). Most notably, multi-turn dialogue reveals a sharp failure mode: early mistakes propagate across turns, driving accuracy down to roughly 50% regardless of model size. These results show that current VLMs are effective for well-defined extraction tasks but remain brittle in interactive, multi-step financial analysis. SCRIBE FINANCE offers a challenging benchmark to measure and drive progress in this high-stakes setting.
CFQA: A Chinese Financial Question Answering Benchmark from Corporate Annual Reports
Tianning Zhu | Mo Liu | Murathan Kurfali
Tianning Zhu | Mo Liu | Murathan Kurfali
We present CFQA, a Chinese financial question answering benchmark constructed from 50 publicly listed companies’ annual reports spanning 2023–2025. The benchmark comprises 500 questions, derived by applying 10 question templates to each source document, and covers five categories: fact extraction, enumeration, comparative calculation, judgment verification, and reasoning analysis. All gold-standard answers are manually annotated and grounded in the source reports. To illustrate benchmark utility, we evaluate a retrieval-augmented generation (RAG) system against a no-retrieval baseline, and introduce a rule-based consistency detector that distinguishes fabricated content from other error types. RAG improves average answer accuracy from 7.53% to 8.07%, with the most consistent gains observed in fact extraction and judgment verification tasks for domain-adapted models. Crucially, by decoupling exact-match accuracy from evidence-support judgments, our detector reveals that despite low absolute scores, RAG architectures successfully constrain model confabulation, exhibiting remarkably low true fabrication rates. However, performance gains in higher-order cognitive tasks, such as comparative calculation and reasoning analysis, remain non-significant across evaluated models, highlighting the boundaries of current retrieval-augmented systems in complex financial reasoning. The dataset, annotation guidelines, and evaluation code are publicly released.
Verifiable Financial Enterprise Question Answering via Inference-Time Grounding and Traceability
Anubha Kabra | Katie Jooyoung Kim | Zhiwei Kou | Helene Sajer | Yimei Fan | Gabriel Martinez Vidiri
Anubha Kabra | Katie Jooyoung Kim | Zhiwei Kou | Helene Sajer | Yimei Fan | Gabriel Martinez Vidiri
Financial enterprise AI systems deployed in high-stakes settings require responses that are verifiable, traceable, and auditable. We introduce a modular, model- and data-agnostic inference-time control framework, together with a deployment-aware evaluation strategy for verifiable financial enterprise question answering. Our method enforces faithfulness at inference time without retraining or changes to retrieval infrastructure. We deploy our method in a production financial enterprise assistant and evaluate it using a combination of intrinsic faithfulness metrics, baseline comparisons, and real-world user feedback. Our approach improves groundedness by 29% over baselines, reduces hallucinations to near-zero levels, and achieves near-perfect document-span traceability. Together, our results demonstrate that modular pipeline design combined with detailed, deployment-aware evaluation provides a practical and effective path toward verifiable financial enterprise QA systems.
Environmental, Social and Governance Sentiment Analysis on Slovene News: A Novel Dataset and Models
Paula Dodig | Boshko Koloski | Katarina Sitar Šuštar | Senja Pollak | Matthew Purver
Paula Dodig | Boshko Koloski | Katarina Sitar Šuštar | Senja Pollak | Matthew Purver
Environmental, Social, and Governance (ESG) considerations are increasingly integral to assessing corporate performance, reputation, and long-term sustainability. Yet, reliable ESG ratings remain limited for smaller companies and emerging markets. We introduce the first publicly available Slovene ESG sentiment dataset and a suite of models for automatic ESG sentiment detection. The dataset, derived from the MaCoCu Slovene news collection, combines large language model (LLM)-assisted filtering with human annotation of company-related ESG content. We evaluate the performance of monolingual (SloBERTa) and multilingual (XLM-R) models, embedding-based classifiers (TabPFN), hierarchical ensemble architectures, and large language models. Results show that LLMs achieve the strongest performance on Environmental (Gemma3-27B, F1-macro: 0.61) and Social aspects (gpt-oss 20B, F1-macro: 0.45), while fine-tuned SloBERTa is the best model on Governance classification (F1-macro: 0.54). We then show in a small case study how the best-preforming classifier (gpt-oss) can be applied to investigate ESG aspects for selected companies across a long time frame.
Not All News Is Equal: Topic- and Event-Conditional Sentiment from Finetuned LLMs for Aluminum Price Forecasting
Alvaro Paredes Amorin | Andre Python | Christoph Weisser
Alvaro Paredes Amorin | Andre Python | Christoph Weisser
By capturing the prevailing sentiment and market mood, textual data has become increasingly vital for forecasting commodity prices, particularly in metal markets. However, the effectiveness of lightweight, finetuned large language models (LLMs) in extracting predictive signals for aluminum prices—and the specific market conditions under which these signals are most informative—remains under-explored. This study generates monthly sentiment scores from English and Chinese news headlines (Reuters, Dow Jones Newswires, and China News Service) and integrates them with traditional tabular data, including base metal indices, exchange rates, inflation rates, and energy prices. We evaluate the predictive performance and economic utility of these models through long-short simulations on the Shanghai Metal Exchange from 2007 to 2024. Our results demonstrate that during periods of high volatility, Long Short-Term Memory (LSTM) models incorporating sentiment data from a finetuned Qwen3 model (Sharpe ratio 1.04) significantly outperform baseline models using tabular data alone (Sharpe ratio 0.23). Subsequent analysis elucidates the nuanced roles of news sources, topics, and event types in aluminum price forecasting
Flipper: An Extended Document-Level Financial Dataset for Training and Evaluation with Annotated Discourse Phenomena
Mariam Nakhlé | Rachel Atherly | Gabriela nicole Gonzalez Saez | Marco Dinarelli | Raheel Qader | Hervé Blanchon
Mariam Nakhlé | Rachel Atherly | Gabriela nicole Gonzalez Saez | Marco Dinarelli | Raheel Qader | Hervé Blanchon
We present a new resource for Machine Translation (MT), namely a training and evaluation dataset containing parallel sections issued from authentic documents in the financial domain. We cover five language pairs: English-French, English-Spanish, English-German, English-Italian and French-Spanish. The total number of parallel sections is 122k and the number of tokens is 118M (source and target combined). MT has improved greatly in recent years, but certain phenomena still cause errors, particularly when context spans beyond a single sentence. Errors can lead to mistranslated pronouns, incorrect gender or number agreement, and inconsistent terminology, which can be especially problematic in high-stakes domains like finance. We therefore construct the dataset at document level (rather than sentence-level alignment) and also produce fine-grained annotations of context-sensitive phenomena. The annotation was performed using preexisting tools and custom scripts. The annotated phenomena are: formality, gender, terminology consistency, verb form and sentence reordering. This aims to improve document-level evaluation of MT models by enabling evaluation solely on texts containing a particular phenomenon of interest. Our primary contribution is the creation and public release of Flipper, a multilingual document-level parallel dataset in the financial domain, designed to support both training and targeted evaluation of context-sensitive machine translation.
TranslateGemma for ES-EN Financial Reports: Exploring Adaptability to Variable-Sized Contexts
Yanco Amor Torterolo Orta | Melina Chatzi | Antonio Moreno-Sandoval
Yanco Amor Torterolo Orta | Melina Chatzi | Antonio Moreno-Sandoval
This paper explores bidirectional financial Machine Translation (MT) between Spanish and English, focusing on the specialized domain of annual reports from IBEX 35 companies. Fine-tuned models are compared against zero-shot scenarios through a series of experiments, testing factors such as prompting strategies and model size. On the one hand, this work studies a combination of existing fine-tuning strategies aimed at improving the adaptability of MT models to variable-sized contexts, and, on the other hand, it analyzes the limitations detected in current evaluation metrics. Results are mixed: fine-tuned models show an improvement in both short and long-context scenarios in traditional metrics, while zero-shot predictions are clearly favored by neural metrics. In fact, reference-free assessment of the source and the human reference received worse scores than the off-the-shelf prediction models. Consequently, fine-tuning on the human-made dataset hardly improves the neural metrics against zero-shot generations. This suggests that neural metrics tend to favor the fluency of MT generations and literalness over creativity, among other technical limitations regarding long-context adaptability. From a practical standpoint, the low Translation Edit Rate (TER) scores suggest that specialized fine-tuning remains the most viable path for companies to implement efficient Machine Translation Post-Editing (MTPE) workflows, given the stylistic alignment.
LabelFusion: Fusing Large Language Models with Transformer Encoders for Robust Financial News Classification
Michael Schlee | Christoph Weisser | Timo Kivimäki | Melchizedek Mashiku | Benjamin Saefken
Michael Schlee | Christoph Weisser | Timo Kivimäki | Melchizedek Mashiku | Benjamin Saefken
Financial news plays a central role in shaping investor sentiment and short-term dynamics in commodity markets. Many downstream financial applications—such as commodity price prediction or sentiment modeling—therefore rely on the ability to automatically identify news articles that are relevant to specific assets. However, obtaining large labeled corpora for financial text classification tasks is costly, and transformer-based classifiers such as RoBERTa often degrade significantly in low-data regimes. Our results show that appropriately prompted out-of-the-box large language models (LLMs) achieve strong performance even in low-data regimes. Furthermore, we propose LabelFusion, a hybrid architecture that combines the output of a prompt-engineered LLM with contextual embeddings produced by a fine-tuned RoBERTa encoder through a lightweight multilayer perceptron (MLP) voting layer. Evaluated on a ten-class multi-label subset of the Reuters-21578 corpus, LabelFusion achieves a macro F1 score of 96.0% and an accuracy of 92.3% when trained on the full dataset, outperforming both standalone RoBERTa (F1 94.6%) and the standalone LLM (F1 93.9%). In low- to mid-data regimes, however, the LLM alone proves surprisingly competitive, achieving an F1 score of 75.9% even in a zero-shot setting and consistently outperforming LabelFusion until approximately 80% of the training data is available. These results suggest that LLM-only prompting represents the preferred strategy under annotation constraints, whereas LabelFusion becomes the most effective solution once sufficient labeled data is available to train the encoder component. The code is available in an anonymized repository.
LLM-as-a-Judge Evaluation of Financial News Articles Generated Based on Factors of Stock Price Fluctuation
Yurina Kosai | Yucheng Xie | Rikuto Tsuchida | Takehito Utsuro
Yurina Kosai | Yucheng Xie | Rikuto Tsuchida | Takehito Utsuro
This paper proposes an LLM-as-a-Judge evaluation framework of stock price fluctuation articles automatically generated based on financial news, corporate disclosures, and stock price fluctuation data. This automatic article generation framework emulates the workflow of human financial journalists by analyzing recent stock price fluctuations and incorporating relevant causal factors extracted from textual and numerical information. In particular, the generation process utilizes news articles and numerical stock price data, including price fluctuation ranges over the past three days. Based on those automatically generated stock price fluctuation articles, this study places particular emphasis on the LLM-as-a-Judge evaluation methodology. We conduct an item wise human evaluation and compare it with the LLM-as-a-Judge automatic metric. We analyze the correlation among these evaluation methods to assess their reliability. Furthermore, through comparisons between zero-shot and few-shot prompting, we examine the effectiveness of the proposed framework and the validity of LLM based evaluation for assessing factual and causal consistency in financial text generation.
The Financial Document Causality Detection Shared Task (FinCausal 2026)
Antonio Moreno-Sandoval | Jordi Porta | Yanco Amor Torterolo Orta | Alexia Stanescu | Melina Chatzi | Sofía Roseti
Antonio Moreno-Sandoval | Jordi Porta | Yanco Amor Torterolo Orta | Alexia Stanescu | Melina Chatzi | Sofía Roseti
The Financial Document Causality Detection shared task (FinCausal) is a competition organized within the Financial Narrative Processing (FNP) workshop series. It aims to identify the causal relationship between a question and its answer in a given financial context. The dataset is built from real annual reports drafted by Spanish IBEX 35 companies and several UK companies. The task includes two subtasks, one in English and one in Spanish. It is formulated as an Extractive Question-Answering (EQA) task in which, given a context (C) and a question (Q), participants must extract the verbatim answer span (A). The 2026 edition introduces several changes to increase task difficulty, including the reformulation of 10% of the questions to require deeper reasoning and a stronger emphasis on multi-step causal chains with three or more elements, achieved by removing overly simple cases and adding 500 new complex fragments per language. Another innovation is the adoption of an LLM-as-a-judge metric on a 1–5 scale, based on a rubric designed to align better with human preferences than Semantic Answer Similarity (SAS) and Exact Match (EM). This edition was hosted as part of the LREC conference in Palma de Mallorca, Spain.
Sheffield NLP at FinCausal 2026: A Comparative Study of RAG Approaches and Fine-Tuning for Causal Q&A in Financial Texts
Aali Abdullah Alqarni | Mark Stevenson | Arif Dwi Laksito
Aali Abdullah Alqarni | Mark Stevenson | Arif Dwi Laksito
This paper describes our approach to the FinCausal 2026 shared task, which addresses causal question answering from financial documents in English and Spanish. We investigated the effectiveness of fine-tuned generative models combined with Retrieval-Augmented Generation (RAG). Our approach compares five retrieval strategies across base and fine-tuned GPT-models (GPT-4.1-mini). RAG-based few-shot selection showed better performance than random sampling, particularly for the base model. In the FinCausal 2026 official run, this approach was ranked first in both the English and Spanish subtasks, obtaining LLM scores of 4.8140 and 4.8131 out of 5, respectively.
Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026
Akash Kumar Gautam | Serhii Hamotskyi | Christian Hänig
Akash Kumar Gautam | Serhii Hamotskyi | Christian Hänig
This paper describes team HSA_CORAL’s submission to the FinCausal 2026 shared task on extracting cause–effect relations from financial narratives via extractive question answering in English and Spanish. We compare three modeling families: (i) encoder-only token tagging with multilingual BERT, (ii) encoder–decoder generation with multilingual BART, and (iii) decoder-only LLMs (Llama 3.1 and GPT variants) using prompt refinement, few-shot demonstrations, and supervised fine-tuning. Across settings, prompting and few-shot examples yield competitive performance, but supervised fine-tuning is the main driver of improvement. Our best system, GPT-4.1 Mini fine-tuned on combined English and Spanish training data, achieves the highest (tied) score on English (score 4.8140) and ranks third on Spanish (score 4.7753) under the shared task’s LLM-as-a-judge metric. Overall, the results highlight the value of task-specific adaptation and multilingual fine-tuning for cross-lingual transfer in financial causality QA.
VERSA: Verbatim Extraction via Rephrasing and Self-Aggregation for Financial Causality
Aldan Jay | Rafael Berlanga | Yoelvis Moreno | Vicent Santamarta
Aldan Jay | Rafael Berlanga | Yoelvis Moreno | Vicent Santamarta
Financial causality detection,the task of identifying and extracting verbatim causal spans from financial narratives, remains a challenging problem in Natural Language Processing (NLP). Large Language Models (LLMs), while powerful reasoners, frequently paraphrase source text or produce imprecise span boundaries when used in zero-shot extraction settings, leading to poor Exact Match scores. In this paper, we present VERSA, our system for the FinCausal 2026 Shared Task, a multi-agent pipeline that integrates two complementary inference strategies: Rephrase-and-Respond (RaR) and Recursive Self-Aggregation (RSA). The pipeline decomposes the extraction task into five sequential stages, each handled by a specialised agent: (1) causal structure analysis, (2) question reformulation via RaR, (3) diverse candidate population generation, (4) iterative refinement through RSA, and (5) verbatim validation with word-boundary alignment. We evaluate our approach on both the English and Spanish subsets of the FinCausal 2026 dataset. An ablation study demonstrates the individual and combined ontributions of RaR and RSA, showing that the full pipeline substantially outperforms a zero-shot baseline in Exact Match and token-level F1.
SpanDiffusion: Flow Matching over Continuous Span Masks for Financial Causal Question Answering
Georg Niess | Roman Kern
Georg Niess | Roman Kern
We present SpanDiffusion, a continuous diffusion approach to extractive causal question answering for the FinCausal 2026 shared task. SpanDiffusion uses two Gaussian masks, continuous signals with peaks at the answer start and end positions, and learns to denoise them from pure noise through a dedicated transformer conditioned on frozen DeBERTa-v3-large embeddings with LoRA adapters (1.6M parameters). By replacing Denoising Diffusion Probabilistic Models (DDPM) with flow matching (rectified flow), we reduce denoising to only 20 Euler steps at inference. A systematic ablation across six diffusion variants and a span-classification baseline shows that LoRA adaptation is the dominant factor (+34 Exact Match points), followed by flow matching (+5.5 EM). However, the standard span classifier (85.8% EM) outperforms our best diffusion model (83.0% EM), suggesting that the denoiser does not yet justify its added complexity. We discuss tradeoffs between the interpretability of diffusion trajectories and classification accuracy.
Improving Verbatim Financial Causality Extraction with Supervised Fine-Tuning and Prompt Repetition
Sanae Attak | Mohammed Salah Chiadmi | Youssef Lamrani Alaoui
Sanae Attak | Mohammed Salah Chiadmi | Youssef Lamrani Alaoui
This paper investigates the application of generative Large Language Models (LLMs) for strict verbatim span extraction. We evaluate our methodology within the FinCausal 2026 shared task. Because generative LLMs optimize next-token probability rather than strict boundaries, they naturally suffer from over-generation and boundary drift in extraction tasks. To address this, we introduce a generalized structural training constraint, extending prompt repetition from a purely inference-time heuristic to a training-time supervision framework. By incorporating duplicated prompts directly into Supervised Fine-Tuning (SFT), we hypothesize that this encourages the model to internalize a form of unidirectional cross-reading behavior, leading to stronger alignment between generated spans and the source context for exact extraction. Evaluating on open-weights (Qwen2.5-14B-Instruct-1M) and proprietary (GPT-4.1-Nano) architectures, we find this soft attention constraint improves Exact Match scores for open models and helps balance cross-lingual performance disparities. Conversely, the proprietary model exhibited sensitivity to prompt duplication, achieving its highest score without repetition. Ultimately, our deterministic SFT approach secured 4th place in the Spanish subtask (4.73) and 6th place in the English subtask (4.70), indicating the viability of structurally simple, natively fine-tuned models compared to complex multi-stage pipelines.
LeedsMEng26: Qwen + Gemini for FinCausal 2026 Causality Detection in Financial Narrative Texts
Zaid Shahrouri | Ayomide Ivienagbor | Idrees Asad | Rijul Shrestha | Yasemin Bal | Zahaab Nadeem
Zaid Shahrouri | Ayomide Ivienagbor | Idrees Asad | Rijul Shrestha | Yasemin Bal | Zahaab Nadeem
This paper presents the LeedsMEng26 system for the FinCausal 2026 shared task (CITATION) on financial causality detection in narrative texts. The task is formulated as extractive question answering over English and Spanish financial reports, where systems must return a verbatim span from the context that answers an abstractive question about a cause or an effect. We propose a two-stage pipeline consisting of candidate span generation followed by span verification and boundary refinement under a strict extractiveness constraint. We evaluate both an extractive RoBERTa-based baseline and instruction-tuned large language models. Results show that Qwen-2.5-1.5B-Instruct is a stronger candidate generator than the RoBERTa baseline, and that a second-stage verifier further improves answer boundary accuracy and overall adequacy. Our best configuration, Qwen-2.5-1.5B-Instruct with Gemini-2.5-flash refinement, achieved an adequacy score of 4.7000 for English and 4.6143 for Spanish. These findings suggest that a modular generation-and-verification pipeline is effective for extractive financial causality detection.
Financial Causal QA via Instruction and Prompt Tuning of Gemma3-12B
Avinash Trivedi | Chindukuri Mallikarjuna
Avinash Trivedi | Chindukuri Mallikarjuna
In this paper we present a novel methodology that harnesses the power of prompt tuning applied directly to Gemma3-12B, a state-of-the-art generative large language model to enhance performance on complex natural language processing challenges. Instead of relying solely on extensive retraining, our approach leverages carefully crafted input prompts to steer the pre-trained Gemma-12B towards generating outputs with superior contextual accuracy and interpretability. Our experimental evaluation employed a composite LLM Score metric that quantifies both semantic coherence and relevance; under this framework, our system (Team Name: Sarang) achieved a score of 4.54, ranking 9th in the shared task. Furthermore, in the competitive task evaluation, our method demonstrated the potential of prompt tuning as a viable alternative to traditional fine-tuning approaches. This study not only demonstrates the practical benefits of integrating prompt engineering with large language models but also opens avenues for future research aimed at further optimizing model performance in domain-specific applications.
QRAFT: QLoRA Retrieval-Augmented Fine-Tuning for Causal Span Extraction in Financial Documents
Bavya Sarda | Pulkit Chatwal | Sonal Dabral
Bavya Sarda | Pulkit Chatwal | Sonal Dabral
Understanding why financial outcomes occur is as important as knowing what they are. Annual reports and regulatory filings are rich with causal reasoning, yet extracting that reasoning automatically remains a difficult problem — one that sits at the intersection of domain expertise, linguistic nuance, and machine comprehension. In this paper, we describe our participation in the English subtask of the Financial Document Causality Detection shared task, FinCausal 2026, where systems are asked to identify verbatim causal spans from financial paragraphs in response to abstractive causal questions. Our approach is grounded in the intuition that a small, well-adapted model with the right inductive biases can outperform a larger but unfocused one. We fine-tune Qwen2.5-4B-Instruct on 2,000 domain-annotated instances using QLoRA, a parameter-efficient technique that enables meaningful adaptation under modest computational resources. Before training, we reformat all instances into the Qwen ChatML instruction template to align the model’s generation behaviour with the verbatim extraction requirement of the task. At inference time, we further guide the model by retrieving the most causally relevant sentence from the context using TF-IDF cosine similarity, providing an explicit local signal before generation. Outputs are produced via greedy decoding to ensure deterministic, source-grounded predictions. Under the official LLM-as-a-judge evaluation framework — which scores responses on a 1–5 adequacy scale based on semantic correctness rather than lexical overlap — our system achieves a score of 4.76 out of 5, placing 4th out of nine teams on the English leaderboard. Our results suggest that combining instruction-tuned fine-tuning with lightweight retrieval is a practical and effective strategy for causal reasoning in specialised financial text.
up
Proceedings fo the Second International Workshop on Eye-Tracking Resources and Evaluation for Human-Aligned NLP
Proceedings fo the Second International Workshop on Eye-Tracking Resources and Evaluation for Human-Aligned NLP
Cengiz Acartürk | Burcu Can | Jamal Nasir | Çağrı Çöltekin
Cengiz Acartürk | Burcu Can | Jamal Nasir | Çağrı Çöltekin
Eye tracking offers unique insights into cognitive processes, making it a promising tool for evaluating machine translation (MT). This study explores the feasibility of using an iPhone 12 camera-based eye tracker with a 14-inch laptop display for conducting translation evaluation in personal workspaces, offering a more accessible and cost-effective alternative to traditional setups. Participants evaluated source sentences, selected translations, and identified problematic words while their gaze metrics were recorded and analyzed. Our findings reveal statistically significant correlations between gaze patterns and preferred translations, as well as increased visual attention to problematic words. These results demonstrate that home-based eye tracking systems are technically sufficient for capturing gaze behavior accurately enough for MT evaluation purposes. A potential practical application is to speed up translation proof-reading using eye tracking technique to automatically mark portions of text that should be attended to and improved based on the gaze pattern during a quick reading.
Cross-Linguistic Analysis of Eye Movement Patterns: Insights from the First Arabic Eye-Tracking Corpus for NLP
Ibtehal Baazeem | Hend Al-Khalifa | Abdulmalik AlSalman
Ibtehal Baazeem | Hend Al-Khalifa | Abdulmalik AlSalman
Eye-tracking corpora have become valuable resources for understanding human reading behavior and developing cognitively-informed NLP models. However, existing resources predominantly focus on left-to-right Latin script languages, leaving a significant gap for morphologically rich, right-to-left languages like Arabic. This paper presents a cross-linguistic analysis of eye movement patterns using the AraEyebility corpus, the first Arabic eye-tracking corpus comprising 57,617 words read by 15 native speakers. We systematically compare gaze metrics across Arabic and established English corpora. Our analysis reveals distinct patterns in fixation duration, saccade length, and regression frequency that reflect Arabic’s unique orthographic properties: cursive script, diacritization, bidirectional reading (text right-to-left, numbers left-to-right), and morphological complexity. The findings demonstrate that Arabic readers exhibit longer mean fixation durations and more frequent regressions compared to English readers, suggesting higher cognitive processing demands. We discuss implications for developing cognitively-aligned NLP models and provide recommendations for future multilingual eye-tracking research. The AraEyebility corpus is publicly available to support Arabic NLP research.
Exploring Cognitively Informed Sentence Simplification with Gaze-Guided Text Generation
Andreas Säuberli | Diego Frassinelli | Barbara Plank
Andreas Säuberli | Diego Frassinelli | Barbara Plank
Automatic text simplification has mostly relied on human judgments when it comes to what is considered easy or difficult to read. Eye movements while reading can offer a more direct and objective signal of processing effort and reading ease. In this paper, we explore gaze-guided text generation (GGTG), an approach to control reading ease in generated texts, and assess its use for sentence simplification. GGTG employs a gaze model that is trained to predict eye-tracking measures such as reading times or regression rates, which are then used to rerank next-token probabilities generated by a language model. We evaluated the approach on an English sentence simplification benchmark and found gains in automatic evaluation metrics, although the simplification operations are mostly limited to the lexical level. Its modular nature also allows GGTG to be combined with other simplification techniques such as prompting or fine-tuning.
Impact of Text Simplification on Eye-Tracking-Based Reading Profiles Across Domains
Oksana Ivchenko | Natalia Grabar
Oksana Ivchenko | Natalia Grabar
Understanding how text readability affects reading behaviour is crucial for improving accessibility and health communication. We analyse sentence-level eye-tracking data from the French Eye-TrAcking (FETA) corpus, which includes original and manually simplified texts from three domains: general, medical, and clinical. Using clustering of fixation-based features, we identify recurrent processing patterns and examine how these patterns change under text simplification. Cluster quality is evaluated using silhouette scores and participant-level bootstrap stability. Simplification does not uniformly reduce reading effort but reorganises processing in domain-dependent ways. Medical texts show strong diversification, general texts moderate diversification, and clinical texts show a reduction in the number of distinct reading profiles. Hence, rather than uniformly facilitating reading, simplification redistributes effort across sentences, underscoring the need for domain-sensitive readability approaches.
This study uses regression analysis of Brazilian Portuguese eye-tracking data to examine variability in reading times across grammatical categories. Mixed-effects models reveal distinct patterns: numerals elicit high individual variability in early-stage reading, while function words (e.g., adpositions, determiners) drive differences in late-stage integration. In contrast, nouns show stable effects. These findings demonstrate that individual differences in reading are systematically linked to specific parts of speech, with numerals and function words as key loci of variability.
CoordiMap: Conceptual Proposition of a new Framework for the Annotation of Verbal Elicitation Paths on Visual Experiment Stimuli and Introduction of the Associated Annotation Tool
Carmen Schacht
Carmen Schacht
Consistent alignment of multi-modal experimental data—such as verbal utterances in elicitation tasks, (static) visual stimuli, and gaze data—presents a challenge in linguistic research. These elicitations often encode information about the visual perception strategies or cognitive processing of the scene. Thus, it is helpful to transform them into a structured, visually grounded format which captures the visual nature of the data, ideally able to be aligned with the corresponding gaze data. To achieve this, the present paper conceptually proposes the annotation framework for verbal elicitation paths as a data type and presents the first release of the associated newly developed CoordiMap annotation tool. The tool enables structured mapping of verbal elicitation data from experimental studies onto the corresponding visual stimuli. Independent of specific paradigms, the tool supports the annotation of verbal utterances in a linearized form based on coordinates directly marked on the image of the stimulus. The format is conceptually inspired by eye-tracking data formats, in which gaze behavior is represented as temporally linearized paths overlaid on the stimulus. The paper motivates the development of the tool and its annotation methodology by theoretical and experimental considerations regarding the relationship between visual perception and language production. As this a work in progress, the functionality of the annotation tool is demonstrated through an exemplary use case.
A Comparative Study Between Mouse and Eye Tracking Signals for Long Romanian Texts
Bogdan Alexandru Gheorghe | Sergiu Nisioi
Bogdan Alexandru Gheorghe | Sergiu Nisioi
Understanding human language processing via eye-tracking (ET) is precise but limited by scalability. Mouse-Tracking (MoTR) offers a cost-effective alternative, yet its viability for long-form reading in languages like Romanian remains underexplored. The primary challenge lies in the motor-induced noise and biomechanical discrepancies between hand and eye movements. Here we show that combining targeted technical enhancements with a Hertz-based velocity transformation allows MoTR to serve as a robust proxy for ET. We evaluate this by training a BERT-enhanced Fusion Model that integrates semantic context to bridge the mechanical gap, achieving an internal consistency of ρ ≈ 0.58 and a cross-modal correlation of ρ ≈ 0.22 in the velocity domain. These results indicate that when properly normalized, manual tracking captures similar cognitive constraints as gaze, with predictive accuracy approaching the empirical bounds of human behavioral variance.
Eye-Contact and Facial Expression Tracking for Assertiveness Training in VR-Based Anti-Bullying Education
Lubomir Ivanov | Anabel Nolasco | Mary Vrahimis
Lubomir Ivanov | Anabel Nolasco | Mary Vrahimis
This paper described the use of eye-contact and facial expression tracking as part of a comprehensive approach to assertiveness training in a VR-based anti-bullying simulation environment. We briefly discuss the psychological foundations of assertiveness and then focus on our approach to tracking the facial expressions and eye-contact that a user maintains while communicating with the virtual bully in the simulation. We also outline additional non-verbal indicators tracked by the software and discuss the dialog system, which drives the simulation. Finally, we outline some ethical considerations, discuss the limitations of our current software prototype, and list future directions for enhancing assertiveness training in anti bullying education.
Predicting Gaze Location without Camera or Eye-Tracker
Saman Rezapoor | Sajad Shirali-Shahreza | Gerald Penn
Saman Rezapoor | Sajad Shirali-Shahreza | Gerald Penn
The task of identifying the location that a user looks at, commonly known as gaze estimation, has various HCI and NLP applications. Traditional gaze estimation methods use special hardware such as eye-trackers or ordinary cameras such as webcams to perform this. However, they are not applicable to the majority of web users either because the user does not have them or does not want to use them due to privacy reasons. In this paper, we propose the idea of using multimodal LLMs to analyze the content of the user’s screen along with mouse location to estimate the gaze location. It primarily uses the results of studies that extract common reading patterns such as the F-pattern and Z-pattern. Our experimental results on The Eye Of The Typer (EOTT) dataset provide promising results for estimating gaze location.
A Survey of Incorporating Gaze Data into Natural Language Processing Models and Applications
Cengiz Acarturk | Burcu Can | Melike Caglayan | Jamal Abdul Nasir | Cagri Coltekin
Cengiz Acarturk | Burcu Can | Melike Caglayan | Jamal Abdul Nasir | Cagri Coltekin
This study presents a survey of research integrating eye-tracking (gaze) data into Language Models (LMs) as a means of cognitively grounding NLP models and applications in human reading behavior. Although contemporary LMs excel at learning statistical patterns from text, they fundamentally lack human-like reading and comprehension capabilities. Incorporating gaze data may offer a window into cognitive processing, yet its impact on LMs remains underexplored. Addressing a persistent bottleneck, namely, the high cost and limited scale of laboratory eye-tracking, we propose a roadmap consisting of three streams of research for advancing this novel research domain: (1) developing cognitive multimodal corpora, (2) leveraging generative models for gaze synthesis to overcome the data bottleneck caused by the high costs of human eye-tracking, and (3) training LMs with gaze-guided attention mechanisms and input augmentation. Furthermore, we illustrate practical applications in readability assessment, educational analytics, and assistive communication, demonstrating how gaze-informed models can enable adaptive technologies. Finally, we critically examine ongoing challenges, including the lack of data standardization, the misalignment between human and machine language processing, and the urgent ethical imperative for privacy-preserving architectures to protect sensitive biometric gaze data, motivating privacy-aware data practices and model designs for scalable deployment.
up
Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes)
Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes)
Isuri Anuradha | Martin Wynne
Isuri Anuradha | Martin Wynne
Integrating TEI Publication, Guided Exploration, and Vector Databases for Semantic Search in the Voci dall’Inferno Project
Angelo Mario Del Grosso | Elvira Mercatanti | Carla Congiu | Marina Riccucci
Angelo Mario Del Grosso | Elvira Mercatanti | Carla Congiu | Marina Riccucci
This paper presents recent advances toward an integrated framework that combines TEI-based digital publishing with embedding-based semantic search to support the preservation, exploration and analysis of Holocaust survivor testimonies. The corpus includes written and oral sources and preserves them within a XML-TEI model supported by an ODD customization that preserves provenance, structure and interpretability. A dedicated web application developed within the eXistdb platform provides guided access to the digital corpus and supports the management, visualization, and exploration of the encoded data. The project aims to investigate a specific research goal: to verify the presence of references to the Divine Comedy by Dante within Holocaust testimonies. To this end, we implement a semantic retrieval component based on SentenceTransformers’ embeddings and a vector database, enabling the discovery of both literal and non-literal Dantean passages within the testimonies. The paper presents the advances achieved toward this objective and the ethical constraints shaping access policies, resulting in a sustainable archive and a reproducible methodology for intertextual research in sensitive historical collections.
Towards Semantic Searching in Diverse Multimodal Collections
Václav Kučera | Martin Bulín | Jan Švec | Pavel Ircing
Václav Kučera | Martin Bulín | Jan Švec | Pavel Ircing
Digital humanities projects increasingly rely on heterogeneous collections of multimodal data, including video testimonies, scanned documents, and photographs. Despite the growing availability of such archives, researchers face challenges in efficiently locating relevant content due to the diversity of formats and the lack of unified retrieval methods. In this work, we present a general framework for semantic search over collections of multiple modalities. The framework integrates specific parsers and transforms all inputs into textual representations leveraging services like automatic speech recognition (ASR), optical character recognition (OCR), and generative-AI-based image captioning. Text is subsequently segmented into overlapping chunks, indexed in a vector database, and enriched through an automatic question generation (AQ) pipeline to create ground-truth queries for evaluation. We evaluate the framework on a constructed dataset derived from Holocaust-related archives, comparing two retrieval strategies (pure vector search vs. hybrid semantic-lexical search) under two chunking scenarios. Results demonstrate that hybrid search consistently outperforms vector-only retrieval, achieving high recall across modalities, and that semantic search is feasible even with diverse and noisy input sources. This framework provides a robust foundation for exploring complex multimodal archives, facilitating access to content that would otherwise remain difficult to discover.
Automatic Transcription of Holocaust Testimonies in Yiddish: Orthographic Comparison and Cross-Domain Validation
Isaac L. Bleaman
Isaac L. Bleaman
The digitization and computational processing of Holocaust testimony interviews are essential for the long-term preservation and accessibility of survivors’ narratives. However, automatic speech recognition (ASR) for Yiddish—the primary language of most Holocaust victims and survivors—remains underdeveloped. This paper introduces the first ASR system for European Yiddish, focused on the Northeastern (“Lithuanian”) dialect and trained on Holocaust survivor testimonies from the Corpus of Spoken Yiddish in Europe (42 hours of speech segments from 60 survivors). A systematic comparison of CTC-based ASR models using transcripts with different orthographic representations reveals that a Hebrew-based phonemic system with precomposed Unicode is optimal, achieving a mean WER of 37.96% compared to 59.40% WER for romanized Yiddish and 99.67% WER (catastrophic failure) for standard Yiddish spelled with decomposed Unicode. Cross-domain testing on Yiddish audiobooks provides additional support for a phonemic representation (27.07% WER, 6.56% CER). Together, the results suggest that automatic transcription developed from oral Holocaust testimonies can support further technological innovation in service of Yiddish-speaking communities.
From Consensus to Split Decisions: ABC-Stratified Sentiment in Holocaust Oral Histories
Daban Q. Jaff
Daban Q. Jaff
Polarity detection becomes substantially more challenging under domain shift, particularly in heterogeneous long-form narratives with complex discourse structure, such as Holocaust oral histories. This paper presents a corpus-scale diagnostic study of off-the-shelf sentiment classifiers on Holocaust oral histories, using three pretrained transformer-based polarity classifiers over a corpus comprising 107,304 utterances and 579,013 sentences. After assembling model outputs, we introduce an agreement-based stability taxonomy (ABC) to stratify inter-model output stability. We report pairwise percent agreement, Cohen’s κ, Fleiss’ κ, and row-normalized confusion matrices to localize systematic disagreement. As an external convergent descriptive signal, we apply a T5-based emotion classifier to stratified samples from each agreement stratum to compare emotion distributions across strata. The combination of multi-model label triangulation and the ABC taxonomy provides a cautious, interpretable framework for characterizing where and how sentiment models diverge in sensitive historical narratives. Inter-model agreement is low to moderate overall and is driven primarily by boundary decisions around neutrality.
EHRI Annotator: A Web-Based Tool for Named Entity Recognition and Linking in Holocaust-Related Texts
Maria Dermentzi
Maria Dermentzi
This paper presents the EHRI Annotator, a web-based tool for multilingual named entity recognition (NER) and entity linking (EL) in Holocaust-related texts. The tool was developed to support services provided by the European Holocaust Research Infrastructure (EHRI), primarily the digital scholarly editions published by EHRI (EHRI Online Editions) by streamlining the process of detecting named entities in documents and linking them to their unique identifiers in EHRI and third-party controlled vocabularies and gazetteers. The EHRI Annotator builds upon previous work on domain-specific NER, taking it a step further to support multilingual EL. The tool adopts a dual entity linking architecture that uses a different matching approach depending on the type of the named entity. It performs semantic matching for entities to be linked to EHRI vocabularies and authority sets which are modestly sized, and string-matching-based retrieval for locations to be linked to the extensive GeoNames gazetteer using a domain-specific relevance weighting. A preliminary evaluation on 264 entities from a manually annotated dataset of Holocaust testimonies yields an Accuracy@5 of 77.7% when it comes to the linking component of the tool. User testing confirms the tool’s usability but also highlights areas for improvement.
The Shape of Testimony: A Scalable Framework for Oral History Archive Comparison
Renana Keydar | Amit Pinchevski | Itamar Trainin
Renana Keydar | Amit Pinchevski | Itamar Trainin
Researchers in Holocaust studies have often distinguished between two styles of oral survivor testimony: the USC Shoah Foundation’s interviews tend to follow a structured, interviewer-guided format, whereas the Yale Fortunoff Video Archive generally favors a more free-form, open-ended style. This distinction has influenced both scholarly research and the development of later archives. In this study, we critically examine that claim by conducting a large-scale computational analysis of more than 1,600 testimonies from both collections. Leveraging discourse segmentation, topic modeling, and large language model (LLM) based analysis, we quantify the “structuredness” level of testimonies through topic coherence, interviewer–survivor dynamics, and the distribution of question types. Our results generally corroborate the structural differences identified in earlier research, while also revealing significant overlaps between the collections, both within individual interviews and across common narrative patterns. This complicates the simple “structured vs. free-form” dichotomy often applied to these oral histories. Beyond revisiting a foundational claim in Holocaust studies, our work provides a scalable, replicable framework for comparative corpus analysis. As a proof of concept, it suggests broader applications for digital oral history, narrative analysis, and the design of citizen-science annotation platforms.
From Oral History to Structured Data: The MalachNER Dataset
Christopher Brückner | Karin Roginer Hofmeister | Jiří Kocián | Pavel Pecina
Christopher Brückner | Karin Roginer Hofmeister | Jiří Kocián | Pavel Pecina
We present MalachNER, a new multilingual dataset for Named Entity Recognition (NER) in testimonies of Holocaust survivors. MalachNER has been sourced from different archives and annotated based on comprehensive domain-specific guidelines refined by a collaboration of international experts. Covering 10 European languages, differs significantly from previously released datasets: It is primarily based on noisy, verbatim transcribed speech, rather than on digitized written documents. These transcripts are characterized, among other challenges, by fillers, dialectal speech, and in-line annotations indicating incomprehensible words, which are not commonly encountered in other datasets. However, large volumes of yet unprocessed oral history make such a dataset a necessity. In addition to the description of the dataset and its annotation guidelines, we show with baseline experiments that MalachNER is complementary with previously released data, and the key to training domain-specific language models that generalize well to written and oral testimony alike, achieving state-of-the-art performance on both types of documents.
Emotions In Oral History Interviews: A Multimodal Approach to Holocaust Testimonies
Nele Mantaj | Vaibhav Agarwal | Ines Matres
Nele Mantaj | Vaibhav Agarwal | Ines Matres
Video interviews with Holocaust survivors and witnesses comprise, to date, the most globally distributed and comprehensive oral history documentation. As survivors among us disappear, these sources are increasingly important to understand the impact of the Holocaust and mechanisms to overcome the trauma experienced. While historians often rely on written transcripts, these omit emotional nuances conveyed through audiovisual cues such as facial expressions, pauses, and eye movements. This article outlines the resources, data-preparation steps, and analytical methods used during a 10-day Digital Humanities Hackathon project to examine emotions in Holocaust testimonies, incorporating video, audio, and text. The group aimed to determine whether audiovisual signals offer meaningful emotional or sentimental information beyond transcripts. To achieve this, the group worked with a sample of 10 interviews facilitated by the US Holocaust Memorial Museum (USHMM); which were separated into video, audio, and textual components for machine processing and realigned side-by-side for analysis. This resulting “cookbook” lays out a workflow, resources, and practical entry points for preparing oral history interviews for multimodal emotion and sentiment annotation, or to aid the detection of emotionally significant moments for deeper examination.
Modeling the Language of Holocaust Survivors’ Testimony with Domain-Adapted Transformers
Christopher Brückner | Jan Lehečka | Jan Švec | Pavel Pecina
Christopher Brückner | Jan Lehečka | Jan Švec | Pavel Pecina
Documents related to the Holocaust increasingly move into the focus of Natural Language Processing research, including the digitization of written text, the automatic transcription of oral archives, and interpretive downstream tasks such as Named Entity Recognition. However, most modern language models are trained primarily on modern text, and thus struggle with historical language, historical entities, and domain-specific terminology. Furthermore, transcribed speech introduces challenges such as transcription errors, noise, filler words, and dialectal speech not often contained in textual datasets. We present XLM-RoBERTa-malach, a text encoder domain-adapted to oral testimonies of Holocaust survivors in seven languages. In addition to descriptions of the data acquisition via Automatic Speech Recognition, data augmentation via Machine Translation, and the continued pretraining of a state-of-the-art multilingual transformer, we evaluate the domain-adapted model on the Named Entity Recognition task. Experiments on this task show superior performance over the general-domain transformer in a multilingual domain-specific setting, including languages not seen during the domain adaptation.
Evaluating Automatic Speech Recognition for Holocaust Testimonies: A Large-Scale Analysis of Whisper Performance on the Fortunoff Video Archive
William J.B. Mattingly | Christy Bailey-Tomecek
William J.B. Mattingly | Christy Bailey-Tomecek
Holocaust testimonies are key primary sources documenting survivors’ experiences, yet many remain inaccessible due to the labor-intensive nature of manual transcription. This paper presents a comprehensive evaluation of OpenAI’s Whisper automatic speech recognition (ASR) system on 1,847 testimonies from the Fortunoff Video Archive for Holocaust Testimonies at Yale University. We assess transcription quality across multiple languages including English, French, German, Hebrew, Yiddish, Ladino, Slovak, and American Sign Language (with English voice-over), using human-reviewed captions as ground truth. Our analysis reveals a mean Word Error Rate (WER) of 15.28%, with 90.9% of testimonies achieving “Fair” or better quality (WER ≤25%). We identify systematic error patterns including challenges with disfluencies, interrupted speech, and language-specific orthographic conventions, particularly in Ladino, where Whisper’s normalization to modern Spanish orthography creates systematic divergences from traditional Judeo-Spanish spelling. For Hebrew and Yiddish, we evaluate specialized models from ivrit-ai and find promising results for heritage language preservation. Our findings demonstrate that current ASR technology can substantially accelerate Holocaust testimony transcription while highlighting the need for domain-specific fine-tuning and post-processing for optimal results.
Cross-Modal Modeling of Emotional and Thematic Trajectories in Holocaust Survivor Oral Histories
Henry Gagnier
Henry Gagnier
Large-scale corpora of Holocaust testimonies preserve vast amounts of historical, emotional, and narrative information, but their size and complexity can make accurate, systematic analysis challenging. This paper presents a cross-modal computational analysis of emotional and thematic trajectories in the CORHOH corpus, containing 500 Holocaust survivor testimonies as a language resource for computational analysis. We segment each testimony into ten segments and apply sentiment analysis, emotion recognition, and topic modeling to each of these segments to reveal how theme and emotion evolve over time in Holocaust testimonies. Results reveal a sharp decline from pre-war life in wartime and camp experiences, with sentiment and emotion remaining negative in post-war segments. Emotion analysis reveals decreasing joy and increasing sadness and fear during segments related to deportation and concentration camps, with limited emotional recovery. Topic modeling identifies coherent themes that align closely with sentiment and emotional patterns. We systematically examine correlations between sentiment, emotion, and topic trajectories, which demonstrate many strong associations between topic and emotion. This work demonstrates that combining sentiment analysis, emotion recognition, and topic modeling can reveal systematic patterns in large oral history corpora, and shows the value of computational approaches for studying historical narratives like the Holocaust.
up
Proceedings of the Second Workshop of Identity Aware AI
Proceedings of the Second Workshop of Identity Aware AI
A Pranav | Valerio Basile | Neele Falk | David Jurgens | Gabriella Lapesa | Anne Lauscher | Soda Marem Lo
A Pranav | Valerio Basile | Neele Falk | David Jurgens | Gabriella Lapesa | Anne Lauscher | Soda Marem Lo
The Point of View of a Sentiment: Towards Clinician Bias Detection in Psychiatric Notes
Alissa A. Valentine | Lauren Lepow | Lili Chan | Alexander Charney | Isotta Landi
Alissa A. Valentine | Lauren Lepow | Lili Chan | Alexander Charney | Isotta Landi
Negative patient descriptions and stigmatizing language can contribute to generating healthcare disparities in two ways: (1) read by patients, they can harm their trust and engagement with the medical center; (2) read by physicians, they may negatively influence their perspective of a future patient. In psychiatry, the patient-clinician therapeutic alliance is a major determinant of clinical outcomes. Therefore, language usage in psychiatric clinical notes may not only create healthcare disparities, but also perpetuate them. Recent advances in natural language processing systems have facilitated the efforts to detect discriminatory language in healthcare. However, such attempts have only focused on the perspectives of the medical center and its physicians. Considering both physicians’ and non-physicians’ subjective points of view is a more equitable approach to identifying harmful language in clinical notes. By leveraging large language models (LLMs), this work aims to characterize potentially harmful language usage in psychiatric notes by identifying the sentiment expressed in sentences describing patients based on the reader’s point of view. First, we curated a psychiatric lexicon containing words commonly used to describe patients in psychiatry. Sentences (N=39) were extracted from clinical text containing psychiatric lexicon at a medical center, with which a set of physicians (N=10) and non-physicians (N=10) annotated them as negative, neutral, or positive. Three LLMs (GPT-3.5, Llama-3.1, and Mistral) used zero-shot/few-shot in-context learning (ICL) approaches to classify the sentiment of the sentences according to the physician or non-physician point of view. Results showed that GPT-3.5 aligned best to physician point of view and Mistral aligned best to non-physician point of view, both with an ICL approach. These results underline the importance of recognizing subjectivity in clinical annotation tasks, not only for improving the note writing process, but also for the quantification, identification, and reduction of bias in computational systems for downstream analyses.
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
Shree Harsha Bokkahalli Satish | Harm Lameris | Olivier Perrotin | Gustav Eje Henter | Eva Szekely
Shree Harsha Bokkahalli Satish | Harm Lameris | Olivier Perrotin | Gustav Eje Henter | Eva Szekely
Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker identity. Because SC is constrained to a single audio stream, it offers a more direct setting for probing biases in speech foundation models than dialogue does. In this work we present the first systematic evaluation of bias in SC, investigating how gender and phonation type (breathy, creaky, end-creak) affect continuation behaviour. We evaluate three recent models: SpiritLM (base and expressive), VAE-GSLM, and SpeechGPT across speaker similarity, voice quality preservation, and text-based bias metrics. Results show that while both speaker similarity and coherence remain a challenge, textual evaluations reveal significant model and gender interactions: once coherence is sufficiently high (for VAE-GSLM), gender effects emerge on text-metrics such as agency and sentence polarity. In addition, continuations revert toward modal phonation more strongly for female prompts than for male ones, revealing a systematic voice-quality bias. These findings highlight SC as a controlled probe of socially relevant representational biases in speech foundation models, and suggest that it will become an increasingly informative diagnostic as continuation quality improves.
Investigating the Automatic Translation of Korean Honorifics
Luis Cihlar | Minh Duc Bui | Kyung eun Park | Manuel Mager | Walter Bisang | Katharina von der Wense
Luis Cihlar | Minh Duc Bui | Kyung eun Park | Manuel Mager | Walter Bisang | Katharina von der Wense
Honorifics encode social hierarchies and relational nuances, making their correct use a culturally sensitive yet challenging aspect of translation. In doing so, they reflect and shape how individuals position themselves and others within a social world. In this work, we investigate how different translation models handle Korean honorifics, both in implicit scenarios, where only the sentence is given, and explicit scenarios. Our findings are as follows: (i) large language models (LLMs) fine-tuned for translation (MTLMs) consistently prefer polite forms more than their instruction-tuned counterparts in both scenarios; (ii) sequence-to-sequence models produce less polite outputs in implicit contexts but shift toward more polite forms when the addressee is explicitly provided; and (iii) both types of LM-based models tend to become more casual when the addressee is known. When compared with human preferences, MTLMs diverge more strongly, exhibiting a systematic overuse of polite forms relative to human judgments.
Balancing the Scales: Reinforcement Learning for Fair Classification
Leon Eshuijs | Shihan Wang | Antske Fokkens
Leon Eshuijs | Shihan Wang | Antske Fokkens
Fairness in classification tasks has traditionally focused on bias removal from neural representations, but recent approaches have shifted towards algorithmic methods that embed fairness into the training process. These methods steer models towards fair performance, preventing potential elimination of valuable information that arises from representation manipulation. Reinforcement Learning (RL), with its ability to learn through interaction and adjust reward functions to encourage desired behaviors, presents a promising approach in this domain. In this paper, we conduct an exploratory evaluation of RL for addressing bias in imbalanced classification by scaling the reward function. We employ the contextual multi-armed bandit framework, adapt three popular RL algorithms, and conduct an extensive empirical evaluation of their relative strengths and limitations. Through this analysis, we contribute meaningful evidence to the ongoing debate between algorithmic and representational fairness approaches.
Evaluating LLMs for Detecting Demographic-Targeted Social Bias: A Comprehensive Benchmark Study
Ayan Majumdar | Feihao Chen | Jinghui Li | Xiaozhen Wang
Ayan Majumdar | Feihao Chen | Jinghui Li | Xiaozhen Wang
Large-scale web-scraped text corpora used to train general-purpose AI models often contain harmful demographic-targeted social biases, creating a regulatory need for data auditing and developing scalable bias-detection methods. Although prior work has investigated biases in text datasets and related detection methods, these studies remain narrow in scope. They typically focus on a single content type (e.g., hate speech), cover limited demographic axes, overlook biases affecting multiple demographics simultaneously, and analyze limited techniques. Consequently, practitioners lack a holistic understanding of the strengths and limitations of recent large language models (LLMs) for automated bias detection. In this study, we conduct a comprehensive benchmark study on English texts to assess the ability of LLMs in detecting demographic-targeted social biases. To align with regulatory requirements, we frame bias detection as a multi-label task of detecting targeted identities using a demographic-focused taxonomy. We then systematically evaluate models across scales and techniques, including prompting, in-context learning, and fine-tuning. Using twelve datasets spanning diverse content types and demographics, our study demonstrates the promise of fine-tuned smaller models for scalable detection. However, our analyses also expose persistent gaps across identity axes and multi-demographic targeted biases, underscoring the need for more effective and scalable detection frameworks.
Queering the Audits: Community-Based Auditing of AI Harms to Queer Communities
Organizers of Queer In AI | A Pranav | Alissa A. Valentine | Alex Markham | Beckett LeClair | Tereza Blazkova | Ekaterina Kornilitsina | Sofie H. Bruun | Gerasimos Spanakis | Anne Lauscher
Organizers of Queer In AI | A Pranav | Alissa A. Valentine | Alex Markham | Beckett LeClair | Tereza Blazkova | Ekaterina Kornilitsina | Sofie H. Bruun | Gerasimos Spanakis | Anne Lauscher
AI systems embed majority-group defaults into training data, evaluation metrics, and category definitions, producing documented harms for queer communities including erasure, misclassification, and discrimination. Standard technical audits often rely on aggregate measures and cannot detect harms that be come visible only through the lived experience of affected communities. We conducted a participatory auditing workshop at EurIPS 2025 where 16 queer community members audited four case studies using the 4Cs harm taxonomy (Content, Conduct, Contact, Contract) applied across the AI lifecycle. Participants used structured worksheets and plenary synthesis to classify harms and trace them to their origins in the development pipeline. Across all four cases, participants traced harms to problem definition and data collection, and they identified contractual structures that extract value from vulnerable populations while providing minimal recourse. These findings illustrate that community-informed auditing surfaces concrete, identity-specific harms that aggregate evaluation methods risk overlooking.
up
Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026
Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026
Simona Frenda | Marco Antonio Stranisci | Shaina Ashraf | Ada Ren | Ioannis Konstas | Usman Naseem
Simona Frenda | Marco Antonio Stranisci | Shaina Ashraf | Ada Ren | Ioannis Konstas | Usman Naseem
Combating Disinformation: Is There No Alternative?
Davide Bassi | Søren Kirkegaard Fomsgaard | Erik Bran Marino | Katarina Laken
Davide Bassi | Søren Kirkegaard Fomsgaard | Erik Bran Marino | Katarina Laken
This position paper critiques the dominance of detection-centered approaches in misinformation research. We argue that the prevailing paradigm treats information disorders as a content-level anomaly to be identified and suppressed, thereby obscuring the structural conditions under which different forms of information disorders emerge and resonate. Drawing on critical anthropology, we propose an alternative “clinical” model: information disorders should be understood not only as informational distortion, but as a syndrome with complex causes embedded in contexts of economic precarity, institutional distrust, and informational inequality. Treating detection as the ends rather than the means of intervention risks misguiding our efforts. Rather than positioning NLP primarily as a tool for boundary enforcement, we outline a reorientation toward structural diagnosis: diversifying data beyond WEIRD contexts, extracting socioeconomic and trust-related signals from discourse, and integrating computational outputs within interdisciplinary causal frameworks. Under this model, detection becomes a means for an epidemiology of discourse, subordinated to the broader objective of cultivating long-term epistemic resilience in our online environments.
In the era of digital communication, the rapid spread of information presents significant challenges to society. This paper provides an in-depth examination of the existing frameworks developed to understand and address these phenomena. More precisely, this paper categorizes and compares various frameworks, including typology-based, process-oriented, impact-oriented, and actor-centric approaches. It highlights the strengths and limitations of each framework type, with a particular focus on their applicability to combat false information in diverse contexts. The paper underscores the importance of adopting a holistic and flexible approach that integrates multiple frameworks and adapts to the evolving nature of technology, particularly AI-driven false and misleading content.
High Accuracy, Low Generalization: Structural Homogeneity and Cross-Dataset Evaluation in Fake-News Benchmarks
Hiram Calvo | Mayte H. Laureano
Hiram Calvo | Mayte H. Laureano
State-of-the-art fake-news classifiers frequently report near-ceiling accuracy on widely used benchmarks such as ISOT, Misinfo, and WELFake. We argue that such results often reflect structural homogeneity and provenance-based separability rather than robust claim-level veracity inference. Anchored in the Information Disorder framework, we analyze how dataset construction operationalizes the notion of “fake” and how this shapes model behavior. We conduct systematic bidirectional cross-dataset experiments across six transfer directions and evaluate performance not only by mean accuracy, but also by variance and directional asymmetry. Results reveal substantial degradation under distribution shift and pronounced transfer asymmetries between dataset pairs. Although not always achieving the highest mean accuracy, affective augmentation combining dimensional (VAD) and categorical (Ekman) representations yields the lowest variance and smallest directional gap, indicating superior cross-domain stability. Our findings expose the disconnect between accuracy-driven benchmarking and construct-valid evaluation. We argue that progress in fake-news detection requires shifting from isolated in-domain optimization toward robustness-oriented, bidirectional, and distribution-aware assessment practices.
Benchmarking Check-Worthiness Models on LLM Generated Claims
Charlie George Roadhouse | Matthew Shardlow | Ashley Williams
Charlie George Roadhouse | Matthew Shardlow | Ashley Williams
The proliferation of large language models (LLMs) has significantly increased the potential for automated dissemination of disinformation, necessitating robust systems for check-worthiness detection. However, existing models are primarily trained on human claims, leaving their performance on machine-generated text largely unexplored. In this paper, we benchmark encoder models (BERT and RoBERTa) and industry accessible tools (ClaimBuster) against LLM-paraphrased claims across three stylistic categories: syntactic restructuring, syntactic complexity and lexical informality. Our results indicate a consistent performance degradation on synthetic claims, particularly on complex and informal claims. We demonstrate that adversarial training significantly improves model resilience, with RoBERTa achieving F1-score gains up to +5.22 on the CheckIt dataset. Finally, SHAP analysis reveals that while base models rely on narrow syntactic heuristics such as active voice, robust models learn to anchor their prediction on core factual entities. These findings highlight the necessity of stylistic-aware training to maintain fact-checking efficacy in an increasingly LLM-populated information landscape.
Media Bias within Information Disorder: Bridging Two Research Communities through a Systematic Review
Francisco-Javier Rodrigo-Ginés | Jorge Chamorro-Padial
Francisco-Javier Rodrigo-Ginés | Jorge Chamorro-Padial
Information disorder research overwhelmingly focuses on fabricated or manipulated content (fake news, deepfakes, propaganda) while comparatively neglecting the most pervasive form of distorted information: media bias. Unlike outright falsehoods, media bias operates within the boundaries of factual reporting, distorting public understanding through framing, omission, and word choice rather than fabrication. This makes it harder to detect, harder to regulate, and paradoxically more influential, since it originates from trusted mainstream sources rather than marginal actors. In this position paper, we argue that media bias should be recognized as a first-class category within information disorder frameworks. Drawing on the Wardle and Derakhshan (2017) taxonomy, communication theory, and a systematic review of over 100 studies on automated media bias detection, we demonstrate that current frameworks inadequately account for the systematic distortion of true content. We present a consolidated taxonomy of media bias types organized by linguistic level, compare detection paradigms across the information disorder and media bias communities, and identify four properties that make media bias uniquely dangerous: its scale, its source credibility, the invisibility of omission, and its cumulative normative effect. We conclude with an integrated research agenda grounded in specific gaps identified through the review.
Reliable News or Propagandist News? A Neurosymbolic Model Using Genre, Topic, and Persuasion Techniques to Improve Robustness in Classification
Géraud Faye | Benjamin Icard | Morgane Casanova | Guillaume Gadek | Guillaume Gravier | Wassila Ouerdane | Celine Hudelot | Sylvain Gatepaille | Paul Égré
Géraud Faye | Benjamin Icard | Morgane Casanova | Guillaume Gadek | Guillaume Gravier | Wassila Ouerdane | Celine Hudelot | Sylvain Gatepaille | Paul Égré
Among news disorders, propagandist news are particularly insidious, because they tend to mix oriented messages with factual reports intended to look like reliable news. To detect propaganda, extant approaches based on Language Models such as BERT are promising but often overfit their training datasets, due to biases in data collection. To enhance classification robustness and improve generalization to new sources, we propose a neurosymbolic approach combining non-contextual text embeddings (fastText) with symbolic conceptual features such as genre, topic, and persuasion techniques. Results show improvements over equivalent text-only methods, and ablation studies as well as explainability analyses confirm the benefits of the added features. Keywords: Information disorder, Fake news, Propaganda, Classification, Topic modeling, Hybrid method, Neurosymbolic model, Ablation, Robustness
This position paper proposes a theory grounded NLP framework for information disorder detection integrating three explicitly connected dimensions: epistemic status, intentionality, and contextual harm. Moving beyond binary fake news classification, we argue that reliable intervention requires structured differentiation between verification outcomes, manipulation indicators, and consequence assessment. We provide concrete annotation schemas with decision rules for ambiguous cases, formal aggregation operators with monotonicity and escalation guarantees, explicit conflict resolution strategies for inconsistent signals, and standardized risk profile templates that translate multidimensional outputs into actionable routing policies. Synthesizing work on harm taxonomies, uncertainty quantification, and automated fact checking pipelines, we introduce an integration layer that preserves interpretability while enabling policy aligned deployment. We further propose a reformed evaluation protocol incorporating conformal prediction for principled abstention, calibration analysis, disagreement modeling, harm weighted metrics, and human uplift assessment to measure real decision support utility rather than standalone classifier accuracy. We position this framework as a conceptual and operational roadmap for structured misinformation assessment, outlining phased validation pathways while acknowledging that empirical validation remains essential future work.
Emotion and Information Disorder in NLP: A Systematic Mapping and Benchmark Blueprint
Renatha Vieira | Alvaro Figueira
Renatha Vieira | Alvaro Figueira
Online misinformation research in NLP has expanded rapidly, including approaches that model affective signals such as sentiment, discrete emotions, and emotion dynamics. However, the Information Disorder framework distinguishes misinformation, disinformation, and malinformation along dimensions of intention, harm, and contextual dependence, which are rarely operationalised in current datasets, tasks, and evaluation protocols. We provide a systematic mapping of 82 studies at the intersection of Information Disorder and emotion-aware NLP (51 model papers, 7 dataset papers, 24 survey/theory papers). Across empirical works (58), veracity-centric supervision dominates (72.4% binary labels), while explicit intention and harm variables appear in only 1.7% each. Evaluation relies mostly on random splits (79.3%), limiting robustness to source and temporal shifts. Emotion is represented in 43.1% of model papers, mostly as static features, with emotion dynamics and audience emotion rare. Based on these findings, we propose an operational taxonomy aligned with Information Disorder and a benchmark blueprint specifying tasks, annotation variables, split strategies, and evaluation protocols to support theory-grounded, comparable progress.
A Multilingual Linguistic Analysis of Human vs LLM-Generated News in a Disinformation Context
Silvia Gargova | Alba Perez-Montero | Elena Lloret Pastor | Paloma Moreda Pozo
Silvia Gargova | Alba Perez-Montero | Elena Lloret Pastor | Paloma Moreda Pozo
The rise of Large Language Models has shifted the Information Disorder landscape toward automated threats. This study investigates the linguistic construction of synthetic news by comparing GPT-5, Gemini 2.5, and Grok 4 across English, Spanish, and Bulgarian. Using multilingual human-authored verified news and disinformation as seeds, we analyze how prompt informativeness and model architecture influence deceptive content production. Our methodology employs five metrics: semantic similarity, factual consistency, readability, lexical richness, and persuasion technique frequency. Our analysis reveals that while prompt scarcity leads to informational loss, LLMs maintain a homogenized stylistic template regardless of input length. Unlike human authors, who intensify rhetorical and emotional markers to drive deceptive intent, LLMs adhere to a neutral register. This study identifies distinct statistical patterns in generated content characterized by hyper-standardized readability and high lexical density (p < 0.001). These features serve as robust “LLM signatures”, enabling a classification accuracy of 96% across English, Spanish, and Bulgarian. These findings suggest that generated disinformation relies on invariant syntactic structures rather than nuanced human rhetoric, providing a framework for detection tools centered on structural patterns rather than content veracity.
This paper aims to contribute to the understanding of information disorder from an epistemological perspective, by analysing the internal/cognitive as well as the external/contextual factors that determine knowledge defects. In this regard, I first compare the concept of information with the properly epistemological concept of knowledge by arguing that the latter makes explicit the two fundamental prescriptive characteristics that beliefs should have—truth and justification—that remain implicit in the epistemically neutral notion of information. Therefore, I provide an externalist account of knowledge according to which a belief is true and justified to the extent that it allows for satisfactory adaptation to the environment, in which the ecological, technological and sociological environment itself becomes an integral part of an extended cognitive system. Based on this epistemological exploration of the notion of information through an externalist conception of knowledge, I suggest that disinformation can be understood as the state of an “ignorant” collective cognitive system, that is, a closed system that establishes interactions only within a virtual environment devoid of any semantic relevance and appeal to rational justification. In conclusion, I point out that, although the digital revolution appears to fuel the expansion of ignorance, it nevertheless poses challenges that allow epistemological reflection itself to renew addressing the crisis of knowledge with extended theoretical resources.
Population Replacement Conspiracy Theories Detection on Telegram and News Headlines: Benchmarking LLMs and BERT Models in Portuguese and Italian
Erik Bran Marino | Renata Vieira
Erik Bran Marino | Renata Vieira
Disinformation has become a serious threat to the democratic stability of Western societies, with various conspiracy theories spreading from fringe spaces to mainstream media and politics. While some of these theories may seem merely absurd and harmless, others pose significant risks. Among the most dangerous are Population Replacement Conspiracy Theories (PRCTs), which promote the false narrative of a deliberate demographic substitution through immigration. Despite their disinformative nature, increasing widespread and documented connections to extremist violence and political polarization, current computational detection models primarily target COVID-19 or general conspiracy theories, lacking specialized annotated corpora and approaches for identifying PRCTs in multilingual contexts. In this work, we present the first systematic benchmark for PRCT detection in Portuguese and Italian.
Mapping Discourse Reframing: A Multi-Layer Network Approach to Italian HPV Vaccine Discourse on X (2010-2024)
Lorella Viola
Lorella Viola
Understanding how online narratives travel through coalitions is critical for identifying information disorder, yet computational analyses often rely on conservative network constructions that erase initially sparse but salient signals. This paper proposes a novel multi-layer framework that captures low-frequency signals of emerging information disorder allowing for locating where online discourse is reframed and amplified over time. The use case is 14 years of Italian discourse on X regarding the Human Papillomavirus (HPV) vaccine across three pivotal epochs (2010–2024). Utilizing hashtag co-occurrence networks, we introduce a dual-layer approach. We first identify robust core discourse coalitions through conservative community detection, revealing a stable prevention-oriented backbone contrasted with increasingly separable skepticism coalitions. We then introduce a ‘coverage’ layer and project fringe hashtags into core coalitions based on weighted connectivity. Using a manually labelled set of skeptical and conspiratorial seed tweets, we demonstrate that this core–coverage projection significantly improves the recovery of long-tail, problematic hashtags while preserving an interpretable coalition structure. Our findings characterize the structural maturation of polarized narratives and provide a methodology for mapping how discourse is reframed and amplified by information disorder over time.
up
Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
Harry Bunt
Harry Bunt
Med2Story Referential: A Domain-Specific Extension of ISO 24617-9 for Clinical Narratives Annotation
Ana Luisa Fernandes | Purificação Silvano | Nuno Guimarães | Luís Filipe Cunha | Rita Rb-Silva | Alipio Mario Jorge
Ana Luisa Fernandes | Purificação Silvano | Nuno Guimarães | Luís Filipe Cunha | Rita Rb-Silva | Alipio Mario Jorge
The semantic annotation of clinical narratives is particularly challenging due to the complexity of medical discourse and the need to integrate linguistic, semantic, and domain-specific information within a unified framework. Existing schemes tend to fall into two categories: general-purpose frameworks, which offer robust linguistic modelling but lack specialised medical representation, and domain-specific schemes, which capture clinical content yet often fail to distinguish fundamental semantic types, especially eventive expressions and referential entities. To address this gap, this study proposes Med2Story Referential, a new extension of the Text2Story annotation scheme (Silvano et al., 2021; Leal et al., 2022) (based on ISO 24617-9: 2019 ) dedicated to referential entities in clinical narratives. Building on previous work that introduced a specialised branch for eventive entities (Fernandes et al., 2025a), and informed by the UMLS Metathesaurus and expert validation from a consultant haematologist, the extension introduces eight referential categories that refine the representation of clinical actors, substances, biological entities, instruments, and documentation. The results show that ISO 24617-9: 2019 can be applied to this type of text; however, several adaptations are required, particularly with regard to the grammatical domain and the inclusion of specialised domain labels. Nonetheless, the annotation experiment conducted to validate our proposal showed that the annotation scheme and its accompanying guidelines enable a comprehensive and detailed representation of both grammatical and medical aspects. Moreover, the results indicate that the scheme can be applied effectively by annotators without medical expertise.
This article focus on a methodology for representing the semantics of polysemous markers whose meanings cannot (or do not have to) be disambiguated, even in context. We name this task (multi-)sense representation and present here the French modal verb devoir as a case study. Specifically, we reframe this task — traditionally treated as a multi-class problem — as a multi-label classification problem to account for instances that remain ambiguous due to contextual and intentional factors. In order to fine-tune our model (CamemBERT), we implement an active learning loop to enhance the annotation process and we demonstrate that combining global and local features yields the best results (F1-micro = 0.83; F1-macro = 0.79). The model is then applied on two distinct corpora, showing that the automatic analysis of devoir’s modal senses provides deeper insights into modal verb usage and facilitates comparisons across corpora differing in medium (spoken vs. written) or genre (e.g. legal discourse). Furthermore, our multi-label approach enables the detection and analysis of double-labeled instances, offering valuable applications, as for example legal discourse interpretation and second language acquisition.
A Frame and Canvas-Based Perspective-Encoding Methodology for Multimodal Semantic Annotation of Classroom Settings
Claudia Ferraz | Ely E. Matos | Frederico Belcavello | Julia Gasparetto | Juliana de Oliveira | Janina Wildfeuer | Tiago Timponi Torrent
Claudia Ferraz | Ely E. Matos | Frederico Belcavello | Julia Gasparetto | Juliana de Oliveira | Janina Wildfeuer | Tiago Timponi Torrent
We propose a methodology for the multimodal semantic annotation of classroom interactions that takes interactional canvases and semantic frames as its core analytical categories. The approach enables the systematic recording of semantic correlations among interactants, communicative modes, and material supports involved in situated meaning-making processes. The methodology encodes participant perspective by relying on the temporal alignment of multiple video recordings of the same instructional event captured from different viewpoints, allowing for the representation of how meaning construction unfolds across perceptual and interactional positions. To operationalize the proposal, we introduce an annotation tool that implements the scheme and supports the integration of multimodal data streams within a unified semantic representation framework. We conclude by discussing the limitations of the current proposal and the possibilities for extending it to other interactional settings.
Pattern Analysis (CPA) procedure developed by Hanks (2004) to manually extract recurrent language patterns from texts, can be automated using LLMs. Specifically, we examine ChatGPT and Gemini performance in the task of semantic type tagging of arguments in 150 Italian sentences realising 30 verb patterns (5 sentences per pattern). We run two experiments. In the first, we prompt ChatGPT to use the CPA ontology (about 200 hierarchically organized semantic types) in the annotation task; we provide the model with 5 sentences per pattern and ask it to assign the most specific type to the argument(s) of each sentence. In the second, we prompt both ChatGPT and Gemini to perform the task without the ontology, and ask the models to assign a single label to the argument(s) of the 5 sentences. Both experiments are performed in a zero-shot setting. We evaluate the results using the existing Italian T-PAS pattern resource as benchmark. Our results show that LLMs perform comparably well on both concrete and abstract type tagging and can therefore be used in a pilot study to support analysts in acquiring verb patterns from text.
Korean Quantification in Abstract Meaning Representation
Kiyong Lee | Chongwon Park | Younggyun Hahm | Harry Bunt | Byongrae Ryu
Kiyong Lee | Chongwon Park | Younggyun Hahm | Harry Bunt | Byongrae Ryu
This paper explores the meaning of quantification in Korean and how it is encoded in Abstract Meaning Representation (AMRg:2019) and an enriched version AMR+ accommodating Uniform Meaning Representation (UMRg:2022) and some of the contextual constraints proposed by Bos(2020). The extension makes five special references: Bunt et al. (2018), Bunt and Lee (2025), Pustejovsky et al. (2019), Bos(2020), and ISO (2025), the main reference. The aim of this paper is threefold. First, it focuses on implementing Korean AMR with the rich specification of QuantML (ISO, 2025) and its partially DRT-based semantics (Kamp and Reyle, 1993). Second, it supports the AMR multilingual development project by exploring methods for constructing a large-scale Korean AMR-annotated corpus. This line of research is necessary because Korean AMR resources remain severely underdeveloped. In addition, Korean’s agglutinative morphology and head-final syntax challenge AMR frameworks that are largely based on the analytic inflectional language English. Third, it advances the current state of the UMR 2026 multilingual shared task by contributing more fine-grained annotations of quantification specified by ISO QuantML for resource domain, individuation, distributivity, and determinacy, as well as by treating coreference and lexical or scope ambiguities in Korean.
Towards Corpus-Based Population and Visualization of ISO 24617-8 Ontology (Short Paper)
Maciej Ogrodniczuk | Dariusz Czerski
Maciej Ogrodniczuk | Dariusz Czerski
This paper presents an extension of the ISO 24617-8 ontology for discourse relations through the integration of corpus-based examples and the development of a dedicated Ontology Viewer. The goal is to bridge the gap between formal ontological representations and practical corpus-based linguistic analysis, making discourse annotation frameworks more accessible to researchers. The proposed approach introduces a method for populating the ISO ontology with instances derived from three corpora (in Polish and English) compliant with the ISO 24617-8 standard. These instances formally connect discourse relations, argument roles, and explicit connectives within a unified semantic model. The Ontology Viewer enables intuitive browsing, filtering, and full-text searching of examples by language, relation type, and connective, offering both a relation-oriented and connective-oriented perspective. The experiment demonstrates the feasibility and effectiveness of this corpus-driven instantiation method and its visualization. The system provides a foundation for future integration of multilingual discourse corpora and contributes to the development of interoperable language resources for the Semantic Web and Natural Language Processing applications.
Tracing Consensus Formation in Meetings: Annotation and Incremental Decision Modelling in the MEET Corpus
Ghazaleh Esfandiari-Baiat | Jens Edlund
Ghazaleh Esfandiari-Baiat | Jens Edlund
We present an incremental annotation scheme and discourse model designed specifically for the study of consensus formation in collaborative meetings. By grounding the representation in observable contributions and enforcing a strict no-lookahead principle, the model provides a tractable way to analyse how decisions emerge over the course of interaction. The resulting structures are intentionally minimal yet expressive enough to capture the evolving task state and support dynamic visualisation and replay of the decision process. A web-based reference implementation of the model demonstrates how the evolving decision state can be inspected and replayed during analysis. Together with a suitable corpus, this framework provides a practical foundation for investigating the multimodal dynamics of collaborative decision-making in professional meetings.
GeoAffect: A Multi-Layer Annotation Schema and Few-Shot LLM Evaluation for Geoaffective Analysis of Literary Texts
Fotini Koidaki | Stergios Chatzykiriakidis
Fotini Koidaki | Stergios Chatzykiriakidis
GeoAffect is an annotation framework that has been especially developed to capture how places are emotionally framed in literary narrative. The project focuses on nineteenth-century Greek prose fiction and brings together named entity recognition with an affect schema that distinguishes experiential, appraisal, and identity-oriented relations to place. The annotation design linked entities, emotion spans, and rhetorical devices, allowing us to model not only sentiment but also forms of belonging, alienation, and longing. To test the schema, we created a manually annotated gold dataset of approximately 360 sentences and evaluated thirteen Large Language Models in a few-shot setting for both entity recognition and affect classification. The results indicate that, with carefully designed prompts and selection strategies, LLMs can support structured geoaffective annotation even in low-resource historical language contexts.
Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: The Case of FrameNet Annotation
Frederico Belcavello | Ely E. Matos | Arthur Lorenzi | Lisandra Bonoto | Livia Pádua Ruiz | Luiz Fernando Pereira | Victor Herbst | Yulla Liquer Navarro | Helen de Andrade Abreu | Lívia Vicente Dutra | Tiago Timponi Torrent
Frederico Belcavello | Ely E. Matos | Arthur Lorenzi | Lisandra Bonoto | Livia Pádua Ruiz | Luiz Fernando Pereira | Victor Herbst | Yulla Liquer Navarro | Helen de Andrade Abreu | Lívia Vicente Dutra | Tiago Timponi Torrent
The use of LLM-based applications as a means to accelerate and/or substitute human labor in the creation of language resources and datasets is a reality. Nonetheless, despite the potential of such tools for linguistic research, an evaluation of their performance and impact on the creation of annotated datasets, especially under a perspectivized approach to NLP, is still missing. This paper contributes to the reduction of this gap by reporting on an extensive evaluation of the (semi-)automatization of FrameNet-like semantic annotation by the use of an LLM-based semantic role labeler. The methodology employed compares annotation time, coverage, and diversity in three experimental settings: manual, automatic, and semi-automatic annotation. Results show that the hybrid, semi-automatic annotation setting leads to increased frame diversity and similar annotation coverage, when compared to the human-only setting, while the automatic setting performs considerably worse in all metrics, except for annotation time, which remains similar.
Annotating Word Meanings over Time: The Trade-off between Scalability, Reliability and Expressivity Power
Pierluigi Cassotti | Nina Tahmasebi
Pierluigi Cassotti | Nina Tahmasebi
Annotating the meanings of a word over time in order to document their emergence or disappearance presents substantial implementation challenges. These difficulties arise for several reasons, notably the need for sufficient expressive power in the annotation paradigm to capture unconventional or rare meanings, as well as issues of scalability related to the number of annotations required. The first challenge is particularly acute in the context of historical texts, where modern annotators must interpret word meanings in sources that are temporally distant and often absent from contemporary dictionaries and language use. The second challenge is inherent to the distribution of word meanings, which tend to occur sparsely and intermittently over long time spans. In this paper, we examine several annotation paradigms, discussing their respective advantages and limitations. We also present a pilot study on English and Swedish. Our results indicate that a usage-sense inventory based annotation paradigm can be adopted in place of a usage-pairs-based approach while maintaining expressivity power and reducing the complexity from quadratic to linear.
Dialogue interactions have varied internal structure, with flow varying, inter alia, in face of both difficulty and agreement. This study investigates eye-gaze in the linguistic progression of interactions. We observe the relation of gaze to illocutionary functions of turns through dialogue acts, and to how turns present “new” or “old” content, through lexical entropy and repetition. Results on the HCRC Map Task corpus, enabled by an event alignment annotation method described, show how gaze is related to linguistic progression. A gaze towards the conversation partner at the end of a turn tends to align with complexity and difficulties being expressed in the turn, while keeping gaze down at the map is more typical of obstacle- and disagreement-free interactions. Addressees who look up or off at the start of a turn show evidence of lexicon adaptation to gaze values.
CATS: An Annotation Scheme of Causality and Temporal Structure
Nana Yu | Purificação Silvano | Luís Filipe Cunha | Alípio Jorge
Nana Yu | Purificação Silvano | Luís Filipe Cunha | Alípio Jorge
This paper presents CATS, a causal and temporal annotation scheme designed to jointly represent causal relations and temporal structures in news texts. The proposed framework integrates components of ISO 24617 Semantic Annotation Framework (SemAF), drawing in particular on Part 1 (Time and Events) (ISO 24617-1: 2012iso24617) and Part 8 (Semantic Relations in Discourse) (ISO 24617-8: 2016iso24617-8). Building on the Text2Story annotation framework (CITATION), the scheme adapts and extends its principles for representing temporal information while introducing new entities and links for modeling causal relations. The resulting annotation model enables the integrated representation of causal arguments, events, temporal relations, and causal signals within a unified structure. By jointly capturing causal and temporal dependencies, CATS provides a resource for studying the interaction between causality and temporality in discourse and supports downstream NLP tasks such as event extraction, temporal ordering, and causal reasoning.
ISO-TimeML Semantics for Interlinking Annotations
Harry Bunt | Alex Chengyu Fang | Kiyong Lee | Volha Petukhova | James Pustejovsky | Purificação Silvano
Harry Bunt | Alex Chengyu Fang | Kiyong Lee | Volha Petukhova | James Pustejovsky | Purificação Silvano
This paper describes a step in the development of a methodology for combining annotation made with different annotation schemes. The methodology, called ‘interlinking’, assumes that different annotations of the same data will contain certain elements that refer to the same entities. This can be represented by a set of ‘identity links’. These links are used for constructing a single, integrated annotation structure at the level of abstract syntax with a semantic interpretation. In this paper we focus on the interlinking of annotations of time and events with ISO-TimeML (ISO 24617-1:2012) and quantification with QuantML (ISO 24617-12:2025).interlinking annotations is in practice only feasible if the respective annotation schemes use the same or convertible representation and interpretation formalisms. Since QuantML and ISO-TimeML use different formalisms and QuantML has a more fully developed semantics than ISO-TimeML, we developed a new, DRT–based semantics for ISO-TimeML which is presented and discussed in this paper.
From Categories to Decisions: A Framework for Attitudinal Analysis of Evaluative Language
Jiamei Zeng | Haitao Wang | Harry Bunt | Xinyu Cao | Min Dong | Tianyong Hao | Kiyong Lee | James Pustejovsky | Laurent Romary | Jianfang Zong | François Claude Rey | Sylviane Cardey | Yangli Jia | Shengqing Liao | Alex Chengyu Fang
Jiamei Zeng | Haitao Wang | Harry Bunt | Xinyu Cao | Min Dong | Tianyong Hao | Kiyong Lee | James Pustejovsky | Laurent Romary | Jianfang Zong | François Claude Rey | Sylviane Cardey | Yangli Jia | Shengqing Liao | Alex Chengyu Fang
The study reported in this paper aims to contribute to the development of an annotation scheme for evaluative language, based on Appraisal Theory, that addresses key sources of classification problems. In particular. it aims to develops a unified annotation scheme that proposes (1) a three-component annotation model comprising Appraiser, Appraised and Appraisal Element, (2) the operationalised distinction between Affect and Appreciation governed by a criterion of experiencer salience and a criterion distinguishing personal emotions from evaluations of conduct, and (3) a decision framework for the Judgement-Appreciation distinction structured on the target and lexis types operating through override conditions and substitution tests. The revised framework is illustrated with examples selected from a corpus of news discourse in English and is designed to be replicable across future Appraisal-based studies of evaluative language.
up
Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26
Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26
Gilles Sérasset | Katerina Gkirtzou | Michael Cochez | Jan-Christoph Kalo
Gilles Sérasset | Katerina Gkirtzou | Michael Cochez | Jan-Christoph Kalo
Linguistic Initialization for Inductive Reasoning in Heterogeneous Knowledge Graphs
Daniele Pasquini | Danilo Croce | Roberto Basili
Daniele Pasquini | Danilo Croce | Roberto Basili
Knowledge Graphs (KGs) provide explicit relational structure, while Large Language Models (LLMs) encode rich semantic knowledge. We propose a lightweight linguistic initialization strategy for heterogeneous link prediction that improves robustness under sparsity and imbalance. For each node, we construct a compact textual view combining intrinsic description and local neighborhood context, encode it with a pre-trained language model, and use the resulting embeddings to initialize a relation-aware GNN. This design preserves standard message passing while providing early semantically meaningful representations. Across multiple imbalance regimes and strict entity-to-entity cold-start settings, the proposed initialization consistently improves over random initialization and reduces degree-dependent degradation. Our results show that semantic grounding can be integrated into heterogeneous GNN pipelines with minimal architectural changes and strong empirical benefits.
OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining
Rian Touchent | Éric de la Clergerie
Rian Touchent | Éric de la Clergerie
We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach has three stages: random walks through ontology graphs capture hierarchical and causal relations between medical codes, a large language model reformulates these walks into fluent textbook-style prose, and the resulting text is used to train ModernCamemBERT, a 149M-parameter French encoder, with two objectives on the same data: masked language modeling and relation prediction between code pairs. On three French medical coding benchmarks (FRACCO, Cantemist-FR, Distemist-FR), OntoBook achieves significant improvements over MLM-only pretraining, with +2.5 micro-F1 on FRACCO and +8.0 micro-F1 on Distemist. We find that alignment between objectives is necessary: misaligned training, where each task uses different data, causes a 30-point degradation. We release 1.3 million LLM-reformulated medical textbooks across three French ontologies (CIM-10, CCAM, ATC) and pretrained model checkpoints.
Conversational Control with Ontologies for Large Language Models: A Lightweight Framework for Constrained Generation
Barbara Gendron | Gael Guibon | Mathieu d’Aquin
Barbara Gendron | Gael Guibon | Mathieu d’Aquin
Conversational agents based on Large Language Models (LLMs) have recently emerged as powerful tools for human-computer interaction. Nevertheless, their black-box nature implies challenges in predictability and a lack of personalization, both of which can be addressed by controlled generation. This work proposes an end-to-end method to obtain modular and explainable control over LLM outputs through ontological definitions of aspects related to the conversation. Key aspects are modeled and used as constraints; we then further fine-tune the LLM to generate content accordingly. To validate our approach, we explore two tasks that tackle two key conversational aspects: the English proficiency level and the polarity profile of the content. Using a hybrid fine-tuning procedure on seven state-of-the-art, open-weight conversational LLMs, we show that our method consistently outperforms pre-trained baselines, even on smaller models. Beyond quantitative gains, the framework remains model-agnostic, lightweight and interpretable, enabling reusable control strategies that can be extended to new domains and interaction goals. This approach enhances alignment with strategy instructions and demonstrates the effectiveness of ontology-driven control in conversational systems.
Is One Token All It Takes? Graph Pooling Tokens for LLM-based GraphQA
Ankit Grover | Lodovico Giaretta | Remi Bourgerie | Sarunas Girdzijauskas
Ankit Grover | Lodovico Giaretta | Remi Bourgerie | Sarunas Girdzijauskas
The integration of Graph Neural Networks (GNNs) with Large Language Models (LLMs) has emerged as a promising paradigm for Graph Question Answering (GraphQA). However, effective methods for encoding complex structural information into the LLM’s latent space remain an open challenge. Current state-of-the-art architectures, such as G-Retriever, typically rely on standard GNNs and aggressive mean pooling to compress entire graph substructures into a single token, creating a severe information bottleneck. This work mitigates this bottleneck by investigating two orthogonal strategies: (1) increasing the bandwidth of the graph-to-LLM interface via multi-token pooling, and (2) enhancing the semantic quality of the graph encoder via global attention mechanisms. We evaluate a suite of hierarchical pruning and clustering-based pooling operators—including Top-k, SAGPool, DiffPool, MinCutPool, and Virtual Node Pooling (VNPool) to project graph data into multiple learnable tokens. Empirically, we demonstrate that while pooling introduces significant instability during soft prompt tuning, the application of Low-Rank Adaptation (LoRA) effectively stabilizes these projections, allowing compressed representations to rival full-graph baselines (achieving ∼73% Hit@1 on WebQSP). Conceptually, we demonstrate that a Graph Transformer with VNPool implementation functions structurally as a single-layer Perceiver IO encoder. Finally, we adapt the FandE (Features and Edges) Score to the generative GraphQA domain. Our analysis reveals that current the GraphQA benchmark suffer from representational saturation, where the target answers are often highly correlated with isolated node features. The implementation of our experiments is available at https://anonymous.4open.science/r/Pool-A85D/README.md.
Advances in Large Language Models (LLMs) have made it possible to convert natural language questions into executable database queries. Text2Cypher focuses on graph databases, converting user questions into queries and providing natural language access to graph-structured data. While significant progress has been made through prompt design, fine-tuning, and iterative refinement, less attention has been given to adaptive test-time strategies that combine multiple generated outputs. In this work, we investigate the impact of confidence-based test-time strategies specifically on the Text2Cypher task by evaluating the model’s traces, which are the sequence of tokens generated during the construction of the query. We show that reasoning models generate diverse query candidates but frequently produce syntactic errors and incomplete structures, limiting executability. On the other hand, instruction-tuned models yield more reliable outputs but lack sufficient diversity for effective confidence-based selection. Further, by tuning diversity parameters such as top-p and temperature, we observe consistent improvements in both query accuracy and execution success. Experiments across multiple instruction-tuned models confirm that combining diversity-controlled generation with confidence-aware inference provides a practical, model-agnostic method for improving query generation.
Integrating Knowledge Graph and Large Language Models for Defining Business Strategies
Eleonora Ghizzota | Alex Jordan | Alessandro Petruzzelli | Lucia Siciliani | Giuseppe Spillo | Pierpaolo Basile | Davide Sola | Giovanni Scarso Borioli | Giovanni Semeraro
Eleonora Ghizzota | Alex Jordan | Alessandro Petruzzelli | Lucia Siciliani | Giuseppe Spillo | Pierpaolo Basile | Davide Sola | Giovanni Scarso Borioli | Giovanni Semeraro
Effective business strategy formulation requires synthesising diverse, often conflicting information sources into coherent action plans. While Large Language Models (LLMs) show potential for processing textual information at scale, their application is limited by hallucinations and a lack of grounding in proprietary data. This paper proposes a methodology that integrates a domain-specific Knowledge Graph (KG) with a GraphRAG pipeline to generate strategic briefing documents, or Primers, which provide a structured overview of a company’s competitive environment. Our approach utilizes an ontology-first framework and Cypher-based graph traversal to capture the relational nature of strategic knowledge beyond simple vector retrieval. Experimental results on a Q&A dataset demonstrate that the Vector + Cypher retrieval strategy significantly improves grounding over LLM-only baselines and outperforms naive vector retrieval in terms of completeness and usefulness. These findings suggest that the synergy of LLMs and structured KGs provides a robust foundation for automated strategic analysis in real-world business scenarios.
GROUNDEDKG-RAG: Grounded Knowledge Graph Index for Long-document Question Answering
Tianyi Zhang | Andreas Marfurt
Tianyi Zhang | Andreas Marfurt
Retrieval-augmented generation (RAG) systems have been widely adopted in contemporary large language models (LLMs) due to their ability to improve generation quality while reducing the required input context length. In this work, we focus on RAG systems for long-document question answering. Current approaches suffer from a heavy reliance on LLM descriptions resulting in high resource consumption and latency, repetitive content across hierarchical levels, and hallucinations due to no or limited grounding in the source text. To improve both efficiency and factual accuracy through grounding, we propose GroundedKG-RAG, a RAG system in which the knowledge graph is explicitly extracted from and grounded in the source document. Specifically, we define nodes in GroundedKG as entities and actions, and edges as temporal or semantic relations, with each node and edge grounded in the original sentences. We construct GroundedKG from semantic role labeling (SRL) and abstract meaning representation (AMR) parses and then embed it for retrieval. During querying, we apply the same transformation to the query and retrieve the most relevant sentences from the grounded source text for question answering. We evaluate GroundedKG-RAG on examples from the NarrativeQA dataset and find that it performs on par with a state-of-the art proprietary long-context model at smaller cost and outperforms a competitive baseline. Additionally, our GroundedKG is interpretable and readable by humans, facilitating auditing of results and error analysis.
Ontology-Guided Synthetic Data Generation for Low-Resource Information Extraction: A Case Study in IT Heritage Domain
Nakanyseth Vuth | Emrick Poncet | Gilles Sérasset | Didier Schwab | Caroline Djambian | Benjamin Lecouteux
Nakanyseth Vuth | Emrick Poncet | Gilles Sérasset | Didier Schwab | Caroline Djambian | Benjamin Lecouteux
Information Extraction (IE) in specialized domains often suffers from a severe cold-start problem due to the high cost of expert annotation. Recent Reverse-IE approaches leverage knowledge graphs to generate synthetic training corpora, but typically assume the availability of an existing knowledge base. In this work, we propose an ontology-driven pipeline for synthetic supervision that removes this requirement. Starting from a formal domain ontology, we introduce a stochastic motif sampling strategy that constructs schema-consistent Knowledge Graph structures with controllable topology, which are then verbalized into natural language. This ontology-first formulation also allows direct control over the data generation process, enabling oversampling of underrepresented entity types or relation patterns. Applied to the IT Heritage domain, our approach produces a fully labeled NER/RE corpus without large-scale manual annotation. Evaluation in a low-resource setting shows that while the synthetic corpus lacks the linguistic diversity of gold data, its scalability produces training sets large enough to alleviate the cold-start problem, making ontology-guided motif generation a practical strategy for domains where gold annotation is limited.
A Clinical SKOS Ontology and Evaluation Benchmark for LLM Query Generation over ICU Knowledge Graphs
Khurrum Ali
Khurrum Ali
Whencliniciansquerydatabasesusingeverydaylanguage—“codebluepatients” or"sugardisease"—LargeLanguage Models must bridge a lexical gap between colloquial speech and formal clinical terminology. While highly capable cloudmodelscanleverageexternalontologieslike SKOStoresolvethese termsviaSPARQLqueries, hospital privacy regulations often mandate the use of air-gapped local LLMs (4–8B parameters). We evaluate query generation across scales (Gemini 2.0 Flash vs. LLaMA 3.1 8B) using ClinSKOS-ICU, a curated ontology of 421 ICU concepts, and ClinNLU, an evaluation benchmark. We identify a critical “Privacy Penalty”: while Gemini achieves 90.2% ontology deferral under an RDF+SKOS architecture, local LLMs exhibit a 100% “Semantic Bypass” vulnerability, hardcodingformaltermsintoqueriesratherthandeferringtothegraph. ToimprovelocalLLMgrounding, weintroduce Architectural Decomposition, a pipeline that restricts the LLM to Grammar-Constrained JSON entity extraction and delegates query generation to deterministic code. This structural pivot entirely eliminates Semantic Bypass (0%) and achieves an 80.4% ontology deferral rate on an 8B model, suggesting that decoupled extraction is highly effective for enforcing W3C semantic compliance on privacy-preserving local hardware.
ReX-GG: A LLM Ensemble Pipeline for Relation-extraction and Graph Generation
Giacomo Magnifico | Eduard Barbu
Giacomo Magnifico | Eduard Barbu
Current LLM ensemble frameworks focus on multi-step setups with additional modules for answer ranking, often opting for token and span analysis rather than structured outputs, leading to heavyweight architectures with potential fail states along the pipeline. Faster, lighter solutions are more vulnerable to hallucination propagation and can lack output control in more complex pipelines. This paper proposes a customisable, lightweight ensemble workflow of coordinated Large Language Models that leverages JSON-structured outputs and anonymous peer-review ranking to mitigate hallucinatory outputs and single-model failure points. The pipeline is demonstrated on a relation extraction task applied to English popular science articles, targeting four ontologically-grounded relation types (strong causation, weak causation, contrastive, and compositional), with semantic node canonicalisation and interactive, colour-coded HTML causal graphs as the final output. Performance is evaluated through an anonymous user study, achieving an average perceived accuracy of 0.778 against a human-annotated gold standard. The modular architecture supports flexible deployment across both API-based and in-house LLM setups, and the full framework is released under an open license to foster reproducibility and collaborative research.
Graph Fusion across Languages Using Large Language Models
Kaung Myat Kyaw | Khush Agarwal | Jonathan Chan
Kaung Myat Kyaw | Khush Agarwal | Jonathan Chan
Combining multiple knowledge graphs (KGs) across linguistic boundaries is a persistent challenge due to semantic heterogeneity and the complexity of graph environments. We propose a framework for cross-lingual graph fusion, leveraging the in-context reasoning and multilingual semantic priors of Large Language Models (LLMs). The framework implements structural linearization by mapping triplets directly into natural language sequences (e.g., [head] [relation] [tail]), enabling the LLM to map relations and reconcile entities between an evolving fused graph and a new candidate graph. Evaluated on the DBP15K dataset, this exploratory study demonstrates that LLMs can serve as a universal semantic bridge to resolve cross-lingual discrepancies. Results show the successful sequential agglomeration of multiple heterogeneous graphs, offering a scalable, modular solution for continuous knowledge synthesis in multi-source, multilingual environments. Our implementation and experimental framework are publicly available in our repository: https://github.com/IC2-Lab-KMUTT/Multilingual-Graph-Fusion
Efficient KG-Augmented RAG with Reusable Graph Community Summaries
Maha Karkout | Maria Andreevna Khodorchenko | Nikolay Alekseevich Butakov | Denis Nasonov
Maha Karkout | Maria Andreevna Khodorchenko | Nikolay Alekseevich Butakov | Denis Nasonov
Retrieval-augmented generation (RAG) performs well for localized factual queries but struggles with complex questions requiring multi-section evidence integration. Graph-based approaches introduce relational structure, yet their practical integration into QA pipelines involves significant query-time overhead. We present a practical KG-augmented RAG (KG-RAG) design that builds a knowledge graph offline with an LLM, converts graph communities into reusable summaries, and retrieves these summaries jointly with textual evidence at query time. We compare dense RAG, pure GraphRAG, and the proposed hybrid on two benchmarks representing complementary retrieval paradigms: QASPER (intra-document reasoning over scientific papers) and ObliQA (cross-document reasoning over regulatory texts). Results show that pure GraphRAG does not consistently outperform dense retrieval, whereas the hybrid configuration systematically improves relevance, correctness, and completeness while maintaining substantially lower latency than full graph-based inference.
Stack2Graph: A Structured Knowledge Representation of Stack Overflow Data for Retrieval-based Question Answering
Lukas Amadeus Kleybolte | Viviana Ventura | Alessandra Zarcone
Lukas Amadeus Kleybolte | Viviana Ventura | Alessandra Zarcone
Community-based platforms like Stack Overflow (SO) offer a vast and diverse source of software development knowledge, combining natural language data with code snippets. Resources built from SO have been widely used to support downstream tasks in software engineering and natural language processing. However, no existing resource fully reconstructs and connects the complete range of information available on SO, leveraging its structure. We introduce Stack2Graph, a large-scale resource that preserves the forum’s structural relationships in a semantically explicit form by combining a knowledge graph with a vector database. This hybrid design captures the intrinsic links between questions, answers, comments, tags, and cross-references, bridging symbolic and vector-based representations to enable structured and multi-hop retrieval. The goal is to make SO knowledge more efficiently accessible for LLM-based systems and easier to integrate into downstream applications. To evaluate its impact, we integrate Stack2Graph into a zero-shot pipeline for multiple-choice question answering on CodeMMLU. Results show that retrieval augmentation particularly benefits mid-sized general-purpose models, with substantial gains in API- and framework-oriented tasks.
LLM-based Atomic Propositions Help Weak Extractors: Evaluation of a Propositioner for Triplet Extraction
Luc Pommeret | Thomas Gerald | Christophe Servan | Sahar Ghannay | Patrick Paroubek | Sophie Rosset
Luc Pommeret | Thomas Gerald | Christophe Servan | Sahar Ghannay | Patrick Paroubek | Sophie Rosset
Knowledge Graph construction from natural language requires extracting structured triplets from complex, information-dense sentences. In this paper, we investigate if the decomposition of text into atomic propositions (minimal, semantically autonomous units of information) can improve the triplet extraction. We introduce MPropositionneur-V2, a small multilingual model covering six European languages trained by knowledge distillation from Qwen3-32B into a Qwen3-0.6B architecture, and we evaluate its integration into two extraction paradigms: entity-centric (GLiREL) and generative (Qwen3). Experiments on SMiLER, FewRel, DocRED and CaRB show that atomic propositions benefit weaker extractors (GLiREL, CoreNLP, 0.6B models), improving relation recall and, in the multilingual setting, overall accuracy. For stronger LLMs, a fallback combination strategy recovers entity recall losses while preserving the gains in relation extraction. These results show that atomic propositions are an interpretable intermediate data structure that complements extractors without replacing them.
Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges
Daniel Schlör | Marius Bohn | Maximilian Wolf | Kevin Bergner | Christian Goldschmied | Andreas Hotho
Daniel Schlör | Marius Bohn | Maximilian Wolf | Kevin Bergner | Christian Goldschmied | Andreas Hotho
Evaluating Large Language Model (LLM) agents on complex multi-step cybersecurity tasks requires structured, reproducible evaluation rubrics. We present BraceGreen, a framework that formalizes Capture-the-Flag (CTF) attack paths as knowledge graphs and uses them as gold-standard rubrics for agentic LLM evaluation. Each node in our knowledge graphs represents an attack step annotated with MITRE ATT&CK tactics, goals, commands, expected outputs, and semantic outcomes, while edges encode prerequisites, dependencies, and alternative paths. Our LangGraph-based evaluation workflow employs LLM-as-judge with chain-of-thought reasoning to semantically compare agent predictions against knowledge graph-encoded alternatives. We contribute a benchmark of 7 CTF machines with knowledge graph annotations, three evaluation modes (command prediction, goal inference, anticipated result), and integration with live machine infrastructure via virtual machines and a MCP server. Our approach bridges the gap between unstructured CTF writeups and graph-structured evaluation rubrics.
We present an end-to-end pipeline for constructing a domain-specific knowledge graph from instructional text using Large Language Model assisted extraction. Applied to the Icelandic Riding Levels, a 602 pages training corpus for riders of the Icelandic Horse, the pipeline produces a hyper-relational knowledge graph of 9,382 nodes and 16,423 edges, where schema-constrained qualifiers preserve the conditional and procedural context that standard triples discard. To evaluate the resulting graph, we introduce the first expert validated question answering benchmark for this domain: 252 questions across four reasoning categories. Comparing Graph-, Text-, and Hybrid-retrieval augmented generation methods, we find that Text-based achieves the highest overall accuracy, but that Graph-based provides the only correct answer for a subset of queries, particularly where the corpus contains competing values for the same fact. A failure analysis traces the majority of Graph-based retrieval errors to context dilution at high-degree hub nodes, an algorithmic limitation in graph traversal. We discuss implications for adaptive retrieval strategies that route queries to the appropriate modality.
The Structure-Content Trade-off in Knowledge Graph Retrieval: A Diagnostic Study of Question Decomposition
Valentin Six | Gaël de Chalendar | Evan Dufraisse
Valentin Six | Gaël de Chalendar | Evan Dufraisse
Large Language Models are increasingly combined with knowledge graphs to support multi-hop factual reasoning. A classical strategy for handling complex questions in such settings is question decomposition, where a question is broken into simpler subquestions to guide retrieval. While decomposition can improve relevance, its impact on the structure and connectivity of the retrieved information, as well as the implications for downstream reasoning, remain unclear. In this work, we present a diagnostic study of the effects of question decomposition on knowledge graph retrieval. We use a simple parametric interpolation between retrieval guided by the original question and its subquestions, allowing us to vary retrieval focus in a controlled manner. By softly anchoring subquestion-level retrieval to the original question, we allow structural properties of the retrieved subgraph to change naturally, without post-hoc enforcement of connectivity. Across different multi-hop QA benchmarks, we observe a consistent structure-content trade-off: subquestion-focused retrieval improves content precision but fragments the retrieved graph, whereas question-focused retrieval preserves structural coherence at the cost of relevance. Downstream QA performance peaks at intermediate settings, where sufficient connectivity emerges while maintaining high relevance. These results highlight the importance of jointly considering content and structure when designing retrieval strategies for reasoning over structured knowledge.
Quantifying Retrieval Quality in GraphRAG: A Schema-Agnostic Approach
Thibaud Vanmechelen | Alexandre Achten | Zaineb Gabsi | Sabri Skhiri
Thibaud Vanmechelen | Alexandre Achten | Zaineb Gabsi | Sabri Skhiri
While LLMs have achieved significant success in natural language tasks, their tendency to hallucinate remains a critical challenge. RAG tries to address this issue by grounding models in external data; however, standard vector-based RAGs often fail when working with highly interconnected datasets. GraphRAG has emerged as a superior alternative in this setting by modelling the relational topology, yet evaluating GraphRAGs remains challenging. Current benchmarks predominantly focus on the final LLM-generated output frequently overlooking the structural accuracy of the underlying retrieval process. In this paper, we propose a novel schema-agnostic framework for the automated generation of synthetic evaluation datasets from KGs. Unlike previous approaches, our framework establishes a rigorous, deterministic ground truth to specifically quantify the retriever performance across nine distinct query categories, including multi-hop and aggregation tasks. We demonstrate the utility of this benchmark by applying it to a biochemical KG and evaluating four diverse retrieval architectures. Our results indicate that agentic, LLM-driven retrievers provide the highest recall and reasoning capacity, effectively navigating complex topologies where other methods struggle. This work provides a robust, scalable methodology for performance tracking, shifting the evaluation of GraphRAG toward a more topologically precise standard.
A Wikidata-Based Framework to Measure Cross-Lingual Bias in Multilingual Large Language Models
Mouloud Iferroudjene | Lisa Poggel | Andrea Schimmenti | Duo Yang | Kanchan Shivashankar | Jan-Christoph Kalo | Marta Boscariol
Mouloud Iferroudjene | Lisa Poggel | Andrea Schimmenti | Duo Yang | Kanchan Shivashankar | Jan-Christoph Kalo | Marta Boscariol
Multilingual large language models (LLMs) are increasingly used for factual question answering, yet their accuracy varies across languages in ways that are difficult to interpret. A central challenge is that many multilingual probing benchmarks conflate multiple factors: the language used to ask the question, the cultural-linguistic context of the entities being queried, and the popularity skew of entities. In our paper, we disentangle these factors by asking: (i) how strongly does the Language of the Question (LoQ) affect factual recall, (ii) does matching LoQ to an entity-associated Language of the Entity (LoE) improve performance, and (iii) do these effects persist when entity popularity is controlled. To this end, we introduce WILA-PopQA, a new Wikidata-grounded benchmark spanning 9 languages with matched popularity profiles, and probe 12 open-weight models of varying sizes and architectures under aligned and misaligned LoQ–LoE conditions. We evaluate models’ answers to 4 types of questions about entity biographical properties in all selected languages. Results show that LoQ is the dominant source of variation. LoQ–LoE alignment does not consistently yield the highest accuracy, and performance depends on the property being asked. These results suggest that prompt language is an actionable experimental factor for multilingual factual evaluation.
Evaluating Large Language Models for Strategic Knowledge Extraction in Capability-Based Planning
Hein C. Kolk | Julia García-Fernández | Julia Bronkhorst | Roos M. Bakker
Hein C. Kolk | Julia García-Fernández | Julia Bronkhorst | Roos M. Bakker
In a security environment that is growing more complex, large national organizations like the police rely on strategic frameworks to guide their decision-making. Frameworks like the Capability Based Planning (CBP) system are used to address this, but require a vast amount of information to function properly. A significant but underused store of information lies within an organization’s own internal flow of documents, like vision statements or annual reports. We tap into this flow by proposing a method to automatically extract relevant strategic entities and structuring them within a knowledge graph. We evaluate the performance of various Large Language Models (LLMs) on a corpus of policy excerpts from the Dutch National Police in extracting relevant strategic entities and linking them to core police capabilities. We employ the novel alternative annotator test (Alt-Test) to determine if an LLM can serve as a reliable substitute for a human domain expert on this highly subjective task. Our evaluation shows that while LLMs cannot fully replace human experts, they prove to be valuable support tools by frequently identifying the same strategic information as the annotators, successfully extracting core entities and linking them to predefined capabilities.
Large Language Models for Knowledge Graph Extraction: A Schema-Constrained Evaluation Framework
Markus Ilves | Eduard Barbu | Jaan Übi
Markus Ilves | Eduard Barbu | Jaan Übi
Large language models enable zero-shot knowledge graph extraction from text, yet evaluation at the level of complete typed graphs remains an open challenge. We present a schema-constrained evaluation framework that combines an explicit ontology of six entity types and 96 relation types with structured generation guided by schema-injected prompts. Supporting both single-step and two-step extraction modes, controlled inference settings, and repeated-run stability analysis, the framework enables systematic benchmarking of LLM-based graph construction under closed ontology constraints. Four large language models Gemini 3 Pro, GPT-5.1, Claude Opus 4.5, and Mistral 7B are evaluated on DocRED using entity and triple F1, schema adherence, and run consistency. Manual review reveals that automatic triple F1 systematically underestimates extraction quality, as a substantial portion of model-predicted triples are textually valid but absent from the incomplete gold annotations. The framework, prompts, and experimental outputs are publicly available for download and experimentation.
up
Proceedings of LANLP: Bridging Ibero and Latin American NLP Communities
Proceedings of LANLP: Bridging Ibero and Latin American NLP Communities
German Rigau Claramunt | Pablo Gamallo | Rafael Muñoz Guillena | Luis Chiruzzo | Eugenio Martínez Cámara
German Rigau Claramunt | Pablo Gamallo | Rafael Muñoz Guillena | Luis Chiruzzo | Eugenio Martínez Cámara
Bridging Ibero and Latin American NLP communities
Eugenio Martínez Cámara | Luis Chiruzzo | Pablo Gamallo | Rafael Muñoz Guillena | German Rigau
Eugenio Martínez Cámara | Luis Chiruzzo | Pablo Gamallo | Rafael Muñoz Guillena | German Rigau
LANLP focuses on community-driven resource development and evaluation for Iberian languages, and diverse Latin American languages (including indigenous and minority languages). We aim to bridge regional communities to share initiatives, corpora and tools. LANLP fills this gap, fostering new contacts between Iberian and Latin American NLP research groups. The goals are to (1) highlight challenges in processing these languages, (2) share novel datasets and models, and (3) catalyze future collaborations and shared tasks. We emphasize both academic rigor and community inclusivity, encouraging contributions from established researchers and grassroots language advocates alike.
An Oral-first Interactive Agentic System for Guaraní Speakers
Samantha Adorno | Akshata Kishore Moharir | Ratna Kandala
Samantha Adorno | Akshata Kishore Moharir | Ratna Kandala
Artificial intelligence systems are often presented as universal, yet their interaction paradigms remain predominantly text-first, limiting alignment with primarily oral languages and communicative practices. Using Guaraní, an official and widely spoken language of Paraguay, as a motivating case, this work examines how language support risks remaining symbolic when spoken interaction is reduced to a speech-to-text interface. We explore an oral-first, multi-agent framing in which turn-taking, repair, shared context, and governance are treated as core components of interaction rather than peripheral features. By separating language understanding from the conversation state and permission mechanisms, the architecture makes conversational structure and control explicit, enabling reasoning over interaction dynamics rather than isolated commands. Framing conversational coordination as a cognitively motivated reasoning problem over shared state connects insights from human dialogue to the design of AI systems that are more interpretable and responsive in oral and low-resource settings.
AI-TraLow: AI-Driven Translation for Low-Resource Languages and Cultures
Antoni Oliver | Maite Melero | Felipe Sanchez-Martinez | Víctor M. Sánchez-Cartagena
Antoni Oliver | Maite Melero | Felipe Sanchez-Martinez | Víctor M. Sánchez-Cartagena
In this paper, we present AI-TraLow, a project dedicated to advancing AI-driven translation for low-resource languages and cultures. The research is structured around three primary objectives: firstly, the development of advanced data curation techniques designed to refine parallel corpora and detect machine-generated content; secondly, the exploration of integrating structured linguistic resources—such as dictionaries and grammatical rules—directly into model prompts and fine-tuning techniques to enhance translation precision; and thirdly, the mitigation of hardware constraints through knowledge distillation to produce efficient models viable for standard desktop environments. By targeting specific linguistic groups, including Iberian varieties (Aranese, Aragonese, and Asturian), Mayan languages, and languages of vulnerable migrant communities, AI-TraLow seeks to foster linguistic diversity and digital inclusion. Ultimately, this initiative delivers open-source tools and models that ensure cultural heritage is both preserved and accessible within the contemporary digital landscape.
The availability of open resources and corpora is a fundamental requirement for research in Natural Language Processing (NLP) and Computational Linguistics; however, languages spoken in Latin America and the Iberian Peninsula, particularly Indigenous, minority, and regional varieties, remain structurally under-resourced and under-represented. This paper presents a historical account of OpenCor (Latin American and Iberian Languages Open Corpora Forum), a community-driven initiative created to promote, document, and discuss open linguistic corpora and lexical resources for these languages. Conceived as a collaborative forum rather than a competitive evaluation venue, OpenCor focuses on data creation, licensing practices, sustainability, and community building. Between 2018 and 2024, OpenCor was organized as a recurring workshop co-located with major conferences, fostering dialogue across countries, institutions, and linguistic traditions. By documenting the initiative’s motivations, organizational trajectory, submission trends, and the diversity of resources presented, this paper aims to preserve institutional memory, highlight the often-invisible labor of corpus development, and provide a reference for future initiatives dedicated to openness and linguistic diversity.
MedicaLLM: LLM-Driven Speech and Language Solutions for Healthcare
Ronghao Pan | Pedro José Vivancos-Vicente | Juan Salvador Castejón-Garrido | Tomás Bernal-Beltrán | Rafael Valencia-Garcia
Ronghao Pan | Pedro José Vivancos-Vicente | Juan Salvador Castejón-Garrido | Tomás Bernal-Beltrán | Rafael Valencia-Garcia
Although healthcare documentation is increasingly dependent on speech-based clinical interactions, general-purpose Automatic Speech Recognition (ASR) and Large Language Models (LLMs) lack the domain adaptation, structured control and interoperability guarantees required in regulated medical environments. These limitations often result in transcription errors, hallucinated content, and limited alignment with standardized coding systems. This paper introduces MedicaLLM, a multilingual, end-to-end framework integrating domain-adapted ASR, LLM-based structured report generation, and ontology-driven semantic enrichment within a modular architecture for clinical documentation. MedicaLLM combines medical interview transcription with structured report generation, summarization, and error correction; Named Entity Recognition (NER); and Medical Entity Linking (MEL) to align with standards such as SNOMED-CT and ICD-10. Deployed as a secure software as a service (SaaS) platform with REST API integration, MedicaLLM aims to reduce the administrative burden, improve the quality of documentation, and enhance semantic interoperability across healthcare systems, all while maintaining computational efficiency and clinical reliability.
mCS-LM: Multimodal Customer Service and Incident Management Systems Based on Large Language Models
Carlos Díaz-Morales | Marcos Checa-Rubio | Tomás Bernal-Beltrán | Ronghao Pan | David Barbáchano | María del Pilar Salas-Zárate | Mario Andrés Paredes-Valverde | Rafael Valencia-Garcia
Carlos Díaz-Morales | Marcos Checa-Rubio | Tomás Bernal-Beltrán | Ronghao Pan | David Barbáchano | María del Pilar Salas-Zárate | Mario Andrés Paredes-Valverde | Rafael Valencia-Garcia
Customer service and incident management increasingly rely on multimodal evidence, combining text, images and audio. However, general-purpose models lack domain grounding, structured output control and reliability guarantees required in regulated enterprise environments, often leading to hallucinated responses and limiting their practical deployment. This paper presents mCS-LM, a multilingual multimodal framework that integrates Large Language Models (LLMs), Visual Language Models (VLMs), Audio Language Models (ALMs) and Retrieval-Augmented Generation (RAG) within a modular and traceable architecture tailored to customer service and incident management. The system introduces complementary processing flows: (i) perception modules for visual and audio understanding aligned with LLM-based reasoning, and (ii) structured report generation from multimodal evidence through supervised fine-tuning using QLoRA and efficient adaptation techniques. To mitigate hallucinations and improve factual reliability, the framework incorporates vector databases and multimodal RAG pipelines that retrieve domain-specific knowledge from external corporate sources. Formal structural schemas and validation mechanisms enforce output consistency and syntactic correctness. The platform is deployed as a web-based system with REST API integration, enabling scalable multimodal interaction across channels such as instant messaging, email and web chat. Experimental results demonstrate that multimodal generative models can be specialized for structured, domain-constrained enterprise tasks while maintaining computational viability and robustness.
SAFEWORDs: un marco reproducible para anonimización conforme al RGPD y evaluación de generación en lenguas cooficiales
Rafael Muñoz Guillena | Manuel Palomar | Elena Lloret | Nuria Fernández
Rafael Muñoz Guillena | Manuel Palomar | Elena Lloret | Nuria Fernández
Los Grandes Modelos de Lenguaje (LLMs) abren oportunidades para el Procesamiento del Lenguaje Natural (PLN) en contextos institucionales, si bien plantean riesgos críticos en entornos regulados y multilingües, especialmente en lo relativo a protección de datos personales, trazabilidad de decisiones y equidad entre lenguas con distinta disponibilidad de recursos. Presentamos SAFEWORDs, proyecto que acaba de iniciarse en el marco del proyecto coordinado “HumanAIze” (Plan Nacional de Inteligencia Artificial 2025, España), que propone un marco reproducible de privacy-by-design y ethics-by-design para la evaluación y alineación de LLMs en las lenguas oficiales de la Península Ibérica (español, catalán, valenciano, gallego y euskera). El marco integra: (i) anonimización automática conforme al RGPD, con protocolos explícitos de detección de fuga residual y verificación adversarial; (ii) transformación orientada a la accesibilidad textual y al lenguaje claro; y (iii) evaluación en el dominio biomédico, donde la sensibilidad de los datos y la precisión terminológica exigen mecanismos adicionales de control generativo. Desde el punto de vista metodológico, se comparan configuraciones zero-shot y few-shot, y se documentan prompts, hiperparámetros y recursos para facilitar la replicabilidad y la gobernanza de recursos. Además de sintetizar resultados de referencia de la literatura para contextualizar métricas y órdenes de magnitud esperables, el trabajo discute implicaciones éticas y limitaciones del enfoque propuesto. La propuesta se alinea con las líneas de trabajo de SEPLN y con los objetivos de LANLP, al establecer protocolos transferibles para el desarrollo de tecnologías lingüísticas confiables en ecosistemas caracterizados por variación dialectal y lenguas infrarepresentadas.
Exploration of Sentence Representations in Spanish BERT-like Models
Gonzalo Herrera | Aiala Rosá | Luis Chiruzzo
Gonzalo Herrera | Aiala Rosá | Luis Chiruzzo
Transformer-based language models, ubiquitous in NLP nowadays, generate internal representations (embeddings) of words and sentences. Yet, systematic comparisons of embedding strategies from various models remain limited. In this work, we evaluate Spanish embeddings from several BERT-like models (BETO, multilingual BERT, XLM-RoBERTa, ROUBERTa) to understand their syntactic and semantic capabilities across layers. We propose novel sentence-level analogy tests to probe generalization. Results show tasks like verb negation or word reordering perform best with embeddings from earlier layers, while nuanced semantic distinctions—such as agent or patient gender—are better captured by deeper layers. Our findings provide guidelines for embedding strategies and offer a foundation for further NLP research.
up
Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026)
Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026)
John P. McCrae | Katerina Gkirtzou | Fahad Khan | Patricia Martin Chozas | Sara Carvalho | Erin Canning
John P. McCrae | Katerina Gkirtzou | Fahad Khan | Patricia Martin Chozas | Sara Carvalho | Erin Canning
Modeling Topics as Linguistic Linked Open Data: A First Attempt Using BERTopic, Ontolex-Lemon and FrAC
Lisa Sophie Albertelli
Lisa Sophie Albertelli
Parliamentary discourse constitutes a key domain in which political actors publicly articulate policy positions and priorities through language. This study investigates debates from the Italian Chamber of Deputies (1948–2006) to identify and analyse latent semantic themes and their evolution using BERTopic-based dynamic topic modeling. The analysis relies on a subset of the ItaParlCorpus (Cova, 2025), a large-scale, machine-readable corpus enriched with temporal, institutional, and political metadata. Beyond topic extraction,this work addresses a largely unexplored challenge: the formalization of topics derived from unsupervised, embedding-based topic modeling as Linked Data entities, adopting a linguistic perspective. Extracted topics are formalized as semantic entities reusing the OntoLex–Lemon model, its FrAC extension and declaring a dedicated ontology to link topics to speeches, speakers, political parties, and temporal information reusing standardized vocabularies and persistent URIs. This integration enables semantic querying through SPARQL, supporting analyses of topic distributions across political actors, parties and illustrating the analytical potential of the proposed approach. Moreover, the study highlights limitations in the formalization of topic modeling outputs, particularly regarding the representation of ambiguous word forms and their alignment with lexical concepts in OntoLex–Lemon.
Towards the LinkEn Knowledge Base. A Neuro-Symbolic Approach to Build a Linked Data Hub for English Lemmas with Large Language Models
Lorenzo Augello | Marco Passarotti
Lorenzo Augello | Marco Passarotti
This paper presents the first core component of LinkEn, a knowledge base of interoperable language resources for English adhering to Linked Open Data principles. With this initial step towards a broader infrastructure, we focus on the development of a lemma-centered hub designed to enable interoperability between distributed lexical resources, corpora, and linguistic annotations. The modeling is inspired by the LiLa Knowledge Base for Latin and the OntoLex-Lemon model, ensuring compatibility with existing lemma-centric knowledge graphs and enabling future cross-linguistic interoperability. Rather than relying solely on manual knowledge graph construction and significant human effort, the lemma bank has been developed through a hybrid neuro-symbolic pipeline that integrates large language models into the generation of RDF data under explicit ontological constraints. This approach combines automated generation with ontology-driven supervision and evaluation, enabling scalable yet controlled construction of structured lexical knowledge. By presenting the first steps towards the LinkEn Knowledge Base, this paper contributes both a new lemma bank for English and an experimental methodology for the semi-automatic creation of Linked Data based knowledge graphs.
Towards a Linguistic Linked Open Data Resource for Italian Cultural Heritage: The Lessico Dei Beni Culturali Corpus
Riccardo Billero
Riccardo Billero
We present an ongoing effort to bridge the Lessico dei Beni Culturali (LBC), a multilingual lexicographic project cov- ering Italian cultural heritage terminology, with the Linguistic Linked Open Data (LLOD) ecosystem. The LBC corpus spans five centuries of art-historical writing, from fifteenth- and sixteenth-century treatises by Alberti, Leonardo, and Vasari to nineteenth-century works by Stendhal and Burckhardt and contemporary tourist guides to Florence, with source texts in several European languages alongside their translations. The resource has already undergone automatic linguistic annotation and term extraction, but lacks structured lexical representation in any standard LLOD formalism. We describe the current state of the resource, identify the main challenges for its publication as Linked Data — including the modelling of culturally-bound terms (realia), historical proper nouns, and multilingual source texts of different registers — and outline a roadmap towards its representation in OntoLex-Lemon (McCrae et al., 2017) and its alignment with existing LLOD resources such as the Getty Vocabularies (Getty Research Institute, 2024a) and Wikidata (Vrandecic and Krötzsch, 2014). By sharing this work with the LLOD community, we expect input on best practices for historical-artistic and cultural heritage lexicons that will raise interoperability between resources from different sources, generating new information and increasing the value of existing data.
Consolidating Syntactically Annotated Corpora with LLOD Technology. An Experiment in the Old Saxon Heliand
Christian Chiarcos | Janine Siewert
Christian Chiarcos | Janine Siewert
The humanities are a vast and highly diverse field – both methodologically and technologically –, so, it is not unsurprising to see independent researchers or projects to work on the same data, and producing complementary, but technically incompatible electronic editions from the same source material. We suggest that existing Linguistic Linked Open Data (LLOD) technology can play a crucial role for performing a post-hoc consolidation of their efforts, illustrated for the Old Saxon (Old Low German) Heliand, a 9th c. gospel harmony previously annotated for different aspects of syntax in three independent research projects and over different versions (editions and manuscripts) of the original text. We describe the derivation of a UD-compliant corpus from the consolidation of the existing annotations. This includes the transformation of the original annotations to corpus-specific CoNLL (TSV) formats, the alignment between the different corpora, and their integration. A particular challenge is the processing of incomplete annotations, as one of the source corpora (Heliand B4) provides non-recursive nominal and clausal chunks only, and another corpus (Heliand DDD) even only sentence boundaries, clause types and parts of speech, but no actual phrasal structures. In this paper, we specifically focus on the application of Fintan (CoNLL-RDF) and SPARQL for performing the necessary graph rewriting operations.
Victim or Assailant? Exploring Narratives through Knowledge Graph Queries
Beatrice Fiumanò | Nicolas Lazzari | Simone Paolo Ponzetto | Valentina Presutti
Beatrice Fiumanò | Nicolas Lazzari | Simone Paolo Ponzetto | Valentina Presutti
Our understanding of social reality is shaped by the specific ways in which that reality is framed by different sources. Analyzing framing means examining how these sources are able to convey particular worldviews by foregrounding or downplaying certain aspects of experience. Current computational approaches address this task by automatically identifying communicative patterns (e.g., topic selection or rhetorical strategies) that characterize individual artifacts. However, they often remain document-bound, overlooking the comparative dimension that enables the uncovering of convergent or conflicting narratives about the same actor, event, or issue. In this paper, we propose DORIS, an ontology that supports both document-level and cross-document framing analysis using SPARQL queries on automatically constructed Knowledge Graphs. We validate the proposed approach through a case study of historical news articles, exploring multiple framings of a real-world event using Fillmore’s Frame Semantics and the FrameNet resource. Code and data are available on GitHub at https://github.com/beatrice-f/DORIS/.
This paper presents LaReS (Latin Represented Speech), a Linked Open Data resource designed to model represented speech in Latin literature and to align the DICES database of direct speeches in Greek and Latin epic with the LiLa Knowledge Base. While DICES provides a rich collection of metadata on direct speech in epic poetry, its operational approach and its relatively shallow conceptual modeling limit its interoperability and extensibility. The modeling strategy implemented in LaReS is based on the separation of the textual level from the narratological dimension. CIDOC CRM and DOLCE+DnS are used to conceptualize the basic notions in the two modules. LaReS now includes 341 speeches in Virgil’s Aeneid, linking 36,782 tokens in LiLa to speech units derived from DICES
We present Open English NameNet, a new large-scale lexical resource that extends Open English Wordnet with named entities derived from Wikidata. While English Wordnet has historically included many proper nouns, its coverage has been incomplete and inconsistent, and encyclopedic knowledge sources have grown rapidly in parallel. To address this gap, we systematically extract and align named entities from Wikidata with the Open English Wordnet hierarchy, ensuring each entity is appropriately placed through instance hypernym relations. Our methodology combines existing WordNet–Wikipedia mappings with Wikidata information and applies domain-specific strategies for people, plants and animals, and languages, to account for structural and semantic differences between the resources. This approach results in the largest English lexical-semantic resource currently available, with extensive coverage and structured integration. We release the resource openly to support the development of lexically and encyclopedically informed language technologies.
Bridging the Gap between Ontologies and Dictionaries: Requirements and Implementation of a New Core for OntoLex-Lemon
John P. McCrae | Jorge Gracia | Fahad Khan | Philipp Cimiano
John P. McCrae | Jorge Gracia | Fahad Khan | Philipp Cimiano
This paper presents the requirements and implementation details for a new core module of the OntoLex-Lemon model, representing the first major evolution of the de-facto standard since its 2016 release. While the original model successfully bridged ontologies and dictionaries through “semantics by reference,” community adoption has identified critical gaps in handling lexicographic structures and retrodigitized resources. We detail a community-driven methodology that identified fifteen key requirements and we present a proposed architecture for the new OntoLex core, which integrates elements from the Lexicography module and addresses both semantic web and lexicography use cases. Further, we improve interoperability with standards like DMLex and TEI-Lex0 while maintaining strict backwards compatibility for existing users of the model.
Linked Open Data for West Nilotic Languages: The NILOMORPH Project
Matteo Pellegrini | Matthew Baerman | Oliver Bond
Matteo Pellegrini | Matthew Baerman | Oliver Bond
In this paper, we present the NILOMORPH project, that aims at describing the complex non-concatenative morphology of West Nilotic languages and reconstructing the dynamics of its evolution from a more straightforward concatenative system. The project adopts techniques from several methodologies and draws on many kinds of data displaying different formats, tagsets and conventions. Data are also multilingual, documenting different West Nilotic varieties, and multimodal, including also audio and video recordings. This makes the process of integration of these data particularly challenging. We first describe how the data can be converted to standard formats such as CLDF and Paralex, to achieve interoperability between resources of the same kind. We then discuss how they can be modelled as Linguistic Linked Open Data in the Resource Description Framework, reusing already existing vocabularies and defining new classes and properties to meet the needs of the project, to also achieve interoperability between resources of different kinds.
A Linguistic Ontology for Constructicography: The Research Constructicon and its Ontology Modules
Elodie Winckel | Peter Uhrig | Stephanie Evert
Elodie Winckel | Peter Uhrig | Stephanie Evert
This paper introduces the Research Constructicon (RCxn), a project developed within the Research Training Group Dimensions of Constructional Space. The training group finances PhD projects in the framework of Construction Grammar (CxG), which views language as a network of form-meaning pairings. The RCxn is designed as a dynamic, community-driven resource that documents linguistic constructions while also capturing the research processes and findings associated with them. The project addresses three core dimensions: (1) the development of a modular ontology to represent constructions, their relationships, and the research surrounding them; (2) the implementation of database populated by researchers’ contributions; and (3) the creation of a web application to visualize and interact with the data. This paper focuses on our work to implement a rich ontology for the RCxn, which has to accommodate diverse research needs, from cross-linguistic comparisons to multimodal analyses, while ensuring flexibility and interoperability. We detail the modular design of the ontology, its alignment with semantic web standards (RDF/OWL), and the integration of existing ontologies (e.g., OLiA, FOAF). The RCxn’s development is iterative, driven by feedback from our diverse group of PhD researchers.
up
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Arturo Montejo-Raez | Cristina Grisot | Joanna Blochowiak | Nikola Ljubešić | Elena Battaner | German Rigau
Arturo Montejo-Raez | Cristina Grisot | Joanna Blochowiak | Nikola Ljubešić | Elena Battaner | German Rigau
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
Taja Kuzman Pungeršek | Peter Rupnik | Ivan Porupski | Vuk Dinić | Nikola Ljubešić
Taja Kuzman Pungeršek | Peter Rupnik | Ivan Porupski | Vuk Dinić | Nikola Ljubešić
Until recently, fine-tuned BERT-like models provided state-of-the-art performance on text classification tasks. With the rise of instruction-tuned decoder-only models, commonly known as large language models (LLMs), the field has increasingly moved toward zero-shot and few-shot prompting. However, the performance of LLMs on text classification, particularly on less-resourced languages, remains under-explored. In this paper, we evaluate the performance of current language models on text classification tasks across several South Slavic languages. We compare openly available fine-tuned BERT-like models with a selection of open-weight and closed-source LLMs across three tasks in three domains: sentiment classification in parliamentary speeches, topic classification in news articles and parliamentary speeches, and genre identification in web texts. Our results show that LLMs demonstrate strong zero-shot performance, often matching or surpassing fine-tuned BERT-like models. Moreover, when used in a zero-shot setup, LLMs perform comparably in South Slavic languages and English. However, we also point out key drawbacks of LLMs, including less predictable outputs, significantly slower inference, and higher computational costs. Due to these limitations, fine-tuned BERT-like models remain a more practical choice for large-scale automatic text annotation.
Exploring the Use of Large Language Models in Critical Discourse Analysis: A Consensus-Based Pilot Study
Emiliano Giovannetti | Francesca Cristiano
Emiliano Giovannetti | Francesca Cristiano
Large Language Models (LLMs) are increasingly used in the social sciences and humanities (SSH) to support the analysis of complex textual data, raising methodological questions about evaluation and interpretive reliability. This paper explores the use of LLMs in Critical Discourse Analysis (CDA), considered here as a paradigmatic case of interpretive research in SSH, through a preliminary consensus-based evaluation framework. The study reports on a pilot experiment conducted on a small, theory-driven corpus of opinion articles addressing the October 7, 2023 attack and its aftermath. An LLM is asked to answer analytically motivated questions targeting different levels of discourse structure. Its responses are compared with annotations produced by multiple human analysts and aggregated through a consensus-based procedure. The results reveal an asymmetry in model performance: while LLMs align well with human consensus on macro- and superstructural features, they struggle with microstructural phenomena involving implicit meaning. These findings support the view of LLMs as epistemic support tools rather than replacements for human interpretation.
LLM Evaluation in Practice: A Review of Metrics, Practitioner Insights, and Lessons Learned
Roos M. Bakker | Marianne Witte-Schaaphok | Julia García-Fernández | Tom Brand | Jens van der Weide | Stephan Raaijmakers
Roos M. Bakker | Marianne Witte-Schaaphok | Julia García-Fernández | Tom Brand | Jens van der Weide | Stephan Raaijmakers
The rapid, widespread adoption of Large Language Models (LLMs) highlights the need to understand their performance, strengths, and limitations. However, evaluating LLMs presents significant challenges due to the broad range of tasks and model capabilities, especially in practice or low-resource settings where benchmark datasets are not available. In text generation tasks, answer diversity has always complicated automatic evaluation, and the enhanced fluency and creativity of LLMs lead to further challenges. Existing metrics and frameworks often fail to account for these complexities. Furthermore, recent research into the replicability of benchmarks has demonstrated serious issues when reproducing historical benchmark results. This paper makes two key contributions: (1) a categorisation of challenges and metrics in LLM evaluation, and (2) lessons learned from practice through a survey and a use case. To this end, a literature study was conducted to identify challenges and metrics in scientific work. A survey among developers working with LLMs provided insights into practical challenges. Furthermore, selected metrics were implemented in a practical use case to gain insights into their strengths and limitations. By combining theoretical analysis with real-world experiences and lessons learned from practice, this work provides an overview and best practices for users evaluating LLM performance.
Is Human–LLM Interaction Culture-Dependent? A Cross-Linguistic NLP Analysis of Student Interviews on AI-Assisted Thesis Writing
Madalina Chitez | Karla Csuros | Dejana Jelena Milićević | Petya Osenova | Stefan Marinov | Teodor Valchev | Nikolay Paev | Otto Kruse | Christian Rapp | Andreea Dinca | Roxana Rogobete | Claudia Doroholschi | Loredana Punga | Anabella Costache | Dumitru Tucan | Cristina Baniceru
Madalina Chitez | Karla Csuros | Dejana Jelena Milićević | Petya Osenova | Stefan Marinov | Teodor Valchev | Nikolay Paev | Otto Kruse | Christian Rapp | Andreea Dinca | Roxana Rogobete | Claudia Doroholschi | Loredana Punga | Anabella Costache | Dumitru Tucan | Cristina Baniceru
This study investigates whether human–LLM interaction in academic writing exhibits cross-cultural variation. Using NLP-informed corpus methods, we analyze nine semi-structured student interviews from three national contexts (Romania, Bulgaria, Switzerland) to examine how AI use is linguistically constructed across three dimensions of epistemic positioning: agency strength, authority dynamics, and discourse-level stance. Results show a strong predominance of distancing and hedging strategies, with AI consistently framed as a functional writing support tool rather than an epistemic authority. At the same time, modest but systematic cross-country differences indicate culturally embedded variation in how students discursively negotiate epistemic responsibility and evaluation in AI-assisted writing practices.
Next Reply Prediction X (NRP-X) Dataset: Linguistic Discrepancies in Naively Generated Content
Simon Münker | Nils Schwager | Kai Kugler | Michael Heseltine | Achim Rettinger
Simon Münker | Nils Schwager | Kai Kugler | Michael Heseltine | Achim Rettinger
The increasing use of Large Language Models (LLMs) as proxies for human participants in social science research presents a promising, yet methodologically risky, paradigm shift. While LLMs offer scalability and cost-efficiency, their “naive” application, where they are prompted to generate content without explicit behavioral constraints, introduces significant linguistic discrepancies that challenge the validity of research findings. This paper addresses these limitations by introducing a novel, history-conditioned reply prediction task on authentic X (formerly Twitter) data, to create a dataset designed to evaluate the linguistic output of LLMs against human-generated content. We analyze these discrepancies using stylistic and content-based metrics, providing a quantitative framework for researchers to assess the quality and authenticity of synthetic data. Our findings highlight the need for more sophisticated prompting techniques and specialized datasets to ensure that LLM-generated content accurately reflects the complex linguistic patterns of human communication, thereby improving the validity of computational social science studies.
Quid est VERITAS? A Modular Framework for Archival Document Analysis
Leonardo Bassanini | Ludovico Biancardi | Alfio Ferrara | Andrea Gamberini | Sergio Picascia | Folco Vaglienti
Leonardo Bassanini | Ludovico Biancardi | Alfio Ferrara | Andrea Gamberini | Sergio Picascia | Folco Vaglienti
The digitisation of historical documents has traditionally been conceived as a process limited to character-level transcription, producing flat text that lacks the structural and semantic information necessary for substantive computational analysis. We present VERITAS (Vision-Enhanced Reading, Interpretation, and Transcription of Archival Sources), a modular, model-agnostic framework that reconceptualises digitisation as an integrated workflow encompassing transcription, layout analysis, and semantic enrichment. The pipeline is organised into four stages—Preprocessing, Extraction, Refinement, and Enrichment—and employs a schema-driven architecture that allows researchers to declaratively specify their extraction objectives. We evaluate VERITAS on the critical edition of Bernardino Corio’s Storia di Milano, a Renaissance chronicle of over 1,600 pages. Results demonstrate that the pipeline achieves a 67.6% relative reduction in word error rate compared to a commercial OCR baseline, with a threefold reduction in end-to-end processing time when accounting for manual correction. We further illustrate the downstream utility of the pipeline’s output by querying the transcribed corpus through a retrieval-augmented generation system, demonstrating its capacity to support historical inquiry.
Do We Still Need Corpora and Corpus Analysis Platforms? Discourse Analysis in Times of LLMs
Julia Krasselt | Dolores Lemmenmeier-Batinić | Philipp Dreesen
Julia Krasselt | Dolores Lemmenmeier-Batinić | Philipp Dreesen
Corpus-based discourse analysis investigates the linguistic construction of societally shared knowledge by iterating between quantitative pattern detection and qualitative interpretation in large text collections. Large Language Models (LLMs) promise to lower practical barriers to such work (e.g., natural-language querying, qualitative coding), yet they also introduce risks that are especially consequential in discourse-analytic settings, where fluent summaries can encourage ungrounded interpretation. This position paper argues that integrating LLMs into corpus analysis platforms is appropriate only insofar as it remains compatible with three epistemic premises of corpus research: (1) transparency of the data basis and traceability of analytical operations; (2) interpretability as evidence-constrained sense-making; and (3) seriality and patternedness as distributional structure and variation. In this opinion paper, we contribute a platform-oriented requirements perspective that translates these premises into design constraints for tool-calling/RAG-style integration, and we outline implementation directions that treat LLMs as an interaction layer over inspectable corpus retrieval and platform-based analysis.
GaelEval: Benchmarking LLM Performance for Scottish Gaelic
Peter Devine | William Lamb | Beatrice Alex | Ignatius Ezeani | Dawn Knight | Mícheál J. Ó Meachair | Paul Rayson | Martin Wynne
Peter Devine | William Lamb | Beatrice Alex | Ignatius Ezeani | Dawn Knight | Mícheál J. Ó Meachair | Paul Rayson | Martin Wynne
Multilingual large language models (LLMs) often exhibit emergent ‘shadow’ capabilities in languages without official support, yet their performance on these languages remains uneven and under-measured. This is particularly acute for morphosyntactically rich minority languages such as Scottish Gaelic, where translation benchmarks fail to capture structural competence. We introduce GaelEval, the first multi-dimensional benchmark for Gaelic, comprising: (i) an expert-authored morphosyntactic MCQA task; (ii) a culturally-grounded translation benchmark and (iii) a large-scale cultural knowledge Q&A task. Evaluating 19 LLMs against a fluent-speaker human baseline (n = 30), we find that Gemini 3 Pro Preview achieves 83.3% accuracy on the linguistic task, surpassing the human baseline (78.1%). Proprietary models consistently outperform open-weight systems, and in-language (Gaelic) prompting yields a small but stable advantage (+2.4pp). On the cultural task, leading models exceed 90% accuracy, though most systems perform worse under Gaelic prompting and absolute scores are inflated relative to the manual benchmark. Overall, GaelEval reveals that frontier models achieve above-human performance on several dimensions of Gaelic grammar, demonstrates the effect of Gaelic prompting and shows a consistent performance gap favouring proprietary over open-weight models.
Argumentation through Discourse Relations and Subjectivity: Introducing FreCaDiS, a French Multi-Genre Corpus
Joanna Blochowiak | Cristina Grisot
Joanna Blochowiak | Cristina Grisot
This paper addresses a crucial yet understudied issue in argumentation studies: the distinction between explanations and justifications, and their interaction with subjectivity. Building on insights from Bex and Walton (2016), who highlight the importance of not conflating explanations with arguments, we propose a corpus-based approach to operationalize this distinction in French. We present FreCaDiS (French Corpus of Causal Connectives, Discourse Relations, and Subjectivity), a novel corpus of French texts annotated for explanatory and justificatory discourse relations and their perceived subjectivity. FreCaDiS comprises excerpts of 2–3 sentences drawn from five distinct genres—SMS, online discussions, blogs, press, and contemporary literature—spanning informal to formal registers. Specifically, we focus on sentences introduced by the connectives parce que and car (“because”) and annotate them along two dimensions: (i) discourse relation (explanation vs. justification) and (ii) subjectivity (subjective vs. objective). The corpus was annotated by three independent human annotators using complementary approaches: a holistic, an intuitive method for subjectivity and a guided, operationalized method for discourse relations. FreCaDiS provides a rich resource for the study of argumentation, causal discourse, causal connectives, and subjective interpretation in French and can support future work in computational argument mining, discourse analysis, and NLP applications.
Text-only Domain Adaptation for Low-Resource ASR Using Large Language Models
William Lamb | Dongge Han | Ondrej Klejch | Peter Bell
William Lamb | Dongge Han | Ondrej Klejch | Peter Bell
Automatic Speech Recognition (ASR) increasingly mediates access to broadcast media, public discourse and cultural archives. For minoritised languages, however, the development of robust ASR systems is constrained by limited and domain-restricted text data. This paper investigates cross-lingual text expansion (XLTE), a method that uses a Large Language Model (LLM) to generate in-domain text in a low-resource language from high-resource language summaries. We further examine whether supervised fine-tuning on a small set of human-authored texts enhances generation quality. Using Scottish Gaelic as a case study, we show that synthetic text generated via fine-tuned XLTE can be used to train an external language model that reduces Word Error Rate (WER) by 24.48% in a previously unseen broadcast domain. Our findings demonstrate that text-only domain adaptation through cross-lingual generation can strengthen speech technology in sparse data settings. Beyond engineering gains, the approach offers a scalable pathway for improving the digital representation, accessibility and sustainability of minoritised-language media and cultural heritage.
Benchmarking LLMs for Aspect-Based Sentiment Classification in Slovene Historical Periodicals
Tina Munda | Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Darja Fišer
Tina Munda | Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Darja Fišer
Historical newspapers present substantial challenges for computational sentiment analysis due to OCR noise, archaic linguistic features, and the absence of domain-specific labeled training data. This paper examines whether instruction-following LLMs can support targeted, mention-level sentiment inference in such conditions. We benchmark four instruction-following LLMs on a manually annotated sample of collective-identity mentions drawn from Slovene historical newspapers. The results provide a benchmark for targeted sentiment classification in OCR-degraded historical Slovene and offer an empirically grounded assessment of the capabilities and limitations of an instruction-tuned LLM in digital humanities research.
Automatic Metrical Scansion of Poetry in a Low-Resource Setting
Pablo Ruiz Fabo | Anxo Alonso Pérez | Pablo Rodríguez Fernández | Pablo Gamallo
Pablo Ruiz Fabo | Anxo Alonso Pérez | Pablo Rodríguez Fernández | Pablo Gamallo
We present the first neural systems for automatic metrical scansion of poetry in Galician, a Romance language close to Portuguese and Spanish. The task is threefold: First, identifying metrical syllables based on lexical ones; both syllable series may differ given metrical licenses modifying a line’s syllable structure to enable stress-related rhythms. Second, identifying stress patterns, and third identifying the metrical syllable count, based on stressed positions. We manually annotated a corpus of 4,287 examples, a first in Galician, and fine-tuned an 8B-parameter LLM specialized in Galician and Portuguese, and two encoder–decoder models: ByT5, a token-free byte-to-byte model, and the multilingual mT5, which includes Galician. We also tested our recent symbolic scansion system. Several fine-tuning setups reached exact per-line accuracy above 90% on our test-set at all three scansion subtasks, using orthographic syllables with explicit stress marks as input. Encoder–decoders performed better than the LLM. The token-free ByT5 was best, particularly when adding the two surrounding lines to the input. The symbolic system (89.9% acc.) managed rare metaplasms infrequent in training data better than the neural ones, and the approaches can be seen as complementary.
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
Qinghao Guan | Yuchen Pan | Donghao Li | Zishi Zhang | Yiyang Chen | Lu Li | Flaminia Canu | Emilia Volkart | Gerold Schneider
Qinghao Guan | Yuchen Pan | Donghao Li | Zishi Zhang | Yiyang Chen | Lu Li | Flaminia Canu | Emilia Volkart | Gerold Schneider
In religion and theology studies, spirituality has garnered significant research attention for the reason that it not only transcends culture but offers unique experience to each individual. However, social scientists often rely on limited datasets, which are basically unavailable online. In this study, we collaborated with social scientists to develop a high-quality multimedia multi-modal datasets, SACRED, in which the faithfulness of classification is guaranteed. Using SACRED, we evaluated the performance of 13 popular LLMs as well as traditional rule-based and fine-tuned approaches. The result suggests DeepSeek-V3 model performs well in classifying such abstract concepts (i.e., 79.19% accuracy in the Quora test set), and the GPT-4o-mini model surpassed the other models in the vision tasks (63.99% F1 score). Purportedly, this is the first annotated multi-modal dataset from online spirituality communication. Our study also found a new type of connectedness which is valuable for communication science studies.
Integrating Knowledge Graphs and Multilingual Scholarly Corpora for Domain-Adaptive LLMs in SSH
Adam Faci | Alessio Miaschi | Anne Combe | Pascal Cuxac | Francesca Frontini | Nicolas Larrousse | Stéphane Pouyllau
Adam Faci | Alessio Miaschi | Anne Combe | Pascal Cuxac | Francesca Frontini | Nicolas Larrousse | Stéphane Pouyllau
The integration of Large Language Models (LLMs) into scientific research workflows, particularly for bibliographic discovery and literature synthesis, raises significant methodological, epistemic and regulatory challenges for the Social Sciences and Humanities (SSH), especially with regard to disciplinary diversity, multilingual access to sources and the evaluation of results. This paper presents an on-going use case developed within the European project LLMs4EU and the ALT-EDIC infrastructure, aimed at adapting foundation models to SSH research practices and supporting tasks such as question answering, comparative document analysis and literature review. The evaluation framework follows the LLMs4EU protocol and encompasses both independent quantitative benchmarking (retrieval, summarisation, traceability and hallucination detection) and a qualitative assessment involving a panel of Digital Humanities experts. By embedding model adaptation within research infrastructures and a structured legal and ethical compliance framework, the use case explores how domain-sensitive and regulation-aware generative AI can support SSH scholarship while preserving reliability and epistemic responsibility.
From One-Hot to Semantic Encoding: Entity Embedding for Small and Heterogeneous Digital Humanities Datasets
Isabelle Gribomont
Isabelle Gribomont
This paper investigates the use of semantic encoding for the analysis of heterogeneous digital literature metadata. Drawing on two databases of Latin American digital literature, Archivo de Literatura Digital en América Latina and the Atlas da Literatura Digital Brasileira, we compare traditional one-hot encoding with a semantically enriched representation derived from feature-value descriptions embedded in a continuous vector space. In contrast to one-hot encoding, which treats categorical values as orthogonal, semantic encoding models accounts for similarity between features, thereby mitigating vocabulary mismatch across databases. We evaluate both approaches using between-group centroid distances, and normalized centrality measures. Our results show that semantic encoding clarifies structural differentiation across genres and might smooth arbitrary differences introduced by differing vocabularies across databases. The findings suggest that semantic representations provide a more interpretable embedding space for small and taxonomically heterogeneous datasets. Beyond technical performance, the study suggests that embedding-based methods can support critical inquiry in digital humanities, enabling the examination of database bias, categorical patterns, and diachronic evolution within a unified semantic framework. Code is available at https://github.com/isag91/semantic-encoding-DH.
Design and Methodological Architecture of a Multilingual Corpus of Interpreter-mediated Public Service Telephone Interactions
Raquel Lazaro Gutierrez | Daniel López Padilla | Jorge Rico | María José Vilella Sánchez | Fernando Manuel Espinoza-Cuadros
Raquel Lazaro Gutierrez | Daniel López Padilla | Jorge Rico | María José Vilella Sánchez | Fernando Manuel Espinoza-Cuadros
Multimodality in Social Sciences and Humanities (SSH) research is often associated with the integration of text and visual data. However, interpreter-mediated telephone interaction presents a different configuration of complexity, where acoustic, temporal, discursive, and pragmatic dimensions converge. This paper presents the design and methodological architecture of PRAGMACOR(Corpus Pragmatics and Telephone Interpreting: Analysis of Face-Threatening Acts, Ref. PID2021-127196NA-I00), a multilingual corpus of interpreter-mediated public service telephone interactions (Chinese–Spanish, English–Spanish, French–Spanish, German–Spanish), as a case study in multimodal and plurilingual SSH infrastructure. The corpus integrates aligned audio recordings, orthographic transcriptions enriched with speech phenomena, temporal segmentation into speech acts, and multilayer pragmatic annotation of Face-Threatening Acts (FTAs), validated through a structured double-annotation and expert review process. Beyond textual data, the infrastructure captures prosodic overlap, turn-taking dynamics, and pragmatic mediation, enabling the study of cross-linguistic transfer and relational negotiation in asymmetrical institutional contexts. Datasets such as PRAGMACOR have proved essential to train LLMs for speech to speech translation (Sakai et al., 2024). Attention is given to the ethical and technical design of the corpus, including local automatic transcription, systematic removal of personal identifiable information, and irreversible voice anonymization through spectral and temporal signal transformation. These procedures ensure both research usability and compliance with responsible data governance principles. By conceptualising interpreter-mediated interaction as an acoustic-discursive multimodal object and plurilingual pragmatic process, this paper argues that PRAGMACOR provides a replicable model for the development of SSH-oriented infrastructures capable of supporting advanced research in multilingual communication, discourse analysis, and future evaluation of language technologies.
Toward Responsible and Epistemically Grounded Multilingual LLMs for Computational Social Science and Humanities
Wajdi Zaghouani
Wajdi Zaghouani
Large language models have rapidly evolved in multilingual competence and reasoning capacity, enabling their integration into Social Sciences and Humanities research workflows. Yet existing evaluation paradigms remain anchored in task-based NLP benchmarks and fail to address interpretive validity, cultural situatedness, and epistemic mediation. This paper reconceptualizes multilingual reasoning LLMs as hermeneutic instruments that actively structure meaning production across linguistic and cultural contexts. Drawing on hermeneutics, philosophy of technology, science and technology studies, multilingual NLP research, and computational social science methodology, we develop a theoretically grounded framework for evaluating multilingual reasoning in Social Sciences and Humanities (SSH) research. We articulate a rigorous experimental protocol with operationalized metrics for cultural alignment, cross-lingual stability, and reasoning faithfulness, along with transparency requirements tailored to interpretive research tasks. The paper contributes a conceptual and methodological foundation for responsible integration of multilingual reasoning LLMs into computational social science infrastructures.
Automatic Evaluation of Multiple-Choice Items for Reading Comprehension: Effects of Question and Distractor Categories
John S. Y. Lee | Yin Poon | Shunjie Wang | Kai Wah Chu
John S. Y. Lee | Yin Poon | Shunjie Wang | Kai Wah Chu
Automatic generation of multiple-choice (MC) items for reading comprehension can support language learning by providing large amounts of practice materials. To enable rapid development of MC generation models, automatic assessment is essential since it is time-consuming to manually evaluate question and distractor quality. Although Text Informativity (TI) has been adopted as an automatic evaluation metric, the ability of Large Language Models (LLMs) to estimate the TI scores of different categories of questions and distractors has not yet been thoroughly analyzed. This paper investigates LLM performance in calculating TI scores for the range of questions and distractors defined in the PIRLS (Progress in International Reading Literacy Study) and STARC (Structured Annotations for Reading Comprehension) frameworks. We show that automatically estimated TI scores may result in systematic preferences for some question and distractor categories, and recommend that TI scores be used for within-category comparisons only.
Reflexive Research with LLMs: Considering the Positionality of Users and Systems
Eleanor L.T. Smith | Luis Morgado da Costa | Antske Fokkens
Eleanor L.T. Smith | Luis Morgado da Costa | Antske Fokkens
Previous work has found that people often perceive computational systems as neutral tools (van Es, 2023), and yet these systems are not developed or deployed within a vacuum. As the popularity of Large Language Models (LLMs) in digital social science and humanities (DSSH) research increases, it is important that we reflect both on our positionality as researchers regarding how we are primed to interact with these systems and the positionality of the systems themselves as defined by their design and training. This paper presents a model of factors and interactions affecting the use of LLMs in DSSH research and argues that explicit discussion of both human biases, which affect how we interact with systems, and the potential biases encoded in systems are needed in conjunction with strong case specific system evaluation when developing methodologically sound applications of LLMs.
Small Can Be Beautiful in LLMs for SSH: a Case for Bulgarian
Kiril Simov | Nikolay Paev | Petya Osenova | Teodor Valchev | Stefan Marinov
Kiril Simov | Nikolay Paev | Petya Osenova | Teodor Valchev | Stefan Marinov
In the paper we present a set of small LLM-based models for solving the basic NLP tasks for Bulgarian - POS tagging, Lemmatization, Dependency parsing, Named Entity Recognition, Named Entity Linking, Event Annotation, among others. In order to create fine-tuned models for these tasks, we first pre-train models using architectures like BERT, Modern-BERT, and T5 with different sizes, over Bulgarian data only. For each of the tasks we report our approach towards the fine-tuning, the results from the experiments and also the evaluation. Then we define a way to visualize the results over HTML documents which contain the analyzed texts. Our rationale are as follows: most, if not all SSH research scenarios, need a reliable processing chains that can be customized with respect to the specific needs. These scenarios also would need proper visualization for human observation. We aim to provide such a basic LLM-based toolkit.
A Multimodal LLM-Based Nutrition Label for Analyzing Social Media Feed Exposure
Tim Gollub | Armin Heidari | Cem Ertürkan | Benno Stein
Tim Gollub | Armin Heidari | Cem Ertürkan | Benno Stein
Algorithmically curated social media feeds shape political exposure, commercial influence, and cultural consumption, yet they remain difficult to study systematically due to limited data access and opaque recommendation mechanisms. We present a research-oriented framework that operationalizes feed-level exposure analysis using a browser extension combined with a server-side multimodal large language model (LLM). The system logs visible posts and their view time, performs zero-shot multimodal classification, and aggregates results into a customizable nutrition label summarizing exposure across analytical categories. It further supports retrieval-grounded conversational querying, dataset export and sharing, and human validation of LLM classifications. Designed as a methodological instrument for Social Sciences and Humanities, the framework enables both observational analysis and experimental research on transparency interventions, while critically examining epistemic, methodological, and ethical implications of LLM-based exposure analysis.
Charting the European LLM Benchmarking Landscape: A New Taxonomy and Registry
Spela Vintar | Mojca Brglez | Taja Kuzman Pungeršek | Nikola Ljubešić
Spela Vintar | Mojca Brglez | Taja Kuzman Pungeršek | Nikola Ljubešić
While new benchmarks for large language models (LLMs) are being developed continuously to catch up with the growing capabilities of new models and AI in general, using and evaluating LLMs in non-English languages remains a poorly-charted landscape. We give a concise overview of recent developments in LLM benchmarking, and then propose a new taxonomy for the categorization of benchmarks that is tailored to multilingual or non-English use scenarios. We further propose a registry of benchmarks implementing the new categorization and documenting benchmarks with a rich set of metadescriptors. While still at a pilot stage, such a registry can lead to a more coordinated development of benchmarks for European languages. We conclude with a review of current trends and advocate for a higher language and culture sensitivity of evaluation methods.
Cross-Lingual Abstractive Keyphrase Generation for Historical Newspapers
Simon Clematide | Jenifer L. Meyer | Juri Opitz | Maud Ehrmann | Kaspar Beelen
Simon Clematide | Jenifer L. Meyer | Juri Opitz | Maud Ehrmann | Kaspar Beelen
We investigate large language models (LLMs) for cross-lingual abstractive keyphrase generation from historical newspapers. The task consists of producing a small set of English keyphrases for articles written in German, French, and Luxembourgish, combining translation, abstraction, and normalization. We conduct a human-centered pilot study comparing model outputs using human selections, LLM-as-judge assessments, and inter-annotator agreement analysis, followed by a medium-scale application to multilingual data from the Impresso corpus. Results show that LLM-generated keyphrases can support semantic enrichment and exploratory analysis of historical collections, while highlighting the subjective and methodologically challenging nature of keyphrase evaluation.
LLM Probe: Evaluating LLMs for Low-Resource Languages
Hailay Kidu Teklehaymanot | Wolfgang Nejdl | Gebrearegawi Gebremariam
Hailay Kidu Teklehaymanot | Wolfgang Nejdl | Gebrearegawi Gebremariam
Despite the rapid progress of large language models (LLMs), their linguistic capabilities in low-resource and morphologically rich languages remain insufficiently understood due to the scarcity of annotated resources and the lack of standardised evaluation frameworks. This paper introduces LLM Probe, a lexicon-based evaluation framework for systematically assessing the linguistic competence of LLMs in low-resource language settings. The framework evaluates models across four dimensions of language understanding: lexical alignment, part-of-speech identification, morphosyntactic probing, and translation fidelity. To demonstrate the framework, we construct a manually annotated benchmark dataset using a low-resource Semitic language as a case study. The dataset consists of bilingual lexicons enriched with linguistic annotations, including part-of-speech categories, grammatical gender, and morphosyntactic features, with high inter-annotator agreement ensuring annotation reliability. We evaluate a diverse set of models spanning causal language models and sequence-to-sequence architectures. The results reveal substantial variation in model performance across linguistic tasks: sequence-to-sequence models generally achieve stronger performance in morphosyntactic analysis and translation quality, while causal models exhibit competitive lexical alignment but weaker translation fidelity. Our findings highlight the importance of linguistically grounded evaluation for understanding the limitations of LLMs in low-resource contexts. We release LLM Probe and the accompanying benchmark dataset as open-source resources to support reproducible benchmarking and to advance the development of more inclusive multilingual language technologies.
up
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Rachele Sprugnoli | Marco Passarotti
Rachele Sprugnoli | Marco Passarotti
We report on the morphological tagging of Old Serbian in the Universal Dependencies framework. To facilitate the manual annotation, we pre-processed the data with the Old Church Slavonic 2.12 UDPipe model. The decision was based on the known similarity of these two languages as well as on the declared performance of this model compared to other models for historical varieties of Slavic languages. With over 3,000 manually annotated tokens, we evaluated the performance of the relevant pre-trained UDPipe2 models of historical Slavic languages. Besides, we also trained and evaluated custom models with UDPipe1 containing the annotated Old Serbian data. We have found that: (1) for this particular domain and amount of training data, the most suitable model is UD Old East Slavic – Birchbark 2.12, although its declared performance is much lower than that of Old Church Slavonic; (2) even 3,000 tokens of Old Serbian increase the performance of UDPipe1 models almost to the level of the Birchbark 2.12 model. The dataset is publicly available at https://doi.org/10.5281/zenodo.19317842.
Tracing Morph Origins in Czech: A Computational Approach to Morph-Level Etymology
Aleš Manuel Manuel Papáček | Zdeněk Žabokrtský
Aleš Manuel Manuel Papáček | Zdeněk Žabokrtský
Modern languages remain connected to ancient ones in multiple ways, including through etymology; for instance, Latin is among the most influential sources of borrowings in (modern) Czech, whether transmitted directly or mediated through other languages. This work focuses on predicting the etymological origin of individual morphs in Czech words. Given morphologically segmented Czech sentences, the task is to determine for each morph whether it is native or borrowed, and if borrowed, to identify the languages through which it entered Czech. Although some linguists have examined etymology at the level of individual morphs rather than whole words (Arkadiev et al., 2015), to our knowledge, no computational work has yet addressed this level of analysis. We created a manually annotated dataset of 300 Czech sentences comprising around 10,000 morphs with morph-level etymology labels, and trained supervised models using character-based and structural features. Our best lightweight system is a feed-forward neural network with a single hidden layer, trained on data augmented with entries from an etymological dictionary, reaching 96.2% F1 on the test set. We also developed and tested several prompting variants for large language models; the best model Claude-Opus-4.5, achieved 97.8% F1. We release the code, prompts, and dataset as open source at https://github.com/ampapacek/MorphemeOrigin.
Uncovering Work from Words: LLM-Based Information Extraction from Historical Petitions
Ellinor Lindqvist | Eva Pettersson | Joakim Nivre
Ellinor Lindqvist | Eva Pettersson | Joakim Nivre
We investigate the extraction and normalisation of phrases describing work from 18th-century Swedish petitions using four LLMs: GPT-4o, Llama-3 70B/8B, and Mixtral-8x7B. Performance is evaluated across four configurations: isolated extraction, isolated normalisation, a staged pipeline, and a combined multitasking setup, using both full and filtered texts (with formal greetings and closing sections removed). While exact phrase matching remains low (F1 < .10), token-level and semantic similarity scores suggest that models consistently locate relevant topical regions. Semantic similarity scores must however be interpreted with caution, since they are often only marginally higher than an average baseline. Results reveal a “multitasking paradox”: combined extraction and normalisation improves phrase location for high-parameter models but degrades normalisation precision. Furthermore, normalisation benefits from the context of a staged pipeline compared to isolated tasks, while text filtering has only marginal effects. Despite a tendency towards over-prediction, qualitative analysis suggests that models can detect plausible work-related expressions missed by human annotators. These findings illustrate the challenges of historical extraction and suggest that hybrid human–machine workflows are a promising approach for enhancing coverage and interpretability in cultural heritage research.
Extracting Volcanological Knowledge from Historical Texts: A Language-Technology Pipeline for Diachronic Geovisualization
Costanza Marini | Gianluca Casagrande | Alessio Palmero Aprosio | Claudia Principe
Costanza Marini | Gianluca Casagrande | Alessio Palmero Aprosio | Claudia Principe
This paper presents the first results of the CorVo project, a transdisciplinary project combining volcanology and computational linguistics to extract and structure volcanological knowledge from historical documents concerning Mount Vesuvius. We introduce the CorVo corpus, a multilingual diachronic corpus of 180 digitized texts (16th–20th centuries), selected to represent the main eruptive scenarios of the volcano. The digitization workflow integrates image pre-processing, OCR, and LLM-based post-correction to address challenges posed by degraded pages, historical typefaces, and orthographic variation. A domain-aware information extraction pipeline was developed to identify both standard toponyms and fine-grained spatial entities, which are typically overlooked by traditional NER systems. Extracted entities undergo human-in-the-loop validation and georeferencing through a dedicated annotation interface supporting multiple spatial geometries. The resulting dataset enables temporally normalized diachronic geovisualization of the textual-spatial footprint of Vesuvian eruptions across centuries.
When Lexicographic Quotations Become a Corpus: To Deduplicate or Not to Deduplicate?
Manuel Favaro | Elisa Guadagnini | Eva Sassolini | Marco Biffi | Simonetta Montemagni
Manuel Favaro | Elisa Guadagnini | Eva Sassolini | Marco Biffi | Simonetta Montemagni
Historical dictionaries are increasingly reused as sources for diachronic language corpora. In this context, lexicographic quotations represent a valuable yet challenging type of data, as they are both editorially curated and diachronically representative. A major issue in their computational reuse is the presence of duplicate and near-duplicate quotations. This paper addresses quotation deduplication in corpora derived from lexicographic resources. We introduce QRD (Quotation Reuse Detection), a multi-stage pipeline designed to identify, compare, and cluster quotations based on graded similarity rather than binary matching. The approach combines string-based similarity measures, iterative threshold analysis, and clustering, enabling both quantitative and qualitative investigation of quotation reuse. Our results show that deduplication in this context cannot be reduced to the automatic elimination of redundant data. The variability observed in the quotations - ranging from OCR-related noise to substantial editorial variation - reflects both technical and structural factors and calls for a more nuanced approach. QRD supports the identification of OCR-related errors and reveals patterns of textual reuse underlying the compilation of the dictionary. We argue that quotation deduplication should be conceived primarily as a task of identification and clustering. This perspective reframes deduplication from a data-cleaning operation into an analytical methodology for historically and editorially curated textual resources.
A New State-of-the-Art BERT Model for Judeo-Arabic
Elisha Rosensweig | Yitzchak Lindenbaum | Hillel Gershuni | Vered Raziel-Kretzmer | Daniel Caine | Avi Shmidman
Elisha Rosensweig | Yitzchak Lindenbaum | Hillel Gershuni | Vered Raziel-Kretzmer | Daniel Caine | Avi Shmidman
We present JABERT, the first BERT model pretrained specifically for historical Judeo-Arabic texts. We demonstrate that JABERT outperforms Arabic and multilingual models on the downstream task of Judeo-Arabic homograph disambiguation. Furthermore, in order to test the latter, we have curated and annotated the first Judeo-Arabic homograph test set. We release both JABERT and the Judeo-Arabic homograph test to the public for unrestricted use.
BEReshiT: an Ancient Hebrew Model based on DictaBERT
Iglika Nikolova-Stoupak | Maxime Amblard | Frédérique Rey
Iglika Nikolova-Stoupak | Maxime Amblard | Frédérique Rey
This project addresses the general absence of Natural Language Processing (NLP) tools when it comes to historical languages as a subset of low-resource languages that is relevant to an array of academic disciplines from linguistics to textual criticism. In particular, we train an Ancient Hebrew language model, BEReshiT, as well as BEReshiT-morph, a submodel for morphological annotation. BEReshiT is achieved through the fine-tuning of DictaBERT, a state-of-the-art model for Modern Hebrew that has also proved useful in Biblical Hebrew tasks. Layer freezing is applied in order to achieve maximal results and gain insight about the adaptation process. In the context of an elaborate cloze test, BEReshiT demonstrates increased performance and notions of the Ancient Hebrew language compared to the source model as well as a selection of additional relevant models. The submodel BEReshiT-morph performs highly on tasks of morphological classification, reaching an F1 score of 0.97 for part-of-speech (POS) tagging. We will release the main and morphological models as well as the datasets used at training and evaluation.
Automatic Detection of Metaphorical Expressions in Classical Japanese Using WLSP-Enhanced BERT
Hang Zhu | Mitoki Ohara | Rei Kikuchi | Kanako Komiya | Masayuki Asahara | Sachi Kato
Hang Zhu | Mitoki Ohara | Rei Kikuchi | Kanako Komiya | Masayuki Asahara | Sachi Kato
Metaphor detection is a fundamental task in natural language processing, yet research on historical languages remains limited. While progress has been made in modern Japanese metaphor detection, classical Japanese texts present unique challenges due to their distinct vocabulary, grammar, and metaphorical patterns. This paper addresses this gap by applying a BERT-based metaphor detection method enhanced with semantic classification information from the Word List by Semantic Principles (WLSP) to classical Japanese texts. We evaluate our approach on CHJ-Metaphor, a newly available corpus featuring metaphor annotations for three medieval Japanese works from the Corpus of Historical Japanese (CHJ). Our method achieves an F1-score of 82.18 through 5-fold cross-validation. Notably, qualitative analysis by domain experts reveals that our model successfully identifies genuine metaphors overlooked during manual annotation, demonstrating its potential as a tool for improving annotation quality in large-scale corpus construction. These results confirm the effectiveness of WLSP-enhanced approaches for metaphor detection in classical Japanese and suggest promising directions for applying similar techniques to other historical languages.
Domain-Aware Error Correction for Citation NER in Medieval Hebrew Responsa
Shmuel LIebeskind | Maayan Zhitomirsky-Geffet | Binyamin Katzoff | Nati Ben-Gigi | Jonathan Schler
Shmuel LIebeskind | Maayan Zhitomirsky-Geffet | Binyamin Katzoff | Nati Ben-Gigi | Jonathan Schler
Citation identification in historical and ancient texts poses challenges that extend beyond surface-level pattern recognition, including implicit references, morphological fusion, and discourse-driven ambiguity. In this work, we address citation Named Entity Recognition (NER) in medieval Hebrew Responsa literature using a modular, LLM-based correction pipeline. Rather than treating large language models as end-to-end predictors, we leverage them as structured components: an initial prompt-based expert tagger, complementary LLM judges for systematic error detection, and domain-aware correction grounded in philological regularities. Our approach requires no end-to-end fine-tuning and only minimal labeled supervision (a small validation set for training a lightweight error-detection classifier), narrowing the performance gap to strong supervised models trained on domain-specific data. The results suggest that explicit error handling and interpretability-driven design offer a promising direction for historical NLP in low-resource settings.
From Lemmas to Links: A Lemma Bank for Ancient Greek
Colin Swaelens | Francesco Mambrini | Marco Passarotti
Colin Swaelens | Francesco Mambrini | Marco Passarotti
This paper introduces the Greek Lemmabank, the core component of the Linking Greek knowledge base, developed according to Linked Open Data principles. Addressing the fragmentation of existing Ancient Greek resources, the lemmabank adopts a descriptive, ontology-driven approach inspired by LiLa (Linking Latin). Lemmas are modelled as canonical forms within an OntoLex-Lemon–compliant framework, preserving alternative canonical solutions and dialectal variation while enabling interoperable linking across heterogeneous datasets. The resource is populated by integrating data from the Ancient Greek WordNet and the Liddell–Scott–Jones lexicon, with additional normalisation and harmonisation to the Universal Dependencies tagset. The resulting dataset establishes a lemma-centric infrastructure for interlinking corpora, lexica, and NLP tools for Ancient Greek.
Across Generations: A Comparative Analysis of NER for Latin Inscriptions from Classical Machine Learning to LLMs
Wenhui Cui | Phillip Benjamin Ströbel
Wenhui Cui | Phillip Benjamin Ströbel
Latin epigraphic texts are a challenging type of historical data for natural language processing (NLP). They are often fragmentary, contain inconsistent spelling, and follow complex Roman naming conventions. This paper investigates Named Entity Recognition (NER) for this domain by comparing several approaches, including feature-based Support Vector Machines, neural models such as BiLSTM and TreeLSTM, pre-trained language models like LatinBERT, fine-tuned Transformer models based on BERT, and large language models used with prompting and supervised fine-tuning. We introduce a manually annotated dataset of 1,000 inscriptions from the Epigraphik-Datenbank Clauss-Slaby, labelled with a fine-grained BIO scheme that captures the internal structure of Roman personal names. Results show that the fine-tuned BERT model achieves the highest performance, with a weighted F1 score of 91.1% and a macro F1 of 68.7%, and clearly outperforms other methods. Additional linguistic features, such as part-of-speech tags and dependency information, yield only limited improvements, likely due to the irregular nature of inscriptional texts. This work provides a new benchmark for NER on Latin inscriptions and offers practical insights into applying modern NLP techniques to historical, non-standardised language.
POS Tagging with Generative LLMs for Historical Germanic Low-Resource Languages: An Evaluation Against Fine-Tuned BERT
Irene Miani | Gregory Darwin | Sara Stymne
Irene Miani | Gregory Darwin | Sara Stymne
Part-of-Speech (POS) tagging is a fundamental task in Natural Language Processing, yet its performance on historical low-resource languages is still underexplored, particularly in the context of large generative models. While recent studies have demonstrated strong results for Large Language Models (LLMs) on modern languages and contemporary low-resource settings, their effectiveness for historical varieties remains unclear. Moreover, genre-specific structural variation, which may substantially affect tagging performance, has received limited attention. This study evaluates the zero- and few-shot POS tagging performance of two generative models on four historical Germanic low-resource languages across two literary genres. Their performance is benchmarked against fine-tuned BERT models. To contextualize the performance on historical data, the models are also evaluated on two modern languages. The results show that fine-tuned encoder models consistently outperform generative models across all settings. The performance of the LLMs on historical languages is substantially lower compared to that on modern languages, suggesting limited representation of these varieties in pretraining data. Furthermore, error analysis reveals structural output inconsistencies in LLM predictions that require additional post-processing. These findings highlight the limitations of zero- and few-shot generative models for historical low-resource POS tagging and underline the importance of task-specific fine-tuning.
From Manuscript to Model: Developing HTR for Medieval Greek
Nicklas Sindlev Andersen | Byron MacDougall | Tariq Yousef | Aglae Pizzone
Nicklas Sindlev Andersen | Byron MacDougall | Tariq Yousef | Aglae Pizzone
We develop and evaluate manuscript-specific text line detection (TLD) and handwritten text recognition (HTR) models for two 14th-century Medieval Greek manuscripts, Vat. gr. 2228 and Phil. gr. 130, comprising 1,356 handwritten pages. From these, we curate and document 36 pages with complete, manually curated text line annotations, together with 10 additional pages with layout annotations only for TLD, forming two manuscript-specific ground truth (GT) datasets. To ensure representative evaluation despite limited annotations, validation splits are optimized for character coverage and distributional similarity using Jensen-Shannon divergence. Using the Transkribus platform, we train manuscript-specific TLD models from scratch and manuscript-specific HTR models, comparing HTR training from scratch with fine-tuning of a publicly available Medieval Greek base model. TLD achieves validation pixel-wise misclassification rates of 5.42% for Vat. gr. 2228 and 8.76% for the more layout-variable Phil. gr. 130. For HTR, fine-tuning consistently outperforms training from scratch. On validation pages with manually curated text line annotations, Vat. gr. 2228 reaches 5.13% character error rate (CER) and 23.66% word error rate (WER), while Phil. gr. 130 reaches 27.13% CER and 65.72% WER after continued training. A supplementary held-out evaluation on Vat. gr. 2228 shows that the fine-tuned model reaches 5.97% CER and 23.52% WER on test pages with manually corrected line polygons, degrading to 12.12% CER and 44.21% WER under automatic TLD-based segmentation. The study also provides a reproducible workflow and evaluation protocol for Medieval Greek HTR under low-resource conditions.
I, RE:Claudius 256: Towards Linking Classical Latin Person Mentions to a Domain-specific Knowledge Base
Marijke Beersmans | Evelien de Graaf | Julie Nijs | Valeria Irene Boano | Alek Keersmaekers | Mark Depauw | Tim Van de Cruys | Margherita Fantoli
Marijke Beersmans | Evelien de Graaf | Julie Nijs | Valeria Irene Boano | Alek Keersmaekers | Mark Depauw | Tim Van de Cruys | Margherita Fantoli
This paper considers Named Entity Linking for person mentions from classical Latin texts to a domain-specific, German language knowledge base, namely Paulys RealencyclopΣdie. Following a methodology similar to (anonymous_reference), we train a transformer-based, retrieval and ranking model (BLINK) first on a general, Wikipedia-derived dataset and subsequently on a more specific dataset, gathered from various sources, linking to our target knowledge base. Results show that while BLINK performs well on mention-entity pairs linked to entities seen during training, it performs significantly worse on mention-entity pairs linking to unseen entities. We provide a detailed error analysis, propose possible exploitation strategies for a human-in-the-loop approach, and identify directions for future improvement.
Capturing Ancient Chinese Sense Induction with Automatic Pipelines
Guan-Yu Tseng | Chunki Lim | Chih-Han Lin | Tung-Le Pan | Yu-Chieh Wang | Lang-Ching Yeh | Shu-Kai Hsieh
Guan-Yu Tseng | Chunki Lim | Chih-Han Lin | Tung-Le Pan | Yu-Chieh Wang | Lang-Ching Yeh | Shu-Kai Hsieh
While the study of diachronic semantic change has advanced alongside recent computational developments, structured lexical resources that reflect semantic evolution remain scarce for many languages, including Ancient Chinese. By systematizing the diachronic transformations within the Chinese Text Project (ctext, a large corpus of Ancient Chinese), we aim to bridge the gap between traditional philological inquiry and contemporary computational linguistics. This study proposes a pipeline that extracts contextualized embeddings from GujiBERT-fan, a language model pre-trained on pre-modern Chinese, and applies dynamic hierarchical clustering to identify distinct senses across historical periods. The pipeline operates at two levels: a global clustering that aggregates data across all periods to capture the full semantic space, and local clustering within each dynasty to reveal period-specific usage patterns. We test the pipeline with a pilot study on the character 手 (shǒu, “hand”) across eight dynastic periods, covering over 185,000 occurrences. The results show that the pipeline can capture the diachronic shift from concrete to abstract senses, demonstrating its potential as a scalable method for mapping semantic evolution in historical languages.
A Computational Evaluation of Syllabic Hypotheses for Rongorongo: Evidence from N-gram Analysis
Evgeniya Korovina
Evgeniya Korovina
The study evaluates the hypothesis that the Rongorongo script of Easter Island functions as a syllabic substitution cipher where one symbol uniquely corresponds to one syllable. Using a genetic algorithm with a fitness function based on Rapa Nui n-gram statistics, we establish a performance baseline on controlled texts. Results show a strong correlation (0.75) between the algorithm’s fitness score and decipherment accuracy, identifying a “noise threshold” at 2,500,000 points. We further demonstrate that cross-corpus genre variance significantly impacts recovery rates, with accuracy dropping by more than half when mismatched linguistic statistics are applied. Application to the CEIPP transliteration yields scores well below the noise threshold (1.0M–1.3M), suggesting a lack of a simple syllabic signal. However, testing the rongopy transliteration produces scores in the “gray zone” (up to 2.7M) and reveals stable mappings for high-frequency glyphs 200 and 006. The consistency of these results across independent inscriptions suggests that while a pure syllabic model is insufficient, specific structural simplifications of the script may capture latent linguistic patterns.
Building a Corpus and Database for Rare and Undeciphered Scripts
Beata Megyesi | Rune Rattenborg | Benedek Láng | Michelle Waldispühl | Mihály Héder
Beata Megyesi | Rune Rattenborg | Benedek Láng | Michelle Waldispühl | Mihály Héder
Historical sources written in rare or undeciphered scripts represent an immense but underexploited part of the world’s cultural and linguistic heritage. Their study is often hindered by fragmentary preservation, non-standard symbol systems, and the absence of interoperable digital resources. While recent advances in imaging, transcription, and computational analysis have improved access to historical texts, most tools rely on large quantities of labeled data and standardized encodings, requirements that are rarely met for rare or unknown writing systems. This paper presents the design and methodology of a new corpus and database dedicated to rare and undeciphered scripts worldwide. The resource integrates high-quality images, transliterations, transcriptions, linguistic annotations, and metadata within a unified data model tailored for low-resource and non-standard scripts. By adhering to FAIR principles and existing standards for linguistic and cultural heritage data, the database enables reproducible, interdisciplinary research across philology, linguistics, cryptology, and computer science. The paper outlines the data collection and digitization workflow, describes the metadata and database architecture, and demonstrates applications in analysis and decipherment.
A Layered Annotation Workflow for Semitic Epigraphy
Tal Bernstein | Shai Gordin | Letizia Cerqueglini
Tal Bernstein | Shai Gordin | Letizia Cerqueglini
This paper presents a layered annotation workflow for the historical linguistic study of Semitic epigraphic texts. Using a curated Phoenician corpus primarily based on Kanaanäische und Aramäische Inschriften (KAI), the system models inscriptions as multi-layered objects that encode graphemic, morphosyntactic, phonological, semantic, and contextual information as independently queryable layers. Annotation is embedded in a structured editorial workflow supporting peer review, expert validation, version tracking, and the representation of variant readings and uncertainty. A case study demonstrates how recurring formulaic constructions can be modeled as morphosyntactic configurations retrievable across inscriptions. Although the current corpus is limited in scope, the data model is language-agnostic, designed for extension to other Semitic epigraphic traditions.
Overview of the Dependency Parsing Task at EvaLatin 2026
Federica Iurescia | Marco Passarotti | Rachele Sprugnoli
Federica Iurescia | Marco Passarotti | Rachele Sprugnoli
This paper presents the organization, methodology, and outcomes of the Dependency Parsing shared task held within the fourth edition of EvaLatin, a campaign dedicated to the evaluation of Natural Language Processing tools for Latin. EvaLatin aims to promote and advance research in language technologies for Latin, fostering the development of robust and linguistically informed computational approaches. The paper provides a detailed description of the data released for the shared task. It also outlines the evaluation framework and metrics adopted for assessing system performance. The results achieved by participating teams are reported and comparatively analyzed, highlighting strengths, limitations, and emerging trends in current approaches to Latin dependency parsing. Finally, the paper discusses the main challenges posed by the task and suggests directions for future research in the field.
THIVLVC: Retrieval Augmented Dependency Parsing for Latin
Luc Pommeret | Thibault Wagret | Jules Deret
Luc Pommeret | Thibault Wagret | Jules Deret
We describe THIVLVC, a two-stage system for the EvaLatin 2026 Dependency Parsing task. Given a Latin sentence, we retrieve structurally similar entries from the CIRCSE treebank using sentence length and POS n-gram similarity, then prompt a large language model to refine the baseline parse from UDPipe using the retrieved examples and UD annotation guidelines. We submit two configurations: one without retrieval and one with retrieval (RAG). On poetry (Seneca), THIVLVC improves CLAS by +17 points over the UDPipe baseline; on prose (Thomas Aquinas), the gain is +1.5 CLAS. A double-blind error analysis of 300 divergences between our system and the gold standard reveals that, among unanimous annotator decisions, 53.3% favour THIVLVC, showing annotation inconsistencies both within and across treebanks.
Overview of the Named Entity Recognition Task at EvaLatin 2026
Valeria Irene Boano | Eleonora Litta | Matteo Romanello
Valeria Irene Boano | Eleonora Litta | Matteo Romanello
This paper describes the organisation and results of the Named Entity Recognition and Classification (NERC) shared task, conducted as part of EvaLatin 2026. The fourth edition of this evaluation campaign for Natural Language Processing on Latin features two shared tasks, i.e. Dependency Parsing and NERC. After introducing the objective of the task and presenting the Ancient Named Entities Special Interest Group, which aims to address the specific challenges that this task presents, this overview details the annotation tagset, the data provided to the participants and their format. The evaluation metrics and the scorer are also described. Finally, the methodology used by each participating team and their results are presented and discussed.
With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient language research by participating in EvaLatin 2026. This paper describes Team uOttawa’s system description and results for the Named Entity Recognition (NER) shared task. The task is divided into two subtasks: coarse-grained NER with 11 classes and fine-grained NER with 28 classes, each evaluated under strict and fuzzy regimes. Through prompt engineering of commercial LLMs gemini-2.5-pro and claude-sonnet-4-5, I show that the underrepresented ancient Latin language can take advantage of cross-lingual transfer learning by using advancements made by the wider LLM development community. Overall, the methods discussed in this report demonstrate very strong results, placing first in both NER subtasks and achieving the best scores across all evaluation metrics and regimes among all submissions.
Overview of EvaHan2026: The First International Evaluation of Ancient Chinese OCR and Layout Analysis
Dongbo Wang | Dongmei Zhu | Jieqiong Li | Chang Liu | Ruifeng Wu | Fan Yang | Zhu Yue | Xin Gao | Zhao Zhixiao | Xue Zhao | Zhongzheng Wu | Liu Liu | Li Bin
Dongbo Wang | Dongmei Zhu | Jieqiong Li | Chang Liu | Ruifeng Wu | Fan Yang | Zhu Yue | Xin Gao | Zhao Zhixiao | Xue Zhao | Zhongzheng Wu | Liu Liu | Li Bin
Ancient Chinese documents are vital for historical research, necessitating high-precision character recognition and layout analysis for digitization. This paper introduces EvaHan2026, the inaugural international shared task for simultaneous optical character recognition and layout parsing of ancient texts. The evaluation framework comprehensively assesses model performance across diverse calligraphic styles and complex structures, including main body text, interlinear annotations, and illustrations. Among thirteen participating teams, four successfully completed all tasks within the closed track. Experimental results reveal that character recognition accuracy reached 97.36% on engraved texts (Test Set A) and 95.71% on handwritten texts (Test Set C) when accounting for character variants. For layout recognition in complex layouts (Test Set B), the best team achieved a peak mean Average Precision (mAP) of 59.41% and an Intersection over Union (loU) of 76.38%. Our analysis indicates that calligraphic variability, layout density, and character variants significantly modulate system performance. Consequently, enhancing robustness within complex layouts and developing synergistic models that integrate textual and structural information remain primary challenges for intelligent interpretation of ancient writings .
A Multi-Stage System for Ancient Chinese OCR and Layout Understanding in the EvaHan2026 Shared Task
KeYan Liang | Meiling Liu
KeYan Liang | Meiling Liu
This paper presents a multi-stage system for the EvaHan2026 shared task, addressing the complex challenges of ancient Chinese optical character recognition (OCR) and layout understanding. For text recognition (Tasks A and C), we adopt parameter-efficient LoRA fine-tuning on the Qwen2.5-VL-7B-Instruct vision-language model (VLM). By directly processing full-resolution long-column images, we preserve critical spatial and contextual integrity without heuristic region cropping. For document layout analysis (Task B), we propose a novel hybrid perception-reasoning paradigm. Instead of relying solely on scaling visual detectors, we decouple localization and understanding: utilizing a YOLO-based ensemble for precise spatial bounding, and casting the VLM as a semantic verifier to eliminate spurious detections. Evaluated on the official unseen test set, our system achieves substantial improvements over the provided baselines, obtaining a 0.0441 Character Error Rate (CER) for printed OCR, a 0.0793 CER for handwritten OCR (including variants), and a 0.5118 mAP@[0.5:0.95] for layout detection. These results demonstrate that integrating VLM-based semantic reasoning into traditional visual detection pipelines is highly effective for multimodal historical document analysis.
A Multi-Modal Recognition Framework for Ancient Books Integrating DoRA-DPO Text Recognition and YOLO Layout Analysis
Chaokun Zhang | Xin Wen | Tongtong Zhou
Chaokun Zhang | Xin Wen | Tongtong Zhou
The digitization and intelligent analysis of ancient Chinese documents face significant challenges due to diverse scripts, complex layouts, and the prevalence of rare characters. We present a comprehensive multi-modal recognition framework developed for the closed-modality track of the EvaHan 2026 Ancient Chinese Document Multi-Modal Recognition Shared Task. Our approach integrates two specialized pipelines to address these complexities. For text recognition (Tasks A and C), we propose a high-precision OCR system based on the domain-adapted Xunzi_Qwen2_VL_7B_Instruct, leveraging DoRA within a two-stage progressive curriculum learning strategy. To further refine character accuracy, DPO is incorporated alongside a dual-adapter architecture for rare character error localization and correction. For layout detection (Task B), we implement DocLayout-YOLO, enhanced by domain-specific pre-training and Mosaic augmentation to achieve efficient NMS-free element detection. Furthermore, a multi-round robust inference strategy, featuring automatic retry mechanisms and multi-prompt brute-force search, is introduced to handle stubborn and degraded samples effectively. Experimental results demonstrate that our proposed framework achieves superior performance across all evaluation metrics, highlighting its robustness and effectiveness in the digital preservation of ancient Chinese heritage.
Beijing Normal University at EvaHan 2026: Enhancing Ancient Chinese Character Recognition and Layout Analysis via VLM Fine-Tuning and Linguistic Post-Processing
Yihuan Yin | Qian Zhao
Yihuan Yin | Qian Zhao
This paper describes the system submitted by the Beijing Normal University (BNU) team for the EvaHan 2026 shared task. We participated in Task A (Printed Text Recognition), Task B (Layout Element Analysis), and Task C (Handwritten Character Recognition). For text recognition (Tasks A and C), we proposed a hybrid pipeline combining supervised fine-tuning (SFT) of Vision-Language Models (VLMs) with a linguistic rule-based post-processing module. In the Open Track, we further explored the use of a general-purpose VLM to correct semantic errors while maintaining visual fidelity to ancient variant characters. For Task B, we adopted a method integrating a VLM with structured prompting strategies. Our system consistently surpassed the official baselines, achieving an F1 score of 94.53% in Task A and 91.33% in Task C, while demonstrating enhanced localization precision in Task B.
A Dual-Modality Framework for Ancient Document Layout Analysis and Text Recognition
Qi Fan | Jieming Hu | Chen Ye
Qi Fan | Jieming Hu | Chen Ye
The digital preservation of ancient Chinese literature requires robust capabilities spanning layout analysis and text recognition. This paper presents a comprehensive framework addressing two fundamental challenges: (1) Layout Element Analysis (Task B) for detecting page elements (text, image, book_edge, seal) amidst degradation, nested structures, and extreme class imbalance; and (2) Text Recognition (Tasks A & C) for end-to-end transcription of printed and handwritten classical documents. For layout analysis, we propose a dual-modality solution. The Closed Modality formulates this as a sequence-to-sequence problem using Vision-Language Models (VLMs), introducing spatial discretization tokenization and a Frequency-Aware Sequential Curriculum Learning framework with dynamic memory replay. The Open Modality presents HistLayout-DETR, a set prediction architecture integrating an Augmented Morphological Encoder and a Polygon Boundary Refinement head. For text recognition, we formulate OCR as a domain-constrained visual language generation task using Qwen2.5-VL with LoRA fine-tuning. We employ structured prompts encoding reading order and Traditional Chinese character preservation across domains. Extensive experiments on the EvaHan 2026 dataset validate our framework’s superiority. In layout analysis, our curriculum-guided paradigm achieves a Macro F1 of 0.7992 and mAP@[.5:.95] of 0.5438. In text recognition, we achieve CERs of 0.0271 on printed and 0.0433 on handwritten texts.
This paper introduces our system proposal and experimental results for the 5th International Evaluation of Ancient Chinese Information Processing (EvaHan 2026). This evaluation focuses on ancient books OCR tasks using multimodal large language models, including three subtasks: Printed Text Recognition (Task A), Layout Element Analysis (Task B), and Handwritten Text Recognition (Task C). To address core challenges such as numerous variant characters, complex handwritten ligatures, dense layout elements, and annotation noise, we propose a Supervised Fine-tuning (SFT) scheme based on data synthesis augmentation and multi-stage curriculum learning. We also optimized the data preprocessing workflow, resolving key issues like repetition mark recognition and annotation quality improvement. We completed a 9:1 train-validation split on the official dataset and verified the effectiveness of our methods through 6 groups of comparative experiments. Finally, we selected the model with the best comprehensive performance for submission. The code and synthetic dataset are available at https://github.com/zhengningch/EvaHan2026-data.
A Parameter-Efficient and Data-Centric Framework for Ancient Chinese Text Recognition and Layout Analysis
Yuchun Meng
Yuchun Meng
This paper presents the system developed for the EvaHan 2026 shared task on Ancient Chinese OCR and Layout Analysis. Participating in the Closed Track, we propose a highly parameter-efficient, data-centric framework based on the Qwen2.5-VL-7B-Instruct multimodal large language model (MLLM). While the official baseline utilizes the same backbone architecture, our approach significantly outperforms it by integrating orientation-aware image preprocessing and expert-constrained adaptive prompt engineering. We employed Low-Rank Adaptation (LoRA) with a minimal rank configuration (Rank=16) to train three independent, task-specific adapters. Our system achieved exceptional results, recording an Overall score of 0.9703 and an F1-score of 97.19% on printed text recognition (Task A)—effectively halving the baseline’s Character Error Rate. On handwritten texts (Task C), we maintained a highly competitive 90.18% F1-score. Furthermore, our model achieved significant progress in layout analysis (Task B), surpassing the baseline’s Macro F1 by 172% (0.4162 vs. 0.1530) and mAP by 37%. These results underscore that embedding explicit document structure and semantic constraints into MLLMs is more effective than simply scaling model parameters.
LVLM Optimization for Ancient Chinese Book Image Analysis with Task-specific Augmentation and Instruction Tuning
Tian Xia | Yulong Liu | Yilin Wang | Yumeng Yang | Dongheng Cai | Yuyang Tan | Menghui Yang
Tian Xia | Yulong Liu | Yilin Wang | Yumeng Yang | Dongheng Cai | Yuyang Tan | Menghui Yang
Ancient Chinese text digitization faces challenges like variant characters and complex layouts. Based on the EvaHan 2026 tasks, this study proposes an LVLM-based framework for printed/handwritten text recognition and layout analysis. To effectively adapt the Qwen2.5-VL-7B-Instruct model, our methodology innovates through a dual-level optimization strategy: distinct augmentation strategies are developed for OCR and layout tasks, while task-specific prompt templates are engineered to decouple text transcription from coordinate prediction. This combined approach significantly enhances overall task proficiency, achieving Character Error Rates of 0.0372 (printed) and 0.0823 (handwritten), alongside a mean average Precision of 0.2933 for layout analysis. Results show general LVLMs underperform in zero-shot ancient text tasks, but fine-tuning with tailored strategies significantly boosts performance and highlights their potential.
Data-Centric Strategies for Ancient Chinese Text Recognition: Augmentation, Annotation Refinement, and Style Transfer in EvaHan 2026
Li Chengfei | Zhang Yunjie | Li XiaoYi | Quan Changshun | Cao Taihe | Liu Bin
Li Chengfei | Zhang Yunjie | Li XiaoYi | Quan Changshun | Cao Taihe | Liu Bin
This paper describes our system for the EvaHan 2026 shared task. We design and experiment with data-centric strategies across three subtasks: printed text OCR (Task A), layout element analysis (Task B), and handwritten text OCR (Task C). Our approach employs systematic data augmentation using 17 transformation strategies, comprehensive manual annotation refinement for layout analysis, and style transfer augmentation for handwritten texts. We use pre-trained Qwen2.5-VL-7B-Instruct with LoRA fine-tuning as the base model. According to the evaluation metrics adopted by the organizers, our system achieves 27.5% and 4.5% CER reduction over official baselines for Tasks A and C respectively. Manual annotation refinement for Task B achieves 205% improvement in Micro F1 and 258% improvement in Macro F1 on the validation set, demonstrating that annotation quality is the primary bottleneck for layout analysis in closed-modality settings.
AnandaSky: A Vision–Language Model for Line-Level Transcription of Historical Sinographic Documents
Colin Brisson | Ayoub Kahfy | Frédéric Constant | Marc Bui
Colin Brisson | Ayoub Kahfy | Frédéric Constant | Marc Bui
We present AnandaSky, a vision–language model for line-level transcription of historical sinographic documents. The model combines a compact high-resolution visual encoder with global attention, 10px patches, uncompressed visual prefix and a Qwen3-0.6B autoregressive decoder. It is trained at scale on 4M annotated lines from documents produced in China and Korea between the 8th and 20th centuries. Across in-domain and held-out public benchmarks, AnandaSky achieves sub-1% CER on five of eight datasets, sets a new state of the art on MTHv2 with 0.92% CER, and shows strong transfer to unseen collections. For EvaHan 2026, full fine-tuning on the organizers’ data to match task-specific annotation conventions reduces CER relative to the official baseline by 5.2% on prints and 12.1% on manuscripts, despite using one-tenth as many parameters.
Multimodal Ancient Document Parsing: Technical Report for EvaHan2026 Competition
Liqi He | Qiwei Li | Ziye Yang | Zuchao Li
Liqi He | Qiwei Li | Ziye Yang | Zuchao Li
We present the multimodal Optical Character Recognition (OCR) and layout analysis methods developed for the EvaHan 2026 competition. Our approach is built upon the Qwen2.5-VL-7B-Instruct architecture and integrates two core strategies: (1) a reinforcement learning alignment pipeline utilizing Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO) to explicitly mitigate hallucination and coordinate instability; and (2) a four-stage curriculum learning framework that synthesizes domain-specific historical artifacts to enhance open-modality generalization. Using this approach, we achieve competitive results, notably reaching a Character Error Rate (CER) of 0.0303 on printed texts (Task A) and 0.0552 on handwritten manuscripts (Task C), as well as an Average Intersection over Union (IoU) of 0.7638 on layout element analysis (Task B).
Multi-Task Learning Trade-offs in Vision–Language Models for Ancient Chinese OCR: An Empirical Analysis of Parameter-Efficient Adaptation
Yuhan Shu | Huizi Zhou
Yuhan Shu | Huizi Zhou
This study evaluates the efficacy of multi-task adaptation in large-scale vision–language models (VLMs), specifically Qwen2.5-VL, for the simultaneous recognition and structural parsing of historical Chinese documents within the EvaHan2026 benchmark. Utilizing a parameter-efficient fine-tuning (PEFT) strategy via LoRA (rank 64), our framework demonstrates superior performance in layout analysis (Task B), achieving an mAP of 0.2802—a 39.6% improvement over the competitive baseline—and a Macro F1 of 0.3609. Conversely, a pronounced performance-utility trade-off is observed in printed OCR (Task A), where the character error rate (CER) escalates from 0.0618 to 0.1100 (+78% relative). This divergence highlights a critical catastrophic forgetting effect induced by gradient interference during multi-task optimization. While handwritten OCR (Task C) remains relatively stable (CER of 0.0963), our findings suggest that although unified VLM architectures excel at high-level structural detection, they encounter significant parameter capacity bottlenecks when concurrently optimizing fine-grained character-level transcription. This analysis highlights the optimization challenges when balancing spatial detection and character recognition in a unified framework.
Building Character(s): Synthetic Data and In-Context Learning Strategies for Few-Shot Ancient Chinese Recognition
Denise Atzori | Marie Bizais-Lillig | Mathias Garnier | Maxime Létoffé | Charles Planque | Tianjie Yin | Chahan Vidal-Gorène
Denise Atzori | Marie Bizais-Lillig | Mathias Garnier | Maxime Létoffé | Charles Planque | Tianjie Yin | Chahan Vidal-Gorène
Ancient Chinese character recognition remains challenging due to severe character imbalance, graphic variants, peculiar layout, degraded printing, and limited annotated data. This paper presents our system for EvaHan 2026, combining synthetic data generation and in-context learning (ICL) across three tasks: line-level text recognition (printed and handwritten) and page layout detection. We introduce UltraGlyph, a synthetic data pipeline recombining glyphs from real data with font-generated characters to improve rare-character coverage, producing 234,528 line images for foundation-model pretraining. We benchmark CRNN, transformer-based OCR, and a suite of vision–language models under a variant-aware ICL framework. On printed text, dedicated OCR systems and top VLMs reach comparable comprehensive scores with around 97% of accuracy; on cursive handwriting, performance drops significantly and is bounded above by 95%, with the best result achieved by Qwen2.5-VL-72B in zero-shot. For layout analysis, YOLO12s achieves the best score with a mAP50 of 75%.
The UD_Latin-PROIEL as Linked Open Data: Integrating a Latin Treebank into the LiLa Knowledge Base
Lucas Consolin Dezotti | Marco Passarotti | Federica Iurescia | Giovanni Moretti
Lucas Consolin Dezotti | Marco Passarotti | Federica Iurescia | Giovanni Moretti
This paper presents the steps taken to integrate data from the UD_Latin-PROIEL treebank into the LiLa Knowledge Base of interoperable linguistic resources for Latin. It describes how the lexical, morphological, syntactic, and citation information from the source was modeled using the Linked Open Data principles as adopted by the LiLa Knowledge Base. The process of linking tokens to the LiLa collection of Latin lemmas is detailed, addressing challenges such as ambiguities, new lemmas, and errors encountered in the source. The outcome is a syntactically annotated textual resource that is interoperable with the (meta)data of other Latin linguistic resources linked within the LiLa Knowledge Base. This integration enables new ways of analyzing linguistic information and using the content as a starting point to explore connections with other interlinked resources. A use case demonstrates this interoperability.
Language Models for the Restoration of Latin Legal Manuscripts
Shibingfeng Zhang | Edoardo Caraffa | Annafelicia Zuffrano | Maddalena Modesti | Giovanni Colavizza
Shibingfeng Zhang | Edoardo Caraffa | Annafelicia Zuffrano | Maddalena Modesti | Giovanni Colavizza
The collection of historical notarial documentation from Bologna is a valuable source, providing deep insights into the city’s institutional, legal, and socio-economic history. However, many of these manuscripts have sustained physical damage during centuries of conservation, rendering the text incomplete. To address this, we explored the restoration of these Latin notary documents using encoder-based pre-trained language models (PLMs) under the assumption that the length of missing text is known by estimation from the physical damage. We address the structural misalignment between the physical lacuna of the manuscript and the subword tokenization schemes of PLMs by designing an iterative decoding strategy to align model predictions with the known physical dimensions of lacuna. We also compared the efficacy of monolingual versus multilingual pre-training. Our strategy significantly outperforms baselines consist of standard decoding methods. Furthermore, stratified analysis across different text sections reveals that while monolingual models achieve better performance in general, multilingual models show a suggestive advantage in lexically dense segments, though this finding is not statistically significant. Overall, the best performance achieved by our method is a Hit@1 rate of 35.47% in the short-span setting and 18.75% in the long-span setting. While fully autonomous restoration remains an open challenge, our system provides a useful assistive tool for paleographers.
Evaluating Hierarchical Aggregation and LLM-Based Matching for Synset Selection in Ancient Greek
Luca Brigada Villa | Marco Passarotti | Chiara Zanchi | Riccardo Ginevra | Erica Fratellini | Eleonora Litta
Luca Brigada Villa | Marco Passarotti | Chiara Zanchi | Riccardo Ginevra | Erica Fratellini | Eleonora Litta
This paper presents a structured framework for WordNet synset selection applied to Ancient Greek lexical material. Starting from synonym definitions extracted from the Liddell–Scott–Jones (LSJ) lexicon, we compare two strategies: hierarchy-driven aggregation via bounded hypernym trees and LLM-based definitional matching with pairwise ranking. Graded human evaluation shows that structure-aware methods provide a robust baseline, particularly for nouns and verbs, while LLM-based reranking does not consistently improve performance, especially for highly ploysemous groups of synonyms. Beyond supporting the development of an Ancient Greek WordNet, the study highlights the methodological portability of the framework to other languages and lexical resources.
Miktub: A Manuscript Dataset of Historical Maltese for Handwritten Text Recognition
Thomas Koppens | Claudia Borg
Thomas Koppens | Claudia Borg
The digitisation of handwritten historical material is essential for preserving cultural heritage and enabling search and computational analysis. For Maltese, historical handwritten resources are scarce, and, to the best of current knowledge, no public handwritten text recognition (HTR) dataset for historical Maltese exists. We introduce a Manuscript Dataset of Historical Maltese (Miktub), collected from the Data Provider: 35 scanned pages transcribed by specialists and converted into a line-level HTR dataset. A key challenge was robust line extraction from heterogeneous pages; fully automatic line segmentation was insufficient, so we developed a semi-automatic pipeline combining horizontal projection profiling with lightweight post-processing and manual refinement to maximise line fidelity. We provide two annotation variants, including a corrected/standardised version (Miktub-COR) designed to improve consistency, accessibility, and downstream learning stability. We benchmark two strong public HTR models, HTR-VT and VAN, and report the best test performance of 4.68% character error rate (CER) and 13.59% word error rate (WER) on Miktub-COR with VAN. We will release Miktub publicly upon acceptance, along with scripts and splits, to support historical Maltese-language technology research.
Smelling the Past: Investigating Historical Models for Olfactory Event Extraction
Teresa Paccosi | Marijn Koolen
Teresa Paccosi | Marijn Koolen
In this paper, we present a series of experiments using historical language models to investigate the impact of pretraining on data that more closely resembles the task domain, focusing on the case study of automatic olfactory event extraction. We tested historical and contemporary pretrained models on the task of extracting olfactory events using a benchmark spanning several centuries. The aim of our research is not only to assess whether historical models can improve performance on this diachronically oriented task, but also to gain deeper insight into the factors influencing model performance through a detailed analysis of performance patterns. We examine potential sources of variation and previously proposed hypotheses to account for lower performance observed in this task, thereby offering a more comprehensive understanding of model behavior in this context.
The paper describes a contemporization effort of a 1.9 million word corpus of Estonian parliament minutes from 100 years ago. The paper describes the corpus of Asutaw Kogu (the Constitutional Assembly) and the main differences of language that require one to contemporize it for modern researchers. The effort is implemented as a work flow that combines a freely available speller lexicon, hand-crafted transformation rules and various corpus-based word lists into finite state transducers. Evaluation on a 53,000 token subset of the corpus showed that 0.02% of text tokens ended up with an incorrect contemporary form, corresponding to 0.05% of the corpus vocabulary. However, if we count only the tokens that actually need changing in the contemporization process, we see that 0.12% end up being incorrect, corresponding to 0.15% of the corpus vocabulary. An additional experiment with generative AI showed that using it as a contemporization tool results in a content-preserving, but more formal version of the original minutes.
Cost-Aware Pre-Annotation Strategies for Nested NER in Historical Latin Notarial Deeds
Charlene Ellul | Vanessa Buhagiar | Claudia Borg | Charlie Abela
Charlene Ellul | Vanessa Buhagiar | Claudia Borg | Charlie Abela
Manual annotation for Named Entity Recognition in historical documents remains expensive and time-consuming, particularly for complex nested entity structures in domain-specific texts such as Latin notarial deeds. Active learning frameworks like the Humanities Entity Recognizer (HER) reduce annotation requirements by iteratively selecting informative samples for expert annotation, but existing sentence-based sampling strategies create unpredictable annotation costs when sentence lengths vary dramatically. We extend the HER to support nested entities through composite BIO label encoding and introduce token-budgeted sample selection to address annotation cost variability. Under token-budgeting, each annotation iteration targets a fixed token budget rather than a fixed sentence count, while Active Curriculum Learning ensures diverse sentence length representation in initial samples. Experiments on seventeenth-century Latin notarial deeds from Malta’s Notarial Registers Archive demonstrate that token-budgeted sampling achieves comparable macro-span F1 to sentence-based sampling while exhibiting more stable learning trajectories across iterations. Additional experiments examining entity-level performance reveal systematic variation by semantic granularity, with higher-level categorical entities achieving stronger recognition than role-based middle-level entities, which depend on discourse context. Our results demonstrate that controlling sample selection at the token level rather than sentence level provides more predictable annotation planning for active learning in historical document corpora with heavy-tailed sentence length distributions.
From Lemmatization to Legal Terminology: Assessing an Hybrid Pipeline on Justinian’s Digest
Paola Marongiu | Eva Sassolini
Paola Marongiu | Eva Sassolini
This paper evaluates a hybrid NLP pipeline for supporting the extraction of Roman legal terminology from Jus- tinian’s Digest. Our goal is not to optimize lemmatization in isolation, but to assess whether integrating a Large Language Model (GPT-4o-mini) as a post-processing component improves lemma quality in ways that are critical for downstream glossary construction. Using LatinPipe as a baseline (F1 = 95.05), we test the integration of GPT-4o-mini under three experimental settings (zero-shot with and without prior lemma information, and few-shot prompting) against a manually annotated gold standard of 3,703 sentences and an expert-validated list of legal Latin technical terms. Results show improvement across all settings, with the best performance achieved in the few-shot configuration. Our analysis shows that the hybrid configuration produces selective improvements, significantly more likely for frequent lemmas and verbs forms, suggesting that the LLM layer primarily assists in resolving morphologically ambiguous inflected forms. Although our experimental conditions may not hold in real-world scenarios, we argue that the main contribution of this work is methodological: demonstrating how evaluation can be aligned with downstream terminological goals, rather than proposing a general-purpose solution to domain-specific lemmatization.
We describe the UppsalaNLP submission to the EvaLatin dependency parsing shared task. We explore using an out-of-the-box parser in combination with multi-treebank training on Latin and multilingual training on other ancient languages. Adding additional languages yields only small gains, but the results vary across treebanks and genres, with the largest positive effect for poetry. Our systems perform best in the shared task for prose but are less competitive for poetry, indicating the need for genre adaptation.
Contextual Probing for Low-Resource Named Entity Recognition in Latin
Maria Mihaela Trusca | Mark Depauw | Violet Soen | Ine de Daele | Kevin Verbruggen | Tim Van de Cruys
Maria Mihaela Trusca | Mark Depauw | Violet Soen | Ine de Daele | Kevin Verbruggen | Tim Van de Cruys
Named Entity Recognition (NER) for low-resource languages remains challenging due to limited annotated data and linguistic characteristics such as rich morphology and flexible word order. In this work, we propose a probing-based method that leverages the contextual knowledge encoded in pretrained language models to detect entities. Our approach uses a substitution strategy in which words in a sentence are replaced, one by one, with candidate entities of predefined entity types, referred to as probes. By measuring how well the probes of a certain entity type fit the surrounding context of the replaced word, we estimate the compatibility between the replaced word and the entity type. The resulting compatibility scores can be used either as a standalone zero-shot NER model or as an auxiliary feature during NER model decoding. We evaluate our method on the Latin dataset provided in the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA). Our system ranked second in the coarse-grained NER task. For the fine-grained NER task, where no training data were available, we relied exclusively on the proposed scoring method without any model training and achieved third place. These results demonstrate that contextual probing can provide an effective signal for NER in low-resource settings.
Classificatio Sine Iactu – That Is, Zero-Shot NERC in Latin
Luisa Ripoll-Alberola | Fernando Nicolás-Flores | Francisco Javier Muñoz Acebes
Luisa Ripoll-Alberola | Fernando Nicolás-Flores | Francisco Javier Muñoz Acebes
This paper presents a zero-shot approach to Named-Entity Recognition and Classification (NERC) in Latin, applied to the EvaLatin shared task. Given the novelty and granularity of the annotation guidelines, which preclude the use of existing annotated resources, we employ the zero-shot model GLiNER2, a general information extraction system capable of CPU-efficient inference, within a cross-lingual pipeline. Latin texts are first translated into English via the Google Translate API, processed by the model, and the resulting annotations are aligned back to the original Latin using word-alignment techniques. Rule-based post-processing addresses labelling inconsistencies and low-confidence predictions. We evaluate two model variants, a large monolingual and a multilingual model, under both strict and fuzzy evaluation. The large model delivers the best results for the coarse-grained task (F1: 0.590 fuzzy), while the multilingual model outperforms it on the fine-grained task (F1: 0.432 fuzzy). Results indicate that multilingual embeddings confer an advantage for fine-grained semantic distinctions, that English embeddings introduce systematic bias in cross-lingual transfer, and that zero-shot NER represents a viable, reproducible baseline for low-resource historical languages. Fine-tuning on guideline-compliant annotated data remains a priority for future work.
Extending omnes flores for the EvaLatin 2026 Dependency Parsing Tasks
Hiroshi Matsuda | Masayuki Asahara
Hiroshi Matsuda | Masayuki Asahara
omnes flores is an NLP framework based on Universal Dependencies (UD) that utilizes multilingual Large Language Models (LLMs), and its default model is trained on data from 40 UD languages comprising 40 treebanks. For the EvaLatin 2026 Dependency Parsing Tasks, we extended the training data of omnes flores by incorporating six public Latin treebanks from UD and trained a dependency parsing model using the extended training data. The dependency parser of omnes flores normally takes a list of word FORM values as input. However, since the EvaLatin 2026 test data includes an UPOS column, we investigated whether incorporating both FORM and UPOS during both training and inference could improve parsing accuracy. Our experiments show that training using both FORM and UPOS improves performance by 0.5-1.0 LAS points on Prose compared with training using only FORM, but decreases performance by 5 points on Poetry.
Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
Thibault Clérice | Rachel Bawden | Anthony Glaise | Ariane Pinche | David Smith
Thibault Clérice | Rachel Bawden | Anthony Glaise | Ariane Pinche | David Smith
Recent advances in Automatic Text Recognition (ATR) have improved access to historical archives, yet a methodological divide persists between palaeographic transcriptions and normalized digital editions. While ATR models trained on more palaeographically-oriented datasets such as CATMuS have shown greater generalizability, their raw outputs remain poorly compatible with most readers and downstream NLP tools, thus creating a usability gap. On the other hand, ATR models trained to produce normalized outputs have been shown to struggle to adapt to new domains and tend to over-normalize and hallucinate. We introduce the task of Pre-Editorial Normalization (PEN), which consists in normalizing graphemic ATR output according to editorial conventions, which has the advantage of keeping an intermediate step with palaeographic fidelity while providing a normalized version for practical usability. We present a new dataset derived from the CoMMA corpus and aligned with digitized Old French and Latin editions using passim. We also produce a manually corrected gold-standard evaluation set. We benchmark this resource using ByT5-based sequence-to-sequence models on normalization and pre-annotation tasks. Our contributions include the formal definition of PEN, a 4.66M-sample silver training corpus, a 1.8k-sample gold evaluation set, and a normalization model achieving a 6.7% CER, substantially outperforming previous models for this task
OldBERTur: Named Entity Recognition for Medieval Icelandic
Pontus Henningsson | Eva Pettersson | Erik Lenas
Pontus Henningsson | Eva Pettersson | Erik Lenas
We present OldBERTur, a Named Entity Recognition (NER) model for Old Icelandic available in two variations, one for normalised texts, and one for diplomatic texts. Using a BERT-based model architecture, we fine-tune an existing BERT language model, and due to training data scarcity, we employ multiple training configurations, including pre-training domain adaptation, sentence-level data resampling, and modern Icelandic data augmentation; achieving a 93 F1 score for normalised texts, and 79 for diplomatic texts. We find that additional training configurations, such as resampling entity-annotated Old Icelandic texts, significantly improve performance in low-resource settings, while the effectiveness of added training configurations diminishes as the available training data increases. Our models can be used to automatically identify and classify person and location names in texts sourced from the rich Icelandic medieval literary tradition. Our models, along with their data and code, are made publicly available to allow for reuse and future research into medieval Scandinavian NLP and beyond.
Neural Machine Translation for Coptic-French: Strategies for Low-Resource Ancient Languages
Nasma Chaoui | Richard Khoury
Nasma Chaoui | Richard Khoury
This paper presents the first systematic study of strategies for translating Coptic into French. Our comprehensive pipeline systematically evaluates: pivot versus direct translation, the impact of pre-training, the benefits of multi-version fine-tuning, and model robustness to noise. Utilizing aligned biblical corpora, we demonstrate that fine-tuning with a stylistically-varied and noise-aware training corpus significantly enhances translation quality. Our findings provide crucial practical insights for developing translation tools for historical languages in general.
up
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Mustafa Jarrar | Mo El-Haj | Amal Haddad | Serin Atiani | Shadi Abudalfa | Terry Regier | Paul Rayson | Khalil Sima’an | Camille Mansour
Mustafa Jarrar | Mo El-Haj | Amal Haddad | Serin Atiani | Shadi Abudalfa | Terry Regier | Paul Rayson | Khalil Sima’an | Camille Mansour
The NakbaEcho Dataset: From Oral Testimonies to a Transcribed Arabic History Corpus
Batool Najeh Balah | Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Batool Najeh Balah | Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
We present NakbaEcho, a dataset derived from Palestinian testimonies about the 1948 Nakba. The resource is constructed from transcribing over 2,180 hours of recorded interviews gathered through the Palestine Remembered Oral History index and linked to multiple repositories, including the Palestinian Oral History Archive (POHA) and YouTube-hosted interviews. We harmonize interview-level metadata and generate timestamp-aligned transcripts from the original Arabic recordings using an automatic transcription pipeline configured for Palestinian Arabic. The dataset includes speaker-labeled segments and auxiliary annotations designed to support downstream research in Arabic speech processing, natural language processing, digital humanities, and oral-history analysis. NakbaEcho contributes a structured computational resource for studying Palestinian oral testimony while expanding the availability of dialectal Arabic materials for speech, text, and social research.
Mining the Pre-1948 Palestinian Press: Unsupervised Keyphrase Extraction and Temporal Discourse Analysis from Five Historical Arabic Newspapers
Basel Barakat | Nizam Barakat
Basel Barakat | Nizam Barakat
The Palestinian Arabic-language press of the late Ottoman and British Mandate periods constitutes a rich but computationally under-explored archive for studying the evolution of political, cultural, and social discourse in pre-1948 Palestine. This paper presents an end-to-end pipeline for extracting and analyzing thematic content from five historically significant Palestinian newspapers: Lisān al-ʿArab (اﻟﻌﺮب ﻟﺴﺎن), Al-Bushrā (اﻟﺒﺸﺮى), Al-Karmil (اﻟﻜﺮﻣﻞ), Al-Difāʿ (اﻟﺪﻓﺎع), and Filasṭīn (ﻓﻠﺴﻄﻴﻦ). We describe (i) the construction of a five-source corpus from scanned newspaper images obtained from archival collections, processed using the Google Cloud Vision OCR (GCV-OCR) API, (ii) the adaptation of KeyBERT with an AraBERT backbone for unsupervised keyphrase extraction, and (iii) a purpose-built Python visualization toolkit that produces keyword-frequency heatmaps, longitudinal trend charts, and ranked bar charts with full Arabic script rendering. Experiments across the five subcorpora show that the pipeline yields topically diverse keyphrases reflecting each newspaper’s editorial orientation and its distinct representation of Palestinian native perspectives—from pan-Arab nationalism and anti-colonial resistance to religious and communal affairs. Temporal analysis reveals event-responsive patterns that align with major historical developments, including the 1936–1939 Arab Revolt and the intensification of sovereignty discourse toward 1948. The pipeline, data format specifications, and visualization code are provided as supplementary material.
Nakba Discourse 2025: A Bilingual Social Media Dataset for Collective Trauma Analysis
Wajdi Zaghouani | Mabrouka Bessghaier | Kais Attia
Wajdi Zaghouani | Mabrouka Bessghaier | Kais Attia
We introduce Nakba Discourse 2025, a bilingual full-year social media dataset capturing Arabic and English discourse about the 1948 Palestinian Nakba across Twitter/X and Facebook from January to December 2025. The corpus contains 70,312 unique posts organized into intersecting sub-corpora by language, sentiment, gender, geography, and platform, with engagement metadata and automatically extracted rhetorical features. Analyses reveal systematic variation in engagement and framing across communities. Per-post engagement is highest in Israel and UK subsets (50.62 and 49.08 average likes respectively), while Arabic-language discourse shows markedly lower per-post engagement. Sentiment distribution is strongly skewed, with negative sentiment posts outnumbering positive ones at an 11:1 ratio (54,424 vs. 4,827 posts). Despite dramatic variation in absolute engagement levels, virality rates remain structurally constant at approximately 10% across all Twitter/X sub-corpora, regardless of language, gender, or geography, pointing to platform-level amplification regularities. Gender analysis reveals that women achieve proportional virality equal to men despite producing roughly one-third the volume of posts. Temporal patterns align with cultural calendars, including Thursday peaks associated with Jumu’ah across Arabic and female subsets, and Sunday peaks in English-language subsets reflecting Western media cycles. The dataset will be released for research use and supports multilingual stance detection, virality modeling, rhetorical analysis, and computational studies of digital political memory.
ChronoLearn: A GRAG LLM-Based System for Structuring and Exploring Historical Narratives
Mohammad O. ALADDASI | Shahd L. Abu Hijleh | Omar Qawasmeh
Mohammad O. ALADDASI | Shahd L. Abu Hijleh | Omar Qawasmeh
ChronoLearn is a KG–LLM framework to structure and ex- plore Arabic historical narratives. It transforms unstructured texts into knowledge graphs using an ETL-based NLP pipeline for entity and re- lation extraction, followed by schema-guided graph construction. The system integrates graph retrieval with LLM generation (GRAG) to pro- duce grounded, explainable narratives and support semantic querying. The approach is evaluated in heterogeneous Palestinian and Jordanian sources, including Nakba-related content, using both quantitative met- rics and comparative analysis. The results demonstrate improved factual grounding and structured reasoning, addressing limitations of text-only approaches in the processing of historical knowledge in Arabic.
Credibility Assessment for Arabic News on the Gaza War: A Hybrid Neural-Symbolic Pipeline
Sanaa Abril | Sihame Mouanid | El habib Ben lahmar | Omar Zahour
Sanaa Abril | Sihame Mouanid | El habib Ben lahmar | Omar Zahour
While misinformation has long circulated online, the Gaza conflict has intensified its visibility and spread across news websites, online portals, and social media, complicating the credibility and long-term curation of conflict-related Arabic records, including historical accounts and written testimonies. This work proposes a hybrid framework for Arabic fake news detection that combines interpretable linguistic cues with contextual semantic representations. The approach integrates fuzzy logic-based handcrafted features capturing exaggerated and sensational linguistic patterns, AraBERT contextual embeddings for semantic understanding, and a CNN-based text feature extractor for local textual patterns. These complementary features are combined into a unified representation for downstream classification. Multiple machine learning and deep learning classifiers are evaluated to identify the most effective detection model. The resulting system is deployed as a real-time web browser plugin, enabling users to automatically assess the credibility of Arabic news content during browsing
Tarikhi: Arabic Temporal Information Extraction from Arabic Historical Documents
Qusay Abdo | Serin Atiani | Tariq Sraiji | Adnan Saeed
Qusay Abdo | Serin Atiani | Tariq Sraiji | Adnan Saeed
Arabic historical books and archival materials contain rich accounts of political, social, and cultural events, yet they remain largely underutilized computationally due to the scarcity of dedicated Arabic information extraction tools. The challenge is amplified in long-form, scanned historical documents, where optical character recognition noise, orthographic variation, and complex narrative structures complicate automatic processing. In this paper, we present Tarikhi, a retrieval-augmented generation framework for structured temporal event extraction from Arabic scanned books. The proposed pipeline integrates high-accuracy optical character recognition, chunking-based processing for long-document handling, Arabic named entity recognition, span refinement, and a retrieval-enhanced attribute extraction module that identifies event dates, locations, and descriptive summaries. Extracted events are consolidated and linked using semantic and temporal similarity measures, and linked through relation classification to construct structured temporal events. Evaluation on a selected part of modern Arabic historical books demonstrates the feasibility of temporal event extraction from long-form Arabic texts, achieving a 75.3% F1-score under dual human verification. Tarikhi represents a step toward scalable temporal knowledge construction for Arabic digital humanities resources.
NAKBA NLP 2026: Shared Task on Arabic Handwritten Manuscript Understanding (Palestine Memory–Omar Al-Saleh Memoir)
Hadi Hamoud | Ahmad Ali Chamseddine | Bilal Shalash | Firas Ben Abid | Mustafa Jarrar | Chadi Abou Chakra | Bernard Ghanem | Fadi A. Zaraket
Hadi Hamoud | Ahmad Ali Chamseddine | Bilal Shalash | Firas Ben Abid | Mustafa Jarrar | Chadi Abou Chakra | Bernard Ghanem | Fadi A. Zaraket
Transcribing historical Arabic manuscripts into machine-readable text is essential for preserving cultural heritage and enabling computational research in the humanities, yet it remains a challenging task due to handwriting variability, page degradation, and the complexity of Arabic script. To advance research in this area, we introduce the NAKBA NLP 2026 shared task on Arabic manuscript understanding, comprising two complementary tracks: a manual transcription track, in which participating teams annotate unlabelled handwritten line images, and an automatic system track for handwritten text recognition (HTR). Both tracks use the Omar Al-Saleh Memoir Collection, a corpus of 6,395 scanned pages and approximately 1.6 million words, written between 1951 and 1965 and provided by the Palestine Memory Project. The dataset, evaluation scripts, and system outputs are publicly available.[https://acr.ps/1L9BaeY] In Subtask 1 (Transcription Track), three teams contributed manual line-level transcriptions; evaluation on hidden ground-truth samples yielded Character Error Rates (CER) between 0.06 and 0.11. In Subtask 2 (Systems Track), seven teams submitted HTR systems. The top-performing system, by Misraj AI, achieved a corpus-level CER of 0.079 and Word Error Rate (WER) of 0.244, outperforming the organiser baseline (CER 0.368, WER 0.691). Rankings shift between corpus-level and per-line evaluation: the 3reeq team achieved the lowest per-line CER (0.082). All contributed transcriptions and system outputs are released under CC-BY-4.0 to support continued research in Arabic manuscript recognition and digital humanities.
StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse
Kholoud Khalil Aldous | Md. Rafiul Biswas | Mabrouka Bessghaier | Shimaa Amer Ibrahim | Kais Attia | Wajdi Zaghouani
Kholoud Khalil Aldous | Md. Rafiul Biswas | Mabrouka Bessghaier | Shimaa Amer Ibrahim | Kais Attia | Wajdi Zaghouani
We present StanceNakba 2026, a shared task on stance detection in polarized social media discourse related to the Palestinian-Israeli conflict, organized as part of Nakba-NLP 2026 at LREC-COLING 2026. The task introduces two subtasks: Subtask A (Actor-Level Stance Detection), which classifies English social media posts as Pro-Palestine, Pro-Israel, or Neutral; and Subtask B (Cross-Topic Stance Detection), which identifies Favor, Against, or Neither stances in Arabic posts toward two conflict-related topics, normalization with Israel and refugee presence in Jordan. The task is grounded in an annotated dataset of 2,606 social media posts. A total of 7 teams participated in Subtask A, and 6 teams in Subtask B. Participating systems primarily fine-tuned Arabic and multilingual transformer-based models, including MARBERT, AraBERT, and DeBERTa-v3 variants, with several teams employing cross-validation, ensemble methods, and topic-conditioned architectures. The best-performing systems achieved a Macro F1 of 0.9620 on Subtask A and 0.8724 on Subtask B, demonstrating that transformer-based approaches are highly effective for conflict-domain stance detection while highlighting persistent challenges in cross-topic generalization and neutral class prediction
The NakbaArchiveClassifier Shared Task on Nakba Image Classification
Alexei Abrahams | Shadi Abudalfa | Mustafa Jarrar | George Mikros
Alexei Abrahams | Shadi Abudalfa | Mustafa Jarrar | George Mikros
The proliferation of social media platforms has significantly reshaped how conflicts are documented, generating large-scale visual records that must be structured to enable meaningful analysis. In this paper, we present the NakbaArchiveClassifier shared task, which focuses on binary classification of infrastructure damage in images from Gaza. This task formed part of the Nakba-NLP Workshop at LREC 2026 and is grounded in an ongoing initiative focused on humanitarian archiving. It utilizes a carefully curated dataset of 2,001 images sourced from Palestinian journalists and content creators on Instagram, spanning the period from October 7, 2023 to December 15, 2025. The objective for participants was to classify whether an image depicts damaged or destroyed infrastructure versus intact structures. This task poses multiple challenges, such as the complexity of real-world conflict imagery, imbalance between classes, and the inherent ambiguity present in many visual scenes. The NakbaArchiveClassifier shared task introduces a new benchmark for analyzing conflict-related visual data and provides valuable resources for advancing research in humanitarian AI, crisis analytics, and Arabic digital humanities.
The NakbaVirality Shared Task on MultimodalVirality Prediction in High-Stakes Discourse
Saad Ezzini | Salima Lamsiyah | Shadi Abudalfa | Samir El-Amrany | Walid Alsafadi
Saad Ezzini | Salima Lamsiyah | Shadi Abudalfa | Samir El-Amrany | Walid Alsafadi
Social media virality significantly shapes public discourse during geopolitical conflicts, where emotionally charged and multimodal content can rapidly gain widespread attention. However, most prior approaches rely on retrospective engagement signals, limiting their usefulness for early prediction. Multimodal virality modeling in high-stakes Arabic discourse remains largely unexplored. We introduce NakbaVirality, a shared task on multimodal virality classification in conflict-related social media posts, organized as part of the Nakba-NLP workshop at LREC 2026. The dataset consists of 2,600 anonymized posts from X and Reddit collected after October 7, 2023, each including text, an associated image, and normalized engagement labels. Participants must classify posts into low, medium, or high virality categories using only textual and visual inputs. The task provides standardized splits, baseline systems, and evaluation using macro-F1 and accuracy. NakbaVirality establishes the first benchmark for multimodal virality prediction in Arabic high-stakes discourse and promotes research on contextual and multimodal modeling for early impact prediction. The shared task attracted 18 participants, who contributed a total of 4 official test phase submissions.
Faisal_Adam at NakbaArchiveClassifier Shared Task: Archival Image Classification for Structural Destruction: A Robust Pipeline Using ResNet-50 and Test-Time Augmentation
Faisal Muhammad Adam | Salisu Aliyu
Faisal Muhammad Adam | Salisu Aliyu
This paper describes our system submission for the Nakba Archive Image Classification task, which requires predicting the presence of structural destruction in historical archival photographs. We framed this as a binary computer vision classification problem (destruction vs. not_destruction). Our system utilizes a pre-trained ResNet-50 convolutional neural network, adapted for binary output, combined with strategic prediction threshold tuning. Evaluated on the unseen final test set, our model achieved a macro F1-score of 0.450 and a balanced accuracy of 0.527, serving as an exploratory baseline that highlights the unique challenges of processing degraded historical imagery.
Doaa Sulaiman at AR-MS NakbaNLP 2026: Faithful Diplomatic Transcription of Arabic Manuscripts Using a Human-Centred Annotation Framework
Doaa Bahjat Sulaiman
Doaa Bahjat Sulaiman
This paper describes my participation in the Human Transcription Track (Subtask 1) of the NAKBA-NLP 2026 Arabic Manuscript Understanding Shared Task, which focuses on historical handwritten documents related to Palestinian Nakba narratives. Participant was asked to manually transcribe approximately 500 cropped line images and to design a comprehensive transcription guideline from scratch. I adopted a faithful diplomatic transcription philosophy that preserves original spelling, punctuation, diacritics, and layout features without editorial normalisation, in order to create research-grade gold-standard data. Building on this philosophy, I developed a 26-convention annotation framework organised into three layers: editorial-structural symbols (11 conventions), faithful-copying rules (12 conventions), and documentation labels (3 types), supported by a four-step quality-control pipeline. My submission achieved full coverage of all 500 assigned lines and attained an official Character Error Rate (CER) of 0.02 and accuracy of 0.98, confirming the high precision of the proposed framework.
KvochurHegel at StanceNakba: Robust Stance Detection with Regularized Natural Language Inference
Minh-Hoang Le
Minh-Hoang Le
Actor-level stance detection over noisy, politically sensitive data can present challenges that standard training procedures fail to handle reliably. This paper presents KvochurHegel, our submission to the StanceNakba 2026 Shared Task, which addresses these challenges by framing stance classification as Natural Language Inference (NLI) to capture actor-level granularity. The official StanceNakba dataset contains high label noise and topic-correlated spurious features, such as texts discussing unrelated global conflicts using in-domain political vocabulary. To handle these conditions within a three-class schema, we construct templates encoding stance hypotheses for specific actors (e.g., “The author expresses support for Palestine”) and introduce a broadened neutral class designed to absorb spurious out-of-domain inputs. A DeBERTa-v3 Cross-Encoder independently evaluates the entailment between the input text and each class-specific hypothesis. Because standard cross-entropy training tends to memorize contradictory annotations under these conditions, we regularize the training procedure with R-Drop and label smoothing. This regularized setup likely contributed to robustness against distribution shifts between the competition’s evaluation phases (the public leaderboard and private test set), allowing our model to improve from a Macro-F1 of 0.9094 to 0.9384 without requiring large generative models, cross-validation, or inference-time ensembling.
KvochurHegel at NakbaArchiveClassifier Shared Task: Nakba Image Classification via ConvNeXt-V2 and Label Smoothing
Minh-Hoang Le
Minh-Hoang Le
This paper presents the KvochurHegel team’s submission to the Nakba Image Classification shared task at the Nakba-NLP 2026 Workshop. The task requires the binary classification of social media images into destruction and not_destruction categories. Given a limited and imbalanced training set of 1,400 images, we utilized a ConvNeXt-V2 Nano backbone combined with extensive data augmentation and label smoothing, prioritizing standard regularization over task-specific architectural modifications. For inference, we applied a 6-view Test-Time Augmentation (TTA) strategy using a hard-voting mechanism. The baseline system achieved a Macro F1-score of 0.8593 and an Accuracy of 0.8706 on the official private test set, ranking 6th out of 16 participating teams.
HCMUS_TheFangs at NakbaArchiveClassifier Shared Task: Foundation Models and Advanced Training Strategies for Conflict Damage Classification
Duy Minh Dao Sy | Trung Kiet Huynh | Nguyen Chi Tran | Phu Quy Nguyen Lam | Phu Hoa Pham
Duy Minh Dao Sy | Trung Kiet Huynh | Nguyen Chi Tran | Phu Quy Nguyen Lam | Phu Hoa Pham
We present our system for the NakbaArchiveClassifier shared task at Nakba-NLP 2026, which requires classifying Instagram images from Gaza as showing destroyed or damaged infrastructure versus intact surroundings. Working with a small, imbalanced dataset (1,400 training images; 1.83:1 class ratio), we conduct a systematic empirical study of six model-training combinations spanning five architecture families: standard CNNs (EfficientNet-B4), self-supervised ViTs (DINOv2-ViT-L), hybrid multi-axis Transformers (MaxViT-Base), masked-image-modelling ViTs (EVA-02-Base), and large-kernel CNNs (UniRepLKNet). For our best performing configuration–MaxViT-Base with focal loss, MixUp, and a rich geometric augmentation pipeline–we provide a detailed component analysis. Our system achieves a macro F1 of 0.899 on the public test set, ranking 1st on the competition leaderboard. We additionally report findings from novel experiments including a Kolmogorov-Arnold Network (KAN) classification head and VLM-regularized training with BLIP-2-generated captions, offering insights into what does and does not transfer to conflict-domain imagery under severe data scarcity.
This paper presents a detailed description of the team’s methodology methodology in participating in Subtask 1 (Transcription Track) of the NAKBA NLP 2026 Shared Task for Arabic Manuscript Understanding. We present a rigorous approach to line-level manual transcription of historical Arabic manuscripts derived from the Omar Al-Saleh memoir collection (1951-1965). Our methodology emphatisez accuracy, consistency, and adherence to diplomatic transcription principles, while addressing the unique palaeographic and physical challenges of Arabic handwriting, such as writing speed, orthographic variation, and the impact of writing tools (e.g., immediate strike-throughs and ink spatter). The work guided by strict transcription guidelines and contextual verification protocols, matching cropped line images with full-page images to resolve ambiguities and automated cropping issues. The team successfully transcribed the entire assigned batch of 500 lines (100% completion rate) across 368 unique pages, producing reference data comprising 6,719 words and 37,646 characters. This effort contributes to providing highly reliable Ground Truth data, serving as an essential foundation for training and evaluating Handwritten Text Recognition (HTR) models for Arabic manuscripts.
Free-Gaza at NakbaArchiveClassifier Shared Task: Towards Distinguishing the Destructive Effect of Nakba: NakbaImage Classification Using Artificial Intelligence Techniques
Nisreen I. R. Yassin | Enas A. Hakim Khalil
Nisreen I. R. Yassin | Enas A. Hakim Khalil
The accounts of the continuing Palestinian Nakba encompass considerable significance. Over the course of the three years of the conflict, millions of photos from social media have been preserved. The preservation and classification of these data through artificial intelligence tools are essential to guarantee their availability, accessibility, and applicability. This paper presents a highly optimized, resource-constrained machine learning pipeline for binary image classification. The system is designed for the NakbaArchiveClassifier Shared Task 2026, which aims to distinguish between destroyed infrastructural images and intact infrastructural images. The system depends on two lightweight EfficientNetB0 networks to build a weighted ensemble system. Using strict hardware limitations of 2GB GPU VRAM, the system achieves an F1-score of 84.16%, which ranked 9th on the leaderboard.
HCMUS_TheFangs at NakbaVirality Shared Task: The Audience is the Message: Escaping the Deep Learning Trap in Conflict-Domain Virality Prediction
Duy Minh Dao Sy
Duy Minh Dao Sy
We present HCMUS_TheFangs’s system for the Nakba-NLP 2026 Virality Shared Task, which achieves Rank #1 on the final leaderboard with a test Macro-F1 of 0.7062, placing first among all competing teams. Our winning system is deliberately simple: a single Community Target Encoding feature - the smoothed historical virality rate of the posting subreddit - combined with TF-IDF text features and an XGBoost classifier. This design emerged from a hard-won insight: virality in conflict reporting is determined not by what is posted but by where it is posted. We spend the majority of this paper showing why this holds. Through 18 ablation experiments we trace the journey from deep learning failure (a cross-attention fusion model scoring 0.2935) to sociological feature engineering. We demonstrate that in conflict domains, deep learning overfits, promotional hashtags negatively correlate with engagement, and visual features are context-dependent modifiers rather than independent signals. Our findings challenge standard practices in multimodal classification and offer a roadmap for predicting virality in highly polarized, community-driven social media environments.
Xin1212 at NakbaVirality Shared Task: Frozen CLIP with Residual Adapter for Multimodal Virality Classification
Xinyan Zhang | Bingzhou Yang
Xinyan Zhang | Bingzhou Yang
We describe our system for the NakbaVirality shared task on multimodal virality classification. Our final approach uses a frozen LAION CLIP backbone, a lightweight residual adapter over fused text–image embeddings, and a small MLP classification head. On our development split, the best configuration (V9) achieves Macro-F1 of 0.5492 and virality-weighted F1 of 0.5252. On the official test submission, our system obtains F1-score 0.4559 and accuracy 0.6089 according to the platform scorer. We provide implementation details, ablations across multiple versions (Baseline–V9), and practical error analysis for reproducibility.
A2NLP at StanceNakba Shared Task: Fine-Tuned AraBERT for Topic-Based Arabic Stance Detection
Alaa Nairat | Aysar Mahmoud Nairat
Alaa Nairat | Aysar Mahmoud Nairat
AbstractThis paper describes A2NLP’s system for Subtask B of the StanceNakba Shared Task, which addresses cross-topic Arabic stance detection. The goal is to classify sentence–topic pairs into pro, against, or neutral labels. We introduce a topic-conditioned prompting strategy built on AraBERTv0.2-Twitter, where each instance is reformulated into a structured prompt that explicitly models the interaction between the sentence and its target topic. The model is trained using 5-fold stratified cross-validation with class-weighted loss to ensure robustness under mild label imbalance. Our final submission achieves a Macro-F1 score of 0.8483 on the official test set, outperforming the AraBERTv2 baseline (0.810) and ranking fifth overall. Ablation analysis confirms that topic-conditioned prompting substantially improves generalization across topics. The findings demonstrate the importance of structured input design and domain-aligned pretraining for reliable stance detection in dialectal Arabic social media discourse.
Ketaba-OCR at AR-MS NakbaNLP 2026: Efficient Adaptation of Vision-Language Models for Handwritten Recognition
Hassan Barmandah | Fatimah Emad Eldin | Khloud Al Jallad | Omer Nacar
Hassan Barmandah | Fatimah Emad Eldin | Khloud Al Jallad | Omer Nacar
This paper presents Ketaba-OCR-LoRA, a system developed for the NakbaNLP 2026 Shared Task on Arabic Manuscript Understanding (Subtask 2), which targets the transcription of the historically significant Omar Al-Saleh Memoir Collection written in Ruq’ah and Naskh scripts. We propose a parameter-efficient adaptation of a publicly available pretrained Arabic-English Handwritten Text Recognition (HRT) model, originally trained on handwritten corpora including the Muharaf dataset. Instead of adapting general Vision-Language Models from scratch, we fine-tune the HRT backbone using Low-Rank Adaptation (LoRA) and 4-bit quantization (QLoRA), reducing memory requirements from 40GB to approximately 8GB. Our final submission combines multiple model variants through a novel Linear+Boost weighted ensemble strategy. Our approach achieves a CER of 0.0819 and WER of 0.2588 on the blind test set (per-line evaluation), ranking 1st on per-line evaluation; on the official corpus-wide leaderboard, we rank 3rd (CER 0.0938, WER 0.2996). This work demonstrates that specialized pretrained HRT models substantially outperform general-purpose Vision-Language Models for Arabic manuscript transcription, and that parameter-efficient fine-tuning provides a practical and reproducible approach for low-resource cultural heritage digitization.
No Overfit at NakbaArchiveClassifier Shared Task: A Swin Transformer-Based System for Destruction Image Classification
Mohamed Fathy Mohamed | Samar Mahmoud Abd El-Mageed | Ensaf Mohamed
Mohamed Fathy Mohamed | Samar Mahmoud Abd El-Mageed | Ensaf Mohamed
Automated destruction identification from visual data plays a critical role in large-scale documentation, humanitarian analysis, and digital archiving of conflict-related events. Within this context, the Nakba-NLP 2026 Workshop introduced a shared task aimed at training and evaluating a binary image classification model to distinguish between destroyed or damaged infrastructure and intact infrastructure. However, the limited dataset size and the visual variability of real-world scenes make this task particularly challenging. This work presents a Swin Transformer–based framework tailored for destruction image classification. The proposed model employs a hierarchical Swin Transformer backbone for robust feature extraction, followed by a multi-layer perceptron classifier for decision-making. To address the limited data issue, transfer learning and a customized training strategy are applied to adapt the model effectively without full end-to-end retraining. Furthermore, a semi-supervised data expansion approach is utilized to enlarge the training set from 1,400 to 10,000 images, improving model generalization and robustness. Experimental results on the official blind test set demonstrate strong performance, achieving an F1-score of 86.55% and an accuracy of 87.81%, ranking 5th in the shared task.
Yafa at StanceNakba: Actor-Level Stance Detection Using Cross-Lingual Approach
Tasnim Zayet | Osama Hamed | Tasneem Duridi
Tasnim Zayet | Osama Hamed | Tasneem Duridi
This paper addresses the problem of actor-level stance detection in English social media posts concerning the Palestinian issue, a subtask of the StanceNakba-2026 Shared Task. The objective is to classify posts into one of three categories: Pro-Palestine, Pro-Israel, or Neutral, which is more challenging than the traditional favor/against/neutral formulations. This study uses a dataset comprising 1,401 posts, collected from X (formerly Twitter) after October 7, 2023, and annotated with one of the three stance labels. As Yafa’s Team, we tried to solve this problem using BERT-based models, which have proven their superiority in similar tasks. Several BERT-based models were fine-tuned and compared, including ARBERT, MARBERT, and PoliBERTweet, among others. Our winning model is the “MARBERT-Y”, where the “Y” comes from Yafa, a MARBERT-based model that has achieved a macro-F1 score of 95% on the test set. We argue this to two main factors: the structured and harsh preprocessing steps applied and the fine-tuning process employed. This indicates that domain-adapted transformer models, i.e., those pretrained on large-scale Twitter data are highly effective for politically stance detection tasks.
The Resistant Word at StanceNakba Shared Task: A Topic-Aware Model for Cross-Topic Stance Detection
Mohamed Bahgat | Doaa Salah | Sarah Yassine
Mohamed Bahgat | Doaa Salah | Sarah Yassine
Cross-topic stance detection in Arabic is the task of identifying whether a text expresses a pro, against, or neutral position toward a given issue, and it is particularly challenging under topic shifts and class imbalance. In Subtask B of the StanceNakba 2026 shared task on Arabic cross-topic stance detection, we are given a Levantine Arabic sentence and one of two topics: “Normalization with Israel” or “Refugee/Immigrant Presence in Jordan,” and we must classify the expressed stance. A central difficulty is the systematic failure of standard fine-tuning to recognize the minority neutral class, driven by majority-class dominance in cross-entropy training and accuracy-based checkpoint selection. To address this, we combine random oversampling with class-weighted cross-entropy loss, and we build an ensemble of four Arabic pre-trained transformers MARBERT, AraBERT Large, XLM-RoBERTa Base, and CAMeL-BERT Mix each trained using Stratified 5-Fold cross-validation. Our final system achieves a macro-F1 of 0.9777 and an accuracy of 97.79% on the evaluation set.
"Hope" at NakbaArchiveClassifier Shared Task: Transfer Learning-Based CNN Models for Infrastructure Damage Detection
Lojien AlKhidir | HebaTalla Abdelhady
Lojien AlKhidir | HebaTalla Abdelhady
This paper describes Team Hope’s system for the NakbaArchiveClassifier Shared Task at Nakba-NLP 2026. The task focuses on binary classification of social media images into two categories: destruction and not_destruction. We evaluated multiple convolutional neural network architectures using transfer learning, including ResNet34, ResNet50, EfficientNet-B0, and a fine-tuned ResNet34 variant with staged training. All models were initialized with ImageNet pretrained weights and fine-tuned on the provided dataset of 2,001 images. The dataset is moderately imbalanced and contains visually diverse Instagram images depicting intact and damaged infrastructure. Our best-performing model, ResNet34 trained for 25 epochs with Adam optimizer and a learning rate of 1e-4, achieved 81% accuracy on the evaluation platform. We provide a comparative analysis of the tested architectures and discuss the impact of model depth, training duration, and class imbalance. Given the political and ethical sensitivity of the dataset, we also include a discussion of responsible AI considerations and potential limitations. Our findings suggest that moderate-depth architectures can generalize effectively in low-resource, contextually complex visual classification tasks.
Al-Warraq at AR-MS NAKBA-NLP 2026: Adapting Vision-Language and Transformer Models for Automatic Manuscript OCR/HTR
Ahmad Edris Youssef | Aya Hafiz Faris | Alhasan Hamood | Zainab Kamil | Jana Alqasem | SARA Ali hamed"
Ahmad Edris Youssef | Aya Hafiz Faris | Alhasan Hamood | Zainab Kamil | Jana Alqasem | SARA Ali hamed"
We present our submission to the NAKBA NLP 2026 Automatic Manuscript OCR/HTR shared task on Arabic manuscripts. The task aims to transcribe manuscript line images into machine-readable Arabic text. Our approach followed an iterative pipeline including model selection, training, error analysis, test-time augmentation, and postprocessing. After evaluating several OCR/HTR models, we selected and trained the most suitable model on the provided manuscript line images and transcriptions. Error analysis showed better character-level performance than word-level performance, which motivated the use of test-time augmentation and text cleaning to improve robustness. The final system achieved a CER of 0.1142 and a WER of 0.378, placing fifth in the shared task. These results show that simple but targeted improvements can support effective Arabic manuscript transcription.
Misraj AI at AR-MS NAKBA-NLP 2026: A State-of-the-Art VLM in Arabic Handwritten Text Recognition
Khalil Hennara | Muhammad Hreden | Zeina Aldallal | Sara Chrouf | Safwan AlModhayan
Khalil Hennara | Muhammad Hreden | Zeina Aldallal | Sara Chrouf | Safwan AlModhayan
Handwritten Text Recognition (HTR) for Arabic presents unique challenges due to the script’s cursive nature, varying writer styles, and morphological complexity. While modern Vision-Language Models (VLMs) have significantly advanced document parsing, their direct application to highly specific cursive domains requires strategic adaptation. This paper details our submission to the Nakba OCR competition, which adapts a 3B-parameter VLM to recognize historical Arabic manuscripts. We employ a progressive training pipeline that utilizes domain-matched data augmentation to bridge the gap between standard printed Arabic OCR and historical handwritten manuscripts. Moving beyond standard decoder-only Supervised Fine-Tuning (SFT), we fine-tune the entire encoder-decoder architecture using differential learning rates. This approach, followed by a final checkpoint merge, allows the model to better resolve the fine visual details of cursive Arabic script. Our final unified model (submitted under the team name Misraj AI) establishes a new state-of-the-art (SOTA) on the Nakba dataset, achieving a Word Er- ror Rate (WER) of 0.24 and a Character Error Rate (CER) of 0.08, and officially securing first place on the leaderboard.
PushingBoundaries at NakbaVirality Shared Task: Recursive Prompt Improvement for Multimodal Virality Classification
Ashhadul Islam | Md Rafiul Biswas | Samir Brahim Belhaouari | Wajdi Zaghouani
Ashhadul Islam | Md Rafiul Biswas | Samir Brahim Belhaouari | Wajdi Zaghouani
This paper describes our participation in the NakbaVirality shared task at the NakbaNLP Workshop (LREC–COLING 2026). We investigate Recursive Prompt Improvement (RPI), an instruction-level optimization strategy for virality classification in high-stakes geopolitical discourse. In this work, we propose a self-supervised approach to iteratively improve the classification prompt without human intervention. We begin with a basic prompt that guides the LLM to perform multi-class classification, incorporating contextual information about the tweets. After obtaining predictions, we identify misclassified tweets and feed them back to the model with an instruction to refine and improve the original classification prompt. This process is repeated over multiple iterations to assess whether performance improves over time. Our results show a remarkable improvement in F1 score from the first iteration to the final one. Although the proposed method does not reach the accuracy of models fine-tuned directly on task-specific data, it demonstrates that iterative, self-supervised prompt refinement can serve as a viable proxy for fine-tuning. By leveraging the model’s own errors as feedback, this approach reduces reliance on computationally expensive training procedures and heavy GPU usage, while preserving much of the adaptability typically associated with fine-tuned models. This paradigm opens promising avenues for resource-efficient model adaptation and suggests new directions for scalable, low-cost performance improvement without traditional fine-tuning. The code has been shared in Github.
U4RASD at StanceNakba Shared Task: Data Augmentation and Auxiliary Objectives for Arabic Stance Detection
Nancy Hamdan | Aya Jouni | Aya Saïd | Fadi Zaraket
Nancy Hamdan | Aya Jouni | Aya Saïd | Fadi Zaraket
This paper describes a submission to Track B of the StanceNakba Shared Task on Arabic cross-topic stance detection in the political domain. We investigate LLM-based data augmentation, auxiliary training objectives including contrastive and multi-task learning, zero-shot prompting, and a preliminary terminology-based clustering approach. Our final system, based on MARBERTv2 with dialect-aware LLM-based augmentation, achieved 86% macro-F1 on the blind test set and ranked 3rd out of 10 teams. Our results show that dialect-aware augmentation substantially improved performance in a low-resource Arabic stance detection setting, while not all auxiliary objectives or clustering-based strategies yielded consistent gains. We release our code at https://acr.ps/1L9B9Tw.
AyahVerse at NakbaArchiveClassifier Shared Task: Architectural Trade-offs and Decision Calibration for Humanitarian Image Classification
Ibad-ur-Rehman Rashid | Akhtar Ali
Ibad-ur-Rehman Rashid | Akhtar Ali
This paper presents our submission to the Nakba-NLP 2026 Shared Task on binary image classification, where the goal is to categorize images of Gaza infrastructure as destroyed or intact. To address the challenges of class imbalance and resource-constrained deployment, we evaluated three convolutional architectures: ResNet50, MobileNetV2, and EfficientNet-B0, combined with a post-hoc threshold optimization step. Our results show that lightweight architectures are competitive with heavier models for this task, with EfficientNet-B0 achieving the highest Test F1-score of 0.85 despite having significantly fewer parameters than ResNet50. We further investigated the effect of input resolution, finding that increasing resolution improved ResNet50’s performance, though it remained below lightweight alternatives. Finally, we demonstrate that shifting the binary decision threshold from the default 0.50 to an optimized 0.45 improved ResNet50’s Test F1 from 0.79 to 0.81 by recovering recall for the minority destroyed class. Notably, this adjustment was only needed for ResNet50, while EfficientNet-B0 and MobileNetV2 performed best at the default 0.50, suggesting that larger models are more prone to majority-class bias. Overall, these results provide a systematic analysis of architectural efficiency and threshold behavior under class imbalance, offering practical insights for damage classification in resource-constrained crisis settings.
mlenthusiast at NakbaArchiveClassifier Shared Task: A Lightweight SVM-Gated Ensemble of EfficientNets for Image Classification
Md. Ajwad Hossain
Md. Ajwad Hossain
Image classification under strict time constraints requires a delicate balance between feature complexity and computational overhead. This paper presents an optimized ensemble methodology developed for the NAKABA competition, focusing on identifying structural destruction. We propose a hybrid architecture that leverages two distinct Convolutional Neural Networks (EfficientNetB0 and EfficientNetB3) as base feature extractors, coupled with a Support Vector Machine (SVM) functioning as a meta-classifier. Instead of standard probability averaging or processing high-dimensional embeddings directly, the Meta-SVM acts as a learned gating mechanism to optimally combine the low-dimensional probability predictions of the base models. This ensures robust performance without the latency of heavier deep learning architectures. Empirical results demonstrate the efficacy of this approach. The model achieved a validation accuracy of 0.884 and a weighted F1-score of 0.885, with a notable F1-score of 0.839 on the challenging ’destruction’ class. On the official NAKABA leaderboard test set, the ensemble maintained strong generalization, achieving an F1-score of 0.831 and an accuracy of 0.845, which secured the 12th position overall and proved the model’s high effectiveness within the competition’s strict operational constraints.
Pixel at NakbaArchiveClassifier Shared Task: ConvNeXt-Based Ensemble for Destruction Detection
Rahaf Jaber
Rahaf Jaber
This paper describes our submission to the Nakba Image Classification Shared Task at the Nakba-NLP 2026 workshop. The task requires binary classification of social media images into two categories: destruction and not_destruction. The dataset includes approximately 1,600 annotated development images and 400 held-out test images, collected from Instagram posts published in Gaza between October 2023 and December 2025. High variability in viewpoint, lighting, and image quality, coupled with the inherent complexities of identifying structural damage in dense urban environments, makes this task particularly challenging. Our system utilizes a pretrained ConvNeXt-Tiny backbone fine-tuned through a stratified 5-fold cross-validation framework. To mitigate class imbalance, we implement a weighted cross-entropy loss function. During the inference phase, we employ an ensemble strategy that averages predictions across all five fold-specific models, and test-time augmentation (TTA) is applied to enhance robustness. The final ensemble achieved a Macro F1-score of 0.8952 and an accuracy of 0.9055 on the official test set. Our results suggest that the integration of modern convolutional architectures with robust ensembling and augmentation strategies provides a reliable baseline for automated destruction detection.
MennaAly at NakbaArchiveClassifier Shared Task: Transfer Learning with ResNet for Historical Image Classification
Menna Aly
Menna Aly
This paper describes our submission to the NakbaArchiveClassifier shared task at Nakba-NLP 2026, co-located with LREC 2026. The task consists of binary image classification, where a model must classify historical images into one of two categories: destruction or not_destruction. We adopt a transfer learning approach based on pretrained residual networks, fine-tuned on the provided training data. To mitigate class imbalance, we incorporate weighted cross-entropy loss during optimization. In the development phase, our ResNet18 model achieved a peak macro F1-score of 0.8137 on the validation set. For the final phase, we trained on the combined training and validation data (1,599 labeled images) and generated predictions for the hidden test set of 402 images. Our final submission achieved a macro F1-score of 0.83228 with an accuracy of 0.84577 on the official evaluation set. These results underscore the effectiveness of lightweight transfer learning approaches for historical image analysis under limited-data conditions, demonstrating that compact residual architectures can achieve competitive performance without complex architectural modifications.
DLRG@ NakbaArchiveClassifier Shared Task: Deep Transfer Learning for Destruction Detection in Nakba Archive Images Using EfficientNet-B3
Ramesh R. Kannan | Ratnavel Rajalakshmi
Ramesh R. Kannan | Ratnavel Rajalakshmi
Automatic identification of destruction in conflict-affected regions is an important task for humanitarian monitoring and historical documentation. Visual analysis of destruction scenes can assist researchers and policy makers in understanding the extent of damage in affected areas. This paper presents a deep learning-based image classification approach for identifying destruction and non-destruction scenes in Nakba-related images. The problem is formulated as binary image classification on Nakba images. A transfer learning approach using EfficientNet-B3 is adopted to learn discriminative visual features from Nakba images. Experimental evaluation shows that the proposed model achieved an Weighted F1-score of 83.87 % and an overall classification accuracy of 85.57 % and secured 10th rank in the competition. The results demonstrate that our proposed pre-trained method can effectively capture structural damage patterns and visual cues associated with destruction scenes. Code: https://github.com/kannanrrk/NakbaImageClassifier
Oblevit at AR-MS NAKBA NLP 2026 Subtask 2: Hybrid CNN–BiLSTM–CTC Framework with Linguistic Refinement for Arabic Handwritten Manuscript Recognition
Reem Juhaysh | Abuelgasim Sami Abusonoun | Sara Ayad
Reem Juhaysh | Abuelgasim Sami Abusonoun | Sara Ayad
Arabic handwritten manuscript recognition is challenging due to the cursive nature of the script, dot ambiguity, and document degradation. In this work, we propose an end-to-end OCR system based on a CNN–BiLSTM–CTC architecture. The model extracts visual features, captures sequential dependencies, and performs alignment-free training. Arabic-specific decoding and post-processing techniques are applied to reduce character and spacing errors. Experimental results show competitive performance in recognizing complex handwritten Arabic text.
Viva_Palestine at StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse
Wafaa El-Kassas | Enas Khalil | Enas El Houby
Wafaa El-Kassas | Enas Khalil | Enas El Houby
Recent research has increasingly focused on user-generated content to clarify opinions expressed in social media discourse. The Actor and Topic-Aware Stance Detection in Public Discourse challenge encourages research on stance detection in polarized social media discourse on the Palestinian–Israeli conflict. The challenge comprises two subtasks: one for actor-level alignments and the other for cross-topic generalization patterns. The StanceNakba2026 task includes two subtasks: (A) Actor-Level Stance Detection in English and (B) Cross-Topic Stance Detection in Arabic. Our team participated in both subtasks with the name “Viva_Palestine”. In Subtask A, the proposed method is based on the Bert-Base-Uncased model and achieved a Macro F1-score of 0.9190, placing 6th out of 13 teams. In Subtask B, the proposed method is based on the MARBERT model and achieved a Macro F1-score of 0.8724 (the top rank in the leaderboard), placing first out of 10 teams. These results show that the proposed modelling method performs well for both entity-specific stance alignment and strong cross-topic generalization.
KUET at StanceNakba Shared Task: StanceMoE: Mixture-of-Experts Architecture for Stance Detection
Abdullah Al Shafi | Md. Milon Islam | Sk. Imran Hossain | K. M. Azharul Hasan
Abdullah Al Shafi | Md. Milon Islam | Sk. Imran Hossain | K. M. Azharul Hasan
Actor-level stance detection aims to determine an author’s expressed position toward specific geopolitical actors mentioned or implicated in a text. Although transformer-based models have achieved relatively good performance in stance classification, they typically rely on unified representations that may not sufficiently capture heterogeneous linguistic signals, such as contrastive discourse structures, framing cues, and salient lexical indicators. This motivates the need for adaptive architectures that explicitly model diverse stance-expressive patterns. In this paper, we propose StanceMoE, a context-enhanced Mixture-of-Experts (MoE) architecture built upon a fine-tuned BERT encoder for actor-level stance detection. Our model integrates six expert modules designed to capture complementary linguistic signals, including global semantic orientation, salient lexical cues, clause-level focus, phrase-level patterns, framing indicators, and contrast-driven discourse shifts. A context-aware gating mechanism dynamically weights expert contributions, enabling adaptive routing based on input characteristics. Experiments are conducted on the StanceNakba 2026 Subtask A dataset, comprising 1,401 annotated English texts where the target actor is implicit in the text. StanceMoE achieves a macro-F1 score of 94.26%, outperforming traditional baselines, and alternative BERT-based variants.
PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection
Md. Shakhoyat Rahman Shujon | MD Jahid Hasan Jim | Md. Milon Islam | Md Rezwanul Haque | Fakhri Karray
Md. Shakhoyat Rahman Shujon | MD Jahid Hasan Jim | Md. Milon Islam | Md Rezwanul Haque | Fakhri Karray
We introduce PAST-TIDE, our stance detection system addressing both subtasks of the StanceNakba Shared Task at NakbaNLP@LREC-COLING 2026. The main idea is statement tuning. We redefine stance as cloze-style masked language modeling (MLM), letting a verbalizer map label words to stance categories through the pre-trained MLM head rather than appending a randomly initialized classification head. We complement this with prototypical contrastive learning, which uses learnable class prototypes for batch-size independent contrastive training, and topic-conditional layer normalization for cross-topic Arabic stance detection. PAST-TIDE achieves macro-F1 scores of 0.75 for Subtask A and 0.74 for Subtask B on the official leaderboard, indicating that minimal architectural additions to a pre-trained model can remain competitive in low-resource settings.
AlSaifTeam at AR-MS NAKBA-NLP 2026: Building Expert-Quality Ground Truth for Arabic Handwritten Manuscripts
Joud Fahad AlSaif | Alhasan Hamood Mohammed | Jana Mohammad Alseed
Joud Fahad AlSaif | Alhasan Hamood Mohammed | Jana Mohammad Alseed
This paper describes our participation in Subtask 1 of the NAKBA NLP 2026 Arabic Manuscript Understanding Shared Task, which focuses on the manual creation of expert-quality, line-level transcriptions for Arabic handwritten manuscripts. To ensure reliable ground truth, we adopt a protocol-driven methodology based on fixed transcription rules, collaborative verification, and confidence-based quality control. The proposed approach aims to improve consistency, reduce annotation bias, and support the creation of trustworthy benchmark resources for future Arabic OCR and HTR research. Keywords:Arabic handwritten manuscripts, ground truth construction, manual transcription, handwritten text recognition, optical character recognition, benchmark enrichment
This paper describes the guidelines that the PalNLP team recursively developed and followed during the transcription of the assigned batch, as part of the AR-MS Shared Task. The team, which consists of a single experienced transcriber, has manually transcribed 500 images of lines from the Omar Al-Saleh Memoir Collection.
Digilians at NakbaVirality Shared Task: Bidirectional Cross-Attention for Multimodal Virality Prediction
Ahmed Eid Hassan | Noureldeen H. Mohamed | Abdelrhman M. Fawzy | Mohamed A. Abdelghany | Ahmed A. Hassan | Ahmed S. Qassim | Fady A. Abd El Sayed | Mohamed H. Mohamed | Rahma M. Mohamed | Arwa M. Abou-Attia | Mayar M. Mohamed | Shahd A. Sawla | Shahd O. Mahmoud
Ahmed Eid Hassan | Noureldeen H. Mohamed | Abdelrhman M. Fawzy | Mohamed A. Abdelghany | Ahmed A. Hassan | Ahmed S. Qassim | Fady A. Abd El Sayed | Mohamed H. Mohamed | Rahma M. Mohamed | Arwa M. Abou-Attia | Mayar M. Mohamed | Shahd A. Sawla | Shahd O. Mahmoud
The NakbaVirality shared task focuses on multimodal virality prediction using a dataset of 2,600 multilingual posts collected from X and Reddit. In this work, we propose a multimodal architecture that combines XLM-RoBERTa for text encoding and a Vision Transformer (ViT) for image representation. The extracted features are aligned through bidirectional cross-attention to capture interactions between textual and visual modalities. To address the class imbalance present in the dataset, we apply focal loss, class weighting, and targeted data augmentation for the minority class. Additionally, layer-wise learning rate scheduling is used to stabilize fine-tuning of the pretrained encoders. Experimental results show that the proposed system achieves an accuracy of 0.6009 on the hidden test set, ranking 4th among 29 participating teams (107 total submissions). These results highlight the effectiveness of cross-modal attention mechanisms for modeling multimodal signals in high-stakes discourse.
Not Gemma at AR-MS NakbaNLP 2026: Mubsir OCR: End-to-End Recognition of Arabic Handwritten Text
Ali Adel Ali | Mona Khaled Ali | Mohamed Emad Sayed | Ibrahim Naser Mostafa
Ali Adel Ali | Mona Khaled Ali | Mohamed Emad Sayed | Ibrahim Naser Mostafa
Historical Arabic handwritten OCR is difficult because of cursive script, fine diacritics, mixed numerals, and degraded media; classical segmentation pipelines compound errors, whereas end-to-end vision-language models can adapt when fine-tuned on in-domain data. We present Mubsir OCR, a systematic evaluation on the NAKBA dataset: an annotated set (15,962 training line crops and 2,095 val lines with ground truth, used for all nine experiments) and a separate blind AR-MS (Subtask 2) set (2,671 images; scores only via official submission). We compare external vs. in-house VLMs (Qwen2.5-VL 3B, Qwen3-VL-4B-Instruct, Gemma3), inference backends (vLLM/bf16 vs. HuggingFace/bf16), training length (16 vs. 32 epochs), and test-time preprocessing (CLAHE+unsharp). Best on the annotated val set: 8.59% CER / 25.87% WER (HuggingFace bf16); the same configuration attains 11.00% CER / 31.26% WER on the blind set. Domain-specific fine-tuning beats general-purpose checkpoints; preprocessing helps only marginally and is not recommended without train-time augmentation.
Latent Narratives at AR-MS NakbaNLP 2026: Reducing Character Errors in Arabic Manuscript Transcription: A CER Oriented System
Sara Abdulmonem Al desouky
Sara Abdulmonem Al desouky
Historic Arabic handwritten texts present significant challenges due to varied handwriting styles, cursive structure, diverse diacritics, and inconsistent character and word sizes. In this work, we introduce Historic-Arabic-OCR, a vision-language OCR system built upon Qari-OCR, which itself is based on Qwen2-VL-2B-Instruct, and further fine- tuned using Low-Rank Adaptation (LoRA) for Arabic manuscript transcription. The proposed approach incorporates contrast enhancement using CLAHE and deterministic decoding strategies to reduce character-level errors. Our model achieves competitive performance, with a Word Error Rate (WER) of 0.28 and a Character Error Rate (CER) of 0.10 on historical Arabic texts, including low-resolution images. The final submitted system uses CLAHE prepro- cessing with deterministic greedy decoding to minimize character-level errors. Keywords: Arabic OCR, Vision-Language Models, Qwen2-VL, LoRA, CER Optimization
shroukgbr at StanceNakba Shared Task: Transformer-Based Ensemble Learning for Actor-Level Stance Detection in Palestinian–Israeli Social Media Discourse
Shrouk Anwar Gbr | Mohamed Ibrahim Ragab
Shrouk Anwar Gbr | Mohamed Ibrahim Ragab
Stance detection has become an essential task for understanding political discourse on social media, particularly in highly polarized contexts where sentiment alone is insufficient to capture author intent. This study addresses stance classification in discussions related to the Palestinian–Israeli conflict by developing transformer-based and ensemble learning approaches for three-class classification: Pr-Palestine, Pro-Israel, and Neutral. Using the StanceNakba 2026 Shared Task dataset, we fine-tune multiple pretrained transformer models, including MARBERT, ARBERT, BERT, RoBERTa, and DeBERTa, and evaluate their performance using stratified cross-validation with macro F1-score as the primary metric. In addition to individual model evaluation, a weighted ensemble combining BERT, RoBERTa, and DeBERTa is proposed to leverage complementary contextual representations. Experimental results show that the ensemble model achieves the best performance with an accuracy and macro F1-score of 0.8905, outperforming specialized Arabic models while maintaining strong class-wise balance. The proposed approach achieved first place on the Codabench leaderboard in both the development and final evaluation phases of the shared task, demonstrating its robustness and effectiveness in real-world stance detection settings.
NU_Hallucinators at NakbaArchiveClassifier Shared Task: A CLIP-Based Approach for Destruction Detection in Historical Image Archives
Salma Khaled Hegazy | Mohamed Ibrahim Ragab
Salma Khaled Hegazy | Mohamed Ibrahim Ragab
This paper presents a CLIP-based transfer learning approach for classifying historical archive images in the Nakba Image Classification Shared Task at the Nakba-NLP 2026 Workshop (LREC 2026). The task involves distinguishing images depicting destroyed or damaged infrastructure from those showing intact scenes using a dataset of 2,001 images collected from Instagram posts published by Palestinian content creators and journalists in Gaza between October 2023 and December 2025. Our method employs the CLIP ViT-B/32 visual encoder with selective fine-tuning of the final transformer block and a lightweight classification head. To address class imbalance, we apply focal loss along with standard data augmentation and threshold optimization. Experimental results show that the proposed model outperforms several CNN baselines and achieves an F1-score of 0.877 on the blind test set, securing 4th place in the shared task.
ZAHIRA BOULANOUAR at NakbaArchiveClassifier Shared Task: Detecting Infrastructure Destruction in Gaza with a ConvNeXt Ensemble
Zahira Boulanouar
Zahira Boulanouar
We present our third-place submission to the Nakba Image Classification Shared Task at LREC-COLING 2026, which requires binary classification of Instagram images from Gaza into destruction (damaged or destroyed infrastructure) versus not_destruction. Our system fine-tunes a ConvNeXt-Tiny backbone within a five-fold stratified cross-validation framework, combining Focal Loss, weighted random sampling, exponential moving average (EMA) weight stabilization, test-time augmentation (TTA), and out-of-fold (OOF) decision threshold calibration. Our system achieves an official test macro F1 of 0.8893 and 90.05% accuracy, placing third among all participants and within 0.02 F1 of the winning system (0.91), demonstrating that a 28M-parameter convolutional architecture with principled training strategies is highly competitive with much larger models.
EGCSS at StanceNakba Shared Task: Cross-Topic Arabic Stance Detection for Two Middle East Issues
Asmaa Qindeel | Toka Khaled | Batool Najeh Balah | Eman Elrefai | Mahmoud Fawzi
Asmaa Qindeel | Toka Khaled | Batool Najeh Balah | Eman Elrefai | Mahmoud Fawzi
Stance detection continues to be an important task sitting at the intersection of Natural Language Processing (NLP) and Computational Social Science (CSS). In this work, we evaluate how different variations of BERT models perform on the cross-topic form of the task. In particular, we inspect their performance on the second subtask of the shared task StanceNakba 2026, where two topics are included, namely Arab Normalization with Israel and The Presence of Refugees in Arab Countries. We find that the best-performing model was bert-base-arabertv02-twitter, and we further improve its performance by providing context about the topic during the training phase, achieving an F1-score of 0.86 and ranking second among the participating teams.
up
Proceedings of the Workshop Neology and Large Language Models
Proceedings of the Workshop Neology and Large Language Models
Giedre Valunaite Oleskeviciene | Voula Giouli | Florentina Armaselu | Chaya Liebeskind | Barbara McGillivray
Giedre Valunaite Oleskeviciene | Voula Giouli | Florentina Armaselu | Chaya Liebeskind | Barbara McGillivray
From 124 Million Tokens to 1,021 Neologisms: A Large-Scale Pipeline for Automatic Neologism Detection
Diego Rossini | Lonneke van der Plas
Diego Rossini | Lonneke van der Plas
We present a scalable, modular pipeline for automatic neologism detection that combines rule-based filtering with LLM classification. The pipeline is grounded in two complementary word-formation frameworks, grammatical and extra-grammatical morphology, which jointly define the scope of what counts as a neologism and inform a four-class classification scheme (NEOLOGISM, ENTITY, FOREIGN, NONE). While designed to be modular and transferable at the architectural level, the pipeline is instantiated on 527 million English-language Reddit posts spanning 2005–2024. From this corpus, we extract 124.6 million unique tokens and reduce them by over 99.99% to yield 1,021 neologism candidates, a set small enough for manual expert verification. Multiple LLMs independently classify each candidate via majority vote, with a final verification step, revealing substantial cross-model disagreement and highlighting the challenge of operationalizing neologism detection at scale. Manual annotation of all 1,021 candidates confirms that 599 (58.7%) are genuine lexical innovations.
Large language models (LLMs) are increasingly deployed to detect, generate, and normalize neologisms across languages. While prior work has examined their capacity to model semantic change and handle temporal drift, insufficient attention has been paid to how training-data asymmetries interact with probabilistic generation mechanisms to structure lexical innovation itself. This paper argues that AI-driven neology is shaped by systematic high-resource bias that privileges dominant languages in the production, stabilization, and dissemination of new lexical items. Drawing on sociolinguistics, language political economy, lexicography, and computational modeling theory, we formalize how distributional imbalance alters innovation likelihood across languages. We introduce a taxonomy of bias types specific to AI-mediated neology, present a probabilistic account of generative reinforcement loops, and illustrate these mechanisms using documented examples from English-Arabic and English-Icelandic language pairs. We derive empirically testable predictions and propose concrete mitigation strategies for lexicographers, language planners, and NLP researchers.
Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models
Nina Hosseini-Kivanani
Nina Hosseini-Kivanani
Large language models (LLMs) are increasingly used for writing assistance in small contact languages, yet it is unclear whether they respect community norms around lexical borrowing and neology. We introduce LexNeo-Bench, a 3,050-instance token-level benchmark derived from LuxBorrow, a large-scale Luxembourgish news corpus, where target tokens are labelled as native or as French, German, or English borrowings. Using this benchmark, we probe three multilingual LLMs across 34 prompt settings on two tasks: borrowing type classification and a binary lexical-innovation proxy (borrowing versus native). Without external context, models perform only slightly above chance on borrowing classification, so we construct a linguistic knowledge graph that encodes donor language, morphological patterns, and lexical analogues, and inject instance-specific subgraphs into the prompt. Knowledge-graph prompts raise borrowing classification accuracy from 25 – 35% up to 71 – 81% and largely close the gap between small and large models, while leaving neology detection difficult and sensitive to few-shot design. Our results show that lexicon-aware prompting is highly beneficial for robust borrowing judgments in low-resource contact languages and that lexical resources can serve as structured context for LLM evaluation. This study was carried out within the ENEOLI COST Action and examines borrowing as a form of lexical innovation in multilingual Luxembourgish data.
Lexical Innovation in Business Colour Idioms: Evidence from Large Language Models in Five Languages
Giedre Valunaite Oleskeviciene | Ágnes Abuczki | Ganit Richter | Berat Ujkani | Vera Moitinho de Almeida | Pedro Madeira
Giedre Valunaite Oleskeviciene | Ágnes Abuczki | Ganit Richter | Berat Ujkani | Vera Moitinho de Almeida | Pedro Madeira
Lexical innovation refers to the process of creating new lexical items, enabling languages to adapt to evolving socio-cultural and material realities. The domains of business, economics, and finance are among the most productive ones of lexical innovation. The present research study lies at the intersection of lexical innovation, idiomaticity, and large language model (henceforth, LLM) research and investigates lexical productivity, semantic shift, and globalization (Anglocentric changes) in business-related colour idioms by comparing human translation and annotation with the output of LLMs. The current experiment involves an initial study carried out for five languages: English (the pivotal one), Albanian (AL), Hebrew (HE), Hungarian (HU), Lithuanian (LT), and Standard European Portuguese (PT). The research results reveal that LLMs show high mutual agreement, but the agreement with humans is lower. The internal consistency of LLMs reflects shared Anglocentric metaphor encoding rather than convergence toward human idiomatic usage. It demonstrates that human expertise remains essential for high-quality idiomatic translation, particularly for culture-specific expressions.
English neologisms, or newly coined words, have previously been shown to emerge in sparser semantic neighborhoods (filling semantic gaps) and near other neologisms (in growing semantic areas). In this work, we investigate where in semantic space Spanish neologisms emerge, and whether this mirrors English neologism development. We find that Spanish neologisms, in comparison to non-neologisms, do indeed appear both nearer to other neologisms and further from non-neologisms. We additionally investigate the prevalence of loanwords from other languages through time in Spanish neologism production and manually assess the topics that appear as loanwords at four years: 1810, 1900, 1950, and 1990. Our findings show that on average, the Spanish neologisms in our dataset have fewer neighboring words in semantic space compared to non-neologisms and tend to cluster more tightly in the semantic space, indicating that patterns of neologism emergence span languages. This suggests that novel methods for neologism detection may be cross-lingually applicable, with these features serving as multilingual predictors of neologism emergence.
Assessing the Pragmatic Competence of LLMs Regarding Novel Discourse Markers in Digital Communication
Ágnes Abuczki | Giedre Valunaite Oleskeviciene
Ágnes Abuczki | Giedre Valunaite Oleskeviciene
The English language is changing faster than before, partly due to the influence of the Internet. Digital language includes a large number of discourse markers (DMs), many of which can be considered innovative. Acronymization, pragmatic specialisation, and compensatory lexical innovation are the most common lexical processes that can be witnessed in the DMs used in computer-mediated communication (CMC). The following novel DMs were identified in recent Twitter chats: lol, tbh, omg, meh, and idk. These DMs perform several functions, such as showing emotions, signalling uncertainty, hesitation, or mitigation. Interpreting these functions may not be an easy or obvious task for AI. The primary aim of the study is to evaluate the pragmatic competence of an LLM, Gemini 3 Pro, regarding the interpretation of these novel DMs. A mixed-method research process was employed: LLM-generated outputs were compared with the findings of the relevant literature, quantitative corpus analysis, and our qualitative human interpretation to assess the model’s analytical usefulness. Gemini 3 Pro was found to show a high level of pragmatic competence in terms of interpreting the functions of DMs, but sometimes tended to overgeneralise, or failed to understand the tone of the text and the intention of the speaker to use a DM.
The growing popularity and misconceptions about conversational AI systems are driving efforts to establish a universally accepted framework for evaluating large language models. Testing large language models on tasks designed to assess human cognitive skills has become widespread. This paper presents the results of a pilot experiment and a comparative evaluation of the ability of OpenAI’s GPT-4.1 and GPT-4.1 mini to detect semantic ambiguity based on the works of Shultz and Pilon (1973) and Zipke et al. (2009). The experiment used a task sheet of 116 items utilising riddles, single sentences, and sentence pairs. It included systematically varied instructions on a four-level scale ranging from no mention of ambiguity to direct mention. Lexical and structural ambiguity were both employed, including surface-structure and deep-structure ambiguity. The results suggest that even advanced models, such as GPT-4.1 and GPT-4.1 mini, tend to consider only one possible meaning of ambiguous sentences. However, the recognition of ambiguity improved quickly when the possibility of ambiguity was explicitly referenced in the instruction. Additionally, the results imply that model size is not directly connected to performance, as GPT-4.1 scored better on lexical ambiguity detection tasks, while GPT-4.1 mini surpassed the larger model in structural ambiguity detection. The findings prove that future research with a more complex experiment design based on the same principles would be beneficial.
LLM-Based Frame and Stance Annotation for 19th-Century Rumour Discourse in US and UK Newspapers
Wanshu Zhang
Wanshu Zhang
Large language models (LLMs) are increasingly used for lexicographic support, neology detection, and semantic categorization, yet their behaviour on historical newspapers remains under-evaluated. This short paper describes an ongoing project that extends a DH2026-accepted two-phase methodology for extracting and tracking rumours in historical newspapers. From large-scale US and UK corpora (PleIAs/US-PD-Newspapers; biglam/hmd_newspapers), the DH workflow produces gold-standard sentence-level rumour instances with proposition-like “rumour content” spans (Rumour_Content/Cleaned_Content) and extraction-pattern metadata. Building on these historically grounded units, we propose an LLM-centered benchmark and analysis pipeline for assigning topical frames and evidential stance to rumour propositions, and for auditing “temporal projection” when models introduce anachronistic modern misinformation framings. For controlled cross-variety comparison we construct a strictly balanced benchmark of 800 instances over two well-attested bins (1840–1859, 1860–1879) and both national varieties (200 per country per bin). We outline prompt conditions (text-only vs time-aware vs historically calibrated) and self-consistency voting to quantify label stability and error modes. A small manually annotated subset supports evaluation, while the main contribution is the benchmark design, prompts, and reproducible protocol enabling community feedback before full-scale results are finalized.
up
Proceedings of the 2nd Workshop on Ecology, Environment, and Natural Language Processing
Proceedings of the 2nd Workshop on Ecology, Environment, and Natural Language Processing
Francesca Grasso | Valerio Basile | Cristina Bosco | Muhammad Okky Ibrohim | Maria Skeppstedt | Manfred Stede
Francesca Grasso | Valerio Basile | Cristina Bosco | Muhammad Okky Ibrohim | Maria Skeppstedt | Manfred Stede
Retrieving Climate Change Disinformation by Narrative
Max Upravitelev | Veronika Solopova | Charlott Jakob | Premtim Sahitaj | Sebastian Möller | Vera Schmitt
Max Upravitelev | Veronika Solopova | Charlott Jakob | Premtim Sahitaj | Sebastian Möller | Vera Schmitt
Climate disinformation evolves faster than the fixed taxonomies used to detect it. Thus, we re-frame narrative detection as a retrieval task: given a narrative’s core message as a query, rank texts from a corpus by alignment with that narrative. This formulation requires no predefined label set and can accommodate emerging narratives. We repurpose three climate disinformation datasets (CARDS, Climate Obstruction, climate change subset of PolyNarrative) for retrieval evaluation and propose SpecFi, a framework that generates hypothetical documents to bridge the gap between abstract narrative descriptions and their concrete textual instantiations. SpecFi uses community summaries from graph-based community detection as few-shot examples for generation, achieving a MAP of 0.505 on CARDS without access to narrative labels. We further introduce narrative variance, an embedding-based difficulty metric, and show via partial correlation analysis that standard retrieval degrades on high-variance narratives (BM25 loses 63.4% of MAP), while SpecFi-CS remains robust (32.7% loss). Our analysis also reveals that unsupervised community summaries converge on descriptions close to expert-crafted taxonomies, suggesting that graph-based methods can surface narrative structure from unlabeled text.
Unsupervised GRI-TCFD Alignment with LLM-Assisted Validation for Climate Disclosure and Greenwashing Risk Analysis
Seyed Alireza Mousavian Anaraki | Danilo Croce | Roberta Costa | Luigi Tiburzi | Armando Calabrese | Roberto Basili
Seyed Alireza Mousavian Anaraki | Danilo Croce | Roberta Costa | Luigi Tiburzi | Armando Calabrese | Roberto Basili
Climate-related corporate disclosures play a central role in sustainable finance and regulatory supervision, but remain difficult to analyze due to their length, unstructured format, and strategic language. While existing NLP approaches have been applied to ESG scoring and greenwashing detection, most operate at the document level and lack explicit alignment with formal reporting standards. We propose a scalable paragraph-level framework for aligning sustainability disclosures with the Global Reporting Initiative (GRI) indicators and the Task Force on Climate-related Financial Disclosures (TCFD) pillars. Our approach combines weak supervision, climate-focused GRI-TCFD mapping, embedding-based semantic similarity, and LLM validation for climate detection. In parallel, we introduce a paragraph-level greenwashing proxy based on commitment intensity, claim specificity, and sentiment polarity. This proxy complements regulatory alignment by capturing linguistic signals associated with potentially symbolic climate communication. The resulting augmented dataset is used to fine-tune ClimateBERT models in both single-task and multi-task settings. Experimental results show that weakly supervised dataset augmentation improves robustness and generalization compared to purely manual training, with further gains in the multi-task configuration. By integrating regulatory semantics, domain-adapted language models, and scalable annotation strategies, this study advances standard-aligned climate disclosure analysis and provides tools directly relevant to climate-related financial risk assessment.
Towards Empowering Consumers through Sentence-level Readability Scoring in German ESG Reports
Benjamin Josef Schüßler | Jakob Prange
Benjamin Josef Schüßler | Jakob Prange
With the ever-growing urgency of sustainability in the economy and society, and the massive stream of information that comes with it, consumers need reliable access to that information. To address this need, companies began publishing so called Environmental, Social, and Governance (ESG) reports, both voluntarily and forced by law. To serve the public, these reports must be addressed not only to financial experts but also to non-expert audiences. But are they written clearly enough? In this work, we extend an existing sentence-level dataset of German ESG reports with crowdsourced readability annotations. We find that, in general, native speakers perceive sentences in ESG reports as easy to read, but also that readability is subjective. We apply various readability scoring methods and evaluate them regarding their prediction error and correlation with human rankings. Our analysis shows that, while LLM prompting has potential for distinguishing clear from hard-to-read sentences, a small finetuned transformer predicts human readability with the lowest error. Averaging predictions of multiple models can slightly improve the performance at the cost of slower inference.
Disambiguating Geographic Names in Biodiversity Occurrence Data: A Retrieval-Augmented Generation Approach
Yanni Jose C. Ella | Monica Ashley R. Laviste | John Michael L. Lastimoso | Wilfred John E. Santiañez | Riza Batista-Navarro | Roselyn Santos Gabud
Yanni Jose C. Ella | Monica Ashley R. Laviste | John Michael L. Lastimoso | Wilfred John E. Santiañez | Riza Batista-Navarro | Roselyn Santos Gabud
The availability of georeferenced coordinates is essential for biodiversity research, as it enables species distribution modeling and supports conservation planning. However, datasets often contain ambiguous or inconsistent geographic names that reduce spatial accuracy and underscore the need for methods that resolve geographic name ambiguity. While traditional named entity linking strategies are well established, they remain limited in low-resource domains, e.g., in biodiversity contexts, due to the scarcity of annotated training data and high lexical ambiguity of local geographic names. This study proposes a Retrieval-Augmented Generation (RAG) framework to automatically disambiguate Philippine seaweed-related geographic names in databases and literature. This approach utilizes a custom knowledge base of gazetteers to support large language models (LLMs) in the task of geospatial disambiguation. With a disambiguation accuracy of 87.8% within a 5 km distance error threshold, our evaluation shows that the RAG-enabled pipeline significantly outperforms standard LLM baselines (Accuracy@5km = 0%), demonstrating the need for external knowledge to resolve geospatial ambiguity.
Recent advances in generative AI have enabled the large-scale production of environmental imagery and descriptions, yet questions remain regarding how such content represents emotion, agency, and responsibility. This study examines how human evaluators respond to AI-generated environmental representations, focusing on sentiment, stance, and argumentation as dimensions of qualitative evaluation. Data were collected from 81 multilingual secondary-school EFL learners in Cyprus, who engaged with AI-generated environmental images and accompanying AI-written descriptions through a sequence of structured tasks. Using qualitative discourse analysis informed by sentiment- and stance-oriented frameworks, the study analyses learner-produced texts to identify affective evaluations, moral positioning, and alignment with or challenge to AI-generated discourse. Findings indicate that participants consistently moved beyond surface-level description to articulate emotional engagement, assign responsibility, and critique omissions in AI-generated content, particularly regarding the representation of human-environment relations. The study contributes to research on human-centered AI evaluation by demonstrating the value of sentiment and stance analysis for assessing AI-generated environmental language, and highlights the potential of educational contexts as sites for examining human interpretive responses to automated discourse.
What Stories Do Language Models Tell About Nature? A Multi Layer Evaluation Framework for Ecological Alignment
Jorge Vallego | Eleanor Tiernan | Mah Rukh | Mariana Roccia | Sabina Fiebig Lord
Jorge Vallego | Eleanor Tiernan | Mah Rukh | Mariana Roccia | Sabina Fiebig Lord
Large language models increasingly generate environmental discourse, yet there is no standardised framework for evaluating the ecological narratives they produce. We introduce a structured prompt corpus and a reproducible multi layer evaluation framework grounded in ecolinguistic theory, operationalising five dimensions of ecological alignment: anthropocentrism, agency attribution, erasure of non human impacts, evaluation of growth, and responsibility framing. The framework integrates human judgement, an ecosophy aligned model judge, and automated semantic metrics, and is applied to outputs from ChatGPT, DeepSeek, and Ecophora, our ecosophy guided model. Ecophora achieves the highest alignment across all layers, with near ceiling judge scores of 159/160 and 142/160, together with the strongest automated composite performance. Divergences between automated metrics and holistic judgement indicate that ecological vocabulary alone does not guarantee ecological reasoning. The proposed framework provides a scalable methodology for benchmarking ecological alignment and assessing narrative shifts in language models.
Ecological Discourse Modeling in a Low-Resource Setting: A Longitudinal Vietnamese Climate Corpus with Comparative Topic Modeling
Huyen Phuong Nguyen
Huyen Phuong Nguyen
Climate change discourse has expanded substantially in recent decades, yet computational analyses remain concentrated on high-resource languages. In this paper, we construct a longitudinal Vietnamese climate news corpus and examine thematic structure and temporal evolution in a lower-resource setting. The corpus comprises 10,401 articles published between 2004 and 2026 and is systematically preprocessed using linguistically informed word segmentation. To ensure domestic relevance, we apply transformer-based Named Entity Recognition and construct a geographically grounded subset of 4,501 Vietnam-focused documents. We analyze this dataset using both Latent Dirichlet Allocation and BERTopic. Results reveal stable thematic dimensions alongside longitudinal shifts from event-driven pollution reporting toward governance- and energy-centered narratives. Embedding-based modeling achieves higher semantic coherence while maintaining comparable topic diversity. The main contribution of this work is thus the compilation of a structured Vietnamese climate corpus and a systematic analysis of discourse evolution in an underrepresented language context.
Greench-v1: distilling SLMs on Greenwashing Detection
Federico Raspanti | Alessandro Pietro Bardelli Bardelli | Simona Scala | İrem Demirtaş | Marilena Di Bari | Michele Filannino
Federico Raspanti | Alessandro Pietro Bardelli Bardelli | Simona Scala | İrem Demirtaş | Marilena Di Bari | Michele Filannino
Validating greenwashing claims in environmental, social, and governance (ESG) reports relies heavily on costly and inconsistent manual review. To address this, this paper introduces Greench-v1, a low-latency small language model (based on Qwen3-4B) that screens ESG text at the paragraph level. The model outputs a three-way classification (Greenwashing Alert, No Greenwashing, Not Relevant) paired with a concise, paragraph-grounded rationale to assist human auditors in triage and validation. The system was trained on a custom dataset of roughly 2,000 paragraphs, adapted from the ClimateBERT corpus. This dataset mitigates class imbalance through controlled paraphrasing of rare positive instances and uses GPT-4o to generate evidence-based justifications. Four training regimes were evaluated: (i) Hard distillation: Supervised fine-tuning on teacher-generated outputs. (ii) Soft distillation: Training the student to match the temperature-scaled logits of a domain-specialized Qwen3-14B teacher. (iii) Group Relative Policy Optimization (GRPO): Reward-based updates driven by exact-match alert generation. (iv) Hybrid GRPO: GRPO initialized from the hard-distilled checkpoint. Distillation and efficient policy optimization significantly improved performance over untuned baselines. Soft distillation and GRPO achieved the strongest results, increasing the “Greenwashing Alert” weighted F1-score by 36.7% and 49.0%, respectively, resulting in a deployable tool for screening large volumes of ESG narratives.
Analyzing Environmental Discourse through Construction-Based Pattern Extraction
Elisa Chierchiello | Eliana Di Palma | Ludovica Pannitto | Cristina Bosco
Elisa Chierchiello | Eliana Di Palma | Ludovica Pannitto | Cristina Bosco
Environmental issues are at the centre of a debate currently taking place across all communication channels. This paper provides an analysis of texts in which these issues are discussed, with the novelty of applying a methodology that enables the extraction and comparison of different narratives and points of view. The texts used in this study are the English Living Planet Reports published biennially by the WWF from 2014 to 2024. The methodology is based on the extraction of constructions – patterns collected in the English constructicon CASA – which allow us to identify differences in the presentation of the issues discussed in the analysed texts. Our results show that this methodology can be very helpful in the comparative analysis of texts to reveal different perspectives, for example, to observe diachronic variations.
Mapping the Historical Ecology of the Cyclades: A Diachronic Natural Language Processing Analysis of Travel Narratives (1700–1920)
Aikaterini Christopoulou | Vassilis Detsis | Basilis Gatos
Aikaterini Christopoulou | Vassilis Detsis | Basilis Gatos
Historical texts can be valuable for the study of a place’s ecological history but reading and extracting information from them can be a tedious and time-consuming task. Natural Language Processing can help in order to extract the most important information of the text in a quick, effective and reproducible way. In this study, travel narratives for the Cyclades Islands from 4 different time periods (1700-1920) have been chosen for analysis. The first step, the quantitative part, includes the semi-automatic detection of geographical entities in the texts and their connection to predefined keywords in order to enable temporal and spatial statistical analysis. The output of this procedure is then inserted in a Retrieval-Augmented Generative Synthesis pipeline in which the text segments with the connected place and keyword are processed by a locally orchestrated Large Language Model. The final output is used for the understanding and interpretation of the original text. Even though the study focuses mainly on the coherence and repeatability of the workflow, an effort is made to interpret and connect the results to the past ecological profile of these islands. The dataset/supplementary material is provided via an open access repository.
Retrieving Floods without Floodlights: Topic Models as Binary Classifiers for Extreme Climate Events in German News
Brielen Madureira | Mariana Madruga de Brito | Andreas Niekler
Brielen Madureira | Mariana Madruga de Brito | Andreas Niekler
In studies of media coverage of extreme climate events, NLP methods have become indispensable for identifying relevant texts in large news databases. Still, enough annotated data to train accurate deep learning-based classifiers from scratch is often not available. Topic Models have the advantage of being both unsupervised and interpretable, but are typically used only for exploratory analysis or data characterisation. In this study, we investigate how to employ Topic Models as binary classifiers for refining the retrieval of relevant news about seven types of extreme climate events in the German media. Our method relies on the posterior distributions estimated by Topic Models to select relevant documents, without modifying their training procedure. Using an annotated sample to guide the evaluation, we show that the probabilities assigned to keywords used to query news databases can also be informative for selecting relevant topics and improve sample precision. We compare our results to a fine-tuned text embedding classifier and an open-weight LLM, discussing observed trade-offs, e.g. the LLM’s lowest precision. Moreover, we show that results are hazard-dependent, which speaks against considering climate events as a single category in NLP tasks.
Why Is This Green? LLM-Based Explanations of Implicit Green Practices in Social Media
Anna Glazkova | Olga Zakharova | Daria Lebedeva
Anna Glazkova | Olga Zakharova | Daria Lebedeva
Identifying green practices in social media is not merely a matter of lexical matching. Many green practices are expressed implicitly, rely on shared background knowledge, or are embedded in broader contextual narratives. In this paper, we investigate how large language models (LLMs) explain expert annotations of green waste management practices and how they rationalize classification errors made by a fine-tuned model (mBART) on a Russian social media corpus (GreenRu). We analyze explanations generated by two LLMs (T-lite and GigaChat) in two settings: (1) explaining gold expert-assigned labels and (2) interpreting erroneous model predictions. Our qualitative and micro-quantitative analysis shows that green practices are frequently inferred through contextual reasoning rather than explicit terminology. Error patterns of mBART reveal overgeneralization, associative misinterpretation (e.g., linking food sharing to waste recycling), and detection of practices where none are present. We further compare explanatory strategies of the two LLMs. T-lite tends to rely on lexical cues and surface markers that may create an impression of a practice, while GigaChat more often reconstructs broader contextual interpretations. Expert feedback highlights limitations of formal textual analysis, sensitivity to missing contextual knowledge, and difficulties in aligning model reasoning with expert conceptual boundaries. Our findings suggest that explanation-based analysis is a productive tool for diagnosing classification errors and refining annotation guidelines. More broadly, the study demonstrates that modeling implicit sustainability discourse requires contextual grounding and deeper semantic integration beyond keyword-based approaches.
Introducing a Green Leaderboard for Sustainable Risk Prediction in Streaming NLP Shared Tasks.
Alba María Mármol-Romero | Adrián Moreno Muñoz | Arturo Montejo-Raez
Alba María Mármol-Romero | Adrián Moreno Muñoz | Arturo Montejo-Raez
Current NLP shared-task evaluations predominantly rank systems by predictive performance, overlooking computational efficiency and environmental impact. This limitation is particularly critical in streaming and early risk detection scenarios, where models operate continuously, and resource consumption accumulates over time. We propose a sustainability-aware evaluation framework for streaming NLP tasks by introducing the Green Early Detection Score (GED), which integrates classification performance, detection timeliness, and carbon emissions. We also present an energy-based variant tailored to on-device early risk detection settings where energy consumption per inference is a key constraint. Applying these metrics to three editions (2023-2025) of the MentalRiskES shared task, we construct the first Green Leaderboard for early risk detection. Our results show that sustainability-aware ranking substantially reshapes system positions, highlighting efficient models that remain undervalued under performance-only evaluation.
Not Everything Is Greenwashing: Limitations of Automatic Analysis of Sustainability Reports, and a Proposal
Maria Pilar Uribe Silva | Rik van Noord | Malvina Nissim
Maria Pilar Uribe Silva | Rik van Noord | Malvina Nissim
Sustainability reports (SRs) are essential for holding companies accountable, and they are required by law. They also serve as a key communication tool through which companies shape their image and disclose non-financial information. However, the rapid growth of these reports, their lack of standardisation, and the frequent use of strategically ambiguous language make it difficult for stakeholders to evaluate whether sustainability claims are genuine or deceptive. Previous work has focused on extracting misleading climate-related content and identifying greenwashing. We argue that this is not enough, because deception does not only appear in overtly false or misleading green claims, and it often emerges through a variety of subtle linguistic strategies. We therefore propose the development of a framework based on deception theories to examine how deceptive language operates in SRs, and we outline the challenges that should be seen as an invitation for future research. Keywords: Deception, Deceptive Language, Sustainability Reports, Greenwashing
up
Proceedings of the the fifth edition of NLPerspectives
Proceedings of the the fifth edition of NLPerspectives
Shiran Dudy | Gavin Abercrombie | Valerio Basile | Elisa Leonardelli | Simona Frenda
Shiran Dudy | Gavin Abercrombie | Valerio Basile | Elisa Leonardelli | Simona Frenda
What is Truth in NLP? Reflecting on Progress, Lessons, and Open Challenges as NLPerspectives turns Five
Gavin Abercrombie
Gavin Abercrombie
This paper reflects on five years of the Workshop on Perspectivist Approaches to NLP (NLPerspectives) and examines how this research community has helped to reconceptualise the notion of ground truth in human-labelled data. As NLP research has increasingly engaged with social and affective tasks, traditional assumptions about annotation reliability–centred on inter-annotator agreement and single ‘gold standard’ labels–have proven insufficient for capturing the genuine diversity of human perspectives. I review the developments that have driven the ‘Perspectivist Turn,’ assess its influence on mainstream NLP practice, and highlight the methodological challenges that arise when modelling disagreement, subjectivity, and annotator variation. In particular, I consider unresolved questions around evaluation paradigms, task formulation, population representation, community norms, and the implications of using pre-trained generative models as classifiers. By synthesising discussions from five years of workshops, keynotes, and related publications, I outline open challenges and propose directions for future work aimed at more rigorous perspectivist NLP. I argue that we should focus on centering minoritised standpoints and caution against viewing potentially harmful interpretations as equally legitimate reactions to ‘subjective’ phenomena.
The Rashomon Wikipedia: A Data-Perspectivist Analysis of Divergent Historical Narratives
Claudiu Creanga | Liviu P. Dinu | Anca Dinu
Claudiu Creanga | Liviu P. Dinu | Anca Dinu
Wikipedia aims to provide a unified, neutral record of history, yet its independent language editions often function as distinct epistemic communities, creating divergent narratives around contested events. This paper investigates cross-lingual historiographical bias by analyzing Wikipedia articles across five languages (Romanian, Hungarian, Russian, Turkish, and English) focusing on three contentious events in Romanian history: the Battle of Posada (1330), the Soviet occupation of Bessarabia (1940), and the Night Attack at Târgoviște (1462). Using human annotators and Large Language Models (LLMs) to classify citation stance and quantify narrative evolution from 2005 to 2024, we identify a phenomenon of “citation isolation”. In the case of the Battle of Posada, only 2 out of 119 citations were shared between language editions, with the Romanian edition exhibiting a 91% pro-national bias compared to the balanced Hungarian edition. Longitudinal analysis reveals that these narratives are volatile and responsive to contemporary geopolitics, evidenced by a significant shift in the Russian framing of Bessarabia in 2024. Finally, we propose a “Peace-Maker” pipeline to automate conflict reconciliation. We demonstrate that while standard prompting leads models to hallucinate consensus, “adversarial” prompting, which explicitly instructs the model to preserve and attribute disagreement, achieves near-perfect neutrality scores.
GSI:detect - A Perspectivist Approach to Gender Stereotypes Identification in Italian
Davide Testa | Sofia Brenna | Manuela Speranza | Gloria Comandini | Stefania Cavagnoli | Bernardo Magnini
Davide Testa | Sofia Brenna | Manuela Speranza | Gloria Comandini | Stefania Cavagnoli | Bernardo Magnini
The deconstruction of gender stereotypes is essential to prevent discrimination, marginalization and gender-based violence. Despite the increasing attention to this issue, research in this field often focuses on explicitly sexist or hateful communication, leaving out all the cases where stereotypes are produced unconsciously or even with apparently positive intentions. Moreover, the identification and analysis of gender stereotypes is often a very subjective task, heavily influenced by the researcher’s background, beliefs and personal sensitivity. In this context GSI:detect, a dataset for gender stereotypes identification in Italian, has been annotated following a perspectivist approach that gives value to the different points of view of four annotators. It has been designed to address (i) the lack of resources focusing on naturally occurring and non-hateful language conveying implicit or ambiguous forms of gender stereotypes, and (ii) the scarcity of datasets that can capture multiple interpretations as well as the inherent variation and disagreement in human perception. Baseline experiments with several LLMs confirm the challenging nature and value of such a linguistic resource, revealing both apparent differences and limitations in performance among the evaluated models, and raising questions about the extent to which current LLMs are suitable for detection and classification tasks in this field. Content warning: Examples taken from the GSI:detect dataset may contain sensitive or potentially distressing content.
It is increasingly recognized that humans do not always agree, and disagreement is inherent in many annotation tasks. However, not all items in a given task elicit the same level of opinion divergence. In this paper, we study the extent to which item-level annotation variation and variation structure can be captured from text features, focusing on inappropriate language detection, including offensive language, hate speech, and toxic language detection. We model annotation variation to assess whether the degree of annotation divergence can be predicted from item-level textual features. We also propose the Opposition Index, a metric that quantifies the extent of opposing stances among annotators based on their Likert ratings.
HurtLens: A Perspectivist Corpus Analysis of Hurtful Language
Samuele D’Avenia | Eliana Di Palma | Marta Marchiori Manerba | Valerio Basile
Samuele D’Avenia | Eliana Di Palma | Marta Marchiori Manerba | Valerio Basile
Offensive language detection systems often rely on majority-aggregated annotations, overlooking the diversity of perspectives that shape how different communities perceive harm. In this contribution, we introduce HurtLens, a perspectivist corpus of hurtful language leveraging four disaggregated datasets which are automatically enriched through HurtLex lemmas, a multilingual resource of offensive and derogatory terms. Using mixed-effects modeling, we investigate how annotators’ sociodemographic backgrounds, the presence of specific types of offensive language (through Hurtlex categories) and their interaction influence offensiveness ratings. Our analysis reveals that offensiveness ratings are influenced both by annotators’ sociodemographic characteristics (particularly when considering them in intersection) and by the presence of specific types of offensive language. Additionally, we identify significant interaction effects showing that different demographic groups vary in their sensitivity to texts containing particular types of offensive language.
We introduce a new metric that quantifies the extent of systematicity of the disagreement between annotators. The metric, called σ, is inspired by Structural Balance Theory and it approximates the clusterability of the annotators of a dataset. Paired with a standard metric of inter-annotator agreement such as Krippendorffs α, σ measures the amount of disagreement which stems from genuine subjective factors as opposed to the amount of disagreement caused by inner features of the annotation task. The metric is applied to over twenty datasets encoding a broad variety of annotations, showing its effectiveness in capturing the systematicity of annotator disagreement and its explanatory value.
Fine-Grained Perspectives: Modeling Explanations with Annotator-Specific Rationales
Olufunke O. Sarumi | Charles Welch | Daniel Braun
Olufunke O. Sarumi | Charles Welch | Daniel Braun
Beyond exploring disaggregated labels for modeling perspectives, annotator rationales provide fine-grained signals of individual perspectives. In this work, we propose a framework for jointly modeling annotator-specific label prediction and corresponding explanations, fine-tuned on the annotators’ provided rationales. Using a dataset with disaggregated natural language inference (NLI) annotations and annotator-provided explanations, we condition predictions on both annotator identity and demographic metadata through a representation-level User Passport mechanism. We further introduce two explainer architectures: a post-hoc prompt-based explainer and a prefixed bridge explainer that transfers annotator-conditioned classifier representations directly into a generative model. This design enables explanation generation aligned with individual annotator perspectives. Our results show that incorporating explanation modeling substantially improves predictive performance over a baseline annotator-aware classifier, with the prefixed bridge approach achieving more stable label alignment and higher semantic consistency, while the post-hoc approach yields stronger lexical similarity. These findings indicate that modeling explanations as expressions of fine-grained perspective provides a richer and more faithful representation of disagreement. The proposed approaches advance perspectivist modeling by integrating annotator-specific rationales into both predictive and generative components.
Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference
Olga Kellert | Sriya Kondury | Candice Koo | Nemika Tyagi | Steffen Eikenberry
Olga Kellert | Sriya Kondury | Candice Koo | Nemika Tyagi | Steffen Eikenberry
Annotation pipelines in Natural Language Processing (NLP) commonly assume a single latent ground truth per instance and resolve disagreement through label aggregation. Perspectivist approaches challenge this view by treating disagreement as potentially informative rather than erroneous. We present a large-scale analysis of graded health-literacy annotations from 6,323 open-ended COVID-19 responses collected in Ecuador and Peru. Each response was independently labeled by multiple annotators using proportional correctness scores, allowing us to analyze the full distribution of judgments rather than aggregated labels. Variance decomposition shows that question-level conceptual difficulty accounts for substantially more variance than annotator identity, indicating that disagreement is structured by the task itself rather than driven by individual raters. Agreement-stratified analyses further reveal that key social-scientific effects, including country, education, and urban-rural differences, vary in magnitude and in some cases reverse direction depending on levels of inter-annotator agreement. These findings suggest that graded health-literacy evaluation contains both epistemically stable and unstable components, and that aggregating across them can obscure important inferential differences. We therefore argue that strong perspectivist modeling is not only conceptually justified but statistically necessary for valid inference in graded interpretive tasks.
SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs
Pietro Bernardelle | Leon Froehling | Stefano Civelli | Gianluca Demartini
Pietro Bernardelle | Leon Froehling | Stefano Civelli | Gianluca Demartini
As increasingly capable large language models (LLMs) emerge, researchers have begun exploring their potential for subjective tasks. While recent work demonstrates that LLMs can be aligned with diverse human perspectives, evaluating this alignment on downstream tasks (e.g., hate speech detection) remains challenging due to the use of inconsistent datasets across studies. To address this issue, in this resource paper we propose a two-step framework: we (1) introduce SubData, an open-source Python library designed for standardizing heterogeneous datasets to evaluate LLMs perspective alignment; and (2) present a theory-driven approach leveraging this library to test how differently-aligned LLMs (e.g., aligned with different political viewpoints) classify content targeting specific demographics. SubData’s flexible mapping and taxonomy enable customization for diverse research needs, distinguishing it from existing resources. We illustrate its usage with an example application and invite contributions to extend our initial release into a multi-construct benchmark suite for evaluating LLMs perspective alignment on natural language processing tasks.
ChatGPT, why can’t anyone afford a house? On the Effects of LLM pre-annotation on Annotator Subjectivity
Emilie Francis | Céline Leuzinger | Ricardo Muñoz Sánchez | Lee D. Gauthier
Emilie Francis | Céline Leuzinger | Ricardo Muñoz Sánchez | Lee D. Gauthier
Large language models (LLMs) have often been proposed as substitutes for human annotators in a variety of tasks. At the same time, there has been increased focus on the role that human subjectivity and perspective plays in data annotation. To avoid eliminating the human role in annotation entirely, the use of LLMs for pre-annotation has been suggested as an alternative approach. In this paper, we explore to which degree this approach affects subjectivity of social media annotation in English. We focus on comments regarding the current status of the housing market and label them for concern level, factors affecting housing affordability, and aspects that authors claim either exacerbate or improve the situation. To investigate this, we design an experiment involving two rounds of annotation: the first, a dataset annotated by humans only; and the second, a dataset with LLM pre-annotations curated by the same human annotators. We observe that the second setting leads to much higher agreement, as well as significant changes in label distribution and co-occurrence. Similar shifts do not appear in the LLM labels. Our findings show that use of LLMs in the annotation process leads to convergence in annotations and, thus, to an erosion of human subjectivity.
An Overview of Current Practices and Recommendations for Working with Stereotypes in NLP
Alessandra Teresa Cignarella | Matteo Pellegrini
Alessandra Teresa Cignarella | Matteo Pellegrini
This article presents a discussion on the main challenges and considerations involved in addressing stereotypes within Natural Language Processing (NLP), and proposes a set of guidelines and recommendations for their treatment in research and resource development. On the one hand, the growing interest in fairness, bias mitigation, and inclusivity has led to an increasing number of studies and datasets dealing with stereotypes; on the other hand, their conceptualization and operationalization remain highly heterogeneous across works. The aim of this article is therefore twofold: (1) to provide a concise yet comprehensive overview of existing annotation schemes highlighting their key features and offering a comparative analysis and (2) to propose a set of tentative guidelines and recommendations to foster clarity when working with stereotypes in NLP. Furthermore, as a case study, we conduct an annotation exercise of a subset of texts from the QUEEREOTYPES dataset, containing stereotypes targeting LGBTQIA+ people, using all labels proposed in prior work to assess their clarity, overlap, and practical usefulness.
Modeling Perspectives in NLP: Parameter-Efficient Perspective Conditioning for Span Extraction and Summarization
Harikrishnan Gurushankar Saisudha | Sabine Bergler
Harikrishnan Gurushankar Saisudha | Sabine Bergler
Understanding text through multiple perspectives is essential in domains such as healthcare community question answering, where answers frequently contain heterogeneous viewpoints, including experiences, suggestions, causes, follow-up questions, and informational claims. We present a unified perspective-conditioned framework for both span identification and perspective-aware summarization on the PerAnsSumm dataset. Our approach introduces explicit perspective signals into transformer models using two parameter-efficient mechanisms: prefix-conditioned representations and perspective-aware attention layers. We first employ multi-label perspective classification to identify relevant viewpoints, which serve as conditioning signals for downstream tasks. For span identification, we model perspective-specific extraction as a conditioned binary sequence labeling problem. For summarization, we guide generation using perspective-enriched encoder representations. Experiments demonstrate that explicit perspective conditioning substantially improves span detection performance while achieving competitive summarization quality. Notably, perspective-aware attention achieves strong results using only a small fraction of the trainable parameters required by full fine-tuning. Our findings highlight the importance of structured viewpoint modeling and show that explicit perspective control enables efficient and interpretable multi-perspective text understanding.
A Pilot Study Investigating Stakeholder Subjectivity in Collaborative Dialog Analysis
Ananya Ganesh | Martha Palmer | Katharina von der Wense
Ananya Ganesh | Martha Palmer | Katharina von der Wense
Qualitative research in education relies on“ground truth" codes or labels generated by having a trained or expert coder code observations in data such as student dialog. Although rigorous validity checks are a part of the coding process, there is limited research investigating how and to what extent, this notion of the ground truth is influenced by inherent task subjectivity. This paper presents a pilot study of task subjectivity centered around the phenomenon of verbal off-task behavior. The context for this study is real-world small-group collaborative conversations among three to five students in a middle-school science classroom. To investigate how stakeholders such as teachers and students show subjectivity in approaching this task, we recruit five teachers from the Prolific online platform, and five students from local middle and high schools as annotators of off-task speech. We show that teachers, students, and expert coders differ in their perception of off-task speech, with some of these differences being systematic. Drawing upon recent research in machine learning and natural language processing, we then outline the potential benefits of collecting and modeling a range of codes that explicitly represent the subjective perspectives of a diverse set of coders.
up
Proceedings of the 1st Workshop on Social Context (SoCon) and the 2nd Workshop on Integrating NLP and Psychology to Study Social Interactions (NLPSI) @ LREC 2026
Proceedings of the 1st Workshop on Social Context (SoCon) and the 2nd Workshop on Integrating NLP and Psychology to Study Social Interactions (NLPSI) @ LREC 2026
Marco Antonio Stranisci | Neele Falk | Sofie Labat | Soda Marem Lo | Aswathy Velutharambath | Sabine Weber | Rossana Damiano | Simona Frenda | Veronique Hoste | Bennett Kleinberg | Roman Klinger | Viviana Patti | Flor Miriam Plaza-del-Arco | Maarten Sap | Seid Muhie Yimam
Marco Antonio Stranisci | Neele Falk | Sofie Labat | Soda Marem Lo | Aswathy Velutharambath | Sabine Weber | Rossana Damiano | Simona Frenda | Veronique Hoste | Bennett Kleinberg | Roman Klinger | Viviana Patti | Flor Miriam Plaza-del-Arco | Maarten Sap | Seid Muhie Yimam
State vs. Trait Anxiety in Causal Language Models
Karin Shistik | Idan-Chaim Cohen | Aviad Elyashar | Ortal Slobodin | Odeya Cohen | Rami Puzis
Karin Shistik | Idan-Chaim Cohen | Aviad Elyashar | Ortal Slobodin | Odeya Cohen | Rami Puzis
Psychological constructs in humans range along a state–trait continuum: traits persist across situations, while states fluctuate with context. Studies have shown that language models exhibit measurable psychological constructs, yet whether these constructs differ in contextual stability, as the state–trait distinction predicts, remains untested. We present the Questionnaire for Causal Language Models (QCLM), a psychometric framework that measures constructs through next-token probability distributions of base models. Applying QCLM to 35 causal language models under vanilla, stress, and neutral conditions, we assess two anxiety instruments targeting opposite ends of the state–trait continuum: STAI-S (state anxiety) and STAI-T (trait anxiety). Paired effect sizes and variance decomposition reveal that state anxiety is more sensitive to stress manipulation than trait anxiety: stimulus type accounts for a larger share of variance in state anxiety, while model identity contributes more to trait anxiety. These results provide empirical evidence that the state–trait distinction extends to language model behavior.
Documenting Rural Gatherings in Aging Japan: Social Context and Language Use in Interaction at a Mobile Supermarket
Haruka Sakai | Rui Sakaida
Haruka Sakai | Rui Sakaida
This paper presents a documentation framework and an exploratory analysis of language use in everyday interactions at rural gatherings in aging Japan, a communicative setting shaped by distinct social contexts that remain largely absent from existing language resources. Drawing on studies of face-to-face encounters, we propose a typology of rural gatherings and examine mobile supermarkets (vehicles that transport and sell daily necessities at scheduled stops in areas that lack fixed retail stores) as a case study. We present a preliminary analysis based on a community-mediated recording methodology. The quantitative findings reveal that conversational hot spots occur immediately following the encounter and transaction phases, indicating that participants experience these encounters as occasions for social connection rather than mere commercial transactions. The qualitative findings from the interaction analysis demonstrate how participants simultaneously manage work and conversation through vocal, bodily, and temporal resources in a social context. We discuss how these findings illuminate dimensions in social contexts that require interdisciplinary investigation beyond what existing language resources currently capture.
This paper argues that, when it comes to modeling language variation and change on the video game streaming platform Twitch, it is necessary to consider “meso-level” communities of practice, i.e. communities of practice that are smaller than the full video game community, yet larger than the usual level of analysis in recent linguistics studies: communities associated with individual Twitch channels. We present a computational method for identifying these linguistically relevant communities of practice and show how this method can be useful for analyzing quantitative patterns of sociolinguistic variation in a corpus composed of the chat transcripts of 15 streamers of the game Elden Ring: Nightreign.
Implicit Cultural Identity Signals in Language: Detection and Effects in Negotiation Dialogue
Bin Han | Danah Yun | James Hale | Jonathan Gratch
Bin Han | Danah Yun | James Hale | Jonathan Gratch
Language conveys cultural identity even when not intentionally disclosed. This study examines cultural signals in task-oriented dialogue using English negotiations from the KODIS dataset. We focus our analysis on participants from four countries: the US, UK, Mexico, and South Korea. Interacting anonymously under identical conditions, we evaluated whether a speaker’s country could be inferred from dialogue by zero-shot LLMs and embedding-based classifiers. Results show that while objective negotiation outcomes remained similar across groups, subjective perceptions varied significantly. Embedding-based models reliably identified country of origin, whereas zero-shot LLM performance dropped under distribution shift. These findings suggest that cultural identity-related signals are embedded in language and may be relevant for analyzing negotiation dialogue.
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
Tatiana Petrova | Stanislav Sokol | Radu State
Tatiana Petrova | Stanislav Sokol | Radu State
Which persuasion strategies, if any, are associated with donation compliance? Answering this requires fine-grained strategy labels across a full corpus and statistical tests corrected for multiple comparisons. We annotate all 10,600 persuader turns in the 1,017-dialogue PersuasionForGood corpus with a taxonomy of 41 strategies in 11 categories, using three open-source large language models (Qwen3:30b, Mistral-Small-3.2, Phi-4). Strategy categories alone explain little variance in donation outcome (pseudo R-squared approximately 0.015, consistent across all three annotators). Guilt Induction is the only strategy significantly associated with lower donation rates (approximately -23 percentage points), an effect that replicates across all three models despite only moderate inter-model agreement. Reciprocity is the most robust positive correlate. Target sentiment and interest predict whether a donation occurs but show at most a weak correlation with donation amount. Logistic regression with sentiment, interest, Guilt Induction, and Reciprocity achieves nearly the same fit (pseudo R-squared = 0.080) as the full model with all strategy categories. These findings suggest that strategy identification alone is insufficient to explain persuasion effectiveness, and that guilt-based appeals may be counterproductive in prosocial settings. We release the fully annotated corpus as a public resource.
Do LLMs Ask the Right Questions? Evaluating GPT-Generated Surveys as Instruments for Measuring Social Attitudes
Tina Behzad | Wenbo Li | Reuben Kline | Klaus Mueller
Tina Behzad | Wenbo Li | Reuben Kline | Klaus Mueller
Understanding human beliefs and social attitudes often relies on carefully designed survey instruments. Recent work has suggested that large language models (LLMs) could automate parts of this process by generating surveys at scale, raising questions about the comparability of such instruments to literature-grounded, human-designed surveys. We present a controlled empirical comparison between GPT-generated surveys and established survey baselines across three social domains: climate change, immigration, and diversity, equity, and inclusion (DEI). GPT-generated surveys were produced using a fixed prompting framework enforcing a 3×3 structure over beliefs, perceptions, and behaviors, while human baselines were assembled from validated instruments to match survey length and construct coverage. We collected responses from U.S.-based participants, who completed both survey types, allowing direct within-subject comparison. We analyze differences in response distributions, clustering behavior, and alignment with self-identified stances. Our results show that GPT-generated surveys capture the same dominant attitudinal divisions as human-designed instruments, while exhibiting differences in the resolution of belief structure and group separation. These findings suggest that LLM-generated surveys are suited for exploratory and large-scale analyses, and can be used to complement expert-designed instruments.
Where Is Politeness in Japanese BERT? A Layerwise Probing and CLS Activation Patching Study
Shusuke Hashimoto | Wenchen Shi
Shusuke Hashimoto | Wenchen Shi
Politeness is a central pragmatic dimension of language use, and Japanese honorifics offer a well-defined testbed for studying whether pretrained encoders represent socially meaningful distinctions. Prior BERT-based work has applied supervised models to Japanese honorific data, but we are not aware of analyses that localize honorific-level information across layers or test causal influence via activation patching in Japanese BERT-style encoders. We study these questions in LineDistilBERT using the KeiCO corpus, which labels sentences with four honorific levels. To isolate pretrained representations while still defining a task predictor, we freeze all encoder parameters and train only a lightweight [CLS] classification head as a minimal readout. We then run layerwise linear probing, training multinomial L2-regularized logistic-regression probes on [CLS] vectors from each layer to quantify linear decodability across depth and to select a best layer on development data. Finally, we test causal leverage with [CLS] activation patching, transplanting donor activations into receiver sentences at selected layers and measuring prediction transitions, logit shifts, and flip rates under standard controls. Overall, honorific level is broadly decodable across layers, and [CLS] interventions can systematically steer the frozen-encoder classifier with strong depth dependence, providing complementary evidence from probing and causal intervention for Japanese politeness in practice.
Rewrite the News: Tracing Editorial Reuse across News Agencies
Soveatin Kuntur | Nina Smirnova | Anna Wroblewska | Philipp Mayr | Sebastijan Razboršek Maček
Soveatin Kuntur | Nina Smirnova | Anna Wroblewska | Philipp Mayr | Sebastijan Razboršek Maček
This paper investigates sentence-level text reuse in multilingual journalism, analyzing where reused content occurs within articles. We present a weakly supervised method for detecting sentence-level cross-lingual reuse without requiring full translations, designed to support automated pre-selection to reduce information overload for journalists (Hołyst et al., 2024). The study compares English-language articles from the Slovenian Press Agency (STA) with reports from 15 foreign agencies (FA) in seven languages, using publication timestamps to retain the earliest likely foreign source for each reused sentence. We analyze 1,037 STA and 237,551 FA articles from two time windows (October 7–November 2, 2023; February 1–28, 2025) and identify 1,087 aligned sentence pairs after filtering to the earliest sources. Reuse occurs in 52% of STA articles and 1.6% of FA articles and is predominantly non-literal, involving paraphrase and compositional reuse from multiple sources. Reused content tends to appear in the middle and end of English articles, while leads are more often original, indicating that simple lexical matching overlooks substantial editorial reuse. Compared with prior work focused on monolingual overlap, we (i) detect reuse across languages without requiring full translation, (ii) use publication timing to identify likely sources, and (iii) analyze where reused material is situated within articles. Dataset and code: https://github.com/kunturs/lrec2026-rewrite-news.
OnCoCo 1.0: A Public Dataset for Fine-Grained Message Classification in Online Counseling Conversations
Jens Albrecht | Robert Lehmann | Aleksandra Poltermann | Eric Rudolph | Philipp Steigerwald | Mara Stieler
Jens Albrecht | Robert Lehmann | Aleksandra Poltermann | Eric Rudolph | Philipp Steigerwald | Mara Stieler
This paper presents OnCoCo 1.0, a new public dataset for fine-grained message classification in online counseling. It is based on a new, integrative system of categories, designed to improve the automated analysis of psychosocial online counseling conversations. Existing category systems, predominantly based on Motivational Interviewing (MI), are limited by their narrow focus and dependence on datasets derived mainly from face-to-face counseling. This limits the detailed examination of textual counseling conversations. In response, we developed a comprehensive new coding scheme that differentiates between 38 types of counselor and 28 types of client utterances, and created a labeled dataset consisting of about 2.800 messages from counseling conversations. We fine-tuned several models on our dataset to demonstrate its applicability. The data and models are publicly available to researchers and practitioners. Thus, our work contributes a new type of fine-grained conversational resource to the language resources community, extending existing datasets for social and mental-health dialogue analysis.
Predicting Social Media User Actions: A Hybrid Approach for Common and Rare Behavior Prediction on Bluesky
Benjamin White | Anastasia Shimorina
Benjamin White | Anastasia Shimorina
Understanding and predicting user behavior on social media platforms is crucial for content recommendation and platform design. While existing approaches focus primarily on common actions like retweeting and liking, the prediction of rare but significant behaviors remains largely unexplored. This paper presents a hybrid methodology for social media user behavior prediction that addresses both frequent and infrequent actions across a diverse action vocabulary. We evaluate our approach on a large-scale Bluesky dataset containing 6.4 million conversation threads spanning 12 distinct user actions across 25 persona clusters. Our methodology combines four complementary approaches: (i) a lookup database system based on historical response patterns; (ii) persona-specific LightGBM models with engineered temporal and semantic features for common actions; (iii) a specialized hybrid neural architecture fusing textual and temporal representations for rare action classification; and (iv) generation of text replies. Our persona-specific models achieve an average macro F1-score of 0.64 for common action prediction, while our rare action classifier achieves 0.56 macro F1-score across 10 rare actions. These results demonstrate that effective social media behavior prediction requires tailored modeling strategies recognizing fundamental differences between action types. Our approach achieved first place in the SocialSim: Social-Media Based Personas challenge organized at the Social Simulation with LLMs workshop at the Conference on Language Modeling (COLM 2025).
The Data Acquisition Framework: Bridging Psychometrics and NLP for Personality Dataset Construction
Lorenz Dumanski | Michael Spranger | Melanie Siegel
Lorenz Dumanski | Michael Spranger | Melanie Siegel
Existing datasets for personality recognition in Natural Language Processing (NLP) suffer from documented quality problems: self-reported labels lacking psychometric validation, limited domain diversity and lack of context. Despite these known limitations, state-of-the-art approaches continue relying on the same datasets due to absence of alternatives. We present the Data Acquisition Framework (DAF), which addresses this gap by systematically translating psychometric questionnaire items into controlled communication scenarios through expert-community validation. DAF-items, validated scenario descriptions with contextual parameters, are deployed via the Automatic Data Acquisition and Annotation Tool (ADAAT). Participants complete personality surveys and engage in scenario-based text interactions with LLM personas configured to the DAF-Item context. This yields communication data with direct, item-level psychometric annotations.
Language Ideologies in a Multilingual Society: An LLM-based Analysis of Luxembourgish News Comments
Emilia Milano | Alistair Plum | Yves Scherrer | Christoph Purschke
Emilia Milano | Alistair Plum | Yves Scherrer | Christoph Purschke
Detecting language ideologies is a valuable yet complex task for understanding how identities are constructed through discourse. In Luxembourg’s multicultural and multilingual society, language ideologies reflect more than simple preferences: they carry deep cultural and social meanings, shaping identities and social belonging. Following recent developments in applying Natural Language Processing tools to linguistics and social science, this paper explores the potential of large language models to assist in the detection of language ideologies. We manually annotate a corpus of user comments in Luxembourgish with predefined ideological categories and then evaluate the performance of large language models under varying prompt conditions to assess their ability to replicate these human annotations. Since Luxembourgish is a small language and poorly represented in the LLMs’ training data, we also investigate whether machine-translating the data to high-resource languages increases performance on the ideology detection task. Our findings suggest that, while LLMs are not yet fully optimized for a multi-class ideological annotation task, they are practical tools to identify language ideological content.
Personality Anchoring for Social Simulation: Linking Personality, Social Behavior, and Interaction Success with LLM Agents
Vahid Sadiri Javadi | Aksa Aksa | Fryderyk Karol Róg | Lucie Flek | Johanne Trippas
Vahid Sadiri Javadi | Aksa Aksa | Fryderyk Karol Róg | Lucie Flek | Johanne Trippas
Social interactions are shaped by the interplay of dispositional traits and situational context, yet systematically investigating how personality configurations between individuals jointly influence social behavior across diverse social contexts remains methodologically challenging. We address this gap by introducing a simulation pipeline adapted from the CHARISMA framework, which employs well-known movie characters and public figures as psychologically grounded agents for multi-LLM social simulation using a method we term personality anchoring. We present a large-scale empirical study examining how dyadic Agreeableness composition influences social interaction outcomes across 1,010 simulated conversations. Our results reveal a monotonic relationship between dyadic Agreeableness composition and shared goal achievement, with Homogeneous-Agreeable pairs achieving success 10 times the rate of Homogeneous-Disagreeable pairs (62% vs. 6%). Behavioral mediation analysis reveals that Agreeableness shapes goal achievement partially through cooperative strategy selection, though it continues to predict outcomes within the same dominant strategy, indicating pathways beyond observable conversational behavior. Robustness analyses confirm high consistency of results across repeated simulations (ICC = 0.89) and stable personality expression across diverse scenarios, validating personality anchoring as a viable operationalization strategy.
up
Proceedings of Learning Non-Literal Expressions with Small Data @ LREC 2026
Proceedings of Learning Non-Literal Expressions with Small Data @ LREC 2026
Markus Egg | Valia Kordoni
Markus Egg | Valia Kordoni
Challenges in Japanese Euphemism Classification: An Analysis of Pretrained Japanese and Multilingual Models
Noriko Takahashi | Whitney Poh | Libby Barak | JIng Peng | Anna Feldman
Noriko Takahashi | Whitney Poh | Libby Barak | JIng Peng | Anna Feldman
Euphemisms present a persistent challenge for NLP because their interpretation depends on pragmatic inference, social norms, and contextual cues rather than surface meaning alone. Although Potentially Euphemistic Terms (PET)-based resources have been developed for several languages, Japanese euphemisms remain computationally unexplored despite their close interaction with honorifics, register variation, and orthographic choice. We introduce JP-PET, the first PET-based dataset for Japanese euphemism classification, comprising 1,672 annotated sentences across 101 PETs and ten semantic domains with register metadata. We evaluate two Japanese monolingual transformer models (Rinna RoBERTa and Tohoku BERT) and the multilingual XLM-R under three controlled PET-level data splits that isolate lexical familiarity and generalization to unseen euphemisms. While models achieve strong performance when PETs are shared between training and test data, performance drops substantially under PET-disjoint conditions, indicating reliance on lexical familiarity. Error analysis reveals systematic challenges in politically conventionalized expressions, metaphor-based euphemisms, and orthographic mitigation strategies. JP-PET provides the first benchmark for studying pragmatic meaning in Japanese NLP.
Steering Pragmatic Interpretation in LLMs: A Diagnostic Evaluation of Few-Shot and Reasoning-Based Prompting for Indirect Speech Acts.
Massimiliano Orsini | Dominique Brunato
Massimiliano Orsini | Dominique Brunato
Pragmatic competence poses a persistent challenge for large language models, as it requires context-dependent inference beyond literal meaning. This study examines whether few-shot prompting can reliably steer LLMs toward appropriate interpretations of indirect speech acts under small-data conditions. Focusing on Italian, we evaluate three LLMs on a small dataset that captures pragmatic ambiguity through graded plausibility judgments. We compare a zero-shot baseline with multiple few-shot prompting configurations that vary in the number and composition of demonstrations, as well as in the presence of explicit pragmatic guidance. Results show that few-shot prompting does not yield robust or monotonic improvements overall. While performance improves substantially for conventionalized indirect speech acts, gains for non-conventionalized indirect speech acts are unstable and limited. In contrast, introducing explicit pragmatic reasoning along with demonstrations through guided chain-of-thought prompting appears more promising. Overall, these findings highlight the limits of example-based steering for pragmatic inference and suggest that explicitly modeling pragmatic reasoning may be a more effective direction in small-data settings.
Injecting Structured Lexicographic Knowledge into LLMs for Non-Literal Expression Disambiguation: A Controlled Study on Croatian
Slobodan Beliga | Ivana Filipović Petrović | Ana Meštrović
Slobodan Beliga | Ivana Filipović Petrović | Ana Meštrović
In potentially idiomatic expressions (PIEs), the same surface form may receive either a literal or an idiomatic interpretation depending on context, making automatic literal–idiomatic disambiguation challenging. This is acute for Croatian, where annotated data and locally runnable generative models are limited. We present a study of Croatian PIE literal–idiomatic disambiguation examining how structured lexicographic knowledge can improve open-weight, decoder-only LLMs without fine-tuning. Using a new expert-annotated concordance dataset – CroPIEs, we compare baseline prompting to inference-time knowledge injection via retrieval-augmented generation (RAG) from a Croatian phraseological dictionary. We isolate the contribution of three knowledge types: definitional knowledge (structured meanings), contextual knowledge as curated prototypical usage examples, and their combination. Results show consistent improvements in macro-F1 for both GaMS-2B-Instruct and GaMS-9B-Instruct models. Definitional knowledge is generally more stable than examples alone, while examples can be effective but less consistent across expressions. The strongest and most reliable gains are obtained when definitions and examples are combined, indicating a synergistic effect between explicit meaning descriptions and contextual cues. Per-class analyses show that injected lexicographic evidence mitigates baseline biases between Literal and Idiomatic predictions, improving decision balance in a low-resource setting with small data of compact, expert-curated lexicographic evidence injected at inference time.
Metaphor Identification in Spanish Oncological Discourse: The Role of Explicit Meaning in Low-Resource Settings
Lucia Pitarch | Jordi Bernad | Gemma Bel-Enguix
Lucia Pitarch | Jordi Bernad | Gemma Bel-Enguix
Metaphor identification remains challenging in specialized and low-resource domains, where large annotated datasets are unavailable and general-domain models often fail to transfer effectively. In this paper, we evaluate FLAVORS-AECC, a Spanish dataset of oncological discourse that provides transparent, instance-level annotations of basic meaning (BM) and contextual meaning (CM) following the Metaphor Identification Procedure (MIP). We test the state-of-the-art Contrast-WSD model under two splits: a random split and a lemma-based split to control for lexical memorization. We compare three configurations: (i) a control model with no meaning information, (ii) manually curated basic meanings, and (iii) first dictionary entry as an approximation of basic meaning. Results show that explicitly modeling meaning contrast substantially improves performance in low-resource settings (from below 0.30 to above 0.50 F1). However, contrary to expectations, manually annotated BM does not consistently outperform first dictionary entries, suggesting that definition length rather than theoretical fidelity may introduce noise. We also find that models perform best on cases with high annotator agreement and that verbs remain the most challenging part of speech. Overall, our findings highlight the importance of linguistically grounded modeling for metaphor detection in specialized domains.
Exploring Detection of Complex, Non-Literal Expressions of Cultural Motifs
Ibrahim H. Alyami | Mark A. Finlayson
Ibrahim H. Alyami | Mark A. Finlayson
Motifs are non-commonplace, recurring narrative elements, often found originally in folk stories and also in modern news, literature, and propaganda. Expressions of motifs in text can be most straightforwardly classified as simple or complex. Simple motif expressions are easy to detect because they almost always appear in a single sentence using the same words as the motif definition itself. However, complex motifs are strongly non-literal and often spread across multiple sentences, thus requiring more context to understand. We propose a baseline system to detect complex motif expressions that have challenged prior work. We used an annotated corpus that identified 992 complex motif expressions of 155 different motifs for training and testing. We tested five different generative approaches that included varying amounts of context: a single sentence baseline (from prior work); a window of 3 or 5 sentences; the entire story; or the entire story with the target sentence identified. We fine-tuned four off-the-shelf open-source LLMs using LoRA under these conditions. Somewhat surprisingly, we report a negative result: our experiments show that in our generative setup more context does not improve the detection of complex motifs. We speculate on why this might be so and identify directions for future research.
Artful Writing, Authentic Emotions: Distinguishing Human-Written from LLM-Generated Metaphors by Annotation and Classification
Michaela Regneri | Nooshin Aghajari | Thomas Kroedel
Michaela Regneri | Nooshin Aghajari | Thomas Kroedel
We analyze differences between human-written and automatically generated metaphors. Using two syntactically standardized datasets containing novel metaphors from poetry and science communication, we generate new figurative expressions with LLMs that describe the same concepts as human-written texts. Using crowdsourcing, we conduct extensive annotation across multiple dimensions (e.g., writing quality and creativity) and ask annotators to judge whether the metaphor was generated automatically. For the poetry set, we also asked annotators for the emotions conveyed by the metaphor. We find that, consistent with prior work, the authorship of scientific metaphors is difficult to determine. However, our results reveal that human-written poetic metaphors stand out by their capacity to convey emotion. We also analyze which types of metaphors are merely perceived as human. Finally, we show that, while human annotators cannot distinguish human from machine metaphors, automated approaches achieve high accuracy in identifying human writers, which suggests substantial differences in text structure.
Creation and Validation of a Monolingual Spanish NLI Dataset for Metaphor Interpretation via Model-in-the-Loop
Alec Sanchez-Montero | Gemma Bel-Enguix | Sergio Luis Ojeda Trueba
Alec Sanchez-Montero | Gemma Bel-Enguix | Sergio Luis Ojeda Trueba
Large Language Models (LLMs) can easily generate fluent text, but assessing whether they truly understand metaphors requires moving beyond English-centric datasets and binary token classification tasks. To test if current state-of-the-art models perform genuine structural alignment and analogical reasoning rather than just echoing statistical token co-occurrence, we introduce a new monolingual Spanish Natural Language Inference (NLI) dataset specifically built for metaphor interpretation. Using a Model-in-the-Loop approach, we reconstruct the literal truth conditions of metaphors sourced from science texts. Before human experts curated the data, we performed an ablation study—evaluated via BERTScore and Cross-Entropy—to test whether explicit symbolic scaffolding improves analogical reasoning. While automated evaluations suggested that forcing models to follow explicit metaphorical rules diminished their fluency and increased text surprisal, human evaluation revealed the opposite: this explicit guidance produced far more accurate and strictly literal outputs. This reveals a limitation in how we evaluate NLU: automated metrics consistently penalize the cognitive ‘heavy lifting’ required to resolve a metaphor, simply because they are built to reward surface-level statistical fluency. By releasing this resource, we aim to shift the focus from surface-level generation to real cognitive alignment and metaphorical understanding in Spanish NLU.
Metonymy, often considered as a figurative trope, is a frequently occurring linguistic phenomenon in which an entity is replaced by a semantically related entity. Named entities are commonly used to refer to associated concepts. For instance, in the sentence India signed a treaty, the geographical name India stands metonymically for the government rather than the physical location. This study develops a hybrid architecture to classify literal and metonymic usages of named entities in Marathi language using small data. The approach integrates Pustejovsky’s Generative Lexicon framework with linguistic features, including part-of-speech tags, named entity labels, and lemmas. The model is evaluated on 890 sentences and achieved F1 scores of 66.98% and 71.97% for literal and metonymic instances, respectively. The study highlights the effectiveness of the features in capturing metonymic contexts, though precision remains a target for improvement. Ablation results confirm that the Formal and Constitutive Qualia roles are the most critical components for detecting metonymic shifts, while the Telic role introduces modest noise in the present corpus. This experiment shows the scope for developing hybrid models for learning non-literal language using small data, which could be beneficial for less-explored and low-resource languages.
Contextualising (Im)plausible Events Triggers Figurative Language
Annerose Eichel | Tonmoy Rakshit | Sabine Schulte im Walde
Annerose Eichel | Tonmoy Rakshit | Sabine Schulte im Walde
This work explores the connection between (non-)literalness and plausibility at the example of subject-verb-object events in English. We design a systematic setup of plausible and implausible event triples in combination with abstract and concrete constituent categories. Our analysis of human and LLM-generated judgments and example contexts reveals substantial differences between assessments of plausibility. While humans excel at nuanced detection and contextualization of (non-)literal vs. implausible events, LLM results reveal only shallow contextualization patterns with a bias to trade implausibility for non-literal, plausible interpretations.
A Novel Dataset and Three Ways to Approach Automatic Metaphor Detection in German Religious Online Forums
Sebastian Reimann | Tatjana Scheffler
Sebastian Reimann | Tatjana Scheffler
In recent years, automatic metaphor detection has received considerable attention within NLP. However, the largest share of research, including most datasets annotated for metaphor, has concentrated on English and a limited set of genres. Automatic metaphor detection for a genre like religious online communication, which is particularly rich in metaphor, remains understudied, in particular since annotated data for this genre is lacking in the first place. This paper aims to close these gaps by offering a novel dataset of posts from German online forums annotated for metaphor, which opens up new research opportunities for automatic metaphor detection for German. Moreover, we present an in-depth exploration in which we evaluate the suitability of different strategies to overcome the relative lack of training data for this task by comparing cross-lingual and cross-genre transfer strategies with the use of LLM prompting. We find that fine-tuning encoder-only language models outperforms the prompting-based approach, that different architectures based on contextual embeddings indeed exhibit considerable differences in their behavior and that smaller in-genre data may be preferable for certain use cases over fine-tuning on larger datasets from different genres.
Decomposing Creativity: Two Small Datasets Combining Originality Ratings and Metaphor Annotations
Emilie Sitter | Sina Zarrieß | Omar Momen | Berenike Herrmann
Emilie Sitter | Sina Zarrieß | Omar Momen | Berenike Herrmann
We introduce MetaphOrig, a small dataset comprising two genre-specific collections of spatial descriptions for the study of linguistic creativity and Non-Literal Expressions (NLEs). The sentence-level spatial descriptions were extracted from two distinct genre- and time-specific source corpora. Both source corpora comprise German texts: literary prose from the 18th to 20th century (KOLIMO) and factual travel reports from the 21st century (Wikivoyage). Along with the spatial descriptions, the dataset contains sentence-level originality ratings obtained through crowdsourcing and from four different LLMs (GPT-5, Qwen2.5-32B-Instruct, Mistral-Small-3.2-24B-Instruct, and Llama-3.2-3B), and word-level metaphor annotations. We provide the MetaphOrig datasets, including all annotations, to the community. The subsets can be used for further research on linguistic creativity or metaphor, either in one specific textual domain or comparatively across the two domains. We conduct an illustrative study on the dataset, treating originality as a proxy of textual creativity. In both subsets, we investigate potential correlations between sentence-level originality ratings and the density of metaphorical expressions within each sentence. We find the correlation to be present only in the KOLIMO subset. A comparison of human and LLM originality ratings shows that this pattern holds for both types of ratings.
up
Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026
Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026
Georg Rehm | Stefan Dietze | Danilo Dessi | Diana Maynard | Sonja Schimmler
Georg Rehm | Stefan Dietze | Danilo Dessi | Diana Maynard | Sonja Schimmler
AstroConcepts: A Large-Scale Multi-Label Classification Corpus for Astrophysics
Atilla Kaan Alkan | Felix Grezes | Sergi Blanco-Cuaresma | Jennifer Lynn Bartlett | Daniel Chivvis | Anna Kelbert | Kelly Lockhart | Alberto Accomazzi
Atilla Kaan Alkan | Felix Grezes | Sergi Blanco-Cuaresma | Jennifer Lynn Bartlett | Daniel Chivvis | Anna Kelbert | Kelly Lockhart | Alberto Accomazzi
Scientific multi-label text classification suffers from extreme class imbalance, where specialized terminology exhibits severe power-law distributions that challenge standard classification approaches. Existing scientific corpora lack comprehensive controlled vocabularies, focusing instead on broad categories and limiting systematic study of extreme imbalance. We introduce AstroConcepts, a corpus of English abstracts from 21,702 published astrophysics papers, labeled with 2,367 concepts from the Unified Astronomy Thesaurus. The corpus exhibits severe label imbalance, with 76 % of concepts having fewer than 50 training examples. By releasing this resource, we enable systematic study of extreme class imbalance in scientific domains and establish strong baselines across traditional, neural, and vocabulary-constrained LLM methods. Our evaluation reveals three key patterns that provide new insights into scientific text classification. First, vocabulary-constrained LLMs achieve competitive performance relative to domain-adapted models in astrophysics classification, suggesting a potential for parameter-efficient approaches. Second, domain adaptation yields relatively larger improvements for rare, specialized terminology, although absolute performance remains limited across all methods. Third, we propose frequency-stratified evaluation to reveal performance patterns that are hidden by aggregate scores, thereby making robustness assessment central to scientific multi-label evaluation. These results offer actionable insights for scientific NLP and establish benchmarks for research on extreme imbalance.
Benchmarking LLMs for ARR Area Assignment: Evidence and Implications for Assignment Strategies
Eileen Bingert | Diego Alves | Stefania Degaetano-Ortlieb
Eileen Bingert | Diego Alves | Stefania Degaetano-Ortlieb
We study how large language models (LLMs) perform at assigning ACL Rolling Review (ARR) areas from paper titles/abstracts. Using 558 papers (ACL/EACL/NAACL, 2020 to 2025), we compare multiple LLMs and prompting schemes (zero/few-shot; with/without ARR keywords; each-category variants) and analyze per-area scores, error overlap, and confusion matrices. One-shot prompting (with OpenAI-gpt-oss-20b) tends to perform best, while injecting ARR keywords often lowers accuracy. Task-bounded areas (e.g., MT, IE, QA, Summarization) are predicted more reliably, whereas broad, cross-cutting labels (e.g., Resources and Evaluation, NLP Applications) are frequently conflated, indicating taxonomy ambiguity rather than solely model limitations. We recommend hierarchical or primary-plus-secondary labels to reduce ambiguity and improve reviewer matching. Our dataset, methods, and findings offer a reproducible baseline for area selection support in ACL workflows.
Benchmarking Retrieval-Augmented Generation for Scientific Knowledge QA in European Portuguese
Jose Matos | Catarina Silva | Hugo Goncalo Oliveira
Jose Matos | Catarina Silva | Hugo Goncalo Oliveira
Retrieval-Augmented Generation (RAG) enables grounding of model outputs in external evidence, but its impact on European Portuguese (pt-PT) scientific question answering (QA) remains unclear. We present a controlled evaluation of RAG on pt-PT knowledge QA across different scientific domains using the Portuguese test split of the Global MMLU Lite dataset. As external evidence, we use a Portuguese scientific literature knowledge base containing over 32,000 documents converted to Markdown. We benchmark five instruction-tuned small language models (4-12B) and compare closed-book baselines against 16 RAG configurations that vary by: (i) dense retriever specialization (multilingual vs. Portuguese-specific), (ii) reranking (on/off), and (iii) number of retrieved chunks (k ∈ 1, 3, 5, 10). Results suggest that RAG gains are model-dependent. Some models improve consistently, others are highly sensitive to retrieval choices, and some degrade under retrieval noise, especially at larger values of k. Findings highlight the importance of model-specific retrieval tuning and ensuring that the retriever and reranker languages and domains align when deploying RAG systems for Portuguese natural scientific language processing.
Beyond Abstracts: A Biomedical MeSH Indexing Corpus Incorporating Summarized Methods Sections
Sujoy Datta | Robert E. Mercer | Xindi Wang
Sujoy Datta | Robert E. Mercer | Xindi Wang
Automated Medical Subject Heading (MeSH) indexing systems rely predominantly on titles and abstracts, while human indexers at the National Library of Medicine examine full-text articles—particularly Methods sections—that often contain crucial experimental terminology absent from abstracts. This information asymmetry limits model performance and prevents detection of methodologically-grounded MeSH descriptors. We introduce a novel biomedical MeSH indexing corpus comprising over one million English biomedical articles, each annotated with title, abstract, journal metadata, publication year, expert-curated MeSH terms, and—uniquely—extractive summaries of Methods sections. Using LLaMA 3 with an iterative re-prompting strategy, we generated high-fidelity summaries. To avoid label leakage, evaluation labels are inferred using journal-specific MeSH frequency profiles rather than gold annotations. This publicly accessible dataset addresses a critical gap in full-text MeSH indexing research. Building upon this resource, we propose an extended multi-channel neural architecture that incorporates Methods-derived representations. Empirical results demonstrate consistent performance gains across both example-based and label-based evaluations, indicating better retrieval of infrequent terms. These findings highlight that procedural knowledge in the Methods section encodes critical semantic cues overlooked by title-abstract only models.
Challenges and Opportunities for NSLP in Scientific Publishing – A Case Study
Thomas Kleinbauer | Michael Didas | Michael Wagner
Thomas Kleinbauer | Michael Didas | Michael Wagner
Research software is not always meant to reach production-grade quality. The same requirements, for instance, regarding performance, security, or reliability that are imposed on professional software do not necessarily apply in the lab. However, recent years have seen an increased interest to bring cutting-edge research results into production. Specialized natural language processing, such as NSLP, is no exception. In this paper, we discuss three real-life challenges in scientific publishing as they relate to NSLP, and also highlight the potential NSLP has to offer in overcoming these challenges. Specifically, we identify issues related to the elicitation of metadata – particularly with respect to what we term the /metadata externality problem –, privacy and data protection laws, and software reliability. This is not a research paper; rather than introducing novel research results, our intention is to contribute to the academic discourse in the NSLP community by providing the perspective of a potential end-user.
ClimateCheck 2026: Scientific Fact-Checking and Disinformation Narrative Classification of Climate-related Claims
Raia Abu Ahmad | Max Upravitelev | Aida Usmanova | Veronika Solopova | Georg Rehm
Raia Abu Ahmad | Max Upravitelev | Aida Usmanova | Veronika Solopova | Georg Rehm
Automatically verifying climate-related claims against scientific literature is a challenging task, complicated by the specialised nature of scholarly evidence and the diversity of rhetorical strategies underlying climate disinformation. ClimateCheck 2026 is the second iteration of a shared task addressing this challenge, expanding on the 2025 edition with tripled training data and a new disinformation narrative classification task. Running from January to February 2026 on the CodaBench platform, the competition attracted 20 registered participants and 8 leaderboard submissions, with systems combining dense retrieval pipelines, cross-encoder ensembles, and large language models with structured hierarchical reasoning. In addition to standard evaluation metrics (Recall@K and Binary Preference), we adapt an automated framework to assess retrieval quality under incomplete annotations, exposing systematic biases in how conventional metrics rank systems. A cross-task analysis further reveals that not all climate disinformation is equally verifiable, potentially implicating how future fact-checking systems should be designed.
This paper describes our submission to the ClimateCheck 2026 shared task on scientific fact-checking of climate-related claims (Task 1) and disinformation narrative classification (Task 2). For Task 1, we use a three-stage pipeline combining BM25 retrieval over 394,269 scientific abstracts, ensemble re-ranking with five fine-tuned BGE cross-encoders aggregated via Reciprocal Rank Fusion, and zero-shot claim verification using the gpt-oss-120b model. For Task 2, we use a zero-shot approach with a custom prompt based on the CARDS taxonomy and the gpt 5.2 model. Our system outperforms the organizers’ baseline across all subtasks. Furthermore, we achieve the best result for the Task 1.1 (retrieval score of 0.466), and Task 1.2 (verification score of 1.183 - F1 + Recall@5), and the third best result for Task 2 (Macro F1 of 0.583).
Comparing LLM-Based Knowledge Graph Extraction Approaches on Literary Studies in Spanish: A Case Study on Orbis Tertius
Federico Cortes
Federico Cortes
Knowledge graph construction from scholarly text increasingly relies on large language models, yet different extraction architectures produce different graphs. Literary studies poses particular challenges: meaning is interpretive rather than factual, and the boundaries of relevant knowledge are determined by hermeneutic frameworks rather than empirical verification. We compare two LLM-based extraction frameworks—entity-anchored extraction (KGGen) and open extraction with schema canonicalization (EDC)—on 472 Spanish-language literary studies articles from Orbis Tertius (1996–2024). Despite fundamental architectural differences, both methods converge on key findings: cultural framing dominates literary discourse by 2.2–2.5× over textual framing (p < .001), and core author networks remain consistent across approaches. The methods diverge in entity composition: KGGen captures more proper names (40.7% vs. 18.7%), while EDC captures more abstract concepts (42.8%) and preserves Spanish predicates with 21,025 semantic definitions. Convergent findings across architecturally different methods merit higher confidence, and we identify methodological considerations for knowledge graph construction from humanities scholarship.
Demystifying Funding: Reconstructing a Unified Dataset of the UK Funding Lifecycle
William Thorne | Rupert Shepherd | Diana Maynard
William Thorne | Rupert Shepherd | Diana Maynard
We present a reconstruction of UKRI’s Gateway to Research (GtR) database that links funding opportunities to their resulting project proposals through panel meeting outcomes. Unlike existing work that focuses primarily on funded projects and their outcomes, we close the complete funding lifecycle by integrating three previously disconnected data sources: the GtR project database, UKRI funding opportunities, and competitive funding decision records across UKRI’s research councils. We describe the technical challenges of data collection, including navigating inconsistent publication formats and restricted access to panel decisions. The resulting dataset enables a holistic interrogation of the entire funding process, from opportunity announcement to research outcomes. We release the database and associated code.
Do Lexical and Contextual Coreference Resolution Systems Degrade Differently under Mention Noise? An Empirical Study on Scientific Software Mentions
Atilla Kaan Alkan | Felix Grezes | Jennifer Lynn Bartlett | Anna Kelbert | Kelly Lockhart | Alberto Accomazzi
Atilla Kaan Alkan | Felix Grezes | Jennifer Lynn Bartlett | Anna Kelbert | Kelly Lockhart | Alberto Accomazzi
We present our participation in the SOMD 2026 shared task on cross-document software mention coreference resolution, where our systems ranked second across all three subtasks. We compare two fine-tuning-free approaches: Fuzzy Matching (FM), a lexical string-similarity method, and Context Aware Representations (CAR), which combines mention-level and document-level embeddings. Both achieve competitive performance across all subtasks (CoNLL F1 of 0.94–0.96), with CAR consistently outperforming FM by 1 point on the official test set, consistent with the high surface regularity of software names, which reduces the need for complex semantic reasoning. A controlled noise-injection study reveals complementary failure modes: as boundary noise increases, CAR loses only 0.07 F1 points from clean to fully corrupted input, compared to 0.20 for FM, whereas under mention substitution, FM degrades more gracefully (0.52 vs. 0.63). Our inference-time analysis shows that FM scales superlinearly with corpus size, whereas CAR scales approximately linearly, making CAR the more efficient choice at large scale. These findings suggest that system selection should be informed by both the noise profile of the upstream mention detector and the scale of the target corpus. We release our code to support future work on this underexplored task.
Do We Need Bigger Models for Science? Task-Aware Retrieval with Small Language Models
Florian Kelber | Matthias Jobst | Yuni Susanti | Michael Färber
Florian Kelber | Matthias Jobst | Yuni Susanti | Michael Färber
Scientific knowledge discovery increasingly relies on large language models, yet many existing scholarly assistants depend on proprietary systems with tens or hundreds of billions of parameters. Such reliance limits reproducibility and accessibility for the research community. In this work, we ask a simple question: do we need bigger models for scientific applications? Specifically, we investigate to what extent carefully designed retrieval pipelines can compensate for reduced model scale in scientific applications. We design a lightweight retrieval-augmented framework that performs task-aware routing to select specialized retrieval strategies based on the input query. The system further integrates evidence from full-text scientific papers and structured scholarly metadata, and employs compact instruction-tuned language models to generate responses with citations. We evaluate the framework across several scholarly tasks, focusing on scholarly question answering (QA), including single- and multi-document scenarios, as well as biomedical QA under domain shift and scientific text compression. Our findings demonstrate that retrieval and model scale are complementary rather than interchangeable. While retrieval design can partially compensate for smaller models, model capacity remains important for complex reasoning tasks. This work highlights retrieval and task-aware design as key factors for building practical and reproducible scholarly assistants.
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
Léane Jourdan | Julien Aubert-Béduchaud | Yannis Chupin | Marah Baccari | Florian Boudin
Léane Jourdan | Julien Aubert-Béduchaud | Yannis Chupin | Marah Baccari | Florian Boudin
Scientific writing is an iterative process that generates rich revision traces, yet publicly available resources typically expose only final or near-final versions of papers. This limits empirical study of revision behaviour and evaluation of large language models (LLMs) for scientific writing. We introduce EarlySciRev, a dataset of early-stage scientific text revisions automatically extracted from arXiv LaTeX source files. Our key observation is that commented-out text in LaTeX often preserves discarded or alternative formulations written by the authors themselves. By aligning commented segments with nearby final text, we extract paragraph-level candidate revision pairs and apply LLM-based filtering to retain genuine revisions. Starting from 1.28M candidate pairs, our pipeline yields 578k validated revision pairs, grounded in authentic early drafting traces. We additionally provide a human-annotated benchmark for revision detection. EarlySciRev complements existing resources focused on late-stage revisions or synthetic rewrites and supports research on scientific writing dynamics, revision modelling, and LLM-assisted editing.
Enhancing Factuality and Transparency in Generative Models for Biomedical Question Answering
Ankita Behura | Siting Liang | Daniel Sonntag
Ankita Behura | Siting Liang | Daniel Sonntag
Biomedical Question Answering (BQA) systems are vital for providing clinicians and researchers with efficient access to large amount of biomedical scientific studies. Existing automated BQA models, however, often rely on complex hybrid architectures to handle diverse question and answer formats, leading to inefficiency and high complexity. While domain-specific generative language models like BioBART offer a unified and simplified alternative capable of producing fluent human-like responses, they are prone to hallucination and lack interpretability, undermining their trustworthiness in critical healthcare domains. To address these limitations, this work introduces an enhanced model that augments BioBART with a pointer network for accurate token copying and a novel Keyphrase Filter (KPF) to guide attention toward critical information during generation. Experimental results on the BioASQ challenge demonstrate that the proposed Pointer-KPF model significantly outperforms the baseline BioBART, particularly on metrics for ideal answers. Furthermore, our evaluation shows that the model enhances transparency: pointer-guided attention heatmaps reveal improved input-output alignment, while keyphrase scores act as saliency maps to identify the most influential input segments. This approach not only reduces hallucination by strengthening textual grounding but also provides crucial insights into the model’s reasoning, thereby increasing confidence and trust in its outputs.
Enhancing Scholarly Knowledge Graphs via Domain-Specific Entity Detection and Linking
Nicolau Duran-Silva | César A. Parra-Rojas | Pablo Accuosto | Julian Moreno-Schneider | Georg Rehm
Nicolau Duran-Silva | César A. Parra-Rojas | Pablo Accuosto | Julian Moreno-Schneider | Georg Rehm
Navigating scholarly content presents important challenges due to the fragmented and heterogeneous nature of research production and outputs. Scholarly Knowledge Graphs offer an efficient means to integrate diverse data sources and consolidate knowledge across outputs in a structured manner. This representation, combined with the grounding of unstructured textual data to well-defined research-related concepts, has great potential for enhancing knowledge discovery and supporting researchers navigating through vast amounts of scientific information. Knowledge extraction capabilities are commonly limited by the availability of large collections of annotated data supporting named-entity recognition (NER) and linking (EL), and the enormous effort that their elaboration entails for domain experts. Recent advances in natural language processing and generative artificial intelligence provide valuable opportunities to reduce the data annotation toll and produce high-quality NER with minimal expert involvement. Here, we present a pipeline for domain-specific NER and EL, leveraging LLMs and knowledge from experts in a human-in-the-loop approach to streamline the annotation process, along with transformer-based models and few-shot techniques. While the application focuses on showcasing four specific domains, the pipeline is designed to be flexible and domain agnostic for scientific fields.
Evaluating Generative Large Language Models for Portuguese Scientific Information Extraction
Tomás Pinto | Catarina Silva | Hugo Goncalo Oliveira
Tomás Pinto | Catarina Silva | Hugo Goncalo Oliveira
Scientific Information Extraction (IE), which identifies entities and their relations from scientific texts, is essential for building Scientific Knowledge Graphs (SciKGs) that encode structured knowledge and enable applications such as semantic search, question answering, and literature reasoning. Large Language Models (LLMs) have shown strong capabilities in processing unstructured text, yet most advances focus on English, with limited exploration for less-resourced languages like Portuguese. The reliability of generative LLMs, including Portuguese-targeted models like the sovereign AMALIA, for structured extraction of scientific knowledge from literature text remains underexplored. We evaluate low- to mid-scale generative LLMs (8–12B parameters) on scientific Named Entity Recognition (NER) and Relation Extraction (RE), using a Portuguese-translated dataset of computer science article abstracts. Overall, our results show moderate performance and indicate that the adaptation strategy has a greater impact than model choice: prompting yields unstable performance and poor RE scores, while fine-tuning consistently improves both NER and RE and reduces cross-model variability. These findings suggest that, at this scale, prompting alone is insufficient for SciKG construction and underscore the need for supervised adaptation. We provide a detailed error analysis and outline directions for advancing Portuguese scientific IE.
From Slides to Chatbots: Enhancing Large Language Models with University Course Materials
Tu Anh Dinh | Philipp Nicolas Schumacher | Jan Niehues
Tu Anh Dinh | Philipp Nicolas Schumacher | Jan Niehues
Large Language Models (LLMs) have advanced rapidly in recent years. One application of LLMs is to support student learning in educational settings. However, prior work has shown that LLMs still struggle to answer questions accurately within university-level computer science courses. In this work, we investigate how incorporating university course materials can enhance LLM performance in this setting. A key challenge lies in leveraging diverse course materials such as lecture slides and transcripts, which differ substantially from typical textual corpora: slides also contain visual elements like images and formulas, while transcripts contain spoken, less structured language. We compare two strategies, Retrieval-Augmented Generation (RAG) and Continual Pre-Training (CPT), to extend LLMs with course-specific knowledge. For lecture slides, we further explore a multi-modal RAG approach, where we present the retrieved content to the generator in image form. Our experiments reveal that, given the relatively small size of university course materials, RAG is more effective and efficient than CPT. Moreover, incorporating slides as images in the multi-modal setting significantly improves performance over text-only retrieval. These findings highlight practical strategies for developing AI assistants that better support learning and teaching, and we hope they inspire similar efforts in other educational contexts.
Generating Research Data Metadata from Their Accompanying README Files
Kotaro Sekido | Yu Watanabe | Koichiro Ito | Shigeki Matsubara
Kotaro Sekido | Yu Watanabe | Koichiro Ito | Shigeki Matsubara
Software repositories have conventionally been used for software development. Recently, they have also served as research data repositories. Research data published in such repositories are frequently accompanied by README files; however, the data frequently lack structured metadata. To address this issue, this paper investigates the feasibility of generating research data metadata from their accompanying README files. First, we analyze the occurrence patterns of metadata-related information in README files. The results of this analysis demonstrated that README files could serve as valuable resources for metadata generation. We then performed an experiment on extracting metadata-related information from README files using large language models (LLMs) and evaluated their performance. The experimental results demonstrated that LLMs could extract metadata-related information with high performance.
Identifying Implicit Research Data References in Paper Citations
Koshi Motegi | Koichiro Ito | Shigeki Matsubara
Koshi Motegi | Koichiro Ito | Shigeki Matsubara
To encourage the public release of research data under open science, it is beneficial to establish mechanisms for evaluating research data based on metrics such as citation counts. In scholarly papers, authors sometimes cite papers that report the creation or release of research data instead of citing the research data themselves. In this paper, as a step toward computing citation counts of research data, we investigate the feasibility of identifying paper citations that refer to research data. We conducted an identification experiment using large language models and evaluated their performance.
Improving Completeness in Deep Research Agents through Targeted Enrichment
Jesse Wonnink | Jakub Zavrel | Paul Groth
Jesse Wonnink | Jakub Zavrel | Paul Groth
Deep research agents, AI systems that autonomously gather, synthesize, and report on complex topics, represent a significant advance in information synthesis, yet ensuring the completeness of their outputs remains an open challenge. A key bottleneck is query generation: current systems decompose research questions into subqueries via prompt engineering alone, offering no formal guarantees on diversity or coverage, which leads to redundant retrieval and gaps in the resulting reports. This paper presents HERO (High Enrichment Retrieval Orchestrator), a hierarchical deep research architecture that addresses this limitation through two complementary mechanisms. First, submodular optimization via a facility location objective provides mathematically grounded control over the relevance–diversity trade-off during query selection, replacing ad-hoc generation with provably diverse query sets. Second, a hierarchical enrichment stage independently analyzes each subquery pipeline’s intermediate synthesis for information gaps and issues targeted follow-up queries, enabling adaptive depth without cross-pipeline interference. We evaluate HERO across academic (ScholarQABench) and general-domain (DeepResearchGym) benchmarks. HERO achieves state-of-the-art coverage (Key Point Recall: 67.63), grounding (Citation F1: 91.57), and presentation quality on DeepResearchGym, and the highest scores on multi-paper synthesis tasks in ScholarQABench.
MioFFAn: An Annotation Software for Formula Formalization with LLM Automation Capabilities
Nicolas Sibuet Ruiz | Horacio Saggion | Riccardo Rossi
Nicolas Sibuet Ruiz | Horacio Saggion | Riccardo Rossi
The automatic translation of mathematical expressions in scientific literature into executable symbolic code—a process we refer to as Formula Formalization—is hindered by a severe scarcity of high-quality, ground-truth datasets specialized for technical scientific domains. In this paper, we present MioFFAn, an open-source, document-centric, and customizable framework designed to facilitate rapid annotation for this task. Building upon the MioGatto architecture, we extend existing features to overcome structural limitations and pivot its scope by introducing specific functionalities for Formula Formalization, such as selection of equations of interest and aided symbolic code specification. By allowing users to configure custom taxonomies and properties for identified symbols, and compatible symbolic operators, we ensure the framework is adaptable to diverse specialized scientific fields. Furthermore, MioFFAn is designed to incorporate partial automation via Large Language Models. By defining a modular set of automated sub-tasks with strict output formats, we enable researchers to iteratively refine automation capabilities and evaluate competing strategies using standard NLP metrics. We specify the current automation methodology and perform a preliminary evaluation that demonstrates to efficacy of this human-in-the-loop approach.
Normalizing Section Names and Structure of Scientific Articles
Nicolau Duran-Silva | Julian Moreno-Schneider | César A. Parra-Rojas | Georg Rehm
Nicolau Duran-Silva | Julian Moreno-Schneider | César A. Parra-Rojas | Georg Rehm
The growing amount of scientific literature has increased the need for automatic methods that can retrieve, process, and exploit scholarly content. In this work, we explore section name normalization and hierarchy prediction for scientific articles using a two-level taxonomy. We compare independent, sequential classification models, and generative large language models on the SASC dataset. Results show that classification approaches, particularly sequential models that employ document-level context, consistently outperform generative methods. Incorporating section content is essential for fine-grained classification, while generative models remain limited in zero-shot settings. Our experiments highlight the importance of structure-aware modelling for large-scale scholarly document processing, and the importance of section normalization for the development of advanced research mapping and research assessment tools.
Retrieval-Augmented LLMs and Encoder Models for Multi-Label Climate Disinformation Narrative Classification
Neda Foroutan | Alexandra Tsiakalou | Vera Schmitt
Neda Foroutan | Alexandra Tsiakalou | Vera Schmitt
The detection of climate misinformation narratives remains challenging due to label imbalance, hierarchical taxonomies, and the multi-label nature of real-world claims. Developing models that can reliably assign fine-grained narrative categories is therefore essential for scalable analysis of climate disinformation. We present our approach to multi-label climate misinformation narrative classification for ClimateCheck@NSLP 2026 Task 2. The task requires assigning one or more narrative categories, defined by the hierarchical CARDS taxonomy, to climate-related claims. We investigate both encoder-based transformers and decoder-only large language models (LLMs), comparing fine-tuning BERT-based models with prompt-based and retrieval-augmented instruction tuning strategies with Qwen3 model. To address data scarcity and label imbalance, we explore targeted augmentation using external CARDS-based resources as well as semantic similarity filtering. Our experiments show that augmentation improves encoder-based models, with ModernBERT achieving competitive performance at low computational cost. However, the strongest results are obtained using retrieval-augmented instruction tuning with Qwen3, which narrows the candidate narrative space prior to prediction. This approach achieves a Macro-F1 score of 59.72% on the official test set, securing second place on the leaderboard. These findings demonstrate the effectiveness of retrieval-guided LLM adaptation for structured multi-label narrative classification while highlighting the continued relevance of efficient encoder-based models.
The Linguist’s Lie Detector: Linguistic Knowledge in Large Language Models
Lucía Catalán Gris | Kim Gerdes | John S. Y. Lee
Lucía Catalán Gris | Kim Gerdes | John S. Y. Lee
We present a benchmark and evaluation pipeline for assessing how well large language models (LLMs) handle linguistic knowledge. Starting from a curated subcorpus of 11 syntax-focused articles published in Glossa: A Journal of General Linguistics (2016–2026), we design a pipeline that (1) segments article text into sentences, (2) extracts atomic, verifiable statements, and (3) classifies them into linguistic categories (language-specific, typological, theoretical, citation, or structural). Each stage is evaluated against human gold annotations produced by three annotators, with inter-annotator agreement measured via Krippendorff’s α and Cohen’s κ. We compare several LLMs on extraction and classification, using BERTScore-style similarity for extraction and macro F1 for classification. Finally, we generate contradictions of the true linguistic statements and test whether LLMs can distinguish true from false claims. On a challenge set of 705 linguistic statements, we compare eight LLMs, with Gemini 3 Flash achieving the highest F1 score of 0.66, indicating that current models possess limited but non-trivial linguistic knowledge.
The Software Mention Detection and Coreference Resolution Shared Task 2026
Sharmila Upadhyaya | Wolfgang Otto | Julia Matela | Frank Krüger | Stefan Dietze
Sharmila Upadhyaya | Wolfgang Otto | Julia Matela | Frank Krüger | Stefan Dietze
Software is referenced in research papers in many different ways: full names, abbreviations, misspellings, versioned names, or indirect references via websites and citations. This makes it hard to link mentions to a single software entity, which in turn limits large-scale analyses and knowledge graph construction. The Software Mention Detection and Coreference Resolution (SOMD) shared task 2026, organized at the Natural Scientific Language Processing (NSLP) workshop at LREC 2026, focuses on clustering software mentions that refer to the same software entity. We provide three subtasks covering gold mentions, automatically extracted mentions, and mentions sampled at scale from large-scale publications. Systems are evaluated with established coreference metrics (MUC, B3, CEAFe) and their CoNLL average. This paper describes the task setup, datasets, evaluation, baseline, and the observed patterns in participant submissions, and outlines future directions for scalable software mention coreference resolution. The shared task was concluded with total five registered participants, with total 43 submissions for all subtasks. Finally, two system papers were submitted with competitive performance against baselines.
Towards Efficient Self-Explainable Climate-Related Claim Verification with Generative Models
Siting Liang | Omar Adjali | Daniel Sonntag
Siting Liang | Omar Adjali | Daniel Sonntag
In this work, we present an empirical investigation into two self-explanatory inference paradigms using pre-trained language models with different sizes, based on our participation in the ClimateCheck@NSLP 2026 shared task on climate-related claim verification. This task aims to address the increasing amount of climate misinformation and disinformation on social media, emphasizing the importance of basing claims on reliable scientific evidence. Our study investigates the impact of different explanation strategies on entailment-based verification performance in scientific claim verification, while analyzing the trade-off between reasoning complexity and computational efficiency.
Transferring Scientific English Pre-Trained Language Models to Multiple Languages Using Cross-Lingual Transfer
Nikolas Ching-Pu Rauscher | Fabio Barth | Georg Rehm
Nikolas Ching-Pu Rauscher | Fabio Barth | Georg Rehm
In this paper, we present a pipeline for domain-adaptive pre-training and cross-lingual transfer of scientific language models from English to non-English languages. Starting from the multilingual scientific corpus SciLaD, we construct a cleaned English pre-training split and continually pre-train a T5-base encoder–decoder model, resulting in EN-T5-Sci. Our model achieves consistent zero-shot improvements on the Global-MMLU English benchmark, outperforming its base model, with particularly strong gains in STEM and Social Sciences. Despite its moderate size, it performs comparably to the much larger BLOOM model on scientific categories. Building on EN-T5-Sci, we transfer scientific knowledge to German, Japanese, Russian, Polish, Spanish, and Portuguese using the WECHSEL method. Our approach reinitializes language-specific embedding layers via aligned static embeddings while retaining the pre-trained Transformer weights, yielding six monolingual scientific T5 models. In zero-shot evaluation in each respective language, the transferred models generally outperform monolingual baselines. These results demonstrate that scientific domain knowledge acquired through English pre-training can be effectively transferred across languages, enabling competitive non-English scientific language models without training large multilingual systems from scratch.
Transformer Encoders with Heuristic-Guided Contrastive Learning for Software Coreference Resolution
Mahmoud Hassan | Dipendra Yadav
Mahmoud Hassan | Dipendra Yadav
This paper describes our system submitted to the Software Mention Detection and Coreference Resolution (SOMD) 2026 shared task, specifically for Subtask 1 (cross-document coreference resolution over gold-standard mentions) and Subtask 2 (cross-document coreference resolution over predicted mentions). The proposed approach employs a SciBERT architecture trained with Supervised Contrastive (SupCon) loss to generate dense mention representations, which are then clustered using Hierarchical Agglomerative Clustering (HAC) with average linkage. Software-aware heuristics are integrated to exploit domain-specific signals such as software name canonicalization and developer disambiguation to adjust pairwise similarity scores before clustering. The system achieved strong performance, with a CoNLL F1 score of 92.18% on coreference resolution over gold-standard mentions and 91.87% on coreference resolution over predicted mentions, showing significant performance of our approach in this area for human annotated and automated systems respectively
UniCite: A Dataset and Unified Hierarchical Taxonomy for Multi-Dimensional Citation Analysis
Amina Mourky | Elena Leitner | Julian Moreno-Schneider | Raia Abu Ahmad | Ekaterina Borisova | Georg Rehm
Amina Mourky | Elena Leitner | Julian Moreno-Schneider | Raia Abu Ahmad | Ekaterina Borisova | Georg Rehm
Research in Citation Context Analysis (CCA) has produced numerous taxonomic schemes that vary from three to 12+ categories, with different granularities and no mappings between frameworks, severely limiting systematic comparison and progress. Despite decades of study, CCA methods have largely relied on fragmented frameworks that treat citation tasks independently, ignoring systematic relationships between function classification, sentiment analysis, and importance assessment. To address these research gaps, we present three integrated contributions. First, we develop UniCite, a two-level taxonomy (six primary functions, 12 subcategories, two orthogonal dimensions) that systematically integrates three existing schemes. Second, we develop a comprehensive dataset of 4,017 citations combining established resources with 1,547 newly extracted citations from 2018-2024 publications, all manually annotated under our unified framework. Third, we demonstrate systematic task relationships through multi-task learning, achieving 21.1% relative improvement in subfunction classification over single-task approaches.
XplaiNLP @ ClimateCheck 2026 Task 2: Comparing Hierarchical Approaches for Fine-Grained Climate Disinformation Narrative Classification
Arthur Hilbert | Jing Yang | Vera Schmitt
Arthur Hilbert | Jing Yang | Vera Schmitt
We present our submission to Task 2 of the ClimateCheck 2026 shared task on Disinformation Narrative Classification which requires assigning climate-contrarian claims to fine-grained disinformation narratives. Using Qwen3-8B as a fixed backbone, we systematically compare data augmentation, prompt engineering and reinforcement learning techniques. Our experiments show that structured reasoning, particularly a chain-of-thought (CoT) prompting strategy aligned with the taxonomy’s hierarchical structure, substantially improves Macro-F1 over both zero-shot baselines and augmentation-based fine-tuning. Our best configuration achieves ∼0.625 Macro-F1, ranking first in Task 2. Our findings demonstrate that carefully designed hierarchical prompting can rival more complex training interventions in low-resource, highly imbalanced narrative classification settings.
up
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Hend Al-Khalifa | Mo El-Haj | Saad Ezzini
Hend Al-Khalifa | Mo El-Haj | Saad Ezzini
Hidden Sentiments: The Impact of Low-level Adversarial Perturbations on Arabic Sentiment Analysis Services
Abdelrahman Hamada Hefny Abdelkader
Abdelrahman Hamada Hefny Abdelkader
Sentiment analysis is one of the most popular applications of supervised machine learning for natural language processing. A common approach for obtaining a dataset to train sentiment analysis models is to extract user posts and comments from social media and other online platforms. However, this content is subject to various types of perturbations that go beyond the target of common preprocessing techniques and may impact the models’ performance. In this paper, a set of six popular corpora used in Arabic sentiment analysis research is analyzed to identify common patterns of character-level perturbations. The samples of three selected corpora were then used to test the performance of the online sentiment analysis services offered by three public cloud providers. This test is done using a clean version of each dataset and four other versions, each perturbed using a different technique. Empirical results indicate that no single sentiment analysis service is superior to others in all cases, and all three services are vulnerable to low-level adversarial attacks which may cause up to a 51% relative drop in macro average F1 score, while maintaining readability.
LLM-Based Financial Sentiment Analysis in Arabic: Evidence from Saudi Markets
Mona H. Albaqawi | Eman M. Albalkhi | Joud A. Albaiti | Enrico Lopedoto
Mona H. Albaqawi | Eman M. Albalkhi | Joud A. Albaiti | Enrico Lopedoto
Investor sentiment significantly influences financial markets, yet Arabic financial sentiment analysis remains limited by linguistic complexity and scarce domain-specific resources. This paper presents an LLM-based framework for large-scale Arabic financial sentiment analysis tailored to the Saudi market. We construct an 84K-sample Arabic Financial Sentiment Corpus integrating official financial news and social media data. The proposed pipeline includes preprocessing, deduplication, entity linking, conditional summarization, and five-class sentiment labeling using a multi-model consensus strategy to enhance reliability. We benchmark multiple large language models against traditional lexicon-based and fine-tuned transformer baselines. GPT-5 achieves the strongest class-balanced performance (Macro-F1 = 0.829), substantially outperforming conventional approaches. For summarization, Allam demonstrates the best trade-off between quality, hallucination control, and cost efficiency. Additional analyses examine cost–quality trade-offs and the impact of summarization on sentiment consistency. The results establish new benchmarks for Arabic financial sentiment classification and demonstrate the effectiveness of scalable LLM-based pipelines for domain-specific Arabic NLP.
Does Translation Preserve Sentiment? An Analysis of Arabic-English Cross-Lingual Classification
Nour Aldin Al Mubarak | Noura Al Moubayed
Nour Aldin Al Mubarak | Noura Al Moubayed
Machine translation is widely used in cross-lingual sentiment analysis, yet the assumption that translation preserves sentiment remains largely unexamined. We present a systematic analysis of translation-induced sentiment shifts across 11,558 samples from three Arabic-English datasets (AJGT, OCLAR, FSA) using three translation models (Helsinki-NMT, GPT-4o-mini, LLaMA-3.1-8B) and a fixed multilingual classifier (XLM-RoBERTa). A substantial proportion of samples experience sentiment shifts after translation, with accuracy drops ranging from less than 1% to nearly 20%. GPT-4o-mini achieves the strongest sentiment preservation, while LLaMA-3.1-8B exhibits both significant distortion and refusal behaviour. Critically, Helsinki-NMT’s successful translation of all samples indicates that LLaMA’s refusals stem from safety policies rather than input untranslatability. We also find that sentiment shift measurements are pipeline-dependent and vary with the classifier used for evaluation. These findings challenge the translate-then-classify paradigm and provide guidance for cross-lingual Arabic NLP systems.
When Bigger Isn’t Better: Evaluating LLMs for Arabic Sentiment Analysis
Mohamed Ibrahim | Abdullah Makki | Youssef Barakat | Nour Samy | Sarah AlHumoud
Mohamed Ibrahim | Abdullah Makki | Youssef Barakat | Nour Samy | Sarah AlHumoud
This study evaluates the performance of a fine-tuned Arabic sentiment transformer (CAMeL-MSA) against eight large language models (LLMs). Using zero-shot prompting across six Arabic sentiment datasets, we compare a specialized, task-specific approach against generalized model capabilities. Results show that the fine-tuned baseline substantially outperformed all LLMs on five of the six datasets in both accuracy and Macro F1-score. While LLMs offer versatility, this comparison highlights the continued practical superiority of task-specific fine-tuning over zero-shot prompting.
GATE-Reranker: A Strong Arabic Cross-Encoder for Document Reranking
Omer Nacar | Omar Elshehy | Mohamed Zaytoon | Khloud Al Jallad
Omer Nacar | Omar Elshehy | Mohamed Zaytoon | Khloud Al Jallad
Arabic information retrieval increasingly relies on multi-stage pipelines in which a fast first-stage retriever produces candidate passages and a neural reranker refines relevance. While transformer cross-encoders deliver strong effectiveness through joint query–passage encoding, multilingual rerankers achieve competitive performance on Arabic benchmarks. However, systematic analysis of calibration, robustness, and deployment behavior in Arabic-specific settings remains limited. We present GATE-Reranker, a compact Arabic cross-encoder initialized from an Arabic semantic embedding backbone and fine-tuned on large-scale mMARCO-style Arabic triplets. The model scores each query–passage pair via full self-attention and a lightweight regression head, enabling plug-and-play second-stage reranking for Arabic search and RAG systems. We evaluate on three Arabic benchmarks covering binary relevance discrimination, controlled multi-negative reranking, and large-scale mMARCO evaluation. While remaining competitive with strong multilingual rerankers in ranking effectiveness, GATE-Reranker demonstrates significantly improved calibration and discriminative behavior. These properties translate into more reliable downstream performance in retrieval and RAG pipelines, while maintaining low GPU memory and latency on a Tesla T4.
How Foundation Models Behave for Arabic Image Captioning?
Khaoula Dahimi | Amel Belabbaci | Hadda Cherroun | Abdelhamid Haouhat
Khaoula Dahimi | Amel Belabbaci | Hadda Cherroun | Abdelhamid Haouhat
Image captioning plays a crucial role in numerous applications, including educational systems. However, ensuring caption quality remains a significant challenge, particularly for morphologically rich, low-resource languages such as Arabic. We investigate an evaluation of Arabic image captioning using state-of-the-art multimodal foundation models. We systematically assess the performance of leading models—Gemini, Gemma, LLaMA, and Fanar. Our evaluation framework employs a diverse set of metrics spanning rule-based, learnable, visually-grounded, and LLM-based approaches to capture semantic accuracy, linguistic fluency, and hallucination detection. Experiments are conducted on two benchmark datasets: Flickr8k-Arabic and JEEM. Our findings reveal significant performance variations across models and evaluation metrics, highlighting the need for Arabic-specific optimization in multimodal architectures.
AlignAR: Generative Sentence Alignment for Arabic–English Parallel Corpora of Legal and Literary Texts
Baorong Huang | Ali Asiri
Baorong Huang | Ali Asiri
High-quality parallel corpora serve as the fundamental backbone for advancements in Machine Translation (MT) research and the development of effective translation pedagogy. Despite this need, robust resources for the Arabic-English language pair remain significantly scarce. Furthermore, existing datasets are often limited by their reliance on simplistic one-to-one sentence mappings, which fail to capture the structural complexities inherent in natural language translation. To address this deficiency, this paper presents AlignAR, a novel generative sentence alignment method, alongside a comprehensive new Arabic–English dataset that juxtaposes simple legal documents with complex literary texts. Our evaluation demonstrates that “Easy” datasets lack the discriminatory power to fully assess alignment methods. By reducing one-to-one mappings within our “Hard” subset, we exposed the limitations of traditional alignment techniques when faced with structural divergence. In contrast, Large Language Model (LLM) based approaches demonstrated superior robustness and adaptability. Specifically, the proposed LLM-based approaches demonstrated better robustness, achieving an overall F1-score of 85.5%, a nearly 9% improvement over previous methods. This study underscores the importance of complex benchmarks and validates the efficacy of generative models in handling the intricacies of bitext alignment. The codes and datasets are available on Github.
Helpful or Harmful? The Dual Role of Linguistic Features in LLM-Based Dialectal Machine Translation
Abdelhalim Hafedh Dahou | Mohamed Amine Cheragui
Abdelhalim Hafedh Dahou | Mohamed Amine Cheragui
Large Language Models (LLMs) have shown promising results in dialectal machine translation, yet the impact of explicit linguistic features remains underexplored. This paper examines whether part-of-speech (POS) tags and diacritization help or hinder LLM-based translation between Algerian dialect (Darija) and Modern Standard Arabic (MSA). Using a linguistically enriched subset of the PADIC dataset, we conduct bidirectional experiments across several frontier and open-weight LLMs, evaluated with automatic metrics and human judgments of adequacy and fluency. Results reveal a dual and asymmetric effect: diacritics can improve adequacy in the MSA → Algerian dialect direction, while POS tags and forced diacritization often introduce noise, especially for Algerian dialect → MSA translation. We further observe a mismatch between traditional overlap-based metrics and human evaluation, suggesting limitations in current evaluation practices. Overall, explicit linguistic augmentation does not consistently benefit LLM-based dialectal translation and must be applied cautiously.
ASCAT: Arabic Scientific Benchmark for Advanced Translation Evaluation
Serry Sibaee | Khloud Al Jallad | Zineb Yousfi | Israa Elhosiny | Yousra Yousra El-Ghawi | Batool Balah | Omer Nacar
Serry Sibaee | Khloud Al Jallad | Zineb Yousfi | Israa Elhosiny | Yousra Yousra El-Ghawi | Batool Balah | Omer Nacar
We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine machine translation and expert post-editing pipeline. Unlike existing Arabic-English corpora that rely on short sentences or single-domain text, ASCAT targets full scientific abstracts averaging 125.3 words (English) and 111.78 words (Arabic), drawn from five scientific domains: physics, mathematics, computer science, quantum mechanics, and artificial intelligence. Each abstract was translated using three complementary architectures generative AI (Gemini), transformer-based models (Hugging Face quickmt-en-ar), and commercial MT APIs (Google Translate, DeepL) and subsequently post-edited by domain experts at the lexical, syntactic, and semantic levels. The resulting corpus contains 67,293 English tokens and 60,026 Arabic tokens, with an Arabic vocabulary of 17,604 unique words reflecting the morphological richness of the language. We benchmark three state-of-the-art LLMs on the corpus GPT-4o-mini (BLEU: 37.07), Gemini-3.0-Flash-Preview (BLEU: 30.44), and Qwen3-235B-A22B (BLEU: 23.68) demonstrating its discriminative power as an evaluation benchmark. ASCAT addresses a critical gap in scientific MT resources for Arabic and is designed to support rigorous evaluation of scientific translation quality and training of domain-specific translation models.
We introduce SHEINfer, a novel task and dataset for inferring product categories from Arabic e-commerce reviews without explicit product mentions. Unlike traditional product classification that relies on product titles or descriptions, our task requires models to deduce product types solely from customer review text, which often contains implicit references through dialectal expressions, quality assessments, and contextual clues. We present a dataset of 801 Arabic reviews from the SHEIN e-commerce website, dual-annotated across 11 product categories with 515 agreed samples achieving moderate inter-annotator agreement (Cohen’s κ = 0.60). Given the relatively small dataset size, we employ 5-fold stratified cross-validation for all models to ensure robust performance estimates. Our experiments compare traditional machine learning approaches (TF-IDF with SVM and Logistic Regression), Arabic transformer models (AraBERT, CAMeLBERT, MARBERT), and large language models (GPT-4o-mini) in zero-shot and few-shot settings. Results show that MARBERT achieves the highest accuracy (0.586 ± 0.026), while TF-IDF with Logistic Regression achieves the best macro F1 (0.417 ± 0.056), indicating better performance across minority categories. GPT-4o-mini demonstrates poor zero-shot performance (0.064 accuracy) with modest improvement in 3-shot settings (0.186 accuracy), indicating that implicit product inference from dialectal Arabic text remains challenging for general-purpose LLMs. Our findings highlight the unique challenges of implicit product classification in Arabic e-commerce and establish benchmarks for future research in this underexplored area.
NAJD-MT: High-Fidelity Saudi Najdi–English Training Data for Bidirectional Neural Machine Translation
Nour Qandos | Samar Essa Ahmed | Omer Nacar | Ahmad Alrabghi | Rahaf Saeed Al Hallay | Aya Hamod | Shaden Alsuhaim
Nour Qandos | Samar Essa Ahmed | Omer Nacar | Ahmad Alrabghi | Rahaf Saeed Al Hallay | Aya Hamod | Shaden Alsuhaim
Dialectal Arabic remains significantly underrepresented in parallel resources for direct machine translation with English, particularly for regional varieties such as Saudi Najdi Arabic. In this work, we introduce NAJD-MT, a systematically constructed Saudi Najdi-English parallel corpus designed for training bidirectional neural machine translation models. Starting from the Saudi Arabic Dialectal Annotated (SADA) dataset, we generate English translations using GPT-4.1 and subsequently apply cross-lingual embedding-based cosine similarity filtering to improve semantic alignment and reduce translation noise. We analyze the impact of varying semantic similarity thresholds on corpus size and downstream translation performance. Using the constructed datasets, we train and evaluate multiple Transformer-based models, including NLLB-200, OPUS-MT, mBART, and AraT5v2, in both Najdi→English and English→Najdi directions. Experimental results demonstrate that stricter semantic filtering (cosine ≥ 0.7) consistently improves translation quality despite reducing dataset size, highlighting that data purity plays a critical role in dialectal machine translation training. Our findings provide a reproducible framework for constructing high-fidelity dialect English parallel corpora and emphasize the importance of semantic alignment filtering in low-resource dialectal settings.
Parsing Arabic Dialects Revisited: New Benchmarks, Models, and Insights
Ahmed Farouk Zakaria Elshabrawy | Go Inoue | Muhammed AbuOdeh | Nizar Habash
Ahmed Farouk Zakaria Elshabrawy | Go Inoue | Muhammed AbuOdeh | Nizar Habash
Parsing dialectal Arabic remains underexplored, with limited progress over the past two decades. Existing Modern Standard Arabic (MSA) parsers perform poorly on dialectal data, motivating the need for dialect-specific approaches. We revisit this task using modern neural models and present new results on Egyptian and Gulf Arabic dependency parsing. We demonstrate that even small amounts of dialectal training data yield substantial improvements in parsing accuracy. Our contributions include: (1) introducing a new annotated dataset for Gulf Arabic, (2) releasing a state-of-the-art multi-variety Arabic parser, and (3) employing dialect identification as a diagnostic tool to better understand how training data affects parsing performance across dialects and test sets.
On LLM Prompting Techniques for Arabic Language Arithmetic Reasoning
Reem Alenezi | Ayed Atallah Salman
Reem Alenezi | Ayed Atallah Salman
Math word problems (MWPs) require complex reasoning to extract mathematical relationships from textual descriptions. While Large Language Models (LLMs) have shown remarkable performance on English mathematical reasoning tasks, their effectiveness on Arabic MWPs remains largely unexplored. This paper introduces three Arabic datasets (AGSM8K, Qudurat, and ArabicMWPs) and evaluates six LLMs using three prompting techniques: Manual Chain-of-Thought (CoT), Zero-shot CoT, and Self-consistency. Performance is assessed using accuracy and BERTScore metrics (precision, recall, F1-score). Our findings demonstrate that GPT-4o with Self-consistency achieves the highest accuracy of 97.65% on AGSM8K. It also obtains a precision of 71.94%, a recall of 71.31%, and an F1-score of 71.50%. The Arabic-specific LLM ALLaM achieves 84.41% accuracy on ArabicMWPs and 43.97% on AGSM8K. Fine-tuning experiments are further conducted on models using Arabic mathematical data. This work addresses the critical gap in Arabic mathematical reasoning resources and provides insights for developing Arabic-capable AI systems. Prompt-engineering methods combined with LLMs are regarded as a strong approach for advancing education and scientific research in solving Arabic mathematical problems.
DIA2 - a Comprehensive and Diverse Diacritized Arabic Corpus for NLP Research
Fatima Dekmak | Shady Elbassuoni | Khaled Shaban | Hazem Hajj | Wassim El-Hajj | Yasmine Abu Adla | Buthaina Alabrash
Fatima Dekmak | Shady Elbassuoni | Khaled Shaban | Hazem Hajj | Wassim El-Hajj | Yasmine Abu Adla | Buthaina Alabrash
The development of Arabic natural language processing (NLP) applications and large language models (LLMs) faces substantial challenges, primarily due to the scarcity of high-quality native Arabic datasets. To address this critical gap, we present DIA2 (a Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus), a novel dataset curated from 28 diverse, carefully selected Arabic sources. DIA2 emphasizes the use of original Arabic text and explicitly avoids machine-translated content. The corpus incorporates substantial amounts of text from books, news articles, and poetry, and employs extensive data preprocessing to support NLP research and LLM development. Our preprocessing pipeline includes rigorous text cleaning, URL- and document-level deduplication, and automatic diacritization, while preserving a gold diacritized subset derived from manually annotated sources. The resulting corpus comprises over 140 GB of high-quality text, containing more than 26 million unique words and 41.9 billion tokens. To evaluate the proposed pipeline, we conducted controlled continued pretraining experiments using Llama3.1-8B on both raw and processed subsets of DIA2. The model trained on processed data consistently outperformed its counterpart across multiple Arabic evaluation benchmarks. These results highlight the positive impact of systematic preprocessing and the utility of DIA2 in empowering native Arabic LLMs and downstream NLP tasks.
CV-18 NER: Augmented Common Voice for Named Entity Recognition from Arabic Speech
Youssef Saidi | Haroun Elleuch | Fethi Bougares
Youssef Saidi | Haroun Elleuch | Fethi Bougares
End-to-end speech Named Entity Recognition (NER) aims to directly extract entities from speech. Prior work has shown that end-to-end (E2E) approaches can outperform cascaded pipelines for English, French, and Chinese, but Arabic remains under-explored due to its morphological complexity, the absence of short vowels, and limited annotated resources. We introduce CV-18 NER, the first publicly available dataset for NER from Arabic speech, created by augmenting the Arabic Common Voice 18 corpus with manual NER annotations following the fine-grained Wojood schema (21 entity types). We benchmark both pipeline systems (ASR + text NER) and E2E models based on Whisper and AraBEST-RQ. E2E systems substantially outperform the best pipeline configuration on the test set, reaching 37.0% CoER (AraBEST-RQ 300M) and 38.0% CVER (Whisper-medium). Further analysis shows that Arabic-specific self-supervised pretraining yields strong ASR performance, while multilingual weak supervision transfers more effectively to joint speech-to-entity learning, and that larger models may be harder to adapt in this low-resource setting. Our dataset and models are publicly released, providing the first open benchmark for end-to-end named entity recognition from Arabic speech. https://huggingface.co/datasets/Elyadata/CV18-NER
ARHAHA 2026: The Shared Task on Arabic Humor Automatic Generation
Ameera Masoud Almasoud | Hend Al-Khalifa | Reem Fahad Alqifari | Nourah Alangari | Manal M. Albahlal
Ameera Masoud Almasoud | Hend Al-Khalifa | Reem Fahad Alqifari | Nourah Alangari | Manal M. Albahlal
Humor generation remains one of the most challenging tasks in natural language processing, particularly in Arabic, where cultural context, dialectal variation, and linguistic nuances are central to comedic effect. In this paper, we present the ARHAHA 2026 shared task on constrained Arabic humor generation. The task requires systems to generate jokes that incorporate a given pair of words while adhering to safety and cultural constraints. We describe the task design, dataset construction, and evaluation framework, which combines automatic validation with human evaluation. Nine teams registered for the shared task; among them, three submitted final system outputs and two provided system description papers. Each participating system generated 1,200 Arabic jokes. For each system, a subset of 300 jokes was selected for evaluation by three independent annotators. The evaluation considered humor quality, originality, lexical constraint compliance, and safety. The results show that participating systems can produce safe and original content. However, generating genuinely humorous outputs remains difficult. The top-performing system was judged humorous in only 5.01% of outputs, highlighting the inherent difficulty of computational humor generation. All three systems maintained very low rates of policy violations and stereotyping, demonstrating the effectiveness of constrained generation for safe content production. However, the very low humor rates indicate a substantial gap between generating fluent, constraint-compliant text and producing genuinely funny content. The top-performing system achieves stronger performance across originality, lexical compliance, and safety, resulting in a final score of 49.25, compared to 44.62 for the second-ranked system and 35.99 for the third-ranked system. These results reveal that humor generation, rather than safety or constraint adherence, is the dominant bottleneck in constrained Arabic humor generation.
The AdabEval 2026 Shared Task on Arabic Politeness Detection
Reem Fahad Alqifari | Hend Al-Khalifa | Nadia Ghezaiel | Maria Bounnit | Hend Hamed Alhazmi | Ameera Masoud Almasoud | Sharefah Ahmed Al-Ghamdi | Noof Abdullah Alfear
Reem Fahad Alqifari | Hend Al-Khalifa | Nadia Ghezaiel | Maria Bounnit | Hend Hamed Alhazmi | Ameera Masoud Almasoud | Sharefah Ahmed Al-Ghamdi | Noof Abdullah Alfear
We present an overview of the AdabEval 2026 shared task, organized as part of the OSACT7 workshop (co-located with LREC 2026). This task introduces the first benchmark suite for politeness detection. It includes two subtasks: Politeness Classification (Subtask A) and Category Prediction (Subtask B). The task focuses on evaluating models’ ability to recognize and categorize politeness phenomena in Arabic text. Evaluation was conducted using an automatic metric (macro F1-score). A total of 28 unique teams participated in the shared task. Of these, 13 teams submitted final system predictions across the two subtasks. The top-performing systems relied primarily on transformer-based architectures. The winning systems achieved macro F1-scores of 0.89 for Subtask A and 0.58 for Subtask B.
GHAD NLP at AdabEval2026: Transformer-Based Approach for Arabic Politeness and Pragmatic Category Classification
Ghada Alfattni | Ghader Kurdi
Ghada Alfattni | Ghader Kurdi
This paper presents our submission to the AdabEval 2026 shared task on Arabic politeness classification and pragmatic category prediction. We explored a range of Arabic-specific and multilingual transformer models and integrated their outputs through an ensemble strategy. Our approach achieved state-of-the-art performance in the shared task, ranking first in both subtasks with a macro-F1 score of 0.89 and an accuracy of 0.93 on subtask A, and a macro-F1 score of 0.58 on subtask B. Although our approach delivered high performance on overall politeness classification, pragmatic category prediction remains more challenging. Despite achieving the top ranking in this subtask, the comparatively lower macro-F1 score suggests that modelling fine-grained pragmatic functions requires further methodological refinement and experimentation.
MOSKA-NLP at AdabEval 2026: Feature-Enriched Ensembling for Arabic Politeness Detection
Nina A. Andriyanova-Almaamary
Nina A. Andriyanova-Almaamary
In this paper, we present our system for subtask A of the AdabEval 2026 shared task, which focuses on classifying Arabic text into Polite, Neutral, and Impolite categories. Politeness detection is challenging because it cannot be inferred from lexical meaning alone. This is prominent in Arabic language, where politeness is often conveyed through formulaic expressions, stylistic cues, and dialectal variations. Our approach follows a three-stage strategy. First, we evaluate five Arabic sentence embedding models based on different pretrained encoders to identify a strong representation backbone. Second, we enrich sentence embeddings with explicit lexical, surface-level, and auxiliary signals derived from external models, including dialect, intent, and sarcasm classifiers. Third, we combine predictions from independently trained models, using weighted probability-level ensembling with class-specific decision thresholds to address class imbalance. Experimental results show that feature-enriched representations consistently outperform embedding-only baselines, with additional gains obtained from calibrated ensembling. The proposed system achieves a macro-F1 score of 0.87 and an accuracy of 93% on the official AdabEval 2026 evaluation for subtask A.
SHU at AdabEval 2026: Category-Aware Fine-Tuning of MARBERT for Arabic Politeness and Pragmatic Function Classification
Alla Zwawi | Stephen Wu
Alla Zwawi | Stephen Wu
This paper describes our submission to the AdabEval 2026 shared task, addressing Subtask A (politeness classification) and Subtask B (multi-label pragmatic category prediction). For Subtask A, we fine-tuned MARBERT using weighted cross-entropy to mitigate class imbalance. For Subtask B, we apply BCEWithLogitsloss with inverse-frequency positive weighting to address the minority categories, and we introduce a category merging strategy to reduce categories’ sparsity and annotation variation. Finally, we propose a stacked architecture where predicted pragmatic categories are injected as auxiliary features into the politeness classifier. Our results demonstrate that dialect-aware modelling, class-imbalance handling, and category-aware stacking improve Macro-F1 across both subtasks, achieving 0.85 for Subtask A and 0.55 for Subtask B on the test set.
Comparative Study of Machine Learning and Transformer-Based Approaches for Arabic Politeness Detection at AdabEval 2026
Mariem Ben Arbia | Ghada Ben Amor | Omar Trigui
Mariem Ben Arbia | Ghada Ben Amor | Omar Trigui
This paper describes our system submitted to the OSACT7 AdabEval shared task on Arabic politeness detection (TaskA). The task requires classifying Arabic texts into three categories: Polite, Impolite, and Neutral. We systematically explore multiple approaches, progressing from classical machine learning baselines using pre-trained embeddings to fine-tuned transformer models. Our best system leverages MARBERT, a transformer model pre-trained on one billion Arabic tweets, fine-tuned with Focal Loss to handle the significant class imbalance present in the dataset (70% Neutral). We additionally experiment with hybrid approaches combining fine-tuned embeddings with gradient-boosted classifiers and ensemble methods. Our best single model achieves a macro F1 score of 0.84 and an accuracy of 0.90 on the validation set, substantially outperforming classical ML baselines (F1 = 0.42).
AdabEval 2026 Task B : Multi-Label Classification of Arabic Politeness Criteria in Social Media Media
Rand Abdullah Alturki
Rand Abdullah Alturki
We address the problem of multi-label classification of politeness and impoliteness criteria in Arabic social media posts, as defined in Subtask B of an Arabic politeness shared task (CITATION) The goal is to assign up to four labels from nine pragmatic categories, including Insult, Criticism, Respect, Prayers, and Hospitality, to each post. We first construct consistent multi-label annotations by mapping heterogeneous criterion strings into the official label set and analyzing their skewed distribution. To mitigate severe class imbalance, especially for rare categories such as Hospitality and Racism/Discrimination, we apply targeted oversampling of minority instances. Our modelling pipeline combines a TF–IDF + Logistic Regression baseline with two transformer-based encoders, MARBERT and AraBERT-twitter, trained for multi-label classification with Focal Loss. We then aggregate model outputs through a weighted ensemble and optimize per-class decision thresholds on a held-out validation set to improve macro-averaged F1. Experiments on the shared-task train/validation split show that the ensemble substantially outperforms the TF–IDF baseline and individual transformers, particularly on underrepresented categories, while maintaining competitive performance on frequent labels.
QIAS 2026: Overview of the Shared Task on Islamic Inheritance Reasoning
Abdessalam Bouchekif | Somaya Eltanbouly | Shahd Gaben | Mohammed Ghaly | Samer Rashwani | MOHAMED Emad | Heba Sbahi
Abdessalam Bouchekif | Somaya Eltanbouly | Shahd Gaben | Mohammed Ghaly | Samer Rashwani | MOHAMED Emad | Heba Sbahi
This paper presents a comprehensive overview of the QIAS 2026 shared task, organized as part of the OSACT7 Workshop and co-located with LREC 2026. The shared task was designed to evaluate the ability of large language models to perform complex reasoning in the religious and legal domain of Islamic inheritance. Unlike conventional question-answering benchmarks, QIAS 2026 focuses on end-to-end reasoning from natural language cases, requiring systems to perform the full inheritance calculation process, from identifying the eligible heirs to assigning the correct share to each beneficiary. To support this evaluation, the task was based on the MAWARITH benchmark, a dataset of 12,500 Arabic inheritance cases annotated with intermediate reasoning steps and final answers. System submissions were evaluated using MIR-E, a multi-step metric that measures performance across the main stages of inheritance reasoning. A total of 16 teams participated in the shared task, investigating a range of approaches, including prompting-based methods, retrieval-augmented generation, and fine-tuning strategies. The results show that Islamic inheritance remains a highly challenging benchmark for current language models, especially in stages that require precise legal interpretation and structured numerical reasoning. This overview summarizes the task design, dataset, evaluation framework, participating systems, and main results.
QU-NLP at QIAS 2026: Multi-Stage QLoRA Fine-Tuning for Arabic Islamic Inheritance Reasoning
Mohammad ALSmadi
Mohammad ALSmadi
Islamic inheritance law (علم المواريث, ilm al-mawarıth) presents a challenging domain for evaluating large language models’ structured reasoning capabilities, requiring multi-step legal analysis, rule-based blocking decisions, and precise fractional calculations. We present QU-NLP’s submission to the QIAS 2026 shared task on Arabic Islamic inheritance reasoning. Our approach employs a multi-stage Quantized Low-Rank Adaptation (QLoRA) fine-tuning strategy on Qwen3-4B: (1) domain adaptation on 3,166 Islamic fatwa records to acquire inheritance terminology and jurisprudential reasoning patterns, followed by (2) task-specific training on 12,000 structured inheritance cases to optimize JSON-formatted output generation. Using 4-bit NF4 quantization with rank-128 LoRA adapters, our model achieves 90% MIR-E (Mawarith Inheritance Reasoning Evaluation) score on the test set, demonstrating competitive performance while requiring minimal computational resources. Our results show that domain-specific pre-adaptation combined with structured output training enables small language models to perform complex legal reasoning tasks effectively comparing to commercial systems such as Gemini-2.5-flash.
PSL at QIAS 2026: Which Models Perform Better in Arabic Inheritance Reasoning?
amine Mohammed | Chahinez Bouchekif
amine Mohammed | Chahinez Bouchekif
This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning. The task evaluates the ability of large language models to solve inheritance cases that require legal interpretation, multi-step reasoning, and precise numerical computation. We compare commercial and open-source models under a unified prompting strategy to assess their effectiveness in structured legal reasoning with minimal task-specific adaptation. Our results show a clear gap in reliability between the two model families. Commercial models demonstrate stronger performance in identifying eligible heirs, applying exclusion rules, and maintaining consistency across reasoning steps. In contrast, open-source models exhibit greater instability, particularly in cases involving dependent legal decisions and fractional share adjustments. The best performance is achieved by Gemini 2.5 Flash, with an MRE of 0.989.
AGS-KSU at QIAS 2026: A Comparative Study of Prompting and LLM Approaches for Structured Islamic Inheritance Reasoning
Hicham Ghazi Sidaoui
Hicham Ghazi Sidaoui
This paper describes our submission to the QIAS 2026 shared task on structured Islamic inheritance reasoning, based on the MAWARITH benchmark (Bouchekif et al., 2026). The task requires multi-step structured prediction for Arabic inheritance cases, including heir identification, blocking, share assignment, adjustment detection, and final distribution, evaluated with the MIR-E metric. We compare four system configurations: a QLoRA fine-tuned Qwen2.5-3B baseline, a multi-stage Fanar-Sadiq pipeline with deterministic validation and post-processing, and two GPT-5.4 prompting setups. On the official test set, the best result was achieved by GPT-5.4 with explicit inheritance rules and development examples used as in-context demonstrations, reaching a MIR-E score of 0.84, compared with 0.76 for a minimal-prompt GPT-5.4 variant. These results suggest that explicit rule conditioning and in-context demonstrations can improve performance in this setup. Since the compared systems vary in model family and prompting strategy, the findings should be interpreted as a comparison of task configurations rather than a controlled model-only comparison.
Silah at QIAS 2026: Fine-Tuning vs. Retrieval-Augmented Generation for Islamic Inheritance Reasoning
Ghader Kurdi | Hanan Justanieah | Hala Justanieah
Ghader Kurdi | Hanan Justanieah | Hala Justanieah
Islamic inheritance is a highly structured and rule-intensive domain that requires precise reasoning. The QIAS 2026 Shared Task introduces a benchmark for evaluating generative artificial intelligence on end-to-end inheritance problem solving. In this paper, we present our team Silah’s participation in the QIAS 2026 shared task, where we compare three approaches: (1) a multi-stage retrieval-augmented, rule-guided pipeline, (2) supervised fine-tuning of generative large language models, and (3) a retrieval-augmented fine-tuning approach. We evaluate several open-source models, including Qwen2.5, Llama, DeepSeek, and Fanar. Our results show that supervised fine-tuning consistently outperforms retrieval-based approaches, with the fine-tuned Fanar-1-9B-Instruct model achieving the best performance (MIR-E = 0.83) and ranking sixth overall in the shared task. These findings suggest that learning implicit reasoning patterns through fine-tuning is more effective than explicit rule injection under current retrieval setups, and emphasize the need for more accurate and minimal rule selection mechanisms in future retrieval-augmented approaches.
KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization
Asma Ali Al Wazrah | Waad Alshammari | Rawan Almatham | Raghad Al-rasheed | Afrah Abdulaziz Altamimi | Rufael Marew | Sawsan Alqahtani | Hanan Aldarmaki | Abdullah I. Alharbi | Abdulrahman Saeed Alshehri | Mohamed Assar | Amal Almazrua | Abdulrahman Alosaimy
Asma Ali Al Wazrah | Waad Alshammari | Rawan Almatham | Raghad Al-rasheed | Afrah Abdulaziz Altamimi | Rufael Marew | Sawsan Alqahtani | Hanan Aldarmaki | Abdullah I. Alharbi | Abdulrahman Saeed Alshehri | Mohamed Assar | Amal Almazrua | Abdulrahman Alosaimy
This paper presents the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization, addressing a persistent challenge in Arabic NLP. The task focuses on transforming speech transcripts into fully diacritized Arabic text by leveraging both the speech signal and its undiacritized transcript. Unlike conventional ASR tasks that focus on transcription, this task integrates acoustic and textual information to improve diacritization accuracy. The shared task consists of two subtasks: (1) Data Contribution, where participants recorded and reviewed speech data through the VoiceWall platform, resulting in 2,160 recordings, and (2) Diacritization, where 5 teams developed systems that generate fully diacritized text from speech and undiacritized transcripts. The dataset includes approximately 5 hours of Modern Standard Arabic (MSA) and multi-dialectal speech with fully diacritized references. Experimental results show that several participant systems outperform the provided baselines, and that incorporating speech information and fine-tuning improves performance compared to text-only approaches. KSAA-2026 shared task establishes a benchmark for multimodal Arabic diacritization and supports the development of robust systems for applications in education, accessibility, and speech-driven text generation.
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
Meshal Abdullah Alamr | Hassan Rshed Alqaeri | Abdullah Aldahlawi
Meshal Abdullah Alamr | Hassan Rshed Alqaeri | Abdullah Aldahlawi
We describe the winning system for Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization. The task requires producing fully diacritized Arabic text from speech audio and undiacritized transcripts, with only 2,327 training samples available and no external data permitted. Our system fine-tunes CATT-Whisper, a character-level multimodal model combining a pretrained CATT text encoder with a frozen Whisper speech encoder. The key to our approach is training regularization: R-Drop consistency regularization, Optuna-optimized hyperparameters with high weight decay, and Focal Loss. At inference, we average 200 stochastic forward passes across four model checkpoints using Monte Carlo Dropout at the softmax probability level. The system achieves 23.26% WER on the primary leaderboard metric (with case endings, including no-diacritic positions), placing 1st among all participants.
TantaArabNLP at KSAA-2026 Task 2: Adapting CATT-Whisper for Arabic Speech Dictation with Automatic Diacritization
Nada Adel Esmaeil | Reda M. Elbasiony | Mohamed T. Faheem
Nada Adel Esmaeil | Reda M. Elbasiony | Mohamed T. Faheem
We present our submission to the KSAA-2026 Shared Task (Subtask 2): Automatic Diacritization of Speech Dictation. Building upon the CATT-Whisper multimodal architecture, which fuses representations from a pre-trained CATT text encoder and the Whisper speech encoder, we fine-tune the model end-to-end on the official shared task training data. To further enhance performance on speech-dictated Arabic text, we apply careful post-processing to the model outputs. Our best submission achieves a Diacritic Error Rate (DER) of 7.04, a Word Error Rate (WER) of 24.39, and a Sentence Error Rate (SER) of 71.65 on the hidden test set, securing 2nd place in the competition. These results demonstrate the effectiveness of adapting a strong multimodal baseline to the speech-aware diacritization setting and highlight the value of task-specific fine-tuning and output refinement for bridging the gap between spoken transcripts and fully diacritized Arabic text.
Fine-Tashkeel at KSAA-2026: A Comprehensive Evaluation of Seq2Seq and Multimodal Approaches for Automatic Diacritization of Arabic Speech Dictation
Hassan Barmandah | Fatimah Emad Eldin | Omer Nacar | Wareef Alzubaidi
Hassan Barmandah | Fatimah Emad Eldin | Omer Nacar | Wareef Alzubaidi
This paper presents the Fine-Tashkeel system for Task 2 of the KSAA-2026 Shared Task on Automatic Diacritization of Speech Dictation. Diacritization of speech-derived Arabic text poses challenges due to dialectal variation, morphological ambiguity, and the absence of acoustic cues in text-only pipelines. Our approach treats diacritization as a character-level sequence-to-sequence task, mapping undiacritized text directly to its fully diacritized form. We evaluate 18 models spanning text-only, ASR-augmented, and fine-tuned configurations, finding that text-only Seq2Seq approaches outperform off-the-shelf multimodal models—a gap we attribute to task mismatch in generic ASR systems rather than an inherent audio limitation. Our best submission, using zero-shot inference without task-specific training, achieved a Diacritic Error Rate (DER) of 10.56%, Word Error Rate (WER) of 34.47%, and Sentence Error Rate (SER) of 79.88%, ranking 5th out of 7 teams. Per-nationality error analysis reveals significant dialectal variation (Egyptian 3.70% vs. Algerian 13.73% DER), and diagnostic analysis confirms that case endings and vowel ambiguity are the primary bottlenecks. Code and evaluation scripts are publicly available.
Eraserhead at OSACT7 Shared Task: ASR Consistency Filtering and Speaker-Adaptive Post-Processing for Arabic Speech Diacritization
Muhammad Abu Horaira | Nahian Chowdhury
Muhammad Abu Horaira | Nahian Chowdhury
Arabic speech diacritization is the task of restoring short vowel marks to undiacritized text derived from speech input. It remains difficult because ASR output can be noisy, dialectal variation is substantial, and speakers often differ in how they realize word-final diacritics. In this paper, we describe our submission to Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization, where our system ranked 4th on the official leaderboard. Our approach builds on a pretrained ASR-aware diacritization model and adds three components: ASR Consistency Filtering, confidence-based ensembling of three checkpoints, and speaker-adaptive post-processing specifically for word-final diacritics. Rather than discarding problematic data, our filtering strategy replaces unreliable ASR transcripts with the undiacritized gold text rather than removing training examples, which makes training more stable. On the official test set, our system achieved a Diacritic Error Rate (DER) of 8.23, a Word Error Rate (WER) of 30.37, and a Sentence Error Rate (SER) of 80.79 under the With Case Endings (WCE), Including No Diacritic (Incl. 0) evaluation setting. It also outperformed the organizers’ fine-tuned Text+ASR baseline in three of the four main evaluation settings.
Abjad AI at KSAA-2026 Shared Task 2: Grouped Speech Conditioning for Arabic Diacritization
Naif Saad Alharthi | Ahmad Ghannam | Faris Alasmary | Kholood Al Tabash | Shouq Sadah | Lahouari Ghouti
Naif Saad Alharthi | Ahmad Ghannam | Faris Alasmary | Kholood Al Tabash | Shouq Sadah | Lahouari Ghouti
We describe Abjad AI’s submission to KSAA-2026 Shared Task 2 on automatic diacritization of Arabic speech dictation. The task requires generating fully diacritized text given speech audio and an undiacritized transcript. Because text-only diacritization cannot resolve ambiguities that are recoverable from the acoustic signal, we propose conditioning a character-level encoder-only Transformer (CATT) (Alasmary et al., 2024) on speech representations. We introduce grouped speech conditioning, which downsamples speech encoder features into a small set of pooled tokens concatenated to the text input, enabling efficient fusion without architectural changes to CATT. We train with a two-phase schedule that first freezes the text encoder, then fine-tunes the full model. Our best system, using Whisper-small (Rad-ford et al., 2022) features with five grouped tokens, achieves a Diacritization Error Rate (DER) of 6.60 and a Word Error Rate (WER) of 18.66 (without case endings, including no-diacritic) on the official test set. Notably, we find that Whisper-small consistently outperforms Whisper-large-v3, suggesting that compact speech representations better suit this fusion setting. This is an extended and revised version of our previous work (Ghannam et al., 2025).
AraSentEval 2026: A Shared Task on Sentiment Analysis and Swapping in Arabic
Saad Ezzini | Shadi Abudalfa | Maram I. Alharbi | Salmane Chafik | Hamzah Luqman | Mo El-Haj | Paul Rayson | Reem Alotaibi
Saad Ezzini | Shadi Abudalfa | Maram I. Alharbi | Salmane Chafik | Hamzah Luqman | Mo El-Haj | Paul Rayson | Reem Alotaibi
Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and Swapping in Arabic (AraSentEval), organized as part of the OSACT7 Workshop at LREC 2026. This shared task consists of two subtasks: Subtask 1 focuses on multi-class and multi-dialect sentiment analysis, requiring models to identify sentiment polarity across various Arabic dialects. Subtask 2 introduces a generative task for Arabic sentiment swap, challenging models to invert sentiment polarity while preserving core semantics. In this overview paper, we present the motivation, dataset creation, and summarize the main findings from participating models.
TTLab at AraSentEval: SARF( صرف) Sentiment Analysis via Root-based Fusion for Multi-Dialectal Arabic
Ali Abusaleh | Bhuvanesh Verma | Alexander Mehler
Ali Abusaleh | Bhuvanesh Verma | Alexander Mehler
Arabic sentiment analysis is challenged by morphological complexity and lexical variation across Arabic dialects, compounded by subjectivity in how speakers and writers express sentiment. In this paper, we present our submission for the AraSentEval 2026 Shared Task on Arabic Dialect Sentiment Analysis. We propose SARF (صرف) a multi-view architectural framework that integrates surface-level context with stemmed and rooted morphological perspectives using a shared MARBERTv2 encoder. Our system employs a hybrid BERT-CNN-BiLSTM-Attention architecture to capture both local sentiment n-grams and global sequential dependencies. Experimental results show that while individual morphological normalization strategies (stemming or rooting) may degrade performance, their joint integration via cross-morphological attention provides robust features across diverse dialects. Our final system achieved a competitive macro-F1-score of 0.9263, ranking 2nd out of 15 participating teams.
A Comparative Study of Arabic Sentiment Swap Models for AraSentEval 2026
Yumna Hamdy | Mohab ElDamhougy | Yomna Eid | Ensaf Hussein
Yumna Hamdy | Mohab ElDamhougy | Yomna Eid | Ensaf Hussein
Sentiment swap is a controlled text generation task that rewrites a sentence by inverting its sentiment polarity while preserving semantic content and fluency. In this paper, we present our system for AraSentEval 2026 Subtask 2 on Arabic sentiment swap, a particularly challenging problem due to Arabic’s rich morphology and dialectal variation. We investigate multiple modeling paradigms, including encoder–decoder and multilingual approaches, and propose an enhanced system that combines targeted data augmentation and ensemble learning. Specifically, we augment underrepresented dialectal patterns to improve robustness and ensemble two Arabic-focused sequence-to-sequence models, AraBART and AraT5v2. Experiments are conducted on the MA’aks parallel dataset under fine-tuned settings. Our system ranked first in AraSentEval 2026 Subtask 2, achieving a BLEU score of 43.0, chrF of 65.36, and sentiment preservation accuracy of 0.7554. The results demonstrate that dialect-aware augmentation together with model ensembling substantially improves sentiment-controlled generation in Arabic and establishes strong baselines for future research in low-resource sentiment manipulation. Keywords: Arabic NLP, sentiment swap, style transfer, AraSentEval, text generation
L3IA at AraSentEval 2026 Subtask 2: LLM-Based Multi-Step Pipeline for Arabic Sentiment Swap
Abdessamad Benlahbib | Hamza Alami | Mohamed M’haouach | Kaouthar Elyoussoufi
Abdessamad Benlahbib | Hamza Alami | Mohamed M’haouach | Kaouthar Elyoussoufi
This paper describes our system submitted to the AraSentEval 2026 Shared Task, Subtask 2: Arabic Sentiment Swap. The task requires rewriting Arabic sentences to invert their sentiment polarity while preserving the core meaning. We propose a multi-step pipeline approach that uses large language models (LLMs). Our method decomposes the sentiment inversion problem into three stages: (1) sentiment expression extraction, where the model identifies all sentiment-bearing words and phrases in the input sentence; (2) opposite expression generation, where each identified expression is replaced by its semantic opposite; and (3) sentence reconstruction, where the final output is assembled to ensure grammatical correctness and natural fluency. Our system achieves 74.3% sentiment style accuracy, 27.22 BLEU, and 55.04 chrF on the official test set.
CasbAI at AraSentEval 2026: Robust Dialectal Arabic Sentiment Classification via Multi-Seed Ensembling and Data Augmentation.
Chaima Abdelaziz | KahinaHouda Saadaoui | Faiza BELBACHIR | Lynda Said Lhadj
Chaima Abdelaziz | KahinaHouda Saadaoui | Faiza BELBACHIR | Lynda Said Lhadj
This paper describes the system we designed for our participation in the AraSentEval 2026 shared task on Arabic dialectal sentiment analysis. We propose a transformer-based approach relying on MARBERT combined with a multi-seed ensemble strategy and several optimization techniques. Our system integrates seven independently trained models with different random initializations and applies Stochastic Weight Averaging (SWA) to improve generalization. To address class imbalance, we augment the training data through dialectal synonym replacement, increasing the dataset size by 13.9% while preserving dialect distribution. In addition, we incorporate Test-Time Augmentation (TTA) and investigate the use of pseudo-labeling based on high-confidence predictions. We report our experiments on the official dataset covering Moroccan, Egyptian, Jordanian, and Saudi dialects, and analyze the contribution of each component through ablation experiments. Our system achieved a macro F1-score of 84.62% on the test set, ranking 3rd among 15 participating teams.
BDSI at AraSentEval Shared Task : A Multi-Transformer Contrastive Learning for Arabic Dialect Sentiment Analysis
Mohamed M’haouach | Kaouthar Elyoussoufi | Abdessamad Benlahbib | Hamza Alami
Mohamed M’haouach | Kaouthar Elyoussoufi | Abdessamad Benlahbib | Hamza Alami
This paper presents our system for the AraSentEval 2026 shared task on Arabic dialect sentiment analysis. We propose a multi-model ensemble combining AraBERTv2 and CAMeLBERT with supervised contrastive learning to improve sentiment classification. The system incorporates dialect-aware preprocessing, class-weighted cross-entropy loss with label smoothing, supervised contrastive loss for enhanced sentence representations, and rule-based post-processing for dialect-specific patterns. Our approach achieves a macro F1-score of 0.83 on the official test set, demonstrating the effectiveness of contrastive learning with pretrained Arabic language models for dialectal sentiment analysis.
Codezone Research Group at AraSentEval Shared Task: Arabic Sentiment Swap beyond Negation Prepending, Benchmarking Multilingual T5 against Large Language Models on the MA’AKS Corpus
Abdulkadir Shehu Bichi | Sarah Yassine
Abdulkadir Shehu Bichi | Sarah Yassine
Abstract We launched ASBN-MT5, the system for Arabic Sentiment Swap, which performs the task of inverting the sentiment of a sentence while keeping the meaning intact. This is a sequence-to-sequence task. We demonstrate ASBN-MT5: mT5, which is a MultiLingual T5 model, fine-tuned on the provided dataset of the AraSentEval 2026 Shared Task. We describe the data as the first of its kind for the Arabic language, as MAAKS is the first manually composed, parallel, cross-linguistic corpus for the Arabic language. With the preliminary results of Sentiment Flip for the task of Sentiment Inversion, we have recorded a rate of 59.5% for positive to negative conversions and 58.5% for negative to positive conversions, while maintaining an average similarity to the original sentences of 0.955. We present the Arabic prompts and a neuro-developmental (Deep Learning) recipe. Due to the evaluation criteria which include Exact Match, Flip Success, Surface Similarity, and Quality of Output, we restrict the use of Prepended Negation as the main technique and recommend the use of LLMs designed for the Arabic language in the near future. Keywords: mT5, sequence-to-sequence, AraSentEval 2026, Arabic NLP, Text Style Transfer, Sentiment Swap
L3IA-Subtask 1 at AraSentEval Shared Task: Multi-Dialect Arabic Sentiment Classification via a Transformer-Based Approach
Mohamed M’haouach | Kaouthar Elyoussoufi | Hamza Alami | Abdessamad Benlahbib
Mohamed M’haouach | Kaouthar Elyoussoufi | Hamza Alami | Abdessamad Benlahbib
This paper presents our system and findings for AraSentEval 2026 Subtask 1 on Arabic Dialect Sentiment Analysis. We propose an automated sentiment classification system grounded in advanced Natural Language Processing (NLP) techniques. The proposed approach leverages pre-trained Transformer-based architectures to categorize textual inputs into three sentiment polarities: positive, negative, and neutral. Initially, a text normalization procedure is applied to unify the orthographic and graphical variations characteristic of the Arabic language. This process is further complemented by repetition reduction techniques, which aim to mitigate textual noise and enhance the overall consistency of the data. Subsequently, the data are adapted to the requirements of the pre-trained models to ensure coherent tokenization. The processed texts are then encoded into numerical representations that serve as inputs during training and evaluation. Finally, we conduct a comprehensive benchmarking study of five Transformer-based architectures to assess their effectiveness. The best-performing experimental setup yielded remarkable results on the AraSentEval 2026 benchmark, achieving a micro-F1 score of 75.96% on the official test set.
University of Tripoli at AraSentEval: Fine-Tuning MARBERTv2 and CAMELBERT for Multi-Dialect Arabic Sentiment Analysis
Abdusalam F. Ahmad Nwesri | Amani Bahlul Sharif | Sarah Farag S. Hmeid
Abdusalam F. Ahmad Nwesri | Amani Bahlul Sharif | Sarah Farag S. Hmeid
This paper presents our contribution to the AraSentEval 2026 shared task, specifically for Subtask 1: Arabic Dialect Sentiment Analysis, hosted at the OSACT7 workshop during LREC 2026. The task focuses on classifying the sentiment (positive, negative, neutral) of text written in four major Arabic dialects: Moroccan, Egyptian, Jordanian, and Saudi. We addressed this by fine-tuning several pre-trained language models, including MARBERTv2 and CAMELBERT, on the provided Multi-Dialect-Sent (MDS-3) dataset. Our best-performing system MARBERTv2, achieved a Macro F1-score of 84.29% on the official test set, securing fourth place among 13 participating teams. Our findings underscore the value of leveraging large pre-trained models tailored to dialectal Arabic for improved sentiment classification in this under-resourced domain.
LinguArabic at AraSentEval 2026: MARBERT for Multi-Dialect Arabic Sentiment Analysis
Norah Saud Alshahrani | Elham Abdullah Al-Qarni | Shatha Hussan Alshomrani
Norah Saud Alshahrani | Elham Abdullah Al-Qarni | Shatha Hussan Alshomrani
Sentiment analysis for Arabic dialects remains challenging due to substantial linguistic variation across dialects and the expansion of informal language in user-generated content. The AraSentEval 2026 shared task introduces a multi-dialect benchmark designed to evaluate sentiment classification systems on real-world Arabic data. In this paper, we present LinguArabic’s submission to the sentiment classification track of AraSentEval 2026. Our approach is based on fine-tuning MARBERT, a transformer model pre-trained on large-scale Arabic social media data that captures diverse dialectal patterns. To improve model robustness, we incorporate a multi-stage preprocessing pipeline that includes text normalization, dialect-aware lexical mapping, and confidence-based prediction adjustment. We specifically investigate the impact of advanced normalization rules in reducing lexical sparsity across various regional dialects. Experimental results show that the proposed system achieves a Macro F1-score of 0.8333 on the offcial evaluation set. Our findings highlight the importance of dialect-aware pretraining and preprocessing strategies for improving sentiment classification performance across diverse Arabic dialects, providing a scalable framework for real-world Arabic NLP applications.
up
Proceedings of the ParlaCLARIN V Workshop on Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora
Proceedings of the ParlaCLARIN V Workshop on Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora
Maria Eskevich | Vincent Vandeghinste | David Bodron
Maria Eskevich | Vincent Vandeghinste | David Bodron
ParlaCAP is an OSCARS Open Science cascading grant project aimed at extending the use of the ParlaMint parliamentary corpora beyond corpus linguistics into the wider Social Sciences and Humanities (SSH). While ParlaMint provides a rich, comparable collection of parliamentary debates and accompanying metadata, its broader uptake has been limited. ParlaCAP addresses this by enriching the data with automatically derived political agendas and sentiment, enabling new forms of comparative political analysis. Using recent advances in multilingual transformer models, the project annotates over 8 million speeches from 28 European parliaments in more than 20 languages. By integrating ParlaMint with the Comparative Agendas Project (CAP) coding schema, ParlaCAP produces a FAIR dataset suitable for cross-national research on interaction of policy, sentiment, and political identity. The enrichments rely on two models, XLM-R-ParlaSent and XLM-R-ParlaCAP, both performing comparably to human annotators. The latter is trained using a teacher–student approach, where GPT-4o-generated labels are used to fine-tune a scalable classifier. The dataset is available via the CROSSDA repository and a user-friendly API. The talk concludes with a series of use cases demonstrating how meaningful insights can be obtained with minimal technical effort.
Quantifying Code-Switching in a Ukrainian Parliamentary Dataset 1990-2021
Olha Kanishcheva | Maria Shvedova
Olha Kanishcheva | Maria Shvedova
Analyzing code-switching – the practice of mixing multiple languages in one discourse – remains a significant task in natural language processing (NLP). This study examines the Ukrainian-Russian bilingual context, focusing on quantifying language alternation in a multilingual dataset. We introduce metrics to assess linguistic boundaries and patterns, specifically addressing the complexities of processing texts where Ukrainian and Russian are used interchangeably, including word-level hybridization. Using a corpus of approximately 200,000 tokens derived from parliamentary transcripts (1990-2021), we apply code-switching metrics to identify frequency and patterns of language use. Our findings provide insights into bilingual communication dynamics and can be used to improve language identification models for mixed-language data.
Ours and Yours: A Discourse Analysis of Political Identity Markers in Slovenian Parliamentary Discourse
Meden Katja
Meden Katja
With recent enrichments of the ParlaMint corpora, new opportunities have emerged for examining a range of political and discursive phenomena. This paper utilises the Slovenian ParlaMint corpus to investigate the construction of political identities through the possessive pronouns ’our’ (slv. naš) and ’your’ (slv. vaš) in Slovenian parliamentary discourse. The analysis uses a corpus-assisted approach, combining text-type, collocation, and keyword analyses of the lemmas naš and vaš. The data are drawn from three subcorpora compiled from speeches of Members of Parliament: (1) Our, containing speeches in which naš occurs; (2) Your, containing speeches in which vaš occurs; and (3) Our&Your, containing speeches in which both lemmas co-occur within the same sentence. The results indicate a clear alignment with established patterns of positive self-representation and negative other-representation. Occurrences of ’our’ are predominantly associated with positively evaluative discourse, with Norms & Values emerging as a central category. In contrast, occurrences of ’your’ are more typically linked to negative sentiment and keywords, with Activities/Discourse identified as the most prominent category. These findings suggest that these possessive pronouns function as markers of ideological positioning and discursive polarisation in Slovenian parliamentary debates.
Representations of Europe and the European Union in Parliamentary Discourse from a Corpus-Assisted Perspective
Anna Kryvenko
Anna Kryvenko
This article examines how Europe and the European Union are represented in parliamentary discourse across three contrasting European political trajectories: the United Kingdom, Slovenia, and Ukraine. Using the ParlaMint 5.0 corpora — uniformly encoded, linguistically annotated, and enriched with sentiment and topic metadata — the study applies a longitudinal, cross-linguistic corpus-assisted discourse approach. Mentions of the EU and Europe in English, Slovenian, and Ukrainian were extracted through targeted queries, supplemented by sentiment profiling, topic distribution analysis, and collocational comparison across three built-in subcorpora (Reference, COVID, COVID,War). The findings show that although the two concepts can overlap, their discursive functions diverge systematically: Europe appears as a broader cultural and geopolitical frame, while the EU attracts more policy-oriented uses. These differences intensify at moments of institutional change or crisis, with sentiment around Europe displaying sharper fluctuations than sentiment around the EU. Cross-national patterns align closely with each country’s EU membership status — past, present, or aspirational — shaping how the EU is invoked, assessed, or contested. The study demonstrates the value of multilingual, longitudinal corpus analysis for tracing the evolution of political concepts and for understanding how parliaments discursively negotiate Europe’s shifting institutional and geopolitical landscape.
Towards ParlaMint-DE: Improving the Interoperability of the GermaParl Corpus of Plenary Protocols of the German Bundestag
Christoph Leonhardt | Andreas Blätte
Christoph Leonhardt | Andreas Blätte
With the number of machine-readable corpora of plenary protocols continuously increasing, concerns about the potentials of harmonisation and shared encoding standards gain prominence. Interoperability of corpora can contribute to innovative research, in particular when comparative analyses are concerned. The ParlaMint encoding schema introduced by CLARIN provides comprehensive guidelines towards this goal. This contribution shows how GermaParl, a large corpus of plenary protocols of the German Bundestag, is transformed from a TEI-inspired XML format to the ParlaMint encoding schema. Based on previous work, this paper presents an adjusted preparation pipeline and discusses challenges of advancing an established resource into a new data format. The prospective ParlaMint-DE corpus will make the plenary debates in Germany from 1949 to 2025 available in a highly interoperable data format. Clear documentation and taxonomies increase the usefulness of the resource in comparative analyses, whereas additional metadata and linguistic annotation broaden its general applicability.
From Transcripts to Insights: A Digital Corpus and Interactive Speech Analysis Platform for Turkish Parliamentary Records
Basak Tepe | Irem Nur Yildirim | Onur Gungor | Susan Uskudarli
Basak Tepe | Irem Nur Yildirim | Onur Gungor | Susan Uskudarli
Turkish parliamentary transcripts constitute a unique longitudinal record of the country’s political, institutional, and linguistic evolution starting from 1920. Yet much of this archive has remained computationally inaccessible due to scanned and analog typewritten transcripts, historical orthography, and heterogeneous formats. We present a unified, machine-readable corpus of the Grand National Assembly of Türkiye (TBMM), comprising 26,648 session transcripts and 1.7 million pages encompassing ten diverse parliamentary entities spanning a century of legislative history. In addition, we introduce an open-access web platform for speech-level analysis of parliamentary debates from 1983 to 2024. The platform integrates named entity recognition, topic modeling, and diachronic semantic shift detection, enabling exploration of discourse patterns across time and parties, including the frequency and thematic focus of speech activities of specific Members of Parliament. By bridging the gap between raw archival scans and modern NLP tools, the dataset and platform support reproducible research in NLP, digital humanities, and computational social science.
Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models
Luigi Curini | Alfio Ferrara | Giovanni Pagano | Sergio Picascia
Luigi Curini | Alfio Ferrara | Giovanni Pagano | Sergio Picascia
Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to digitise Italian parliamentary speeches have relied on traditional Optical Character Recognition pipelines, resulting in transcription errors and limited semantic annotation. In this paper, we propose a pipeline based on Vision-Language Models for the automatic transcription, semantic segmentation, and entity linking of Italian parliamentary speeches. The pipeline employs a specialised OCR model to extract text while preserving reading order, followed by a large-scale Vision-Language Model that performs transcription refinement, element classification, and speaker identification by jointly reasoning over visual layout and textual content. Extracted speakers are then linked to the Chamber of Deputies knowledge base through SPARQL queries and a multi-strategy fuzzy matching procedure. Evaluation against an established benchmark demonstrates substantial improvements both in transcription quality and speaker tagging.
Beyond OCR: Structural Segmentation and Speaker Attribution in Historical Italian Parliamentary Debates
Claudia Corbetta | Samuele Mazzei | Alessio Palmero Aprosio
Claudia Corbetta | Samuele Mazzei | Alessio Palmero Aprosio
Historical parliamentary debates are essential for longitudinal political and linguistic research, yet much early material remains available only as scanned images. In the Italian context, proceedings from 1848–1996 lack large-scale, structurally annotated, machine-readable representations. This paper addresses the challenge of transforming historical Italian parliamentary debates into structured corpora by moving beyond plain Optical Character Recognition (OCR) toward functional block segmentation and speaker attribution. We present detailed annotation guidelines and a manually annotated dataset of 300 randomly sampled pages. Two approaches are compared: (i) direct multimodal Large Language Model (LLM) annotation and (ii) a modular pipeline combining OCR with LLM-based structural reconstruction under zero-shot and few-shot prompting. Evaluation on a held-out test set shows that separating transcription from structural reasoning improves performance, with few-shot prompting yielding the most reliable results. The study demonstrates the feasibility of integrating LLM-based reasoning into historical parliamentary digitisation workflows.
Computational Political Landscape of the Netherlands and Prime Minister Schoof’s Position
Wessel Ledder | Iris Hendrickx
Wessel Ledder | Iris Hendrickx
This study presents a computational model of the Dutch political landscape during the Schoof government period, constructed using debate speeches from the House of Representatives. We construct a two-dimensional representation of the Dutch political landscape by fine-tuning a BERT model on parliamentary debate speeches and applying dimensionality reduction techniques to the resulting embeddings. We evaluate the validity of this model by comparing it to an independently developed model from an external research institute, finding that both models reveal similar patterns along the socio-economic left–right dimension. We also examine content patterns and word frequency distributions in targeted samples located at distinct regions of the landscape to interpret the model. We further evaluate the stability of the landscape to ensure that the observed patterns are not driven by random variation. Finally, we position Prime Minister Schoof within this computational landscape. Schoof was intended to be a neutral Prime Minister without any party affiliation that would represent the coalition parties of the government equally. Our analysis will show whether Schoof was indeed neutral in his statements or not.
up
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)
Haithem Afli | Houda Bouamor | Wajdi Zaghouani | Sahar Ghannay | Shehenaz Hossain
Haithem Afli | Houda Bouamor | Wajdi Zaghouani | Sahar Ghannay | Shehenaz Hossain
From News Streams to Narrative Intelligence Briefs: LLM-Assisted Political Discourse Analysis in the Hungarian 2026 Pre-Election Context
Ekaterina Loginova | Maksim Ermakov | Stephan Khramov
Ekaterina Loginova | Maksim Ermakov | Stephan Khramov
Civil-society organisations, journalists, and fact-checkers monitoring elections require scalable ways to convert high-volume political news into actionable narrative intelligence, yet most NLP pipelines stop at classification outputs that are difficult to operationalise. We present a methodology-driven case study assessing whether large language models, constrained by an explicit analytical schema and multi-stage validation, can reliably transform Hungarian pre-election news into structured narrative intelligence briefs. Using RSS-scraped content from 21 Hungarian-language sources (574 election-relevant articles), we implement a multi-stage pipeline: (1) per-article extraction of narrative event frames grounded in the Narrative Policy Framework (actor–action–target with role assignment and causal claims) and manipulation techniques from the SemEval propaganda taxonomy; (2) embedding-based clustering of narrative frames with domain classification; and (3) constrained brief generation producing five structured sections—narrative summary, character map, manipulation profile, escalation assessment, and counter-strategy—where counter-strategies are grounded in verified external sources via curated contextual cards and constrained by evidence-based de-escalation principles. We evaluate brief quality through dual-track evaluation combining three human domain experts and three LLM judges on a single brief, with a scaled 29-brief LLM-as-judge assessment, and document key failure modes across a defined taxonomy. We conclude with implications for trustworthy human-in-the-loop political NLP and the practical limits of LLM-assisted narrative intelligence.
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
Valeria Pastorino | Jasivan Alex Sivakumar | Nafise Sadat Moosavi
Valeria Pastorino | Jasivan Alex Sivakumar | Nafise Sadat Moosavi
The growing complexity and diversity of news coverage have made framing analysis a crucial yet challenging task in computational social science. Traditional approaches, including manual annotation and fine-tuned models, remain limited by high annotation costs, domain specificity, and inconsistent generalisation. Instruction-based large language models (LLMs) offer a promising alternative, yet their reliability for framing analysis remains insufficiently understood. In this paper, we conduct a systematic evaluation of several LLMs, including GPT-3.5/4, FLAN-T5, and Llama 3, across zero-shot, few-shot, and explanation-based prompting settings. Focusing on domain shift and inherent annotation ambiguity, we show that model performance is highly sensitive to prompt design and prone to systematic errors on ambiguous cases. Although LLMs, particularly GPT-4, exhibit stronger cross-domain generalisation, they also display systematic biases, most notably a tendency to conflate emotional language with framing. To enable principled evaluation under real-world topic diversity, we introduce a new dataset of out-of-domain news headlines covering diverse subjects. Finally, by analysing agreement patterns across multiple models on existing framing datasets, we demonstrate that cross-model consensus provides a useful signal for identifying contested annotations, offering a practical approach to dataset auditing in low-resource settings.
From Cairo to Cape Town: How African Twitter Shapes the Global Palestine-Israel Narrative
Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
African Twitter users are active shapers of the Palestine-Israel conversation but their contribution remains relatively understudied. Using 132.5K geo-located tweets from 2020 to 2023 and 451-term list of keywords in 33 languages, we identify three patterns in this context: (1) broad participation (Egypt supplies 43% of posts, yet Nigeria, South Africa, Kenya and Ghana contribute more than a third); (2) multilingual predominantly pro-Palestine amplification across Arabic, English, French, Swahili and other tongues, with 8% of tweets left “undetermined” by Twitter’s language detector; and (3) a humanitarian framing that centers civilian harm through hashtags such as #GazaUnderAttack and #PalestenianLivesMatter. We outline design implications for language-agnostic interfaces, low-friction source verification and cross-movement recommendation tools that foreground African epistemologies in global civic-tech systems.
Multimodal Analysis of State-Funded News Coverage of the Israel–Hamas War on YouTube Shorts
Daniel Miehling | Sandra Kübler
Daniel Miehling | Sandra Kübler
YouTube Shorts have become central to news consumption on the platform, yet research on how geopolitical events are represented in this format remains limited. To address this gap, we present a multimodal pipeline that combines automatic transcription, aspect-based sentiment analysis (ABSA), and semantic scene classification. The pipeline is first assessed for feasibility and then applied to analyze short-form coverage of the Israel–Hamas war by state-funded outlets. Using over 2,300 conflict-related Shorts and more than 94,000 visual frames, we systematically examine war reporting across major international broadcasters. Our findings reveal that the sentiment expressed in transcripts regarding specific aspects differs across outlets and over time, whereas scene-type classifications reflect visual cues consistent with real-world events. Notably, smaller domain-adapted models outperform large transformers and even LLMs for sentiment analysis, underscoring the value of resource-efficient approaches for humanities research. The pipeline serves as a template for other short-form platforms, such as TikTok and Instagram, and demonstrates how multimodal methods, combined with qualitative interpretation, can characterize sentiment patterns and visual cues in algorithmically driven video environments.
Characterization of User Engagement in Electronic News Media: A Case Study for India
Abhishek Kumar | Abhik Jana | Debi Prosad Dogra
Abhishek Kumar | Abhik Jana | Debi Prosad Dogra
Online news platforms have become central spaces for political discourse, playing a critical role in shaping public opinion and democratic participation. In large-scale democracies such as India, the combination of extensive user engagement, ideological polarization, and automated participation raises significant concerns regarding trust, transparency, and manipulation. This paper presents a large-scale empirical study of political discourse on a prominent Indian news platform, analyzing over 21,000 news articles and more than 1.5 million user comments. We investigate how ideological bias, sentiment dynamics, and non-organic user behavior interact to shape engagement patterns. Our methodology integrates hybrid article bias classification, large-scale sentiment analysis, heuristic-based bot detection, coordinated behavior analysis, and a focused examination of super-active users. In addition, we compare rule-based stance inference with large language model (LLM)-based stance classification to assess trade-offs between computational efficiency and contextual accuracy. The results reveal systematic sentiment skew, disproportionate influence by super-active and bot-like users, and coordinated campaigns aligned with specific political narratives. We conclude by discussing the implications of these findings for trust in online political discourse and reflecting on the dual role of generative AI as both an analytical tool and a potential vector for manipulation.
Analyzing Political Stances on Twitter/X in the Lead-up to the 2024 U.S. Election
Hazem Ibrahim | Farhan Kamrul Khan | Yasir Zaki | Talal Rahwan
Hazem Ibrahim | Farhan Kamrul Khan | Yasir Zaki | Talal Rahwan
Social media platforms play a pivotal role in shaping public opinion and amplifying political discourse, particularly during elections. However, the same dynamics that foster democratic engagement can also exacerbate polarization. To better understand these challenges, here, we investigate the ideological positioning of tweets related to the 2024 U.S. Presidential Election. To this end, we analyze 1,235 tweets from key political figures and 63,322 replies, and classify ideological stances into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. Using a classification pipeline involving three large language models (LLMs)—GPT-4o, Gemini-Pro, and Claude-Opus—and validated by human annotators, we explore how ideological alignment varies between candidates and constituents. We find that Republican candidates author significantly more tweets in criticism of the Democratic party and its candidates than vice versa, but this relationship does not hold for replies to candidate tweets. Furthermore, we highlight shifts in public discourse observed during key political events. By shedding light on the ideological dynamics of online political interactions, these results provide insights for policymakers and platforms seeking to address polarization and foster healthier political dialogue.
Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus
John E. Ortega | Rodolfo Joel Zevallos | Fabrício Carraro
John E. Ortega | Rodolfo Joel Zevallos | Fabrício Carraro
We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruvian Constitution using three state-of-the-art text-to-speech (TTS) architectures: XTTS v2, F5-TTS, and DiFlow-TTS. Our models are trained on independent Spanish and Quechua speech datasets with heterogeneous sizes and recording conditions, and leverage bilingual and multilingual TTS capabilities to improve synthesis quality in both languages. By exploiting cross-lingual transfer, our framework mitigates data scarcity in Quechua while preserving naturalness in Spanish. We release trained checkpoints, inference code, and synthesized audio for each constitutional article, providing a reusable resource for speech technologies in indigenous and multilingual contexts. This work contributes to the development of inclusive TTS systems for political and legal content in low-resource settings
Parliamentary debate constitutes a central arena of political power, shaping legislative outcomes and public discourse. Incivility within this arena signals political polarization and institutional conflict. This study presents a systematic investigation of incivility in the German Bundestag by examining calls to order (CtO; plural: CtOs) as formal indicators of norm violations. Despite their relevance, CtOs have received little systematic attention in parliamentary research. We introduce a rule-based method for detecting and annotating CtOs in parliamentary speeches and present a novel dataset of German parliamentary debates spanning 72 years that includes annotated CtO instances. Additionally, we develop the first classification system for CtO triggers and analyze the factors associated with their occurrence. Our findings show that, despite formal regulations, the issuance of CtOs is partly subjective and influenced by session presidents and parliamentary dynamics, with certain individuals disproportionately affected. An insult towards individuals is the most frequent cause of CtO. In general, male members and those belonging to opposition parties receive more calls to order than their female and coalition-party counterparts. Most CtO triggers were detected in speeches dedicated to governmental affairs and actions of the presidency. The CtO triggers dataset is available at: https://github.com/kalawinka/cto_analysis.
Operationalising the "Right to Be Forgotten" in LLMs: A Lightweight Sequential Unlearning Framework for Privacy-Aligned Deployment in Politically Sensitive Environments
Esen Kurt | Haithem Afli
Esen Kurt | Haithem Afli
Large Language Models (LLMs) are increasingly deployed in politically sensitive environments, where memorisation of personal data or confidential content raises regulatory concerns under frameworks such as the GDPR and its “right to be forgotten”. Translating such legal principles into large-scale generative systems presents significant technical challenges. We introduce a lightweight sequential unlearning framework that explicitly separates retention and suppression objectives. The method first stabilises benign capabilities through positive fine-tuning, then applies layer-restricted negative fine-tuning to suppress designated sensitive patterns while preserving general language competence. Experiments on the SemEval-2025 LLM Unlearning benchmark demonstrate effective behavioural suppression with minimal impact on factual accuracy and fluency. GPT-2 exhibits greater robustness than DistilGPT-2, highlighting the role of model capacity in privacy-aligned adaptation. We position sequential unlearning as a practical and reproducible mechanism for operationalising data erasure requirements in politically deployed LLMs.
Exploring Two Decades of Parliamentary Speeches on the Use of Narratives
Matti Wiegmann | Jürgen Neyer | Benno Stein
Matti Wiegmann | Jürgen Neyer | Benno Stein
Political scientists are interested in changes in political discourse over time. However, the topics of interest, such as the changes in support for or understanding of certain narratives, are often ill-defined and require deliberation, which prevents most lexical or metadata-based methods of temporal aggregation. To enable a diachronic analysis, we propose to model such settings as a series of binary document classification tasks – which current reasoning LLMs can adequately solve – and aggregate the decisions into a temporal signal. Specifically, we propose to use LLMs to classify if a parliament speech is in support of either of two narratives, and we use the monthly count of positives per narrative to track the support over time. We show that the classification is sufficiently accurate and use it to create detailed time series data showing support for the selected narratives in speeches given in the European Parliament from 2006 to 2023. The method is developed in close collaboration with political scientists and is considered an ideal starting point for diachronic analyses of political decision-making processes by domain experts.
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
Taja Kuzman Pungeršek | Peter Rupnik | Daniela Širinić | Nikola Ljubešić
Taja Kuzman Pungeršek | Peter Rupnik | Daniela Širinić | Nikola Ljubešić
This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method for building domain-specific policy topic classifiers. Applying the Comparative Agendas Project (CAP) schema to the multilingual ParlaMint corpus of over 8 million speeches from 28 parliaments of European countries and autonomous regions, we follow a teacher-student framework in which a high-performing large language model (LLM) annotates in-domain training data and a multilingual encoder model is fine-tuned on these annotations for scalable data annotation. We show that this approach produces a classifier tailored to the target domain. Agreement between the LLM and human annotators is comparable to inter-annotator agreement among humans, and the resulting model outperforms existing CAP classifiers trained on manually-annotated but out-of-domain data. In addition to the CAP annotations, the ParlaCAP dataset offers rich speaker and party metadata, as well as sentiment predictions coming from the ParlaSent multilingual transformer model, enabling comparative research on political attention and representation across countries. We illustrate the analytical potential of the dataset with three use cases, examining the distribution of parliamentary attention across policy topics, sentiment patterns in parliamentary speech, and gender differences in policy attention.
A Vocabulary Analysis of News Articles in Relation to the Political Orientation of Their Source and Their Thematic
Laurà ̈ne Cave | Gaël Lejeune
Laurà ̈ne Cave | Gaël Lejeune
Understanding how political orientation influences lexical choices is essential for detecting bias and framing in news media. In this paper, we present a computational framework for identifying nouns whose interpretation varies across politically divergent newspapers. Using a large corpus of French news articles published in 2024, we categorize texts by topics and political orientation. We use contextual embeddings to cluster occurrences of nouns to detect semantic variations and dissimilarity among sources. This allows us to map semantic distances between newspapers and identify polarized or editorially marked lexical choices. Our results show that topics, polysemy, and editorial priorities contribute differently to lexical divergence. We discuss these findings and highlight how contextual embeddings can help reveal semantic biases that would remain invisible through frequency-based methods. We conclude by outlining perspectives for improving topic classification and the clustering method, exploring alternative divergence measures, conducting a qualitative analysis of our results, and extending the framework to other languages or genres.
Posts Talk Policy, Stories Don’t: Policy-Issue Detection on Instagram with Fine-Tuned Transformers and Prompted LLMs
Michael Achmann-Denkler | Mario Haim | Christian Wolff
Michael Achmann-Denkler | Mario Haim | Christian Wolff
Policy issues are central to election campaigns, yet systematic analyses of issue communication on Instagram remain scarce — particularly for ephemeral Stories. We develop and evaluate automated methods for detecting the binary presence of policy issues in Instagram posts and Stories from the 2021 German federal election. Drawing on a gold-standard dataset of 1,357 annotated documents across three textual channels (captions, OCR-extracted image text, and speech transcripts), we compare a fine-tuned German transformer (GBERT) with multiple LLM prompting strategies (zero-shot, few-shot, retrieval-augmented). Both approaches prove effective: GBERT achieves a cross-validated macro F1 of 0.90, closely matched by GPT-o3 under few-shot prompting (0.88). Substantively, policy visibility varies far more by content format than by party: 70% of posts contain policy references compared to only 17% of Stories, a pattern that holds consistently across all eight parties. An exploratory topic model confirms that parties reproduce familiar issue-ownership profiles within the subset of policy-relevant texts. Our results establish binary issue detection as a feasible foundation for studying policy communication in multimodal, ephemeral social media environments.
Mobilize, Inform, Interact: Classifying Political Calls-to-Action Types on Instagram
Michael Achmann-Denkler | Clara Helmig | Jakob Fehle | Mario Haim | Christian Wolff
Michael Achmann-Denkler | Clara Helmig | Jakob Fehle | Mario Haim | Christian Wolff
Calls-to-action (CTAs) are central to digital campaigning, yet computational research has largely focused on binary detection only. We address CTA type classification in German Instagram campaign texts (posts and ephemeral stories), distinguishing Support, Inform, Interact, and No CTA. With limited annotated data, we benchmark a fine-tuned GBERT model against GPT models using zero-shot, few-shot, and retrieval-augmented few-shot prompting in a multi-label setup. Both approaches reach similar performance in five-fold cross-validation (macro-F1 ca. 0.79), with persistent difficulty on the rare Interact category. As a proof of concept, we apply the selected setup to the 2021 federal election corpus and show that parties varied not only in overall CTA use but also in how they balanced appeals across posts versus stories. The results demonstrate the feasibility of CTA type classification with modest data and position retrieval-augmented prompting as a practical alternative to supervised fine-tuning.
Hate Speech and Hate Crime: A Cross-Disciplinary Analysis of Xenophobia in Greece
Maria Gavriilidou | Vasiliki Georgiadou | Lamprini Rori | Maria Pontiki
Maria Gavriilidou | Vasiliki Georgiadou | Lamprini Rori | Maria Pontiki
This paper investigates the correlation of hate speech and hate crime, in a inter-disciplinary approach, using a computational hate speech and hate crime detection method in a socio-political science framework, coupling Natural Language Processing with Political Sciences. The study focuses on Greece in the turbulent period from 2015 to 2022 (a period marked by economic, refugee, foreign policy, and pandemic crises); it analyzes tweets to discern linguistic patterns used to verbally attack predefined target groups consisting of ethnic and religious minorities in the country. Furthermore, it investigates hate crimes reported in the press, against the same target groups and during the same period and proceeds to examine correlations between xenophobic attitudes expressed verbally through social media, and those manifested as physical attacks in real life.
Navigating Global AI Regulation: A Multi-Jurisdictional Retrieval-Augmented Generation System
Courtney Ford | Ojas Rane | Susan Leavy
Courtney Ford | Ojas Rane | Susan Leavy
Navigating AI regulation across jurisdictions is increasingly difficult for policymakers, legal professionals, and researchers. To address this, we present a multi-jurisdictional Retrieval-Augmented Generation system for global AI regulation. Our corpus includes 241 documents across 73 jurisdictions, ranging from formal legislation like the EU AI Act to unstructured policy documents such as national AI strategies. The system makes three technical contributions: type-specific chunking that preserve legal structure across heterogenous documents; conditional retrieval routing with entity detection and metadata for legal citations; and priority-based re-ranking to boost enacted legislation over policy and secondary sources. Evaluation of 50 queries reveals strong performance across both single-entity and multi-jurisdictional questions, achieving 0.87 average faithfulness and 0.84 average answer relevancy. Single-entity queries achieve 0.86 average faithfulness and 0.92 average answer relevancy, while multi-jurisdictional comparison queries achieve 0.88 average faithfulness and 0.75 average answer relevancy. These findings highlight the effectiveness of domain-specific retrieval strategies for navigating complex, heterogenous regulatory corpora.
Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation
Maryam Fooladi | Federico Bottino
Maryam Fooladi | Federico Bottino
Sentiment analysis remains the dominant computational approach for evaluating political news, yet its ability to capture the rhetorical complexity of political discourse is increasingly questioned. This paper presents a systematic comparison between a transformer-based sentiment classifier (RoBERTa) and a Large Language Model-based multi-dimensional framing analysis framework across 50 political news articles from 17 international outlets. While RoBERTa classifies 70% of articles as neutral and reduces political discourse to a three-way polarity scale, the LLM-based framework captures 13 numerical dimensions including bias direction and intensity, manipulation indicators (cherry-picking, loaded language, false equivalence), sensationalism, and communicative intent. Our correlation analysis reveals only weak-to-moderate relationships between sentiment polarity and framing dimensions (maximum Pearson r = 0.38, p < 0.01), demonstrating that these approaches measure fundamentally different properties of political text. Through case studies, we show that sentiment-neutral articles can exhibit extreme manipulation patterns, while highly negative articles may reflect factual reporting on inherently negative events. These findings argue for moving beyond sentiment as a proxy for media quality, toward multi-dimensional frameworks that can reveal the rhetorical strategies invisible to polarity-based analysis. All data and analysis code will be made available upon acceptance.
Evaluating the Abilities of LLMs and SpeechLMs in Discovering Implicit Contents of Italian Political Speeches
Lorenzo Gregori | Walter Paci | Alessandro Panunzi
Lorenzo Gregori | Walter Paci | Alessandro Panunzi
This research investigates the pragmatic competence of Large Language Models (LLMs) in interpreting implicit meanings within Italian political discourse. Using the IMPAQTS-PIDMM dataset, which is a multimodal benchmark derived from the 2.5-million-token IMPAQTS corpus, the experiment evaluates how effectively models identify tendentious content such as presuppositions and implicatures. The study compares the performance of text-only LLMs against speech-based models (SpeechLMs) that process both audio and transcriptions to determine if acoustic cues enhance understanding. The results reveal that text-only models significantly outperform multimodal variants, with Qwen2.5-72B achieving the highest global accuracy of 0.863. Surprisingly, the inclusion of audio did not improve performance, as SpeechLMs like GPT-4o-mini-audio-preview and Qwen2-Audio-7B-Instruct obtained lower accuracy scores and a higher frequency of missed answers compared to their text-only equivalents. Across all tested architectures, models generally demonstrated a superior ability to process presuppositions over implicatures.
Reproducibility under Threat: Proposing a Framework for Reliable LLM-Research in Psychology and Computational Social Science
Kevin Dirk Kiy | Alexander Porshnev | Dounia Lakhzoum | Dermot Lynott | Diarmuid O’Donoghue | Manokamna Singh
Kevin Dirk Kiy | Alexander Porshnev | Dounia Lakhzoum | Dermot Lynott | Diarmuid O’Donoghue | Manokamna Singh
The integration of artificial intelligence (AI), particularly large language models (LLMs), into research across the social sciences has accelerated innovation but also introduced significant challenges to reproducibility - a cornerstone of scientific integrity. In this review of scientific practices, we examine the reproducibility crisis in AI-driven research with a focus on psychology, identifying common pitfalls, reviewing proposed solutions, and advocating for best practices. Common pitfalls in current practices in the social sciences are identified and highlighted through synthesized research scenarios, such as: (1) using inaccessible datasets or language models with restricted access, (2) treating black-box API outputs as stable observations ignoring updates and hidden changes, (3) producing single runs for measurements instead of stochastic draws for aggregated performances, (4) failing to report full LLM version, prompting, and sampling parameters, and (5) opaque training and fine-tuning of LLMs. Our recommended practices include precisely documenting the model used, fixing all inference parameters, using automation and scripts to control prompts, context, and outputs, and standardizing the environment and API conditions. By embracing transparency and methodological rigour, we can transform the challenges of AI-driven research into opportunities for more robust and impactful science, ensuring that innovation never comes at the cost of credibility.
How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment
Kokil Jaidka | Saifuddin Ahmed
Kokil Jaidka | Saifuddin Ahmed
This study analyzes a publicly released dataset from a discontinued field experiment on Reddit’s r/ChangeMyView. The intervention, conducted by unknown, external researchers and halted following ethical backlash, involved undisclosed AI-generated accounts engaging users in live debate. After public disclosure, Reddit authorized moderators to release an archive of the AI-generated comments, creating a rare opportunity to examine how large language models operated in an identity-rich deliberative forum without disclosure. We conduct a structured content analysis of this corpus, evaluating identity performance, authority signaling, alignment strategies, and activation of cognitive heuristics. Identity targeting or adoption appears in over two-thirds of comments, alignment moves and authority claims in nearly all of them, and cognitive-bias triggers—particularly confirmation bias, representativeness, and availability—in the large majority. These patterns co-occur systematically, composing a rhetorical architecture calibrated for persuasive efficiency rather than authentic deliberative participation. Compared against human-authored CMV counter-arguments, the agents inverted the typical distribution on every dimension: denser authority use, more adversarial alignment, and heavier reliance on external citation over experiential grounding. In such environments, distinctions between authentic and synthetic epistemic standing grow increasingly opaque—an asymmetry that disclosure mandates alone cannot address. The results point toward auditing frameworks capable of assessing how AI systems structure credibility, not merely whether they are present. Our dataset is available at https://github.com/kokiljaidka/UnauthorizedRedditCMVPosts
An Examination of the Party Leanings of Large Language Models
Lars Bungum | Charles Huang | Jari Bakken
Lars Bungum | Charles Huang | Jari Bakken
This paper examines the party leanings of international and Norwegian large language models. The experiments are two-fold; first they are asked to answer the question of a Valgomat–an election affiliation guide–as a neutral observer, and secondly as if it were a paying party member of the parties in the data. Results show that the neutral prompting show centrist leanings, whereas models struggle with mimicking party members. Models with additional training on Norwegian text performed better on this task.
Attitude Identification through Parameter-Efficient Fine-Tuning
Mariia Anisimova | Gabriella Lapesa | Sarka Zikanova
Mariia Anisimova | Gabriella Lapesa | Sarka Zikanova
We investigate automatic attitude detection in UN Security Council speeches using adapters. Following Martin and White’s Appraisal Theory, we identify three types of evaluative language: affect (emotional responses such as hope or concern), judgement (ethical evaluations of behavior), and appreciation (valuations of objects or situations). Training only 0.95% of BERT-large’s parameters, adapters achieve F1 scores ranging from 0.76 (affect) to 0.46 (appreciation), approaching full fine-tuning performance while enabling rapid task-specific experimentation. Differences in observed evaluation metrics mirror the pattern of the human inter-annotator agreement. This correlation suggests that computational difficulty reflects genuine linguistic ambiguity. Affect benefits from conventionalized diplomatic expressions, while appreciation faces context-dependent evaluation and severe class imbalance. Analysis demonstrates that evaluative intensity varies systematically across diplomatic contexts, with implications for corpus design in specialized discourse analysis.
Do LLMs Transfer Political Framing across Languages? A Cross-Lingual Analysis of LLM-Generated Discourse
Nooredeen Awwad | Ebtihal Ismail Enfes | Kolawole John Adebayo
Nooredeen Awwad | Ebtihal Ismail Enfes | Kolawole John Adebayo
As large language models (LLMs) increasingly mediate political information across linguistic contexts, concerns emerge regarding cross-lingual consistency in political framing. We investigate whether multilingual LLMs generate systematically different rhetorical and semantic frames when prompted in English versus Arabic on the politically salient issue of migration. Focusing on two widely used models, i.e., GPT-4o and Jais-13B, we implement a controlled prompt design (N = 800 generations; 400 per language), to isolate language as the primary experimental variable. We introduce a mixed method evaluation framework that combines lexical frame analysis, statistical association testing, and qualitative discourse analysis. Our results show a significant association between language and framing distribution (χ2 = 43.32, p = 2.11 × 10−9). While security-oriented framing is prominent in both languages, English generations exhibit substantially higher rates of institutional and legislative framing, whereas Arabic generations show greater concentration in security and communitarian discourse. These findings indicate that input language acts as a conditioning signal that systematically modulates political framing within multilingual LLMs, even under controlled semantic prompts. We conceptualize this phenomenon as cross-lingual framing drift and discuss its implications for multilingual alignment, political bias evaluation, and global information ecosystems. We conclude by outlining an evaluative protocol for detecting language-conditioned asymmetries in generative models. We make all data, code, and experimental settings publicly available at: https://github.com/NRAwwad/-A-Cross-Lingual-Analysis-of-Political-Framing-in-English-and-Arabic.git.
When Neutral Turns Negative: Cross-Domain Failure Modes in Hinglish Political Sentiment Analysis
Rahul Chennuru | Kolawole John Adebayo
Rahul Chennuru | Kolawole John Adebayo
Sentiment analysis models are increasingly deployed to analyze political discourse, yet strong in-domain performance does not guarantee robustness under domain shift. We study cross-domain generalization in Hinglish (Hindi–English code-mixed) sentiment analysis by evaluating a fine-tuned XLM-RoBERTa classifier, trained on 29,000 general-domain Hinglish sentences, on a curated benchmark of politically oriented Hinglish text. While the model achieves 92.02% accuracy in-domain, performance drops to 71.83% under political domain shift. Error analysis reveals a pronounced directional bias with 48.9% of neutral political statements misclassified as negative, indicating a systematic neutrality-to-negative shift. In addition, 87.5% of incorrect predictions are assigned confidence scores above 95%, pointing to severe miscalibration under distribution shift. We further compare these results against an instruction-tuned large language model (Llama 3.3), which achieves 90.85% zero-shot accuracy and 94.37% accuracy with contextual prompting, while substantially reducing neutrality bias. Our findings indicate the need for domain-aware evaluation, calibration diagnostics, and explicit reporting of failure modes when deploying sentiment models in politically sensitive settings.
A comprehensive framework was developed to detect political bias in Arabic news articles, with a case study focusing on media reporting of the Palestinian issue. The methodology integrates MARBERT contextual embeddings with classical and deep learning classifiers, including SVM, Logistic Regression, Random Forest, and LSTM. The scalability of data processing was ensured through Apache Spark for potential real-time deployment. Experimental results showed that fine-tuned MARBERT embeddings combined with LSTM achieved the highest classification accuracy of 0.87, along with notable improvements in F1-scores across the pro, against, and neutral categories. These findings highlight the effectiveness of domain-specific fine-tuning of transformer models for political bias classification. The study also addressed class imbalance using SMOTE and class weighting strategies, and assessed feature robustness using multiple vectorization techniques.
PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs
Shariar Kabir | Yue Dong | Kevin Esterling
Shariar Kabir | Yue Dong | Kevin Esterling
Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspective by examining how ideological tendencies vary across topics and how consistently models maintain their positions, a property we refer to as stability. To capture this dimension, we propose PReSS (Political Response Stability under Stress), an automated black-box framework that evaluates LLMs by jointly considering model and topic context, categorizing responses into four stance types: stable-left, unstable-left, stable-right, and unstable-right. Applying PReSS to 9 widely used LLMs across 19 political topics reveals substantial variation in stance stability; for instance, a model that is left-leaning overall can exhibit stable-right behavior on certain topics. This highlights the importance of topic-aware and fine-grained evaluation of political ideologies of LLMs. Moreover, stability has practical implications for controlled generation and model alignment: interventions such as debiasing or ideology reversal should explicitly account for stance stability. Our empirical analyses reveal that when models are prompted or fine-tuned to adopt the opposite ideology, unstable topic stances are more likely to change, whereas stable ones resist modification. Thus, treating stability as a moderating factor provides a principled foundation for understanding, evaluating, and guiding interventions in politically sensitive model behavior.
Large Language Models Unpack Complex Political Opinions through Target-Stance Extraction
Ozgur Togay | Javier Garcia-Bernardo | Florian Kunneman | Anastasia Giachanou
Ozgur Togay | Javier Garcia-Bernardo | Florian Kunneman | Anastasia Giachanou
Political polarization emerges from a complex interplay of beliefs about policies, figures, and issues. However, most computational analyses reduce discourse to coarse partisan labels, overlooking how these beliefs interact. This is especially evident in online political conversations, which are often nuanced and cover a wide range of subjects, making it difficult to automatically identify the target of discussion and the opinion expressed toward them. In this study, we investigate whether Large Language Models (LLMs) can address this challenge through Target-Stance Extraction (TSE), a recent natural language processing task that combines target identification and stance detection, enabling more granular analysis of political opinions. For this, we construct a dataset of 1,084 Reddit posts from r/NeutralPolitics, covering 138 distinct political targets and evaluate a range of proprietary and open-source LLMs using zero-shot, few-shot, and context-augmented prompting strategies. Our results show that the best models perform comparably to highly trained human annotators and remain robust on challenging posts with low inter-annotator agreement. These findings demonstrate that LLMs can extract complex political opinions with minimal supervision, offering a scalable tool for computational social science and political text analysis.
Can Large Language Models Facilitate Qualitative Political Narrative Analysis?
Luke Stephens | Clare Llewellyn | Lauren Rogers | Constantine Kyritsopoulos | Arman Prangere | Feiteng Long | Peyton Snyder | Laura Cram
Luke Stephens | Clare Llewellyn | Lauren Rogers | Constantine Kyritsopoulos | Arman Prangere | Feiteng Long | Peyton Snyder | Laura Cram
This study evaluates whether Large Language Models (LLMs) can facilitate qualitative political narrative analysis by comparing outputs from four models—Mistral, Llama, ChatGPT-4o, and DeepSeek—against narrative analyses written by expert scholars. Using European Union State of the Union speeches (2010–2023), we examine migration and solidarity narratives through semantic and lexical similarity metrics alongside systematic validation. The narrative scholars demonstrate strong semantic alignment despite differences in wording, establishing a benchmark for interpretive consistency. Across both topics, the models produce lexical and semantic similarity scores that are broadly comparable to those observed between the scholars themselves, with differences at these levels often marginal. However, similarity metrics do not provide the full picture. Validation reveals model-specific weaknesses that are not captured by lexical or semantic alignment alone, including factual errors, over-structural abstraction, and difficulty engaging less salient narrative threads. These findings demonstrate that LLMs can produce narratives that align closely with human outputs in semantic and lexical similarity, yet these measures alone are insufficient to assess interpretive quality.
The integration of large language models into political discourse analysis creates new opportunities for comparative research, policy analysis, and civic technology, while introducing material risks for democratic accountability. This paper argues that cultural adaptation is a prerequisite for trustworthy deployment of large language models in political communication across diverse linguistic and institutional contexts. Current systems remain shaped by English dominant data, uneven multilingual coverage, and assumptions grounded in a narrow range of political institutions and discourse conventions, producing systematic errors when applied across cultures. We formalize cultural adaptation across translation, discourse, and ontology levels, identify recurring cultural failure modes in political NLP, and propose an operational evaluation matrix grounded in cultural fidelity, calibration, and democratic safety. Building on political text analysis, sociotechnical auditing, and cross cultural pragmatics, we outline methodological pathways including participatory dataset development, culturally aware transfer learning, and benchmark design that makes cultural adaptation empirically measurable. We conclude by clarifying governance constraints and scope conditions under which culturally adaptive political NLP can support democratic legitimacy.
Automated Analysis of Global AI Safety Initiatives: A Taxonomy-Driven LLM Approach
Takayuki Semitsu | Naoto Kiribuchi | Kengo Zenitani
Takayuki Semitsu | Naoto Kiribuchi | Kengo Zenitani
We present an automated crosswalk framework that compares an AI safety policy document pair under a shared taxonomy of activities. Using the activity categories defined in Activity Map on AI Safety as fixed aspects, the system extracts and maps relevant activities, then produces for each aspect a short summary for each document, a brief comparison, and a similarity score. We assess the stability and validity of LLM-based crosswalk analysis across public policy documents. Using five large language models, we perform crosswalks on ten publicly available documents and visualize mean similarity scores with a heatmap. The results show that model choice substantially affects the crosswalk outcomes, and that some document pairs yield high disagreements across models. A human evaluation by three experts on two document pairs shows high inter-annotator agreement, while model scores still differ from human judgments. These findings support comparative inspection of policy documents.
up
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Maciej Ogrodniczuk | Petya Osenova | Tanja Wissik
Maciej Ogrodniczuk | Petya Osenova | Tanja Wissik
PressMint: Towards Interoperable Corpora of Historical Newspapers
Tomaž Erjavec | Matyáš Kopp | Maciej Ogrodniczuk | Petya Osenova | German Rigau
Tomaž Erjavec | Matyáš Kopp | Maciej Ogrodniczuk | Petya Osenova | German Rigau
This paper presents Project X (name anonymized for review), an ongoing initiative to compile a multilingual, comparable, annotated, translated, and interoperable collection of European historical newspaper corpora. Spanning 17 countries and covering 15 languages, the project addresses a key shortcoming of existing newspaper resources: their lack of interoperability, which limits cross-lingual and transnational research. Building on the infrastructure and experience of the ParlaMint projects, the project adapts established encoding guidelines, validation workflows, and open-source tools to historical newspaper data. We outline the overall project architecture, the corpus encoding scheme, and the GitHub-based framework supporting collaborative development and quality control. The paper further describes the sample linguistic annotation pipeline, including OCR correction, text normalisation, and annotation within the Universal Dependencies framework, with attention to challenges posed by historical language varieties. The resulting FAIR, openly available corpora are intended to support comparative, diachronic research across the humanities and social sciences.
This article presents the Polish contribution to the PressMint project, a CLARIN initiative aimed at creating a pan-European, multilingual corpus of historical newspapers. The Polish dataset consists of three subcorpora spanning 110 years (1830–1939). The first two components are drawn from the Microcorpus of Nineteenth-Century Polish (its short press texts and journalistic texts subcorpora), each containing 200 samples of brief news items and journalistic articles from diverse periodicals. The third component, the InterWar Corpus, covers the period 1918–1939 and comprises approximately 6.5 million words from complete newspaper issues, representing the territory of the interwar Republic of Poland. The authors argue for the scholarly value of historical press, highlighting its precise chronological dating as a key advantage for diachronic research despite challenges such as heterogeneous content and anonymous authorship. The conversion pipeline maps source metadata to a standardized TEI format and enriches texts with linguistic annotation using the Hydra NLP tool, providing lemmatization, part-of-speech tagging (mapped to Universal Dependencies), dependency parsing, and named entity recognition. The resulting openly accessible dataset enables cross-linguistic comparison and distant reading of historical press materials on a European scale.
PressMint QuickCheck: Operationalising Readiness Diagnostics for Interoperable Historical Newspaper Corpora
Elena Battaner Moro | Almudena Caballos Villar | María Cuevas Riaño | Marina Miguez Lamanuzzi | Dolores Romero López
Elena Battaner Moro | Almudena Caballos Villar | María Cuevas Riaño | Marina Miguez Lamanuzzi | Dolores Romero López
PressMint QuickCheck is a lightweight, reproducible readiness diagnostic for historical newspaper collections. Given a candidate dataset (ZIP export or IIIF manifests), it detects which components are present, identifies interoperability-critical metadata gaps, and applies lightweight OCR sanity checks. It produces three standardised artefacts: a human-readable readiness report, a minimal normalised manifest (CSV), and a tentative v1 scorecard (suitability_score 0-4) for prioritisation across collections. The workflow is delivered as a Colab-first notebook (no installation required). A key design decision treats content_language and metadata_language declarations as first-class interoperability signals, reflecting the multilingual scope of PressMint and ParlaMint corpora projects.
We present a new European Portuguese corpus of newspapers from the 19th and early 20th centuries, integrated in the recent PressMint project, whose goal is to provide a set of comparable newspaper corpora for European languages in that time frame. We discuss the raw data that was previously available, as well as new data specifically compiled for the project, and the challenges involving OCR, text recognition, and different orthographical norms. We describe the pipeline setup for XML encoding and annotation, partially based on work developed for the ParlaMint corpora. The corpus is currently under development and will be made freely available at the end of the project, as part of the PressMint corpora.
Historical Newspapers in the General Regionally Annotated Corpus of Ukrainian (GRAC): Current State and PressMint Integration Prospects
Maria Shvedova | Arsenii Lukashevskyi
Maria Shvedova | Arsenii Lukashevskyi
This paper presents the historical newspaper collection of the General Regionally Annotated Corpus of Ukrainian (GRAC) and outlines its prospective integration into the PressMint infrastructure. The collection comprises 117 newspaper titles published before 1950, totaling 23.6 million tokens, and reflects the political fragmentation, regional variation, and orthographic diversity of Ukrainian-language press from the late nineteenth to mid-twentieth century. We describe the corpus composition, temporal and geographic distribution, and metadata architecture. Special attention is given to morphosyntactic annotation challenges arising from the old Western Ukrainian orthography (Zhelekhivka), as well as issues related to annotating historical texts using the rule-based TagText parser and neural UDPipe2 models. The paper compares GRAC’s vertical format and metadata system with the TEI-based PressMint standard, identifying technical and conceptual harmonization challenges. Integrating GRAC newspapers into PressMint will facilitate comparative research on language policy, regional standardization, and media discourse within a broader European context.
CLARIAH-ES PressMint: Building Interoperable Corpora of Historical Press in Spain
Ainara Estarrona | Aritz Farwell | German Rigau | Xabier Goenaga
Ainara Estarrona | Aritz Farwell | German Rigau | Xabier Goenaga
This paper describes CLARIAH-ES’s contribution to PressMint in Spain as a distributed effort across regional nodes (e.g., Catalonia, Madrid, Basque Country, Galicia, Canary Islands, Alicante), each developing manageable corpora in partnership with key repositories such as ARCA, Patrimonio Digital Complutense, Euskariana, Jable, Galiciana, and the BVMC periodicals portal. A central technical challenge is heterogeneous legacy OCR quality, motivating experiments with AI/LLM-assisted OCR renewal, normalization layers, and linguistic enrichment (e.g., NER and entity linking). This effort is situated alongside ongoing dissemination and the EOSC Mesh “historical newspapers” use-case work aimed at scalable discovery, access, and federated computation over interoperable historical press data.
Towards an Interoperable Corpus of Austrian Historical Newspapers: The case of PressMint-AT
Tanja Wissik | Jona Hassenbach | Hannes Pirker | Claudia Resch | Stefan Resch
Tanja Wissik | Jona Hassenbach | Hannes Pirker | Claudia Resch | Stefan Resch
In this paper the PressMint-AT project is presented, which aims to create a historical newspaper corpus based on the Wiener Abendpost. The quality of automatic text recognition (ATR) is a key factor in creating historical newspaper corpora. Therefore, the performance of established ATR tools, multimodal large language models (LLMS), and existing full-text transcriptions provided by the Austrian National Library via ANNO is evaluated in order identify the most suitable approach for the PressMint-AT project. Even though recent research has demonstrated promising results for OCR tasks using multimodal LLMs, the experiments presented in this paper show, that PERO OCR achieves the best performance for the PressMint-AT dataset.
A Growing Literature of the Public Sphere: Fiction in Danish Newspapers (1666–1850)
Pascale Feldkamp | Alie Lassche | Rie Eriksen | Kit Morgenstjerne | Kristoffer Nielbo | Johan Heinsen | Yuri Bizzoni
Pascale Feldkamp | Alie Lassche | Rie Eriksen | Kit Morgenstjerne | Kristoffer Nielbo | Johan Heinsen | Yuri Bizzoni
Digitized literary corpora of the 19th century largely focus on standalone volumes, sidelining the broader and more diverse literary production of the period. Fiction published in less enduring formats – such as novellas and serialized pieces in newspapers – remains underexplored, particularly for low-resource languages like Danish, despite the growing availability of digitized newspaper archives. This paper addresses that gap by identifying and tagging fiction in Danish newspapers (1666–1850). We (1) present a manually annotated dataset of 1,831 articles with both binary (fiction/nonfiction) and fine-grained subcategories (travelogue, biography, essay), and (2) evaluate a document-embedding classifier that achieves an F1-score of up to 0.89 for the fiction/nonfiction distinction. Building on this pipeline, we further provide two resources for future research: (a) fiction probability scores for nearly five million newspaper articles (n=4,898,084), and (b) a small, cleaned, and curated subset of newspaper fiction (n=139), intended as a growing resource.
PressMint is a CLARIN initiative that aims to build multilingual, comparable and interoperable corpora of historical newspapers. For Hungarian, the main challenge is not a lack of material but fragmentation: newspapers are distributed across several portals, with heterogeneous metadata, access paths and OCR quality. This extended abstract reports the current status of the Hungarian PressMint subcorpus, focusing on the 19th century and the early 20th century (roughly 1800–1920). We describe two project artefacts already used in practice: a structured source inventory and a validation-driven repository. We summarise source scouting across Europeana, Hungaricana, OSZK–EPA, DiFMOE and related portals, including a curated 12-title Hungaricana manual-download pilot list with explicit target coverage periods. We then outline a reproducible pipeline for acquisition, OCR, layout analysis and conversion to PressMint-compatible TEI with facsimile linkage. Finally, we specify near-term deliverables for a first Hungarian release candidate and the evaluation steps planned for OCR and layout processing.
Towards a Bulgarian Historical Newspaper Corpus – Construction of Reading Order over the Text in Searchable PDFs
Nikolay Paev | Stefan Marinov | Ivan Kratchanov | Petya Osenova | Kiril Simov
Nikolay Paev | Stefan Marinov | Ivan Kratchanov | Petya Osenova | Kiril Simov
The determine the reading order of the text extracted from a searchable PDF produced by an OCR software from an old newspaper is the first task in the process of preparation of corpora of old newspapers. In the paper we present an algorithm for generation of reading order of black selected from the corresponding PDF. Also we performed a tuning of the parameters of the algorithm. The optimization provides 10 % improvement.
A Survey of the Digitisation of German Newspapers in interwar Lithuania (1918–1940)
Lina Plaušinaitytė | Heike Zinsmeister
Lina Plaušinaitytė | Heike Zinsmeister
This paper presents a survey of the preservation and digitisation status of the German-language press published in interwar Lithuania which existed between 1918 and 1940. Produced within this newly established and ethnically diverse republic which was operating in accordance with the European Minority Protection Regime, German-language newspapers and other periodicals formed a relevant part of the country’s multilingual press. They represent an interesting yet underexplored resource for historical and linguistic research. The survey summarises bibliographic information and the results of earlier digitisation projects. It further addresses challenges for optical character recognition (OCR) of newspaper facsimiles. Although systematic digitisation remains future work, the paper identifies major challenges for OCR within this collection in particular in relation to typographic variation.
Toward Interoperable and Scalable Representations of Complex Heterogeneous Digitized Historical Media
Pauline Conti | Simon Clematide | Maud Ehrmann
Pauline Conti | Simon Clematide | Maud Ehrmann
The value of digitized historical media archives for computational historical research is now well established, yet an underexplored challenge concerns data management itself: how to represent and process, at scale, complex primary sources that vary widely in digitization granularity, refinement quality, and archival organization and curation practices. This paper presents the data representation framework designed for large-scale processing and indexing of historical newspapers and radio broadcasts developed within the Impresso project. Grounded in a structured characterization of the heterogeneity found in digitized historical media collections, it identifies the distinct dimensions along which collections diverge and the challenges they pose for a unified representation and processing. The framework navigates the competing demands of machine learning pipelines requiring uniform and lightweight document representations, information retrieval systems requiring well-defined indexable content units, user-facing interfaces requiring fidelity to original sources, and the need to return semantically enriched data to archival holders in interoperable formats. We describe the design principles guiding the framework and discuss how it reconciles these constraints across highly heterogeneous collections into a unified and research-ready corpus.
Data Matters: Looking for High-Quality Corpora to Build Robust and Reliable Models for Humanists
Jaione Macicior-Mitxelena | Ana García-Serrano
Jaione Macicior-Mitxelena | Ana García-Serrano
The digitization of Spanish historical newspapers poses significant challenges due to low scan quality, typographical diversity, complex layouts and linguistic variation from contemporary Spanish. While advances in Optical Character Recognition (OCR) and layout-aware models offer promising results, their effectiveness strongly depends on the quality and consistency of the underlying training corpora. This work focuses on corpus construction and evaluation for historical document processing. Two experiments were conducted. In the first corpus los101 was used, a manually curated and structurally annotated subcorpus derived from historical Spanish newspapers, designed to ensure coherent ground truth under heterogeneous real-world conditions. This corpus enables systematic experimentation across OCR and document layout analysis tasks. In a second experimental phase, we apply an additional layout-focused corpus characterized by structural regularity and consistent page organization, allowing us to isolate the impact of layout homogeneity on segmentation performance. State-of-the-art OCR models and a layout detection model are evaluated as validation instruments to assess corpus adequacy rather than as primary contributions. Quantitative and qualitative analyses based on (1) relationship between annotation quality, (2) structural variability, and (3) model behavior, show that heterogeneous corpora challenge both transcription and segmentation stability, while layout-consistent data significantly improves structural detection reliability.
Thematic Landscapes of the Past: Analysing Slovene Historical Periodicals With Topic Modeling
Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Tina Munda | Darja Fiser
Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Tina Munda | Darja Fiser
This paper explores the thematic landscapes of three Slovene historical periodicals—Slovenka, Slovenec, and Slovenski narod—from the sPeriodika corpus, a comprehensive collection of Slovene press published between 1771 and 1914. Using BERTopic, we analyse the thematic profiles of these periodicals, enriched with diachronic perspectives. Our study examines the thematic commonalities and specificities of the selected periodicals, highlighting their distinct political orientations, target audiences, and the increasing nationalist polarisation in public discourse. This work contributes to digital humanities by demonstrating the potential of modern topic modelling techniques, such as BERTopic, to advance historical and cultural research.
up
Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026
Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026
Muzi Matfunjwa | Mmasibidi Setaka | Rooweither Mabuya | Menno van Zaanen
Muzi Matfunjwa | Mmasibidi Setaka | Rooweither Mabuya | Menno van Zaanen
A Morpho-Syntactically Annotated Corpus of Ògè Folk Narratives with a Focus on Nominal Structure
Priscilla Adenuga
Priscilla Adenuga
This paper presents a manually annotated morpho-syntactic corpus of Ògè, an under-resourced indigenous language spoken in Nigeria. The corpus consists of ten folk narratives (approximately 4,667 tokens) collected for the investigation of nominal structure. Annotation is expert-driven and includes token-level part-of-speech tagging together with a structured Determiner Phrase (DP) classification framework designed to capture language-specific nominal configurations. The scheme distinguishes between bare nouns and modified noun phrases, reflecting a central structural property of Ògè: noun forms remain morphologically stable across contexts, while modifiers exhibit formal and positional variation contributing to reference, specificity, and discourse prominence. The DP classification layer encodes both simple and complex nominal constructions, enabling systematic analysis of internal phrase structure. Designed as a reusable digital resource, the corpus supports morphosyntactic tagging, noun phrase boundary detection, and modeling of nominal structure in low-resource NLP settings. The annotated dataset will be made publicly available through the SADiLaR repository. This work demonstrates how descriptive linguistic analysis can inform annotation design and provides a replicable framework for developing structured resources for under-resourced African languages. Keywords: Ògè, low-resource NLP, annotated corpus, nominal structure, African languages
Extension of Linguistic Resources for South African Languages: Part-of-Speech Annotated Domain-Specific Data
Tanja Gaustad | Roald Eiselen | Cindy Arlene McKellar
Tanja Gaustad | Roald Eiselen | Cindy Arlene McKellar
In this paper, we present part-of-speech (POS) annotated domain-specific data for nine South African languages. The data has been sourced from five different domains (two academic domains, Caps and theses, two non-academic domains, news and magazines, and one fiction domain, novels), uniformly pre-processed, automatically POS-tagged and then corrected by linguistic experts. The widely used NCHLT government data sets (Eiselen and Puttkammer, 2014) have also been re-tagged with the current tag sets and manually corrected. Both the new domain-specific data sets and the re-tagged NCHL data sets have been uploaded into a public repository. To illustrate the characteristics of the domain data in comparison to government data, we include and discuss data statistics, namely type-token ration (TTR), tokens per sentence and out-of-vocabulary (OOV) rates, as well as POS tagging results with a baseline tagger trained on NCHLT data and applied to the different domains for all languages. Both the data statistics and the POS results clearly show that the domain data is significantly different to government data: For all domains and languages, the tagging accuracy decreases significantly compared to testing on in-domain government data. Also, POS results for the two domains with the highest OOV rates for all languages (Caps and novels) are much lower than for the other domains. These findings emphasise the need for more diverse data resources which in turn will aid in the development of more domain-independent language technologies.
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
Pericles Adjovi | Prasenjit Mitra | Roald Eiselen
Pericles Adjovi | Prasenjit Mitra | Roald Eiselen
Large language models (LLMs) are trained on data contributed by low-resource language communities, including curated datasets such as MasakhaNER and MAFAND-MT, yet the linguistic knowledge encoded in these models remains accessible only through commercial APIs. This paper investigates whether strategic prompting can extract usable text data from LLMs for two West African languages: Hausa (Afroasiatic, approximately 80 million speakers) and Fongbe (Niger-Congo, approximately 2 million speakers). We systematically compare six elicitation task types: creative writing, functional text, structured knowledge, dialogue, topic-switching probes, and constrained generation across two commercial LLMs (GPT-4o Mini and Gemini 2.5 Flash). Generated outputs are evaluated on linguistic accuracy, lexical diversity, domain coverage, and code-switching rates through automatic metrics assessment. Our findings reveal that elicitation strategy significantly affects output quality and that optimal strategies differ by language: Hausa benefits from volume-maximizing tasks such as functional text and dialogue, while Fongbe requires constraint-heavy prompts that enforce monolingual output. GPT-4o Mini extracts 6–41x more usable target-language words per API call than Gemini, though Gemini achieves higher language purity for Fongbe on constrained tasks. We provide a practical framework for low-resource language communities to maximize usable data extraction from LLMs and release all generated corpora and code.
Comparing Source Language Selection Strategies for Multi-Source Cross-Lingual Transfer to African Languages
Tewodros Kederalah Idris | Roald Eiselen | Prasenjit Mitra
Tewodros Kederalah Idris | Roald Eiselen | Prasenjit Mitra
Cross-lingual transfer learning enables building NLP systems for low-resource languages by leveraging data from higher-resource languages. A critical but understudied question for African languages is: which source languages should be selected for multi-source transfer? We present a systematic comparison of four source language selection strategies: random selection (baseline), genetic distance based on language family trees, geographic distance based on speaker locations, and embedding similarity from multilingual models. We evaluate these strategies on Named Entity Recognition, Part-of-Speech tagging, and sentiment analysis across five typologically diverse African target languages (Hausa, Yoruba, Swahili, Igbo, Kinyarwanda) using three multilingual models. We further investigate how the number of source languages affects transfer performance. Our experiments reveal that no single strategy dominates across tasks: geographic distance leads on sequence labeling tasks while embedding similarity is most effective for sentiment analysis, and all informed strategies consistently outperform random selection.
In this work we introduce a collection of monolingual embedding models for ten South African languages in four different architectures. To determine the quality of the embedding models we evaluate the embeddings on two sequence-labelling tasks, namely Part-of-Speech (POS) tagging and Named Entity Recognition (NER). Languages are grouped into conjunctive (isiNdebele, isiXhosa, isiZulu, and Siswati), disjunctive (Sepedi, Sesotho, Setswana, Tshivenḓa, and Xitsonga), and Afrikaans to establish the influence of training data set size and typology on the quality of the different embeddings. To isolate representation effects we train BiLSTM-CRF taggers, while keeping the architecture, data splits, and training budget fixed, varying only the input imbedding representations, namely GloVe, fastText, Flair, and RoBERTa. In our experiments, GloVe lags behind fastText, Flair, and the transformer-based models, confirming that static word-level vectors are less suited to morphologically complex, low-resource languages. Subword-aware embeddings such as fastText remain a reliable and computationally efficient baseline, while Flair is the most competitive overall across both POS tagging and NER tasks.
Improving Amharic Information Retrieval with Translative and Multi-Agent Debate Retrieval Augmented Generation
Abel Alemu Jotie | Prasenjit Mitra
Abel Alemu Jotie | Prasenjit Mitra
Retrieval-augmented generation (RAG) has been used to improve the accuracy and transparency of outputs produced by large language models (LLMs) by integrating external knowledge; however, applying RAG to low-resource languages presents unique challenges, including poor embedding representations, low retrieval quality, and semantic gaps caused by the scarcity of digital documents. In this research, we address these challenges for a selected low-resource language, Amharic, by using translative and debate-based RAG techniques to improve retrieval and reasoning. This paper outlines the key problems and research gaps in applying RAG to low-resource languages and introduces a method to enhance RAG performance for Amharic. Additionally, we introduce the first comprehensive Amharic Retrieval-Augmented Generation Benchmark (ARGB), designed to capture grammatical, cultural, and writing-system-specific constraints of the Amharic language. ARGB evaluates not only retrieval and generation quality, but also noise robustness, counterfactual robustness, negative rejection, and multi-source information integration, providing a holistic assessment of RAG capabilities. The dataset, which spans a wide range of categories, is evaluated using multiple evaluation metrics. Furthermore, we demonstrate that, using our dataset, translation-based and debate-based methods substantially improve various aspects of RAG pipeline assessment in the Amharic language. This work aims to improve the reliability, accessibility, and inclusiveness of AI systems for Amharic speakers while providing a scalable framework for other low-resource languages. Current progress on the code and benchmark can be found on this GitHub link: link.
Less can be More: Towards a Parameter-Efficient Fine-Tuning of Wav2Vec2 XLSR for Low-Resource Cape Verdean Creole ASR
Mateus Neves Andrade | Mouhamadou Lamine Ba | Idy Diop | Arlindo Oliveira da Veiga
Mateus Neves Andrade | Mouhamadou Lamine Ba | Idy Diop | Arlindo Oliveira da Veiga
Automatic Speech Recognition (ASR) for low-resource languages remains challenging due to limited annotated data and high linguistic variability. In this work, we investigate parameter-efficient fine-tuning strategies for Cape Verdean Creole ASR using the Wav2Vec 2.0 XLSR model. We evaluate the impact of structured layer freezing on model performance, training stability, and computational efficiency. Experiments conducted on a newly curated Santiago-dialect dataset show that full fine-tuning achieves the best absolute performance (WER 0.212, CER 0.120). However, several freezing configurations achieve comparable recognition performance while substantially reducing the number of trainable parameters and exhibiting more stable convergence. These results highlight a trade-off between adaptability and efficiency, showing that selective freezing can serve as an effective regularization strategy in low-resource settings. This work provides practical insights into parameter-efficient adaptation for under-resourced Creole languages.
From Script to Semantics: Prompting Strategies for African NLI
Anuj Tiwari | Terry Oko-odion | Hannah Nwokocha
Anuj Tiwari | Terry Oko-odion | Hannah Nwokocha
Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains underexplored especially under pure prompting without fine-tuning. We present a systematic study of prompting strategies for Natural Language Inference (NLI) in Swahili, Yoruba, and Hausa using the AfriXNLI benchmark. We evaluate five prompting strategies Baseline (zero-shot), Script-Aware, Language Specific, Contrastive, and Native-Label Self-Translation (NL-STP) across two mid-sized open weight models (Llama3.2-3B and Gemma3-4B). To isolate the effect of prompt design, the effect of few-shot examples and Chain-of-Thought reasoning is eliminated in our study. We find a significant difference in performance of class wise across strategies with highly neutral class collapse and high prediction skew in some configurations. Contrastive prompting proves to be the most reliable and steadily improving strategy over language and model and has better balance of class behavior and balance of overall accuracy gains. Notably, well-constructed prompts are sufficient to beat more powerful baselines that are provided with few-shot prompts and Chain-of-Thought prompts. We have found that prompt formulation is essential to multilingual NLI with low-resource languages and that language aware decision structuring can be used to meaningfully enhance robustness in resource challenged settings.
HaYo: Repurposing DiaSafety Dataset for Dialogue Safety Evaluation in Hausa and Yoruba
Tunde Oluwaseyi Ajayi | Bolade Deborah Ashaolu | Falalu Ibrahim Lawan | Daud Olamide Abolade | Amina Imam Abubakar | Oluwatosin Ayomide Akinrinde | Murja Sani Gadanya | Omodolapo Dorcas Ashaolu | Abubakar Khalid Auwal | Adewumi Awujoola | Shamsuddeen Umaru Adamu | Israel Olawole Ashaolu | Mihael Arcan | Paul Buitelaar
Tunde Oluwaseyi Ajayi | Bolade Deborah Ashaolu | Falalu Ibrahim Lawan | Daud Olamide Abolade | Amina Imam Abubakar | Oluwatosin Ayomide Akinrinde | Murja Sani Gadanya | Omodolapo Dorcas Ashaolu | Abubakar Khalid Auwal | Adewumi Awujoola | Shamsuddeen Umaru Adamu | Israel Olawole Ashaolu | Mihael Arcan | Paul Buitelaar
Research efforts aimed at detecting unsafe dialogues have resulted in creation of benchmark datasets and models for evaluation. The benchmarks mostly exist in English and other high resourced languages. In order to address the challenge of unavailability of dialogue safety evaluation dataset in Hausa and Yorùbá, we repurporse DiaSafety dataset to develop HaYo dataset, by providing contextualised human annotation of dialogues in DiaSafety. We provide dialogues in Hausa and Yorùbá, obtained by human translation of dialogues in the DiaSafety dataset, to raters who are native speakers. The dialogues are annotated as Unsafe or Safe. We evaluate seven models with moderation, conversational or multilingual capabilities in terms of F1 Score. Using McNemar test, we observe that the predictions of GPT-4.1 and Gemma-3-12b-it on HaYo are statistically significant at p < 0.05. In our evaluation with instructions in English, we observe lower F1 scores in six out of the seven models, comparing the performance on DiaSafety and HaYo labels. The model predictions were inconsistent with the labels in the HaYo dataset when instructions and dialogues were provided in Hausa and Yorùbá. Compared to providing instructions in English, the issues range from responses in unspecified languages to underperformance in terms of F1 score. We plan to release the HaYo dataset to the public to promote dialogue safety research, especially in under-resourced languages.
Reclaiming African Voices: Surveying Indigenous Writing Systems for Inclusive NLP
Mamady Traore | Ngoc Tan Le | Fatiha Sadat
Mamady Traore | Ngoc Tan Le | Fatiha Sadat
Multilingual NLP has expanded rapidly through large-scale pretraining and cross-lingual transfer, yet this progress remains structurally uneven across writing systems. This survey reframes multilingual NLP around scripts rather than languages, arguing that writing systems constitute an under-theorized axis of computational inequality. Focusing on African scripts — Indigenous (Vai, Ge’ez, Tifinagh), modern (ADLaM, N’Ko), and adapted Arabic-based (Ajami)—we analyze how script properties interact with digital infrastructure, tokenization, and downstream task performance. We organize the literature across four analytical layers: infrastructural (Unicode and input systems), representational (segmentation efficiency and vocabulary allocation), functional (task-level disparities), and epistemic (evaluation bias and the “low-resource” framing). Synthesizing evidence from 47 studies, we show that performance gaps across scripts arise primarily from engineering design choices rather than intrinsic linguistic complexity. We conclude by outlining a research agenda for native multiscript foundation models, including script-aware scaling laws, tokenizer equity metrics, and evaluation reform. We argue that multiscript equity is not a peripheral concern but a structural precondition for genuine multilingual inclusion
Getting Close to Cloze: Investigating Language Model and Human Cloze-test Performance in Afrikaans
Susan Lotz | Rik van Noord | Gertjan van Noord
Susan Lotz | Rik van Noord | Gertjan van Noord
Models that can estimate the readability of a given text automatically are a valuable resource for any language. There are however many languages for which such models do not work well or simply do not exist yet. In this paper, we lay the groundwork for developing a high-quality application for Afrikaans by having encoder-only language models (LMs) complete a set of cloze tests already completed by humans. Strong correlation between the cloze-test performance of humans and an LM is an indication that the LM could possibly serve as a proxy for human participants. We show that the output of models trained on (some) Afrikaans correlates reasonably well with human answers, underscoring the potential of LMs to be used in automatic readability assessment. A more fine-grained analysis confirms that the correlation is not driven by only a few strongly correlating word classes, but spread relatively evenly over all word classes. We further establish by means of a manual evaluation that, in cases where the cloze-test performance of humans and an LM correlate strongly because both were wrong, LM answers tend to be further off than human answers for the same cloze items. It is noteworthy that the model with the best correlation, afRoBERTa (r=0.62; Spearman’s ρ=0.62), is neither the most accurate nor the largest model, but a model trained on Afrikaans only, showing the benefit of small, monolingual LMs compared to large, multilingual models for specific purposes.
The Hundzula Retreat-Based Infrastructure Model for African Natural Language Processing
Johannes Sibeko | Seani Rananga | Neo N. Putini | Dan Masethe
Johannes Sibeko | Seani Rananga | Neo N. Putini | Dan Masethe
The development of Natural Language Processing (NLP) resources for African indigenous languages remains constrained by limited data availability, fragmented expertise, and a lack of sustainable, locally grounded infrastructures for enabling language research. While much existing work focuses on producing discrete resources such as corpora or lexicons, less attention has been paid to the social, institutional, and methodological conditions that enable such resources to be created, maintained, and sustained. This paper presents the Hundzula Retreat for NLP and Linguistics as a retreat-based resource infrastructure model that addresses these constraints. We conceptualise Hundzula not as a once-off event, but as a structured, upstream research infrastructure that facilitates human capacity development, interdisciplinary collaboration between linguistics and NLP, ethical data practices, and the early-stage incubation of language resources for African indigenous languages. Drawing on evidence from multiple iterations of the retreat, we describe the design principles, workflows, and governance mechanisms that support resource development, including training pipelines, human-in-the-loop methodologies, and collaborative project formation. Rather than focusing on already formalised outputs, the paper foregrounds the infrastructural conditions that make such outputs possible within under-resourced contexts. In doing so, the paper shifts attention from outputs to the enabling ecosystems required for their production. We argue that retreat-based infrastructures constitute an essential but under-recognised category of language resources and demonstrate how the Hundzula model can be adapted and replicated in other low-resourced language contexts. The paper contributes a transferable framework for sustainable NLP resource development grounded in African linguistic realities.
Creative Commons (CC) licenses are prevalent in African natural language processing (NLP) corpus releases, but their compatibility implications are rarely examined systematically. CC-BY-SA and CC-BY-NC cannot be combined in a single published dataset; a NoDerivs (ND) clause prohibits redistribution of tokenised or annotated derivatives. This paper presents an empirical audit of license provenance across more than twenty corpus families used in African NLP, applying established compatibility rules to three case-study languages: Kituba/Munukutuba, Zarma, and Moore. Four failure modes are documented with primary-source evidence: outright prohibition (JW300, removed from OPUS after a legal audit confirmed a Terms of Service violation); composite license misrepresentation (WAXAL, whose CC-BY 4.0 claim is contradicted by its HuggingFace dataset card); a ND restriction not reflected in the CC-BY label (Tanzil); and data persistence failure (the Congolese Radio Corpus, where 402 of 405 source URLs are no longer accessible). A due diligence checklist and a survey of legally compliant enrichment opportunities conclude the paper. We argue that lawful data use is an ethical baseline: for African language communities with limited institutional recourse, license violations are not only legal risks but ethical failures that compound existing power asymmetries.
up
Proceedings of the Sixth Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with various forms of cognitive/psychiatric/developmental impairments in cooperation with the MENTAL.ai consortium
Proceedings of the Sixth Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with various forms of cognitive/psychiatric/developmental impairments in cooperation with the MENTAL.ai consortium
Dimitrios Kokkinakis | Charalambos Themistocleous | Gaël Dias | Kathleen C. Fraser | Fredrik Öhman | Sebastião Pais
Dimitrios Kokkinakis | Charalambos Themistocleous | Gaël Dias | Kathleen C. Fraser | Fredrik Öhman | Sebastião Pais
Multilingual Cognitive Impairment Detection in the Era of Foundation Models
Damar Hoogland | Boshko Koloski | Jaya Caporusso | Tine Kolenik | Senja Pollak | Christina Manouilidou | Matthew Purver
Damar Hoogland | Boshko Koloski | Jaya Caporusso | Tine Kolenik | Senja Pollak | Christina Manouilidou | Matthew Purver
We evaluate cognitive impairment (CI) classification from transcripts of speech in English, Slovene, and Korean. We compare zero-shot large language models (LLMs) used as direct classifiers under three input settings—transcript-only, linguistic-features-only, and combined—with supervised tabular approaches trained under a leave-one-out protocol. The tabular models operate on engineered linguistic features, transcript embeddings, and early or late fusion of both modalities. Across languages, zero-shot LLMs provide competitive no-training baselines, but supervised tabular models generally perform better, particularly when engineered linguistic features are included and combined with embeddings. Few-shot experiments focusing on embeddings indicate that the value of limited supervision is language-dependent, with some languages benefiting substantially from additional labelled examples while others remain constrained without richer feature representations. Overall, the results suggest that, in small-data CI detection, structured linguistic signals and simple fusion-based classifiers remain strong and reliable signals.
The Icelandic Language Biobank: Data Collection through a Clinical Analysis Platform
Iris Nowenstein | Naizeth Núñez Macías | Gunnar Thor Örnólfsson | Stefán Ólafsson | Bryndís Bergþórsdóttir | Iðunn Kristínardóttir | Hinrik Hafsteinsson
Iris Nowenstein | Naizeth Núñez Macías | Gunnar Thor Örnólfsson | Stefán Ólafsson | Bryndís Bergþórsdóttir | Iðunn Kristínardóttir | Hinrik Hafsteinsson
Recent work on clinical applications of language technology shows considerable potential for people with speech and language symptoms and disorders, including for the diagnosis and monitoring of diseases and disorders as well as the development of novel communication aids. This has resulted in a variety of digital health tools becoming accessible, including personalized automatic speech recognition for disordered speech and the monitoring of disease progression in neurodegeneration through language samples. Currently, these tools are almost exclusively accessible to speakers of high-resource languages. A major hurdle for small, lower-resourced language communities in this context is the creation of clinical language corpora. We describe ongoing efforts to build the necessary infrastructure for clinical speech and language data collection in Iceland through the Icelandic Language Biobank, a resource that leverages collaboration with clinicians and robust linguistically-informed data collection against data scarcity.
Disfluencies and ASR Performance on Swedish Spontaneous Speech from the ‘Trip to Stockholm’ Discourse Narrative Task
Dimitrios Kokkinakis | Herbert Lange | Ricardo Muñoz Sánchez
Dimitrios Kokkinakis | Herbert Lange | Ricardo Muñoz Sánchez
Automatic Speech Recognition (ASR) offers a scalable and cost-efficient alternative to manual transcription and is becoming increasingly relevant in clinical contexts, particularly for the detection of cognitive decline and mental health assessment. However, current ASR-systems still struggle with spontaneous speech, particularly when processing disfluencies, pauses, and speaker variability that often carry diagnostic value. This study evaluates state-of-the-art open ASR models targeting Swedish using recordings from the “Trip to Stockholm” discourse narrative task which elicits ecologically valid, cognitively demanding speech. Recognition quality is assessed using various metrics, alongside an analysis of linguistic and technical sources of error focused on disfluencies. Our findings show that disfluency-related phenomena degrade recognition performance. Possible post-processing strategies can improve specific error patterns emerging for filled pauses, word repetitions, and self-corrections. The results illustrate both the advances and ongoing limitations of ASR for spontaneous Swedish speech, emphasizing the need for models explicitly trained, or fine-tuned, on disfluent data to ensure robustness in clinical and research applications.
ALBA: An Automated Framework for Benchmarking Clinical Language Biomarkers against Standardized Corpora
Charalambos Themistocleous | Brielle C. Stark
Charalambos Themistocleous | Brielle C. Stark
Patients with diverse neurocognitive conditions frequently exhibit measurable language deficits that serve as biomarkers for differential diagnosis and therapy decision making. Discourse analysis can offer reliable ecological measures of human communication, yet manual discourse analysis is cumbersome. Recent advances in automated analysis software provide quick and easy extraction of raw language metrics in the clinic. Nevertheless, transforming these measures into actionable clinical insights remains a significant challenge. The aim of this paper is to present the Automated Language Biomarker Application (ALBA), an integrated framework developed within the Open Brain AI ecosystem to bridge the gap between feature extraction and clinical interpretation. ALBA provides clinicians with a robust statistical infrastructure to benchmark individual patient measures against standardized, large-scale clinical corpora. By utilizing a shared elicitation and processing pipeline, the application ensures that user-provided data are directly comparable to population norms for conditions including Aphasia, Mild Cognitive Impairment (MCI), Dementia, and other neurological conditions. The system implements adaptive statistical logic, employing one-sample t-tests and robust non-parametric alternatives to provide real-time significance testing and dynamic visualizations (box, bar, and violin plots). By automating the comparison of “Language Signatures” against healthy controls and specific clinical phenotypes, ALBA facilitates rapid, evidence-based decision-making in both research and rehabilitation contexts.
Cognitive decline refers to the gradual loss of thinking abilities, including memory, attention, reasoning, and problem-solving. It can be a normal part of ageing or a symptom of conditions like dementia or Alzheimer’s disease when it significantly interferes with daily life. Early diagnosis is crucial, as timely intervention can slow progression and improve quality of life. Emerging approaches such as Digital Linguistic Biomarkers, subtle changes in speech and language patterns captured through digital tools, offer a promising, non-invasive way to detect early signs of cognitive decline before more obvious symptoms appear and perform massive population screening. In this position paper, we contend that the prevailing paradigm for the automatic detection of cognitive decline, primarily relying on classifiers that analyse subjects’ linguistic productions at a single point in time, is not the most effective approach. Instead, we advocate for a paradigm shift toward longitudinal analyses that track linguistic patterns over decades. To support this perspective, we present an experiment in which we compile and analyse a long-term corpus of spontaneous speech productions from well-known individuals, enabling insights into cognitive changes across extended time spans.
Benchmarking NLP-supported Language Sample Analysis for Swiss Children’s Speech
Anja Ryser | Yingqiang Gao | Sarah Ebling
Anja Ryser | Yingqiang Gao | Sarah Ebling
Language sample analysis (LSA) is a process that complements standardized psychometric tests for diagnosing, for example, developmental language disorder (DLD) in children. However, its labor-intensive nature has limited its use in speech-language pathology practice. We introduce an approach that leverages natural language processing (NLP) methods not based on commercial large language models (LLMs) applied to transcribed speech data from 119 children in the German-speaking part of Switzerland with typical and atypical language development. This preliminary study aims to identify optimal practices that support speech-language pathologists in diagnosing DLD more efficiently with active involvement of human specialists. Preliminary findings underscore the potential of integrating locally deployed NLP methods into the process of semi-automatic LSA.
Resource-Efficient LLMs for Depression Symptoms Screening: Performance and Limitations in Zero Shot Setting
Muhammad Rizwan | Jure Demšar
Muhammad Rizwan | Jure Demšar
Depression is the leading cause of global disability and early detection is crucial for effective intervention. Recent advances in large language models (LLMs) offer potential for analyzing text to identify depression symptoms. This work investigates the zero-shot capability of LLMs to recognize nine DSM5 depression symptoms from short-text inputs. We evaluated eight open LLMs with model sizes ranging from 1.5B to 14B parameters using a clinically annotated dataset and assessed both overall agreement and symptom-level performance. Results indicate that while smaller models exhibit limited clinical accuracy, the Qwen 2.5-7B model achieves substantial performance with a Cohen’s Kappa of 0.603 and a Macro F1 score of 0.648. Notably, a performance plateau between the 7B and 14B Qwen variants suggests that model scaling alone does not guarantee improved symptom-level classification, establishing Qwen 2.5-7B as a resource-efficient model. Further analysis of the best-performing model revealed strengths in identifying salient symptoms like suicidal thoughts, but limitations in recognizing core symptoms such as depressed mood and anhedonia. Misclassification analysis reveals that the model frequently misclassifies posts expressing ’depressed mood’ as ’no symptom’ or vice versa, often overlooking indicators of irritability or social withdrawal. These findings suggest that resource-efficient LLMs can support preliminary symptom screening in zero shot settings, but there is risk of overlooking clinically important symptoms without fine-tuning.
CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis
Jinyuan Xu | Tian Lan | Xintao Yu | Xue He | Hezhi Zhang | Ying Wang | Mathieu Valette | Pierre Magistry | Lei Li
Jinyuan Xu | Tian Lan | Xintao Yu | Xue He | Hezhi Zhang | Ying Wang | Mathieu Valette | Pierre Magistry | Lei Li
Depression is a pressing global public health issue, yet publicly available Chinese-language resources for depression risk detection remain scarce and largely focus on binary classification. To address this limitation, we release CNSocialDepress, a benchmark dataset for depression risk detection on Chinese social media. The dataset contains 44,178 posts from 233 users; psychological experts annotated 10,306 depression-related segments. CNSocialDepress provides binary risk labels along with structured, multidimensional psychological attributes, enabling interpretable and fine-grained analyses of depressive signals. Experimental results demonstrate the dataset’s utility across a range of NLP tasks, including structured psychological profiling and fine-tuning large language models for depression detection. Comprehensive evaluations highlight the dataset’s effectiveness and practical value for depression risk identification and psychological analysis, thereby providing insights for mental health applications tailored to Chinese-speaking populations.
Depression Detection in Modern Greek
Vivian Stamou | George Mikros | George Markopoulos | Spyridoula Varlokosta
Vivian Stamou | George Mikros | George Markopoulos | Spyridoula Varlokosta
Despite advancements in NLP-based mental health screening, research remains predominantly English-centric, leaving under-resourced languages insufficiently explored. This study investigates depression detection in Modern Greek social media through a series of experiments. We benchmark traditional machine learning (ML) models against transformer architectures (GreekBERT, GreekSocialBERT, mBERT, and XLM-R) under two settings: a topic-oriented control corpus and a high-similarity stress-test contrasting a gold case of a depressed user with a matched control. Transformer models consistently outperform ML models (F1 = 0.95) but offer limited interpretability. To address this limitation, we incorporate LIWC-derived psycholinguistic features with SHAP explanations to examine model behavior in relation to established linguistic markers. The analysis reveals linguistic patterns consistent with depressive symptoms, such as reduced work-related engagement, social withdrawal, and the motivational deficits characteristically linked to anhedonia in clinical literature. Overall, the results provide a baseline for depression detection in Modern Greek and underscore the importance of grounding automated screening in clinically interpretable evidence.
Profiling Psychopathic Behavior Using Machine Learning
Avi Treistman | Tehilla David | Sivan Levi | Dror Mughaz
Avi Treistman | Tehilla David | Sivan Levi | Dror Mughaz
Psychopathy is a complex personality disorder characterized by persistent deficits in empathy and manipulative behavior. Traditional diagnostic methods often rely on subjective clinical assessments, which are susceptible to deception. This research proposes an objective, non-invasive computational framework for profiling psychopathic traits using Natural Language Processing (NLP) and Machine Learning. We developed a systematic pipeline utilizing transcribed interviews from confirmed criminal psychopaths and a balanced control group. To address data sparsity and noise, we employed the Dynamic Variance Thresholding (DyVaT) algorithm to construct a semantically dense vocabulary of over 1,300 features. The methodology integrates advanced preprocessing, TF-IDF vectorization, and synonym-based data augmentation to enhance model generalization. Among the evaluated classifiers, a Linear Support Vector Machine (SVM) achieved the highest performance, with an accuracy of 0.8081 and an F1-score of 0.7957. Our findings demonstrate the efficacy of linguistic biomarkers and feature importance analysis in distinguishing psychopathic speech patterns. This study provides a scalable methodology for early screening and diagnostics, with significant implications for forensic psychology, security, and ethical AI deployment in mental health.
Developing Annotation Guidelines for CSAM Prevention Interventions: Psychosocial Risk and Protective Factors Grounded in Research and Clinical Practice
Vera Czehmann | Christine Hovhannisyan | Lena Elisabeth Hoffmann | Paula Busch | Ibrahim Baroud | Sebastian Möller | Roland Roller | Hannes Gieseler | Lisa Raithel
Vera Czehmann | Christine Hovhannisyan | Lena Elisabeth Hoffmann | Paula Busch | Ibrahim Baroud | Sebastian Möller | Roland Roller | Hannes Gieseler | Lisa Raithel
This work discusses sexual offending, specifically child sexual abuse material (CSAM), in the context of prevention. We introduce a domain-specific, span-level annotation scheme and guidelines to identify psychosocial risk and protective factors in therapist-led, anonymous chat interventions with voluntarily help-seeking individuals concerned about their pedophilic interests and the risk of CSAM use. The scheme is grounded in previous research and clinical experience, and intended for within-intervention guidance and longitudinal tracking, rather than actuarial risk scoring. Annotating a pilot subset (8 clients, 31 sessions), inter-annotator agreement was moderate but improved after calibration, which is consistent with the linguistic and clinical ambivalence present in the data. We track a session-wise Protective Ratio, i.e., the share of protective factors among all coded factors, and examine its behaviour over time during the intervention and around self-reported relapse within clients. In exploratory automation, LLM-based span extraction outperforms BERT baselines but overall performance remains limited by small data and mixed-evidence spans. While complete anonymisation of the corpus is in progress, we release the label scheme, guidelines, and non-sensitive artefacts of our analyses.
Automatic Detection of Direct and Self-Repetitions in Naturalistic Speech Recordings of French- and Dutch-Speaking Autistic Children
Federica Beccaria | Marie Kolenberg | Pierre Labendzki | Inge Zink | Mikhail Kissine
Federica Beccaria | Marie Kolenberg | Pierre Labendzki | Inge Zink | Mikhail Kissine
This study investigates the use of cosine similarity measures across syntactic, lexical, and semantic vector repre- sentations to detect repetitions in the spontaneous speech of autistic children. It focuses on direct repetitions (i.e., immediate verbatim repetitions of linguistic output produced by another individual) and self-repetitions (i.e., within-speaker recurrence). The performance of similarity-based methods is then compared with state-of-the-art black-box classification models based on BERT, trained on the same data. Using spontaneous speech data from French- and Dutch- speaking autistic children, the results show that lexical and semantic similarity provide reliable cues for identifying self-repetitions, achieving high precision and recall, with F1-scores exceeding 83%, comparable to those obtained by BERT-based models. In contrast, direct repetitions are more difficult to detect using similarity-based approaches, with BERT models clearly outperforming them and reaching F1-scores above 73%. Across all conditions, syntactic similarity consistently underperforms relative to lexical and semantic measures. These findings highlight the strengths and limitations of similarity-based approaches and suggest directions for future research, particularly in improving the detection of direct repetitions and assessing the cross-linguistic generalizability of these methods.
up
Proceedings of the Joint Workshop on Readability and Text Simplification (READIxTSAR) @ LREC 2026
Proceedings of the Joint Workshop on Readability and Text Simplification (READIxTSAR) @ LREC 2026
Matthew Shardlow | Thomas François | Raquel Amaro | Jorge Baptista | Rémi Cardon | Eugénio Ribeiro | Horacio Saggion | Regina Stodden | Amalia Todirascu | Rodrigo Wilkens
Matthew Shardlow | Thomas François | Raquel Amaro | Jorge Baptista | Rémi Cardon | Eugénio Ribeiro | Horacio Saggion | Regina Stodden | Amalia Todirascu | Rodrigo Wilkens
Revisiting German Complex Word Identification: Contextualized LLMs and Feature Injection
Thorben Schomacker | Seid Muhie Yimam | Chris Biemann | Marina Tropmann-Frick
Thorben Schomacker | Seid Muhie Yimam | Chris Biemann | Marina Tropmann-Frick
Complex word identification (CWI) is essential in text simplification, yet work on German CWI remains comparatively limited. To address this gap, we investigate the capabilities of three state-of-the-art LLMs and compare them to previously proposed baseline systems. We fine-tune the LLMs in three setups: (i) using the target expression only, (ii) using the target expression together with its sentence-level context, and (iii) using the context and injection of classical machine learning features. Our results show that while pretrained-only LLMs fall short, fine-tuned LLMs set new benchmarks for both binary and probabilistic CWI. In addition, embedding the target in its context sentence improves performance, whereas feature injection has no clearly measurable effect. All models in this paper are trained on the probabilistic CWI task and additionally evaluated on the binary task; thus, we publish a single model that supports both evaluation views We released all accompanying resources (https://github.com/tschomacker/german-cwi-llm) and model checkpoints (https://huggingface.co/collections/tschomacker/german-cwi-llm).
Book Complexity Level Assignment in French and Portuguese
Jorge Baptista | David Antunes | Wafa Aissa | Julien Zakhia Doueihi | Hanh Trang Tran Pham | Eugénio Ribeiro | Thomas François | Raquel Amaro
Jorge Baptista | David Antunes | Wafa Aissa | Julien Zakhia Doueihi | Hanh Trang Tran Pham | Eugénio Ribeiro | Thomas François | Raquel Amaro
Selecting reading materials that are appropriate for adults with low literacy skills remains a central challenge in Adult Learning contexts. This challenge becomes particularly acute when the unit of analysis is not a short passage but a full book or long-form text, where internal heterogeneity in lexical, syntactic, and discourse-level properties make global readability estimation non-trivial. In practice, librarians, educators, and publishers often need to make decisions about the suitability of books for specific learner populations without access to complete texts, relying instead on partial excerpts or limited samples, and personal intuition.
Taming CATS: Controllable Automatic Text Simplification through Instruction Fine-Tuning with Control Tokens
Hanna Hubarava | Yingqiang Gao
Hanna Hubarava | Yingqiang Gao
Controllable Automatic Text Simplification (CATS) produces user-tailored outputs, yet controllability is often treated as a decoding problem and evaluated with metrics that are not reflective to the measure of control. We observe that controllability in ATS is significantly constrained by data and evaluation. To this end, we introduce a domain-agnostic CATS framework based on instruction fine-tuning with discrete control tokens, steering open-source models to target readability levels and compression rates. Across three model families with different model sizes (Llama, Mistral, Qwen; 1-14B) and four domains (medicine, public administration, news, encyclopedic text), we find that smaller models (1-3B) can be competitive, but reliable controllability strongly depends on whether the training data encodes sufficient variation in the target attribute. Readability control (FKGL, ARI, Dale-Chall) is learned consistently, whereas compression control underperforms due to limited signal variability in the existing corpora. We further show that standard simplification and similarity metrics are insufficient for measuring control, motivating error-based measures for target-output alignment. Finally, our sampling and stratification experiments demonstrate that naive splits can introduce distributional mismatch that undermines both training and evaluation.
PLABA-EVAL: A Multi-Dimensional, In-Context Sentence Readability Dataset for Medical Text
Kexin Bian | Su-Youn Yoon | Mamoru Komachi
Kexin Bian | Su-Youn Yoon | Mamoru Komachi
We present an in-context framework for assessing readability that separates reading difficulty into multiple subjective dimensions. Participants read biomedical abstracts with full-document access and provide sentence-level ratings of Processing Ease and Perceived Understanding, followed by an open-book multiple-choice comprehension check. Using this protocol, we release PLABA-EVAL, a dataset of 78 biomedical abstracts and expert plain-language adaptations (609 sentences), annotated by three independent raters per document. Analyses show that Ease and Understanding are strongly related but not interchangeable, and that perceived understanding aligns more closely with open-book comprehension performance. We provide baseline linguistic analyses for both dimensions, illustrating how the dataset supports work on readability, simplification, and sentence-level difficulty modeling.
Automatic Extraction of Textual and Phonemic Complexity for French Cued Speech
Magali Norré | Brigitte Bigi | Núria Gala | Ludivine Javourey Drevet | Thomas François
Magali Norré | Brigitte Bigi | Núria Gala | Ludivine Javourey Drevet | Thomas François
This article presents the results of an analysis of a written corpus with the view of automatically generating it in French Cued Speech (CS). CS is a communication system developed for people with hearing impairment to complement speech reading at the phonetic level using hands. This visual communication mode uses handshapes in different positions near the face in combination with the mouthshape (called ’cues’ or ’keys’) to make the phonemes of spoken language look different from each other. Despite many studies demonstrating its benefits, there are few resources available for learning and practicing it, especially in French. As part of a wider project aimed at creating an online learning platform with automatically generated videos using an augmented reality system displaying a virtual coding, we propose to identify, extract, and analyze 41 textual and phonemic features that might be more complex to (de)code in French CS. For the automatic extraction of complexity, several tools are used: FABRA for readability, SPPAS for phonetization and CS key generation. The results show some strong correlations between readability features, few between phonemic variables, and few between the two types. An initial model is proposed for selecting texts to be recorded for learning French CS.
Can LLMs Control Readability? A Multi-Dimensional Evaluation Framework for CEFR-Controlled Arabic Generation
Nour Rabih | Chatrine Qwaider | Ted Briscoe
Nour Rabih | Chatrine Qwaider | Ted Briscoe
While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a multi-dimensional evaluation framework for Common European Framework of Reference for Language (CEFR)-controlled Arabic text generation, assessing whether instruction-following LLMs can serve as reliable generators for adaptive language learning. Our framework integrates controlled prompting, automatic readability prediction using a validated Taha-19 model, lexical constraint validation, and syntactic complexity profiling. Results show that structured prompting substantially improves CEFR alignment. In particular, CEFR-guided prompting with lexical constraints achieves the highest conformity to reference linguistic profiles (0.91 cosine similarity) and near-perfect agreement with predicted readability levels (0.99), while unconstrained prompting exhibits weak control. These findings establish an empirical foundation for integrating readability-aware Arabic text generation into adaptive educational systems.
Lexical Conditioning of Model’s Distribution through Uncertainty-gated Soft-Mixing of Probabilities
Michele Papucci | Giulia Venturi | Felice Dell’Orletta
Michele Papucci | Giulia Venturi | Felice Dell’Orletta
We present Uncertainty-Gated Lexical Decoding (UGLD), a decoding-time framework for fine-grained lexical control in Large Language Models (LLMs) that explicitly addresses the trade-off between controllability and fluency. UGLD adaptively scales intervention through an entropy-based gating mechanism derived from the model’s predictive distribution, activating control when uncertainty is high and limiting interference when predictions are confident. The method supports both promotion toward and against predefined vocabularies. We evaluate UGLD in Italian on two open-weight LLMs (ANITA 8B and Qwen 3 4B) across paraphrasing and free-text generation settings, considering Simple Vocabulary Conditioning and Jargon Reduction scenarios. Automatic evaluation shows consistent improvements in lexical coverage over standard decoding strategies, while human evaluation confirms that fluency is preserved under controlled intervention.
A Comparative Study of Multilingual Fine-tuning and Prompting for Automatic Text Readability Classification in Galician
Sandra Rodríguez Rey | Marcos Garcia
Sandra Rodríguez Rey | Marcos Garcia
Despite advancements in automatic readability assessment, low-resource languages such as Galician remain under-explored. This study addresses this gap by presenting a comparative study of readability assessment techniques in Galician, including fine-tuning of encoder models as well as prompting strategies using large generative models. Due to the scarcity of native Galician resources, neural machine translation was employed to generate synthetic Galician data. The analysis begins with BERT-based monolingual models trained on the synthetic data. For multilingual models, the impact of using original versus translated data was compared in order to assess the effects of translation-based augmentation. Finally, several LLMs were evaluated using zero-shot and few-shot prompting methods. The results indicate that generative models are not yet competitive with encoder models tuned for text classification in Galician, and that data generated through machine translation improves the performance of monolingual models but has little effect on multilingual models.
In this paper, we investigate the impact of increasing context lengths (one to five paragraphs) on plan-following accuracy in plan-guided text simplification. Plan-guided models simplify text according to sentence-level operation labels such as copy, rephrase, split, and delete. Previous work fine-tunes BART with target reading-level and sentence-level operation tokens to perform this task. We find that BART’s plan-following accuracy on Newsela-auto drops significantly as context increases from one to five paragraphs. This means that the model becomes less reliable with longer contexts, and the quality of its outputs decreases. To address this, we propose replacing the fine-tuned BART models with a prompting-based approach using instruction-tuned Qwen models. We find that this approach not only maintains robust plan-following across all context lengths, but even at the longest context length still exceeds BART’s performance at the shortest. We further provide ablations on model size and model family, showing that a minimum model capacity is required for the approach to work and that it transfers across LLM families.
LLM-Generated Stories for Students with Significant Cognitive Disabilities: Promise, Gaps, and Evaluation Framework
Pragati Maheshwary | Ananya Ganesh | Shamya Karumbaiah
Pragati Maheshwary | Ananya Ganesh | Shamya Karumbaiah
Students with significant cognitive disabilities (SCD) require specially designed accessible stories for reading comprehension assessments, yet creating such content is labor-intensive and difficult to scale. This preliminary study investigates whether large language models (LLMs) can generate short accessible stories for alternate assessment system. Using an 8-fold cross-validation design, we generated 120 stories with GPT-4o via one-shot prompting with human-written exemplars and evaluated them against a test set comprising 7 expert-human written stories as baselines across three dimensions: simplicity, fluency & coherence, and thematic adherence. Cross-validation results show that generated stories meet surface-level simplicity targets, with approximately two-thirds falling within the human baseline range for readability metrics. However, generated stories exhibited a systematic coherence gap where only 5% fell within the human range for adjacent sentence similarity, a pattern consistent across all folds. Thematic adherence was moderate, with adequate diversity across stories. These findings suggest LLMs can serve as a drafting tool within accessible content generation pipelines, but human expert review remains essential to ensure coherence, testability, and alignment with quality standards required for high-stakes alternate assessments.
Evaluating Transformer Model Family Representations Through Automated Essay Scoring
Akchay Ozten | Rodrigo Wilkens
Akchay Ozten | Rodrigo Wilkens
Large Language Models have become central to Automated Essay Scoring (AES), typically through fine-tuned transformer encoders or prompt-based applications of decoder models. However, the representational capacity of decoder models as frozen embedding extractors remains largely unexplored. In this paper, we present a controlled comparison between encoder and decoder transformer embeddings for prompt-agnostic AES. Using regression models, we evaluate frozen representations across two English datasets. We analyzed scaling effects and the impact of integrating explicit linguistic features in hybrid configurations. Our results show that decoder embeddings consistently outperform encoder embeddings in embedding-only settings, with gains generalizing across holistic essay scoring and proficiency prediction. Scaling effects are modest, and hybrid models that combine contextual embeddings with linguistic features yield further improvements. Notably, frozen decoder embeddings achieve performance competitive with a fine-tuned BERT. These findings highlight the importance of representation-level properties in essay scoring.
Proficiency-Controlled Text Simplification in European Portuguese: A Preliminary Study using Prompting Approaches
Eugénio Ribeiro | David Antunes | Nuno Mamede | Jorge Baptista
Eugénio Ribeiro | David Antunes | Nuno Mamede | Jorge Baptista
This paper presents a preliminary study on proficiency-controlled text simplification in European Portuguese using multiple prompting strategies. We focus on the iRead4Skills dataset, which defines four complexity levels targeted at adult native speakers with low literacy. Specifically, we simplify 40 texts from the highest complexity level into three easier levels (plain, easy, and very easy), corresponding approximately to Common European Framework of Reference for Languages (CEFR) levels B1, A2, and A1. We evaluate zero-shot and few-shot prompting configurations, exploring the impact of CEFR anchoring, explicit meaning-preservation instructions, and example-based guidance. Automatic evaluation relies on a fine-tuned proficiency classifier and semantic similarity metrics, including BERTScore and document embeddings. The results show that while exact target-level accuracy remains below 40%, target-or-below accuracy reaches up to 61.39%, indicating that the model generally simplifies texts but struggles to consistently match precise proficiency targets. Human evaluation confirms the overall trends observed automatically, while highlighting the subjectivity inherent to proficiency assessment and meaning preservation. Our findings suggest that prompt engineering alone is insufficient for robust proficiency control in European Portuguese, motivating future work on model adaptation and improved evaluation protocols.
Automatic Text Simplification for French Medical Documents with LLMs: The Role of Target Audience and Genre
Rémi Cardon | A. Seza Doğruöz
Rémi Cardon | A. Seza Doğruöz
Medical information is hard for non-specialists to understand, despite its importance for treatment success. Automatic text simplification (ATS) rewrites complex documents into simpler versions, with effectiveness measured through ATS evaluation metrics and readability metrics. A key challenge in ATS is calibrating simplification to match the reading abilities of specific target audiences, as different populations have different comprehension needs. Since socio-demographic factors such as education level and health literacy are known to correlate with reading abilities, we hypothesize that large language models (LLMs) may be able to adjust their simplification strategies when provided with descriptions of target audiences. In this study, we investigate how LLMs simplify French medical documents when prompted with socio-demographic characteristics of target patients. We compare this approach with prompts based on language proficiency levels (CEFR) to determine whether LLMs respond differently to explicit proficiency levels versus implicit audience descriptions. Our experiments with five LLMs on three types of French medical documents show that CEFR prompts produce greater readability variation (particularly for Llama-3.1-8B), while socio-demographic factors yield more homogeneous outputs. Text genre also considerably impacts LLM outputs for ATS.
A Learner-Oriented Annotated Resource of French Multiword Expressions for Text Adaptation in Foreign Language Reading
Anna Kalinina | Thomas François | Hélène Vassiliadou | Amalia Todirascu
Anna Kalinina | Thomas François | Hélène Vassiliadou | Amalia Todirascu
This article presents a learner-oriented annotated lexical resource of French multiword expressions (MWEs) designed to support text adaptation in foreign language reading. MWEs, including idioms and collocations, pose major comprehension challenges for learners because their meaning often cannot be inferred compositionally or depends on conventional lexical constraints. To address this issue, the study extends the existing verbal MWE database by integrating nominal and verbal MWEs annotated according to a linguistically grounded typology distinguishing idioms, opaque collocations, and transparent collocations. The resource was developed through a multi-step methodology combining automatic extraction from pedagogical corpora, manual annotation using decision-tree-based guidelines, and CEFR level assignment based on corpus distribution. The resulting dataset includes approximately 2,700 expressions enriched with detailed linguistic and learner-relevant metadata. Annotation campaigns involving native and non-native annotators showed moderate agreement, reflecting the gradient nature of phraseological opacity. By linking phraseological complexity with learner proficiency, this resource provides a reproducible framework for modeling MWE difficulty. It offers valuable support for text adaptation, readability assessment, and the development of NLP-based educational tools, contributing to improved accessibility of French texts for language learners.
A Meta-evaluation of Automatic Metrics for Elaborative Simplification
Abdullah Alshatti | Steven Schockaert | Fernando Alva-Manchego
Abdullah Alshatti | Steven Schockaert | Fernando Alva-Manchego
Elaborative simplification aims to improve the readability of texts by adding content that helps the readers. However, evaluating these elaborations remains challenging due to their subjective nature and the lack of suitable annotated datasets. To support the evaluation of elaborative simplification models, we introduce a new dataset with human ratings of elaborations generated by Large Language Models (LLMs), focusing on two quality criteria: cohesion and informativeness. Using these human judgments as a reference, we conduct a meta-evaluation of existing automatic evaluation approaches, with a focus on LLM-as-a-judge strategies. Our experiments suggest that evaluations made by smaller LLMs correlate poorly with human judgments, while larger models with structured prompting exhibit higher agreement. Informativeness evaluation proved to be challenging due to its subjectivity, as evidenced by the low inter-annotator agreement compared to cohesion.
Readability Measures in Automatic Text Simplification: Is Simplification Quality a Coherent Construct?
Rémi Cardon | A. Seza Doğruöz
Rémi Cardon | A. Seza Doğruöz
Readability is a central concept in automatic text simplification (ATS), yet the two fields have largely developed in parallel, with limited cross-fertilization. While prior work has studied correlations between automatic evaluation metrics and human judgment in ATS, the correlations between these two aspects and readability measures have not received systematic attention. We address this gap by investigating to what extent readability measures align with both human judgment and automatic metrics in ATS. Using two English datasets annotated with human judgments (SimplicityDA at the sentence level and D-Wikipedia at the document level), we compute 1,066 linguistic features (covering lexical diversity, lexical sophistication, syntactic sophistication, and cohesion) and eight traditional readability formulas, and correlate them against human scores and standard ATS metrics (BLEU, SARI, BERTScore, LENS, D-SARI). Our results show that readability measures correlate poorly with both human judgment and automatic metrics across both levels. The meaning preservation criterion consistently yields the highest correlation values, while simplicity and fluency criteria remain low. We also find systematic differences between sentence-level and document-level simplification in terms of which features are most informative: type-token ratio features are predictive at the sentence level but not at the document level, while corpus-frequency features show the opposite pattern. These findings point to a broader issue: ATS lacks a shared theoretical construct for simplification quality, and the three main approaches to its assessment (human judgment, readability measures, and automatic metrics) do not consistently converge.
Understanding whether proficiency is encoded as structured knowledge rather than inferred from surface correlates is critical for interpreting and applying LLMs in educational contexts. We investigate whether multilingual large language model (LLM) embeddings encode language proficiency as a structured recoverable dimension rather than merely supporting predictive classification. Using the UniversalCEFR benchmark, which spans 13 languages and the full proficiency range from A1 to C2, we evaluate the frozen LLM embedding space in two complementary ways. First, we test whether proficiency levels can be predicted directly from frozen embeddings across languages and model variants. The results show that embeddings without task-specific fine-tuning consistently support CEFR classification. Variation in results is strongly associated with the amount of annotated data and language family, suggesting that data availability and cross-linguistic structure matter more than architectural differences. Second, we examine how CEFR levels are organized inside embedding space. We find that texts from lower to higher proficiency levels align along a consistent ordered direction, with higher levels systematically positioned further along this gradient. Distances between levels increase proportionally to their ordinal gap (e.g., A1 vs. C2 is farther apart than B1 vs. B2), indicating a continuous proficiency continuum rather than arbitrary clusters. Together, these findings show that CEFR is not only predictable from multilingual LLM embeddings but is also internally structured as an ordered representational dimension.
up
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Felix Morger | Nikolai Ilinykh | Barbara Scalvini | Simon Dobnik | Dana Dannélls
Felix Morger | Nikolai Ilinykh | Barbara Scalvini | Simon Dobnik | Dana Dannélls
Lost in Translation: Repurposing semantic similarity benchmarks for evaluating lexical-semantic consistency in LLM-based machine translation
Quin Ye | Jelke Bloem
Quin Ye | Jelke Bloem
We propose and demonstrate a repurposing of the lexical similarity benchmark Multi-SimLex and the SimLex-999 family of resources for assessing the cross-lingual lexical-semantic consistency of multilingual large language models. While originally gathered for evaluating word embedding models, the parallel nature of the word pairs enables their use in machine translation settings. Using a manually verified subset of 500 word pairs from the Multi-SimLex dataset, we evaluate models’ ability to assess semantic similarity and perform translation between English and Mandarin through zero-shot prompting. We compare BLOOMZ and GPT-4’s similarity ratings against human-annotated benchmarks and examine translation consistency using our and other metrics, with GPT-4 showing stronger human alignment. As SimLex-999 and Multi-SimLex together cover a range of at least 25 languages, this approach has the potential to be extended to many language pairs including ones that don’t involve English, though it requires some manual checks.
Bridging the Low Resource Gap in Historical Cryptology: A Multilingual Diachronic Synthetic Dataset for Reproducible Cryptanalysis
Micaella Bruton | Meriem Beloucif | Beáta Megyesi
Micaella Bruton | Meriem Beloucif | Beáta Megyesi
Many NLP tasks suffer from limited aligned supervision in the target domain. Historical cipher decryption represents an extreme case: aligned plaintext–ciphertext pairs are scarce, access to decrypted archives is restricted, and prior work often relies on synthetic data that is neither released nor evaluated for realism. This limits reproducibility and obscures whether models trained on artificial benchmarks transfer to archival conditions. We introduce HistCiph, the first publicly available multilingual collection of historically grounded plaintext–ciphertext datasets for classical ciphers. Spanning ten languages and multiple centuries, the collection combines diachronically balanced historical plaintext with independently generated homophonic substitution keys and controlled transcription noise. Synthetic generation is explicitly constrained by documented properties of historical ciphers, including multi-homophone allocation and variable-length codes. We validate the datasets using information-theoretic diagnostics—entropy, redundancy, frequency masking, and unicity distance—showing that ciphertext distributions approach theoretical bounds while preserving cross-linguistic variation. HistCiph provides a reproducible benchmark for neural decryption and alignment, and illustrates a principled framework for empirically grounded synthetic data generation in low-resource NLP.
Cultural Grounding in Swedish: Extending an Everyday Knowledge Benchmark for LLMs
Meriem Beloucif | Johan Sjons
Meriem Beloucif | Johan Sjons
Benchmarks for evaluating Large Language Models (LLMs) on everyday knowledge across cultures and languages are increasingly used to assess cultural competence and contextual understanding. However, many multilingual extensions rely primarily on translated question–answer pairs, limiting their ability to capture locally grounded variation. In this work, we present a Swedish extension of an existing cross-cultural everyday knowledge benchmark in which questions are translated into Swedish, and answers are individually collected from five participants, coming from diverse social and professional backgrounds. This design enables us to capture culturally situated, naturally produced responses rather than transferred or translated answer templates. We document the translation protocol, annotators, and agreement analysis, and examine variation across annotators as a signal of culturally contingent knowledge. We evaluate several state-of-the-art multilingual and instruction-tuned LLMs against the aggregated human responses and analyze model performance. Our results reveal that while models often approximate prototypical answers, they struggle with culturally specific nuances and intra-cultural variation. The Swedish extension provides a resource for studying culturally grounded evaluation and highlights the importance of human-generated local answers when benchmarking LLMs across languages.
Entity Linking for Faroese Using Large Language Models with Web Search
Annika Simonsen | Iben Nyholm Debess | Hafsteinn Einarsson
Annika Simonsen | Iben Nyholm Debess | Hafsteinn Einarsson
Entity linking connects text mentions to knowledge bases. For low-resource languages, entity linking has typically not been a research priority, as named entity recognition and knowledge base creation must first be addressed. We present the first study of entity linking for Faroese, a North Germanic language with approximately 70,000 speakers. Unlike traditional systems that rely on separate candidate retrieval and ranking components, we employ an end-to-end approach using GPT-5 with integrated web search. Our method prompts the model to directly identify and link named entities to Wikipedia pages through a three-tier fallback strategy: Faroese Wikipedia, English Wikipedia, and finally any available Wikipedia. We evaluate our approach on 1,010 manually annotated examples from a Faroese NER dataset, analyzing entity mentions across Person, Location, Organization, and Miscellaneous types. Human evaluation shows our system achieves 87.5% precision and 87.3% recall, with particularly strong performance on locations (93-95% precision, 92-95% recall). Persons are more challenging (86-88% precision, 72-83% recall). The majority of links (76.5%) point to Faroese Wikipedia, demonstrating the model’s ability to leverage language-specific knowledge bases. A Wikipedia API search baseline without any LLM achieves F1 = 0.57–0.60 on the same evaluation data, confirming that the LLM’s contextual reasoning provides substantial gains over simple search. We validate our approach across three models (GPT-5, Gemini 3 Flash, GPT-5.4 Mini), achieving F1 scores of 0.74–0.87 and confirming that the method generalizes across providers. This work establishes initial performance benchmarks for Faroese entity linking and demonstrates the viability of LLM-based approaches for low-resource languages.
From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene
Mojca Brglez | Spela Vintar
Mojca Brglez | Spela Vintar
Large language models are demonstrating increasing capabilities, excelling at benchmarks once considered very difficult. As their capabilities grow, there is a need for more challenging evaluations that go beyond surface-level linguistic competence. The latter involves not only syntax and semantics but also pragmatics, i.e., understanding situational meaning shaped by context and linguistic and cultural norms. To contribute to this line of research, we introduce SloPragEval and SloPragMega, the first pragmatics understanding benchmarks for Slovene, comprising 405 multiple-choice questions. We discuss the difficulties of translation, describe the campaign to establish a human baseline, and report pilot evaluations with LLMs. Our results indicate that current models have substantially improved in their understanding of nuanced language but may still fail to infer implied speaker meaning in non-literal utterances, especially those that are culture-specific. We also observe a significant gap between proprietary and open-source models. Finally, we argue that benchmarks targeting nuanced language understanding and knowledge of the target culture must be designed with care, preferably constructed from native data, and validated with human responses.
SdQuAD: A Large Benchmark Question Answering Dataset for Low-resource Sindhi Language
Wazir Ali | Muhammad Rafay Shaikh | Nadia Ali | Amar Rehman
Wazir Ali | Muhammad Rafay Shaikh | Nadia Ali | Amar Rehman
Question answering (QA) datasets are crucial for developing and evaluating monolingual and multilingual language models, yet low-resource languages like Sindhi lack open-source QA resources. We introduce SdQuAD, a novel open-source textual QA dataset for the low-resource Sindhi language, comprising 15,000 QA pairs meticulously annotated by native speakers using the Label Studio platform. Sourced from diverse domains, including news, history, science, geography, business, and tourism, SdQuAD supports both extractive and abstractive QA tasks while capturing Sindhi’s linguistic and topical diversity. We assess annotation quality using span-level agreement and evaluate extractive performance with Exact Match (EM), F1 score, and a TF-IDF baseline. Additionally, we fine-tune mBERT, XLM-R, and mT5 models on SdQuAD, benchmarking their performance to demonstrate the dataset’s utility.
LLMs as Assistants for Data Annotation: Addressing Disagreement and Supporting Expert Processes
Mark Andrade | Bláithín Heffernan | Abigail Walsh | Sheila Castilho
Mark Andrade | Bláithín Heffernan | Abigail Walsh | Sheila Castilho
This paper investigates the potential of Large Language Models to assist human annotation pipelines, with a particular focus on supporting the development of expert-informed annotation guidelines for document-level content categorisation. We present three experiments exploring distinct roles for LLMs in annotation: as annotators, as domain experts assisting in disagreement resolution, and as analysts of annotator discussions. Using GPT-4.5 and Claude Sonnet 4, we evaluate LLM-generated annotation guidelines for a document-level classification tasks in terms of coverage, applicability, and usefulness. Preliminary results are mixed-to-positive, with evidence that LLMs can provide useful support across different stages of the annotation pipeline, particularly when supplied with rich contextual information such as prior human annotations and annotator discussions. However, their effectiveness remains sensitive to prompting strategies and input configuration.
Annotation Quality in Aspect-Based Sentiment Analysis: A Case Study Comparing Experts, Students, Crowdworkers, and Large Language Models
Niklas Donhauser | Jakob Fehle | Nils Constantin Hellwig | Markus Weinberger | Udo Kruschwitz | Christian Wolff
Niklas Donhauser | Jakob Fehle | Nils Constantin Hellwig | Markus Weinberger | Udo Kruschwitz | Christian Wolff
Aspect-Based Sentiment Analysis (ABSA) enables fine-grained opinion analysis by identifying sentiments toward specific aspects or targets within a text. While ABSA has been widely studied for English, research on other languages such as German remains limited, largely due to the lack of high-quality annotated datasets. This paper examines how different annotation sources influence the development of German ABSA. To this end, an existing dataset is re-annotated by experts to establish a ground truth, which serves as a reference for evaluating annotations produced by students, crowdworkers, Large Language Models (LLMs), and experts. Annotation quality is compared using Inter-Annotator Agreement (IAA) and its impact on downstream model performance for different ABSA subtasks. The evaluation focuses on Aspect Category Sentiment Analysis (ACSA) and Target Aspect Sentiment Detection (TASD). We apply State-of-the-Art (SOTA) methods for ABSA, including BERT-, T5-, and LLaMA-based approaches to assess performance differences, spanning fine-tuning and in-context learning with instruction prompts. The findings provide practical insights into trade-offs between annotation reliability, and efficiency, offering guidance for dataset construction in under-resourced Natural Language Processing (NLP) scenarios.
Cross-Lingual Mathematical Reasoning in LLMs: Evaluating Performance on Icelandic vs. English Problems
Hafsteinn Einarsson
Hafsteinn Einarsson
We investigate whether large language models (LLMs) exhibit performance differences when solving mathematical problems presented in a low-resource language (Icelandic) versus a high-resource language (English). Using 847 multiple-choice problems from the Icelandic Mathematics Competition corpus (STAK), we evaluate two state-of-the-art models (Gemini-3-Flash-Preview and GPT-5.4-mini) in both multiple-choice (MC) and open-ended (OE) formats, with correctness determined by a three-judge quorum (Gemini-3-Flash, GPT-5.4-mini, Claude Sonnet 4.6) achieving 97.6% unanimous agreement. Our results reveal significant cross-lingual performance gaps that vary by model: Gemini-3-Flash shows a consistent English advantage of 2.4–10.0 percentage points across both evaluation modes, while GPT-5.4-mini exhibits no significant language effects. Notably, GPT-5.4-mini demonstrates a substantial MC deficit, achieving only 42% in that format despite reaching 69-71% accuracy on OE problems. Analysis of answer patterns reveals a strong option position bias in GPT-5.4-mini, with systematic over-selection of option B and under-selection of option D. These findings suggest that language does affect LLM mathematical reasoning for some models, but the effect is model-dependent and interacts with evaluation format, with implications for deploying LLMs in educational contexts for speakers of low-resource languages.
Struct2Unstruct: Creating Tender NER Datasets from Structured Procurement Records using Large Language Models
Asim Abbas | Mark Lee | Niloofer Shanavas | Venelin Kovatchev | Mubashir Ali
Asim Abbas | Mark Lee | Niloofer Shanavas | Venelin Kovatchev | Mubashir Ali
Named Entity Recognition (NER) in the tender and procurement domain is critical for tasks such as contract monitoring, supplier analysis, and compliance tracking. However, unlike general-purpose NER, no open-source datasets exist for Tender NER, largely due to data sensitivity and confidentiality restrictions. This scarcity limits the development of automated entity extraction models. To address this gap, we propose struct2unstruct, a data preparation pipeline that generates and annotates tender-specific datasets using large language models (LLMs). Starting from structured procurement data published by the Singapore government (2015–2021) available in English language, we employ Llama-3 to generate synthetic tender narratives in multiple writing styles, ensuring each contains at least one tender-related entity. Post-processing steps correct inconsistencies in dates, symbols, and entity formats. Entities are then annotated using a BIO tagging scheme through deterministic alignment with structured fields, followed by expert validation to ensure accuracy. This study focuses on data preparation and evaluation, not model training. The resulting dataset provides a scalable resource for future Tender NER research in low-resource environments. By releasing both the dataset and pipeline as open-source resources, we establish a foundation for advancing domain-adapted information extraction and automated tender entity recognition.
Link Prediction for Event Logs in the Process Industry
Anastasia Zhukova | Thomas Walton | Christian E. Lobmüller | Bela Gipp
Anastasia Zhukova | Thomas Walton | Christian E. Lobmüller | Bela Gipp
In the era of graph-based retrieval-augmented generation (RAG), link prediction is a significant preprocessing step for improving the quality of fragmented or incomplete domain-specific data for the graph retrieval. Knowledge management in the process industry uses RAG-based applications to optimize operations, ensure safety, and facilitate continuous improvement by effectively leveraging operational data and past insights. A key challenge in this domain is the fragmented nature of event logs in shift books, where related records are often kept separate, even though they belong to a single event or process. This fragmentation hinders the recommendation of previously implemented solutions to users, which is crucial in the timely problem-solving at live production sites. To address this problem, we develop a record linking model, which we define as a cross-document coreference resolution (CDCR) task. Record linking adapts the task definition of CDCR and combines two state-of-the-art CDCR models with the principles of natural language inference (NLI) and semantic text similarity (STS) to perform link prediction. The evaluation shows that our record linking model outperformed the best versions of our baselines, i.e., NLI and STS, by 28 (11.43 p) and 27.4 (11.21 p), respectively. Our work demonstrates that common NLP tasks can be combined and adapted to a domain-specific setting of the German process industry, improving data quality and connectivity in shift logs.
We create high-quality datasets for LLM evaluation of logical reasoning skills across nine different languages, which have been manually checked by fluent speakers. The datasets consist of so-called zebra puzzles, and we analyse different ways of tuning the difficulty of the puzzles to fit modern LLMs. This includes the size of the puzzle (number of objects and number of clues), as well as a novel addition of red herring clues containing only irrelevant information. We show that presence of red herrings indeed makes the puzzles significantly harder for the models, and we find puzzle sizes 2×3 and 4×5 are sufficiently challenging for GPT-4o mini (a non-reasoning model) and o3-mini (a reasoning model), respectively. We analyse whether LLM performance of these are sensitive to the language, the cultural sensitivity of the puzzle theme, and the choice of clue types. These analyses are conducted with English and Danish, where we show that there is no significant difference for either of these three aspects, at least for the OpenAI models GPT-4o mini and o3-mini, chosen as representative non-reasoning and reasoning models, respectively. We publish the datasets for each of the nine languages for the identified sizes 2×3 and 4×5. We also publish the code used to generate the puzzles, which can be used to extend the benchmark into more languages.
Progressing beyond Art Masterpieces or Touristic Clichés: how to assess your LLMs for cultural alignment?
António Branco | João Ricardo Silva | Nuno Marques | Luis M. S. Gomes | Ricardo Campos | Raquel Sequeira | Sara Nerea | Rodrigo Silva | Miguel Marques | Rodrigo Duarte | Artur Putyato | Diogo Folques | Tiago Valente
António Branco | João Ricardo Silva | Nuno Marques | Luis M. S. Gomes | Ricardo Campos | Raquel Sequeira | Sara Nerea | Rodrigo Silva | Miguel Marques | Rodrigo Duarte | Artur Putyato | Diogo Folques | Tiago Valente
Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention - often framed in terms of cultural bias - until recently there has been limited work on the design and development of datasets for cultural assessment. Here, we review existing approaches to such datasets and identify their main limitations. To address these issues, we propose design guidelines for annotators and report on the construction of a dataset built according to these principles. We further present a series of contrastive experiments conducted with this dataset. The results demonstrate that our design yields test sets with greater discriminative power, effectively distinguishing between models specialized for a given culture and those that are not, ceteris paribus.
Evaluating Large Language Model-based Natural Language Generation for Modular Dialog systems
Vincent Emmerling | Christoph Kowalski | Amelie Sophie Robrecht-Hilbig | Stefan Kopp
Vincent Emmerling | Christoph Kowalski | Amelie Sophie Robrecht-Hilbig | Stefan Kopp
While many dialogue systems currently use end-to-end solutions, modular systems offer greater control, sustainability, and more human-like dialogue. This makes them relevant especially when aiming to study human behavior patterns in interactions or applying them to sensitive domains. In this paper, we develop an automated metric to measure the quality of an LLM-based NLG-component in a modular system based on the hallucination tendency and linguistic quality. We apply the metric to various language models and usage techniques and, based on the results, discuss the conditions a model must meet in order to be a good candidate for an NLG-component in a real-time capable dialogue system. Although such automated metrics cannot replace a real interaction study, they help to compare potential approaches of the individual modules. Therefore, they are indispensable when developing and testing modules in isolation. One advancement of the introduced metrics is that it is developed and tested on a German dataset, showing challenges when working with languages other than English and discrepancies to the abilities of Generative AI assumed in current state-of-the-art literature.
JobResQA: Semi-Automatic Multilingual Benchmark Creation for LLM Machine Reading Comprehension on Résumés and Job Descriptions
Casimiro Pio Carrino | Paula Estrella | Rabih Zbib | Carlos Escolano | Jose A. R. Fonollosa
Casimiro Pio Carrino | Paula Estrella | Rabih Zbib | Carlos Escolano | Jose A. R. Fonollosa
We present a methodology for building privacy-preserving multilingual QA benchmarks in low-resource and sensitive domains, demonstrated through JobResQA, a multilingual MRC benchmark over synthetic HR documents. The dataset comprises 581 QA pairs across 105 synthetic résumé-job description pairs in five languages (English, Spanish, Italian, German, and Chinese), with questions spanning four types based on document source (intra vs. cross-document) and reasoning complexity (single-hop vs. multi-hop). We propose a privacy-preserving synthetic data pipeline applicable to other sensitive domains, with controlled demographic attributes (via placeholders) enabling future bias studies. Our cost-effective, human-in-the-loop translation pipeline based on TEaR methodology incorporates MQM error annotations and selective post-editing. Baseline evaluations across multiple open-weight LLM families using LLM-as-judge reveal higher performance on English and Spanish but substantial degradation for other languages, highlighting critical cross-lingual MRC gaps. Our pipeline, where LLMs act as synthesizers, translators, and evaluators under human oversight, constitutes a reusable methodology for resource creation and a case study in evaluation-integrity challenges of LLM-era benchmark construction.
Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese
Wajdi Zaghouani | Kholoud Khalil Aldous | Yicheng Gao
Wajdi Zaghouani | Kholoud Khalil Aldous | Yicheng Gao
When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break down. These systems struggle to cross linguistic and cultural boundaries, leaving models exposed to adversarial prompts that exploit Chinese-specific evasion techniques, including Pinyin romanization, character decomposition, Internet slang, and hedging tone. To address this gap, we introduce ChiSafe-PAS (Chinese Safety Pilot Annotation Set), a human-annotated benchmark of 1,897 adversarial Chinese prompts spanning four high-stakes domains: self-harm and violence, drugs and illicit trade, fraud, and satire. Of these, 1,544 entries carry complete gold-standard annotations: a 3-class response label (refuse, redirect, respond), a nine-category obfuscation taxonomy, a risk-level rating, and annotator rationale. We describe the dataset design, annotation process, and obfuscation taxonomy in detail. Our primary goal is practical: to give the research community a high-quality, culturally grounded resource for benchmarking LLM safety alignment. In doing so, we engage three broader tensions in the field: the blurring boundary between training and evaluation data, the need for domain coverage grounded in real-world risk, and the limits of scale as a substitute for cultural expertise.
Most hallucination evaluations focus on English, leaving it unclear whether findings transfer to lower-resource languages. We investigate faithfulness hallucinations, defined as model-generated content that is fluent and plausible but diverges from the provided input or is internally inconsistent. Leveraging the multilingual MultiWikiQA dataset, we utilize the LettuceDetect framework to create synthetic hallucination datasets for 21 European languages, which are then used to create token-level hallucination classifiers. In this work, we present evaluations of model hallucinations on a selection of languages: English, Danish, German, and Icelandic. Using these classifiers, we evaluate the hallucination rates for Qwen3-0.6B, Qwen3-14B, Gemma-3-12B-IT, cogito-v1-preview-qwen-32B, and cogito-v1-preview-llama-70B. Our classifiers reveal notably higher hallucination rates for Qwen3-0.6B (up to 60% of answers containing at least one hallucination, peaking in Icelandic) and generally lower rates for larger models, with cogito-v1-preview-qwen-32B and cogito-v1-preview-llama-70B performing best on most languages. Hallucination rates are consistently higher for lower-resource languages, particularly Icelandic.
Exploring the similarities and differences between VLM-driven and traditional OCR for Historical Swedish Data
Martin Johansson | Selma Waginder | Dana Dannélls
Martin Johansson | Selma Waginder | Dana Dannélls
Recent Swedish OCR efforts rely primarily on traditional OCR methods, including deep CNN–LSTM hybrid neural networks and transformer-based models. Some approaches have also demonstrated the applicability of VLM-driven OCR to historical material. However, to date, no studies have examined in depth the performance of VLM-based OCR on historical Swedish sources. In this paper, we ask: How do transformers and VLMs differ in character- and word-level recognition performance across typefaces, and what qualitative differences can be observed in their error patterns? We show that fine-tuned versions of the Alibaba Cloud Qwen3-VL-8B-Instruct and Qwen3-VL-2B-Instruct, combined with a simple repetition-trimming step, outperform conventional OCR systems. Remaining errors are primarily attributable to challenges associated with the Blackletter typeface and formatting issues, such as missing or extra line breaks, characters, and spaces. Even when characters are correctly recognized, formatting inconsistencies can substantially increase transcription error rates.
up
Proceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion
Proceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion
Eleni Efthimiou | Stavroula-Evita Fotinea | Thomas Hanke | Julie A. Hochgesang | Johanna Mesch | Marc Schulder
Eleni Efthimiou | Stavroula-Evita Fotinea | Thomas Hanke | Julie A. Hochgesang | Johanna Mesch | Marc Schulder
Capturing Methodology for Generating Synthetic and 3D Training Data in Catalan Sign Language (LSC): The Case of Verbal Agreement
Gemma Barberà | Inés Broto Clemente | Xavier Vinaixa Roselló | Roger Cassany Viladomat
Gemma Barberà | Inés Broto Clemente | Xavier Vinaixa Roselló | Roger Cassany Viladomat
This paper proposes a hybrid methodology to generate high-quality synthetic data. Unlike other approaches based purely on generative Artificial Intelligence, which may suffer from hallucinations or inconsistent movements, this project uses 3D biomechanics and kinematics algorithms that enforce the anatomical constraints of the human body to ensure physically plausible movements. The aim of this research is to demonstrate that it is possible to synthetically expand the dataset. In particular, this paper focuses on verb agreement, a grammatical domain which is known for its morphological and articulatory complexity. By concentrating on the possible configurations of the movements in signing space when expressing different person agreeing verbal forms, we aim to capture real movements to extract physical parameters and apply them as logical rules —similar to those of a video game engine— to automatically synthesize thousands of new conjugations from infinitives with complete anatomical precision. Beyond spatial conjugation, the methodology further augments data through procedural variation of prosody and body morphology.
Leveraging Unannotated Sign Language Data via a Robust Data Augmentation Method for Contrastive Representation Learning
Ariel Basso Madjoukeng | Pierre Poitier | Belise Edith Kenmogne | Adelaide Couplet | Margaux Leleu | Frenay Benoit
Ariel Basso Madjoukeng | Pierre Poitier | Belise Edith Kenmogne | Adelaide Couplet | Margaux Leleu | Frenay Benoit
Contrastive learning is a deep learning paradigm that allows the learning of useful representations without annotations. In many fields, including sign language recognition (SLR), contrastive approaches have proven to be very effective for developing pretrained models. To learn representations, they generate augmented variants of an instance through augmentation techniques and then maximize their similarities. The quality of the learned representations is strongly correlated with the augmentations used during training. In several fields, specialized augmentations have been developed and adopted. However, in SLR, we observed two trends: contrastive-based SLR approaches often rely on augmentations that are not realistic for the application (e.g., vertical flip, excessive rotations); specialized augmentation methods lack robustness. Hence, when they are used as a starting point for contrastive algorithms, the learned representations are often irrelevant, and sometimes sensitive. These issues considerably affect the accuracy of SLR models on downstream tasks. In response, this paper proposes a robust augmentation method specially designed for contrastive approaches applied to SLR. The results show an improvement in accuracy during linear evaluation and semi-supervised learning with only 30% of annotations.
The SMILE Continuous DSGS Corpus: A Resource for Longitudinal Exploration of Continuous Swiss German Sign Language
Alessia Battisti | Katja Tissi | Sandra Sidler-Miserez | Sarah Ebling
Alessia Battisti | Katja Tissi | Sandra Sidler-Miserez | Sarah Ebling
This paper presents the SMILE Continuous DSGS Corpus, a longitudinal dataset that allows for investigating how hearing adults acquire Swiss German Sign Language as a second language. It includes recordings of sign language learners and native signer controls collected at four points over a period of 18 months and annotated for manual and non-manual components, errors, and sentence-level acceptability. The resource provides high-quality, synchronized video suitable for both linguistic and automatic sign language processing research, for example, supporting studies of interlanguage development and training of automatic sign language recognition models. We present here an exploratory analysis of the learner subcorpus using Bayesian mixed-effects modeling. The corpus and accompanying annotations are available for research purposes under a Creative Commons license (CC BY-NC-SA 4.0).
The In-Car Sign Language Corpus (ICSL): A Multi-Modal Resource for Constrained-Space Sign Language Recognition
Raviteja Boddu | Guilherme Vieira Leite | Joed Lopes da Silva | Ângelo Benetti | Isabela Barbieri | Natália de Melo Afonso | Thyago Santos | Hélio Pedrini | Felipe Venâncio Barbosa | José Mario De Martino | Munir Georges | Alessandro Zimmer
Raviteja Boddu | Guilherme Vieira Leite | Joed Lopes da Silva | Ângelo Benetti | Isabela Barbieri | Natália de Melo Afonso | Thyago Santos | Hélio Pedrini | Felipe Venâncio Barbosa | José Mario De Martino | Munir Georges | Alessandro Zimmer
This paper addresses the challenges of using sign language within shared mobility services, such as taxis, carpools, or ride-sharing platforms. The use of sign language recognition (SLR) in real-world, confined environments, specifically vehicle interiors remains largely unexplored. To motivate research in this area, we present the In-Car Sign Language (ICSL) dataset for Brazilian Sign Language (Libras), with the long-term goal of improving public transport accessibility for the Deaf and Hard-of-Hearing community. The dataset consists of: (1) high-precision laboratory motion capture (MoCap) data to establish an idealized linguistic baseline and (2) real-world multi-modal in-car recordings captured using a 2D camera and 3D Time-of-Flight sensors. The dataset provides a basis for comparative analyses between synthesized signing avatar animations and recorded real signing interpreter videos, which enable future research into robust “in-the-wild” SLR models and domain adaptation. We describe in detail the use cases, the setup, the data collection protocol, and the metadata structure of the corpus. In total, we recorded a multimodal dataset exceeding 1.5 million frames, comprising the synchronized multimodal streams described above featuring Libras users across various in-car scenarios. The corpus is provided with gloss annotation of lexical signs and non-lexical sign language elements specially designed to support the training and evaluation of deep neural networks for constrained space recognition. In-vehicle signing offers a technically significant example of a constrained, occluded, and non-frontal environment. While recognizing the diverse communication strategies already employed by the Deaf community, identifying automotive-specific limitations provides a useful stepping stone for research into enhancing in-car accessibility and passenger quality of life.
This study is a computer vision analysis of 4.5 hours of video data from 40 signers in the Swedish Sign Language Corpus, aiming to evaluate the reliability of classifying 1) who the main signer is at any given time during dyadic conversation, and 2) the dominant hand (i.e., handedness) of each signer. First, the distance moved by the hands of each signer is used to compare the manual activity between a) the two signers to determine whose hands are more active, and b) the hands of each signer to determine which hand is more likely to be dominant. Second, the height of the hands is used to compare their prominence in signing space between a) the two signers to determine whose hands are more prominent, and b) the hands of each signer to determine which hand is more likely to be dominant. The results show that while both distance and height approaches can reliably classify – individually or combined – the main signer in any segment of a conversation, the height approach is better at determining the overall handedness (right- or left-dominant) of signers. For the handedness classification, the optimal method turns out to be a two-step approach, first classifying the main signer per segment, then using only signer-relevant segments to classify handedness.
SignGPT and the Visual Language Toolkit
Matt Brown | Oline Ranum | Edward Fish | Heidi Proctor | Bencie Woll | Richard Bowden | Kearsy Cormier
Matt Brown | Oline Ranum | Edward Fish | Heidi Proctor | Bencie Woll | Richard Bowden | Kearsy Cormier
SignGPT’s Visual Language Toolkit (VLTK) aims to remove fundamental barriers to large scale sign language modelling by developing data-driven, linguistically grounded methods for continuous sign language recognition. We first identify fundamental issues around the ecological validity of potential data sources (e.g. broadcast media with interpreted signing or captions, scraping of social media). We contrast these with the currently highly resource-intensive development of curated sign language corpora based on linguistic principles. The VLTK addresses this scarcity of high quality sign language data by providing semi-automated glossing and other recognition tools, driving large scale corpus expansion without sacrificing linguistic principles. Unlike prior systems that rely on sparse glossing, the project integrates dense temporal annotation, non-manual and non-lexical feature tracking, and transformer-based architectures to capture the multimodal and spatial structure of signing. By aligning machine vision innovation with linguistic insights and community-embedded evaluation, SignGPT establishes a foundation for robust and extensible sign language models.
Nonmanual markers, such as head and eyebrow movements, eye blinks, and mouth shapes, are an important part of natural languages, both spoken and signed. Recent developments in computer vision have made it possible to extract facial and body landmark positions, as well as head-rotation measures, from 2D video recordings, which can be further processed to analyse the kinematics of nonmanual articulators. In this paper, we present an R-based workflow for processing raw outputs of computer vision toolkits with the goal of producing reliable and interpretable kinematic measurements of nonmanual articulators.
A Small Model for Big Articulators: Sign Language Detection With a Tiny Machine Learning Model
Frederick Chan | Gina-Anne Levow | Qi Cheng
Frederick Chan | Gina-Anne Levow | Qi Cheng
This paper introduces a small (1,013 parameter) machine learning model for sign language detection in videos of isolated American Sign Language (ASL) signs. Our model aims to alleviate the time-consuming nature of producing sign clips for psycholinguistic study stimuli, sign dictionaries, and sign databases. Given a video where the signer starts from a resting position, signs a sign, and returns to the resting position for an arbitrary number of repetitions, the model detects frames in which signing occurs that can be used to segment video into clips of individual signs. We train and evaluate our model on data with precise coding of signing onset and offset from ASL-LEX 2.0, so that our model’s annotations are suitable for psycholinguistics research. The model works on both real signs and pseudosigns, two types of stimuli needed for certain psycholinguistic studies. Our model’s small size compared to the state-of-the-art (100K parameters or more) enables quick, bulk processing even on resource-constrained hardware. It achieves this by computing Instantaneous Visual Change (IVC), a 1D measure of changes in brightness in the input video, extracting features from the IVC-over-time signal with a convolution, and classifying the video frames as signing or non-signing with three neural layers.
"A Sacred Bird Called the Phoenix". Auditing the most-used Parallel Corpus for German Sign Language Recognition and Translation
Vera Czehmann | Shakib Yazdani | Yasser Hamidullah | Fabrizio Nunnari | Eleftherios Avramidis
Vera Czehmann | Shakib Yazdani | Yasser Hamidullah | Fabrizio Nunnari | Eleftherios Avramidis
This paper presents an empirical audit of the widely used RWTH-PHOENIX-2014T corpus, examining its suitability as a benchmark for sign language recognition and translation. Through human annotation of the training set and extensive sign-to-text back translation of the test set, we provide detailed statistics that indicate substantial quality issues, including information loss and lexical errors. Automatic scores comparing human sign-to-text back translations to the original speech transcribed references are remarkably low, suggesting strong translationese effects and substantial paraphrasing, revealing limitations of lexical metrics in adequately scoring translation quality. Replacing the original speech-transcribed references with human sign-to-text back translations while scoring existing sign language translation systems reveals the lack of robustness of system evaluation with lexical metrics against this test set. Our findings highlight risks associated with relying on this corpus for model evaluation and call for more rigorous, linguistically grounded evaluation practices in sign language technology research. The back-translated test set and error annotations are made publicly available.
Diffusion-Based 3D Sign Language Motion Anonymization: A Feasibility Study on Balancing Identity Confusion and Semantic Preservation
Zixuan Dai | Shinji Sako
Zixuan Dai | Shinji Sako
Sign language motions contain individual-specific kinematic features. As the engineering applications of sign language become more widespread, privacy protection of sign language data has emerged as a new challenge. This paper proposes a diffusion model-based approach for sign language motion anonymization. The proposed framework combines conditional diffusion processes with adversarial training to transform identity features while preserving semantic information. For the design and preliminary validation of the proposed model, we conduct a proof-of-concept experiment using a subset of 22 signers from the ASL100 dataset of WLASL, which demonstrates the feasibility of the proposed approach for sign language anonymization.
GeoQuery-LSFB: A French Belgian Sign Language Corpus with Procedural Semantic Annotations
Liesbet De Vos | Laurence Meurant | Paul Van Eecke | Katrien Beuls
Liesbet De Vos | Laurence Meurant | Paul Van Eecke | Katrien Beuls
Procedural semantic representations describe the meaning of natural language expressions in terms of computer programs that can be evaluated against images, databases, knowledge graphs or other external resources. While resources annotated with procedural semantic representations already exist for a variety of spoken languages, such resources are still lacking entirely for signed languages. In this paper, we introduce GeoQuery-LSFB as a signed language extension to the multilingual GeoQuery corpus. Concretely, we have complemented each procedural semantic annotation from the original corpus with a corresponding French Belgian Sign Language (LSFB) expression that was phonetically transcribed from video recordings following the HamNoSys convention and annotated with French ID-glosses. The GeoQuery-LSFB corpus constitutes a new resource for a low-resource language and offers for the first time the possibility to study, from an onomasialogical perspective, a signed language along a diverse variety of spoken languages.
The De-Sign Platform: An Online Psychometric Tool for Dementia Screening of Deaf Older Adults in two Sign Languages, GSL and ÖGS
Athanasia-Lida Dimou | Theodore Goulas | Marianna Tsatali | Tarsita Ntova | Doris Hoffmann-Lamplmair | Stavroula-Evita Fotinea | Eleni Efthimiou | Birgit Teichmann | Magda Tsolaki | Joanna Atkinson | Bencie Woll
Athanasia-Lida Dimou | Theodore Goulas | Marianna Tsatali | Tarsita Ntova | Doris Hoffmann-Lamplmair | Stavroula-Evita Fotinea | Eleni Efthimiou | Birgit Teichmann | Magda Tsolaki | Joanna Atkinson | Bencie Woll
This article presents the De-Sign platform, a web-based psychometric tool specifically designed for screening dementia in Deaf older adults (50+) who use Austrian Sign Language and Greek Sign Language hereinafter ÖGS and GSL respectively. The limited access to dementia services for these populations is primarily attributed to a scarcity of healthcare professionals fluent in sign language. Hence, enhancing access to relevant diagnostic services has become a priority. Currently, there is a significant lack of screening tools specifically developed to identify early signs of dementia that are compatible with national sign languages. To address this issue, the De-Sign Erasmus+ (2022-2025) project has employed suitable psychometric instruments that are adapted to the cultural contexts and linguistic norms of Deaf communities in Austria and Greece. The only existing Cognitive Screening Test (CST) for British Sign Language (BSL), used for diagnosing dementia in Deaf older adults, was initially adapted from English by Atkinson et al. (2015). The De-Sign platform hosts a cognitive screening test in ÖGS and GSL. Both were linguistically and culturally adapted from the BSL-CST test, providing two web-based versions of a psychometric tool that enables dementia screening within these populations.
Feature Analysis of MoCap Data for Optimised Sign Language Processing
Yves A. Duppen | Mirella De Sisto | Ifigeneia Mavridou | Phillip Brown | Lisa Lepp | Dimitar Shterionov
Yves A. Duppen | Mirella De Sisto | Ifigeneia Mavridou | Phillip Brown | Lisa Lepp | Dimitar Shterionov
Despite the rapid advances in AI and its impact on machine translation (MT), when it comes to sign language (SL) processing and MT, there is a big bottleneck – the lack of substantial quantities of quality signed data suitable for developing SLMT models. Marker-based motion capturing (MoCap) is a technique for tracing and recording the body movements (including hands and figures) in 3D space with high precision and has been widely used in SL research. MoCap data is of high representative accuracy, making it very suitable for analysing movement patterns and articulatory features. However, it is also very complex – a recording of a single sign may contain more than 240 entries over 156 features making it difficult for processing. In this paper we analyse MoCap data aiming to understand which captured features are of high importance. Consecutively, we optimise the MoCap data representation, reducing the number of features, and assess how this feature- reduced data impacts sign classification task. We organise MoCap features based on their importance and show how models trained on feature-reduced representations outperform those developed on the complete feature set.
Leveraging Text-side Augmentation For Sign Language Translation
Diandra Fabre | Julie Lascar | Julie Halbout | Markarit Vartampetian
Diandra Fabre | Julie Lascar | Julie Halbout | Markarit Vartampetian
Sign language translation faces significant challenges due to the scarcity of annotated data and the inherent complexity of sign languages. This paper presents a method to improve sign-to-text translation models by augmenting data on the text side. We conduct experiments using two state-of-the-art models on two publicly available datasets: PHOENIX-2014T for German Sign Language and Mediapi-RGB for French Sign Language. Our main contributions are : (1) augmenting the training sets of both datasets on the text side using a generative model, (2) evaluating the impact of paraphrasing on BLEU and BLEURT scores, and (3) analyzing the impact of paraphrasing on translation outputs. We observed a significant improvement in translation for both languages. This suggests that adding variability to the training dataset through paraphrasing can lead to better generalization of the models. These results are comparable to state-of-the-art methods that use more complex approaches, such as Visual-Language fine-tuning, to improve translation.
The Construction of the CORALSE Corpus, Now and Beyond: A Tool for Documenting Spanish Sign Language
Ana Fernández Soneira | María C. Bao-Fente | Rayco H. González-Montesino | Inmaculada C. Báez-Montero
Ana Fernández Soneira | María C. Bao-Fente | Rayco H. González-Montesino | Inmaculada C. Báez-Montero
The main objective of this paper is to present the experience of building the CORALSE corpus and to discuss the challenges that arise when attempting to provide a comprehensive description of a sign language. To this end, we address the following questions, drawing on the data obtained in the completed phases of the CORALSE project as well as on the foundational principles guiding the project’s third phase. THE CORALSE CORPUS TODAY: How have we developed a linguistic corpus of sign language?, What steps have we taken in developing the CORALSE corpus?, Which informants have we recorded and what criteria have guided their selection? THE CORALSE CORPUS IN THE FUTURE: Which (native) languages do we prioritise when selecting informants?, How do the perspectives of reference signers, interpreters, educators, and psycholinguists contribute to a more complete understanding of a sign language? Corpus linguistics is understood as a set of methodologies designed to study language through collections of digitised texts. Its development over recent decades—initially driven by advances in computing and, subsequently, by the emergence of the internet—represents one of the most significant transformations in contemporary linguistic research. The projects CORALSE: Annotated Inter-university Corpus of Spanish Sign Language and Textual Typology, Registers and Styles in Spanish Sign Language: New Data for the Expansion of the CORALSE Corpus adopt a corpus linguistics approach to collect, analyse and describe a representative sample of Spanish Sign Language (LSE). We also reflect on the types of linguistic data that are truly necessary to document the actual use of Spanish Sign Language.
Recently, the Norwegian Sign Language Corpus has been published, and it includes language data from over 100 signers from around Norway. Collecting and building such multimodal signed language corpora have important implications for both research and deaf communities. However, consideration is needed to protect the personal nature of signed language data, while also making a long-term resource that is as accessible as possible to various community, research, and professional stakeholders. In addition, the potential exploitation of corpus resources by commercial and other interests, which are not necessarily aligned with the deaf community itself, must also be deliberated. Here, these seemingly opposing issues and the ethics that surround them are discussed. Current best practices in Open Science (including FAIR and CARE data principles), along with ethical discussions raised by scholars working, for example, in Deaf Studies, are shown to be important in navigating this complex research data landscape.
Generations in the DGS Corpus: Evolving Outreach Activities and Cross-Generational Stories on Social Media in a Long-Term Corpus Project
Anike Fiedler | Marc Schulder | Julian Bleicken | Annika Herrmann
Anike Fiedler | Marc Schulder | Julian Bleicken | Annika Herrmann
Social media has become a powerful tool for research projects, community outreach, science communication, and to recruit participants. Due to its differences to traditional media and presentation modes, it provides a particular focus on producing very concise content that is entertaining and accessible while staying informative. In this paper, we describe how the long-term project DGS-Korpus, creators of a corpus and dictionary of German Sign Language, evolved its outreach strategies over time. One unique aspect of its unusually long project run-time of nineteen years is that it has involved several cases of multiple family members participating in the project at different points in time, resulting in cross-generational participation. The paper describes how the project’s social media campaign uses these cross-generational connections to illustrate important aspects of the project, such as its relevance for cultural heritage and language identity, the different ways that members of the German deaf community were and are involved in the project, and its relevance to interpersonal connections.
Formalising Sign Language Depiction, Characterising Categories and Measuring Iconicity with AZee
Michael Filhol | Emmanuella Martinod
Michael Filhol | Emmanuella Martinod
This paper deals with depiction in (French) Sign Language, the formal account AZee can provide, and how it compares, validates or simplifies the linguistic notions of classifiers and iconic structures. It reports on a partial encoding work on “Mocap1”, a corpus with a high density of depicting structures, following the same method that led to the first AZee reference corpus “40 brèves”. The approach does not postulate classifiers or iconic structures as entities separate from lexical signs, and nonetheless manages to model the corpus data. We discuss the entailed possibility to rediscover some of the useful categories, and if so define them from AZee’s premises. We also specify how a formal metric can be specified to measure iconicity in signed data. While this paper is of linguistic interest as it compares to existing theories, it also provides a concrete step to covering depicting discourse with AZee, therefore enable automatic SL animation of depiction.
Extracting Signs from Weakly Aligned Sign Language Corpora: A Study on LSF and LSM
Lorena de la Garza | Julie Halbout | Julie Lascar | Niels Martinez | Arturo Curiel | Michèle Gouiffès | Annelies Braffort
Lorena de la Garza | Julie Halbout | Julie Lascar | Niels Martinez | Arturo Curiel | Michèle Gouiffès | Annelies Braffort
This paper presents a framework for the automatic annotation of sign language data across different recording conditions, including original and interpreted content. The proposed approach integrates weak alignment, sign segmentation, and multiple instance learning with a contrastive loss. The resulting annotations are subsequently refined and filtered to enhance their reliability. Our method was applied to two historically related sign languages, French Sign Language (LSF) and Mexican Sign Language (LSM). This led to the creation of two signaries, comprising approximately 2k categories in LSF (25k occurrences) and 41 categories in LSM (1k occurrences). Both resources provide valuable support for future research in artificial intelligence and linguistics, particularly for comparative analyses between the two languages. A seminal analysis is presented as part of this paper.
An Annotation Formalism for a French–LSF Bilingual Corpus Supporting Sign Language Generation
Sylvie Gibet | Clément Reverdy | Pierre-François Marteau
Sylvie Gibet | Clément Reverdy | Pierre-François Marteau
This paper introduces an annotation formalism for bilingual corpora of written French and French Sign Language (LSF), based on a manually-produced, expert transcription of LSF video data. The formalism captures the grammatical specificities of LSF, including spatial and iconic mechanisms, while explicitly encoding features that support motor programs for animated signing avatars. We propose a parameterized gloss-based approach, called PGloss-LSF, which integrates syntactic and semantic structures alongside motion features critical for accurate sign synthesis. We illustrate the framework with examples drawn from our bilingual corpus. The annotation process is incremental, ensuring internal consistency and computational tractability through a two-step evaluation: a qualitative assessment aligning generated signs with the annotation language, and a quantitative evaluation via automatic translation using large language models. By bridging the linguistic specificities of sign language with the computational requirements of sign synthesis, this work advances the integration of sign language corpora into multilingual resources and contributes to the standardization of sign language technologies.
A Pose-Based Pipeline for Annotation of Headshakes in Sign Language Corpora
Gustaf Gren | Nikolaus Riemer Kankkonen
Gustaf Gren | Nikolaus Riemer Kankkonen
This paper introduces a pose-based pipeline designed to support scalable annotation of headshakes in sign language corpora. Motivated by the scarcity of annotated datasets and the need for quantitative typological research, the study evaluates whether automated detection can reduce human annotation effort. The system operates on yaw trajectories extracted with MediaPipe Holistic and uses sliding-window segmentation with neural sequence models (LSTM/CNN) to surface candidate segments for review. Training and evaluation are conducted on a subset of the German Sign Language (DGS) Corpus annotated to target grammatical headshakes functioning as negation rather than for every instance of headshakes. On the DGS dataset the best performing LSTM model achieves an F2-score of 0.45, recall of 0.63. Despite the narrow annotation scope, the pipeline reduces search space: annotators need review only 13% of frames to recover 87% of labeled instances. Error analysis indicates that many false positives correspond to plausible head movements excluded by the annotation criteria. A pilot transfer to Swedish Sign Language shows reduced effectiveness without adaptation, underscoring the need for alignment in cross-lingual transfer scenarios.
Learning to Spot Signs from Named Entities. A study on French Sign Language.
Julie Halbout | Annelies Braffort | Michèle Gouiffès | Diandra Fabre | Julie Lascar
Julie Halbout | Annelies Braffort | Michèle Gouiffès | Diandra Fabre | Julie Lascar
French Sign Language (LSF) is a low-resourced language, with few available corpora, most of which being only partially annotated. Previous work on other sign languages has explored automatic sign annotation using subtitles as weak supervision, existing signaries, or mouthing cues. This paper focuses on the corpus Matignon-LSF, by first leveraging lexical token spotting then by studying Named Entities (locations, companies, persons). Accounting for the Named entities enables the automatic detection of 30% to 100% more signs per class and improves the spotting of rare signs. In addition, this work provides insights into the signing of named entities and contributes resources for improving LSF-to-French translation models.
The Iterative Development and Evaluation Framework for Kazakh-Russian Signing Avatars Targeted to Native Deaf Signers
Alfarabi Imashev | Tohid Alizadeh
Alfarabi Imashev | Tohid Alizadeh
Nowadays, existing research predominantly focuses on already well-researched sign languages. However, the most extensive studies of sign language in Kazakhstan, which adhere to international standards, started about a decade ago. Native deaf signers in Kazakhstan can often suffer from insufficient educational opportunities, which may result in limited reading proficiency too. Sometimes, deaf signers can recognize letters and read words, but they may not fully understand the overall concept and need to break it down into a sequence of simpler ideas to comprehend it better. Consequently, signing avatars have the potential to interpret internet statements, movie subtitles, or YouTube videos, and this sign language production may increase accessibility and improve communication between deaf and hearing individuals, as well as between humans and avatars. An equally critical challenge is how to develop a tool that will help deaf signers evaluate the performance, appearance, and naturalness of signing avatars without relying on written text across all sign languages, particularly in underserved communities. This paper outlines the iterative development of the Kazakh-Russian Sign Language interpreting avatar, ongoing improvements to the evaluation instrument, and a comparative analysis of this instrument with another evaluation method designed to attain the same objective.
Movement Coherence in High Visual Load Environments: Implications for Attention in Mixed-Hearing Classes
Mert Inan | Saki Imai | Anna Marshall | Tessa Karel | Malihe Alikhani
Mert Inan | Saki Imai | Anna Marshall | Tessa Karel | Malihe Alikhani
Signed interpretation in movement based instruction creates high visual load environments in which spoken language, sign language, and physical demonstration compete for the same perceptual channel. We present a participatory multimodal observational study of mixed hearing movement and mindfulness classes in which Deaf, Hard of Hearing, and hearing participants practice together. Based on synchronized video recordings and instructor interviews, we examine how alignment across demonstration, signed instruction, and bodily execution is achieved and restored in real time. Drawing on theories of grounding, repair, and sign language interaction, we conceptualize movement coherence as alignment across these parallel streams and describe how breakdowns trigger observable attention shifts and distributed repair across participants, interpreters, and instructors. Across sessions, we identify recurrent coordination strategies including peer checking, freeze and scan, interpreter repositioning, tactile cueing, and pacing adjustment. Our findings provide an empirically grounded account of grounding under attentional constraint in inclusive embodied settings, with implications for sign language interpretation, multimodal discourse, and the design of accessible movement instruction. This paper includes deidentified materials derived from recorded sessions, including selected keyframes, structured interactional annotations, and anonymized instructor and participant survey responses.
A Comparative Analysis of Traditional and Contemporary Visual Features for Computational Annotation of Irish Sign Language
Sarmad Khan | Simon McLoughlin | Irene Murtagh
Sarmad Khan | Simon McLoughlin | Irene Murtagh
Automatic annotation of sign language data is critical for advancing linguistic research and developing sign language technologies, yet it remains a major bottleneck due to the inherently motion-based and multi-modal nature of signing. Irish Sign Language, like many sign languages, presents challenges for computational annotation and sign language processing due to limited annotated corpora and the inherent difficulty of reliably annotating movement, trajectories, and coarticulation across manual and non-manual articulators. This paper presents an automated computational framework for gloss-level annotation support in Irish Sign Language, designed to assist scalable corpus annotation by learning motion-related cues directly from sign language videos. Using ELAN-aligned segments from the Signs of Ireland Corpus, we compare contemporary self-supervised visual representations with traditional pose-based features derived from explicit skeletal tracking, evaluating three feature configurations: DINOv2, MediaPipe, and multi-modal fusion. Our results show that self-supervised visual embeddings achieve the highest average accuracy 86.12%, outperforming both multi-modal fusion 84.28% and pose-based representations 76.74%. This indicates that recent visual models can implicitly encode linguistically relevant motion information, including articulator movement and transitional dynamics, reducing the need for explicit landmark extraction in practical annotation pipelines. Overall, this work provides empirical guidance and a deployable computational framework to support computational annotation and enrichment of sign language corpora.
HeSLEx: A novel online questionnaire for heritage sign language research
Evgeniia Khristoforova | Roman Poryadin
Evgeniia Khristoforova | Roman Poryadin
In this paper, we present Heritage Sign Language Experience (HeSLEx), a novel online sociolinguistic questionnaire adapted from the heritage spoken language survey HeLEx (Tomić et al., 2023) to provide a standardized community profile prior to data collection. Heritage sign languages are minority sign languages used by Deaf signers in migration contexts and thus offer a unique window on bilingualism in the visual modality. HeSLEx is designed to be visual-first: most content is delivered as videos featuring signing in Russian Sign Language (RSL) by community member in a JavaScript/jsPsych interface. To accommodate heterogeneous RSL comprehension, each video includes optional Russian and German text hidden behind a “Show text” button. HeSLEx adds sign-specific modules, including participant and parental hearing status; modality-appropriate proficiency ratings (signing/comprehension for RSL and German Sign Language; reading/writing for Russian and German); educational histories and language(s) of instruction; interactional contexts central to Deaf life (including Deaf clubs); and Deaf-centered identity and language-attitude measures. Many items use slider scales to yield continuous predictors. The tool is designed to be adaptable to other sign language pairs in the framework of heritage language research and beyond.
Comparison of Low Bitrate Quantizers for Encoding Swedish Sign Language
Anna Klezovich | Johanna Mesch | Gustav Eje Henter | Jonas Beskow
Anna Klezovich | Johanna Mesch | Gustav Eje Henter | Jonas Beskow
This paper investigates the bitrate–distortion trade-off of different discrete representations for Swedish Sign Language (STS) using the STS Mocap v1 motion capture dataset. We compare the K-Means algorithm with the Residual Vector Quantized Variational Autoencoder (RQ-VAE) to determine how efficiently each method preserves salient motion information at low bitrates. The results show that RQ-VAE consistently achieves lower reconstruction error than K-Means at matching bitrates, particularly for body motion, and better preserves the signing space volume. We further demonstrate that quantized representations can serve as conditioning for a flow-matching generative model, producing plausible but still imperfect sign sequences at low bitrates. These findings highlight the advantages of vector quantized models for efficient sign language motion encoding.
Exploring Aspects of Spontaneous Signing in the DGS Corpus
Maria Kopf | Reiner Konrad | Gabriele Langer | Marc Schulder | Lutz König
Maria Kopf | Reiner Konrad | Gabriele Langer | Marc Schulder | Lutz König
Most use of sign language is spontaneous, unplanned, embedded in a one-to-one situation and transient. General sign language corpora aim at such naturalistic data. Thus it can be expected that they include phenomena of spontaneous language similar to the ones described for spontaneous speech in vocal languages: that is, (dis)fluencies such as pauses, hesitations, errors, false starts and repairs as well as discourse markers. In this paper we explore which of the known phenomena of spontaneous language from previous research on vocal and sign languages could be identified in the DGS Corpus using the annotations at hand. We describe our search strategies, consider additional annotation tiers for spontaneous language, and provide examples for the phenomena identified.
Two-Handed Signs and Handedness: Phonological Implications for Sign Language Structure
Lisa Lepp | Mirella De Sisto | Dimitar Shterionov
Lisa Lepp | Mirella De Sisto | Dimitar Shterionov
Handedness —the use of one versus two hands in sign production— has traditionally been discussed in relation to dominance and symmetry conditions, yet it remains underrepresented in formal phonological models of sign languages. This paper argues that handedness constitutes a core phonological parameter that directly influences the structure and interaction of movement, handshape, location, and orientation. Building on hierarchical and dependency-based approaches, we propose an adapted phonological dependency model that explicitly integrates handedness in the representation of manual articulators. In one-handed signs, features are specified for a single active hand. In two-handed signs, feature distribution is constrained by symmetry and dominance conditions, which regulate whether the hands must share features or may differ in a structurally restricted way. This structural encoding accounts for variation phenomena such as weak add, weak prop, and weak drop as constrained adjustments within the phonological system. From a technical perspective, this refinement suggests more formal restrictiveness and empirical discriminability within the feature geometries, reduced representational ambiguity, and improved empirical testability across theoretical, corpus-based, and computational implementations, strengthening the interface between phonological theory and sign language technology.
Over the past six decades, a variety of systems have been developed for representing sign language forms, from Stokoe Notation (Stokoe, 1960) to SignWriting (Sutton, 1999) and lexical database schemas. Each was designed with specific goals and applications, leading to a fragmented landscape of representations. To enable greater interoperability and data sharing among sign language users and researchers, we propose a robust approach to translating between notation systems. As a first step in this direction, we introduce a formal mapping framework between HamNoSys and the SL CatForm coding schema, describe its implementation, and present empirical evidence of its performance. An extensive evaluation of mapping mismatches revealed improvements to the mapping logic needed to further advance the HNS2CF mapping tool. However, the initial version of the system already achieves an overall accuracy of 76.7% and an in-depth analysis reveals that many apparent mismatches stem from annotator disagreement rather than mapping errors, indicating that the tool’s actual accuracy is even higher. These results demonstrate the feasibility and promise of establishing mapping mechanisms across sign representation systems.
Emotion Recognition in German Sign Language with Facial Action Units
Cristina Luna Jimenez | Lennart Eing | Sergio Esteban Romero | Tanja Schneeberger | Patrick Gebhard | Fabrizio Nunnari | Elisabeth Andre
Cristina Luna Jimenez | Lennart Eing | Sergio Esteban Romero | Tanja Schneeberger | Patrick Gebhard | Fabrizio Nunnari | Elisabeth Andre
Emotion Recognition research in Sign Languages is still in its infancy. Still today, there exists a lack of knowledge about appropriate annotation guidelines and the impact that facial expressions, body postures and head positions have in recognizing emotions while signing, considering that sign language encompasses manual and non-manual cues with linguistic purposes. In this article, we present an acquisition protocol to record acted emotions in German Sign Language under four scenarios (High-Valence and High-Arousal, High-Valence and Low Arousal, Low-Valence and High-Arousal, and Low-Valence and Low-Arousal). The goal is to provide a reference dataset to explore the use of machine learning techniques for an automated classification of emotions in sign language utterances. As a baseline reference, we trained static models with features extracted from the facial muscle activations. The best model achieved an accuracy of 68.84% and a F1 of 67.96% with a random forest trained on the statistics extracted from Action Units. These results highlight the importance of facial expression in sign language, not only for carrying linguistic information but also for transmitting emotions. Results also indicate challenges in detecting emotions in the High-Valence and Low Arousal scenario, which suggests future investigation lines to explore.
DGS-BIGEKO: A Dataset for Hypothetical Emergency Scenarios in German Sign Language
Cristina Luna Jimenez | Lennart Eing | Daksitha Senel Withanage Don | Marco González | Fabrizio Nunnari | Pamela Perniss | Patrick Gebhard | Elisabeth Andre
Cristina Luna Jimenez | Lennart Eing | Daksitha Senel Withanage Don | Marco González | Fabrizio Nunnari | Pamela Perniss | Patrick Gebhard | Elisabeth Andre
In this article, we describe DGS-BIGEKO, a sign language dataset containing a conversation in a crisis scenario signed by a professional interpreter in German Sign Language (DGS). The dataset comprises 14 sentences with common questions and answers from protocols occurring in emergency call scenarios translated into DGS. Additionally, the dataset contains signs for an additional 108 concepts that are relevant to emergency call scenarios. The dataset is intended to support research in sign language linguistics and sign language machine translation by providing resources in a very specific domain, where no previous resources are available in DGS. The dataset is freely available for research purposes at the following address: https://doi.org/10.5281/zenodo.18458557
Perceptual Validation of 3D Pose, Guided Sign Language Synthesis
Ezekiel Maina | Lilian Wanzare | James Obuhuma
Ezekiel Maina | Lilian Wanzare | James Obuhuma
Sign language corpora face a structural tension between open-access requirements and the irreducible biometric identity embedded in visual, gestural data. While 3D pose estimation enables signer-agnostic abstraction, the representational adequacy of pose-based modeling for preserving linguistic structure remains underexplored. This paper introduces a perceptually-grounded kinematic modeling framework that formalizes 3D landmark sequences as an intermediate linguistic representation and validates their adequacy through avatar-mediated synthesis and large-scale human evaluation. Using 30370 gloss-level Kenyan Sign Language (KSL) segments derived from the AI4KSL corpus, we construct normalized 3D motion trajectories via MediaPipe Holistic. These trajectories are retargeted to parameterized avatars through a constrained kinematic mapping that preserves non-manual marker geometry and articulatory timing. We define a dual evaluation paradigm combining geometric fidelity metrics (PCK=92.7%, OKS=0.88, PCP=91.5%, PDJ>85.3%) with perceptual constructs measured across a statistically powered Deaf participant cohort (N=384). Results demonstrate a strong predictive relationship between structural joint precision and perceived gesture clarity (r=0.76, p<.01), suggesting that linguistic adequacy is partially recoverable from normalized kinematic structure. Furthermore, representational diversity in avatar instantiation significantly increases perceived inclusivity without degrading intelligibility. These findings establish pose-based motion abstraction not merely as an anonymization technique but as a viable corpus-level modeling layer for ethically sustainable language in motion.
The Displacement-Velocity Dissociation in Sign Language Learning: Kinematic Signatures of Event Structure in Novice ÖGS Signers
Evguenia A. Malaia | Julia Krebs | Eric Harbour | Julia Martetschläger | Hermann Schwameder | Dietmar Roehm | Ronnie B. Wilbur
Evguenia A. Malaia | Julia Krebs | Eric Harbour | Julia Martetschläger | Hermann Schwameder | Dietmar Roehm | Ronnie B. Wilbur
This study investigates how adult learners acquire linguistically contrastive movement patterns in Austrian Sign Language (ÖGS), focusing on the telic/atelic distinction predicted by the Event Visibility Hypothesis. Telic verbs (bounded events) are produced by proficient Deaf signers with shorter duration and temporally precise, low-entropy velocity profiles, whereas atelic verbs (unbounded processes) show more continuous motion. Using 3D motion capture (300 Hz), we compared 8 novice learners (6–12 weeks of instruction) with 6 proficient Deaf signers across 71 verbs. Linear mixed-effects models revealed a dissociation between gross movement patterning and fine-grained velocity profile structure in learner productions. Learners correctly reproduced the proportional path-length contrast between telic and atelic verbs, replicating the gross spatial distinction of proficient signers. However, temporal marking of the telic/atelic contrast was underproduced: learners showed a significantly smaller duration difference between verb types than proficient signers, while total path length did not differ significantly between verb types or groups. Temporal control showed significant between-group differences: learners exhibited elevated sample entropy, with non-proficient velocity profiles within individual sign productions, though spatial consistency across trials (STI) was comparable to that of proficient signers. Peak velocity did not differ between groups, suggesting that learners can reach target speeds but cannot yet modulate temporal structure reliably. These findings support distinct learning trajectories for gross movement patterning and fine-grained motion complexity, and demonstrate that velocity profile structure within signs constitutes a core linguistic target in sign language learning.
CEFR-Based Assessment in Sign Languages: The Case of LSE and Perspectives for LIS
Maria Grazia Marrocu
Maria Grazia Marrocu
This study examines the application of the Common European Framework of Reference for Languages (CEFR) to Spanish Sign Language (LSE) in a university context, with reference to the Italian situation (Council of Europe, 2020). In Spain, CEFR descriptors are already integrated into academic programmes for the assessment of LSE, whereas in Italy the context remains uneven due to the lack of shared criteria for the teaching and assessment of Italian Sign Language (LIS). The research project, conducted jointly by Ca’ Foscari University of Venice and Rey Juan Carlos University of Madrid, adopts a longitudinal and comparative design focusing on the first three CEFR proficiency levels (A1, A2, B1) of LSE among L2M2 learners (second language second modality). A mixed-methods approach combining classroom observations, self-assessment instruments, and standardised assessment rubrics is used to analyse the alignment between students’ self-assessments and instructors’ external evaluations, with particular attention to linguistic and metacognitive awareness. The findings show increasing accuracy in self-assessment as proficiency develops, alongside recurring issues such as the overestimation of receptive skills and the underestimation of productive competence. These results highlight the need for targeted assessment interventions and contribute to the development of CEFR-consistent evaluation practices for sign languages.
Improving phonological distance measures for signs: the CatFormCompare tool
Hope E. Morgan | Amy Isard | Anh Dang
Hope E. Morgan | Amy Isard | Anh Dang
This paper describes the CatFormCompare tool, designed to enable the comparison of phonological content between pairs of signs, especially in larger datasets. With this tool and a schema for coding categorical form (the SL CatForm coding schema), a pipeline is created that allows a feedback mechanism for advancing research—specifically by directly addressing one of the hard problems in sign language phonology: how to extract true minimal pairs from datasets coded for categorical form? Solving this problem would simultaneously improve phonological distance measurements for sign languages because it would mean that the units for measuring distance are grounded in the linguistic structure of the language and not simply a by-product of the coding system. Here we report on the tool and the first evaluation of its functioning.
Assisting Corpus Annotation: Automatic BIO-Tagging of Clause-Like Units in Polish Sign Language. A Pilot Study on Corpus Data
Piotr Mostowski | Anna Kuder | Joanna Wójcicka
Piotr Mostowski | Anna Kuder | Joanna Wójcicka
The creation of large-scale sign language corpora is often bottlenecked by the labour-intensive process of multi-layered annotation that requires manual analysis. One of the annotation steps is the challenging and time-consuming task of segmenting continuous signing into clause-like-units (CLUs). In this paper, we propose an automated segmentation framework for Polish Sign Language (PJM) designed to support manual annotation. To detect sentence boundaries, we adapt the Multi-Stage Temporal Convolutional Network (MS-TCN) architecture, enhanced with a Channel Attention mechanism, to effectively fuse multimodal skeleton features (hands, body, and face) extracted via MediaPipe. We evaluate the model on a diverse subset of the PJM Corpus (40 video files, 25 signers), containing nearly 16,000 manually annotated clauses prior to the start of this study. The proposed method achieves a Segmental F1-score of 75.43% at IoU = 0.10 and 57.52% at IoU = 0.50, demonstrating a strong capability in localising sentence boundaries. Furthermore, ablation studies reveal that fusing manual kinematics with non-manual prosodic cues (face) yields a significant performance gain (+13.6 pp) over unimodal baselines, empirically confirming the linguistic necessity of incorporating both manual and non-manual articulators in the process of sentence delimitation. The solution offers a viable means for reducing CLU annotation time by automatically generating high-quality clause boundary proposals.
Introducing VISTA-SL: A Multilingual e-Learning Platform for Deaf and Hearing Learners of Sign Languages
Irene Murtagh | Marc Schulder | Annika Herrmann | Liona Paulus | Julian Bleicken | Kostas Blekos | Athanasios Konstantakopoulos | Klimis Antzakas | Dimitrios Kosmopoulos | Eva Valls | Ricardo Marques | Josep Blat | Konstantinos Karampidis | Ben Elsendoorn
Irene Murtagh | Marc Schulder | Annika Herrmann | Liona Paulus | Julian Bleicken | Kostas Blekos | Athanasios Konstantakopoulos | Klimis Antzakas | Dimitrios Kosmopoulos | Eva Valls | Ricardo Marques | Josep Blat | Konstantinos Karampidis | Ben Elsendoorn
This article introduces the VISTA-SL project, which aims to create an integrated e-learning platform for four European sign languages: German Sign Language, Greek Sign Language, Irish Sign Language, and Dutch Sign Language. Designed as a complement to face-to-face classes, the VISTA-SL platform will combine expertise in sign language education and education technologies to provide an adaptive and interactive learning environment suitable for deaf, hard of hearing and hearing users seeking to learn a sign language, whether it constitutes their first language or not. Building on a co-ordinated curriculum that covers vocabulary, grammar and Deaf culture materials, the platform will provide video material presented by deaf L1 signers, together with games and gamification features to motivate learning, while also providing several assistive technologies. By leveraging cutting edge language processing and computer vision approaches, the platform will provide augmented reality feedback, 3D avatars and an LLM-based virtual instructor, as part of the learning environment. VISTA-SL is developed in collaboration with end-user focus groups, comprising deaf, hard of hearing and hearing individuals. This will serve to ensure that the educational platform aligns with the expectations and needs of its intended users.
Evaluation of Pose Estimation Systems for Sign Language Translation
Catherine O’Brien | Gerard Sant | Mathias Müller | Sarah Ebling
Catherine O’Brien | Gerard Sant | Mathias Müller | Sarah Ebling
Many sign language translation (SLT) systems operate on pose sequences instead of raw video to reduce input dimensionality, improve portability, and partially anonymize signers. The choice of pose estimator is often treated as an implementation detail, with systems defaulting to widely available tools such as MediaPipe Holistic or OpenPose. We present a systematic comparison of pose estimators for pose-based SLT, covering widely used baselines (MediaPipe Holistic, OpenPose) and newer whole-body/high-capacity models (MMPose WholeBody, OpenPifPaf, AlphaPose, SDPose, Sapiens, SMPLest-X). We quantify downstream impact by training a controlled SLT pipeline on RWTH-PHOENIX-Weather 2014 where only the pose representation varies, evaluating with BLEU and BLEURT. To contextualize translation outcomes, we analyze temporal stability, missing hand keypoints, and robustness to occlusion using higher-resolution videos from the Signsuisse dataset. SDPose and Sapiens achieve the best translation performance (BLEU ~11.5), outperforming the common MediaPipe baseline (BLEU ~10). In occlusion cases, Sapiens is correct in all tested instances (15/15), while OpenPifPaf fails in nearly all (1/15) and also yields the weakest translation scores. Estimators that frequently leave out hand keypoints are associated with lower BLEU/BLEURT. We release code that can be used not only to reproduce our experiments, but also considerably lowers the barrier for other researchers to use alternative pose estimators.
Designing a Data Model for a Diachronic Sign Language Database: A Case Study of Nineteenth-Century Bohemian Sources
Lenka Okrouhlíková
Lenka Okrouhlíková
Diachronic research on sign languages is limited by the fragmentary and heterogeneous nature of historical documentation. Eighteenth- and nineteenth-century printed texts and manuscripts contain valuable lexical data, but their descriptions vary in precision, terminology, and representational conventions. This paper proposes a structured data model for a diachronic sign language database designed to systematise such archival materials. The proposed model adopts a multi-layered architecture that separates primary evidence from analytical interpretation, distinguishes attested from inferred sign parameters, applies graded confidence levels, and encodes structural, iconic, and metaphorical properties in parallel layers. Detailed source metadata ensures traceability and explicit representation of uncertainty. The model is illustrated through sign attestations drawn from nineteenth century Bohemian sources. The case study demonstrates that even fragmentary records, most commonly documented in dictionaries and pedagogical materials through written descriptions or illustrations, can be systematically represented within a unified data model suitable for structured comparison and diachronic analysis. The proposed model may also provide a methodological basis for comparable work on other European sign languages.
A Video-Based Reverse Dictionary for Sign Language Using Gesture Similarity
Batyrbek Orazumbekov | Daniyal Bayanov | Aruzhan Kaltay | Anara Sandygulova
Batyrbek Orazumbekov | Daniyal Bayanov | Aruzhan Kaltay | Anara Sandygulova
Sign language recognition systems are usually modeled as classification systems that map gesture videos to pre-defined glosses. But these systems do not allow similarity searches, where a user can search for similar gestures without knowing the corresponding gloss. This paper presents a pose-based video-to-video search framework for isolated signs, which acts as a reverse gesture dictionary. The system employs keypoints on the skeletal structure instead of RGB images. Two architectures are proposed for modeling temporal information: an encoder with self-attention in a Transformer architecture and a Spatial-Temporal Graph Convolutional Network (ST-GCN). The embedding space is optimized using metric learning objectives, including supervised contrastive learning and ArcFace angular margin loss. The performance of the retrieval system is evaluated on the WLASL dataset using ranking metrics like Recall@K and mean Average Precision (mAP). Experiments reveal that the temporal modeling using the Transformer architecture is an improvement over the graph-based modeling approach in the low-shot learning scenario. The attention-based temporal pooling approach further enhances the ranking quality, with the best-performing model achieving an mAP of 0.237 on the WLASL validation set. Cross-dataset evaluation on a 226-label AUTSL dataset reveals non-trivial generalization performance on the unseen dataset, despite training only on the WLASL dataset.
Norwegian Sign Language: Overview of Resources and Experiments with Automatic SignWriting Transcription
Elisabeth Othamar | Yves Scherrer
Elisabeth Othamar | Yves Scherrer
Norwegian Sign Language (NTS) remains an under-resourced sign language despite its official recognition in Norway since 2022. The limited availability of structured, reusable, and publicly accessible datasets continues to hinder both linguistic research and the development of sign language technologies such as recognition and translation systems. This paper presents an overview of existing datasets and potential data sources for NTS, categorizing them by accessibility, format, and suitability for computational research. We further discuss legal, ethical, and practical considerations related to data reuse, including copyright and privacy constraints. In addition, we report on a series of pilot experiments exploring alternative data acquisition strategies, including dictionary videos, SignWriting resources, and broadcast news material. These preliminary experiments explore whether automatic SignWriting transcription can serve as an intermediate representation for NTS, and examine its potential role in sign identification within continuous signing. The aim of this work is both to document ongoing efforts and to support future initiatives toward the sustainable development of NTS resources.
Long-Term Sign Language Data Crowdsourcing Through Collaborative Lexicons
Pierre Poitier | Jérôme Fink | Ariel Basso Madjoukeng | Adelaide Couplet | Margaux Leleu | Benoît Frénay
Pierre Poitier | Jérôme Fink | Ariel Basso Madjoukeng | Adelaide Couplet | Margaux Leleu | Benoît Frénay
While there exists a multitude of different sign languages (SLs) across the world, Deaf communities often lack the digital tools required to document and process their languages. In this work, we introduce Mot-Signe (MOSI), an application designed in close collaboration with actors from the French Belgian Deaf community. Our tool enables users to search for French Belgian Sign Language (LSFB) translations or to propose new ones by recording signs themselves. This crowdsourcing approach facilitates the collection of SL data in the wild, enriching the available documentation on LSFB and proposing an innovative response to the data scarcity issue inherent to sign language processing. To evaluate the sustainability of this community-driven data collection, a longitudinal user study was conducted. Following its public release, MOSI demonstrated significant real-world adoption, enabling the collection of over 3,000 distinct LSFB signs. Notably, MOSI captures highly valuable linguistic variations and specialized vocabulary often absent from traditional corpora.
Effect of Data Augmentation with Multi-View Perspectives of Signers on the DGS-Fabeln-1 Dataset
Fabian Renner | Daksitha Withanage Don | Elisabeth Andre | Cristina Luna Jimenez
Fabian Renner | Daksitha Withanage Don | Elisabeth Andre | Cristina Luna Jimenez
Sign languages constitute the principal form of communication for deaf communities across the globe. Nevertheless, the development of reliable Continuous Sign Language Translation (CSLT) systems is constrained by the lack of sufficient data and models able to handle spatio-temporal information. In this article, we explore the effect of adding multiview perspectives of the signer to the training set as data augmentation using the UniSign framework for the DGS-Fabeln-1 dataset. Our results reveal that increasing dataset size and using multiple camera perspectives significantly improve performance, with the best configurations achieving BLEU-4 scores of 4.20%. These results provide a competitive baseline for the DGS-Fabeln-1 dataset and guidance for further optimizations of CSLT systems.
Lost in Expression: Diagnosing Systemic Challenges with Non-Manual Generalization in Sign Language Understanding Tasks
Dmitriy Sazonov | Sevgi Z. Gurbuz | Evie A. Malaia
Dmitriy Sazonov | Sevgi Z. Gurbuz | Evie A. Malaia
Incorporation of non-manual information is one of the most challenging aspects of Sign Language Understanding (SLU), as these features contribute to the semantic, syntactic, and pragmatic structure of signed communication as a critical feature of compositional meaning at sign, phrase and sentence level. Despite their key linguistic role, non-manuals are often an afterthought in SLU model and dataset design, with many recent models still neglecting to implement non-manual analysis or evaluate how articulators beyond the hands are contributing to the model prediction. In this work, we identify and analyze the challenges relating to recognition of non-manuals and generalization of their linguistic roles encountered by SLU models, offering new explanations for failures to properly model non-manual behavior. We perform a case study on the subtasks of Continuous Sign Language Recognition and Sign Language Translation by applying the Uni-Sign model to Isharah-1000, a Saudi Sign Language dataset. Using controlled partitioning and feature attribution, we further analyze model behavior and failure cases. With this work we hope to set the stage for the creation of diagnostic frameworks for generalization of non-manuals.
The SignBeach Dataset of Dutch Sign Language (NGT) signs
Annika Schiefner | Gomèr Otterspeer | Beyza Sümer | Floris Roelofsen
Annika Schiefner | Gomèr Otterspeer | Beyza Sümer | Floris Roelofsen
This paper presents the SignBeach dataset, including 1401 lexical signs from Dutch Sign Language (NGT). The items in this dataset represent everyday vocabulary appropriate for primary school children and are part of a larger research project, investigating sign learning in a digital environment. Each sign is presented by four deaf signers in a controlled studio environment. For each item, high quality video recordings are available from five synchronised cameras, providing rich multi-view visual input suitable for linguistic analysis and the development of computer vision pipelines. In addition, we provide three types of computational derivatives: keypoint estimates using MediaPipe, handshape estimates using HaMeR, and 3D body reconstructions using SAM 3D Body. Signs are aligned with lexical entries in the NGT Signbank to provide interoperability of the database with other NGT resources. We outline the construction of the dataset and provide information on opportunities for reuse, for example in the context of psycholinguistic studies or in the context of sign language technology. All materials are available for non-commercial reuse under a CC BY-NC 4.0 license.
Comparing Computer Vision Instruments for Eye Blink Analysis
Margaux Susman | Carla Miquel Blasco | Jan Bulla
Margaux Susman | Carla Miquel Blasco | Jan Bulla
We compared four tools for analyzing blink velocity and amplitude, examining how MediaPipe, OpenFace, InsightFace, and 3DDFA compare in terms of blink analysis. Building on previous findings that different tools yield different results (Kuznetsova and Kimmelman, 2024), we explored their fixed-effect estimates across linguistic versus non-linguistic blinks, within non-linguistic blinks (eye watering blinks versus gaze-direction-change blinks), and within linguistic blinks (prosodic/turn-taking blinks, sign-aligned/list-marking blinks and backchanneling blinks), while controlling for head pose (Pitch, Roll, Yaw). Using mixed-effects linear models on annotated French Sign Language data, we found tool-specific patterns: consistent negative effects for InsightFace and MediaPipe, but positive. effects for 3DDFA. In addition, the influence of head pose varied across models (Pitch is strongly positive in MediaPipe but negative in InsightFace and some 3DDFA models; Roll and Yaw also switch importance across tools). These discrepancies highlight methodological biases that can distort linguistic interpretations.
Grounding Sign Language Representation Learning in Phonology
Toon Vandendriessche | Mathieu De Coster | Joni Dambre
Toon Vandendriessche | Mathieu De Coster | Joni Dambre
Sign language recognition systems are commonly trained using gloss-level supervision, treating signs as holistic lexical units. While effective for classification, such approaches entangle sub-lexical structure and fail to capture the phonological parameters that govern sign formation, limiting interpretability, robustness, and cross-lingual transfer. In this work, we propose a phonologically informed representation learning architecture that explicitly structures the latent space according to linguistic principles. Grounded in the Dependency Model – a phonological model used to describe Flemish Sign Language (VGT) – our hierarchical architecture disentangles parameter-specific subspaces for handshape and location and is trained with multi-label phoneme supervision. To evaluate whether phonological information is directly encoded in the geometry of the embedding space, we introduce a non-parametric probing method that measures neighbourhood consistency across increasing scales. We show that conventional gloss-based networks achieve reasonable performance only for very small neighbourhoods, reflecting incidental visual similarity. In contrast, our disentangled representations maintain stable performance for larger neighbourhoods. This behaviour indicates that phonological structure is preserved across broader regions of the space, yielding more coherent and robust embeddings. Together, our results show that explicit phonological supervision – and crucially, disentangled representation learning – provides a principled foundation for interpretable and transferable sign language representations. Keywords: Sign Language, Machine Learning
Towards Integrating Pose Estimation with Neuroimaging for the Analysis of Signed Language Video Stimuli
Sébastien Vandenitte | Doris Hernández | Jarkko Keränen | Tommi Jantunen | Anna Puupponen
Sébastien Vandenitte | Doris Hernández | Jarkko Keränen | Tommi Jantunen | Anna Puupponen
We present our project revisiting the video stimuli of an EEG study in Finnish Sign Language to ask whether kinematic properties of the videos impacted their processing by study participants. For each stimulus, an average measure of brain responses across participants is computed. To analyse movement properties in the video stimuli, we rely on MediaPipe for pose estimation. We subsequently report on our project to perform an exploratory analysis of the kinematic properties of the videos which may affect their processing. We focus on several landmarks: the signer’s right and left wrists, nose, and upper torso. Our goal is to obtain a kinematic profile of each stimulus video using several average kinematic variables: velocity and acceleration for all selected landmarks, distance between the wrists, and surface covered by the triangular area defined by the left hand, the right hand, and the nose. We conclude by discussing the potential benefits and limitations of this methodological approach.
KWIC view on Constructed Action (CA) and its Collocates in German Sign Language (DGS) – Possibilities and Limitations
Sabrina Wähl
Sabrina Wähl
Constructed action (CA) is a phenomenon that is used in signed discourse to show the actions of a referent (cf. Cormier et al., 2015; for DGS, cf. Fischer and Kollien, 2010). To achieve this, the signer adopts the role of the referent. Most studies use retellings as their data base (e.g. Herrmann and Pendzich, 2018; Cormier et al., 2015). Consequently, there is less research on CA and its use in data that is not influenced by stimuli. Though there is a considerable number of studies on CA, the phenomenon is still not well understood. One possible way to understand this multifaceted phenomenon better is to analyse collocations in conversations. In spoken language lexicography concordance lines – also known as keyword in context (KWIC) – have proven to be a useful tool in the analysis of collocations. The data used in this study are Free conversations in the Public DGS Corpus. This paper explores the possibilities and limitations of concordance lines as a tool to analyse collocational behaviour of CA. It also presents preliminary results regarding CA and its collocates, which may be explored further in the future.
Beyond BLEU: Linguistic Invisibility and Interactional Repair Sequence in End-to-End Sign Language Translation
Zirui Wang | Mayumi Bono
Zirui Wang | Mayumi Bono
Recent advances in end-to-end sign language translation (SLT) have achieved benchmark performance, yet little is known about whether these systems preserve the multi-channel linguistic structures that are essential for real-world communication. We argue that current optimization and evaluation practices create a form of linguistic invisibility, where interactionally decisive non-manual signals (NMS) are systematically underrepresented despite high translation scores.To empirically examine this issue, we analyze an interactional repair sequence from a Japanese Sign Language (JSL) conversational corpus as a diagnostic probe. Combining qualitative interactional analysis with kinematic measurements, we demonstrate a consistent manual–mouth decoupling pattern in which semantic resolution is carried primarily by mouthing while manual articulation remains largely constant. We show that such cross-channel contrast is unlikely to be preserved under current end-to-end training objectives that prioritize global motion similarity. Based on these findings, we argue that progress in SLT should be evaluated not only by sequence-level accuracy but also by the preservation of linguistically contrastive structures, motivating the development of diagnostic, multi-channel evaluation protocols for future SLT benchmarks. We therefore propose incorporating multi-channel diagnostic evaluation sets and decoupling-sensitive metrics into future SLT benchmarking frameworks, providing a pathway toward models that achieve both high performance and linguistic structural visibility.
Continuous Sign Language Recognition using Multimodal Input and Handshape-aware Boundary Detection
Mingyu Zhao | Zhanfu Yang | Yang Zhou | Zhaoyang Xia | Can Jin | Xiaoxiao He | Shuhang Lin | Carol Neidle | Dimitri Metaxas
Mingyu Zhao | Zhanfu Yang | Yang Zhou | Zhaoyang Xia | Can Jin | Xiaoxiao He | Shuhang Lin | Carol Neidle | Dimitri Metaxas
This paper employs a multimodal approach for continuous sign recognition by first using ML for detecting the start and end frames of signs in videos of American Sign Language (ASL) sentences, and then by recognizing the segmented signs. For improved robustness, we use 3D skeletal features extracted from sign language videos to take into account the convergence of sign properties and their dynamics that tend to cluster at sign boundaries. Another focus of this paper is the incorporation of information from 3D hand configuration for boundary detection. To detect handshapes normally expected at the beginning and end of signs, we pretrain a handshape classifier for detection of 87 linguistically defined canonical handshape categories using a dataset that we created by integrating and normalizing several existing datasets. A multimodal fusion module is then used to unify the pretrained sign video segmentation framework and handshape classification models. Finally, the estimated boundaries are used for sign recognition, where the recognition model is trained on a large database containing both citation-form isolated signs and signs pre-segmented (based on manual annotations) from continuous signing—as such signs often differ a bit in certain respects. We evaluate our method on the ASLLRP corpus and demonstrate significant improvements over previous work.
up
Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE)
Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE)
Erhard Hinrichs | Joakim Nivre | Petya Osenova | James Pustejovsky | Claus Zinn
Erhard Hinrichs | Joakim Nivre | Petya Osenova | James Pustejovsky | Claus Zinn
Degrees of Subjectivity and Their Repercussions in Conversation. The View from Online Interactions
Gonzalo Freijedo Aduna | Anastasia Giannakidou | Alda Mari
Gonzalo Freijedo Aduna | Anastasia Giannakidou | Alda Mari
We present the Annotated Reddit Conversation Corpus (ARCC), an English-language dataset of online discussions annotated for Speech Acts and Functional Dependence Relations, designed to investigate how varying degrees of subjectivity influence conversational dynamics and interaction patterns. At the speech act level, we distinguish factual from opinion statements and further classify opinions along a five-degree scale of subjectivity. Functional Dependence Relations capture how segments relate to preceding ones. Analyses show that opinion-discussion contexts feature frequent inter-subjective opinions eliciting explicit agreement and disagreement, while information-exchange contexts exhibit less subjective opinions with responses like answers or requests for clarification. We further demonstrate that a transformer model can predict the subjectivity scale with promising performance. The corpus and annotation guidelines are made available to support future research on opinion expression and automated dialogue analysis.
Constraints on Linking Element Choice in German Nominal Compounding: A Large-Scale Corpus Study
Maksim Shmalts
Maksim Shmalts
The N+N compound class is the largest and the most productive class of compounds in German. A significant number of N+N compounds insert a so-called linking element from a large inventory. The linker choice is notoriously irregular; instead of rules, it is governed by a set of constraints that can only limit this choice based on morphological, phonological, sometimes semantic and lexical properties of the first constituent. While constraints on linking element choice in German nominal compounding are extensively researched and well-documented, no large-scale corpus study has ever been reported on the subject of their empirical application. The present work aims at filling in this gap by conducting an extensive corpus study on potential and actual applicability of these constraints. The study summarizes 64 constraints collected from the relevant literature and obtains applicability statistics for 39 of them over 280k+ German N+N compounds. The study both confirms most of the evidence from previous literature and suggests novel evidence on German nominal compounding. It additionally highlights the importance of structured linguistic data for large-scale empirical studies.
Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian
Giuseppe Samo | Paola Merlo
Giuseppe Samo | Paola Merlo
This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (BLMs), structured datasets designed to probe linguistic knowledge of underlying patterns across sentence sets. We compare structured templates instantiated with natural sentences extracted from Universal Dependencies to structured templates of synthetic sentences. Experiments show that while models achieve ceiling performance when trained and tested on synthetic datasets, they do not reliably generalize to natural sentences. In contrast, models trained on natural data exhibit robust performance across both natural and synthetic test suites, demonstrating their superior ability to capture abstract linguistic patterns. These results corroborate the value of natural data and of structured set ups in linguistic evaluation for probing LLMs’ syntactic and semantic knowledge.
Subevent Structure as a Predictor of Entity Identity Change in Procedural Text
Kyeongmin Rim | James Pustejovsky
Kyeongmin Rim | James Pustejovsky
We test whether the subevent structure encoded in VerbNet-GL predicts entity identity change in procedural text, using only the verb’s lexical specification and no training data. From each VN verb class’s SEMANTICS block we extract an aspectual classification and an I/O count, yielding a predicted dynamic event topology (DET). On the observation side, ten large language models (LLMs) annotate per-entity dynamic object mode (DOM) labels over ∼3100 OpenPI steps, from which we derive observed DET for comparison. The VN-only predictor achieves 67.4% precision (F1 = 0.35) for transformation, showing that formal subevent structure carries genuine predictive signal only for the most common topology, but low performance for other topologies. For events VN predicts as having no result state, only 24% are confirmed as no-change by silver, indicating that the remaining outcomes arise from the argument side of the composition. These results provide empirical evidence that event semantics is distributed across predicate and argument: the VN supplies the subeventual skeleton, but is not sufficient to determine the final outcome.
In this paper we introduce an example-based method for exploring dependency treebanks that is based on principles of vector symbolic architectures. It leverages key properties of this framework to provide fast and flexible search capabilities, since all combinations of query parameters can be compared with a given parse tree in parallel via a single vector operation. The framework also allows for graded similarity and the natural integration of various kinds of information, such as word embeddings. After some background on the framework and an explanation of our implementation, we provide a few examples of the system’s output and draw comparisons to similar applications.
Modular Neural Machine Translation with a Semantic Pivot - Pilot Study Using AMR
Wenyang Gao | Yaxuan Li | Yunxin Bao | Shulin Huang | Yue Zhang
Wenyang Gao | Yaxuan Li | Yunxin Bao | Shulin Huang | Yue Zhang
Neural machine translation (NMT) has become the predominant approach for automated translation, yet conventional models trained on extensive bilingual datasets exhibit critical limitations, including quadratic scaling of training data, sensitivity to out-of-distribution inputs, and a lack of interpretability. Inspired by the classical “translation pyramid” concept, which advocates for translation via a semantic pivot (interlingua), this work explores the integration of Abstract Meaning Representation (AMR) as a structured semantic intermediary to decouple translation into comprehension (source-to-AMR) and generation (AMR-to-target) phases. We conduct a pilot study using a strong AMR parser to create a multilingual silver-standard AMR corpus from the United Nations Parallel Corpus, training modular semantic understanding and generation components for each language. Experimental results demonstrate that our approach achieves an average improvement of 3% in robustness and over 15% in generalization compared to traditional Seq2Seq baselines. Analysis suggests that enhancing semantic parsing and generation accuracy could bridge the gap to conventional NMT systems. To our knowledge, this is the first work to integrate AMR as a semantic pivot in NMT, offering enhanced transparency, scalability, and robustness. This study underscores the potential of semantic-driven translation frameworks and provides a foundation for future research in interpretable, resource-efficient multilingual systems.
We introduce Gutenberg+, a temporally more faithful version of the Project Gutenberg (PG) corpus, one of the most widely used resources for diachronic text analysis. Despite its popularity, the PG corpus contains a major yet overlooked flaw: around 15% of its entries are collections (e.g., anthologies of books, letters, or poems) rather than atomic works, which distorts temporal analyses since such collections may span multiple decades. We present an automatic method to detect and split these collections into their constituent works, producing a finer-grained and temporally consistent corpus. We further re-annotate publication years using LLM-based retrieval-augmented generative methods, demonstrating the potential of LLMs to enhance structured linguistic resources. To illustrate the utility of Gutenberg+, we conduct a small-scale diachronic case study on negation, showing that our refined corpus captures more nuanced cross-linguistic variation than the original PG data. Finally, we release the corpus in UIMA format with full metadata and linguistic annotations, providing a standardized resource for future research on diachronic language change.
We report on a new lemmatisation system for Norwegian, which is a particularly challenging language with two written standards, Bokmål and Nynorsk, that both have a lot of optionality. Our system covers both varieties and consists of a neural model that classifies words into rewrite rule classes that produce their lemma, as well as a large-scale computational lexicon of Norwegian that gives all possible inflections of a large part of the Norwegian vocabulary. We test different ways of combining these components. When evaluated with pure string-matching against the lemmas in the gold data, all systems perform approximately at the same level (99.1-99.2% on Bokmål and 98.5-98.6% on Nynorsk), but detailed error analysis shows that the computational lexicon reduces the number of true errors by more than half (reaching 99.6% accuracy on Bokmål and 99.3% on Nynorsk), as opposed to “surface errors” like using a different, but equally acceptable spelling variant of the correct lemma.
Quantification is common in text, but it is underrepresented in AMR. Quantification also stresses the common conjunctive interpretation of AMR graphs, since universal quantification introduces scope-taking structure and variable binding that cannot be captured as a flat list of conjuncts. We propose an enriched AMR that supports quantificational meaning while keeping AMR’s graph backbone. At the predicate level, we add QuantML features, such as domain restriction, determinacy, distributivity, and involvement. At the discourse level, we add contextual constraints that encode scope and other discourse-sensitive conditions. The two levels follow the UMR architecture and are linked by shared identifiers. We map the enriched graphs to two-block logical forms: a minimal model of events and participants, plus a constraint block that relates them.
Improving Slovene Language Models for Lexicographic Question Answering through Continued Pretraining and Instruction Fine-Tuning
Timotej Knez | Slavko Zitnik
Timotej Knez | Slavko Zitnik
This paper presents a two-stage training approach to improve the performance of Slovene large language models on lexicographic question-answering tasks. We developed a comprehensive lexical pretraining corpus containing 356,294 Slovene word entries. We constructed the corpus by converting structured data from multiple lexicographic sources into markdown format. Additionally, we created a question-answering dataset with 10,485 QA pairs from diverse sources, including automatically generated questions, a linguistic advisory portal, and community forums. Using the Slovenian GaMS model (based on Gemma 2 9B) and GaMS 3 model (based on Gemma 3 12B), we performed continued pretraining on the lexical corpus, followed by instruction fine-tuning with our QA dataset combined with translated general-domain questions. We compared results to different model configurations. Our results demonstrate significant improvements (text similarity increasing from 0.226 to 0.542, BERTScore F1 of 0.915) in answering Slovene lexicographic questions, validating the effectiveness of domain-specific continued pretraining for low-resource languages.
Structured Partial Predictability in Non-Concatenative Morphology: The Case of Tashlhiyt Berber
John Alderete | Hamza Sellami
John Alderete | Hamza Sellami
Non-concatenative morphology poses a persistent challenge for NLP, yet structured quantitative resources for Amazigh (Berber) languages remain scarce. We present the first large-scale computational study of Tashlhiyt Berber plural formation, drawing on a richly annotated dataset of 1,185 noun paradigms with phonological, morphological and semantic features. We decompose the plural system into macro-level word-formation strategies and micro-level stem mutations, and evaluate predictability across ten target domains using linguistic feature models, N-gram baselines, and Bi-LSTM neural models. Results reveal a structured split: linguistic features decisively outperform neural models on systematic macro-level strategies (e.g., +44.5pp F1), while Bi-LSTMs better capture lexically idiosyncratic patterns. Rather than supporting a categorical rule/memory divide, this complementarity reveals gradient layers of regularity within a single morphological system. These findings demonstrate the value of linguistically informed annotation for probing morphological complexity in low-resource, typologically diverse languages. All data, code, and models are publicly available.
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
Matej Klemen | Tjaša Arčon | Luka Terčon | Marko Robnik-Sikonja | Kaja Dobrovoljc
Matej Klemen | Tjaša Arčon | Luka Terčon | Marko Robnik-Sikonja | Kaja Dobrovoljc
Empirical grammar research has become increasingly data-driven, but the systematic analysis of annotated corpora still requires substantial methodological and technical effort. We explore how agentic large language models (LLMs) can streamline this process by reasoning over annotated corpora and producing interpretable, data-grounded answers to linguistic questions. We introduce an agentic framework for corpus-grounded grammatical analysis that integrates concepts such as natural-language task interpretation, code generation, and data-driven reasoning. As a proof of concept, we apply it to Universal Dependencies (UD) corpora, testing it on multilingual grammatical tasks inspired by the World Atlas of Language Structures (WALS). The evaluation spans 13 word-order features and over 170 languages, assessing system performance across three complementary dimensions – dominant-order accuracy, order-coverage completeness, and distributional fidelity – which reflect how well the system generalizes, identifies, and quantifies word-order variations. The results demonstrate the feasibility of combining LLM reasoning with structured linguistic data, offering a first step toward interpretable, scalable automation of corpus-based grammatical inquiry.
The L2 Network: A CEFR-Aligned Knowledge Graph for Grammar Domain Modeling
Luisa Ribeiro-Flucht | Xiaobin Chen
Luisa Ribeiro-Flucht | Xiaobin Chen
Large language models have renewed interest in the role of structured linguistic data for applications that require controllable, interpretable, and pedagogically aligned language generation. This need is especially visible in intelligent language tutoring, where grammar cannot be modeled as a flat inventory of patterns alone, but must also capture their relations and functions they realize. We present the L2 Network, a machine-readable knowledge graph of CEFR A1-A2 English grammar that encodes formal patterns, functions, and typed relations between them. The resource is grounded in established pedagogical reference materials, combining form inventory and progression information from the English Grammar Profile with a functional layer derived from CEFR descriptors. We further report content validation of the form-function mappings through expert annotation, including agreement analysis and a consensus-filtered core release. The resulting graph provides an explicit schema for representing pedagogically relevant grammatical knowledge and supports downstream uses such as learner modeling, adaptive task selection, and controlled generation in dialogue-based ICALL systems.
Paraphrase Acquisition via Bilingual Pivoting Based on Neural Word Alignment
Risa Kondo | Seiji Sugiyama | Tomoyuki Kajiwara | Takashi Ninomiya
Risa Kondo | Seiji Sugiyama | Tomoyuki Kajiwara | Takashi Ninomiya
We utilize neural word alignment to improve the quality of paraphrase databases in English and Japanese. For large-scale paraphrase acquisition, previous studies have employed a framework of bilingual pivoting based on word alignment on bilingual parallel corpora. Naturally, the quality of paraphrases acquired by bilingual pivoting depends on the performance of word alignment. Previous studies based on statistical word alignment have limitations in the quality of acquired paraphrases because they do not consider word meaning. This study employs a more sophisticated neural approach for word alignment in bilingual pivoting to enhance the quality of paraphrase acquisition. Experimental results revealed that our paraphrase databases outperformed existing ones in both internal and external evaluations.
Using syntax for the semantic representation of sentences
Iskandar Boucharenc | Eve Sauvage | Thomas Gerald | Julien Tourille | Sabrina Campano | Cyril Grouin | Sophie Rosset
Iskandar Boucharenc | Eve Sauvage | Thomas Gerald | Julien Tourille | Sabrina Campano | Cyril Grouin | Sophie Rosset
Deep learning methods in natural language processing often rely on statistical methods to tokenize texts before vectorization. This segmentation produces lexical subunits offering great flexibility. However, the reuse of identical tokens across words with different meanings can favor representations based on surface form rather than on linguistic information, especially semantics. This mismatch between semantics and surface form can lead to undesirable effects in language processing. To limit the influence of form on the semantics of vector representations, we propose an intermediate representation based on syntactic parsing that is more compact and more faithful to word meaning.
Modeling Word-Internal Structures: Morphological Segmentation Across 58 Languages
Vojtěch John | Zdeněk Žabokrtský | Benjamin Reeves
Vojtěch John | Zdeněk Žabokrtský | Benjamin Reeves
We present the largest multilingual experiment to date on word-to-morph segmentation, covering 58 typologically diverse languages. We describe a newly compiled collection of linguistically annotated resources for the task, providing broad coverage and enabling systematic cross-lingual evaluation. Second, we train two neural models on surface morphological segmentation, achieving 81% average word accuracy on the original datasets, slightly outperforming previous methods. Experiments on custom test sets reveal substantial variation in performance, highlighting the need for further harmonization and more robust multilingual approaches.
DiNoS: Creating a Data-Driven German Noun Phrase Lexicon from Universal Dependencies
Jacob Lee Suchardt | Ronja Laarmann-Quante
Jacob Lee Suchardt | Ronja Laarmann-Quante
To foster investigations of noun phrase (NP) inflection in German at scale, this paper introduces DiNoS (Distributional Noun Structure), a data-driven lexicon of NP heads, which includes statistical information on the dependents and the morphosyntactic features of their original in-context appearances. We make available the source code for the extraction of NPs from CoNLL-U treebanks, which includes rule-based heuristics to improve feature annotation coverage and ensures a homogeneous lemmatisation strategy across treebanks. While the resulting JSON-based lexicon is suitable for no-code interaction for non-experts, it is further supported by a toolkit for the automatic calculation of, and access to, various statistical overviews. In this paper, we present the heuristics employed to extract NP datasets from the German Universal Dependencies’ Hamburg Dependency and GSD treebanks. In addition, we provide a preview of the emerging DiNoS lexica’s properties and discuss some implications of noun and determiner word form ambiguity for NP complexity.
Figurative, Polysemous, Conventional: Designing a Dataset of Regular Metaphor
Anna Temerko | Pablo Gamallo | Marcos Garcia
Anna Temerko | Pablo Gamallo | Marcos Garcia
Metaphor, a figure of speech and a cognitive device, offers a powerful way to explain one conceptual domain in terms of another. Particularly successful metaphorical mappings are conventionalized through frequent use and lose their creative quality. They become sense extensions of polysemous words. Our dataset project captures such metaphors with ten regular polysemy patterns that manifest repetitively in the meaning structures of English words. Regular metaphor, unlike its counterpart regular metonymy, has not previously received a dedicated dataset, and we intend to close this gap. The dataset under construction features naturalistic sentences extracted from a general language corpus and is manually annotated with sense labels for metaphorically extended polysemes. Its intended use is to support linguistic, cognitive, and computational investigations into patterns of meaning in polysemy, while accounting for its complexity, regularity, continuity, and heterogeneity. We see neural language models as an excellent experimental ground for such research because they are able to show both distributional (continuous) and symbolic (discrete) behavior in language processing and representation. In this paper, we reflect on how these systems tally.
Domain-Specific Considerations in the Preparation of Specialized Corpora: A Case Study on a Corpus of German Sermons
Cora Haiber | Adam Roussel | Stefanie Dipper
Cora Haiber | Adam Roussel | Stefanie Dipper
We present a new corpus of contemporary German sermons and describe the steps taken in its preparation. We apply a semi-automatic approach to sentence segmentation, tokenization, and lemmatization, utilizing annotation guidelines that are specialized to this domain. In the process of preparing these data, we find that state-of-the-art tools for these tasks still make problematic errors, especially with non-standard data, despite apparently very high performance on common benchmarks. We obtain test scores of F1 = 96.69 % for sentence segmentation, F1 = 99.99 % for tokenization, and acc = 64.00 % for lemmatization with our domain-adapted models and show that domain-adaptation improves performance over state-of-the-art models for the token and sentence segmentation tasks.
Semantic, Syntactic, Lexical: What Makes QA Augmentation Work in Limited Quantity?
Benedictus Kent Rachmat | Thomas Gerald | Takuya Nakamura | Zheng Zhang | Cyril Grouin
Benedictus Kent Rachmat | Thomas Gerald | Takuya Nakamura | Zheng Zhang | Cyril Grouin
Data augmentation is a common fix in domains where training data is scarce or difficult to collect, such as specialized medical or any other domain specific applications. In question answering (QA), most studies report headline accuracy while saying little about the quality of the synthetic data. Here, quality goes beyond fluent rewording: augmented items must remain faithful to the supporting evidence and preserve the original answerability. We study three augmentation families lexical, syntactic, and semantic edits generated with LLaMA 3.1 70B, and analyze how these edits affect model behavior. To mirror low-resource settings, we focus on subsets of SQuADv2 (general) and PubMedQA (biomedical, domain specific). We report Exact Match (EM)/F1 alongside quality diagnostics, yielding a fuller picture than accuracy alone. Our results show that augmentation behaves differently across domains and scales. In SQuADv2, augmented variants maintain performance on par with baselines, showing that added diversity mostly does not harm model quality, whereas in PubMedQA semantic edits bring improvements under extreme scarcity and support stronger performance as supervision grows.
Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations
Anna Nedoluzhko | Sarka Zikanova | Jiri Mirovsky | Milan Straka | Eva Hajicova
Anna Nedoluzhko | Sarka Zikanova | Jiri Mirovsky | Milan Straka | Eva Hajicova
As previous research on annotator disagreement in discourse phenomena has shown, understanding text coherence varies considerably from one individual to another. To explore this phenomenon, we created two corpora with multiple annotations of Czech texts, accompanied by annotators’ explanations of their choices. The first corpus consists of 1,024 contexts annotated in parallel by three annotators. It captures differences in the identification of coreference across various text types and grammatical-semantic categories, including pronouns, full noun phrases, and anaphoric adverbials. The second corpus comprises 512 contexts, annotated in parallel by five annotators, and focuses on identifying discourse relations in attributive and non-attributive constructions. Both corpora achieve a comparable inter-annotator agreement of approximately 60–65%. For coreference annotation, agreement tends to be lower in cases where automatic coreference resolution models disagree, suggesting that when the models disagree, the examples tend to be more difficult or ambiguous for human annotators to interpret. The annotators’ comments, both for coreference and discourse relations, further reveal differences in interpretation, varying levels of confidence in text understanding, and individual reading strategies.
up
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Nina Hosseini-Kivanani | Alessio Brutti | Marco Matassoni | Sandipana Dowerah | Davide Liga | Christoph Schommer
Nina Hosseini-Kivanani | Alessio Brutti | Marco Matassoni | Sandipana Dowerah | Davide Liga | Christoph Schommer
Adapting Foundational ASR Models to Efik: An Empirical Study of an Extremely Low-Resource Tonal Language
Offiong Bassey Edet | Stephen Orok Duke | Enoima Essien Umoh | Benjamin Okon Nyong | Andrew Asuquo Nkpanam
Offiong Bassey Edet | Stephen Orok Duke | Enoima Essien Umoh | Benjamin Okon Nyong | Andrew Asuquo Nkpanam
Automatic Speech Recognition (ASR) has significantly transformed human-computer-interaction and natural language processing. However, many African spoken languages, including Efik, remain severely underrepresented in ASR research. This paper investigates the adoption of state-of-the-art foundational ASR models such as XLS-R and Whisper through fine-tuning for Efik, a low-resource tonal language and empirically evaluates their performance. We curate a 3-hour Efik speech dataset and conduct a comparative evaluation using standard ASR metrics. We further augmented the XLS-R CTC model with a 3-gram KenLM language model trained on an Efik text corpus. Experimental results show that XLS-R-300M + KenLM achieves a word error rate (WER) of 10.86% and a character error rate (CER) of 3.16%, substantially outperforming both the baseline XLS-R (WER: 29.2%, CER: 6.4%) and Whisper across noisy and multi-speaker conditions. These findings suggest that lightweight CTC models augmented with language model integration offer a more robust and practical approach for extremely low-resource tonal languages than larger sequence-to-sequence models.
PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions
Sicheng Jin | Dipankar Srirag | Aditya Joshi
Sicheng Jin | Dipankar Srirag | Aditya Joshi
While modern Automatic Speech Recognition (ASR) systems achieve high accuracy on benchmark corpora, their performance often degrades when there is real-world variability. This work focuses on variability arising due to accented, spontaneous, and domain-specific speech. In particular, we introduce PAper REading DAtaset (PAREDA), a first-of-its-kind multi-accent speech dataset consisting of discussions on academic Natural Language Processing (NLP) papers between speakers with Australian, Indian-English, and Chinese English accents. Each session elicits a spontaneous monologue (a summary of a paper’s abstract) and a non-monologue (a question-and-answer session between participants), resulting in a corpus rich with technical jargon and conversational phenomena. We evaluate the performance of SOTA ASR models on PAREDA, analysing the impact of accent mixing and increased speech rate. Our results show that, in the zero-shot setting, models perform worse, confirming the dataset’s challenging nature. However, fine-tuning on PAREDA significantly reduces the Word Error Rate (WER), demonstrating that our dataset captures linguistic characteristics often missing from existing corpora. PAREDA serves as a valuable new resource for building and evaluating more robust and inclusive ASR systems for specialised, real-world applications.
Say Again? The Limits of Whisper with Conversation. A Case Study on the KIParla Corpus.
Martina Simonotti | Ludovica Pannitto | Caterina Mauri | Adriano Ferraresi | Gabriele Carioli
Martina Simonotti | Ludovica Pannitto | Caterina Mauri | Adriano Ferraresi | Gabriele Carioli
This study investigates how Whisper handles interactional phenomena in spontaneous Italian conversation, focusing on backchannels, repairs, and filled pauses. We compare standard Word Error Rate (WER) optimization with a decoding strategy that explicitly rewards the preservation of interactional events. Results show that decoding choices have limited impact on overall accuracy, while recognition remains strongly phenomenon-dependent, suggesting structural limitations in the handling of interactional phenomena, with systematic linearization of repairs and frequent suppression of short conversational items.
Word Error Rate (WER) remains the standard metric in automatic speech recognition (ASR) evaluation, yet it does not capture higher-level linguistic distinctions such as prosody. This article examines how three state-of-the-art open-source ASR models (Whisper, Meta’s MMS, and GigaAM) handle the distinction between Russian polar questions and assertions. Russian is particularly suitable for this investigation because polar questions can be marked either morphologically (li, razve) or purely intonationally, without changes in word order. Using audio stimuli from a controlled psycholinguistic experiment, I compare human classification performance in two experimental studies with ASR transcriptions, taking sentence-final punctuation as a proxy for prosodic interpretation. While human participants show near-ceiling accuracy, the ASR models perform inconsistently, especially on intonationally marked questions. Additional contextual cues improve performance in some cases but also reveal instability across conditions. The results demonstrate that evaluating punctuation provides insights beyond WER and allows a more fine-grained view of how current ASR systems encode prosodic and grammatical information.
Quantizing Whisper: How Design Choices Affect ASR Performance
Arthur Söhler | Julian Irigoyen | Andreas Søeborg Kirkedal
Arthur Söhler | Julian Irigoyen | Andreas Søeborg Kirkedal
Large speech recognition models like OpenAI’s Whisper achieve high accuracy but are difficult to deploy in resource-constrained environments due to their high memory and computational demands. This matters for low-resource and on-device settings, where compute and memory constraints often limit the practical use and evaluation of ASR systems. To address this, we present a unified, cross-library evaluation of post-training quantization (PTQ) on Whisper-small, comparing supported configurations across quantization scheme, method, granularity, and bit-width. Our study is based on four libraries—PyTorch, Optimum-Quanto, HQQ, and bitsandbytes. Experiments on LibriSpeech test-clean and test-other show that dynamic int8 quantization with Optimum-Quanto offers the best trade-off, reducing model size by 57% while lowering Word Error Rate below the baseline. Additional experiments on Whisper-base and Whisper-tiny confirm these trends, though with more pronounced degradation at lower bit-widths. Static quantization performed worse, likely due to the absence of efficient low-bit implementations for operations such as LayerNorm and Softmax. More aggressive formats (e.g., nf4, int3) achieved up to 71% compression at the cost of accuracy in acoustically challenging conditions. Our results demonstrate that carefully chosen PTQ methods can substantially reduce model size and inference cost without retraining, enabling efficient deployment of Whisper on constrained hardware.
"OK Aura, Be Fair with Me": Demographics-Agnostic Training for Bias Mitigation in Wake-up Word Detection
Fernando López | Paula Delgado-Santos | Pablo Gómez | David Solans | Jordi Luque
Fernando López | Paula Delgado-Santos | Pablo Gómez | David Solans | Jordi Luque
Voice-based interfaces are widely used; however, achieving fair Wake-up Word detection across diverse speaker populations remains a critical challenge due to persistent demographic biases. This study evaluates the effectiveness of demographics-agnostic training techniques in mitigating performance disparities among speakers of varying sex, age, and accent. We utilize the OK Aura database for our experiments, employing a training methodology that excludes demographic labels, which are reserved for evaluation purposes. We explore (i) data augmentation techniques to enhance model generalization and (ii) Knowledge Distillation of pre-trained foundational speech models. The experimental results indicate that these demographics-agnostic training techniques markedly reduce demographic bias, leading to a more equitable performance profile across different speaker groups. Specifically, one of the evaluated techniques achieves a Predictive Disparity reduction of 39.94% for sex, 83.65% for age, and 40.48% for accent when compared to the baseline. This study highlights the effectiveness of label-agnostic methodologies in fostering fairness in Wake-up Word detection.
Scalable Expansion of Multilingual Speech LLMs for ASR: A Continual Learning Approach
Lorenzo Concina | Marco Matassoni | Alessio Brutti
Lorenzo Concina | Marco Matassoni | Alessio Brutti
Speech Large Language Models have recently enabled the processing of spoken language by coupling powerful language models (LLMs) with pre-trained speech encoders. However, their multilingual scalability remains limited, particularly for low - resource and unseen languages, while naïve fine- tuning often triggers catastrophic forgetting of previously learned languages. This work investigates how Continual Learning (CL) can be used to sustainably expand multilingual Speech LLMs. We first demonstrate that multilingual projectors can be efficiently bootstrapped to new languages , even with extremely small datasets, but at the cost of severe degradation on the original supported languages. To address this, we adopt rehearsal-based CL strategies and show that interleaving even small amounts of replay data effectively stabilizes multilingual performance. Through extensive ablations, we quantify the minimum rehearsal budget required to prevent forgetting and identify fragile languages that require more targeted reinforcement. We further evaluate sequential acquisition of four linguistically diverse languages (Ukrainian, Japanese, Thai, and Vietnamese), revealing the trade -offs between buffer size and long- term stability. Finally, based on these empirical observations, we propose a Fragility-Based Sampling heuristic as a pathway to allocate rehearsal data more efficiently by tiering languages according to their stability thresholds. Our findings provide a practical roadmap for scalable, resource-efficient multilingual expansion of Speech LLMs, enabling inclusive ASR systems that can grow over time without sacrificing prior knowledge.
Responsible Benchmarking of Fairness for Automatic Speech Recognition
Felix E. Herron | Ange Richard | François Portet | Alexandre Allauzen | Solange Rossato
Felix E. Herron | Ange Richard | François Portet | Alexandre Allauzen | Solange Rossato
Many studies have shown automatic speech processing (ASR) systems have unequal performance across speaker groups (SG’s). However, the manner in which such studies arrive at this conclusion is inconsistent. To pave the way for more reliable results in future studies, we lay out best practices for benchmarking ASR fairness based on literature from machine learning fairness, social sciences, and speech science. We then perform a case study on the Fair-speech benchmark, applying aforementioned best practices, and discuss how failing to do so can result in erroneous conclusions. On the whole, we advocate for as fine-grained an analysis as possible, taking into account as many variables as are available, in order to eschew dataset-level bias.
Addressing Accent Disparities in Automatic Speech Recognition: A Comparative Study of Single and Two-Step Adaptation
Mykhailo Danilevskyi | Fernando Perez-Tellez | Jelena Vasic
Mykhailo Danilevskyi | Fernando Perez-Tellez | Jelena Vasic
Automatic speech recognition (ASR) systems often exhibit uneven performance across accents, raising concerns about fairness and bias. This study investigates the impact of model fine-tuning strategies on ASR performance and accent-related disparities. We conduct a controlled empirical evaluation of two adaptation approaches—single-step and two-step fine-tuning—using pretrained Whisper (small) and Wav2Vec2-XLSR-53 models on African-accented English speech from the AfriSpeech-200 dataset, covering Yoruba, Igbo, Swahili, and Hausa accents. Both fine-tuning strategies substantially reduced mean word error rate (WER) for all models. However, these improvements did not translate into consistent reductions in accent-related performance gaps. When analysed separately across general and clinical subsets, WER gaps often increased due to uneven gains across accents. Although two-step fine-tuning provided modest improvements over single-step adaptation, its impact on reducing disparities remained limited. These findings indicate that fine-tuning primarily optimises performance without effectively addressing systematic bias across speaker groups, even when models are specialised for individual accents. This highlights the limitations of per-accent specialisation as a practical bias mitigation strategy.
Investigating Speaker Pronunciation Variability in Speech Embeddings: Speaker and L1 Effects on French as a Second Language
Maxime Fily | Martine Adda-Decker | Guillaume Wisniewski
Maxime Fily | Martine Adda-Decker | Guillaume Wisniewski
Speech variation between native and non-native speakers of French is addressed with a low-resource method based on a frame-wise comparison of wav2vec2 acoustic embeddings, using fine-grained phonetic transcriptions by expert annotators as baseline. z-normalisation and t-normalisation are explored to assess what the embeddings contain in terms of phonetically analysable information. We explore non-supervised methods for solving basic speech-related research questions. Adapting Dynamic Time Warping to speech embeddings, we compare phonologically similar recordings of sentences read-aloud by native vs. non-native speakers of French. The question is whether XLSR-53 embeddings are more robust than MFCCs to inter-speaker vs. intra-speaker variability for same words. Then we investigate whether native speaker productions are more stable than those of non-native speakers. Results suggest that the model allows phonetically meaningful correlative analyses. Working on the raw embeddings shows however that the representations are not speaker-independent, so with a view to address issues in relationship with L2 pronunciation variability, we show that t-normalisation brings us a way to separate fluency and accuracy effects in L2-speech. This shows that wav2vec2 encapsulates time-dependent phonetic information in the embeddings, including speaker accent which can not easily be disentangled from speaker ID.
What LID Systems Say About Dialectal Variation. The Case of Yiddish, Quechua and Mande
Johanna Cordova | Eric Jordan | Valentina Fedchenko
Johanna Cordova | Eric Jordan | Valentina Fedchenko
This study investigates the ability of speech-based language identification (LID) systems to handle dialectal variation in low-resource settings and explores whether classification outcomes correspond to phonetic proximity and can serve as an exploratory tool for dataset quality. We collected corpora for three macrolanguages Mande, Quechuan, and Yiddish, each presenting distinct internal variation, and evaluated three types of models: GMM, Whisper, and Wav2vec2-based architectures. Models were tested both within language families and across the entire multilingual dataset to assess generalization. Layer-wise classifiers built on wav2vec2-XLSR embeddings were used to identify the layers most sensitive to phonetic or phonological features. Results show that simple GMM models can generalize well in small, highly similar datasets, while Whisper-based classifiers tend to overfit, particularly on closely related dialects. Wav2vec2-XLSR (layer 12 + MLP) captures better fine phonetic and prosodic distinctions, suggesting that embeddings encode nuanced pronunciation cues. For datasets with more diverse sources like Quechua, Whisper demonstrates better generalization. Overall, LID classifiers can both reveal linguistic patterns and highlight dataset quality issues, with model architecture and layer-specific representations shaping performance.
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
Vrunda Nileshkumar Sukhadia | Shammur Absar Chowdhury
Vrunda Nileshkumar Sukhadia | Shammur Absar Chowdhury
Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We present HArnESS, an Arabic-centric self-supervised speech model family trained from scratch with iterative self-distillation, together with lightweight student variants that offer strong accuracy-efficiency trade-offs on Automatic Speech Recognition (ASR), Dialect Identification (DID), and Speech Emotion Recognition (SER). Our approach begins with a large bilingual Arabic-English teacher and progressively distills its knowledge into compressed student models while preserving Arabic-relevant acoustic and paralinguistic representations. We further study PCA-based compression of the teacher supervision signal to better match the capacity of shallow and thin students. Compared with HuBERT and XLS-R, HArnESS consistently improves performance on Arabic downstream tasks, while the compressed models remain competitive under substantial structural reduction. These results position HArnESS as a practical and accessible Arabic-centric SSL foundation for real-world speech applications.
Automatic Speech Recognition (ASR) evaluation has traditionally relied on Word Error Rate (WER), a metric that treats all errors equally and obscures critical failure modes. In this paper, we present a fine-grained human evaluation of Meta’s recently released OmniASR system on Saudi Arabic dialects using the SADA dataset. Three trained annotators evaluated 103 audio samples, producing 264 annotations across two dimensions (comprehensibility and naturalness) while categorizing errors using a novel 10-category Arabic-specific error taxonomy. OmniASR achieved a mean WER of 42.2% and mean comprehensibility of 3.62/5, but exhibited a bimodal performance pattern: 32.6% of transcriptions achieved perfect scores while 21.2% were essentially unusable. Error analysis reveals that hallucinations and deletions have the greatest negative impact on comprehensibility (−1.64 and −1.57 points respectively), roughly 6× more damaging than named entity errors. Importantly, WER correlates only moderately with human comprehensibility ratings (r = −0.679), explaining just 46% of variance in human judgments. These findings demonstrate the limitations of WER as a sole evaluation metric and highlight the need for human-centered, error-type-aware evaluation frameworks for Arabic ASR systems.
SpeechLM for Automatic Speech Recognition in Low-resource Languages
Md Abdur Razzaq Riyadh | Eneko Agirre | Eva Navas | Claudia Borg
Md Abdur Razzaq Riyadh | Eneko Agirre | Eva Navas | Claudia Borg
Multi-modal Speech Language Models (SpeechLMs) are a recent advancement in natural language processing. These SpeechLMs are instruction-tuned and optimized for general tasks. Their usefulness for Automatic Speech Recognition (ASR), particularly in relatively low-resource scenarios, remains largely understudied. This work developed SpeechLM for ASR in Basque and Maltese and studied the impact of language-adapted Large Language Model (LLM) and speech encoder within the SpeechLM for ASR. Using supervised learning, we fine-tuned LLaMA-Omni, a SpeechLM, for ASR. We have conducted comprehensive hyperparameter tuning and experimented with language-adapted SpeechLM components to improve performance and evaluated our best models on in-distribution datasets for both languages and an out-of-distribution dataset for Basque. LLaMA-Omni achieved 8.09% WER in Basque and 25.65% WER for Maltese on average across multiple test splits. The in-distribution results show that SpeechLM outperforms a fine-tuned ASR system under specific constraints, whereas it underperforms the baseline model on out-of-distribution Basque, indicating weaker overall robustness. We also find that a language-adapted LLM within SpeechLM improves in out-of-distribution settings when compared to the off-the-shelf LLM within SpeechLM.
Improving Low-resource ASR Using Bilingual Fine-tuning with Language Identification: A Cross-linguistic Evaluation
Reihaneh Amooie | Yun Hao | Wietse de Vries | Jelske Dijkstra | Matt Coler | Martijn Wieling
Reihaneh Amooie | Yun Hao | Wietse de Vries | Jelske Dijkstra | Matt Coler | Martijn Wieling
This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographically diverse language pairs, covering a range of language families and writing systems. To distinguish the two languages, during training, we pre-pend each input text with a language identification token. At inference, the model jointly predicts both the language and transcription from the speech input alone. As texts for which the language is incorrectly determined show low ASR performance, we also conduct a follow-up experiment in which the language identification token is provided both during training and inference. Our results show that bilingual fine-tuning can be beneficial when language identification accuracy is high, and that in cases where language identification performance is low, including the language identification token at inference helps to improve ASR performance.
Leveraging Speech Models for Audio-based Lexical Retrieval in Dictionaries: The Case of the Teochew Language
Siman Chen | Ilaine Wang | Maxime Fily | Pierre Magistry
Siman Chen | Ilaine Wang | Maxime Fily | Pierre Magistry
This study presents our attempt on applying Query by Example - Spoken Term Detection methodologies to a real-world, low-resource scenario: building an audio-based query functionality for the diasporan Teochew dictionary WhatTCSay. This functionality enables users to retrieve dictionary entries without prior knowledge of the writing systems in Teochew, thereby enhancing the accessibility of the dictionary and facilitating language revitalization efforts within Teochew communities. To address the retrieval task, we investigate two approaches: (i) an ASR-based approach using text-to-text matching, and (ii) a Dynamic Time Warping (DTW)-based acoustic framework for audio-to-audio retrieval. In the first approach, we compare an automatic romanization of the spoken query against the gold romanization from the dictionary; in the second, we directly match the user’s spoken query against audio recordings from the dictionary pronounced by a native speaker. Retrieval performance is evaluated using recall at rank k. Results show that text-to-text matching achieves better performance than audio-to-audio matching; however, the two approaches were not optimized under fully comparable conditions, as the ASR-based approach benefited from additional optimization, which was not equally available for the DTW method.
Stage-Aware Cross-Lingual Transfer for Faroese ASR: When and Which Languages Matter
Dávid í Lág | Barbara Scalvini | Carlos Daniel Mena | Jón Guðnason
Dávid í Lág | Barbara Scalvini | Carlos Daniel Mena | Jón Guðnason
Automatic speech recognition (ASR) for low-resource languages remains challenging due to limited labeled data. Although multilingual models and the inclusion of related auxiliary languages enable cross-lingual transfer, it is still unclear how introducing cross-lingual information at different training stages-pre-training versus fine-tuning-affects downstream performance. Prior work largely treats transfer as a single-stage optimization problem without disentangling stage effects. We present a stage-aware analysis of cross-lingual transfer for Faroese ASR using related auxiliary languages and Wav2Vec 2.0 XLS-R models. We systematically compare two complementary adaptation pipelines: (i) cross-lingual supervised fine-tuning and (ii) cross-lingual continuous pre-training prior to fine-tuning. Both strategies are evaluated under a unified setup with controlled model architectures, balanced representation of auxiliary languages, and identical evaluation protocols. Results demonstrate that cross-lingual transfer is stage-dependent. Supervised adaptation optimizes in-domain accuracy, while pretraining-level adaptation enhances robustness and reduces Character Error Rate (CER). Auxiliary language effects vary across pipelines, reinforcing the idea that transfer effectiveness depends on when and how cross-lingual information is introduced. Comparisons with large-scale multilingual ASR models highlight trade-offs between model scale and explicit, small-scale domain-aware adaptation. These findings suggest that effective cross-lingual transfer for Faroese low-resource ASR is inherently stage-dependent rather than a single-step design choice.
Doing More with Less: Determining Optimal Pre-training Model for Irish Automatic Speech Recognition through Multi-step Fine-tuning
Caoilfhionn Ní Dheoráin | Ruth Holmes | Nicholas Evans | Thomas Laurent | Anthony Ventresque | Ellen Rushe
Caoilfhionn Ní Dheoráin | Ruth Holmes | Nicholas Evans | Thomas Laurent | Anthony Ventresque | Ellen Rushe
In recent years, there has been an upsurge in research on automatic speech recognition (ASR) for low-resource languages. Particularly, transfer learning using multi-lingual models has become a popular remedy for the lack of available datasets for target languages. However, given the complexities associated with each individual language, we argue it is unlikely that a single multi-lingual pre-training model will provide equal performance gains across all languages. We also recognise the important, and insufficiently studied influence that the specific pre-training dataset has on the performance of the model. In this paper, using the Irish language as a case study, we propose a more directed, incremental form of pre-training which we term multi-step fine-tuning. This method accounts for the complex relationships between the language and dataset features of the source pre-training and target datasets. We show multi-step fine-tuning improves performance over simple multi-lingual fine-tuning alone, and we investigate factors leading to certain pre-trained models achieving better results through linguistic and dataset similarity measures. This research also investigates the uniformity of the performance gains across different demographics. We show that the optimal pre-training strategy can differ between demographics suggesting that more careful pre-training dataset selection is necessary to ensure equitable outcomes in practice.
Blank-Aware Decoding for Transcript-Free Phoneme Alignment in Low-Resource Languages and Dialects
Domenico De Cristofaro | Barbara Plank | Alessandro Vietti
Domenico De Cristofaro | Barbara Plank | Alessandro Vietti
We present a blank-aware decoding approach for transcript-free phoneme alignment with CTC-based speech foundation models, designed to improve annotation bootstrapping in low-resource languages. While CTC models provide frame-level phoneme posteriors without requiring transcripts, greedy decoding produces blank-dominated and temporally unstable segmentations that are difficult to correct manually. Our approach introduces two training-free blank-resolution strategies operating directly on CTC logits: (i) confidence-ratio substitution, which promotes competitive non-blank hypotheses relative to the blank symbol, and (ii) recursive context adjustment, which enforces local contextual consistency within blank spans. Experiments on English (TIMIT) and on Sardinian and Tyrolean dialect corpora show consistent improvements in boundary F1 prediction, phoneme duration regularity, and segmentation stability over greedy CTC decoding. Although absolute boundary deviations remain higher than transcript-conditioned aligners, the resulting alignments are structurally coherent and suitable for manual correction. A post-hoc phoneme-class analysis further reveals systematic asymmetries in blank resolution, highlighting complementary roles of local acoustic evidence and contextual cues, and outlining prominising venues for future improvements.
On the Role of Encoder Depth: Pruning Whisper and LoRA Fine-Tuning in SLAM-ASR
Ganesh Pavan Kartikeya Bharadwaj Kolluri | Michael Kampouridis | Ravi Shekhar
Ganesh Pavan Kartikeya Bharadwaj Kolluri | Michael Kampouridis | Ravi Shekhar
Automatic speech recognition (ASR) has advanced rapidly in recent years, driven by large-scale pretrained models and end-to-end architectures such as SLAM-ASR. A key component of SLAM-ASR systems is the Whisper speech encoder, which provides robust acoustic representations. While model pruning has been explored for the full Whisper encoder–decoder architecture, its impact within the SLAM-ASR setting remains under-investigated. In this work, we analyze the effects of layer pruning in the Whisper encoder when used as the acoustic backbone of SLAM-ASR. We further examine the extent to which LoRA-based fine-tuning can recover performance degradation caused by pruning. Experiments conducted across three Whisper variants (Small, Medium, Large-v2), three languages representing distinct resource levels (Danish, Dutch, English), and over 200 training runs demonstrate that pruning two encoder layers causes only 2–4% WER degradation, and that combining this pruning with LoRA adaptation consistently outperforms the unpruned baseline while reducing total parameters by 7–14%. Moreover, our error analysis reveals that LoRA primarily compensates through the language model’s linguistic priors, reducing total word errors by 18.2%, with substitution errors showing the largest reduction. However, for low-resource Danish, LoRA introduces increased insertion errors, indicating that compensation effectiveness depends on the LLM’s pre-existing language proficiency and available training data.
TaLK-Corpus: A Regionally Diverse Evaluation Set for Sri Lankan Tamil Speech
Adsajan Thillainathan | Nishanthini Kanthakumar | Nivethiga Rasan | Kengatharaiyer Sarveswaran
Adsajan Thillainathan | Nishanthini Kanthakumar | Nivethiga Rasan | Kengatharaiyer Sarveswaran
This paper introduces the TaLK Corpus, the first speech benchmark corpus for Sri Lankan Tamil Automatic Speech Recognition (ASR) covering speech from 22 administrative districts of Sri Lanka. The corpus contains 1 hour and 33 minutes of speech from 22 native speakers (one per district) and includes rich metadata on demographics, location history, recording conditions, and domain information, along with transcriptions in Tamil script and the International Phonetic Alphabet (IPA). Standardised preprocessing (16 kHz mono WAV format) and segmentation using Silero Voice Activity Detection (VAD) resulted in 1,214 utterances. All recordings were manually transcribed by trained linguists, and MD5-based file naming used to ensure data integrity and consistency. TaLK corpus enables district-wise benchmarking of ASR systems and supports dialect-sensitive evaluation. We establish baseline results for multilingual models (Whisper Large-V3 and Facebook’s MMS) in zero-shot settings. The evaluation reveals substantial performance disparities across districts, highlighting the impact of regional phonological variation in low-resource Sri Lankan Tamil. Although Whisper Large-V3 outperforms MMS overall, it shows considerable variability, with mean Word Error Rates ranging from 0.672 to 0.903 across districts. These findings demonstrate strong regional effects even within a single model. By releasing TaLK-Corpus under the CC-BY-NC 4.0 licence, we aim to support dialect-robust ASR research and foster inclusive speech technologies for Sri Lankan Tamil-speaking communities.
up
Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
Çağrı Çöltekin | Kaja Dobrovoljc
Çağrı Çöltekin | Kaja Dobrovoljc
Probing the Dynamics of Syntactic Ability Acquisition Throughout LLM Pretraining
Hiroshi Matsuda | Masayuki Asahara
Hiroshi Matsuda | Masayuki Asahara
In this research, we introduce LoRA probing, a lightweight approach for observing how core syntactic abilities emerge during LLM pretraining. Leveraging OLMo-2’s public intermediate checkpoints, we trace learning curves across 24 pretraining stages on 33 Universal Dependencies languages by fine-tuning LoRA with step-by-step parsing instructions and a simple tabular output. To fit the relatively short context length of the OLMo-2, we design a compact 2-step-no-form prompt template and this matches the baseline in average accuracy while halving the context length and substantially increasing throughput, enabling efficient large-scale evaluation. Token Recall surpasses 0.9 within the first 1–2K pretraining steps, indicating that stable output formatting emerges early. Despite OLMo-2-7B’s English-centric pretraining, LAS exceeds 80 points in 29 of 33 languages; however, relations such as iobj and csubj show delayed onset and instability across many languages. LoRA probing thus provides a practical, reproducible lens on the cross-lingual dynamics of syntactic acquisition during LLM pretraining.
Which languages are "hot", and which are "cool"? Using Universal Dependencies for large-scale comparisons of subject expression
Natalia Levshina
Natalia Levshina
This study uses Universal Dependencies to investigate subject omission across fifty-six news corpora and twenty geographic varieties of English. Building on McLuhan’s "hot–cool" distinction, Hall’s LC–HC continuum, and Bisang’s notion of overt vs. hidden complexity, it tests whether subject omission rates reflect degrees of contextual reliance. The results broadly support these theories: low-context, "hot" languages, such as German, Dutch and Swedish, show low omission rates, while high-context, "cool" languages, such as Japanese, Korean and Chinese, show higher rates. English behaves as a "hot" language but exhibits internal variation across varieties, with Southeast Asian varieties exhibiting more omission than African ones. The study provides large-scale quantitative evidence while highlighting the need for further theoretical and methodological refinement, particularly regarding the role of word order and agreement.
This paper aims to give a preliminary analysis of focus marking constructions in Tigrinya within the Universal Dependency (UD) framework. We identify three types of constructions that we consider clefting. All three involve a copula placed after the focus element, the UD root, and a subordinate verb form representing the presupposed information. Tigrinya has three possibilities for this subordinate verb, which we annotate as csubj, advcl, and xcomp. In each, we use the subrelation cleft to indicate the common structure and function of the different cleft types on a par with the use of subrelations pass for passive in different languages.
Coconstructions in Spoken Data: UD Annotation Guidelines and First Results
Ludovica Pannitto | Kaja Dobrovoljc Zor | Sylvain Kahane | Elena Battaglia | Bruno Guillaume | Caterina Mauri | Eleonora Zucchini
Ludovica Pannitto | Kaja Dobrovoljc Zor | Sylvain Kahane | Elena Battaglia | Bruno Guillaume | Caterina Mauri | Eleonora Zucchini
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns — including collaborative coconstructions proper, wh-question answers, and backchannels — in spoken language treebanks within the Universal Dependencies framework. Two representations are proposed: a speaker-based representation following the segmentation into speech turns, and a dependency-based representation with dependencies across speech turns. New propositions are also put forward to distinguish between reformulations and repairs, and to promote elements in unfinished phrases.
Verifying the Menzerath-Altmann law in the verbal domain in 180 languages
Pegah Faghiri | Kim Gerdes | Sylvain Kahane
Pegah Faghiri | Kim Gerdes | Sylvain Kahane
We present a large-scale evaluation of the Menzerath-Altmann law (MAL) in the verbal domain across 180 languages, using the Universal Dependencies (UD) treebank collection (v2.17). MAL predicts that as the number of constituents of a linguistic unit increases, their average size decreases. We propose a robust metric to estimate the MAL effect across corpora of widely varying sizes and define threshold-based categories to classify languages along a MAL preference cline. Crucially, we analyse the preverbal and postverbal domains separately, in addition to the standard bilateral MAL, and control for potential sampling bias by comparing results across language families (Indo-European vs. non-Indo-European) and syntactic types (VO, OV and no dominant order). Our results confirm MAL as a typologically widespread preference but not an absolute universal: several languages display a trivial or even opposite (anti-MAL) tendency. Furthermore, we uncover a significant asymmetry between the two sides of the verb: the MAL effect is stronger in the postverbal domain, while anti-MAL is stronger in the preverbal domain. VO languages tend to show a stronger MAL preference postverbally, whereas OV languages do so preverbally. These findings challenge the widespread assumption that length-based ordering constraints apply symmetrically on both sides of the verb and contribute new cross-linguistic evidence to the debate on the interaction between dependency length minimization and constituent size.
Comparing Dependency Distances of Esperanto and Other Languages in a Multi-Lingual Parallel Corpus
Masanori Oya
Masanori Oya
This study attempts to integrate Esperanto into the research of dependency distance, based on a small-scale multi-lingual parallel corpus of Manifesto de Prago, annotated with UD relations, to contribute to the development of research on Esperanto as a natural language. The mean dependency distance and the distribution of dependency distances of Esperanto are not significantly different from the majority of the 21 languages in the corpus, thus partially supporting the claim that Esperanto is no less natural than other natural languages.
Negation of Turkic non-verbal clauses: Analysis and Universal Dependencies Implementation
Nikolett Mus | Furkan Akkurt | Bermet Chontaeva | Soudabeh Eslami | Sardana Ivanova | Çağrı Çöltekin | Jonathan N. Washington | Gulnura Dzhumalieva | Aida Kasieva
Nikolett Mus | Furkan Akkurt | Bermet Chontaeva | Soudabeh Eslami | Sardana Ivanova | Çağrı Çöltekin | Jonathan N. Washington | Gulnura Dzhumalieva | Aida Kasieva
The paper examines the grammatical behavior of the negative element used to negate predicates in non-verbal clauses in three Turkic languages: Azerbaijani, Kyrgyz, and Turkish. We focus on its interaction with verbal copulas, subject agreement, and the distribution of agreement suffixes, as well as its position within the predicate phrase. The study draws on both previously described corpus data and newly collected examples. Across all three languages, agreement features are realised on the negative element only in the absence of an overt copula. The agreement morphology involved is identical to that found with nominal, adjectival, and adverbial predicates. In all the languages examined, the negative element remains within the predicate phrase; thus, its position is syntactically constrained. At the same time, we observe differences among the languages in the degree to which the position of the negator is fixed within the predicate. In Turkish and Azerbaijani, regardless of which element of the nominal predicate it negates, the negator invariably follows the predicate. In Kyrgyz, by contrast, it consistently appears immediately after the element it negates within the predicate. These patterns suggest that the negative element behaves syntactically as a phrasal operator associated with non-verbal predicates. For annotation purposes, we therefore propose analysing the negative element as a negation modifier, assigning it the POS tag ADV and the dependency relation advmod:neg.
Towards Universal Dependencies for L2 Learners of Modern Greek: Annotation and Challenges
Christina Klironomou | Thelka Pasparaki | Arianna Masciolini | Alexandros Tantos | Despoina Ourania Touriki | Konstantinos Tsiotskas | Eleni Tsourilla
Christina Klironomou | Thelka Pasparaki | Arianna Masciolini | Alexandros Tantos | Despoina Ourania Touriki | Konstantinos Tsiotskas | Eleni Tsourilla
This paper focuses on annotating the Greek Learner Corpus in Universal Dependencies (UD). It presents the annotation process, development of guidelines and evaluation of the attempted annotation of two annotators. This work is part of a larger annotation project which aims to compile a sizeable learner treebank that can be used to promote research on second language acquisition and its automatic processing.
A Comparative Linguistic Analysis of Ottoman and Modern Turkish through UD Treebanks
Enes Yılandiloğlu
Enes Yılandiloğlu
While the linguistic shifts between Ottoman and modern Turkish are well-documented qualitatively, quantitative analyses remain scarce. This study addresses this by conducting a comparative computational analysis using two Universal Dependencies treebanks: OTA-DUDU for Ottoman Turkish and TR-BOUN for modern Turkish. By employing descriptive statistics and a log-likelihood ratio test, we demonstrate the change and quantify the magnitude of diachronic variation. The analysis yields three primary statistical findings. First, our data reveals a 77% compliance rate with labial vowel harmony for suffixes, while this value is 98% in modern Turkish. This discrepancy can be explained by the presence of rounding in Ottoman Turkish, which disappears in modern Turkish. On the other hand, the compliance rate of palatal vowel harmony is quite high for both languages, 96% for Ottoman Turkish and 99% for modern Turkish. Second, some suffixes, such as the converb -(y)Ip and the dative infinitive -mAyA, changed by reducing their allomorphs in modern Turkish. Third, we demonstrate that Arabic and Persian pluralization rules, which constituted 28% of plural nouns in Ottoman Turkish, lost their pluralizing function in modern Turkish, although the words remain with singular meaning.
The Southern Bantu language family contains languages with so-called conjunctive orthographies and disjunctive orthographies. In languages with conjunctive orthographies, such as isiZulu, orthographic words correspond to linguistic words, whereas in languages with disjunctive orthographies, prefix morphemes of verbs and other predicates are written as disjunct, orthographic words. When developing Universal Dependencies treebanks, the basic principle is to consider syntactic (linguistic) words, but for languages with agglutinating morphology, it has been argued that this reduces the informativeness of the treebank. In this paper we investigate this claim by analysing and measuring the effects of annotating universal dependencies on the basis of orthographic words on two morphosyntactically parallel treebanks for isiZulu and Sepedi.
CoBra: A Compound Branching Resource for Nominal Triconstituent Compounds in English and German
Carmen Schacht | Isabell Landwehr | Diana Davidson | Konrad Grabowski | Magdalena Meiser | Sophia Wiedmann
Carmen Schacht | Isabell Landwehr | Diana Davidson | Konrad Grabowski | Magdalena Meiser | Sophia Wiedmann
We present CoBra, a resource containing triconstituent nominal compounds in English and German. This addresses an understudied aspect of compound processing, since research and resources in psycholinguistics and NLP have mostly focused on two-constituent compounds. In addition, our resource covers both general and scientific language, allowing for a register-informed perspective on compounds. It provides syntactic and semantic annotation of compound structure, in particular of the branching direction (i.e. the internal embedding structure, the Compound Branching) and the semantic relationship between constituents. Annotations are implemented using extensions of Universal Dependencies (UD) labels. To explore applications of our new resource, we also conduct a pilot study investigating the relationship between semantic transparency and branching direction. Our results indicate that there is indeed a correlation. Overall, our resource contributes to gaining a more detailed understanding of the structure and processing of morphologically complex words within the UD framework.
SE Constructions Revisited: Focus on Treebanks for Romance Languages
Verginica Barbu Mititelu | Elena Irimia | Adriana S. Pagano | Roxana Ciolaneanu | Ioana Buhnila
Verginica Barbu Mititelu | Elena Irimia | Adriana S. Pagano | Roxana Ciolaneanu | Ioana Buhnila
We analyze the current annotation of SE constructions, i.e. verbal constructions marked by the clitic se and its cognates across five Romance languages (French, Italian, Portuguese (European and Brazilian), Romanian and Spanish) in several Universal Dependencies treebanks (version 2.17). We discuss the morphologic, syntactic and semantic characteristics of such constructions in each of the languages considered, both from a theoretical perspective and from that of existing annotation. To address inconsistencies in the data and strengthen Universal Dependencies as a scaffold for the automatic conversion of morphosyntactic annotation into semantic representations (Uniform Meaning Representation), we propose a clear distinction between argumental and non-argumental uses of the reflexive clitic, and outline systematic ways to implement this distinction in annotation guidelines. We also examine how some of the reported inconsistencies are handled in the treebanks under study and discuss the extent to which these practices can be extended to other treebanks, within the same or across different languages.
Say "No" to Missing Polarity: A Negation Enrichment of Porttinari UD Treebank
Isaac Souza de Miranda Junior | Oto Araújo Vale | Marie-Catherine de Marneffe
Isaac Souza de Miranda Junior | Oto Araújo Vale | Marie-Catherine de Marneffe
Negation is a central phenomenon in linguistics: every language has some way of expressing the difference between an affirmative sentence and a negative one (Horn and Wansing, 2025). However, the treatment of negation remains uneven in Natural Language Processing (Jimenez-Zafra et al., 2017; Jiménez-Zafra et al., 2020). This paper presents the enrichment of a Brazilian Portuguese corpus with negation-related morphological information within the Universal Dependencies (UD) framework (Nivre et al., 2020; de Marneffe et al., 2021). We enrich the Porttinari-base corpus (Duran et al., 2023) by systematically adding the UD morphological features Polarity=Neg and PronType=Neg for 18 negation-related lexical items. The enrichment only modifies the morphological features, leaving tokenization and dependency structure unchanged. To evaluate the computational results of this enrichment, we present an experiment using the Brazilian Portuguese parser PortParser (Lopes and Pardo, 2024), which we trained both on the original Porttinari-base data (Duran et al., 2023) and on our enriched version. Our results show that after enrichment, the parser’s performance remains stable, and the newly introduced features are being learned.
The Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across the UD Languages
Kim Gerdes
Kim Gerdes
Dependency length minimization (DLM) is a well-documented processing universal, but previous studies report a single mean dependency distance (MDD) per language, obscuring variation across syntactic relation types. We analyze 122 languages in UD and SUD (version 2.17), showing that DLM operates on two distinct levels. Grammar-driven optimization targets functional dependencies (det, case, aux), which are universally short (mean 1.71, σ=0.33) and invariant across typologically diverse languages. Processing-driven optimization operates on lexical dependencies (nsubj, obj, obl), which are longer (mean 2.87), highly variable (σ=0.63), and constrained by word-order typology. This asymmetry holds in SUD despite reversed head direction (r=0.92). We conclude that "the grammar does the work" of minimization by scaffolding sentences with local functional attachments, leaving processing pressures to determine the ordering of lexical heads.
Greenberg’s Universal 45 in Universal Dependencies: Gender Distinctions and Annotation Challenges
Antoni Brosa-Rodriguez | M. Dolores Jimenez Lopez
Antoni Brosa-Rodriguez | M. Dolores Jimenez Lopez
This paper revisits and extends Greenberg’s Universal 45 on gender distinctions using Universal Dependencies 2.17, comprising 339 treebanks across 186 languages. A systematic analysis of morphosyntactic patterns confirms the implicational hierarchy (singular > plural gender marking), with 98.6 % conformity in pronominal categories. Only two potential exceptions are detected, both with minimal occurrences and likely attributable to annotation errors. Extending the analysis beyond pronouns to 13 UPOS categories shows that core categories maintain near-perfect compliance, while peripheral categories exhibit higher violation rates, primarily driven by annotation inconsistencies rather than genuine linguistic exceptions. A total of 90 treebanks display gender-number features in traditionally invariable categories (e.g., adpositions, conjunctions, adverbs), indicating annotation issues such as prepositional contraction handling, homophone merging, and erroneous feature assignment. The study establishes a replicable computational methodology for large-scale typological validation, highlighting both the potential of corpus-based approaches and key limitations, including genealogical sampling biases, annotation heterogeneity despite universal schemas, and the false sense of comparability across treebanks.
A Proposal for a More Universal Annotation of Relative Clauses in Universal Dependencies
Santiago Herrera | Sylvain Kahane
Santiago Herrera | Sylvain Kahane
This paper proposes a new encoding of relative clauses compatible with the Universal Dependencies annotation scheme for syntactic treebanks. After showing that the current guidelines are based on the main strategy of relativization in European languages, we show that it cannot be easily extended to other relativization strategies, especially head-internal relatives clauses. The criteria for the POS of relativizers are discussed. We apply our annotation to an important variety of strategies in Mandarin, Japanese, Bambara, German, Basque, Turkish, Latin, Beja, Gbaya, and Wolof, as well as to various strategies in English, including participial clauses
MesoTree: Annotated Linguistic Resources for Quantitative Comparative Linguistic Analysis and NLP in Mesoamerica
Robert Pugh | Francis Tyers | Robert Henderson
Robert Pugh | Francis Tyers | Robert Henderson
One aspect of descriptive and documentary linguistic materials that is becoming increasingly important in the information age is that they be searchable, quantifiable, and comparable. In this paper, we describe an effort to create morphosyntactically-annotated corpora for a number of under-served Mesoamerican languages using Universal Dependencies. We describe the Mesoamerican linguistic area and languages involved in the project, the training and annotation process, and give a status report on the current state of the corpora. Finally, we describe a comparitive syntax experiment and train UD parsing models on the data, demonstrating the usefulness of UD for facilitating quantitative, comparative linguistic research.
Nonprototypical Predication and Nonpredicational Clauses in Universal Dependencies
Joakim Nivre | William Croft | Andre Coneglian
Joakim Nivre | William Croft | Andre Coneglian
To assess whether the framework of Universal Dependencies (UD) is compatible with findings from linguistic typology, we need to systematically review how UD represents linguistic constructions and how it handles the range of morphosyntactic variation attested across languages. In this paper, we present such a review focusing on nonprototypical predication and nonpredicational clauses. We find that, while nonprototypical predication is generally handled well in the UD framework, nonpredicational clauses are not discussed as such in the guidelines and often have to be annotated in a way that does not reflect their special information packaging functions. We briefly discuss ways in which the UD framework could be extended in order to better capture these functions.
To assess whether the framework of Universal Dependencies (UD) is compatible with findings from linguistic typology, we need to systematically review how UD represents linguistic constructions and how it handles the range of morphosyntactic variation attested across languages. In this paper, we present the results of such a review focusing on complex predicates. We arrive at distinct findings regarding the two main types of complex predicates. The UD framework can well accommodate eventive complex predicates, particularly serial verbs, and more grammaticalized forms of complex predicates, such as voice and TAMP auxiliaries, with the exception of incorporating strategies. However, the guidelines for stative complex predicates could be revised based on the typology of morphosyntactic strategies. We briefly discuss possible ways in which UD can be extended to better capture these strategies.
Bringing Information Structure to Universal Dependencies
Nikolett Mus | Andrew Dyer | Claudia Corbetta | Sylvain Kahane
Nikolett Mus | Andrew Dyer | Claudia Corbetta | Sylvain Kahane
The aim of the paper is to present a first attempt at annotating Information Structure roles in syntactic treebanks of the Universal Dependencies collection, discussing theoretical considerations and practical methodological questions while presenting our core annotation principles. We focus on constructions in which Topic or Focus is overtly marked through (morpho)syntactic means. The proposed annotation is illustrated using examples from five languages: Wolof, Japanese, Tundra Nenets, Hungarian, and Italian.
Introducing Universal Dependencies for Sardinian: the UD ContSar Treebank
Nicoletta Puddu | Manuela Sanguinetti | Luigi Talamo
Nicoletta Puddu | Manuela Sanguinetti | Luigi Talamo
This paper introduces the first steps towards the creation of a novel resource for contemporary Sardinian within the Universal Dependencies framework. Sardinian is a Romance language spoken in Sardinia, an island belonging to the Italian Republic and located in the center of the western Mediterranean. It is a minority and endangered language, traditionally transmitted mainly orally, and characterized by a multiplicity of varieties (usually grouped into two macro-varieties Logudorese and Campidanese), all recognized as part of the Sardinian linguistic continuum. These varieties share basic morphosyntactic features, while presenting differences at the lexical level and in the realization of specific constructions. This internal variation can be particularly challenging with regard to the normalization of lemmas and the linguistic characterization of certain phenomena. The development of the treebank therefore aims to provide an annotated resource for contemporary Sardinian that takes into account the specificities of the different varieties, using Universal Dependencies to represent them within a unified theoretical framework, in order to facilitate both linguistic analysis and automatic processing. The present paper thus describes some linguistic characteristics of Sardinian and the attempts to encode them within the UD framework. Finally, we present the results of our evaluation of an NLP pipeline for Sardinian, trained on our corpus, for the Stanford Stanza parser.
In this paper we describe TEITOK’s approach to parallel aligned treebanks: a framework that supports word-level aligned, as well as multiple level text alignment (text, paragraph, sentence). It also provides various ways a visualizing aligned data, including a newly introduced parallel visualisation for dependency trees, with mouse-over aligned highlighting. The also is a parallel search function under development, that allows queries and statistics that are both tree capable and multi-level alignment capable, including word-level alignment.
Cross-Dialectal Transfer for Low-Resource Arabic: The Tunisian Arabic Dependency Treebank
Amal Aissaoui
Amal Aissaoui
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties. The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction. This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency. This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
From Treebank Metadata to Sentence-Level Genre in Universal Dependencies: A Reproducible, Versioned Resource
Egon Stemle
Egon Stemle
We release a sentence-level genre layer for Universal Dependencies as a separate, joinable dataset, computed across UD revisions and linked back to the underlying treebanks via a release-aware composite key comprising treebank, split, sent_id, and UD release metadata. The annotations are derived rather than authoritative and are accompanied by provenance and uncertainty indicators, enabling downstream users to choose appropriate precision-coverage trade-offs and to re-run the pipeline as UD evolves. To support both parity tracking and deployment-oriented interpretation, we report results under two complementary regimes: a fixed-partition setting aligned with earlier protocols, and a language-grouped 10-fold generalisation setting that highlights cross-language heterogeneity and anchor sparsity as operational constraints. The resulting resource is intended to make genre a practical control variable for UD-based experimentation, including genre-stratified evaluation and training data selection for POS tagging and parsing, where performance varies substantially across text types. Finally, we note that reduced genre spaces aligned with recurring robustness profiles (e.g. transcribed speech versus interactional web/social text versus edited prose/news) appear pragmatically useful, but should be treated as a community coordination task implemented through explicit, versioned mapping tables.
Towards a Universal Dependency Corpus for Old Saxon (Old Low German)
Christian Chiarcos | Janine Siewert
Christian Chiarcos | Janine Siewert
Among the West Germanic languages of the first millenium C.E. (Old English, Old Low Franconian/Old Dutch, Old High German, and – although much later – Old Frisian), Old Saxon occupies a special role both linguistically – in that it represents a middle ground in the dialect spectrum between Old English at one extreme and Old High German on the other –, and in terms of material quality, in that it is attested with considerable amounts of coherent text (unlike Old Low Franconian) which is not only particularly old (unlike, especially, Old Frisian), but also original (i.e., not translated, a rarity in the attested Old English and Old High German material). It is thus a language central to the understanding of the emergence of several modern major languages, incl. English, Dutch and German, and has been studied intensely, albeit – so far – not in the context of the Universal Dependencies. This paper addresses this gap and describes the introduction of (a) a manually annotated test corpus of Old Saxon, (b) a highly reusable conversion pipeline for converting the Penn bracketing syntax of the Penn Historical Corpora (and the Old Saxon Heliand) to UD, and (c) the evaluation of the latter against the manual annotations.
Syntactic annotation is time- and resource-consuming, especially for historical and heterogeneous data. The Universal Dependencies (UD) framework provides a stable and cross-linguistically consistent annotation scheme, offering a crucial backbone for diachronic corpus studies. However, ensuring internal consistency within historical UD treebanks remains challenging due to syntactic variation and parser errors. We address this issue for Medieval and Classical French by integrating valency information into our corrections to support UD treebank maintenance. Valency frames were extracted from the Profiterole treebank (v. 2.7) and used to enrich OFrLex with structured valency information for Medieval French. Existing lexical resources such as Lefff are also exploited for Contemporary French. These valency frames are used to detect and correct inconsistencies in automatically annotated data through batch operations, thereby reinforcing UD guideline compliance and improving annotation coherence across diachronic stages. Preliminary experiments on Medieval French and exploratory annotation of Classical French data suggest that lexicon-informed error mining can reduce manual revision effort while strengthening the diachronic continuity enabled by the UD framework.
Syntax is the Key to Semantics: Combining Universal Dependencies and Abstract Meaning Representation
Johannes Heinecke
Johannes Heinecke
This paper presents a new Abstract Meaning Representation (AMR) dataset using sentences which are already translated in 21 languages and annotated in dependency syntax (Parallel Universal Dependencies, PUD) and . We annotated the English version of the 1000 sentences available in the PUD dataset. Since PUD provides syntactic annotations, this new AMR dataset (PUD-AMR) allows comparisons of syntactic phenomena and their corresponding semantics. We finally provide a first analysis of parallels between syntactic and semantic structures.
Extending Retag to Conversion Error Detection: A Case Study on SynTagRus Morphology
Andrei Movsesian | Daniil Timchenko
Andrei Movsesian | Daniil Timchenko
Linguistically annotated corpora are often converted between annotation schemes, but errors introduced during conversion can compromise their reliability. While annotation error detection is a well-studied topic, conversion error detection remains largely unexplored. We adapt the Retag method, which is traditionally used for finding annotation errors, to identify conversion errors by comparing model performance on original and converted versions of the same corpus, aligned at the token level. Applying this approach to the SynTagRus corpus converted to Universal Dependencies, we achieve high-precision detection of conversion errors in morphological annotation. Our analysis reveals systematic errors in distinction of auxiliary verbs, pronouns, numerals, and multi-word named entities, and uncovers previously undocumented annotation inconsistencies between different sections of the corpus. The method can be applied to any converted dataset for which an aligned source is available, providing an efficient way to target conversion errors for manual correction without exhaustive inspection.
Is the framework of Universal Dependencies (UD) compatible with findings from linguistic typology about constructions in the world’s languages? To address this question, we need to systematically review how UD represents these constructions, and how it handles the range of morphosyntactic variation attested across languages. In this paper, we present the results of such a review focusing on speech act constructions. We find that UD currently lack mechanisms for systematically capturing speech act constructions and briefly discuss ways in which this can be remedied.
up
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Girish Nath Jha | Kalika Bali | Sobha L | Devendr Kumar
Girish Nath Jha | Kalika Bali | Sobha L | Devendr Kumar
Many classical languages have well-studied traditions of poetic meter which enforce constraints on a poem in terms of syllable and phoneme patterns. Such advanced literary forms offer opportunities for probing deeper reasoning and language understanding in Large Language Models (LLMs) and their ability to follow strict pre-requisites and rules in generating text. In this paper, we introduce MetricalARGS, the first taxonomy of poetry-related NLP tasks designed to evaluate LLMs on metrical poetry across four dimensions: Analysis, Retrieval, Generation, and Support. We discuss how these tasks relate to existing NLP tasks, addressing questions around datasets and evaluation metrics. Taking the metrical poetry of Telugu language as our example, we illustrate how the taxonomy can be used with LLMs in practice through a quantitative and qualitative evaluation. MetricalARGS highlights the broader possibilities for understanding the capabilities and limitations of today’s LLMs through the lens of metrical poetry. We believe MetricalARGS can also serve as a reference taxonomy for studying and comparing metrical poetry across Indian languages as a starting point, and can be extended to other languages with established metrical poetry traditions.
Semi-automatic Approach for Tamil Discourse Relation Annotation
Frances Yung | Enosh Peter Ponraj | Vera Demberg
Frances Yung | Enosh Peter Ponraj | Vera Demberg
Discourse relations (DRs) specify the logical relations between text spans and are essential for modeling extended discourse. Resources annotated with DRs can help train large language models (LLMs) to recognize and generate these relations more naturally. However, there is currently no open-source DR-annotated resource for Tamil. Annotation is particularly challenging because many Tamil discourse connectives are realized as morphologically complex suffixes rather than standalone tokens, often involving phonological alternations. In this work, we present a DR-annotated dataset for Tamil based on the PDTB framework. We adopt a semi-automatic pipeline: 1) projection of automatic English discourse annotations onto Tamil in a parallel corpus; 2) lexical normalization using a morphological analyzer; and 3) manual verification of each instance. The resulting resource contains approximately 7;200 explicit DR annotations and a lexicon of 450 Tamil discourse connectives. The annotated data is available for download at https://anonymous.4open.science/r/Tamil-Semi-Automatic-Discourse-Relation-Dataset/.
Konkani Daan: A Community-Driven Culturally Grounded Speech Corpus for Low-Resource ASR
Milind Shivolkar | Vaibhav Gawas | Jyoti Pawar
Milind Shivolkar | Vaibhav Gawas | Jyoti Pawar
Indian languages are deeply embedded in cultural traditions, oral narratives, regional lexicons, and socially grounded communicative practices. However, existing speech resources and large multilingual ASR models often underrepresent culturally rich and naturally occurring speech varieties. In this paper, we introduce Konkani Daan, a community-driven initiative for collecting culturally grounded speech data for the Konkani language. The corpus currently comprises over 43.9 hours of 16 kHz speech recordings contributed through a web-based participatory platform. We evaluate a strong multilingual baseline, AI4Bharat IndicConformer-600M, in zero-shot mode on the Konkani Daan development set (379 utterances), achieving Word Error Rate (WER) of 46.46% and Character Error Rate (CER) of 15.47%, indicating substantial domain and cultural mismatch. Through qualitative error analysis, we identify systematic challenges including compound word segmentation, numeric normalisation, named entity distortion, and orthographic variation. Our findings demonstrate that culturally dense community speech exposes systematic limitations in multilingual ASR systems and motivates normalisation-aware and culturally informed evaluation strategies.
Is Literal Annotation Enough? Building an Annotation Framework for Metonymic Named Entities in Marathi
Pratibha Dongare
Pratibha Dongare
Named Entity Recognition (NER) has been a core task of natural language processing (NLP) since the Message Understanding Conferences (MUCs). Data annotation plays a crucial role in this task. However, existing annotation studies often rely on the literal sense of entities. Such annotations may lead to inconsistencies, while resolving ambiguity introduced by figurative tropes like metonymy. For example, in India won the series, India refers to a sports team instead of a geographic location. Understanding such non-literal senses is crucial for various NLP applications such as Question Answering, Information Extraction, etc. By addressing this gap, this study presents an annotation framework and detailed guidelines for annotating metonymic readings of named entities in Marathi, an Indo-Aryan language spoken in the central-western region of India. The study uses news corpus from various domains. It presents a two-tiered annotation framework for annotating conventional metonymies in Marathi language. Further, it describes the annotation framework applied to a corpus of 1,279 Marathi sentences. The result shows the inadequacy of literal-only annotation as 53.6% of named entity spans have metonymic readings. This study makes a crucial contribution for resource development for low-resource languages that share similar linguistic structures and cultural contexts. The paper describes the framework with necessary examples, challenges and concludes with a future scope.
Bengali-English and Hindi-English Code Mixed Speech Data with Disfluencies
Anuran Mitra | Tapabrata Mondal | Anirvan Chakravarty | Sivaji Bandyopadhyay
Anuran Mitra | Tapabrata Mondal | Anirvan Chakravarty | Sivaji Bandyopadhyay
Spontaneous speech in multilingual communities such as India frequently combines code-switching (CS) and disfluencies, yet existing Bengali–English and Hindi–English speech corpora largely consist of fluent or scripted utterances. This limits their suitability for developing and evaluating automatic speech recognition (ASR) systems intended for real conversational settings, particularly in micro-resource scenarios. We introduce BEHE-CMDisfl, a synthetic speech corpus that explicitly integrates disfluency phenomena within Bengali–English and Hindi–English code-mixed (CM) utterances. The textual content was generated using prompting strategies with large language models (LLMs) to encourage controlled switching and varied disfluency patterns, including filled pauses, repetitions, and restarts. The utterances were subsequently synthesized using Indic Parler text-to-speech (TTS) system. To demonstrate usability, we establish a reproducible GMM–HMM baseline for Bengali–English ASR using Kaldi on a 1.3-hour subset of the corpus. In our experiments, improvements were mainly observed after ensuring consistency in the pronunciation lexicon and applying phonetic normalization, with the best setup reaching a word error rate (WER) of 37.74%. A closer look at the decoded transcripts suggests that filled pauses and repetitions are not automatically collapsed, but appear in the output, indicating that the disfluency cues present in the synthetic speech are captured during recognition.
Konkani is a low-resource Indo-Aryan language spoken along the western coast of India, characterized by significant dialectal variation, multi-script usage, and limited standardized computational resources. This paper presents a consolidated and analysis-ready lexical resource derived from the Konkani Wordnet, built under the IndoWordNet framework. The resource comprises 32,370 synsets, 37,719 unique lexical entries, 32,370 glosses, and 33,318 example sentences, enriched with pronunciations, semantic relations, and illustrative examples. We describe the systematic extraction, normalization, and structural integration of wordnet data, resolving identifier inconsistencies and ensuring semantic coherence across distributed lexical files. To demonstrate the practical utility of this resource, we present an API-based bilingual vocabulary exercise generation system that leverages shared synset identifiers to automatically produce semantically aligned Hindi–Konkani word pairs for e-learning applications. The resulting resource enhances accessibility, reproducibility, and computational readiness for NLP tasks, while providing a foundational infrastructure for developing technology-driven teaching and learning tools for Konkani.
Development of Speech Corpus for Low-Resource Language- a Case of Sanskrit
Devendr Kumar | Girish Nath Jha | Khalid Choukri
Devendr Kumar | Girish Nath Jha | Khalid Choukri
This paper presents a comprehensive framework for the development of a speech corpus for Sanskrit, designed to facilitate advances in Automatic Speech Recognition (ASR) and AI/ML research. The proposed corpus comprises over 107 hours of transcribed speech data, collected from diverse Sanskrit sources through a systematic and scalable pipeline. We detail the end-to-end methodology adopted for corpus creation, encompassing web crawling, data sanitization, audio downloading, and transcription alignment. Particular emphasis is placed on the methodological rigor applied at each stage, including source selection, preprocessing for quality assurance, transcription protocols, and forced alignment techniques. The paper further addresses the unique complexities inherent to Sanskrit, spanning its phonetic richness, intricate morphological structure, and distinctive syntactic patterns. By systematically addressing these dimensions, the resulting 107-hour corpus aims to serve as a foundational resource for speech technology research in Sanskrit.
The Shabd Portal – Searchable Lexical Resources for Indian Languages by Government of India
Mercy Lalrohluo Hmar | Girish Nath Jha | Dhananjay Singh
Mercy Lalrohluo Hmar | Girish Nath Jha | Dhananjay Singh
The shabd portal ( https://shabd.education.gov.in ) of Commission for Scientific and Technical Terminology (CSTT), a subordinate office under the Ministry of Education, Department of Higher Education, Government of India (GOI) is a data server designed and developed by Prof. Girish Nath Jha, former Chairman CSTT, featuring all the standardized scientific and technical glossaries of CSTT in digital searchable mode. The aim is to launch a central repository for the terminologies prepared in Indian Languages, thus enriching the language bank of India enabling user friendly and free access to standardized terminology. This website is available in 22 Indian languages. The data covers several domains of science, humanities, engineering, medical science and agriculture subjects. The data is dynamic with regular updates in various domains. Users can search the equivalents of terms in Indian languages and submit their feedback for those equivalents prepared by CSTT. The unique feature of the search platform is that users have various options for search, based on languages, subjects, dictionary type and language pairs. The user can also choose to search in a specific glossary or the entire collection which includes about 471 glossaries having about (29,56,125 headwords).
POS Tagging in Low-Resource Maithili Language: Specific Challenges and Nuances
Shivani Priya | Shruti Jha | Urmila Jha | Girish Nath Jha | Deepali Tiwari | Jyoti Raj
Shivani Priya | Shruti Jha | Urmila Jha | Girish Nath Jha | Deepali Tiwari | Jyoti Raj
Abstract Part-of-Speech (POS) tagging is a key step in Natural Language Processing (NLP), laying the groundwork for more advanced syntactic and semantic tasks. Despite Maithili’s status as an Indo-Aryan language with a rich literary tradition and official recognition in India, computational resources for it are still very limited. In this paper, the creation of an annotated corpus of 25,000 sentences drawn from the fields of health, tourism, and administration is described with the hierarchical tagset currently used for Maithili. This paper also indicates that standard tagsets, typically adapted from English or Hindi, fail to capture the linguistic nuances of Maithili. This underestimates the need for a dedicated tagging framework that considers characteristics like vocative particles, verbal nuances, honorific complexities. Keywords: Parts of Speech, Natural Language Processing, Maithili, annotation
Preserving Civilisation Memory: A Digital Humanities Approach to the Ramayan
Shashank Tiwari | Girish Nath Jha
Shashank Tiwari | Girish Nath Jha
Abstract The Ramayana takes a leading role in the list of the most important texts in the world literature, with a multiplicity of textual traditions of unparalleled numbers of more than three hundred variants throughout South and Southeast. Stone manuscripts, bamboo manuscripts, and palm leaf manuscripts have been passed down in palm-leaf codices, inscriptions on temple walls, in highly illustrated folios and through generations of oral performance. But the corpus now faces serious challenges due to the destruction of the environment, material frailty, fragmentation, and the scope of modern script recognition methods. The current analysis examines how digitization projects are re-defining the preservation of Ramayana in a heritage system that is networked across the world. In this paper, the author critically assesses the work of large-scale projects thru the use of a qualitative research design that has been conducted between the years 2003 and 2026, including the National Mission for Manuscripts (NMM), the digital reunification of the Mewar Ramayana, and efforts by southeast Asian countries to document adaptations, like the Reamker by Cambodia). It predicts imaging standards, metadata formatting policies, integration of optical character recognition (OCR) and digital access structures, and struggles with the problem of multi-script complexity (Grantha, Devanagari, Kawi), partial coverage of variant texts, and infrastructural inequities. The results support that digitization has a significant positive impact on scholarly accessibility and comparative research but the advantages are unexpressed, especially relating to oral traditions. To make the endeavor sustainable preservation, interoperable standards must be adopted, the script recognition with the help of AI should be encouraged, the community should be involved, and cross-border collaboration institutionalized to protect the long-term cultural viability of the Ramayana. Keywords: Ramayana, Manuscript Preservation, Digitization of Cultural Heritage, Digital Humanities, Textual Transmission, Palm-Leaf Manuscripts, Grantha and Kawi Scripts, Metadata Standards (METS/XML), IIIF Interoperability, AI-Assisted Philology, Intangible, Cultural Heritage, Archival Sustainability, Cultural Heritage Informatics, Open Access Repositories, Civilizational Memory.
Naamah: A Large Scale Synthetic Sanskrit NER Corpus via DBpedia Seeding and LLM Generation
Annarao Kulkarni | Akhil Rajeev P
Annarao Kulkarni | Akhil Rajeev P
The digitization of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition (NER). While recent methodologies utilize generic Large Language Models (LLMs) for data augmentation, these approaches remain prone to error and often lack the reasoning depth required for classical grammar. In this work, we introduce Naamah, a high quality silver standard Sanskrit NER dataset comprising 102,942 sentences. We propose a methodology that combines entity extraction from DBpedia with the generative capabilities of a 24B parameter hybrid reasoning model to create grammatically natural and synthetically diverse training data. We utilize this dataset to benchmark two transformer architectures: the massive multilingual XLM RoBERTa and the parameter efficient IndicBERTv2. Our experiments reveal a key insight: while both models scale well with synthetic data, IndicBERTv2 qualitatively outperforms XLM RoBERTa in entity identification and classification. On a fixed split of 92,647 train and 10,295 validation examples, IndicBERTv2 achieves the best validation F1 of 0.9615, outperforming XLM R’s 0.9506 while remaining substantially lighter for deployment. We demonstrate that the generic tokenizer of XLM R fractures Sanskrit terms, whereas the domain adapted tokenizer of IndicBERTv2 preserves semantic integrity.
IndEuph-170: Benchmarking Cultural Pragmatics through Euphemism Detection in Indian English
Debamita Samajdar
Debamita Samajdar
Large Language Models (LLMs) have shown remarkable proficiency in standard English benchmarks, yet their ability to navigate the sociopragmatic cues of non-Western English varieties remains underexplored. This paper introduces IndEuph-170, a novel benchmark dataset focused on Indian English (IndE) euphemisms — expressions whose roots lie in local social hierarchies, politeness norms, and cultural taboos (e.g., "setting," "loose character," "suitable boy"). IndEuph-170 comprises 170 curated IndE sentences, against which the performance of two distinct architectures was evaluated: a fine-tuned BART model and GPT-4. The findings reveal a significant "cultural gap". While GPT-4 achieves 82.5% accuracy, it struggles with authoritative and punitive nuances. BART achieves 55.3% accuracy but exhibits a high rate of false positives by over-classifying general Indianisms as euphemisms. The paper argues that current multilingual benchmarks such as MME (Fu et al., 2025) and GLUE (Wang et al., 2018) fail to capture these dialectal pragmatics, and that a culturally-aware evaluation framework for Global Englishes is necessary.
Integrating Syntactic and Discourse Signals through Multi-Encoder Fusion in NMT for Low-Resource Indian Language Pairs
Sobha Lalitha Devi | Vijay Sundar Ram | Pattabhi RK Rao
Sobha Lalitha Devi | Vijay Sundar Ram | Pattabhi RK Rao
Neural Machine Translation (NMT) for low-resource Indian language pairs such as Hindi–Tamil and Tamil–Malayalam remains challenging due to morphological richness, syntactic divergence, and limited availability of high-quality parallel corpora. While Transformer-based architectures achieve strong performance in high-resource settings, they often struggle to model syntactic structure and discourse-level dependencies in low-resource scenarios, resulting in errors in agreement, word order, and pronoun translation. In this work, we propose a linguistically informed multi-encoder fusion framework that explicitly incorporates syntactic and discourse signals into NMT. Experiments conducted on Hindi–Tamil and Tamil–Malayalam parallel corpora demonstrate consistent improvements over strong Transformer baselines in BLEU and ChrF scores, along with gains in pronoun translation accuracy and agreement consistency. The results highlight the effectiveness of explicit linguistic integration for improving NMT in low-resource Indian language settings.
NE-LID: A Fast and Accurate Language Identification System for Northeast Indian Languages
Badal Nyalang
Badal Nyalang
Language identification (LID) is crucial for natural language processing systems, yet Northeast Indian languages remain severely underserved by existing multilingual LID models. We present NE-LID, a fast and accurate language identification system specifically designed for eleven languages of Northeast India. Built using character n-gram features with fastText, NE-LID achieves 99.09% accuracy on a balanced test set, significantly outperforming existing multilingual systems including GlotLID (73.12%), OpenLID (42.03%), IndicLID (39.30%), and LangDetect (24.33%). Our model processes predictions in 0.084 milliseconds on average, enabling real-time applications. We demonstrate that character-level modeling outperforms transformer-based approaches for script-diverse, low-resource languages
Integrating Cultural Wisdom and Digital Technologies for Children’s Moral and Emotional Development
Ms Garima | Girish Nath Jha
Ms Garima | Girish Nath Jha
The influence of technology in children’s education is increasing rapidly in the digital age, but with it the challenge of how to develop the cultural and moral development of children in a balanced manner in the technological environment. In traditional societies, moral and cultural teachings have often been imparted through religious and philosophical texts, memorization, interpretation, and oral tradition. This research presents an AI-based value-based learning framework, which aims to make cultural and ethical teachings more structured, simple, and technologically accessible to children. The study provides a brief analysis of memory-based teaching systems prevalent in various religious traditions and presents a model based on verses from the Bhagavad Gita as an example. The proposed system includes data generation and processing, simplified interpretation, semantic understanding, pronunciation analysis and interactive learning facilities based on selected material from cultural texts. The study indicates that through AI and modern technologies, traditional cultural knowledge can be delivered to children in a more effective and engaging form, developing new possibilities for reinforcing their moral and cultural development.