InFACT: Benchmarking LLM Explanations Against Institutional Reasoning for Deliberation-Aware Fact-Checking

Diana Constantina Hoefels


Abstract
Explainability in deliberation-support NLP is usually evaluated through post-hoc rationales or model-internal attribution methods, and only rarely against explicit institutional reasoning procedures. We introduce , a Romanian corpus of professional fact-checking reports that preserves the workflow of editorial epistemic arbitration, namely claim articulation, contextualisation, verification scope, evidence-based verification narrative, and calibrated conclusion. contains 789 raw reports from factual.ro and a processed benchmark release of 788 instances after removal of a singleton non-standard verdict label. Beyond six-way verdict prediction, we position as a benchmark for LLM explanation alignment, where models must generate short explanations that can be compared directly to gold institutional reasoning. We evaluate primarily with instruction-tuned LLMs, reporting full-corpus experiments for open-weight models and a matched pilot comparison with GPT-4 Turbo. The resulting evidence shows that verdict prediction and institutional explanation alignment are not the same capability: models that improve verdict accuracy do not necessarily preserve institutional calibration or produce explanations that align with professional verification narratives. These results support the central claim of the paper, namely that measures not only whether a model reaches a verdict, but also whether it does so in a manner that resembles documented public reasoning.
Anthology ID:
2026.delite-1.3
Volume:
Proceedings of The 2nd Workshop on Language-driven Deliberation Technology
Month:
May
Year:
2026
Address:
Mallorca, Spain
Editors:
Lucas Anastasiou, Katarina Boland, Anna De Liddo, Neele Falk, Annette Hautli-Janisz, Gabriella Lapesa, Julia Romberg
Venues:
DELITE | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
18–28
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-delite-03
DOI:
10.63317/5bp93rjt27hy
Bibkey:
Cite (ACL):
Diana Constantina Hoefels. 2026. InFACT: Benchmarking LLM Explanations Against Institutional Reasoning for Deliberation-Aware Fact-Checking. In Proceedings of The 2nd Workshop on Language-driven Deliberation Technology, pages 18–28, Mallorca, Spain. Association for Computational Linguistics.
Cite (Informal):
InFACT: Benchmarking LLM Explanations Against Institutional Reasoning for Deliberation-Aware Fact-Checking (Hoefels, DELITE 2026)
Copy Citation: