Assessing Logical Coherence of LLMs via Fine-Grained NLI

Jon Felix Apaolaza Larraya, Begoña Altuna, Aitor Soroa, Inigo Lopez-Gazpio


Abstract
Natural Language Inference (NLI) is a long-standing probe of models’ reasoning capabilities, yet it remains unclear how state-of-the-art systems represent and combine logical clauses in a way that supports robust generalization. We study directional effects in deductive NLI and introduce causal coherence, an evaluation paradigm that tests whether predictions remain consistent when the directionality of inference is reversed. Using fine-grained minimal-pair phrase data from PhrasIS, we evaluate encoder, decoder, and encoder–decoder transformers and analyze their behavior under both standard and manipulated settings. Our results show that models frequently fail to maintain logical stability when directionality varies, indicating shallow pattern matching rather than genuine clause composition. We formalize soft and hard causal coherence to disentangle directional consistency from correctness, and we provide an error analysis that highlights systematic failures involving semantic relations. Our findings suggest that deductive causal reasoning and coherence remain missing components in current transformer architectures, and that addressing them is necessary for reliable NLI.
Anthology ID:
2026.lrec-1.423
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5431–5444
Language:
External URL:
https://lrec.elra.info/lrec2026-main-423
DOI:
10.63317/4prei82n6ev9
Bibkey:
Cite (ACL):
Jon Felix Apaolaza Larraya, Begoña Altuna, Aitor Soroa, and Inigo Lopez-Gazpio. 2026. Assessing Logical Coherence of LLMs via Fine-Grained NLI. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5431–5444, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Assessing Logical Coherence of LLMs via Fine-Grained NLI (Apaolaza Larraya et al., LREC 2026)
Copy Citation: