For What Reason? Interpreting Models’ Encoding of Causation and Antithesis

Abhidip Bhattacharyya, Shira Wein


Abstract
Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality. In this work, we investigate how instruction-tuned Transformer models (LLaMA and Mistral) encode discourse relations, with a focus on the somewhat opposed relations on causation and antithesis. Framing the task as a next-token prediction task and applying a suite of interpretability techniques to test model internals, our findings show that certain early layers make predictive decisions at mid-sequence tokens, while some mid-level layers finalize their decisions closer to the last token. Most of the remaining layers primarily propagate earlier decisions rather than actively influencing them. Additionally, we observe that some layers exhibit a preference for one answer over alternatives, suggesting asymmetric representation of discourse-based reasoning.
Anthology ID:
2026.sigdial-1.17
Volume:
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Month:
August
Year:
2026
Address:
Atlanta, Georgia, USA
Editors:
Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
Venue:
SIGDIAL
SIG:
SIGDIAL
Publisher:
Association for Computational Linguistics
Note:
Pages:
237–253
Language:
URL:
https://aclanthology.org/2026.sigdial-1.17/
DOI:
Bibkey:
Cite (ACL):
Abhidip Bhattacharyya and Shira Wein. 2026. For What Reason? Interpreting Models’ Encoding of Causation and Antithesis. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 237–253, Atlanta, Georgia, USA. Association for Computational Linguistics.
Cite (Informal):
For What Reason? Interpreting Models’ Encoding of Causation and Antithesis (Bhattacharyya & Wein, SIGDIAL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.sigdial-1.17.pdf