CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

Jungmin Yun, June Hyoung Kwon, Youngbin Kim


Abstract
Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks two critical vulnerabilities: (1) reliance on internal parametric knowledge rather than adherence to the provided context, and (2) exploitation of dataset shortcuts, such as single-document cues or type-matching, that diminish the need for genuine evidence aggregation across multiple documents. We introduce CRiT-QA (Counterfactual Reasoning with Traps), a dataset explicitly designed to address both limitations. To neutralize reliance on memorized knowledge and enforce strict context dependency, CRiT-QA transforms factual reasoning chains with counterfactual entities. Furthermore, it injects multi-anchor distractor chains, plausible but incorrect reasoning paths that diverge at different hops. These traps require models to follow the entire reasoning process rather than exploiting shallow heuristics. Our experiments show that LLMs exhibit substantial performance degradation on CRiT-QA compared to standard datasets, exposing their vulnerability to counterfactual conditions and distractor traps. CRiT-QA thus serves as a rigorous diagnostic tool for evaluating genuine multi-hop reasoning and provides a foundation for developing more reliable, evidence-grounded LLMs.
Anthology ID:
2026.lrec-1.410
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5246–5255
Language:
External URL:
https://lrec.elra.info/lrec2026-main-410
DOI:
10.63317/2jvxj7kwecuo
Bibkey:
Cite (ACL):
Jungmin Yun, June Hyoung Kwon, and Youngbin Kim. 2026. CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5246–5255, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps (Yun et al., LREC 2026)
Copy Citation: