Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision

Dawei Zhu; Xiyu Wei; Guangxiang Zhao; Wenhao Wu; Haosheng Zou; Junfeng Ran; XWang; Lin Sun; Xiangzheng Zhang; Sujian Li (李素建)

doi:10.18653/v1/2025.findings-emnlp.170

Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision

Dawei Zhu, Xiyu Wei, Guangxiang Zhao, Wenhao Wu, Haosheng Zou, Junfeng Ran, XWang, Lin Sun, Xiangzheng Zhang, Sujian Li

Abstract

Recent advances in Large Language Models (LLMs) have highlighted the challenge of handling long-context tasks, where models need to reason over extensive input contexts to aggregate target information. While Chain-of-Thought (CoT) prompting has shown promise for multi-step reasoning, its effectiveness for long-context scenarios remains underexplored. Through systematic investigation across diverse tasks, we demonstrate that CoT’s benefits generalize across most long-context scenarios and amplify with increasing context length. Motivated by this, we propose a process-supervised framework that teaches models to generate high-quality reasoning paths for enhanced long-context performance. Our framework incorporates a self-sampling mechanism to bootstrap reasoning paths and a novel quality assessment protocol specifically designed for long-context scenarios. This protocol evaluates both answer correctness and process reliability, with the latter decomposed into source faithfulness and intrinsic consistency components for efficient and accurate assessment. Experimental results on various long-context benchmarks demonstrate the effectiveness of our approach, achieving significant improvements over outcome supervision baselines on both in-domain tasks (+13.6/+3.8 points for LLaMA/Qwen on MuSiQue) and cross-domain generalization (+9.3/+8.1 points on average across diverse QA tasks). Our code, data and trained models will be released upon acceptance.

Anthology ID:: 2025.findings-emnlp.170
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 3197–3211
Language:
URL:: https://aclanthology.org/2025.findings-emnlp.170/
DOI:: 10.18653/v1/2025.findings-emnlp.170
Bibkey:
Cite (ACL):: Dawei Zhu, Xiyu Wei, Guangxiang Zhao, Wenhao Wu, Haosheng Zou, Junfeng Ran, XWang, Lin Sun, Xiangzheng Zhang, and Sujian Li. 2025. Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 3197–3211, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision (Zhu et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-emnlp.170.pdf
Checklist:: 2025.findings-emnlp.170.checklist.pdf

PDF Cite Search Checklist Fix data