Kaiyuan Zhang
Author directoryPapers on this page may belong to the following people: Kaiyuan Zhang, Kaiyuan Zhang, Kaiyuan Zhang, Kaiyuan Zhang
2025
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
Chenghao Yang | Yinbo Luo | Zhoufutu Wen | Qi Chu | Tao Gong | Longxiang Liu | Kaiyuan Zhang | Jianpeng Jiao | Ge Zhang | Wenhao Huang | Nenghai Yu
Findings of the Association for Computational Linguistics: EMNLP 2025
Chenghao Yang | Yinbo Luo | Zhoufutu Wen | Qi Chu | Tao Gong | Longxiang Liu | Kaiyuan Zhang | Jianpeng Jiao | Ge Zhang | Wenhao Huang | Nenghai Yu
Findings of the Association for Computational Linguistics: EMNLP 2025
Large Language Models (LLMs), e.g. ChatGPT, have been widely adopted in real-world dialogue applications. However, LLMs’ robustness, especially in handling long complex dialogue sessions, including frequent motivation transfer, sophisticated cross-turn dependency, is criticized all along. Nevertheless, no existing benchmarks can fully reflect these weaknesses. We present MARS-Bench, a Multi-turn Athletic Real-world Scenario Dialogue Benchmark, designed to remedy the gap. MARS-Bench is constructed from play-by-play text commentary so to feature realistic dialogues specifically designed to evaluate three critical aspects of multi-turn conversations: ultra multi-turn, interactive multi-turn, and cross-turn tasks. Extensive experiments on MARS-Bench also reveal that closed-source LLMs significantly outperform open-source alternatives, explicit reasoning significantly boosts LLMs’ robustness on handling long complex dialogue sessions, and LLMs indeed face significant challenge when handling motivation transfer and sophisticated cross-turn dependency. Moreover, we provide mechanistic interpretability on how attention sinks due to special tokens lead to LLMs’ performance degradation when handling long complex dialogue sessions based on attention visualization experiment in Qwen2.5-7B-Instruction.
System Report for CCL25-Eval Task 8: Structured ICD Coding with LLM-Augmented Learning and Group-specific Classifiers
Bo Wang | Kaiyuan Zhang | Chong Feng | Ge Shi | Jinhua Ye | Jiahao Teng | Shouzhen Wang | Fanqing Meng | Changsen Yuan | Yan Zhuang
Proceedings of the 24th China National Conference on Computational Linguistics (CCL 2025)
Bo Wang | Kaiyuan Zhang | Chong Feng | Ge Shi | Jinhua Ye | Jiahao Teng | Shouzhen Wang | Fanqing Meng | Changsen Yuan | Yan Zhuang
Proceedings of the 24th China National Conference on Computational Linguistics (CCL 2025)
“The International Classification of Diseases (ICD) provides a standardized framework for encoding diagnoses, serving critical roles in clinical scenarios. Automatic ICD coding aims to assign formalized diagnostic codes to medical records for documentation and analysis, which is challenged by an extremely large and imbalanced label space, noisy and heterogeneous clinical text,and the need for interpretability. In this paper, we propose a structured multi-class classification framework that partitions diseases into clinically coherent groups, enabling group-specific dataaugmentation and supervision. Our method combines input compression with generative and discriminative fine-tuning strategies tailored to primary and secondary diagnoses, respectively.On the CCL2025-Eval Task 8 benchmark for Chinese electronic medical records, our approach ranked first in the final evaluation.”