Jin Ke
Author directory2026
MdEval: Massively Multilingual Code Debugging
Shukai Liu | Linzheng Chai | Jian Yang | Jiajun Shi | He Zhu | Liran Wang | Jin Ke | Wei Zhang | Hualei Zhu | Shuyue Guo | Tao Sun | Jiaheng Liu | Yunlong Duan | Yu Hao | Liqun Yang | Guanglin Niu | Ge Zhang | Zhoujun Li
Findings of the Association for Computational Linguistics: ACL 2026
Shukai Liu | Linzheng Chai | Jian Yang | Jiajun Shi | He Zhu | Liran Wang | Jin Ke | Wei Zhang | Hualei Zhu | Shuyue Guo | Tao Sun | Jiaheng Liu | Yunlong Duan | Yu Hao | Liqun Yang | Guanglin Niu | Ge Zhang | Zhoujun Li
Findings of the Association for Computational Linguistics: ACL 2026
Code large language models (LLMs) have made significant progress in code debugging by directly generating the correct code based on the buggy code snippet. Programming benchmarks, typically consisting of buggy code snippets and their associated test cases, are used to assess the debugging capabilities of LLMs. However, many existing benchmarks primarily focus on Python and are often limited in terms of language diversity (e.g., DebugBench and DebugEval). To advancethe field of multilingual debugging with LLMs, we propose the first massively multilingual debugging benchmark, which includes 3.9K test samples of 20 programming languages and covers the automated program repair (APR) task, the bug localization(BL) task, and the bug identification (BI) task. In addition, we introduce the debugging instruction corpora MdEval-Instruct by injecting bugs into the correct multilingual queries and solutions (xDebugGen). Further, a multilingual debugger xDebugCoder trained on MdEval-Instruct as a strong baseline specifically to handle bugs of a wide range of programming languages (e.g. “Missing Mut” in language Rust and “Misused Macro Definition” in language C). Our extensive experiments on MdEval reveal a notable performance gap between open-source and closed-source LLMs (e.g., GPT and Claudeseries), highlighting huge room for improvement in multilingual code debugging scenarios.
2025
CodeArena: Evaluating and Aligning CodeLLMs on Human Preference
Jian Yang | Jiaxi Yang | Wei Zhang | Jin Ke | Yibo Miao | Lei Zhang | Liqun Yang | Zeyu Cui | Yichang Zhang | Zhoujun Li | Binyuan Hui | Junyang Lin
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jian Yang | Jiaxi Yang | Wei Zhang | Jin Ke | Yibo Miao | Lei Zhang | Liqun Yang | Zeyu Cui | Yichang Zhang | Zhoujun Li | Binyuan Hui | Junyang Lin
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
We present CodeArena to emulate the complexity/diversity of real-world coding tasks, spanning 40 categories and 44 PLs. A 20B diverse synthetic instruction corpus is created by scaling instructions to help Qwen2.5-SynCoder achieve SOTA performance. Abstract: Code large language models (codeLLMs) have made significant strides in code generation. Most previous code-related benchmarks, which consist of various programming exercises along with the corresponding test cases, are used as a common measure to evaluate the performance and capabilities of code LLMs. However, the current code LLMs focus on synthesizing the correct code snippet, ignoring the alignment with human preferences, where the query should be sampled from the practical application scenarios and the model-generated responses should satisfy the human preference. To bridge the gap between the model-generated response and human preference, we present a rigorous human-curated benchmark CodeArena to emulate the complexity and diversity of real-world coding tasks, where 397 high-quality samples spanning 40 categories and 44 programming languages, carefully curated from user queries. Further, we propose a diverse synthetic instruction corpus SynCode-Instruct (nearly 20B tokens) by scaling instructions from the website to verify the effectiveness of the large-scale synthetic instruction fine-tuning, where Qwen2.5-SynCoder totally trained on synthetic instruction data can achieve top-tier performance of open-source code LLMs. The results find performance differences between execution-based benchmarks and CodeArena. Our systematic experiments of CodeArena on 40+ LLMs reveal a notable performance gap between open SOTA code LLMs (e.g. Qwen2.5-Coder) and proprietary LLMs (e.g., OpenAI o1), underscoring the importance of the human preference alignment.
M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
Jiaheng Liu | Ken Deng | Congnan Liu | Jian Yang | Shukai Liu | He Zhu | Peng Zhao | Linzheng Chai | Yanan Wu | Jin Ke | Ge Zhang | Zekun Moore Wang | Guoan Zhang | Yingshui Tan | Bangyu Xiang | Zhaoxiang Zhang | Wenbo Su | Bo Zheng
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Jiaheng Liu | Ken Deng | Congnan Liu | Jian Yang | Shukai Liu | He Zhu | Peng Zhao | Linzheng Chai | Yanan Wu | Jin Ke | Ge Zhang | Zekun Moore Wang | Guoan Zhang | Yingshui Tan | Bangyu Xiang | Zhaoxiang Zhang | Wenbo Su | Bo Zheng
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Repository-level code completion has drawn great attention in software engineering, and several benchmarks have been introduced. However, existing repository-level code completion benchmarks usually focus on a limited number of languages (<5), which cannot evaluate the general code intelligence abilities across different languages for existing code Large Language Models (LLMs). Besides, the existing benchmarks usually report overall average scores of different languages, where the fine-grained abilities in different completion scenarios are ignored. Therefore, to facilitate the research of code LLMs in multilingual scenarios, we propose a massively multilingual repository-level code completion benchmark covering 18 programming languages (called M2RC-EVAL), and two types of fine-grained annotations (i.e., bucket-level and semantic-level) on different completion scenarios are provided, where we obtain these annotations based on the parsed abstract syntax tree. Moreover, we also curate a massively multilingual instruction corpora M2RC-INSTRUCT dataset to improve the repository-level code completion abilities of existing code LLMs. Comprehensive experimental results demonstrate the effectiveness of our M2RC-EVAL and M2RC-INSTRUCT.
Search
Fix author
Co-authors
- Jian Yang 3
- Linzheng Chai 2
- Zhoujun Li 2
- Jiaheng Liu 2
- Shukai Liu 2
- Liqun Yang 2
- Ge Zhang 2
- Wei Zhang 2
- He Zhu 2
- Zeyu Cui 1
- Ken Deng 1
- Yunlong Duan 1
- Shuyue Guo 1
- Yu Hao 1
- Binyuan Hui 1
- Junyang Lin 1
- Congnan Liu 1
- Yibo Miao 1
- Guanglin Niu 1
- Jiajun Shi 1
- Wenbo Su 1
- Tao Sun 1
- Yingshui Tan 1
- Liran Wang 1
- Zekun Moore Wang 1
- Yanan Wu 1
- Bangyu Xiang 1
- Jiaxi Yang 1
- Guoan Zhang 1
- Lei Zhang 1
- Yichang Zhang 1
- Zhaoxiang Zhang 1
- Peng Zhao 1
- Bo Zheng 1
- Hualei Zhu 1