He Zhu
Author directoryOther people with similar names: He Zhu
Unverified author pages with similar names: He Zhu
2026
MdEval: Massively Multilingual Code Debugging
Shukai Liu | Linzheng Chai | Jian Yang | Jiajun Shi | He Zhu | Liran Wang | Jin Ke | Wei Zhang | Hualei Zhu | Shuyue Guo | Tao Sun | Jiaheng Liu | Yunlong Duan | Yu Hao | Liqun Yang | Guanglin Niu | Ge Zhang | Zhoujun Li
Findings of the Association for Computational Linguistics: ACL 2026
Shukai Liu | Linzheng Chai | Jian Yang | Jiajun Shi | He Zhu | Liran Wang | Jin Ke | Wei Zhang | Hualei Zhu | Shuyue Guo | Tao Sun | Jiaheng Liu | Yunlong Duan | Yu Hao | Liqun Yang | Guanglin Niu | Ge Zhang | Zhoujun Li
Findings of the Association for Computational Linguistics: ACL 2026
Code large language models (LLMs) have made significant progress in code debugging by directly generating the correct code based on the buggy code snippet. Programming benchmarks, typically consisting of buggy code snippets and their associated test cases, are used to assess the debugging capabilities of LLMs. However, many existing benchmarks primarily focus on Python and are often limited in terms of language diversity (e.g., DebugBench and DebugEval). To advancethe field of multilingual debugging with LLMs, we propose the first massively multilingual debugging benchmark, which includes 3.9K test samples of 20 programming languages and covers the automated program repair (APR) task, the bug localization(BL) task, and the bug identification (BI) task. In addition, we introduce the debugging instruction corpora MdEval-Instruct by injecting bugs into the correct multilingual queries and solutions (xDebugGen). Further, a multilingual debugger xDebugCoder trained on MdEval-Instruct as a strong baseline specifically to handle bugs of a wide range of programming languages (e.g. “Missing Mut” in language Rust and “Misused Macro Definition” in language C). Our extensive experiments on MdEval reveal a notable performance gap between open-source and closed-source LLMs (e.g., GPT and Claudeseries), highlighting huge room for improvement in multilingual code debugging scenarios.
2025
OAgents: An Empirical Study of Building Effective Agents
He Zhu | Tianrui Qin | King Zhu | Heyuan Huang | Yeyi Guan | Jinxiang Xia | Hanhao Li | Yi Yao | Ningning Wang | Pai Liu | Tianhao Peng | Xin Gui | Li Xiaowan | Yuhui Liu | Xiangru Tang | Jian Yang | Ge Zhang | Xitong Gao | Yuchen Eleanor Jiang | Changwang Zhang | Jun Wang | Jiaheng Liu | Wangchunshu Zhou
Findings of the Association for Computational Linguistics: EMNLP 2025
He Zhu | Tianrui Qin | King Zhu | Heyuan Huang | Yeyi Guan | Jinxiang Xia | Hanhao Li | Yi Yao | Ningning Wang | Pai Liu | Tianhao Peng | Xin Gui | Li Xiaowan | Yuhui Liu | Xiangru Tang | Jian Yang | Ge Zhang | Xitong Gao | Yuchen Eleanor Jiang | Changwang Zhang | Jun Wang | Jiaheng Liu | Wangchunshu Zhou
Findings of the Association for Computational Linguistics: EMNLP 2025
Recently, Agentic AI has become an increasingly popular field of research. However, we argue that current practices on agent research are far from standard, rigorous scientific research, which makes it hard to conduct apples-to-apples comparisons among and against existing methods. As a result, it is still obscure how different design choices in an agent framework impact its effectiveness, and measuring progress on agent research remains very hard. In this work, we conduct a systematic empirical study on the GAIA benchmark to investigate the impact of different popular design choices within key agent components in a fair and rigorous way. To begin with, we find that the lack of a standard evaluation protocol makes previous works, even the open-sourced ones, not reproducible, and the variance between different random runs is often non-negligible. Therefore, we first introduce a more robust evaluation protocol to make comparisons more stable. Our empirical study then unveils which components and designs, as well as correlations between these designs, are the keys for building effective agents, while others are not and redundant, despite seemingly making sense. With the insights gained from our empirical study, we build and open-source OAgents, a new foundation agent framework that achieves state-of-the-art performance among open-source projects, providing a good starting point and guidelines for building effective agents. More importantly, supports various design choices for agent components in a modularized way, facilitating future scientific research on Agentic AI.
M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
Jiaheng Liu | Ken Deng | Congnan Liu | Jian Yang | Shukai Liu | He Zhu | Peng Zhao | Linzheng Chai | Yanan Wu | Jin Ke | Ge Zhang | Zekun Moore Wang | Guoan Zhang | Yingshui Tan | Bangyu Xiang | Zhaoxiang Zhang | Wenbo Su | Bo Zheng
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Jiaheng Liu | Ken Deng | Congnan Liu | Jian Yang | Shukai Liu | He Zhu | Peng Zhao | Linzheng Chai | Yanan Wu | Jin Ke | Ge Zhang | Zekun Moore Wang | Guoan Zhang | Yingshui Tan | Bangyu Xiang | Zhaoxiang Zhang | Wenbo Su | Bo Zheng
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Repository-level code completion has drawn great attention in software engineering, and several benchmarks have been introduced. However, existing repository-level code completion benchmarks usually focus on a limited number of languages (<5), which cannot evaluate the general code intelligence abilities across different languages for existing code Large Language Models (LLMs). Besides, the existing benchmarks usually report overall average scores of different languages, where the fine-grained abilities in different completion scenarios are ignored. Therefore, to facilitate the research of code LLMs in multilingual scenarios, we propose a massively multilingual repository-level code completion benchmark covering 18 programming languages (called M2RC-EVAL), and two types of fine-grained annotations (i.e., bucket-level and semantic-level) on different completion scenarios are provided, where we obtain these annotations based on the parsed abstract syntax tree. Moreover, we also curate a massively multilingual instruction corpora M2RC-INSTRUCT dataset to improve the repository-level code completion abilities of existing code LLMs. Comprehensive experimental results demonstrate the effectiveness of our M2RC-EVAL and M2RC-INSTRUCT.
Search
Fix author
Co-authors
- Jiaheng Liu 3
- Jian Yang 3
- Ge Zhang 3
- Linzheng Chai 2
- Jin Ke 2
- Shukai Liu 2
- Ken Deng 1
- Yunlong Duan 1
- Xitong Gao 1
- Yeyi Guan 1
- Xin Gui 1
- Shuyue Guo 1
- Yu Hao 1
- Heyuan Huang 1
- Yuchen Eleanor Jiang 1
- Hanhao Li 1
- Zhoujun Li 1
- Congnan Liu 1
- Pai Liu 1
- Yuhui Liu 1
- Guanglin Niu 1
- Tianhao Peng 1
- Tianrui Qin 1
- Jiajun Shi 1
- Wenbo Su 1
- Tao Sun 1
- Yingshui Tan 1
- Xiangru Tang 1
- Jun Wang 1
- Liran Wang 1
- Ningning Wang (王宁宁) 1
- Zekun Moore Wang 1
- Yanan Wu 1
- Jinxiang Xia 1
- Bangyu Xiang 1
- Li Xiaowan 1
- Liqun Yang 1
- Yi Yao 1
- Changwang Zhang 1
- Guoan Zhang 1
- Wei Zhang 1
- Zhaoxiang Zhang 1
- Peng Zhao 1
- Bo Zheng 1
- Wangchunshu Zhou 1
- Hualei Zhu 1
- King Zhu 1