Wenting Zhao
Author directoryOther people with similar names: Wenting Zhao
Unverified author pages with similar names: Wenting Zhao
2026
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
Yibo Wang | Congying Xia | Wenting Zhao | Jiangshu Du | Chunyu Miao | Zhongfen Deng | Philip S. Yu | Chen Xing
Findings of the Association for Computational Linguistics: ACL 2026
Yibo Wang | Congying Xia | Wenting Zhao | Jiangshu Du | Chunyu Miao | Zhongfen Deng | Philip S. Yu | Chen Xing
Findings of the Association for Computational Linguistics: ACL 2026
Unit test generation has become a promising and important use case for Large Language Models (LLMs). However, existing evaluation benchmarks for LLM unit test generation primarily focus on function- or class-level code (single-file) rather than on more practical, challenging multi-file codebases.To address this limitation, we propose MultiFileTest, a multi-file-level benchmark for unit test generation covering Python, Java, and JavaScript. MultiFileTest features 20 high-quality, moderate-sized projects per language. We evaluate eleven frontier LLMs on MultiFileTest, and the results show that most tested LLMs exhibit moderate performance on MultiFileTest, highlighting the benchmark’s inherent difficulty.We also conduct a thorough error analysis, which shows that even advanced LLMs, such as Gemini 3.0 Pro, exhibit basic yet critical errors, including executability and cascade errors. Motivated by this observation, we further evaluate these frontier LLMs under manual error-fixing and self-error-fixing scenarios to assess their potential when equipped with error-fixing mechanisms.Our dataset is available at MultiFileTest.
2024
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
Jiangshu Du | Yibo Wang | Wenting Zhao | Zhongfen Deng | Shuaiqi Liu | Renze Lou | Henry Peng Zou | Pranav Narayanan Venkit | Nan Zhang | Mukund Srinath | Haoran Ranran Zhang | Vipul Gupta | Yinghui Li | Tao Li | Fei Wang | Qin Liu | Tianlin Liu | Pengzhi Gao | Congying Xia | Chen Xing | Cheng Jiayang | Zhaowei Wang | Ying Su | Raj Sanjay Shah | Ruohao Guo | Jing Gu | Haoran Li | Kangda Wei | Zihao Wang | Lu Cheng | Surangika Ranathunga | Meng Fang | Jie Fu | Fei Liu | Ruihong Huang | Eduardo Blanco | Yixin Cao | Rui Zhang | Philip S. Yu | Wenpeng Yin
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Jiangshu Du | Yibo Wang | Wenting Zhao | Zhongfen Deng | Shuaiqi Liu | Renze Lou | Henry Peng Zou | Pranav Narayanan Venkit | Nan Zhang | Mukund Srinath | Haoran Ranran Zhang | Vipul Gupta | Yinghui Li | Tao Li | Fei Wang | Qin Liu | Tianlin Liu | Pengzhi Gao | Congying Xia | Chen Xing | Cheng Jiayang | Zhaowei Wang | Ying Su | Raj Sanjay Shah | Ruohao Guo | Jing Gu | Haoran Li | Kangda Wei | Zihao Wang | Lu Cheng | Surangika Ranathunga | Meng Fang | Jie Fu | Fei Liu | Ruihong Huang | Eduardo Blanco | Yixin Cao | Rui Zhang | Philip S. Yu | Wenpeng Yin
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Claim: This work is not advocating the use of LLMs for paper (meta-)reviewing. Instead, wepresent a comparative analysis to identify and distinguish LLM activities from human activities. Two research goals: i) Enable better recognition of instances when someone implicitly uses LLMs for reviewing activities; ii) Increase community awareness that LLMs, and AI in general, are currently inadequate for performing tasks that require a high level of expertise and nuanced judgment.This work is motivated by two key trends. On one hand, large language models (LLMs) have shown remarkable versatility in various generative tasks such as writing, drawing, and question answering, significantly reducing the time required for many routine tasks. On the other hand, researchers, whose work is not only time-consuming but also highly expertise-demanding, face increasing challenges as they have to spend more time reading, writing, and reviewing papers. This raises the question: how can LLMs potentially assist researchers in alleviating their heavy workload?This study focuses on the topic of LLMs as NLP Researchers, particularly examining the effectiveness of LLMs in assisting paper (meta-)reviewing and its recognizability. To address this, we constructed the ReviewCritique dataset, which includes two types of information: (i) NLP papers (initial submissions rather than camera-ready) with both human-written and LLM-generated reviews, and (ii) each review comes with “deficiency” labels and corresponding explanations for individual segments, annotated by experts. Using ReviewCritique, this study explores two threads of research questions: (i) “LLMs as Reviewers”, how do reviews generated by LLMs compare with those written by humans in terms of quality and distinguishability? (ii) “LLMs as Metareviewers”, how effectively can LLMs identify potential issues, such as Deficient or unprofessional review segments, within individual paper reviews? To our knowledge, this is the first work to provide such a comprehensive analysis.
Search
Fix author
Co-authors
- Zhongfen Deng 2
- Jiangshu Du 2
- Yibo Wang 2
- Congying Xia 2
- Chen Xing 2
- Philip S. Yu 2
- Eduardo Blanco 1
- Yixin Cao 1
- Lu Cheng 1
- Meng Fang 1
- Jie Fu 1
- Pengzhi Gao 1
- Jing Gu 1
- Ruohao Guo 1
- Vipul Gupta 1
- Ruihong Huang 1
- Cheng Jiayang 1
- Haoran Li 1
- Tao Li 1
- Yinghui Li 1
- Fei Liu 1
- Qin Liu 1
- Shuaiqi Liu 1
- Tianlin Liu 1
- Renze Lou 1
- Chunyu Miao 1
- Surangika Ranathunga 1
- Raj Sanjay Shah 1
- Mukund Srinath 1
- Ying Su 1
- Pranav Narayanan Venkit 1
- Fei Wang 1
- Zhaowei Wang 1
- Zihao Wang 1
- Kangda Wei 1
- Wenpeng Yin 1
- Haoran Ranran Zhang 1
- Nan Zhang 1
- Rui Zhang 1
- Henry Peng Zou 1