Jie Wu
Other people with similar names: Jie Wu, Jie Wu, Jie Wu
Unverified author pages with similar names: Jie Wu
2026
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
Keke Lian | Wang Bin | Lei Zhang | Libo Chen | Junjie Wang | Ziming Zhao | Yujiu Yang | Miaoqian Lin | Haotong Duan | Haoran Zhao | Shuang Liao | Mingda Guo | Quan Jiazheng | Yilu Zhong | Chenhao He | Chen Zichuan | Jie Wu | Haoling Li | Zhaoxuan Li | Jiongchi Yu | Hui LI | Dong Zhang
Findings of the Association for Computational Linguistics: ACL 2026
Keke Lian | Wang Bin | Lei Zhang | Libo Chen | Junjie Wang | Ziming Zhao | Yujiu Yang | Miaoqian Lin | Haotong Duan | Haoran Zhao | Shuang Liao | Mingda Guo | Quan Jiazheng | Yilu Zhong | Chenhao He | Chen Zichuan | Jie Wu | Haoling Li | Zhaoxuan Li | Jiongchi Yu | Hui LI | Dong Zhang
Findings of the Association for Computational Linguistics: ACL 2026
The increasing adoption of large language models (LLMs) in software engineering necessitates rigorous security evaluation of their generated code. However, existing benchmarks often lack relevance to real-world AI-assisted programming scenarios, making them inadequate for assessing the practical security risks associated with AI-generated code in production environments. To address this gap, we introduce A.S.E (AI Code Generation Security Evaluation), a repository-level evaluation benchmark designed to closely mirror real-world AI programming tasks, offering a comprehensive and reliable framework for assessing the security of AI-generated code. Our evaluation of leading LLMs on A.S.E reveals several key findings. In particular, current LLMs still struggle with secure coding. The complexity in repository-level scenarios presents challenges for LLMs that typically perform well on snippet-level tasks. Moreover, a larger reasoning budget does not necessarily lead to better code generation. These observations offer valuable insights into the current state of AI code generation and help developers identify the most suitable models for practical tasks. They also lay the groundwork for refining LLMs to generate secure and efficient code in real-world applications.
2025
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
Jiani Guo | Zuchao Li | Jie Wu | Qianren Wang | Yun Li | Lefei Zhang | Hai Zhao | Yujiu Yang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jiani Guo | Zuchao Li | Jie Wu | Qianren Wang | Yun Li | Lefei Zhang | Hai Zhao | Yujiu Yang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Large Language Models (LLMs), constrained by limited context windows, often face significant performance degradation when reasoning over long contexts. To address this, Retrieval-Augmented Generation (RAG) retrieves and reasons over chunks but frequently sacrifices logical coherence due to its reliance on similarity-based rankings. Similarly, divide-and-conquer frameworks (DCF) split documents into small chunks for independent reasoning and aggregation. While effective for local reasoning, DCF struggles to capture long-range dependencies and risks inducing conflicts by processing chunks in isolation. To overcome these limitations, we propose ToM, a novel Tree-oriented MapReduce framework for long-context reasoning. ToM leverages the inherent hierarchical structure of long documents (e.g., main headings and subheadings) by constructing a DocTree through hierarchical semantic parsing and performing bottom-up aggregation. Using a Tree MapReduce approach, ToM enables recursive reasoning: in the Map step, rationales are generated at child nodes; in the Reduce step, these rationales are aggregated across sibling nodes to resolve conflicts or reach consensus at parent nodes. Experimental results on 70B+ LLMs show that ToM significantly outperforms existing divide-and-conquer frameworks and retrieval-augmented generation methods, achieving better logical coherence and long-context reasoning.
Teaching Your Models to Understand Code via Focal Preference Alignment
Jie Wu | Haoling Li | Xin Zhang | Xiao Liu | Yangyu Huang | Jianwen Luo | Yizhen Zhang | Zuchao Li | Ruihang Chu | Yujiu Yang | Scarlett Li
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jie Wu | Haoling Li | Xin Zhang | Xiao Liu | Yangyu Huang | Jianwen Luo | Yizhen Zhang | Zuchao Li | Ruihang Chu | Yujiu Yang | Scarlett Li
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Preference learning extends the performance of Code LLMs beyond traditional supervised fine-tuning by leveraging relative quality comparisons. In existing approaches, a set of n candidate solutions is evaluated based on test case success rates, with the candidate demonstrating a higher pass rate being labeled as positive and its counterpart with a lower pass rate as negative. However, because this approach aligns entire failing code blocks rather than pinpointing specific errors, it lacks the granularity necessary to capture meaningful error-correction relationships. As a result, the model is unable to learn more informative error-correction patterns. To address these issues, we propose Target-DPO, a new preference alignment framework that mimics human iterative debugging to refine Code LLMs. Target-DPO explicitly locates error regions and aligns the corresponding tokens via a tailored DPO algorithm. To facilitate it, we introduce the CodeFlow dataset, where samples are iteratively refined until passing tests, with modifications capturing error corrections. Extensive experiments show that a diverse suite of Code LLMs equipped with Target-DPO achieves significant performance gains in code generation and improves on challenging tasks like BigCodeBench. In-depth analysis reveals that Target-DPO yields fewer errors. Code, model and datasets are in: https://github.com/JieWu02/Target-DPO.
Search
Fix author
Co-authors
- Yujiu Yang 3
- Haoling Li 2
- Zuchao Li 2
- Wang Bin 1
- Libo Chen 1
- Ruihang Chu 1
- Haotong Duan 1
- Jiani Guo 1
- Mingda Guo 1
- Chenhao He 1
- Yangyu Huang 1
- Quan Jiazheng 1
- Hui LI 1
- Scarlett Li 1
- Yun Li 1
- Zhaoxuan Li 1
- Keke Lian 1
- Shuang Liao 1
- Miaoqian Lin 1
- Xiao Liu 1
- Jianwen Luo 1
- Junjie Wang 1
- Qianren Wang 1
- Jiongchi Yu 1
- Dong Zhang 1
- Lefei Zhang 1
- Lei Zhang 1
- Xin Zhang 1
- Yizhen Zhang 1
- Hai Zhao 1
- Haoran Zhao 1
- Ziming Zhao 1
- Yilu Zhong 1
- Chen Zichuan 1