Tianyu Liu
Author directoryOther people with similar names: Tianyu Liu, Tianyu Liu, Tianyu Liu
Unverified author pages with similar names: Tianyu Liu
2026
Multi-Docker-Eval: A ‘Shovel of the Gold Rush’ Benchmark on Automatic Environment Building for Software Engineering
Kelin Fu | Tianyu Liu | Zeyu Shang | Yingwei MA | Jiaheng Liu | Jian Yang | Kaigui Bian
Findings of the Association for Computational Linguistics: ACL 2026
Kelin Fu | Tianyu Liu | Zeyu Shang | Yingwei MA | Jiaheng Liu | Jian Yang | Kaigui Bian
Findings of the Association for Computational Linguistics: ACL 2026
Automated environment configuration is a critical bottleneck in scaling software engineering (SWE) automation. To provide a reliable evaluation standard for this task, we present Multi-Docker-Eval benchmark. It includes 40 real-world repositories spanning 9 programming languages and measures both success in achieving executable states and efficiency under realistic constraints. Our extensive evaluation of state-of-the-art LLMs and agent frameworks reveals key insights: (1) the overall success rate of current models is low (F2P at most 37.7%), with environment construction being the primary bottleneck; (2) model size and reasoning length are not decisive factors, and open-source models like DeepSeek-V3.1 and Kimi-K2 are competitive in both efficiency and effectiveness; (3) agent framework and programming language also have significantly influence on success rate. These findings provide actionable guidelines for building scalable, fully automated SWE pipelines.
2025
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
Jinsheng Huang | Liang Chen | Taian Guo | Fu Zeng | Yusheng Zhao | Bohan Wu | Ye Yuan | Haozhe Zhao | Zhihui Guo | Yichi Zhang | Jingyang Yuan | Wei Ju | Luchen Liu | Tianyu Liu | Baobao Chang | Ming Zhang
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Jinsheng Huang | Liang Chen | Taian Guo | Fu Zeng | Yusheng Zhao | Bohan Wu | Ye Yuan | Haozhe Zhao | Zhihui Guo | Yichi Zhang | Jingyang Yuan | Wei Ju | Luchen Liu | Tianyu Liu | Baobao Chang | Ming Zhang
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, often assessed through multiple-choice questions (MCQs) that include an image, a question, and several options. However, many benchmarks used for such evaluations suffer from systematic biases. Remarkably, Large Language Models (LLMs) without any visual perception capabilities achieve non-trivial performance, undermining the credibility of these evaluations. To address this issue while maintaining the efficiency of MCQ evaluations, we propose MMEVALPRO, a benchmark designed to avoid Type-I errors through a trilogy evaluation pipeline and more rigorous metrics. For each original question from existing benchmarks, human annotators augment it by creating one perception question and one knowledge anchor question through a meticulous annotation process. MMEVALPRO comprises 2,138 question triplets, totaling 6,414 distinct questions. Two-thirds of these questions are manually labeled by human experts, while the rest are sourced from existing benchmarks (MMMU, ScienceQA, and MathVista). Compared with the existing benchmarks, our experiments with the latest LLMs and LMMs demonstrate that MMEVALPRO is more challenging (the best LMM lags behind human performance by 31.73%, compared to an average gap of 8.03% in previous benchmarks) and more trustworthy (the best LLM trails the best LMM by 23.09%, whereas the gap for previous benchmarks is just 14.64%). Our in-depth analysis explains the reason for the large performance gap and justifies the trustworthiness of evaluation, underscoring its significant potential for advancing future research.
LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
Bofei Gao | Zefan Cai | Runxin Xu | Peiyi Wang | Ce Zheng | Runji Lin | Keming Lu | Dayiheng Liu | Chang Zhou | Wen Xiao | Tianyu Liu | Baobao Chang
Findings of the Association for Computational Linguistics: ACL 2025
Bofei Gao | Zefan Cai | Runxin Xu | Peiyi Wang | Ce Zheng | Runji Lin | Keming Lu | Dayiheng Liu | Chang Zhou | Wen Xiao | Tianyu Liu | Baobao Chang
Findings of the Association for Computational Linguistics: ACL 2025
In recent progress, mathematical verifiers have achieved success in mathematical reasoning tasks by validating the correctness of solutions generated by policy models. However, existing verifiers are trained with binary classification labels, which are not informative enough for the model to accurately assess the solutions. To mitigate the aforementioned insufficiency of binary labels, we introduce step-wise natural language feedback as rationale labels, that is, the correctness of each step and the detailed explanations. In this paper, we propose Math-Minos, a natural language feedback-enhanced verifier by constructing automatically generated training data and a two-stage training paradigm for effective training and efficient inference. Our experiments reveal that a small set of natural language feedback can significantly boost the performance of the verifier in both verification and reinforcement learning and also significantly alleviates the data-demanding problems of the reward model with an over 700% data efficiency improvement.
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
Feifan Song | Shaohang Wei | Wen Luo | Yuxuan Fan | Tianyu Liu | Guoyin Wang | Houfeng Wang
Findings of the Association for Computational Linguistics: ACL 2025
Feifan Song | Shaohang Wei | Wen Luo | Yuxuan Fan | Tianyu Liu | Guoyin Wang | Houfeng Wang
Findings of the Association for Computational Linguistics: ACL 2025
Large Language Models (LLMs) require alignment with human preferences to avoid generating offensive, false, or meaningless content. Recently, low-resource methods for LLM alignment have been popular, while still facing challenges in obtaining both high-quality and aligned content. Motivated by the observation that the difficulty of generating aligned responses is concentrated at the beginning of decoding, we propose a novel framework, Weak-to-Strong Decoding (WSD), to enhance the alignment ability of base models by the guidance of a small aligned model. The small model first drafts well-aligned beginnings, followed by the large base model to continue the rest, controlled by a well-designed auto-switch mechanism. We also collect a new dataset, GenerAlign, to fine-tune a small-sized Pilot-3B as the draft model, which effectively enhances different base models under the WSD framework to outperform all baseline methods, while avoiding degradation on downstream tasks, termed as the alignment tax. Extensive experiments are further conducted to examine the impact of different settings and time efficiency, as well as analyses on the intrinsic mechanisms of WSD in depth.
Towards A Better Initial Policy Model For Scalable Long-CoT Reinforcement Learning
Bofei Gao | Yejie Wang | Yibo Miao | Ruoyu Wu | Feifan Song | Longhui Yu | Tianyu Liu | Baobao Chang
Findings of the Association for Computational Linguistics: ACL 2025
Bofei Gao | Yejie Wang | Yibo Miao | Ruoyu Wu | Feifan Song | Longhui Yu | Tianyu Liu | Baobao Chang
Findings of the Association for Computational Linguistics: ACL 2025
Long-CoT reasoning combined with reinforcement learning for large language models demonstrates remarkable performance and scalability. However, we observe that the initial policy model could significantly influence the final performance as well as the token efficiency. Additionally, there is a lack of systematic guidelines for obtaining a better initial policy model. To bridge this gap, we initiate a comprehensive investigation by activating the initial model using a variety of datasets with different data volumes and reasoning patterns. Then, we conduct a thorough analysis and comparison of the RL process for different initial models from the perspectives of upper bounds, diversity, and token efficiency, providing a deeper understanding and insight into the long-CoT RL. Based on our empirical results, we propose a systematic guideline and a novel Re-RFT method for constructing a better RL start point. Our experiment results based on the 14B model surpass the DeepSeek-R1-Distill-Qwen-14B by an average of 4.6%, demonstrating our approach’s effectiveness and superiority.
CodeV: Issue Resolving with Visual Data
Linhao Zhang | Daoguang Zan | Quanshun Yang | Zhirong Huang | Dong Chen | Bo Shen | Tianyu Liu | Yongshun Gong | Huang Pengjie | Xudong Lu | Guangtai Liang | Lizhen Cui | Qianxiang Wang
Findings of the Association for Computational Linguistics: ACL 2025
Linhao Zhang | Daoguang Zan | Quanshun Yang | Zhirong Huang | Dong Chen | Bo Shen | Tianyu Liu | Yongshun Gong | Huang Pengjie | Xudong Lu | Guangtai Liang | Lizhen Cui | Qianxiang Wang
Findings of the Association for Computational Linguistics: ACL 2025
Large Language Models (LLMs) have advanced rapidly in recent years, with their applications in software engineering expanding to more complex repository-level tasks. GitHub issue resolving is a key challenge among these tasks. While recent approaches have made progress on this task, they focus on textual data within issues, neglecting visual data. However, this visual data is crucial for resolving issues as it conveys additional knowledge that text alone cannot. We propose CodeV, the first approach to leveraging visual data to enhance the issue-resolving capabilities of LLMs. CodeV resolves each issue by following a two-phase process: data processing and patch generation. To evaluate CodeV, we construct a benchmark for visual issue resolving, namely Visual SWE-bench. Through extensive experiments, we demonstrate the effectiveness of CodeV, as well as provide valuable insights into leveraging visual data to resolve GitHub issues.
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web
Hongcheng Guo | Wei Zhang | Junhao Chen | Yaonan Gu | Jian Yang | Junjia Du | Shaosheng Cao | Binyuan Hui | Tianyu Liu | Jianxin Ma | Chang Zhou | Zhoujun Li
Findings of the Association for Computational Linguistics: ACL 2025
Hongcheng Guo | Wei Zhang | Junhao Chen | Yaonan Gu | Jian Yang | Junjia Du | Shaosheng Cao | Binyuan Hui | Tianyu Liu | Jianxin Ma | Chang Zhou | Zhoujun Li
Findings of the Association for Computational Linguistics: ACL 2025
Recently, advancements in large multimodal models have led to significant strides in image comprehension capabilities. Despite these advancements, there is a lack of a robust benchmark specifically for assessing the image‐to‐web conversion proficiency of these large models. It is essential to ensure the integrity of the web elements generated, which comprise both visible and invisible categories. Previous evaluation methods (e.g., BLEU) are notably susceptible to significant alterations due to the presence of invisible elements. Furthermore, it is crucial to measure the layout information of web pages—i.e., the positional relationships between elements—which has been overlooked by prior work. To address these challenges, we have curated and aligned a benchmark of images and corresponding web codes (IW-bench). Specifically, we propose Element Accuracy, which tests the completeness of elements by parsing the Document Object Model (DOM) tree. We also introduce Layout Accuracy to analyze positional relationships by converting the DOM tree into a common subsequence. In addition, we design a five‐hop multimodal Chain‐of‐Thought prompting strategy for improved performance, consisting of: 1) SoM prompt injection, 2) inferring elements, 3) inferring layout, 4) inferring web code, and 5) reflection. Our benchmark comprises 1,200 image–code pairs with varying levels of difficulty. We have conducted extensive experiments on existing large multimodal models, providing insights into their performance and identifying areas for improvement in the image‐to‐web domain.
Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction Tuning
Jian Yang | Wei Zhang | Yibo Miao | Shanghaoran Quan | Zhenhe Wu | Qiyao Peng | Liqun Yang | Tianyu Liu | Zeyu Cui | Binyuan Hui | Junyang Lin
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Jian Yang | Wei Zhang | Yibo Miao | Shanghaoran Quan | Zhenhe Wu | Qiyao Peng | Liqun Yang | Tianyu Liu | Zeyu Cui | Binyuan Hui | Junyang Lin
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Recent advancement in code understanding and generation demonstrates that code LLMs fine-tuned on a high-quality instruction dataset can gain powerful capabilities to address wide-ranging code-related tasks. However, most previous existing methods mainly view each programming language in isolation and ignore the knowledge transfer among different programming languages. To bridge the gap among different programming languages, we introduce a novel multi-agent collaboration framework to enhance multilingual instruction tuning for code LLMs, where multiple language-specific intelligent agent components with generation memory work together to transfer knowledge from one language to another efficiently and effectively. Specifically, we first generate the language-specific instruction data from the code snippets and then provide the generated data as the seed data for language-specific agents. Multiple language-specific agents discuss and collaborate to formulate a new instruction and its corresponding solution (A new programming language or existing programming language), To further encourage the cross-lingual transfer, each agent stores its generation history as memory and then summarizes its merits and faults. Finally, the high-quality multilingual instruction data is used to encourage knowledge transfer among different programming languages to train Qwen2.5-xCoder. Experimental results on multilingual programming benchmarks demonstrate the superior performance of Qwen2.5-xCoder in sharing common knowledge, highlighting its potential to reduce the cross-lingual gap.
Search
Fix author
Co-authors
- Baobao Chang (常宝宝) 3
- Jian Yang 3
- Bofei Gao 2
- Binyuan Hui 2
- Yibo Miao 2
- Feifan Song 2
- Wei Zhang 2
- Chang Zhou 2
- Kaigui Bian 1
- Zefan Cai 1
- Shaosheng Cao 1
- Dong Chen 1
- Junhao Chen 1
- Liang Chen 1
- Lizhen Cui 1
- Zeyu Cui 1
- Junjia Du 1
- Yuxuan Fan 1
- Kelin Fu 1
- Yongshun Gong 1
- Yaonan Gu 1
- Hongcheng Guo 1
- Taian Guo 1
- Zhihui Guo 1
- Jinsheng Huang 1
- Zhirong Huang 1
- Wei Ju 1
- Zhoujun Li 1
- Guangtai Liang 1
- Junyang Lin 1
- Runji Lin 1
- Dayiheng Liu 1
- Jiaheng Liu 1
- Luchen Liu 1
- Keming Lu 1
- Xudong Lu 1
- Wen Luo 1
- Yingwei MA 1
- Jianxin Ma 1
- Qiyao Peng 1
- Huang Pengjie 1
- Shanghaoran Quan 1
- Zeyu Shang 1
- Bo Shen 1
- Guoyin Wang 1
- Houfeng Wang 1
- Peiyi Wang (王培懿) 1
- Qianxiang Wang 1
- Yejie Wang 1
- Shaohang Wei 1
- Bohan Wu 1
- Ruoyu Wu 1
- Zhenhe Wu 1
- Wen Xiao 1
- Runxin Xu 1
- Liqun Yang 1
- Quanshun Yang 1
- Longhui Yu 1
- Jingyang Yuan 1
- Ye Yuan 1
- Daoguang Zan 1
- Fu Zeng 1
- Linhao Zhang 1
- Ming Zhang 1
- Yichi Zhang 1
- Haozhe Zhao 1
- Yusheng Zhao 1
- Ce Zheng 1