Li Zhang
Author directoryPapers on this page may belong to the following people: Li Zhang, Li Zhang, Li Zhang, Li Zhang, Li Zhang (AWS), Li Zhang (Birmingham), Li Zhang (Google), Li Zhang (Google), Li Zhang (IBM-china), Li Zhang (Nankai), Li Zhang (Newcastle, UK), Li Zhang (State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications), Li Zhang (Teesside University), Li Zhang (China Telecom Research Institute), Li Zhang (UC San Diego), Li Zhang (UK), Li Zhang (University of Pennsylvania), Li Zhang (Wuhan)
2026
R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning
Minggui He | Yilun Liu | Shimin Tao | Hongyong Zeng | Jian Zhang | Yuanchang Luo | Li Zhang | Daimeng Wei | Weibin Meng | Osamu Yoshie
Transactions of the Association for Computational Linguistics, Volume 14
Minggui He | Yilun Liu | Shimin Tao | Hongyong Zeng | Jian Zhang | Yuanchang Luo | Li Zhang | Daimeng Wei | Weibin Meng | Osamu Yoshie
Transactions of the Association for Computational Linguistics, Volume 14
Despite recent breakthroughs in reasoning-enhanced large language models (LLMs), incorporating inference-time reasoning into application tasks such as machine translation (MT), where human translators naturally employ structured, multi-layered reasoning chain-of-thoughts (CoTs), is yet un-derexplored. Existing methods either design a fixed CoT tailored for a specific MT sub-task (e.g., literature translation), or rely on synthesizing CoTs unaligned with humans and supervised fine-tuning (SFT) prone to overfitting, limiting their adaptability to diverse translation scenarios. This paper introduces R1-Translator (R1-T1), a novel framework to achieve inference-time reasoning for general MT via reinforcement learning (RL) with human-aligned CoTs comprising six common patterns. Our approach pioneers three innovations: (1) verifying reasoning-based translation in various MT scenarios (e.g., multilingual MT, domain MT) unseen from the training phase; (2) formalizing six expert-curated CoT templates that mirror hybrid human strategies like context-aware paraphrasing and round-trip translation; and (3) enabling more flexible CoTs through an RL stage after cold-start. Both human and automatic evaluation results indicate a steady translation quality improvement in a total of 10+ languages and 40+ translation directions on Flores-101 test set and four domain-specific MT tasks, especially on the languages unseen from training.
2025
Documentation Retrieval Improves Planning Language Generation
Renxiang Wang | Li Zhang
Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics
Renxiang Wang | Li Zhang
Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics
Certain strong LLMs have shown promise for zero-shot formal planning by generating planning languages like PDDL. Yet, performance of most open-source models under 50B parameters has been reported to be close to zero due to the low-resource nature of these languages. We significantly improve their performance via a series of lightweight pipelines that integrates documentation retrieval with modular code generation and error refinement. With models like Llama-4-Maverick, our best pipeline improves plan correctness from 0% to over 80% on the common BlocksWorld domain. However, while syntactic errors are substantially reduced, semantic errors persist in more challenging domains, revealing fundamental limitations in current models’ reasoning capabilities.
Data Interpreter: An LLM Agent for Data Science
Sirui Hong | Yizhang Lin | Bang Liu | Bangbang Liu | Binhao Wu | Ceyao Zhang | Danyang Li | Jiaqi Chen | Jiayi Zhang | Jinlin Wang | Li Zhang | Lingyao Zhang | Min Yang | Mingchen Zhuge | Taicheng Guo | Tuo Zhou | Wei Tao | Robert Tang | Xiangtao Lu | Xiawu Zheng | Xinbing Liang | Yaying Fei | Yuheng Cheng | Yongxin Ni | Zhibin Gou | Zongze Xu | Yuyu Luo | Chenglin Wu
Findings of the Association for Computational Linguistics: ACL 2025
Sirui Hong | Yizhang Lin | Bang Liu | Bangbang Liu | Binhao Wu | Ceyao Zhang | Danyang Li | Jiaqi Chen | Jiayi Zhang | Jinlin Wang | Li Zhang | Lingyao Zhang | Min Yang | Mingchen Zhuge | Taicheng Guo | Tuo Zhou | Wei Tao | Robert Tang | Xiangtao Lu | Xiawu Zheng | Xinbing Liang | Yaying Fei | Yuheng Cheng | Yongxin Ni | Zhibin Gou | Zongze Xu | Yuyu Luo | Chenglin Wu
Findings of the Association for Computational Linguistics: ACL 2025
Large Language Model (LLM)-based agents have excelled in various domains but face significant challenges when applied to data science workflows due to their complex, multi-stage nature. Current LLM-based agents struggle with non-linear relationships, recursive dependencies, implicit data- and logic-dependent reasoning, and managing extensive context. In this paper, we introduce Data Interpreter, an LLM-based agent that addresses these challenges through hierarchical graph-based modeling to represent the complexity and a progressive strategy for step-by-step verification, refinement, and consistent context management. Extensive experiments confirm the effectiveness of Data Interpreter. On InfiAgent-DABench, it boosts performance by 25% (from 75.9% to 94.9%), and on machine learning and open-ended tasks, it lifts accuracy from 88% to 95% and from 60% to 97%, respectively. Moreover, our method surpasses state-of-the-art baselines by 26% on the MATH dataset. We will release the code upon publication.
Evaluating the Impact of LLM-guided Reflection on Learning Outcomes with Interactive AI-Generated Educational Podcasts
Vishnu Menon | Andy Cherney | Elizabeth B. Cloude | Li Zhang | Tiffany D. Do
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Vishnu Menon | Andy Cherney | Elizabeth B. Cloude | Li Zhang | Tiffany D. Do
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study examined whether embedding LLM-guided reflection prompts in an interactive AI-generated podcast improved learning and user experience compared to a version without prompts. Thirty-six undergraduates participated, and while learning outcomes were similar across conditions, reflection prompts reduced perceived attractiveness, highlighting a call for more research on reflective interactivity design.
Search
Fix author
Co-authors
- Jiaqi Chen 1
- Yuheng Cheng 1
- Andy Cherney 1
- Elizabeth B. Cloude 1
- Tiffany D. Do 1
- Yaying Fei 1
- Zhibin Gou 1
- Taicheng Guo 1
- Minggui He 1
- Sirui Hong 1
- Danyang Li 1
- Xinbing Liang 1
- Yizhang Lin 1
- Bang Liu 1
- Bangbang Liu 1
- Yilun Liu 1
- Xiangtao Lu 1
- Yuanchang Luo 1
- Yuyu Luo 1
- Weibin Meng 1
- Vishnu Menon 1
- Yongxin Ni 1
- Robert Tang 1
- Shimin Tao 1
- Wei Tao 1
- Jinlin Wang 1
- Renxiang Wang 1
- Daimeng Wei 1
- Binhao Wu 1
- Chenglin Wu 1
- Zongze Xu 1
- Min Yang 1
- Osamu Yoshie 1
- Hongyong Zeng 1
- Ceyao Zhang 1
- Jian Zhang 1
- Jiayi Zhang 1
- Lingyao Zhang 1
- Xiawu Zheng 1
- Tuo Zhou 1
- Mingchen Zhuge 1