Zheng Li
Author directoryOther people with similar names: Zheng Li, Zheng Li, Zheng Li
Unverified author pages with similar names: Zheng Li
2026
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
Zheng Li | Qingxiu Dong | Jingyuan Ma | Di Zhang | Kai Jia | Zhifang Sui
Findings of the Association for Computational Linguistics: ACL 2026
Zheng Li | Qingxiu Dong | Jingyuan Ma | Di Zhang | Kai Jia | Zhifang Sui
Findings of the Association for Computational Linguistics: ACL 2026
Recently, large reasoning models demonstrate exceptional performance on various tasks. However, reasoning models always consume excessive tokens even for simple queries, leading to resource waste and prolonged user latency. To address this challenge, we propose SelfBudgeter - a self-adaptive reasoning strategy for efficient and controllable reasoning. Specifically, we first train the model to self-estimate the required reasoning budget based on the query. We then introduce budget-guided GRPO for reinforcement learning, which effectively maintains accuracy while reducing output length. Experimental results demonstrate that SelfBudgeter dynamically allocates budgets according to problem complexity, achieving an average response length compression of 61% on math reasoning tasks while maintaining accuracy. Furthermore, SelfBudgeter allows users to see how long generation will take and decide whether to continue or stop. Additionally, users can directly control the reasoning length by setting token budgets upfront.
HAUNTATTACK: When Attack Follows Reasoning as a Shadow
Jingyuan Ma | Rui Li | Zheng Li | Junfeng Liu | Heming Xia | Lei Sha | Zhifang Sui
Findings of the Association for Computational Linguistics: ACL 2026
Jingyuan Ma | Rui Li | Zheng Li | Junfeng Liu | Heming Xia | Lei Sha | Zhifang Sui
Findings of the Association for Computational Linguistics: ACL 2026
Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However, the enhancement of reasoning abilities and the exposure of internal reasoning processes introduce new safety vulnerabilities. A critical question arises: when reasoning becomes intertwined with harmfulness, will LRMs become more vulnerable to jailbreaks in reasoning mode? To investigate this, we introduce HauntAttack, a novel and general-purpose black-box adversarial attack framework that systematically embeds harmful instructions into reasoning questions. Specifically, we modify key reasoning conditions in existing questions with harmful instructions, thereby constructing a reasoning pathway that guides the model step by step toward unsafe outputs. We evaluate HauntAttack on 11 LRMs and observe an average attack success rate of over 70%, achieving up to 13 percentage points of absolute improvement over the strongest prior baseline. Our further analysis reveals that even advanced safety-aligned models remain highly susceptible to reasoning-based attacks, offering insights into the urgent challenge of balancing reasoning capability and safety in future model development.
2024
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
Yixin Yang | Zheng Li | Qingxiu Dong | Heming Xia | Zhifang Sui
Findings of the Association for Computational Linguistics: ACL 2024
Yixin Yang | Zheng Li | Qingxiu Dong | Heming Xia | Zhifang Sui
Findings of the Association for Computational Linguistics: ACL 2024
Understanding the deep semantics of images is essential in the era dominated by social media. However, current research works primarily on the superficial description of images, revealing a notable deficiency in the systematic investigation of the inherent deep semantics. In this work, we introduce DEEPEVAL, a comprehensive benchmark to assess Large Multimodal Models’ (LMMs) capacities of visual deep semantics. DEEPEVAL includes human-annotated dataset and three progressive subtasks: fine-grained description selection, in-depth title matching, and deep semantics understanding. Utilizing DEEPEVAL, we evaluate 9 open-source LMMs and GPT-4V(ision). Our evaluation demonstrates a substantial gap between the deep semantic comprehension capabilities of existing LMMs and humans. For example, GPT-4V is 30% behind humans in understanding deep semantics, even though it achieves human-comparable performance in image description. Further analysis reveals that LMM performance on DEEPEVAL varies according to the specific facets of deep semantics explored, indicating the fundamental challenges remaining in developing LMMs.