Chenglong Wang
Other people with similar names: Chenglong Wang
Unverified author pages with similar names: Chenglong Wang
2026
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
Yifu Huo | Chenglong Wang | Ziming Zhu | Shunjie Xing | Peinan Feng | Tongran Liu | Qiaozhi He | Tian Hua Zhou | Changxiaojia | JingBo Zhu | Zhengtao Yu | Tong Xiao
Findings of the Association for Computational Linguistics: ACL 2026
Yifu Huo | Chenglong Wang | Ziming Zhu | Shunjie Xing | Peinan Feng | Tongran Liu | Qiaozhi He | Tian Hua Zhou | Changxiaojia | JingBo Zhu | Zhengtao Yu | Tong Xiao
Findings of the Association for Computational Linguistics: ACL 2026
Reinforcement learning (RL) has emerged as a promising paradigm for training reasoning-oriented models by leveraging rule-based reward signals. However, RL training typically tends to improve single-sample success rates (i.e., Pass@1) while offering limited exploration of diverse reasoning trajectories, which is crucial for multi-sample performance (i.e., Pass@k). Our preliminary analysis reveals that this limitation stems from a fundamental squeezing effect, whereby probability mass is excessively concentrated on a narrow subset of high-reward trajectories, restricting genuine exploration and constraining attainable performance under RL training. To address this issue, in this work, we propose Steering Probability Squeezing (SPS), a training paradigm that interleaves conventional RL with inverse reinforcement learning (IRL). SPS treats on-policy rollouts as demonstrations and employs IRL to explicitly reshape the induced trajectory distribution, thereby enhancing exploration without introducing external supervision. Experiments on five commonly used reasoning benchmarks demonstrate that SPS can enable better exploration and improve Pass@k. Beyond algorithmic contributions, we provide an analysis of RL learning dynamics and identify an empirical upper bound on Pass@k, shedding light on intrinsic exploration limits in RL-based reasoning models. Our findings suggest that alternating between RL and IRL offers an effective pathway toward extending the exploration capacity of reasoning-oriented large language models.
SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
Chenglong Wang | Canjia Li | Xingzhao Zhu | Yifu Huo | Huiyu Wang | Weixiong Lin | Yun Yang | Qiaozhi He | Tian Hua Zhou | Changxiaojia | JingBo Zhu | Tong Xiao
Findings of the Association for Computational Linguistics: ACL 2026
Chenglong Wang | Canjia Li | Xingzhao Zhu | Yifu Huo | Huiyu Wang | Weixiong Lin | Yun Yang | Qiaozhi He | Tian Hua Zhou | Changxiaojia | JingBo Zhu | Tong Xiao
Findings of the Association for Computational Linguistics: ACL 2026
Due to the dynamically evolving nature of real-world query streams, relevance models struggle to generalize to practical search scenarios. A sophisticated solution is self-evolution techniques. However, in large-scale industrial settings with massive query streams, this technique faces two challenges: (1) informative samples are often sparse and difficult to identify, and (2) pseudo-labels generated by the current model could be unreliable. To address these challenges, in this work, we propose a Self-Evolving Relevance Model approach (SERM), which comprises two complementary multi-agent modules: a multi-agent sample miner, designed to detect distributional shifts and identify informative training samples, and a multi-agent relevance annotator, which provides reliable labels through a two-level agreement framework. We evaluated SERM on a large-scale industrial platform, which serves billions of user requests daily. Experimental results demonstrate that SERM can achieve significant performance gains through iterative self-evolution, as validated by extensive offline multilingual evaluations and online testing.
On the Emotion Understanding of Synthesized Speech
Yuan Ge | Haishu Zhao | AoKai Hao | Junxiang Zhang | Bei Li | Xiaoqian Liu | Chenglong Wang | Jianjin Wang | Bingsen Zhou | Bingyu Liu | JingBo Zhu | Zhengtao Yu | Tong Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yuan Ge | Haishu Zhao | AoKai Hao | Junxiang Zhang | Bei Li | Xiaoqian Liu | Chenglong Wang | Jianjin Wang | Bingsen Zhou | Bingyu Liu | JingBo Zhu | Zhengtao Yu | Tong Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding results a plausible reward or evaluation metric for assessing emotional expressiveness in speech synthesis. In this work, we critically examine this assumption by systematically evaluating Speech Emotion Recognition (SER) on synthesized speech across datasets, discriminative and generative SER models, and diverse synthesis models. We find that current SER models can not generalize to synthesized speech, largely because speech token prediction during synthesis induces a representation mismatch between synthesized and human speech. Moreover, generative Speech Language Models (SLMs) tend to infer emotion from textual semantics while ignoring paralinguistic cues. Overall, our findings suggest that existing SER models often exploit non-robust shortcuts rather than capturing fundamental features, and paralinguistic understanding in SLMs remains challenging.
RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment
Yingfeng Luo | Hongyu Liu | DingYang Lin | Kaiyan Chang | Chenglong Wang | Bei Li | Quan Du | Tong Xiao | JingBo Zhu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
Yingfeng Luo | Hongyu Liu | DingYang Lin | Kaiyan Chang | Chenglong Wang | Bei Li | Quan Du | Tong Xiao | JingBo Zhu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
Large Language Models (LLMs) have achieved remarkable performance in Machine Translation (MT), but deploying them at scale remains prohibitively expensive. A widely adopted remedy is the hybrid system paradigm, which balances cost and quality by serving most requests with a small model and selectively routing a fraction to a large model. However, existing routing strategies often rely on heuristics, external predictors, or absolute quality estimation, which fail to capture whether the large model actually provides a worthwhile improvement over the small one. In this paper, we formulate routing as a budget allocation problem and identify marginal gain, i.e., the large model’s improvement over the small model, as the optimal signal for budgeted decisions. Building on this, we propose RouteLMT (routing for LLM-based MT), an efficient in-model router that predicts this expected gain by probing the small translator’s prompt-token representation, without requiring external models or hypothesis decoding. Extensive experiments demonstrate that our RouteLMT outperforms heuristics, quality/difficulty estimation baselines, achieving a superior quality–budget Pareto frontier. Furthermore, we analyze regression risks and show that a simple guarded variant can mitigate severe quality losses.
2025
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
Yifu Huo | Chenglong Wang | Qiren Zhu | Shunjie Xing | Tong Xiao | Chunliang Zhang | Tongran Liu | JingBo Zhu
Findings of the Association for Computational Linguistics: EMNLP 2025
Yifu Huo | Chenglong Wang | Qiren Zhu | Shunjie Xing | Tong Xiao | Chunliang Zhang | Tongran Liu | JingBo Zhu
Findings of the Association for Computational Linguistics: EMNLP 2025
Preference optimization methods like DPO have achieved remarkable performance in LLM alignment. However, the evaluation for these methods relies on a single response and overlooks other potential outputs, which could also be generated in real-world applications within this hypothetical space. To address this issue, this paper presents a Hypothesis-based PrEference-aware AnaLysis Framework (HEAL), a novel evaluation paradigm that formulates preference alignment as a re-ranking process within hypothesis spaces. The framework incorporates two complementary metrics: ranking accuracy for evaluating ordinal consistency and preference strength correlation for assessing continuous alignment. To facilitate this framework, we develop UniHypoBench, a unified hypothesis benchmark constructed from diverse instruction-response pairs. Through extensive experiments based on HEAL, with a particular focus on the intrinsic mechanisms of preference learning, we demonstrate that current preference learning methods can effectively capture preferences provided by proxy models while simultaneously suppressing negative samples. These findings contribute to preference learning research through two significant avenues. Theoretically, we introduce hypothesis space analysis as an innovative paradigm for understanding preference alignment. Practically, HEAL offers researchers robust diagnostic tools for refining preference optimization methods, while our empirical results identify promising directions for developing more advanced alignment algorithms capable of comprehensive preference capture.
Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
Kaiyan Chang | Yonghao Shi | Chenglong Wang | Hang Zhou | Chi Hu | Xiaoqian Liu | Yingfeng Luo | Yuan Ge | Tong Xiao | JingBo Zhu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Kaiyan Chang | Yonghao Shi | Chenglong Wang | Hang Zhou | Chi Hu | Xiaoqian Liu | Yingfeng Luo | Yuan Ge | Tong Xiao | JingBo Zhu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Test-Time Scaling (TTS) is a promising approach to progressively elicit the model’s intelligence during inference. Recently, training-based TTS methods, such as continued reinforcement learning (RL), have further surged in popularity, while training-free TTS methods are gradually fading from prominence. However, the additional computation overhead of training amplifies the burden on test-time scaling.In this paper, we focus on training-free TTS methods for reasoning. We first design Conditional Step-level Self-refinement, a fine-grained sequential scaling method guided by process verification. On top of its effectiveness, we further combine it with other classical parallel scaling methods at the step level, to introduce a novel inference paradigm called Hybrid Test-Time Scaling. Extensive experiments on five instruction-tuned LLMs across different scales (3B-14B) and families demonstrate that hybrid strategy incorporating various training-free TTS methods at a fine granularity has considerable potential for expanding the reasoning performance boundaries of LLMs.
Search
Fix author
Co-authors
- Tong Xiao (肖桐) 6
- JingBo Zhu (朱靖波) 6
- Yifu Huo 3
- Kaiyan Chang 2
- Changxiaojia 2
- Yuan Ge 2
- Qiaozhi He 2
- Bei Li 2
- Tongran Liu 2
- Xiaoqian Liu 2
- Yingfeng Luo 2
- Shunjie Xing 2
- Zhengtao Yu (余正涛) 2
- Tian Hua Zhou 2
- Quan Du 1
- Peinan Feng 1
- AoKai Hao 1
- Chi Hu 1
- Canjia Li 1
- DingYang Lin 1
- Weixiong Lin 1
- Bingyu Liu 1
- Hongyu Liu 1
- Yonghao Shi 1
- Huiyu Wang 1
- Jianjin Wang 1
- Yun Yang 1
- Chunliang Zhang 1
- Junxiang Zhang 1
- Haishu Zhao 1
- Bingsen Zhou 1
- Hang Zhou 1
- Qiren Zhu 1
- Xingzhao Zhu 1
- Ziming Zhu 1