SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

Yifu Huo; Chenglong Wang; Ziming Zhu; Shunjie Xing; Peinan Feng; Tongran Liu; Qiaozhi He; Tian Hua Zhou; Changxiaojia; JingBo Zhu (朱靖波); Zhengtao Yu (余正涛); Tong Xiao (肖桐)

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

Yifu Huo, Chenglong Wang, Ziming Zhu, Shunjie Xing, Peinan Feng, Tongran Liu, Qiaozhi He, Tian Hua Zhou, Changxiaojia, JingBo Zhu, Zhengtao Yu, Tong Xiao

Abstract

Reinforcement learning (RL) has emerged as a promising paradigm for training reasoning-oriented models by leveraging rule-based reward signals. However, RL training typically tends to improve single-sample success rates (i.e., Pass@1) while offering limited exploration of diverse reasoning trajectories, which is crucial for multi-sample performance (i.e., Pass@k). Our preliminary analysis reveals that this limitation stems from a fundamental squeezing effect, whereby probability mass is excessively concentrated on a narrow subset of high-reward trajectories, restricting genuine exploration and constraining attainable performance under RL training. To address this issue, in this work, we propose Steering Probability Squeezing (SPS), a training paradigm that interleaves conventional RL with inverse reinforcement learning (IRL). SPS treats on-policy rollouts as demonstrations and employs IRL to explicitly reshape the induced trajectory distribution, thereby enhancing exploration without introducing external supervision. Experiments on five commonly used reasoning benchmarks demonstrate that SPS can enable better exploration and improve Pass@k. Beyond algorithmic contributions, we provide an analysis of RL learning dynamics and identify an empirical upper bound on Pass@k, shedding light on intrinsic exploration limits in RL-based reasoning models. Our findings suggest that alternating between RL and IRL offers an effective pathway toward extending the exploration capacity of reasoning-oriented large language models.

Anthology ID:: 2026.findings-acl.865
Volume:: Findings of the Association for Computational Linguistics: ACL 2026
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 17472–17489
Language:
URL:: https://aclanthology.org/2026.findings-acl.865/
DOI:
Bibkey:
Cite (ACL):: Yifu Huo, Chenglong Wang, Ziming Zhu, Shunjie Xing, Peinan Feng, Tongran Liu, Qiaozhi He, Tian Hua Zhou, Changxiaojia, JingBo Zhu, Zhengtao Yu, and Tong Xiao. 2026. SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models. In Findings of the Association for Computational Linguistics: ACL 2026, pages 17472–17489, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models (Huo et al., Findings 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.findings-acl.865.pdf
Checklist:: 2026.findings-acl.865.checklist.pdf

PDF Cite Search Checklist Fix data