CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models

Peyman Hosseini, Ondrej Bohdal, Taha Ceritli, Ignacio Castro, Matthew Purver, Mete Ozay, Umberto Michieli


Abstract
Test-Time Reinforcement Learning (TTRL) has shown promise in adapting foundation models for complex tasks at test time, resulting in large performance improvements. TTRL leverages an elegant two-phase sampling strategy: first, multi-sampling derives a pseudo-label via majority voting, while subsequent downsampling and reward-based fine-tuning encourages the model to explore and learn diverse valid solutions, with the pseudo-label modulating the reward signal. Meanwhile, In-Context Learning has been widely explored at inference time to enhance model performance without weight updates. However, TTRL’s two-phase sampling strategy under-utilizes contextual guidance, which can potentially improve pseudo-label accuracy in the initial exploitation phase while regulating exploration in the second. To address this, we propose Context-Guided TTRL (CG-TTRL), integrating context dynamically into both sampling phases and propose a method for efficient context selection for on-device applications. Our evaluations on mathematical and scientific QA benchmarks show CG-TTRL outperforms TTRL (e.g. additional 7% relative accuracy improvement over TTRL), while boosting efficiency by obtaining strong performance after only a few steps of Test-Time Training (e.g. 8% relative improvement rather than 1% over TTRL after 3 steps).
Anthology ID:
2026.tacl-1.85
Volume:
Transactions of the Association for Computational Linguistics, Volume 14
Month:
Year:
2026
Address:
Cambridge, MA
Venue:
TACL
SIG:
Publisher:
MIT Press
Note:
Pages:
1883–1898
Language:
URL:
https://aclanthology.org/2026.tacl-1.85/
DOI:
10.1162/tacl.a.783
Bibkey:
Cite (ACL):
Peyman Hosseini, Ondrej Bohdal, Taha Ceritli, Ignacio Castro, Matthew Purver, Mete Ozay, and Umberto Michieli. 2026. CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models. Transactions of the Association for Computational Linguistics, 14:1883–1898.
Cite (Informal):
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models (Hosseini et al., TACL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.tacl-1.85.pdf