IREASONER: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models

Meghana Sunil, Manikandarajan Venmathimaran, Muthu Subash Kavitha


Abstract
Recent work shows that large multimodal models (LMMs) can self-improve from unlabeled data via self-play and intrinsic feedback. Yet existing self-evolving frameworks mainly reward final outcomes, leaving intermediate reasoning weakly constrained despite its importance for visually grounded decision making. We propose IREASONER, a self-evolving framework that improves an LMM’s implicit reasoning by explicitly eliciting chain-of-thought (CoT) and rewarding its internal agreement. In a Proposer–Solver loop over unlabeled images, IREASONER augments outcome-level intrinsic rewards with a trajectory-aware signal defined over intermediate reasoning steps, providing learning signals that distinguish reasoning paths leading to the same answer without ground-truth labels or external judges. Starting from Qwen2.5-VL-7B, IREASONER yields up to +2.1 points across diverse multimodal reasoning benchmarks under fully unsupervised post-training. We hope this work serves as a starting point for reasoning-aware self-improvement in LMMs in purely unsupervised settings.
Anthology ID:
2026.findings-acl.1468
Volume:
Findings of the Association for Computational Linguistics: ACL 2026
Month:
July
Year:
2026
Address:
San Diego, California, United States
Editors:
Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
29361–29372
Language:
URL:
https://aclanthology.org/2026.findings-acl.1468/
DOI:
10.18653/v1/2026.findings-acl.1468
Bibkey:
Cite (ACL):
Meghana Sunil, Manikandarajan Venmathimaran, and Muthu Subash Kavitha. 2026. IREASONER: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models. In Findings of the Association for Computational Linguistics: ACL 2026, pages 29361–29372, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):
IREASONER: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models (Sunil et al., Findings 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.findings-acl.1468.pdf
Checklist:
 2026.findings-acl.1468.checklist.pdf