Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing

Xin Guo; Zhiheng Xi; Yiwen Ding; Yitao Zhai; Xiaowei Shi; Xunliang Cai; Tao Gui; Qi Zhang; Xuan-Jing Huang (黄萱菁)

Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing

Xin Guo, Zhiheng Xi, Yiwen Ding, Yitao Zhai, Xiaowei Shi, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang

Abstract

Self-improvement has emerged as a mainstream paradigm for advancing the reasoning capabilities of large vision–language models (LVLMs), where models explore and learn from successful trajectories iteratively. However, we identify a critical imbalance during this process: the model readily generates high-quality trajectories for simple queries (i.e., head data) but struggles with complex ones (i.e., tail data). This bias drives the optimization to disproportionately prioritize simple reasoning skills, while inhibiting the acquisition of complex capabilities. As iterations progress, this imbalance becomes more acute—a dynamic we term the "Matthew effect", ultimately stalling performance gains. To mitigate this, we approach head-tail re-balance during the exploration-and-learning process from two perspectives: distribution-reshaping and trajectory-resampling. Extensive experiments on Qwen2-VL-7B-Instruct and InternVL2.5-4B models across visual reasoning tasks demonstrate that our methods consistently improve visual reasoning capabilities, outperforming vanilla self-improvement baselines by an average of 3.86 points.

Anthology ID:: 2026.acl-long.1010
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 22104–22121
Language:
URL:: https://aclanthology.org/2026.acl-long.1010/
DOI:
Bibkey:
Cite (ACL):: Xin Guo, Zhiheng Xi, Yiwen Ding, Yitao Zhai, Xiaowei Shi, Xunliang Cai, Tao Gui, Qi Zhang, and Xuanjing Huang. 2026. Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 22104–22121, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing (Guo et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.1010.pdf
Checklist:: 2026.acl-long.1010.checklist.pdf

PDF Cite Search Checklist Fix data