Bo Li
Other people with similar names: Bo Li (Tsinghua, Baidu), Bo Li (Chinese Academy of Sciences), Bo Li, Bo Li (Xi'an Jiaotong University), Bo Li (Hebei), Bo Li, Bo Li (BeiHang), Bo Li (Chinese Academy of Sciences), Bo Li (NUS, Google), Bo Li (Vanderbilt, UIUC)
Unverified author pages with similar names: Bo Li
2026
MGPO: Thinking with Images via Multi-Turn Grounding-Based Reinforcement Learning
Xinyu Huang | Yuhao Dong | Weiwei Tian | Bo Li | Rui Feng | Ziwei Liu
Findings of the Association for Computational Linguistics: ACL 2026
Xinyu Huang | Yuhao Dong | Weiwei Tian | Bo Li | Rui Feng | Ziwei Liu
Findings of the Association for Computational Linguistics: ACL 2026
State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous visual tokens, many of which are irrelevant to the downstream task. In this paper, we propose Multi-turn Grounding-based Policy Optimization (MGPO), an end-to-end reinforcement learning (RL) framework that enables LMMs to iteratively focus on key visual regions by automatically cropping sub-images, based on model-predicted grounding coordinates within a multi-turn conversation framework. Compared to supervised fine-tuning (SFT), which requires costly additional grounding annotations, our approach highlights that LMMs can emerge robust grounding abilities during the RL training process, leveraging only a binary reward function derived from the correctness of the final answer. Additionally, we observe that LMMs struggle to autonomously trigger visual grounding during the rollout process. To address this cold start problem, we design a multi-turn conversational template and restrict policy loss computation to model outputs generated across multiple dialogue rounds, thereby promoting stable optimization. Extensive experiments demonstrate that, when trained on standard visual-question-short answering data without grounding annotations, MGPO effectively elicits stronger grounding capabilities compared to GRPO, leading to 5.4% improvement on in-distribution MME-Realworld and 5.2% improvement on the challenging out-of-distribution (OOD) V* Bench. Notably, MGPO post-training on Qwen2.5-VL-7B with 21K samples surpasses OpenAI’s o1 and GPT-4o models on the OOD V* Bench.
Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional Videos
Kairui Hu | Penghao Wu | Fanyi Pu | Wang Xiao | Xiang Yue | Bo Li | Yuanhan Zhang | Ziwei Liu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Kairui Hu | Penghao Wu | Fanyi Pu | Wang Xiao | Xiang Yue | Bo Li | Yuanhan Zhang | Ziwei Liu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Humans acquire knowledge through three cognitive stages: perceiving information, comprehending knowledge, and adapting knowledge to solve novel problems. Videos serve as an effective medium for knowledge acquisition, facilitating a natural progression through these learning stages. However, existing video benchmarks fail to evaluate the knowledge acquisition capabilities of Large Multimodal Models (LMMs). To address this gap, we introduce Video-MMMU, a multi-modal, multi-discipline, multi-track benchmark that evaluates LMMs’ ability to acquire knowledge from college-level, educational videos. Video-MMMU features a collection of 300 videos and 900 human-annotated questions across six disciplines, evaluating knowledge acquisition through stage-aligned question-answer pairs: Perception, Comprehension, and Adaptation. Beyond measuring final accuracy, Video-MMMU proposes the performance gain metric that quantifies an LMM’s learning gain from video, shifting the focus of evaluation from absolute performance to learning efficiency. Our evaluation reveals a substantial gap between human learners and current LMMs, highlighting the need to improve models’ ability to learn and adapt knowledge from video content.
MMSearch-R1: Incentivizing LMMs to Search
Jinming Wu | Zihao Deng | Wei Li | Yiding Liu | Bo You | Bo Li | Zejun MA | Ziwei Liu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Jinming Wu | Zihao Deng | Wei Li | Yiding Liu | Bo You | Bo Li | Zejun MA | Ziwei Liu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Robust deployment of large multimodal models (LMMs) in real-world scenarios requires access to external knowledge sources, given the complexity and dynamic nature of real-world information. Existing approaches such as retrieval-augmented generation (RAG) and prompt engineered search agents rely on rigid pipelines, often leading to inefficient or excessive search behaviors. We present MMSearch-R1, the first end-to-end reinforcement learning framework that enables LMMs to perform on-demand, multi-turn search in real-world Internet environments. Our framework integrates both image and text search tools, allowing the model to reason about when and how to invoke them guided by an outcome-based reward with a search penalty. To support training, We collect a multimodal search VQA dataset through a semi-automated pipeline that covers diverse visual and textual knowledge needs and curate a search-balanced subset with both search-required and search-free samples, which proves essential for shaping efficient and on-demand search behavior. Extensive experiments on knowledge-intensive and info-seeking VQA tasks show that our model not only outperforms traditional RAG-based baselines of the same model size, but also matches the performance of a larger RAG-based model while reducing search calls by over 30%. We further analyze key empirical findings to offer actionable insights for advancing research in multimodal search.
2025
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Kaichen Zhang | Bo Li | Peiyuan Zhang | Fanyi Pu | Joshua Adrian Cahyono | Kairui Hu | Shuai Liu | Yuanhan Zhang | Jingkang Yang | Chunyuan Li | Ziwei Liu
Findings of the Association for Computational Linguistics: NAACL 2025
Kaichen Zhang | Bo Li | Peiyuan Zhang | Fanyi Pu | Joshua Adrian Cahyono | Kairui Hu | Shuai Liu | Yuanhan Zhang | Jingkang Yang | Chunyuan Li | Ziwei Liu
Findings of the Association for Computational Linguistics: NAACL 2025
The advances of large foundation models necessitate wide-coverage, low-cost, and zero-contamination benchmarks. Despite continuous exploration of language model evaluations, comprehensive studies on the evaluation of Large Multi-modal Models (LMMs) remain limited. In this work, we introduce LMMS-EVAL, a unified and standardized multimodal benchmark framework with over 50 tasks and more than 10 models to promote transparent and reproducible evaluations. Although LMMS-EVAL offers comprehensive coverage, we find it still falls short in achieving low cost and zero contamination. To approach this evaluation trilemma, we further introduce LMMS-EVAL LITE, a pruned evaluation toolkit that emphasizes both coverage and efficiency. Additionally, we present Multimodal LIVEBENCH that utilizes continuously updating news and online forums to assess models’ generalization abilities in the wild, featuring a low-cost and zero-contamination evaluation approach. In summary, our work highlights the importance of considering the evaluation trilemma and provides practical solutions to navigate the trade-offs in evaluating large multi-modal models, paving the way for more effective and reliable benchmarking of LMMs.
LIME: Less Is More for MLLM Evaluation
King Zhu | Qianbo Zang | Shian Jia | Siwei Wu | Feiteng Fang | Yizhi Li | Shuyue Guo | Tianyu Zheng | Jiawei Guo | Bo Li | Haoning Wu | Xingwei Qu | Jian Yang | Ruibo Liu | Xiang Yue | Jiaheng Liu | Chenghua Lin | Hamid Alinejad-Rokny | Min Yang | Shiwen Ni | Wenhao Huang | Ge Zhang
Findings of the Association for Computational Linguistics: ACL 2025
King Zhu | Qianbo Zang | Shian Jia | Siwei Wu | Feiteng Fang | Yizhi Li | Shuyue Guo | Tianyu Zheng | Jiawei Guo | Bo Li | Haoning Wu | Xingwei Qu | Jian Yang | Ruibo Liu | Xiang Yue | Jiaheng Liu | Chenghua Lin | Hamid Alinejad-Rokny | Min Yang | Shiwen Ni | Wenhao Huang | Ge Zhang
Findings of the Association for Computational Linguistics: ACL 2025
Multimodal Large Language Models (MLLMs) are measured on numerous benchmarks like image captioning, visual question answer, and reasoning. However, these benchmarks often include overly simple or uninformative samples, making it difficult to effectively distinguish the performance of different MLLMs. Additionally, evaluating models across many benchmarks creates a significant computational burden. To address these issues, we propose LIME (Less Is More for MLLM Evaluation), a refined and efficient benchmark curated using a semi-automated pipeline. This pipeline filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding. Our experiments show that LIME reduces the number of samples by 76% and evaluation time by 77%, while it can more effectively distinguish different models’ abilities. Notably, we find that traditional automatic metrics like CIDEr are insufficient for evaluating MLLMs’ captioning performance, and excluding the caption task score yields a more accurate reflection of overall model performance. All code and data are available at https://anonymous.4open.science/r/LIME-49CD
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
Yiming Jia | Jiachen Li | Xiang Yue | Bo Li | Ping Nie | Kai Zou | Wenhu Chen
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Yiming Jia | Jiachen Li | Xiang Yue | Bo Li | Ping Nie | Kai Zou | Wenhu Chen
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Vision-Language Models have made significant progress on many perception-focused tasks. However, their progress on reasoning-focused tasks remains limited due to the lack of high-quality and diverse training data. In this work, we aim to address the scarcity of reasoning-focused multimodal datasets. We propose VisualWebInstruct, a novel approach that leverages search engines to create a diverse and high-quality dataset spanning multiple disciplines, including mathematics, physics, finance, and chemistry, etc. Starting with a meticulously selected set of 30,000 seed images, we employ Google Image Search to identify websites containing similar images. We collect and process HTML data from over 700K unique URLs. Through a pipeline of content extraction, filtering, and synthesis, we construct a dataset of approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs and the remaining comprising text-based QA pairs. Models fine-tuned on VisualWebInstruct demonstrate significant performance improvements: (1) fine-tuning on Llava-OV results in 10-20 absolute points improvement across benchmarks, and (2) fine-tuning from MAmmoTH-VL yields a 5 absolute points gain across benchmarks. Our best model, MAmmoTH-VL2, achieves the best known performance with SFT without RL within the 10B parameter class on MMMU-Pro (40.7), MathVerse (42.6), and DynaMath (55.7). These results highlight the effectiveness of our dataset in enhancing the reasoning capabilities of vision-language models for complex multimodal tasks.
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Jiawei Guo | Tianyu Zheng | Yizhi Li | Yuelin Bai | Bo Li | Yubo Wang | King Zhu | Graham Neubig | Wenhu Chen | Xiang Yue
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Jiawei Guo | Tianyu Zheng | Yizhi Li | Yuelin Bai | Bo Li | Yubo Wang | King Zhu | Graham Neubig | Wenhu Chen | Xiang Yue
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Open-source multimodal large language models (MLLMs) have shown significant potential in a broad range of tasks. However, their reasoning capabilities remain constrained by existing instruction-tuning datasets, which were predominately repurposed from academic datasets such as VQA, AI2D, and ChartQA. These datasets target simplistic tasks, and only provide phrase-level answers without any intermediate rationales.To address these challenges, we introduce a scalable and cost-effective method to construct a large-scale multimodal instruction-tuning dataset with rich intermediate rationales designed to elicit CoT reasoning. Using only open models, we create a dataset containing 12M instruction-response pairs to cover diverse reasoning-intensive tasks.Experiments demonstrate that training MLLMs on our dataset not only significantly improves reasoning capabilities, achieving state-of-the-art performance on benchmarks such as MathVerse (+8.1%), MMMU-Pro (+7%), and MuirBench (+13.3%), but also gains improvements of up to 4% on non-reasoning-based benchmarks.
Search
Fix author
Co-authors
- Ziwei Liu 4
- Xiang Yue 4
- Wenhu Chen 2
- Jiawei Guo 2
- Kairui Hu 2
- Yizhi Li 2
- Fanyi Pu 2
- Yuanhan Zhang 2
- Tianyu Zheng 2
- King Zhu 2
- Hamid Alinejad-Rokny 1
- Yuelin Bai 1
- Joshua Adrian Cahyono 1
- Zihao Deng 1
- Yuhao Dong 1
- Feiteng Fang 1
- Rui Feng 1
- Shuyue Guo 1
- Wenhao Huang 1
- Xinyu Huang 1
- Shian Jia 1
- Yiming Jia 1
- Chunyuan Li 1
- Jiachen Li 1
- Wei Li 1
- Chenghua Lin 1
- Jiaheng Liu 1
- Ruibo Liu 1
- Shuai Liu 1
- Yiding Liu 1
- Zejun Ma 1
- Graham Neubig 1
- Shiwen Ni 1
- Ping Nie 1
- Xingwei Qu 1
- Weiwei Tian 1
- Yubo Wang 1
- Haoning Wu 1
- Jinming Wu 1
- Penghao Wu 1
- Siwei Wu 1
- Wang Xiao 1
- Jian Yang 1
- Jingkang Yang 1
- Min Yang 1
- Bo You 1
- Qianbo Zang 1
- Ge Zhang 1
- Kaichen Zhang 1
- Peiyuan Zhang 1
- Kai Zou 1