Jiaqi Wang
Other people with similar names: Jiaqi Wang, Jiaqi Wang, Jiaqi Wang
Unverified author pages with similar names: Jiaqi Wang
2026
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
Dianyi Wang | Wei Song | Yikun Wang | Siyuan Wang | Kaicheng Yu | Zhongyu Wei | Jiaqi Wang
Findings of the Association for Computational Linguistics: ACL 2026
Dianyi Wang | Wei Song | Yikun Wang | Siyuan Wang | Kaicheng Yu | Zhongyu Wei | Jiaqi Wang
Findings of the Association for Computational Linguistics: ACL 2026
Typical large vision-language models (LVLMs) apply autoregressive supervision primarily to textual responses, without fully exploiting causal learning over rich visual inputs. As a result, these models often emphasize vision-to-language alignment while potentially overlooking fine-grained visual information. While prior work has explored autoregressive image generation, effectively leveraging autoregressive visual supervision to enhance image understanding remains an open challenge. In this paper, we introduce Autoregressive Semantic Visual Reconstruction (ASVR), which enables joint learning of visual and textual modalities within a unified autoregressive framework. ASVR trains models to autoregressively reconstruct the semantic content of input images, which consistently enhances multimodal comprehension. Notably, we show that even when provided with continuous image features as input, models can effectively reconstruct discrete semantic tokens, resulting in stable and consistent improvements across various multimodal understanding benchmarks. ASVR delivers significant performance gains and scalability across varying data scales, visual input, visual supervision and model architectures. In particular, ASVR generally improves baselines by 2-3% across 14 multimodal benchmarks.
GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization
Yikun Wang | Yibin Wang | Dianyi Wang | Zimian Peng | Qipeng Guo | Dacheng Tao | Jiaqi Wang
Findings of the Association for Computational Linguistics: ACL 2026
Yikun Wang | Yibin Wang | Dianyi Wang | Zimian Peng | Qipeng Guo | Dacheng Tao | Jiaqi Wang
Findings of the Association for Computational Linguistics: ACL 2026
Recent progress in large language models (LLMs) has boosted mathematical reasoning, yet geometry remains challenging where auxiliary construction is often essential. Prior methods either underperform or depend on very large models (e.g., GPT-4o), making them costly. We argue that reinforcement learning with verifiable rewards (e.g., GRPO) can train smaller models to couple auxiliary construction with solid geometric reasoning. However, naively applying GRPO yields unconditional rewards, encouraging indiscriminate and sometimes harmful constructions. We propose Group Contrastive Policy Optimization (GCPO), an RL framework with two components: (1) Group Contrastive Masking, which assigns positive/negative construction rewards based on contextual utility, and (2) a Length Reward that encourages longer reasoning chains. On top of GCPO, we build GeometryZero, an affordable family of geometry reasoning models that selectively use auxiliary construction. Experiments on Geometry3K and MathVista show GeometryZero consistently outperforms RL baselines (e.g., GRPO, ToRL).
VideoPro: Adaptive Program Reasoning for Long Video Understanding
Chenglin Li | Feng Han | Yikun Wang | Ruilin Li | Shuai Dong | Haowen Hou | Haitao Li | Qianglong Chen | Feng Tao | Jingqi Tong | Yin Zhang | Jiaqi Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Chenglin Li | Feng Han | Yikun Wang | Ruilin Li | Shuai Dong | Haowen Hou | Haitao Li | Qianglong Chen | Feng Tao | Jingqi Tong | Yin Zhang | Jiaqi Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Understanding long videos remains challenging due to the sparsity of visual evidence relevant to a given query. Prior work has explored program-based visual grounding, typically relying on executable programs generated by auxiliary large language models. However, when scaling to long videos, existing approaches face several critical limitations: (1) frame-centric vision modules are often insufficient for long video processing; (2) naively applying program-based reasoning to all queries incurs considerable computational overhead; and (3) errors arising from low-confidence predictions and imperfect program execution are difficult to recover from. To address these challenges, we propose VideoPro, a unified framework that enables VideoLLMs to adaptively reason over long videos and refine their predictions through executable programs. VideoPro first performs adaptive reasoning, dynamically determining whether a query can be resolved directly by the native VideoLLM or requires explicit multi-step program reasoning. For complex queries, the model decomposes the task into executable programs that invoke specialized vision modules for precise temporal and semantic grounding. To further improve robustness, VideoPro incorporates a self-refinement mechanism that leverages execution feedback and confidence signals to correct erroneous executions and refine low-confidence reasoning programs. By tightly integrating adaptive reasoning with self-refinement, VideoPro consistently outperforms prior methods across multiple long-video understanding benchmarks, yielding an average 6.7% improvement for Qwen3-VL-8B.
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Junbo Niu | Zheng Liu | Zhuangcheng Gu | Bin Wang | Linke Ouyang | Zhiyuan Zhao | Tao Chu | Tianyao He | Fan Wu | Qintong Zhang | Zhenjiang Jin | Guang Liang | Rui Zhang | Wenzheng Zhang | Yuan Qu | Zhifei Ren | Yuefeng Sun | Zirui Tang | Boyu Niu | Yuanhong Zheng | Dongsheng Ma | Ziyang Miao | Hejun Dong | Siyi Qian | Junyuan Zhang | Fangdong Wang | Jingzhou Chen | Xiaomeng Zhao | Liqun Wei | Wei Li | Shasha Wang | RuiLiang Xu | Yuanyuan Cao | Lu Chen | Qianqian Wu | Huaiyu Gu | Lindong Lu | Dechen Lin | Shenguanlin | Xuanhe Zhou | Linfeng Zhang | Yuhang Zang | Xiaoyi Dong | Jiaqi Wang | Bo Zhang | Lei Bai | Pei Chu | Weijia Li | Jiang Wu | Lijun Wu | Zhenxiang Li | Guangyu Wang | Zhongying Tu | Chao Xu | Kai Chen | Bowen Zhou | Dahua Lin | Wentao Zhang | Conghui He
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
Junbo Niu | Zheng Liu | Zhuangcheng Gu | Bin Wang | Linke Ouyang | Zhiyuan Zhao | Tao Chu | Tianyao He | Fan Wu | Qintong Zhang | Zhenjiang Jin | Guang Liang | Rui Zhang | Wenzheng Zhang | Yuan Qu | Zhifei Ren | Yuefeng Sun | Zirui Tang | Boyu Niu | Yuanhong Zheng | Dongsheng Ma | Ziyang Miao | Hejun Dong | Siyi Qian | Junyuan Zhang | Fangdong Wang | Jingzhou Chen | Xiaomeng Zhao | Liqun Wei | Wei Li | Shasha Wang | RuiLiang Xu | Yuanyuan Cao | Lu Chen | Qianqian Wu | Huaiyu Gu | Lindong Lu | Dechen Lin | Shenguanlin | Xuanhe Zhou | Linfeng Zhang | Yuhang Zang | Xiaoyi Dong | Jiaqi Wang | Bo Zhang | Lei Bai | Pei Chu | Weijia Li | Jiang Wu | Lijun Wu | Zhenxiang Li | Guangyu Wang | Zhongying Tu | Chao Xu | Kai Chen | Bowen Zhou | Dahua Lin | Wentao Zhang | Conghui He
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
We introduce MinerU2.5, a 1.2B-parameter document parsing vision-language model that achieves state-of-the-art recognition accuracy while maintaining exceptional computational efficiency. Our approach employs a coarse-to-fine, two-stage parsing strategy that decouples global layout analysis from local content recognition. In the first stage, the model performs efficient layout analysis on downsampled images to identify structural elements, circumventing the computational overhead of processing high-resolution inputs. In the second stage, guided by the global layout, it performs targeted content recognition on native-resolution crops extracted from the original image, preserving fine-grained details in dense text, complex formulas, and tables. To support this strategy, we developed a comprehensive data engine that generates diverse, large-scale training corpora for both pretraining and fine-tuning. Ultimately, MinerU2.5 demonstrates strong document parsing ability, achieving state-of-the-art performance on multiple benchmarks, surpassing both general-purpose and domain-specific models across various recognition tasks, while maintaining significantly lower computational overhead.
2025
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Yuhang Zang | Xiaoyi Dong | Pan Zhang | Yuhang Cao | Ziyu Liu | Shengyuan Ding | Shenxi Wu | Yubo Ma | Haodong Duan | Wenwei Zhang | Kai Chen | Dahua Lin | Jiaqi Wang
Findings of the Association for Computational Linguistics: ACL 2025
Yuhang Zang | Xiaoyi Dong | Pan Zhang | Yuhang Cao | Ziyu Liu | Shengyuan Ding | Shenxi Wu | Yubo Ma | Haodong Duan | Wenwei Zhang | Kai Chen | Dahua Lin | Jiaqi Wang
Findings of the Association for Computational Linguistics: ACL 2025
Despite the promising performance of Large Vision Language Models (LVLMs) in visual understanding, they occasionally generate incorrect outputs. While reward models (RMs) with reinforcement learning or test-time scaling offer the potential for improving generation quality, a critical gap remains: publicly available multi-modal RMs for LVLMs are scarce, and the implementation details of proprietary models are often unclear. We bridge this gap with InternLM-XComposer2.5-Reward (IXC-2.5-Reward), a simple yet effective multi-modal reward model that aligns LVLMs with human preferences. To ensure the robustness and versatility of IXC-2.5-Reward, we set up a high-quality multi-modal preference corpus spanning text, image, and video inputs across diverse domains, such as instruction following, general understanding, text-rich documents, mathematical reasoning, and video understanding. IXC-2.5-Reward achieves excellent results on the latest multi-modal reward model benchmark and shows competitive performance on text-only reward model benchmarks. We further demonstrate three key applications of IXC-2.5-Reward: (1) Providing a supervisory signal for RL training. We integrate IXC-2.5-Reward with Proximal Policy Optimization (PPO) yields IXC-2.5-Chat, which shows consistent improvements in instruction following and multi-modal open-ended dialogue; (2) Selecting the best response from candidate responses for test-time scaling; and (3) Filtering outlier or noisy samples from existing image and video instruction tuning training data.
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
Yubo Ma | Jinsong Li | Yuhang Zang | Xiaobao Wu | Xiaoyi Dong | Pan Zhang | Yuhang Cao | Haodong Duan | Jiaqi Wang | Yixin Cao | Aixin Sun
Findings of the Association for Computational Linguistics: ACL 2025
Yubo Ma | Jinsong Li | Yuhang Zang | Xiaobao Wu | Xiaoyi Dong | Pan Zhang | Yuhang Cao | Haodong Duan | Jiaqi Wang | Yixin Cao | Aixin Sun
Findings of the Association for Computational Linguistics: ACL 2025
Despite the strong performance of ColPali/ColQwen2 in Visualized Document Retrieval (VDR), its patch-level embedding approach leads to excessive memory usage. This empirical study investigates methods to reduce patch embeddings per page while minimizing performance degradation. We evaluate two token-reduction strategies: token pruning and token merging. Regarding token pruning, we surprisingly observe that a simple random strategy outperforms other sophisticated pruning methods, though still far from satisfactory. Further analysis reveals that pruning is inherently unsuitable for VDR as it requires removing certain page embeddings without query-specific information. Turning to token merging (more suitable for VDR), we search for the optimal combinations of merging strategy across three dimensions and develops Light-ColPali/ColQwen2. It maintains 98.2% of retrieval performance with only 11.8% of original memory usage, and preserves 94.6% effectiveness at 2% memory footprint. We expect our empirical findings and resulting Light-ColPali/ColQwen2 offer valuable insights and establish a competitive baseline for future efficient-VDR research.
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
Xiangyu Zhao | Shengyuan Ding | Zicheng Zhang | Haian Huang | Maosong Cao | Weiyun Wang | Jiaqi Wang | Xinyu Fang | Wenhai Wang | Guangtao Zhai | Haodong Duan | Hua Yang | Kai Chen
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Xiangyu Zhao | Shengyuan Ding | Zicheng Zhang | Haian Huang | Maosong Cao | Weiyun Wang | Jiaqi Wang | Xinyu Fang | Wenhai Wang | Guangtao Zhai | Haodong Duan | Hua Yang | Kai Chen
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Recent advancements in open-source multi-modal large language models (MLLMs) have primarily focused on enhancing foundational capabilities, leaving a significant gap in human preference alignment. This paper introduces OmniAlign-V, a comprehensive dataset of 200K high-quality training samples featuring diverse images, complex questions, and varied response formats to improve MLLMs’ alignment with human preferences. We also present MM-AlignBench, a human-annotated benchmark specifically designed to evaluate MLLMs’ alignment with human values. Experimental results show that finetuning MLLMs with OmniAlign-V, using Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO), significantly enhances human preference alignment while maintaining or enhancing performance on standard VQA benchmarks, preserving their fundamental capabilities.
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
Shuangrui Ding | Zihan Liu | Xiaoyi Dong | Pan Zhang | Rui Qian | Junhao Huang | Conghui He | Dahua Lin | Jiaqi Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Shuangrui Ding | Zihan Liu | Xiaoyi Dong | Pan Zhang | Rui Qian | Junhao Huang | Conghui He | Dahua Lin | Jiaqi Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Creating lyrics and melodies for the vocal track in a symbolic format, known as song composition, demands expert musical knowledge of melody, an advanced understanding of lyrics, and precise alignment between them. Despite achievements in sub-tasks such as lyric generation, lyric-to-melody, and melody-to-lyric, etc, a unified model for song composition has not yet been achieved. In this paper, we introduce SongComposer, a pioneering step towards a unified song composition model that can readily create symbolic lyrics and melodies following instructions. SongComposer is a music-specialized large language model (LLM) that, for the first time, integrates the capability of simultaneously composing lyrics and melodies into LLMs by leveraging three key innovations: 1) a flexible tuple format for word-level alignment of lyrics and melodies, 2) an extended tokenizer vocabulary for song notes, with scalar initialization based on musical knowledge to capture rhythm, and 3) a multi-stage pipeline that captures musical structure, starting with motif-level melody patterns and progressing to phrase-level structure for improved coherence. Extensive experiments demonstrate that SongComposer outperforms advanced LLMs, including GPT-4, in tasks such as lyric-to-melody generation, melody-to-lyric generation, song continuation, and text-to-song creation. Moreover, we will release SongCompose, a large-scale dataset for training, containing paired lyrics and melodies in Chinese and English.
Search
Fix author
Co-authors
- Xiaoyi Dong 4
- Kai Chen 3
- Haodong Duan 3
- Dahua Lin 3
- Yikun Wang 3
- Yuhang Zang 3
- Pan Zhang 3
- Yuhang Cao 2
- Shengyuan Ding 2
- Conghui He 2
- Yubo Ma 2
- Dianyi Wang 2
- Lei Bai 1
- Maosong Cao 1
- Yixin Cao 1
- Yuanyuan Cao 1
- Jingzhou Chen 1
- Lu Chen 1
- Qianglong Chen 1
- Pei Chu 1
- Tao Chu 1
- Shuangrui Ding 1
- Hejun Dong 1
- Shuai Dong 1
- Xinyu Fang 1
- Huaiyu Gu 1
- Zhuangcheng Gu 1
- Qipeng Guo 1
- Feng Han 1
- Tianyao He 1
- Haowen Hou 1
- Haian Huang 1
- Junhao Huang 1
- Zhenjiang Jin 1
- Chenglin Li 1
- Haitao Li 1
- Jinsong Li 1
- Ruilin Li 1
- Wei Li 1
- Weijia Li 1
- Zhenxiang Li 1
- Guang Liang 1
- Dechen Lin 1
- Zheng Liu 1
- Zihan Liu 1
- Ziyu Liu 1
- Lindong Lu 1
- Dongsheng Ma 1
- Ziyang Miao 1
- Boyu Niu 1
- Junbo Niu 1
- Linke Ouyang 1
- Zimian Peng 1
- Rui Qian 1
- Siyi Qian 1
- Yuan Qu 1
- Zhifei Ren 1
- Shenguanlin 1
- Wei Song 1
- Aixin Sun 1
- Yuefeng Sun 1
- Zirui Tang 1
- Dacheng Tao 1
- Feng Tao 1
- Jingqi Tong 1
- Zhongying Tu 1
- Bin Wang 1
- Fangdong Wang 1
- Guangyu Wang 1
- Shasha Wang 1
- Siyuan Wang (王思远) 1
- Weiyun Wang 1
- Wenhai Wang 1
- Yibin Wang 1
- Liqun Wei 1
- Zhongyu Wei (魏忠钰) 1
- Fan Wu 1
- Jiang Wu 1
- Lijun Wu 1
- Qianqian Wu 1
- Shenxi Wu 1
- Xiaobao Wu 1
- Chao Xu 1
- RuiLiang Xu 1
- Hua Yang 1
- Kaicheng Yu 1
- Guangtao Zhai 1
- Bo Zhang 1
- Junyuan Zhang 1
- Linfeng Zhang 1
- Qintong Zhang 1
- Rui Zhang 1
- Wentao Zhang 1
- Wenwei Zhang 1
- Wenzheng Zhang 1
- Yin Zhang 1
- Zicheng Zhang 1
- Xiangyu Zhao 1
- Xiaomeng Zhao 1
- Zhiyuan Zhao 1
- Yuanhong Zheng 1
- Bowen Zhou 1
- Xuanhe Zhou 1