Kun Zhang
Author directoryOther people with similar names: Kun Zhang, Kun Zhang, Kun Zhang (Inria Saclay-Île-de-France), Kun Zhang (University of Chinese Academy of Sciences), Kun Zhang (University of Science and Technology of China)
Unverified author pages with similar names: Kun Zhang
2026
FactVerse: A Benchmark for Factual Consistency in Interleaved Image–Text Generation
Yubo Shan | Kun Zhang | Qiming Xu | Liping Cao | Yingying Cao | Jian Zhang | Yu Wang | Jingyuan Li | Yuanzhuo Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yubo Shan | Kun Zhang | Qiming Xu | Liping Cao | Yingying Cao | Jian Zhang | Yu Wang | Jingyuan Li | Yuanzhuo Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Interleaved multimodal understanding and generation—where models can interactively comprehend and produce images and text in arbitrary orders—has emerged as a key research direction in generative Multimodal Large Language Models(MLLMs). Such interleaved image–text content plays an increasingly important role in information dissemination. However, the compounded persuasive power of multimodal narratives also raises the risk of factual misinformation. Despite this, existing benchmarks lack effective mechanisms to evaluate factual consistency in interleaved image–text content. To bridge this gap, we introduce FactVerse, a benchmark dedicated to evaluating factual consistency in interleaved image-text generation. FactVerse comprises 3,000 human-verified instances across four categories and 50 domains, supporting both English and Chinese. We also establish a multi-dimensional evaluation framework designed to rigorously assess factual consistency. Experiments demonstrate that our framework achieves high alignment with human judgments, significantly outperforming existing evaluation methods. Furthermore, our analysis reveals systematic deficiencies in current models, offering critical insights for future design.
2025
TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding
Kun Zhang | Liqiang Niu | Zhen Cao | Fandong Meng | Jie Zhou
Findings of the Association for Computational Linguistics: EMNLP 2025
Kun Zhang | Liqiang Niu | Zhen Cao | Fandong Meng | Jie Zhou
Findings of the Association for Computational Linguistics: EMNLP 2025
Text-rich images are ubiquitous in real-world applications, serving as a critical medium for conveying complex information and facilitating accessibility.Despite recent advances driven by Multimodal Large Language Models (MLLMs), existing benchmarks suffer from limited scale, fragmented scenarios, and evaluation protocols that fail to fully capture holistic image understanding.To address these gaps, we present TIU-Bench, a large-scale, multilingual benchmark comprising over 100,000 full-image annotations and 22,000 rigorously validated question-answer (QA) pairs that span 18 subtasks across diverse real-world scenarios.TIU-Bench introduces a novel full-image structured output format that jointly models geometric, textual, and relational information, enabling fine-grained evaluation of perception and reasoning capabilities. Furthermore, we propose a two-stage understanding framework named T2TIU, which first generates a structured representation of the entire image and subsequently conducts reasoning on this representation to address complex visual-textual queries.Extensive experiments on 10 state-of-the-art generative models highlight the challenges and opportunities in advancing text-rich image understanding.Our benchmark and framework provide a comprehensive platform for developing and evaluating next-generation multimodal AI systems.