Ming Cheng
Author directoryOther people with similar names: Ming Cheng
Unverified author pages with similar names: Ming Cheng
2026
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
Wenhao You | Xingjian Diao | Wenjun Huang | Chunhui Zhang | Keyi Kong | Weiyi Wu | Chiyu Ma | Zhongyu Ouyang | Tingxuan Wu | Ming Cheng | Soroush Vosoughi | Jiang Gui
Findings of the Association for Computational Linguistics: ACL 2026
Wenhao You | Xingjian Diao | Wenjun Huang | Chunhui Zhang | Keyi Kong | Weiyi Wu | Chiyu Ma | Zhongyu Ouyang | Tingxuan Wu | Ming Cheng | Soroush Vosoughi | Jiang Gui
Findings of the Association for Computational Linguistics: ACL 2026
While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Audio-Visual Question Answering (Music AVQA) particularly underscores this, presenting unique challenges with its continuous, densely layered audio-visual content, intricate temporal dynamics, and the critical need for domain-specific knowledge. Through a systematic analysis of Music AVQA datasets and methods, this paper identifies that specialized input processing, architectures incorporating dedicated spatial-temporal designs, and music-specific modeling strategies are critical for success in this domain. Our study provides valuable insights for researchers by highlighting effective design patterns empirically linked to strong performance, proposing concrete future directions for incorporating musical priors, and aiming to establish a robust foundation for advancing multimodal musical understanding. We aim to encourage further research in this area and provide a GitHub repository of relevant works: https://github.com/WenhaoYou1/Survey4MusicAVQA.
2025
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
Xingjian Diao | Chunhui Zhang | Weiyi Wu | Zhongyu Ouyang | Peijun Qing | Ming Cheng | Soroush Vosoughi | Jiang Gui
Findings of the Association for Computational Linguistics: NAACL 2025
Xingjian Diao | Chunhui Zhang | Weiyi Wu | Zhongyu Ouyang | Peijun Qing | Ming Cheng | Soroush Vosoughi | Jiang Gui
Findings of the Association for Computational Linguistics: NAACL 2025
Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models face inherent limitations due to their finite internal capacity, which restricts their ability to process extended temporal sequences—an essential requirement for comprehensive video and audio analysis. To overcome these challenges, we introduce a specialized cognitive module, temporal working memory (TWM), which aims to enhance the temporal modeling capabilities of MFMs. It selectively retains task-relevant information across temporal dimensions, ensuring that critical details are preserved throughout the processing of video and audio content. The TWM uses a query-guided attention approach to focus on the most informative multimodal segments within temporal sequences. By retaining only the most relevant content, TWM optimizes the use of the model’s limited capacity, enhancing its temporal modeling ability. This plug-and-play module can be easily integrated into existing MFMs. With our TWM, nine state-of-the-art models exhibit significant performance improvements across tasks such as video captioning, question answering, and video-text retrieval. By enhancing temporal modeling, TWM extends the capability of MFMs to handle complex, time-sensitive data effectively. Our code is available at https://github.com/xid32/NAACL_2025_TWM.
A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labelling
Shiyu Ji | Farnoosh Hashemi | Joice Chen | Juanwen Pan | Weicheng Ma | Hefan Zhang | Sophia Pan | Ming Cheng | Shubham Mohole | Saeed Hassanpour | Soroush Vosoughi | Michael Macy
Findings of the Association for Computational Linguistics: EMNLP 2025
Shiyu Ji | Farnoosh Hashemi | Joice Chen | Juanwen Pan | Weicheng Ma | Hefan Zhang | Sophia Pan | Ming Cheng | Shubham Mohole | Saeed Hassanpour | Soroush Vosoughi | Michael Macy
Findings of the Association for Computational Linguistics: EMNLP 2025
Rhetorical strategies are central to persuasive communication, from political discourse and marketing to legal argumentation. However, analysis of rhetorical strategies has been limited by reliance on human annotation, which is costly, inconsistent, difficult to scale. Their associated datasets are often limited to specific topics and strategies, posing challenges for robust model development. We propose a novel framework that leverages large language models (LLMs) to automatically generate and label synthetic debate data based on a four-part rhetorical typology (causal, empirical, emotional, moral). We fine-tune transformer-based classifiers on this LLM-labeled dataset and validate its performance against human-labeled data on this dataset and on multiple external corpora. Our model achieves high performance and strong generalization across topical domains. We illustrate two applications with the fine-tuned model: (1) the improvement in persuasiveness prediction from incorporating rhetorical strategy labels, and (2) analyzing temporal and partisan shifts in rhetorical strategies in U.S. Presidential debates (1960–2020), revealing increased use of affective over cognitive argument in U.S. Presidential debates.
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
Xingjian Diao | Weiyi Wu | Keyi Kong | Peijun Qing | Xinwen Xu | Ming Cheng | Soroush Vosoughi | Jiang Gui
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Xingjian Diao | Weiyi Wu | Keyi Kong | Peijun Qing | Xinwen Xu | Ming Cheng | Soroush Vosoughi | Jiang Gui
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Visual Question Answering (VQA) is increasingly used in diverse applications ranging from general visual reasoning to safety-critical domains such as medical imaging and autonomous systems, where models must provide not only accurate answers but also explanations that humans can easily understand and verify. Prototype-based modeling has shown promise for interpretability by grounding predictions in semantically meaningful regions for purely visual reasoning tasks, yet remains underexplored in the context of VQA. We present ProtoVQA, a unified prototypical framework that (i) learns question-aware prototypes that serve as reasoning anchors, connecting answers to discriminative image regions, (ii) applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant, and (iii) supports both answering and grounding tasks through a shared prototype backbone. To assess explanation quality, we propose the Visual–Linguistic Alignment Score (VLAS), which measures how well the model’s attended regions align with ground-truth evidence. Experiments on Visual7W show that ProtoVQA yields faithful, fine-grained explanations while maintaining competitive accuracy, advancing the development of transparent and trustworthy VQA systems.
Learning Sparsity for Effective and Efficient Music Performance Question Answering
Xingjian Diao | Tianzhen Yang | Chunhui Zhang | Weiyi Wu | Ming Cheng | Jiang Gui
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
Xingjian Diao | Tianzhen Yang | Chunhui Zhang | Weiyi Wu | Ming Cheng | Jiang Gui
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
Music performances, characterized by dense and continuous audio as well as seamless audio-visual integration, present unique challenges for multimodal scene understanding and reasoning. Recent Music Performance Audio-Visual Question Answering (Music AVQA) datasets have been proposed to reflect these challenges, highlighting the continued need for more effective integration of audio-visual representations in complex question answering. However, existing Music AVQA methods often rely on dense and unoptimized representations, leading to inefficiencies in the isolation of key information, the reduction of redundancy, and the prioritization of critical samples. To address these challenges, we introduce Sparsify, a sparse learning framework specifically designed for Music AVQA. It integrates three sparsification strategies into an end-to-end pipeline and achieves state-of-the-art performance on the Music AVQA datasets. In addition, it reduces training time by 28.32% compared to its fully trained dense counterpart while maintaining accuracy, demonstrating clear efficiency gains. To further improve data efficiency, we propose a key-subset selection algorithm that selects and uses approximately 25% of MUSIC-AVQA v2.0 training data and retains 70–80% of full-data performance across models.