Yong Wang
Author directoryOther people with similar names: Yong Wang, Yong Wang
Unverified author pages with similar names: Yong Wang
2026
CaM-HG: Causal-Enhanced MoE and Hypergraphs Network for Incomplete Multimodal Emotion Recognition in Conversations
Mingjian Yang | Yong Wang | Peng Liu | Wen Yin
Findings of the Association for Computational Linguistics: ACL 2026
Mingjian Yang | Yong Wang | Peng Liu | Wen Yin
Findings of the Association for Computational Linguistics: ACL 2026
Multimodal Emotion Recognition in Conversation (MERC) relies on integrating heterogeneous signals, yet real-world modality missingness frequently disrupts these systems. We contend that missingness is not merely a loss of data fidelity but a rupture of the fine-grained inter-modal causal chains essential for reasoning. Existing methods, which primarily focus on statistical reconstruction, often fail to bridge these logical gaps, effectively leaving semantic holes. To address this, we propose the Causal-Enhanced Mixture-of-Experts and Hypergraph Network (CaM-HG), employing a “restore-then-mine” paradigm. First, a Causal-Enhanced MoE module conditions experts on historical context to synthesize missing features that are both realistic and causally consistent, thereby patching the broken topology. Subsequently, an Asymmetric Causal Dynamic Hypergraph mines high-order correlations from the restored graph while enforcing strict temporal causality. Experiments on IEMOCAP, CMU-MOSI, and CMU-MOSEI show consistent improvements in terms of WAF1 and accuracy over strong baselines, e.g., surpassing SOTA benchmarks by 1.43% and 1.25% on IEMOCAP. The source code is included in the supplementary material.
2025
Cross-MoE: An Efficient Temporal Prediction Framework Integrating Textual Modality
Ruizheng Huang | Zhicheng Zhang | Yong Wang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Ruizheng Huang | Zhicheng Zhang | Yong Wang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
It has been demonstrated that incorporating external information as textual modality can effectively improve time series forecasting accuracy. However, current multi-modal models ignore the dynamic and different relations between time series patterns and textual features, which leads to poor performance in temporal-textual feature fusion. In this paper, we propose a lightweight and model-agnostic temporal-textual fusion framework named Cross-MoE. It replaces Cross Attention with Cross-Ranker to reduce computational complexity, and enhances modality-aware correlation memorization with Mixture-of-Experts (MoE) networks to tolerate the distributional shifts in time series. The experimental results demonstrate a 8.78% average reduction in Mean Squared Error (MSE) compared to the SOTA multi-modal time series framework. Notably, our method requires only 75% of computational overhead and 12.5% of activated parameters compared with Cross Attention mechanism. Our codes are available at https://github.com/Kilosigh/Cross-MoE.git