Bridging the Temporal Gap in Multimodal LLMs: Deeply Stacking Temporal Tokens for Audio-Visual Speech Recognition

Liyong Wang, Junliang Xing, Tianyu Hu, Jianfei Jiang, Jihuai Zhao, Huimin Ma


Abstract
Audio-Visual Speech Recognition enhances speech recognition robustness in noisy conditions by leveraging visual cues. However, current Multimodal LLMs suffer from a fundamental temporal gap. This gap is characterized by limited fine-grained temporal modeling in vision encoders and progressive temporal semantic degradation throughout the deep layers of LLM decoders. To bridge this gap, we propose a novel framework that deeply stacks temporal tokens across both the encoding and decoding stages. Specifically, we enhance the vision encoder with a temporal-aware attention module and temporal rotary positional embeddings to precisely capture the sequential evolution and dynamics of lip movements. Furthermore, we stack hierarchical temporal tokens that incorporate temporally enriched features into multiple layers of the LLM decoder in a bottom-up manner. Extensive experiments on the LRS2 and LRS3 benchmarks demonstrate that our approach achieves high efficiency and firm performance, outperforming existing supervised, self-supervised, and LLM-based methods by 6.1% on LRS2 and 7.8% on LRS3.
Anthology ID:
2026.findings-acl.1381
Volume:
Findings of the Association for Computational Linguistics: ACL 2026
Month:
July
Year:
2026
Address:
San Diego, California, United States
Editors:
Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
27748–27759
Language:
URL:
https://aclanthology.org/2026.findings-acl.1381/
DOI:
10.18653/v1/2026.findings-acl.1381
Bibkey:
Cite (ACL):
Liyong Wang, Junliang Xing, Tianyu Hu, Jianfei Jiang, Jihuai Zhao, and Huimin Ma. 2026. Bridging the Temporal Gap in Multimodal LLMs: Deeply Stacking Temporal Tokens for Audio-Visual Speech Recognition. In Findings of the Association for Computational Linguistics: ACL 2026, pages 27748–27759, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):
Bridging the Temporal Gap in Multimodal LLMs: Deeply Stacking Temporal Tokens for Audio-Visual Speech Recognition (Wang et al., Findings 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.findings-acl.1381.pdf
Checklist:
 2026.findings-acl.1381.checklist.pdf