Yuhao Zhang
Author directoryOther people with similar names: Yuhao Zhang, Yuhao Zhang
Unverified author pages with similar names: Yuhao Zhang
2025
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
Yuhao Zhang | Xiangnan Ma | Kaiqi Kou | Peizhuo Liu | Weiqiao Shan | Benyou Wang | Tong Xiao | Yuxin Huang | Zhengtao Yu | JingBo Zhu
Findings of the Association for Computational Linguistics: ACL 2025
Yuhao Zhang | Xiangnan Ma | Kaiqi Kou | Peizhuo Liu | Weiqiao Shan | Benyou Wang | Tong Xiao | Yuxin Huang | Zhengtao Yu | JingBo Zhu
Findings of the Association for Computational Linguistics: ACL 2025
The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long sequences, called cross-lingual (CL). We propose the unit language to overcome the two modeling challenges. The unit language can be considered a text-like representation format, constructed using n-gram language modeling. We implement multi-task learning to utilize the unit language in guiding the speech modeling process. Our initial results reveal a conflict when applying source and target unit languages simultaneously. We propose task prompt modeling to mitigate this conflict. We conduct experiments on four languages of the Voxpupil dataset. Our method demonstrates significant improvements over a strong baseline and achieves performance comparable to models trained with text.
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
Weiqiao Shan | Yuang Li | Yuhao Zhang | Yingfeng Luo | Chen Xu | Xiaofeng Zhao | Long Meng | Yunfei Lu | Min Zhang | Hao Yang | Tong Xiao | JingBo Zhu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Weiqiao Shan | Yuang Li | Yuhao Zhang | Yingfeng Luo | Chen Xu | Xiaofeng Zhao | Long Meng | Yunfei Lu | Min Zhang | Hao Yang | Tong Xiao | JingBo Zhu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Connecting audio encoders with large language models (LLMs) allows the LLM to perform various audio understanding tasks, such as automatic speech recognition (ASR) and audio captioning (AC). Most research focuses on training an adapter layer to generate a unified audio feature for the LLM. However, different tasks may require distinct features that emphasize either semantic or acoustic aspects, making task-specific audio features more desirable. In this paper, we propose Prompt-aware Mixture (PaM) to enhance the Speech LLM that uses multiple audio encoders. Our approach involves using different experts to extract different features based on the prompt that indicates different tasks. Experiments demonstrate that with PaM, only one Speech LLM surpasses the best performances achieved by all single-encoder Speech LLMs on ASR, speaker number verification, and AC tasks. PaM also outperforms other feature fusion baselines, such as concatenation and averaging.
Soundwave: Less is More for Speech-Text Alignment in LLMs
Yuhao Zhang | Zhiheng Liu | Fan Bu | Ruiyu Zhang | Benyou Wang | Haizhou Li
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yuhao Zhang | Zhiheng Liu | Fan Bu | Ruiyu Zhang | Benyou Wang | Haizhou Li
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Existing end-to-end speech large language models (LLMs) usually rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth. We focus on two fundamental problems between speech and text: the representation space gap and sequence length inconsistency. We propose Soundwave, which utilizes an efficient training strategy and a novel architecture to address these issues. Results show that Soundwave outperforms other advanced speech LLMs in speech translation and AIR-Bench speech tasks with only a fraction of the training data. Further analysis shows that Soundwave still retains its intelligence during conversation.