The EZ-AI System for Formosa Speech Recognition Challenge 2025

Yu-Sheng Tsao; Hung-Yang Sung; An-Ci Peng; Jhih-Rong Guo; Tien-Hong Lo

The EZ-AI System for Formosa Speech Recognition Challenge 2025

Yu-Sheng Tsao, Hung-Yang Sung, An-Ci Peng, Jhih-Rong Guo, Tien-Hong Lo

Abstract

This study presents our system for Hakka Speech Recognition Challenge 2025. We designed and compared different systems for two low-resource dialects: Dapu and Zhaoan. On the Pinyin track, we gain boosts by leveraging cross-lingual transfer-learning from related languages and combining with self-supervised learning (SSL). For the Hanzi track, we employ pretrained Whisper with Low-Rank Adaptation (LoRA) fine-tuning. To alleviate the low-resource issue, two data augmentation methods are experimented with: simulating conversational speech to handle multi-speaker scenarios, and generating additional corpus via text-to-speech (TTS). Results from the pilot test showed that transfer learning significantly improved performance in the Pinyin track, achieving an average character error rate (CER) of 19.57%, ranking third among all teams. While in the Hanzi track, the Whisper + LoRA system achieved an average CER of 6.84%, earning first place among all. This study demonstrates that transfer learning and data augmentation can effectively improve recognition performance for low-resource languages. However, the domain mismatch seen in the media test set remains a challenge. We plan to explore in-context learning (ICL) and hotword modeling in the future to better address this issue.

Anthology ID:: 2025.rocling-main.57
Volume:: Proceedings of the 37th Conference on Computational Linguistics and Speech Processing (ROCLING 2025)
Month:: November
Year:: 2025
Address:: National Taiwan University, Taipei City, Taiwan
Editors:: Kai-Wei Chang, Ke-Han Lu, Chih-Kai Yang, Zhi-Rui Tam, Wen-Yu Chang, Chung-Che Wang
Venue:: ROCLING
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 476–480
Language:
URL:: https://aclanthology.org/2025.rocling-main.57/
DOI:
Bibkey:
Cite (ACL):: Yu-Sheng Tsao, Hung-Yang Sung, An-Ci Peng, Jhih-Rong Guo, and Tien-Hong Lo. 2025. The EZ-AI System for Formosa Speech Recognition Challenge 2025. In Proceedings of the 37th Conference on Computational Linguistics and Speech Processing (ROCLING 2025), pages 476–480, National Taiwan University, Taipei City, Taiwan. Association for Computational Linguistics.
Cite (Informal):: The EZ-AI System for Formosa Speech Recognition Challenge 2025 (Tsao et al., ROCLING 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.rocling-main.57.pdf

PDF Cite Search Fix data