Improving Speech Recognition with Jargon Injection

Minh-Tien Nguyen, Dat Phuoc Nguyen, Tuan-Hai Luu, Xuan-Quang Nguyen, Tung-Duong Nguyen, Jeff Yang


Abstract
This paper introduces a new method that improves the performance of Automatic speech recognition (ASR) engines, e.g., Whisper in practical cases. Different from prior methods that usually require both speech data and its transcription for decoding, our method only uses jargon as the context for decoding. To do that, the method first represents the jargon in a trie tree structure for efficient storing and traversing. The method next forces the decoding of Whisper to more focus on the jargon by adjusting the probability of generated tokens with the use of the trie tree. To further improve the performance, the method utilizes the prompting method that uses the jargon as the context. Final tokens are generated based on the combination of prompting and decoding. Experimental results on Japanese and English datasets show that the proposed method helps to improve the performance of Whisper, specially for domain-specific data. The method is simple but effective and can be deployed to any encoder-decoder ASR engines in actual cases. The code and data are also accessible (https://shorturl.at/nMsaY).
Anthology ID:
2024.sigdial-1.42
Volume:
Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Month:
September
Year:
2024
Address:
Kyoto, Japan
Editors:
Tatsuya Kawahara, Vera Demberg, Stefan Ultes, Koji Inoue, Shikib Mehri, David Howcroft, Kazunori Komatani
Venue:
SIGDIAL
SIG:
SIGDIAL
Publisher:
Association for Computational Linguistics
Note:
Pages:
490–499
Language:
URL:
https://aclanthology.org/2024.sigdial-1.42
DOI:
10.18653/v1/2024.sigdial-1.42
Bibkey:
Cite (ACL):
Minh-Tien Nguyen, Dat Phuoc Nguyen, Tuan-Hai Luu, Xuan-Quang Nguyen, Tung-Duong Nguyen, and Jeff Yang. 2024. Improving Speech Recognition with Jargon Injection. In Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 490–499, Kyoto, Japan. Association for Computational Linguistics.
Cite (Informal):
Improving Speech Recognition with Jargon Injection (Nguyen et al., SIGDIAL 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.sigdial-1.42.pdf