Jeff Yang
2024
Improving Speech Recognition with Jargon Injection
Minh-Tien Nguyen
|
Dat Phuoc Nguyen
|
Tuan-Hai Luu
|
Xuan-Quang Nguyen
|
Tung-Duong Nguyen
|
Jeff Yang
Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue
This paper introduces a new method that improves the performance of Automatic speech recognition (ASR) engines, e.g., Whisper in practical cases. Different from prior methods that usually require both speech data and its transcription for decoding, our method only uses jargon as the context for decoding. To do that, the method first represents the jargon in a trie tree structure for efficient storing and traversing. The method next forces the decoding of Whisper to more focus on the jargon by adjusting the probability of generated tokens with the use of the trie tree. To further improve the performance, the method utilizes the prompting method that uses the jargon as the context. Final tokens are generated based on the combination of prompting and decoding. Experimental results on Japanese and English datasets show that the proposed method helps to improve the performance of Whisper, specially for domain-specific data. The method is simple but effective and can be deployed to any encoder-decoder ASR engines in actual cases. The code and data are also accessible (https://shorturl.at/nMsaY).
Search