Lexicon Design for Transcription of Spontaneous Voice Messages
Michal Gishri | Vered Silber-Varod | Ami Moyal
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Building a comprehensive pronunciation lexicon is a crucial element in the success of any speech recognition engine. The first stage of lexicon design involves the compilation of a comprehensive word list that keeps the Out-Of-Vocabulary (OOV) word rate to a minimum. The second stage involves providing optimized phonemic representations for all lexical items on the list. The research presented here focuses on the first stage of lexicon design ― word list compilation, and describes the methodologies employed in the collection of a pronunciation lexicon designed for the purpose of American English voice message transcription using speech recognition. The lexicon design used is based on a topic domain structure with a target of 90% word coverage for each domain. This differs somewhat from standard approaches where probable words from textual corpora are extracted. This paper raises four issues involved in lexicon design for the transcription of spontaneous voice messages: the inclusion of interjections and other characteristics common to spontaneous speech; the identification of unique messaging terminology; the relative ratio of proper nouns to common words; and the overall size of the lexicon.
Database Adaptation for Speech Recognition in Cross-Environmental Conditions
Oren Gedge | Christophe Couvreur | Klaus Linhard | Shaunie Shammass | Ami Moyal
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Creation of Spoken Hebrew Databases
Tami Rannon | Ofra Golani | Anat Goren | Sherrie Shammass | Ami Moyal
Proceedings of the Second International Conference on Language Resources and Evaluation (LREC’00)
- Michal Gishri 1
- Vered Silber-Varod 1
- Oren Gedge 1
- Christophe Couvreur 1
- Klaus Linhard 1
- show all...