Benefit of a Class-based Language Model for Real-time Closed-captioning of TV Ice-hockey Commentaries

Jan Hoidekr, J.V. Psutka, Aleš Pražák, Josef Psutka


Abstract
This article describes the real-time speech recognition system for closed-captioning of TV ice-hockey commentaries. Automatic transcription of TV commentary accompanying an ice-hockey match is usually a hard task due to the spontaneous speech of a commentator put often into a very loud background noise created by the public, music, siren, drums, whistle, etc. Data for building this system was collected from 41 matches that were played during World Championships in years 2000, 2001, and 2002 and were transmitted by the Czech TV channels. The real-time closed-captioning system is based on the class-based language model designed after careful analysis of training data and OOV words in new (till now unseen) commentaries with the goal to decrease an OOV (Out-Of-Vocabulary) rate and increase recognition accuracy.
Anthology ID:
L06-1375
Volume:
Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06)
Month:
May
Year:
2006
Address:
Genoa, Italy
Editors:
Nicoletta Calzolari, Khalid Choukri, Aldo Gangemi, Bente Maegaard, Joseph Mariani, Jan Odijk, Daniel Tapias
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2006/pdf/615_pdf.pdf
DOI:
Bibkey:
Cite (ACL):
Jan Hoidekr, J.V. Psutka, Aleš Pražák, and Josef Psutka. 2006. Benefit of a Class-based Language Model for Real-time Closed-captioning of TV Ice-hockey Commentaries. In Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06), Genoa, Italy. European Language Resources Association (ELRA).
Cite (Informal):
Benefit of a Class-based Language Model for Real-time Closed-captioning of TV Ice-hockey Commentaries (Hoidekr et al., LREC 2006)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2006/pdf/615_pdf.pdf