Creation and Validation of Large Lexica for Speech-to-Speech Translation Purposes
Hanne Fersøe | Elviira Hartikainen | Henk van den Heuvel | Giulio Maltese | Asuncíon Moreno | Shaunie Shammass | Ute Ziegenhain
Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04)
This paper presents specifications and requirements for creation and validation of large lexica that are needed in automatic Speech Recognition (ASR), Text-to-Speech (TTS) and statistical Speech-to-Speech Translation (SST) systems. The prepared language resources are created and validated within the scope of the EU-project LC-STAR (Lexica and Corpora for Speech-to-Speech Translation Components) during years 2002-2005. Large lexica consisting of phonetic, suprasegmental and morpho-syntactic content will be provided with well-documented specifications for 13 languages. A short summary of the LC-STAR project itself is presented. Overview about the specification for the corpora collection and word extraction as well as the specification and format of the lexica are presented. Particular attention is paid to the validation of the produced lexica and the lessons learnt during pre-validation. The created and validated language resources will be available via ELRA/ELDA.
Database Adaptation for Speech Recognition in Cross-Environmental Conditions
Oren Gedge | Christophe Couvreur | Klaus Linhard | Shaunie Shammass | Ami Moyal
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
- Hanne Fersøe 1
- Elviira Hartikainen 1
- Henk van den Heuvel 1
- Giulio Maltese 1
- Asunción Moreno 1
- show all...