Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset

Nick Rossenbach, Robin Schmitt, Tina Raissi, Simon Berger, Larissa Kleppel, Ralf Schlüter


Abstract
The recently published Loquacious dataset aims to be a replacement for established English automatic speech recognition (ASR) datasets such as LibriSpeech or TED-Lium. The main goal of Loquacious dataset is to provide properly defined training and test partitions across many acoustic and language domains, with an open license suitable for both academia and industry. To further promote the benchmarking and usability of this new dataset, we present additional resources in the form of n-gram language models (LMs), a grapheme-to-phoneme (G2P) model and pronunciation lexica, with open and public access. Utilizing those additional resources we show experimental results across a wide range of ASR architectures with different label units and topologies. Our initial experimental results indicate that the Loquacious dataset offers a valuable study case for a variety of common challenges in ASR.
Anthology ID:
2026.lrec-1.462
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5839–5848
Language:
External URL:
https://lrec.elra.info/lrec2026-main-462
DOI:
10.63317/4zsvhm25r7zf
Bibkey:
Cite (ACL):
Nick Rossenbach, Robin Schmitt, Tina Raissi, Simon Berger, Larissa Kleppel, and Ralf Schlüter. 2026. Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5839–5848, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset (Rossenbach et al., LREC 2026)
Copy Citation: