The Added Value of Metadata and Annotations: Evidence from Two Large-Scale, Naturalistic Corpus Studies

Anisia Popescu, Johanna Cronenberg, Ioana Vasilescu, Ioana Chitoran, Lori Lamel, Martine Adda-Decker


Abstract
This paper presents two case studies that highlight both the challenges and benefits of working with large-scale, naturalistic phonetic data. Our aim is to encourage researchers not to shy away from phonetic data found “in the wild”, even when such data are messy, noisy, or incomplete – because they can yield robust, novel insights beyond the reach of controlled laboratory studies. We focus on challenges that are endemic to large corpora, including degraded audio quality, sparse or inconsistent annotations, and missing speaker metadata. By comparing two corpus-based studies that diverge in methodology and statistical design, we show how different approaches can mitigate these limitations while still extracting meaningful patterns.
Anthology ID:
2026.lrec-1.455
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5767–5775
Language:
External URL:
https://lrec.elra.info/lrec2026-main-455
DOI:
10.63317/4c4triganae3
Bibkey:
Cite (ACL):
Anisia Popescu, Johanna Cronenberg, Ioana Vasilescu, Ioana Chitoran, Lori Lamel, and Martine Adda-Decker. 2026. The Added Value of Metadata and Annotations: Evidence from Two Large-Scale, Naturalistic Corpus Studies. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5767–5775, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
The Added Value of Metadata and Annotations: Evidence from Two Large-Scale, Naturalistic Corpus Studies (Popescu et al., LREC 2026)
Copy Citation: