Transcription Accuracy in the Icelandic Gigaword Corpus: Evaluating Automatic and Manual Annotation

Johanna Mechler, Lilja Björk Stefánsdóttir, Anton Karl Ingason


Abstract
This paper aims to compare automatic and manually corrected annotation data in the Icelandic Gigaword Corpus. We focus on the variable use of Stylistic Fronting (SF) in Icelandic, an optional movement of words or phrases, which indicates a more formal style. Examining SF rates across time, we find that manual coding results in slightly lower SF rates than automatic coding. This difference can be explained by the different sources used in the coding process: For automatic coding, written transcripts compiled by parliament employees are used, and for manual correction, coding relies on audio files of the parliament speeches. Importantly, both types of coding are well suited to trace changing patterns of SF over a span of 16 years, suggesting that the automatic feature extraction reliably reflects the speeches that have been transcribed.
Anthology ID:
2026.lrec-1.373
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
4757–4764
Language:
External URL:
https://lrec.elra.info/lrec2026-main-373
DOI:
10.63317/4f2rpzig5h8p
Bibkey:
Cite (ACL):
Johanna Mechler, Lilja Björk Stefánsdóttir, and Anton Karl Ingason. 2026. Transcription Accuracy in the Icelandic Gigaword Corpus: Evaluating Automatic and Manual Annotation. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 4757–4764, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Transcription Accuracy in the Icelandic Gigaword Corpus: Evaluating Automatic and Manual Annotation (Mechler et al., LREC 2026)
Copy Citation: