Evaluating Data Augmentation Strategies for Training Spanish Misspelling Detection Models

Manuel Castillo-Sancho, Jordi Porta, Asunción Gómez-Pérez


Abstract
This paper evaluates three data augmentation strategies for training misspelling detection models in Spanish. Using the Spanish CORRSIC corpus of naturally occurring misspellings, we compare three misspelling generation methods: random perturbations, keyboard-based errors, and a statistical model derived from empirical edit patterns encoded as weighted finite-state transducers. We also analyze two word selection strategies (random and length-based) and two augmentation configurations designed to balance data diversity and reduce spurious correlations. This study shows that the statistical model produces misspellings most similar to real data, showing the lowest Jensen–Shannon divergence (0.148 nats) with the empirical distribution. In downstream detection experiments, performance improves with training size, and differences between word selection strategies remain minimal. Overall, the results highlight the value of statistically grounded misspelling generation for realistic and effective data augmentation in spell-checking tasks in Spanish.
Anthology ID:
2026.cawl-1.7
Volume:
Proceedings of the Third Workshop on Computation and Written Language (CAWL 2026) @ LREC 2026
Month:
June
Year:
2026
Address:
Palma de Mallorca, Spain
Editor:
Kyle Gorman
Venues:
CAWL | WS
SIG:
SIGWrit
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
71–78
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-cawl-07
DOI:
10.63317/3mw3y4ovzwsz
Bibkey:
Cite (ACL):
Manuel Castillo-Sancho, Jordi Porta, and Asunción Gómez-Pérez. 2026. Evaluating Data Augmentation Strategies for Training Spanish Misspelling Detection Models. In Proceedings of the Third Workshop on Computation and Written Language (CAWL 2026) @ LREC 2026, pages 71–78, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
Evaluating Data Augmentation Strategies for Training Spanish Misspelling Detection Models (Castillo-Sancho et al., CAWL 2026)
Copy Citation: