Align and Shine: Building High-quality Sentence-aligned Corpora for Multilingual Text Simplification

Luis Kenji Hilasaca Sanchez, Nouran Khallaf, Serge Sharoff


Abstract
Text simplification plays a crucial role in improving the accessibility and comprehensibility of written information for diverse audiences, including language learners and readers with limited literacy. Despite its importance, large-scale, high-quality datasets for training and evaluating text simplification models remain scarce for languages other than English. This paper reports an experimental study on the collection and processing of crowd-sourced simplification data to construct a corpus suitable for both training and testing text simplification systems across multiple languages (Catalan, English, French, Italian and Spanish). We report mechanisms for sentence-level alignment from document-level data. The resulting dataset of the aligned sentence pairs is publicly available.
Anthology ID:
2026.bucc-1.8
Volume:
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Reinhard Rapp, Ayla Rigouts Terryn, Serge Sharoff, Pierre Zweigenbaum
Venues:
BUCC | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
62–71
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-bucc-08
DOI:
10.63317/55pt8xqgkge6
Bibkey:
Cite (ACL):
Luis Kenji Hilasaca Sanchez, Nouran Khallaf, and Serge Sharoff. 2026. Align and Shine: Building High-quality Sentence-aligned Corpora for Multilingual Text Simplification. In Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC), pages 62–71, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
Align and Shine: Building High-quality Sentence-aligned Corpora for Multilingual Text Simplification (Hilasaca Sanchez et al., BUCC 2026)
Copy Citation: