A Recipe for Adapting Multilingual Embedders to OCR-Error Robustness and Historical Texts

Andrianos Michail, Stylianos Psychias, Juri Opitz, Simon Clematide


Abstract
Modern multilingual text embedding models excel at semantic search on contemporary text but their performance degrades measurably on digitized historical documents. This issue is especially pronounced for underrepresented languages such as Luxembourgish, where historical materials combine evolving spelling conventions with OCR artifacts absent from standard training data. To address these challenges, we introduce OCR M-GTE, a pair of multilingual embedding models adapted for OCR robustness and historical texts, and show that the observed degradation can be mitigated through a simple multi-step training procedure tailored to historical variants and OCR noise. We evaluate the models on standard semantic search tasks, simulated OCR degradation, and genuine historical collections, observing consistent improvements under OCR-induced noise and on genuine historical data while maintaining comparable performance on clean modern text. Our ablation findings suggest that multilingual embedding models can be effectively adapted to perform robust cross-lingual search in heterogeneous European digitized corpora. We release our adapted models, code, and datasets under the AGPL-3.0 license: https://github.com/impresso/ocr-robust-multilingual-embeddings
Anthology ID:
2026.lrec-1.71
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
929–935
Language:
External URL:
https://lrec.elra.info/lrec2026-main-071
DOI:
10.63317/29rfx5wcyz3z
Bibkey:
Cite (ACL):
Andrianos Michail, Stylianos Psychias, Juri Opitz, and Simon Clematide. 2026. A Recipe for Adapting Multilingual Embedders to OCR-Error Robustness and Historical Texts. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 929–935, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
A Recipe for Adapting Multilingual Embedders to OCR-Error Robustness and Historical Texts (Michail et al., LREC 2026)
Copy Citation: