Adja-French Parallel Corpus: A New Resource for Machine Translation of a West African Under-Resourced Language

Josue Frejus Godeme, Rolando Coto-Solano


Abstract
We present the first parallel text corpus for Adja machine translation, an under-resourced Gbe language spoken by approximately 1,000,000 people in Benin and Togo. The corpus contains 10,000 French-Adja sentence pairs, providing a foundation for machine translation research. We establish baseline results using fine-tuned NLLB and ByT5 models, achieving a chrF++ of 28 in the French→Adja direction, and up to a chrF++ of 34 in the Adja→French direction. This work represents the first public machine translation resource for Adja. It provides benchmarks for future studies on this under-resourced West African language. The dataset is available at https://huggingface.co/datasets/JosueG/french-adja-parallel-corpus.
Anthology ID:
2026.lrec-1.299
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3742–3749
Language:
External URL:
https://lrec.elra.info/lrec2026-main-299
DOI:
10.63317/5bk4g2k7mpmu
Bibkey:
Cite (ACL):
Josue Frejus Godeme and Rolando Coto-Solano. 2026. Adja-French Parallel Corpus: A New Resource for Machine Translation of a West African Under-Resourced Language. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3742–3749, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Adja-French Parallel Corpus: A New Resource for Machine Translation of a West African Under-Resourced Language (Godeme & Coto-Solano, LREC 2026)
Copy Citation: