Cheick Tidiani Cissé


2026

Neural machine translation for extremely low-resource languages faces compounding challenges: scarce parallel data, orthographic inconsistency, and absence of quality metadata for principled training. We present Kumatigi, a quality-annotated French-Bambara corpus combining systematic curation with data augmentation strategies tailored to Bambara. We provide 67k quality-scored pairs that enable targeted data filtering and address pervasive orthographic normalization issues in existing resources. Our dual-dataset generation framework strategically exploits round-trip translation, producing synthetic pairs for fluency reinforcement alongside back-translated pairs that preserve authentic vocabulary for coverage expansion. We further introduce linguistically-motivated augmentation techniques addressing Bambara’s orthographic variability, improving model robustness for real-world text. Experiments with LoRA-based fine-tuning demonstrate consistent improvements across automatic metrics, with our full system achieving up to +3–4 BLEU over strong baselines. Data generation and augmentation strategies contribute +1-2 BLEU beyond high-quality parallel data alone. Human evaluation by native speakers confirms these automatic improvements align with substantial gains in translation adequacy and fluency, with our best model approaching human reference translation quality. Our methodology provides a reproducible framework applicable to other under-resourced languages facing similar data challenges.
Search
Co-authors
    Venues
    Fix author