Sakayo Toadoum Sari
2026
Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives
Sakayo Toadoum Sari | Nelly Robin | Michelle Auzanneau | Lakhdar Sais | Véronique Petit | Marie Veniard | Said Jabbour | Fabien Delorme
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Sakayo Toadoum Sari | Nelly Robin | Michelle Auzanneau | Lakhdar Sais | Véronique Petit | Marie Veniard | Said Jabbour | Fabien Delorme
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Migrants traversing geographically distinct routes such as the Trans-Saharan and Balkan corridors often recount strikingly parallel lived experiences: police violence, smuggler exploitation, dangerous crossings, and family separation. We introduce the task of experiential intertextuality detection: automatically identifying shared experiential echoes across migration narratives without requiring annotated training data. From 108 French migration narratives spanning both corridors, we automatically generate sentence pairs and score them using annotation-free methods: lexical baselines, sentence embeddings, POS-based structural features, a migration-specific theme lexicon, context-aware narrative features, and zero-shot LLM scoring with Qwen2.5-7B and Mistral-7B under three prompting strategies. We validate all methods against 816 expertannotated intertextuality judgments (interannotator Krippendorff’s α=0.27). Our results reveal that all surface, structural, and embedding methods correlate only weakly with expert judgments (r≤0.30); Qwen2.5-7B zero-shot achieves the best single-method correlation (r=0.38); few-shot examples degrade Qwen but dramatically improve Mistral; narrative position significantly predicts intertextuality, with departure-phase pairs showing the highest experiential echoes; and a supervised hybrid combining all 31 features achieves r=0.45, a 21% improvement over the best individual method.
2024
AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages
Jiayi Wang | David Ifeoluwa Adelani | Sweta Agrawal | Marek Masiak | Ricardo Rei | Eleftheria Briakou | Marine Carpuat | Xuanli He | Sofia Bourhim | Andiswa Bukula | Muhidin Mohamed | Temitayo Olatoye | Tosin Adewumi | Hamam Mokayed | Christine Mwase | Wangui Kimotho | Foutse Yuehgoh | Anuoluwapo Aremu | Jessica Ojo | Shamsuddeen Hassan Muhammad | Salomey Osei | Abdul-Hakeem Omotayo | Chiamaka Chukwuneke | Perez Ogayo | Oumaima Hourrane | Salma El Anigri | Lolwethu Ndolela | Thabiso Mangwana | Shafie Abdi Mohamed | Ayinde Hassan | Oluwabusayo Olufunke Awoyomi | Lama Alkhaled | Sana Al-Azzawi | Naome A. Etori | Millicent Ochieng | Clemencia Siro | Samuel Njoroge | Eric Muchiri | Wangari Kimotho | Lyse Naomi Wamba Momo | Daud Abolade | Simbiat Ajao | Iyanuoluwa Shode | Ricky Macharm | Ruqayya Nasir Iro | Saheed S. Abdullahi | Stephen E. Moore | Bernard Opoku | Zainab Akinjobi | Abeeb Afolabi | Nnaemeka Obiefuna | Onyekachi Raphael Ogbu | Sam Brian | Verrah Akinyi Otiende | Chinedu Emmanuel Mbonu | Sakayo Toadoum Sari | Yao Lu | Pontus Stenetorp
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Jiayi Wang | David Ifeoluwa Adelani | Sweta Agrawal | Marek Masiak | Ricardo Rei | Eleftheria Briakou | Marine Carpuat | Xuanli He | Sofia Bourhim | Andiswa Bukula | Muhidin Mohamed | Temitayo Olatoye | Tosin Adewumi | Hamam Mokayed | Christine Mwase | Wangui Kimotho | Foutse Yuehgoh | Anuoluwapo Aremu | Jessica Ojo | Shamsuddeen Hassan Muhammad | Salomey Osei | Abdul-Hakeem Omotayo | Chiamaka Chukwuneke | Perez Ogayo | Oumaima Hourrane | Salma El Anigri | Lolwethu Ndolela | Thabiso Mangwana | Shafie Abdi Mohamed | Ayinde Hassan | Oluwabusayo Olufunke Awoyomi | Lama Alkhaled | Sana Al-Azzawi | Naome A. Etori | Millicent Ochieng | Clemencia Siro | Samuel Njoroge | Eric Muchiri | Wangari Kimotho | Lyse Naomi Wamba Momo | Daud Abolade | Simbiat Ajao | Iyanuoluwa Shode | Ricky Macharm | Ruqayya Nasir Iro | Saheed S. Abdullahi | Stephen E. Moore | Bernard Opoku | Zainab Akinjobi | Abeeb Afolabi | Nnaemeka Obiefuna | Onyekachi Raphael Ogbu | Sam Brian | Verrah Akinyi Otiende | Chinedu Emmanuel Mbonu | Sakayo Toadoum Sari | Yao Lu | Pontus Stenetorp
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Despite the recent progress on scaling multilingual machine translation (MT) to several under-resourced African languages, accurately measuring this progress remains challenging, since evaluation is often performed on n-gram matching metrics such as BLEU, which typically show a weaker correlation with human judgments. Learned metrics such as COMET have higher correlation; however, the lack of evaluation data with human ratings for under-resourced languages, complexity of annotation guidelines like Multidimensional Quality Metrics (MQM), and limited language coverage of multilingual encoders have hampered their applicability to African languages. In this paper, we address these challenges by creating high-quality human evaluation data with simplified MQM guidelines for error detection and direct assessment (DA) scoring for 13 typologically diverse African languages. Furthermore, we develop AfriCOMET: COMET evaluation metrics for African languages by leveraging DA data from well-resourced languages and an African-centric multilingual encoder (AfroXLM-R) to create the state-of-the-art MT evaluation metrics for African languages with respect to Spearman-rank correlation with human judgments (0.441).
Search
Fix author
Co-authors
- Saheed S. Abdullahi 1
- Daud Abolade 1
- David Ifeoluwa Adelani 1
- Tosin Adewumi 1
- Abeeb Afolabi 1
- Sweta Agrawal 1
- Simbiat Ajao 1
- Zainab Akinjobi 1
- Sana Al-Azzawi 1
- Lama Alkhaled 1
- Anuoluwapo Aremu 1
- Michelle Auzanneau 1
- Oluwabusayo Olufunke Awoyomi 1
- Sofia Bourhim 1
- Eleftheria Briakou 1
- Sam Brian 1
- Andiswa Bukula 1
- Marine Carpuat 1
- Chiamaka Chukwuneke 1
- Fabien Delorme 1
- Salma El Anigri 1
- Naome A. Etori 1
- Ayinde Hassan 1
- Xuanli He 1
- Oumaima Hourrane 1
- Ruqayya Nasir Iro 1
- Said Jabbour 1
- Wangari Kimotho 1
- Wangui Kimotho 1
- Yao Lu 1
- Ricky Macharm 1
- Thabiso Mangwana 1
- Marek Masiak 1
- Chinedu Emmanuel Mbonu 1
- Muhidin Mohamed 1
- Shafie Abdi Mohamed 1
- Hamam Mokayed 1
- Stephen E. Moore 1
- Eric Muchiri 1
- Shamsuddeen Hassan Muhammad 1
- Christine Mwase 1
- Lolwethu Ndolela 1
- Samuel Njoroge 1
- Nnaemeka Obiefuna 1
- Millicent Ochieng 1
- Perez Ogayo 1
- Onyekachi Raphael Ogbu 1
- Jessica Ojo 1
- Temitayo Olatoye 1
- Abdul-Hakeem Omotayo 1
- Bernard Opoku 1
- Salomey Osei 1
- Verrah Akinyi Otiende 1
- Véronique Petit 1
- Ricardo Rei 1
- Nelly Robin 1
- Lakhdar Sais 1
- Iyanuoluwa Shode 1
- Clemencia Siro 1
- Pontus Stenetorp 1
- Marie Veniard 1
- Lyse Naomi Wamba Momo 1
- Jiayi Wang 1
- Foutse Yuehgoh 1