@inproceedings{ajayi-etal-2026-hayo,
title = "{H}a{Y}o: Repurposing {D}ia{S}afety Dataset for Dialogue Safety Evaluation in {H}ausa and {Y}oruba",
author = "Ajayi, Tunde Oluwaseyi and
Ashaolu, Bolade Deborah and
Lawan, Falalu Ibrahim and
Abolade, Daud Olamide and
Abubakar, Amina Imam and
Akinrinde, Oluwatosin Ayomide and
Gadanya, Murja Sani and
Ashaolu, Omodolapo Dorcas and
Auwal, Abubakar Khalid and
Awujoola, Adewumi and
Adamu, Shamsuddeen Umaru and
Ashaolu, Israel Olawole and
Arcan, Mihael and
Buitelaar, Paul",
editor = "Matfunjwa, Muzi and
Setaka, Mmasibidi and
Mabuya, Rooweither and
van Zaanen, Menno",
booktitle = "Proceedings of Resources for {A}frican Indigenous Languages ({RAIL}) 2026 @ {LREC} 2026",
month = may,
year = "2026",
address = "Palma, Mallorca (Spain)",
publisher = "ELRA Language Resources Association (ELRA)",
url = "https://aclanthology.org/2026.rail-1.9/",
doi = "10.63317/2dot9b59z24c",
pages = "84--95",
abstract = "Research efforts aimed at detecting unsafe dialogues have resulted in creation of benchmark datasets and models for evaluation. The benchmarks mostly exist in English and other high resourced languages. In order to address the challenge of unavailability of dialogue safety evaluation dataset in Hausa and Yor{\`u}b{\'a}, we repurporse DiaSafety dataset to develop HaYo dataset, by providing contextualised human annotation of dialogues in DiaSafety. We provide dialogues in Hausa and Yor{\`u}b{\'a}, obtained by human translation of dialogues in the DiaSafety dataset, to raters who are native speakers. The dialogues are annotated as Unsafe or Safe. We evaluate seven models with moderation, conversational or multilingual capabilities in terms of F1 Score. Using McNemar test, we observe that the predictions of GPT-4.1 and Gemma-3-12b-it on HaYo are statistically significant at p {\ensuremath{<}} 0.05. In our evaluation with instructions in English, we observe lower F1 scores in six out of the seven models, comparing the performance on DiaSafety and HaYo labels. The model predictions were inconsistent with the labels in the HaYo dataset when instructions and dialogues were provided in Hausa and Yor{\`u}b{\'a}. Compared to providing instructions in English, the issues range from responses in unspecified languages to underperformance in terms of F1 score. We plan to release the HaYo dataset to the public to promote dialogue safety research, especially in under-resourced languages."
}<?xml version="1.0" encoding="UTF-8"?>
<modsCollection xmlns="http://www.loc.gov/mods/v3">
<mods ID="ajayi-etal-2026-hayo">
<titleInfo>
<title>HaYo: Repurposing DiaSafety Dataset for Dialogue Safety Evaluation in Hausa and Yoruba</title>
</titleInfo>
<name type="personal">
<namePart type="given">Tunde</namePart>
<namePart type="given">Oluwaseyi</namePart>
<namePart type="family">Ajayi</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Bolade</namePart>
<namePart type="given">Deborah</namePart>
<namePart type="family">Ashaolu</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Falalu</namePart>
<namePart type="given">Ibrahim</namePart>
<namePart type="family">Lawan</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Daud</namePart>
<namePart type="given">Olamide</namePart>
<namePart type="family">Abolade</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Amina</namePart>
<namePart type="given">Imam</namePart>
<namePart type="family">Abubakar</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Oluwatosin</namePart>
<namePart type="given">Ayomide</namePart>
<namePart type="family">Akinrinde</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Murja</namePart>
<namePart type="given">Sani</namePart>
<namePart type="family">Gadanya</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Omodolapo</namePart>
<namePart type="given">Dorcas</namePart>
<namePart type="family">Ashaolu</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Abubakar</namePart>
<namePart type="given">Khalid</namePart>
<namePart type="family">Auwal</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Adewumi</namePart>
<namePart type="family">Awujoola</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Shamsuddeen</namePart>
<namePart type="given">Umaru</namePart>
<namePart type="family">Adamu</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Israel</namePart>
<namePart type="given">Olawole</namePart>
<namePart type="family">Ashaolu</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Mihael</namePart>
<namePart type="family">Arcan</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Paul</namePart>
<namePart type="family">Buitelaar</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<originInfo>
<dateIssued>2026-05</dateIssued>
</originInfo>
<typeOfResource>text</typeOfResource>
<relatedItem type="host">
<titleInfo>
<title>Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026</title>
</titleInfo>
<name type="personal">
<namePart type="given">Muzi</namePart>
<namePart type="family">Matfunjwa</namePart>
<role>
<roleTerm authority="marcrelator" type="text">editor</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Mmasibidi</namePart>
<namePart type="family">Setaka</namePart>
<role>
<roleTerm authority="marcrelator" type="text">editor</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Rooweither</namePart>
<namePart type="family">Mabuya</namePart>
<role>
<roleTerm authority="marcrelator" type="text">editor</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Menno</namePart>
<namePart type="family">van Zaanen</namePart>
<role>
<roleTerm authority="marcrelator" type="text">editor</roleTerm>
</role>
</name>
<originInfo>
<publisher>ELRA Language Resources Association (ELRA)</publisher>
<place>
<placeTerm type="text">Palma, Mallorca (Spain)</placeTerm>
</place>
</originInfo>
<genre authority="marcgt">conference publication</genre>
</relatedItem>
<abstract>Research efforts aimed at detecting unsafe dialogues have resulted in creation of benchmark datasets and models for evaluation. The benchmarks mostly exist in English and other high resourced languages. In order to address the challenge of unavailability of dialogue safety evaluation dataset in Hausa and Yorùbá, we repurporse DiaSafety dataset to develop HaYo dataset, by providing contextualised human annotation of dialogues in DiaSafety. We provide dialogues in Hausa and Yorùbá, obtained by human translation of dialogues in the DiaSafety dataset, to raters who are native speakers. The dialogues are annotated as Unsafe or Safe. We evaluate seven models with moderation, conversational or multilingual capabilities in terms of F1 Score. Using McNemar test, we observe that the predictions of GPT-4.1 and Gemma-3-12b-it on HaYo are statistically significant at p \ensuremath< 0.05. In our evaluation with instructions in English, we observe lower F1 scores in six out of the seven models, comparing the performance on DiaSafety and HaYo labels. The model predictions were inconsistent with the labels in the HaYo dataset when instructions and dialogues were provided in Hausa and Yorùbá. Compared to providing instructions in English, the issues range from responses in unspecified languages to underperformance in terms of F1 score. We plan to release the HaYo dataset to the public to promote dialogue safety research, especially in under-resourced languages.</abstract>
<identifier type="citekey">ajayi-etal-2026-hayo</identifier>
<identifier type="doi">10.63317/2dot9b59z24c</identifier>
<location>
<url>https://aclanthology.org/2026.rail-1.9/</url>
</location>
<part>
<date>2026-05</date>
<extent unit="page">
<start>84</start>
<end>95</end>
</extent>
</part>
</mods>
</modsCollection>
%0 Conference Proceedings
%T HaYo: Repurposing DiaSafety Dataset for Dialogue Safety Evaluation in Hausa and Yoruba
%A Ajayi, Tunde Oluwaseyi
%A Ashaolu, Bolade Deborah
%A Lawan, Falalu Ibrahim
%A Abolade, Daud Olamide
%A Abubakar, Amina Imam
%A Akinrinde, Oluwatosin Ayomide
%A Gadanya, Murja Sani
%A Ashaolu, Omodolapo Dorcas
%A Auwal, Abubakar Khalid
%A Awujoola, Adewumi
%A Adamu, Shamsuddeen Umaru
%A Ashaolu, Israel Olawole
%A Arcan, Mihael
%A Buitelaar, Paul
%Y Matfunjwa, Muzi
%Y Setaka, Mmasibidi
%Y Mabuya, Rooweither
%Y van Zaanen, Menno
%S Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026
%D 2026
%8 May
%I ELRA Language Resources Association (ELRA)
%C Palma, Mallorca (Spain)
%F ajayi-etal-2026-hayo
%X Research efforts aimed at detecting unsafe dialogues have resulted in creation of benchmark datasets and models for evaluation. The benchmarks mostly exist in English and other high resourced languages. In order to address the challenge of unavailability of dialogue safety evaluation dataset in Hausa and Yorùbá, we repurporse DiaSafety dataset to develop HaYo dataset, by providing contextualised human annotation of dialogues in DiaSafety. We provide dialogues in Hausa and Yorùbá, obtained by human translation of dialogues in the DiaSafety dataset, to raters who are native speakers. The dialogues are annotated as Unsafe or Safe. We evaluate seven models with moderation, conversational or multilingual capabilities in terms of F1 Score. Using McNemar test, we observe that the predictions of GPT-4.1 and Gemma-3-12b-it on HaYo are statistically significant at p \ensuremath< 0.05. In our evaluation with instructions in English, we observe lower F1 scores in six out of the seven models, comparing the performance on DiaSafety and HaYo labels. The model predictions were inconsistent with the labels in the HaYo dataset when instructions and dialogues were provided in Hausa and Yorùbá. Compared to providing instructions in English, the issues range from responses in unspecified languages to underperformance in terms of F1 score. We plan to release the HaYo dataset to the public to promote dialogue safety research, especially in under-resourced languages.
%R 10.63317/2dot9b59z24c
%U https://aclanthology.org/2026.rail-1.9/
%U https://doi.org/10.63317/2dot9b59z24c
%P 84-95
Markdown (Informal)
[HaYo: Repurposing DiaSafety Dataset for Dialogue Safety Evaluation in Hausa and Yoruba](https://aclanthology.org/2026.rail-1.9/) (Ajayi et al., RAIL 2026)
ACL
- Tunde Oluwaseyi Ajayi, Bolade Deborah Ashaolu, Falalu Ibrahim Lawan, Daud Olamide Abolade, Amina Imam Abubakar, Oluwatosin Ayomide Akinrinde, Murja Sani Gadanya, Omodolapo Dorcas Ashaolu, Abubakar Khalid Auwal, Adewumi Awujoola, Shamsuddeen Umaru Adamu, Israel Olawole Ashaolu, Mihael Arcan, and Paul Buitelaar. 2026. HaYo: Repurposing DiaSafety Dataset for Dialogue Safety Evaluation in Hausa and Yoruba. In Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026, pages 84–95, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).