MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, Salomey Osei
Abstract
We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1- Anthology ID:
- 2021.tacl-1.66
- Volume:
- Transactions of the Association for Computational Linguistics, Volume 9
- Month:
- Year:
- 2021
- Address:
- Cambridge, MA
- Editors:
- Brian Roark, Ani Nenkova
- Venue:
- TACL
- SIG:
- Publisher:
- MIT Press
- Note:
- Pages:
- 1116–1131
- Language:
- URL:
- https://aclanthology.org/2021.tacl-1.66
- DOI:
- 10.1162/tacl_a_00416
- Bibkey:
- Cite (ACL):
- David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, et al.. 2021. MasakhaNER: Named Entity Recognition for African Languages. Transactions of the Association for Computational Linguistics, 9:1116–1131.
- Cite (Informal):
- MasakhaNER: Named Entity Recognition for African Languages (Adelani et al., TACL 2021)
- Copy Citation:
- PDF:
- https://aclanthology.org/2021.tacl-1.66.pdf
- Video:
- https://aclanthology.org/2021.tacl-1.66.mp4
Export citation
@article{adelani-etal-2021-masakhaner, title = "{M}asakha{NER}: Named Entity Recognition for {A}frican Languages", author = "Adelani, David Ifeoluwa and Abbott, Jade and Neubig, Graham and D{'}souza, Daniel and Kreutzer, Julia and Lignos, Constantine and Palen-Michel, Chester and Buzaaba, Happy and Rijhwani, Shruti and Ruder, Sebastian and Mayhew, Stephen and Azime, Israel Abebe and Muhammad, Shamsuddeen H. and Emezue, Chris Chinenye and Nakatumba-Nabende, Joyce and Ogayo, Perez and Anuoluwapo, Aremu and Gitau, Catherine and Mbaye, Derguene and Alabi, Jesujoba and Yimam, Seid Muhie and Gwadabe, Tajuddeen Rabiu and Ezeani, Ignatius and Niyongabo, Rubungo Andre and Mukiibi, Jonathan and Otiende, Verrah and Orife, Iroro and David, Davis and Ngom, Samba and Adewumi, Tosin and Rayson, Paul and Adeyemi, Mofetoluwa and Muriuki, Gerald and Anebi, Emmanuel and Chukwuneke, Chiamaka and Odu, Nkiruka and Wairagala, Eric Peter and Oyerinde, Samuel and Siro, Clemencia and Bateesa, Tobius Saul and Oloyede, Temilola and Wambui, Yvonne and Akinode, Victor and Nabagereka, Deborah and Katusiime, Maurice and Awokoya, Ayodele and MBOUP, Mouhamadane and Gebreyohannes, Dibora and Tilaye, Henok and Nwaike, Kelechi and Wolde, Degaga and Faye, Abdoulaye and Sibanda, Blessing and Ahia, Orevaoghene and Dossou, Bonaventure F. P. and Ogueji, Kelechi and DIOP, Thierno Ibrahima and Diallo, Abdoulaye and Akinfaderin, Adewale and Marengereke, Tendai and Osei, Salomey", editor = "Roark, Brian and Nenkova, Ani", journal = "Transactions of the Association for Computational Linguistics", volume = "9", year = "2021", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/2021.tacl-1.66", doi = "10.1162/tacl_a_00416", pages = "1116--1131", abstract = "We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1", }
<?xml version="1.0" encoding="UTF-8"?> <modsCollection xmlns="http://www.loc.gov/mods/v3"> <mods ID="adelani-etal-2021-masakhaner"> <titleInfo> <title>MasakhaNER: Named Entity Recognition for African Languages</title> </titleInfo> <name type="personal"> <namePart type="given">David</namePart> <namePart type="given">Ifeoluwa</namePart> <namePart type="family">Adelani</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Jade</namePart> <namePart type="family">Abbott</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Graham</namePart> <namePart type="family">Neubig</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Daniel</namePart> <namePart type="family">D’souza</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Julia</namePart> <namePart type="family">Kreutzer</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Constantine</namePart> <namePart type="family">Lignos</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Chester</namePart> <namePart type="family">Palen-Michel</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Happy</namePart> <namePart type="family">Buzaaba</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Shruti</namePart> <namePart type="family">Rijhwani</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Sebastian</namePart> <namePart type="family">Ruder</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Stephen</namePart> <namePart type="family">Mayhew</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Israel</namePart> <namePart type="given">Abebe</namePart> <namePart type="family">Azime</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Shamsuddeen</namePart> <namePart type="given">H</namePart> <namePart type="family">Muhammad</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Chris</namePart> <namePart type="given">Chinenye</namePart> <namePart type="family">Emezue</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Joyce</namePart> <namePart type="family">Nakatumba-Nabende</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Perez</namePart> <namePart type="family">Ogayo</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Aremu</namePart> <namePart type="family">Anuoluwapo</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Catherine</namePart> <namePart type="family">Gitau</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Derguene</namePart> <namePart type="family">Mbaye</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Jesujoba</namePart> <namePart type="family">Alabi</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Seid</namePart> <namePart type="given">Muhie</namePart> <namePart type="family">Yimam</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Tajuddeen</namePart> <namePart type="given">Rabiu</namePart> <namePart type="family">Gwadabe</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Ignatius</namePart> <namePart type="family">Ezeani</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Rubungo</namePart> <namePart type="given">Andre</namePart> <namePart type="family">Niyongabo</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Jonathan</namePart> <namePart type="family">Mukiibi</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Verrah</namePart> <namePart type="family">Otiende</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Iroro</namePart> <namePart type="family">Orife</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Davis</namePart> <namePart type="family">David</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Samba</namePart> <namePart type="family">Ngom</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Tosin</namePart> <namePart type="family">Adewumi</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Paul</namePart> <namePart type="family">Rayson</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Mofetoluwa</namePart> <namePart type="family">Adeyemi</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Gerald</namePart> <namePart type="family">Muriuki</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Emmanuel</namePart> <namePart type="family">Anebi</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Chiamaka</namePart> <namePart type="family">Chukwuneke</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Nkiruka</namePart> <namePart type="family">Odu</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Eric</namePart> <namePart type="given">Peter</namePart> <namePart type="family">Wairagala</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Samuel</namePart> <namePart type="family">Oyerinde</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Clemencia</namePart> <namePart type="family">Siro</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Tobius</namePart> <namePart type="given">Saul</namePart> <namePart type="family">Bateesa</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Temilola</namePart> <namePart type="family">Oloyede</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Yvonne</namePart> <namePart type="family">Wambui</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Victor</namePart> <namePart type="family">Akinode</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Deborah</namePart> <namePart type="family">Nabagereka</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Maurice</namePart> <namePart type="family">Katusiime</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Ayodele</namePart> <namePart type="family">Awokoya</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Mouhamadane</namePart> <namePart type="family">MBOUP</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Dibora</namePart> <namePart type="family">Gebreyohannes</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Henok</namePart> <namePart type="family">Tilaye</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Kelechi</namePart> <namePart type="family">Nwaike</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Degaga</namePart> <namePart type="family">Wolde</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Abdoulaye</namePart> <namePart type="family">Faye</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Blessing</namePart> <namePart type="family">Sibanda</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Orevaoghene</namePart> <namePart type="family">Ahia</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Bonaventure</namePart> <namePart type="given">F</namePart> <namePart type="given">P</namePart> <namePart type="family">Dossou</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Kelechi</namePart> <namePart type="family">Ogueji</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Thierno</namePart> <namePart type="given">Ibrahima</namePart> <namePart type="family">DIOP</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Abdoulaye</namePart> <namePart type="family">Diallo</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Adewale</namePart> <namePart type="family">Akinfaderin</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Tendai</namePart> <namePart type="family">Marengereke</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Salomey</namePart> <namePart type="family">Osei</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <originInfo> <dateIssued>2021</dateIssued> </originInfo> <typeOfResource>text</typeOfResource> <genre authority="bibutilsgt">journal article</genre> <relatedItem type="host"> <titleInfo> <title>Transactions of the Association for Computational Linguistics</title> </titleInfo> <originInfo> <issuance>continuing</issuance> <publisher>MIT Press</publisher> <place> <placeTerm type="text">Cambridge, MA</placeTerm> </place> </originInfo> <genre authority="marcgt">periodical</genre> <genre authority="bibutilsgt">academic journal</genre> </relatedItem> <abstract>We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1</abstract> <identifier type="citekey">adelani-etal-2021-masakhaner</identifier> <identifier type="doi">10.1162/tacl_a_00416</identifier> <location> <url>https://aclanthology.org/2021.tacl-1.66</url> </location> <part> <date>2021</date> <detail type="volume"><number>9</number></detail> <extent unit="page"> <start>1116</start> <end>1131</end> </extent> </part> </mods> </modsCollection>
%0 Journal Article %T MasakhaNER: Named Entity Recognition for African Languages %A Adelani, David Ifeoluwa %A Abbott, Jade %A Neubig, Graham %A D’souza, Daniel %A Kreutzer, Julia %A Lignos, Constantine %A Palen-Michel, Chester %A Buzaaba, Happy %A Rijhwani, Shruti %A Ruder, Sebastian %A Mayhew, Stephen %A Azime, Israel Abebe %A Muhammad, Shamsuddeen H. %A Emezue, Chris Chinenye %A Nakatumba-Nabende, Joyce %A Ogayo, Perez %A Anuoluwapo, Aremu %A Gitau, Catherine %A Mbaye, Derguene %A Alabi, Jesujoba %A Yimam, Seid Muhie %A Gwadabe, Tajuddeen Rabiu %A Ezeani, Ignatius %A Niyongabo, Rubungo Andre %A Mukiibi, Jonathan %A Otiende, Verrah %A Orife, Iroro %A David, Davis %A Ngom, Samba %A Adewumi, Tosin %A Rayson, Paul %A Adeyemi, Mofetoluwa %A Muriuki, Gerald %A Anebi, Emmanuel %A Chukwuneke, Chiamaka %A Odu, Nkiruka %A Wairagala, Eric Peter %A Oyerinde, Samuel %A Siro, Clemencia %A Bateesa, Tobius Saul %A Oloyede, Temilola %A Wambui, Yvonne %A Akinode, Victor %A Nabagereka, Deborah %A Katusiime, Maurice %A Awokoya, Ayodele %A MBOUP, Mouhamadane %A Gebreyohannes, Dibora %A Tilaye, Henok %A Nwaike, Kelechi %A Wolde, Degaga %A Faye, Abdoulaye %A Sibanda, Blessing %A Ahia, Orevaoghene %A Dossou, Bonaventure F. P. %A Ogueji, Kelechi %A DIOP, Thierno Ibrahima %A Diallo, Abdoulaye %A Akinfaderin, Adewale %A Marengereke, Tendai %A Osei, Salomey %J Transactions of the Association for Computational Linguistics %D 2021 %V 9 %I MIT Press %C Cambridge, MA %F adelani-etal-2021-masakhaner %X We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1 %R 10.1162/tacl_a_00416 %U https://aclanthology.org/2021.tacl-1.66 %U https://doi.org/10.1162/tacl_a_00416 %P 1116-1131
Markdown (Informal)
[MasakhaNER: Named Entity Recognition for African Languages](https://aclanthology.org/2021.tacl-1.66) (Adelani et al., TACL 2021)
- MasakhaNER: Named Entity Recognition for African Languages (Adelani et al., TACL 2021)
ACL
- David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, et al.. 2021. MasakhaNER: Named Entity Recognition for African Languages. Transactions of the Association for Computational Linguistics, 9:1116–1131.