Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language Resources

Abdisalam Badel, Ting Zhong, Wenxin Tai, Fan Zhou


Abstract
Despite the growing use of the Somali language in various online domains, research on Somali language information retrieval remains limited and primarily relies on query translation due to the lack of a dedicated corpus. To address this problem, we collaborated with language experts and natural language processing (NLP) researchers to create an annotated corpus for Somali information retrieval. This corpus comprises 2335 documents collected from various well-known online sites, such as hiiraan online, dhacdo net, and Somali poetry books. We explain how the corpus was constructed, and develop a Somali language information retrieval system using a pseudo-relevance feedback (PRF) query expansion technique on the corpus. Note that collecting such a data set for the low-resourced Somali language can help overcome NLP barriers, such as the lack of electronically available data sets. Which, if available, can enable the development of various NLP tools and applications such as question-answering and text classification. It also provides researchers with a valuable resource for investigating and developing new techniques and approaches for Somali.
Anthology ID:
2023.emnlp-main.462
Volume:
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
Month:
December
Year:
2023
Address:
Singapore
Editors:
Houda Bouamor, Juan Pino, Kalika Bali
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
7463–7469
Language:
URL:
https://aclanthology.org/2023.emnlp-main.462
DOI:
10.18653/v1/2023.emnlp-main.462
Bibkey:
Cite (ACL):
Abdisalam Badel, Ting Zhong, Wenxin Tai, and Fan Zhou. 2023. Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language Resources. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7463–7469, Singapore. Association for Computational Linguistics.
Cite (Informal):
Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language Resources (Badel et al., EMNLP 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.emnlp-main.462.pdf
Video:
 https://aclanthology.org/2023.emnlp-main.462.mp4