Benchmarking Text Embedding Models for South African Languages

Ockert de Villiers, Roald Eiselen


Abstract
In this work we introduce a collection of monolingual embedding models for ten South African languages in four different architectures. To determine the quality of the embedding models we evaluate the embeddings on two sequence-labelling tasks, namely Part-of-Speech (POS) tagging and Named Entity Recognition (NER). Languages are grouped into conjunctive (isiNdebele, isiXhosa, isiZulu, and Siswati), disjunctive (Sepedi, Sesotho, Setswana, Tshivenḓa, and Xitsonga), and Afrikaans to establish the influence of training data set size and typology on the quality of the different embeddings. To isolate representation effects we train BiLSTM-CRF taggers, while keeping the architecture, data splits, and training budget fixed, varying only the input imbedding representations, namely GloVe, fastText, Flair, and RoBERTa. In our experiments, GloVe lags behind fastText, Flair, and the transformer-based models, confirming that static word-level vectors are less suited to morphologically complex, low-resource languages. Subword-aware embeddings such as fastText remain a reliable and computationally efficient baseline, while Flair is the most competitive overall across both POS tagging and NER tasks.
Anthology ID:
2026.rail-1.5
Volume:
Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Muzi Matfunjwa, Mmasibidi Setaka, Rooweither Mabuya, Menno van Zaanen
Venues:
RAIL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
41–51
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-rail-05
DOI:
10.63317/28nicfxsyx92
Bibkey:
Cite (ACL):
Ockert de Villiers and Roald Eiselen. 2026. Benchmarking Text Embedding Models for South African Languages. In Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026, pages 41–51, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Benchmarking Text Embedding Models for South African Languages (de Villiers & Eiselen, RAIL 2026)
Copy Citation: