MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions

Abdullatif Köksal; Marion Thaler; Ayyoob Imani; Ahmet Üstün; Anna Korhonen; Hinrich Schütze

doi:10.1162/tacl.a.18

MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions

Abdullatif Köksal, Marion Thaler, Ayyoob Imani, Ahmet Üstün, Anna Korhonen, Hinrich Schütze

Abstract

Instruction tuning enhances large language models (LLMs) by aligning them with human preferences across diverse tasks. Traditional approaches to create instruction tuning datasets face serious challenges for low-resource languages due to their dependence on data annotation. This work introduces a novel method, Multilingual Reverse Instructions (MURI), which generates high-quality instruction tuning datasets for low-resource languages without requiring human annotators or pre-existing multilingual models. Utilizing reverse instructions and a translation pipeline, MURI produces instruction-output pairs from existing human-written texts in low-resource languages. This method ensures cultural relevance and diversity by sourcing texts from different native domains and applying filters to eliminate inappropriate content. Our dataset, MURI-IT, includes more than 2 million instruction-output pairs across 200 languages. Evaluation by native speakers and fine-tuning experiments with mT5 models demonstrate the approach’s effectiveness for both NLU and open-ended generation. We publicly release datasets and models at https://github.com/akoksal/muri.

Anthology ID:: 2025.tacl-1.48
Volume:: Transactions of the Association for Computational Linguistics, Volume 13
Month:
Year:: 2025
Address:: Cambridge, MA
Venue:: TACL
SIG:
Publisher:: MIT Press
Note:
Pages:: 1032–1055
Language:
URL:: https://aclanthology.org/2025.tacl-1.48/
DOI:: 10.1162/tacl.a.18
Bibkey:
Cite (ACL):: Abdullatif Köksal, Marion Thaler, Ayyoob Imani, Ahmet Üstün, Anna Korhonen, and Hinrich Schütze. 2025. MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions. Transactions of the Association for Computational Linguistics, 13:1032–1055.
Cite (Informal):: MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions (Köksal et al., TACL 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.tacl-1.48.pdf

PDF Cite Search Fix data