Generation of Instruction and Preference Dataset for Improving Japanese Instruction Following in LLMs

Kei Moriyama, Takashi Kodama, Kouta Nakayama


Abstract
Instruction following, the ability to generate text that aligns with human intent, is a core capability of large language models (LLMs) for real-world applications. Instruction tuning is widely used to obtain this capability, but it requires large amounts of annotated data. To reduce the labor and cost of large-scale annotation, data augmentation using LLMs has been proposed as a promising approach. As this approach has primarily been applied to English datasets, its effectiveness in other languages, such as Japanese, remains unclear. In this paper, we propose an automatic pipeline for generating instruction and preference datasets in Japanese. The instruction dataset is created by expanding a manually annotated dataset using an LLM. The preference dataset is then constructed by adding LLM-generated negative examples to the instruction dataset. To ensure the quality of the datasets, instructions and responses are evaluated using LLM-as-a-Judge and ROUGE-L. Experimental results using supervised fine-tuning and direct preference optimization demonstrate that these synthetic datasets improve the instruction-following capability in Japanese.
Anthology ID:
2026.lrec-1.111
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
1435–1454
Language:
External URL:
https://lrec.elra.info/lrec2026-main-111
DOI:
10.63317/3w8ceszaj7m9
Bibkey:
Cite (ACL):
Kei Moriyama, Takashi Kodama, and Kouta Nakayama. 2026. Generation of Instruction and Preference Dataset for Improving Japanese Instruction Following in LLMs. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 1435–1454, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Generation of Instruction and Preference Dataset for Improving Japanese Instruction Following in LLMs (Moriyama et al., LREC 2026)
Copy Citation: