Constructing a Japanese Claim Decomposition Dataset for Fact-Checking of LLM-Generated Texts

Miwa Masano, Ribeka Keyaki, Atsushi Keyaki, Rei Minamoto, Kaito Horio, Hirokazu Kiyomaru, Kouta Nakayama, Hideyuki Tachibana, Daisuke Kawahara


Abstract
Since texts generated by large language models (LLMs) may contain misinformation (hallucinations), develop- ing fact-checking systems capable of assessing their veracity has become increasingly important. One of the mainstream approaches to fact-checking is the claim-based one, which first decomposes a generated text into claims, i.e., independent and atomic units of information. Each claim is then used as a query to retrieve supporting evidence, and a verdict is predicted for each claim-evidence pair. Conducting fact-checking at the claim level enhances the explainability of verification results. However, achieving highly accurate verification requires that the text be decomposed into claims at an appropriate level of granularity. To address this, we constructed a dataset for Japanese claim decomposition. As part of this dataset construction, we design detailed guidelines for claim decomposition, ensuring that the extracted claims are in a form useful for fact-checking and that the decomposition rules mitigate annotator variability. Quantitative evaluation confirmed that the constructed dataset is of high quality. Additionally, experiments on prompt-based claim decomposition using the constructed dataset demonstrated that adding high-quality few-shot examples and guidelines to prompts improved performance.
Anthology ID:
2026.lrec-1.186
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
2362–2375
Language:
External URL:
https://lrec.elra.info/lrec2026-main-186
DOI:
10.63317/5nsactrpnuu6
Bibkey:
Cite (ACL):
Miwa Masano, Ribeka Keyaki, Atsushi Keyaki, Rei Minamoto, Kaito Horio, Hirokazu Kiyomaru, Kouta Nakayama, Hideyuki Tachibana, and Daisuke Kawahara. 2026. Constructing a Japanese Claim Decomposition Dataset for Fact-Checking of LLM-Generated Texts. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 2362–2375, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Constructing a Japanese Claim Decomposition Dataset for Fact-Checking of LLM-Generated Texts (Masano et al., LREC 2026)
Copy Citation: