Judging the LLM Judges: A Human-Centric Validation of LLM-Generated Training Data for Software Retrieval

Ogtay Hasanov, Saad Ezzini


Abstract
As Large Language Models (LLMs) increasingly generate training data for downstream machine learning systems, the quality of this synthetic data becomes a critical security concern. Low-quality synthetic training data can silently poison retrieval systems deployed in security-sensitive contexts such as software issue triage, user support, and threat intelligence matching. We present a multi-dimensional quality assessment protocol for LLM-generated synthetic training data and apply it to a case study involving 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications. We evaluate 400 stratified samples using an LLM judge (GPT-4o-mini) along a five-point rubric, find that 10.5% of generated reviews fail to meaningfully capture their source issues, and identify systematic failure patterns concentrated in developer-internal issues (continuous integration, refactoring) and sarcastic persona framings. To validate the LLM judge against human annotation, we compute Cohen’s Kappa between one human rater, a second independent human rater, and the LLM judge on 20 stratified reviews. Our results highlight the need for hybrid human-AI protocols when assessing synthetic data quality for security-critical applications.
Anthology ID:
2026.nlpaics-1.20
Volume:
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Month:
June
Year:
2026
Address:
Alicante, Spain
Editors:
Ruslan Mitkov, Rafael Muñoz, Elena Lloret, Tharindu Ranasinghe, Ernesto L. Estevanell-Valladares, Salima Lamsiyah, Andrés Montoyo, Saad Ezzini
Venue:
NLPAICS
SIG:
Publisher:
Department of Languages and Information Systems, University of Alicante
Note:
Pages:
184–188
Language:
URL:
https://aclanthology.org/2026.nlpaics-1.20/
DOI:
Bibkey:
Cite (ACL):
Ogtay Hasanov and Saad Ezzini. 2026. Judging the LLM Judges: A Human-Centric Validation of LLM-Generated Training Data for Software Retrieval. In Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security, pages 184–188, Alicante, Spain. Department of Languages and Information Systems, University of Alicante.
Cite (Informal):
Judging the LLM Judges: A Human-Centric Validation of LLM-Generated Training Data for Software Retrieval (Hasanov & Ezzini, NLPAICS 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.nlpaics-1.20.pdf