Human-in-the-Loop Gemini Item Generation for Large-Scale MCQ Banks

Melchor Sánchez-Mendiola, Milton García-Lima, Elibidú Ortega-Sánchez, Manuel García-Minjares, Enrique Buzo-Casanova


Abstract
This work-in-progress compares 2,405 Spanish-language MCQs drafted by customized Gemini Gems or faculty in a Mexican university. Expert committees reviewed all items. Gemini drafts showed lower validation-friction in three of four areas and stronger blueprint adherence, while faculty drafts offered richer contextualization. Findings support structured AI generation with human validation.
Anthology ID:
2026.aimecon-wip.12
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
86–92
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.12/
DOI:
Bibkey:
Cite (ACL):
Melchor Sánchez-Mendiola, Milton García-Lima, Elibidú Ortega-Sánchez, Manuel García-Minjares, and Enrique Buzo-Casanova. 2026. Human-in-the-Loop Gemini Item Generation for Large-Scale MCQ Banks. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 86–92, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Human-in-the-Loop Gemini Item Generation for Large-Scale MCQ Banks (Sánchez-Mendiola et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.12.pdf