From Validation to Resilience: Sustaining Measurement Quality in AI-Assisted Enemy Item Identification

Ye Ma


Abstract
Enemy items are item pairs that must not appear on the same test form. An automatic enemy-identification method using large language models (LLMs) has been deployed in operation. This study asks a question: how is measurement quality sustained as conditions shift after deployment? Using operational data from certification exams, two analyses examine the factors that affect the method’s resilience. Analysis 1 isolates model-version and prompt updates: changing the model with the prompt held constant reduced recall from 0.75 to 0.54, while prompt refinement with the model held constant recovered it to 0.82. It also shows that most model–reviewer disagreements are edge cases and that the human standard is itself variable, with four experts spanning 0.70 to 0.91 in recall. Analysis 2 reports a disruption case: the established method works effectively on a professional-level exam but not on a foundational-level exam. A complementary content-tag based method was added to the existing method in response, with human review surfacing the disruption and validating the fix. Responsible LLM deployment in assessment requires monitoring with labeled data, testing before deployment, and evaluation systems tied to each use case, so that validity, reliability, and fairness are sustained rather than certified once.
Anthology ID:
2026.aimecon-sessions.9
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
76–81
Language:
URL:
https://aclanthology.org/2026.aimecon-sessions.9/
DOI:
Bibkey:
Cite (ACL):
Ye Ma. 2026. From Validation to Resilience: Sustaining Measurement Quality in AI-Assisted Enemy Item Identification. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 76–81, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
From Validation to Resilience: Sustaining Measurement Quality in AI-Assisted Enemy Item Identification (Ma, AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-sessions.9.pdf