Transformer-Aided Detection of Gaming in Constructed-Response English Language Assessments

Melanie Sharif, Scott Hellman, Martha Bellows, Sue Lottridge


Abstract
We compare handcrafted features, frozen transformer embeddings, and ensembles across three item types and two regimes, testing pooled versus specialist models on a gaming detection task. Ensembles perform best (ROC-AUC 0.95, 𝜅=0.75), except for rarest item type under unseen prompts. Labels reflect review detections, motivating reference-conditioned recall and blind re-review.
Anthology ID:
2026.aimecon-wip.21
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
159–167
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.21/
DOI:
Bibkey:
Cite (ACL):
Melanie Sharif, Scott Hellman, Martha Bellows, and Sue Lottridge. 2026. Transformer-Aided Detection of Gaming in Constructed-Response English Language Assessments. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 159–167, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Transformer-Aided Detection of Gaming in Constructed-Response English Language Assessments (Sharif et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.21.pdf