Autoscoring Anticlimax: A Meta-analytic Understanding of AI’s Short-answer Shortcomings and Wording Weaknesses

Michael Hardy


Abstract
We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies. We estimate the factors that contribute to LLM performance, quantifying the disconnect between human and LLM difficulties in SAS and provide recommendations for developers working on SAS for schoolchildren.
Anthology ID:
2026.aimecon-main.13
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
117–138
Language:
URL:
https://aclanthology.org/2026.aimecon-main.13/
DOI:
Bibkey:
Cite (ACL):
Michael Hardy. 2026. Autoscoring Anticlimax: A Meta-analytic Understanding of AI’s Short-answer Shortcomings and Wording Weaknesses. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 117–138, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Autoscoring Anticlimax: A Meta-analytic Understanding of AI’s Short-answer Shortcomings and Wording Weaknesses (Hardy, AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.13.pdf