An Evaluation of Agreement and Uncertainty in LLM-Based Automated Essay Scoring

Yiting Yao, Yanyun Yang, Huan (Hailey) Kuang


Abstract
This study evaluates score agreement and uncertainty estimation in large language model (LLM)-based automated essay scoring. Results showed that LLMs tended to be overconfident and highly consistent in their assigned scores, and majority voting did not yield stronger agreement with human scores than deterministic scoring.
Anthology ID:
2026.aimecon-main.31
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
280–287
Language:
URL:
https://aclanthology.org/2026.aimecon-main.31/
DOI:
Bibkey:
Cite (ACL):
Yiting Yao, Yanyun Yang, and Huan (Hailey) Kuang. 2026. An Evaluation of Agreement and Uncertainty in LLM-Based Automated Essay Scoring. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 280–287, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
An Evaluation of Agreement and Uncertainty in LLM-Based Automated Essay Scoring (Yao et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.31.pdf