Two Uses of Human-scored Anchors: Few-shot Direct Scoring and Anchor-based Comparative Grading

Prashreet Poudel, Huayue Gu, Collin Lynch, Zhikai Gao


Abstract
This study evaluates direct few-shot scoring and anchor based comparative judgment for LLM essay grading. Using human scored anchor essays, we compare their accuracy and stability. Although both approaches produce competitive scores, comparative judgments demonstrate better stability and consistency, suggesting it as a better alternative for grading writing assessment.
Anthology ID:
2026.aimecon-main.51
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
458–465
Language:
URL:
https://aclanthology.org/2026.aimecon-main.51/
DOI:
Bibkey:
Cite (ACL):
Prashreet Poudel, Huayue Gu, Collin Lynch, and Zhikai Gao. 2026. Two Uses of Human-scored Anchors: Few-shot Direct Scoring and Anchor-based Comparative Grading. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 458–465, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Two Uses of Human-scored Anchors: Few-shot Direct Scoring and Anchor-based Comparative Grading (Poudel et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.51.pdf