Composite Scores vs. Preference Rankings: Measuring Architecture Effects in LLM Feedback

Harvey Ngoe Kolle, Carrie Demmans Epp, Amna Liaqat, Maria Cutumisu


Abstract
AI-generated feedback is often judged by quality ratings or preference rankings but rarely both. Comparing a multi-agent system, a single-agent system, and human feedback on student writing, we show that the two evaluation approaches can support different reported conclusions even when their underlying effects are nearly identical. These differences have consequences for how AI-generated feedback should be evaluated.
Anthology ID:
2026.aimecon-wip.30
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
231–234
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.30/
DOI:
Bibkey:
Cite (ACL):
Harvey Ngoe Kolle, Carrie Demmans Epp, Amna Liaqat, and Maria Cutumisu. 2026. Composite Scores vs. Preference Rankings: Measuring Architecture Effects in LLM Feedback. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 231–234, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Composite Scores vs. Preference Rankings: Measuring Architecture Effects in LLM Feedback (Kolle et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.30.pdf