Maria Cutumisu
Author directory2026
Composite Scores vs. Preference Rankings: Measuring Architecture Effects in LLM Feedback
Harvey Ngoe Kolle | Carrie Demmans Epp | Amna Liaqat | Maria Cutumisu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Harvey Ngoe Kolle | Carrie Demmans Epp | Amna Liaqat | Maria Cutumisu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
AI-generated feedback is often judged by quality ratings or preference rankings but rarely both. Comparing a multi-agent system, a single-agent system, and human feedback on student writing, we show that the two evaluation approaches can support different reported conclusions even when their underlying effects are nearly identical. These differences have consequences for how AI-generated feedback should be evaluated.