Developing an LLM Tutor Quality Evaluation Scale

Michael Xiao, Zewei Tian, Alex Liu, Lief Esbenshade, Min Sun


Abstract
We introduce a unified metric set for evaluating LLM tutors by inductively coding 177 recent literature-based metrics and scoring tutoring dialogues with three LLM judges. Exploratory factor analysis identified five constructs and a strong general factor. We present the resulting 10-category, 41-item framework for confirmatory analysis and human validation.
Anthology ID:
2026.aimecon-sessions.15
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
124–145
Language:
URL:
https://aclanthology.org/2026.aimecon-sessions.15/
DOI:
Bibkey:
Cite (ACL):
Michael Xiao, Zewei Tian, Alex Liu, Lief Esbenshade, and Min Sun. 2026. Developing an LLM Tutor Quality Evaluation Scale. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 124–145, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Developing an LLM Tutor Quality Evaluation Scale (Xiao et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-sessions.15.pdf