Lief Esbenshade
Author directory2026
Developing an LLM Tutor Quality Evaluation Scale
Michael Xiao | Zewei Tian | Alex Liu | Lief Esbenshade | Min Sun
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Michael Xiao | Zewei Tian | Alex Liu | Lief Esbenshade | Min Sun
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
We introduce a unified metric set for evaluating LLM tutors by inductively coding 177 recent literature-based metrics and scoring tutoring dialogues with three LLM judges. Exploratory factor analysis identified five constructs and a strong general factor. We present the resulting 10-category, 41-item framework for confirmatory analysis and human validation.
2025
Implementation Considerations for Automated AI Grading of Student Work
Zewei Tian | Alex Liu | Lief Esbenshade | Shawon Sarkar | Zachary Zhang | Kevin He | Min Sun
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Zewei Tian | Alex Liu | Lief Esbenshade | Shawon Sarkar | Zachary Zhang | Kevin He | Min Sun
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
19 K-12 teachers participated in a co-design pilot study of an AI education platform, testing assessment grading. Teachers valued AI’s rapid narrative feedback for formative assessment but distrusted automated scoring, preferring human oversight. Students appreciated immediate feedback but remained skeptical of AI-only grading, highlighting needs for trustworthy, teacher-centered AI tools.