Xintian Gao
Author directory2026
Can LLMs Judge Pedagogy? Assessing Conversational AI STEM Tutoring with AI-as-a-Judge
Xintian Gao | Qian Shen | Xin Li
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Xintian Gao | Qian Shen | Xin Li
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Large language models are increasingly used as automated judges in education, yet their ability to score pedagogical quality in AI tutor responses to K-12 STEM student inquiries remains underexplored. This study evaluates whether two LLM-based scorers, Nemotron-3-Super-120B-A12B and GPT-OSS-120B, can approximate human judgment of single-turn Socratic style responses to STEM inquiries. Using a four-dimension rubric adapted from the CPS-R Questioning and Thinking subscale, we compare human and model ratings. Results show mixed reliability: agreement is stronger for more observable instructional features such as Cognitive Demand, but weaker for more interpretive dimensions, especially Encouraging Metacognition and Differentiation. Chance-corrected reliability is also sensitive to skewed score distributions, as shown by a base-rate effect in Differentiation. A mixed-effects analysis further reveals that the two LLM scorers differ systematically, with larger divergence on elementary-level items. We also observe prompt-adherence failures in generated tutoring responses, where some outputs briefly violate the instruction to avoid direct answers before returning to a Socratic response. Overall, the findings suggest that LLMs can assist with large-scale pedagogical evaluation, but human oversight remains necessary for nuanced instructional assessment and for maintaining Socratic response style.