Can LLMs Judge Pedagogy? Assessing Conversational AI STEM Tutoring with AI-as-a-Judge

Xintian Gao, Qian Shen, Xin Li


Abstract
Large language models are increasingly used as automated judges in education, yet their ability to score pedagogical quality in AI tutor responses to K-12 STEM student inquiries remains underexplored. This study evaluates whether two LLM-based scorers, Nemotron-3-Super-120B-A12B and GPT-OSS-120B, can approximate human judgment of single-turn Socratic style responses to STEM inquiries. Using a four-dimension rubric adapted from the CPS-R Questioning and Thinking subscale, we compare human and model ratings. Results show mixed reliability: agreement is stronger for more observable instructional features such as Cognitive Demand, but weaker for more interpretive dimensions, especially Encouraging Metacognition and Differentiation. Chance-corrected reliability is also sensitive to skewed score distributions, as shown by a base-rate effect in Differentiation. A mixed-effects analysis further reveals that the two LLM scorers differ systematically, with larger divergence on elementary-level items. We also observe prompt-adherence failures in generated tutoring responses, where some outputs briefly violate the instruction to avoid direct answers before returning to a Socratic response. Overall, the findings suggest that LLMs can assist with large-scale pedagogical evaluation, but human oversight remains necessary for nuanced instructional assessment and for maintaining Socratic response style.
Anthology ID:
2026.aimecon-main.69
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
612–621
Language:
URL:
https://aclanthology.org/2026.aimecon-main.69/
DOI:
Bibkey:
Cite (ACL):
Xintian Gao, Qian Shen, and Xin Li. 2026. Can LLMs Judge Pedagogy? Assessing Conversational AI STEM Tutoring with AI-as-a-Judge. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 612–621, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Can LLMs Judge Pedagogy? Assessing Conversational AI STEM Tutoring with AI-as-a-Judge (Gao et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.69.pdf