Xintian Gao

Author directory

2026

Large language models are increasingly used as automated judges in education, yet their ability to score pedagogical quality in AI tutor responses to K-12 STEM student inquiries remains underexplored. This study evaluates whether two LLM-based scorers, Nemotron-3-Super-120B-A12B and GPT-OSS-120B, can approximate human judgment of single-turn Socratic style responses to STEM inquiries. Using a four-dimension rubric adapted from the CPS-R Questioning and Thinking subscale, we compare human and model ratings. Results show mixed reliability: agreement is stronger for more observable instructional features such as Cognitive Demand, but weaker for more interpretive dimensions, especially Encouraging Metacognition and Differentiation. Chance-corrected reliability is also sensitive to skewed score distributions, as shown by a base-rate effect in Differentiation. A mixed-effects analysis further reveals that the two LLM scorers differ systematically, with larger divergence on elementary-level items. We also observe prompt-adherence failures in generated tutoring responses, where some outputs briefly violate the instruction to avoid direct answers before returning to a Socratic response. Overall, the findings suggest that LLMs can assist with large-scale pedagogical evaluation, but human oversight remains necessary for nuanced instructional assessment and for maintaining Socratic response style.