Towards evaluating teacher performance in a GenAI teaching simulation of a science discussion

Beata Beigman Klebanov, Jamie N. Mikeska, Mengxuan Zhao, Catherine Flynn, Devon Fetrow, Shreyashi Halder, Rutuja Ubale, Tricia Maxwell, Michael Suhan


Abstract
GenAI can power simulated student agents that provide opportunities for educators to engage in core teaching practices, such as leading a small group argumentation-based science discussion. To realize the potential of such simulations and support teacher reflection and learning, it is necessary to provide participants with timely feedback on their performance in the simulation. This study investigates systems for automated evaluation of and feedback on teacher performance in a simulation along the dimension of making use of student ideas to move the discussion forward. We address three research questions: (a) How well do models fine-tuned on transcripts of teacher performance in a matching human-puppeteered teaching simulation (that served as the model during the development of the GenAI one) perform in evaluating transcripts from the GenAI teaching simulation? (b) How well does a system using a few-shot LLM perform on the same task? (c) How do educators perceive the quality and usefulness of the automatically generated feedback? The findings underscore the importance of a rigorous evaluation of automated evaluation and feedback systems.
Anthology ID:
2026.aimecon-main.64
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
571–580
Language:
URL:
https://aclanthology.org/2026.aimecon-main.64/
DOI:
Bibkey:
Cite (ACL):
Beata Beigman Klebanov, Jamie N. Mikeska, Mengxuan Zhao, Catherine Flynn, Devon Fetrow, Shreyashi Halder, Rutuja Ubale, Tricia Maxwell, and Michael Suhan. 2026. Towards evaluating teacher performance in a GenAI teaching simulation of a science discussion. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 571–580, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Towards evaluating teacher performance in a GenAI teaching simulation of a science discussion (Beigman Klebanov et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.64.pdf