Shreyashi Halder
Author directory2026
Data-lean fine-tuning of models for evaluating teacher performance in a GenAI-led elicitation simulation
Beata Beigman Klebanov | Andrew Hoang | Jamie Mikeska | Benny Longwill | Sanjna Kashyap | Shreyashi Halder | Aakanksha Bhatia
Proceedings of the 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026)
Beata Beigman Klebanov | Andrew Hoang | Jamie Mikeska | Benny Longwill | Sanjna Kashyap | Shreyashi Halder | Aakanksha Bhatia
Proceedings of the 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026)
Recent advances in the capabilities of conversational agents based on large language models make them a very promising tool for role playing K-12 students in order to train educators in conversational teaching practices, such as eliciting student thinking, explaining disciplinary content, and facilitating a classroom discussion. In fact, such simulations can and have been developed relatively quickly and without data to machine-learn from – neither classroom data nor human-simulated data. To enhance the usefulness and effectiveness of such teaching simulations, it is necessary to provide pedagogically sound, timely, and personalized feedback to the educator about their simulation performance. In this study, we present experiments on fine-tuning models to evaluate educator performance in an elicitation teaching simulation. The models are developed with data collected during usability testing of the simulation and evaluated on real user data. We show that even with relatively little fine-tuning data, robust performance can be obtained
Evaluating Educators’ Instructional Skills on Making Content Explicit in Instruction in a Generative AI Teaching Simulation
Shreyashi Halder | Jamie N. Mikeska | Heather Jorgenson
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Shreyashi Halder | Jamie N. Mikeska | Heather Jorgenson
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This paper examines elementary educators’ instructional skills as they practice eliciting student thinking in a generative AI (GenAI) mathematics teaching simulation. Study findings indicate that educators were able to elicit some aspects of a GenAI student’s conceptual understanding and misunderstanding. Implications for using GenAI teaching simulations as a practice space and formative assessment tool to build and evaluate educators’ instructional skills are addressed.
Towards evaluating teacher performance in a GenAI teaching simulation of a science discussion
Beata Beigman Klebanov | Jamie N. Mikeska | Mengxuan Zhao | Catherine Flynn | Devon Fetrow | Shreyashi Halder | Rutuja Ubale | Tricia Maxwell | Michael Suhan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Beata Beigman Klebanov | Jamie N. Mikeska | Mengxuan Zhao | Catherine Flynn | Devon Fetrow | Shreyashi Halder | Rutuja Ubale | Tricia Maxwell | Michael Suhan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
GenAI can power simulated student agents that provide opportunities for educators to engage in core teaching practices, such as leading a small group argumentation-based science discussion. To realize the potential of such simulations and support teacher reflection and learning, it is necessary to provide participants with timely feedback on their performance in the simulation. This study investigates systems for automated evaluation of and feedback on teacher performance in a simulation along the dimension of making use of student ideas to move the discussion forward. We address three research questions: (a) How well do models fine-tuned on transcripts of teacher performance in a matching human-puppeteered teaching simulation (that served as the model during the development of the GenAI one) perform in evaluating transcripts from the GenAI teaching simulation? (b) How well does a system using a few-shot LLM perform on the same task? (c) How do educators perceive the quality and usefulness of the automatically generated feedback? The findings underscore the importance of a rigorous evaluation of automated evaluation and feedback systems.
2025
Generative AI Teaching Simulations as Formative Assessment Tools within Preservice Teacher Preparation
Jamie N. Mikeska | Aakanksha Bhatia | Shreyashi Halder | Tricia Maxwell | Beata Beigman Klebanov | Benny Longwill | Kashish Behl | Calli Shekell
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Jamie N. Mikeska | Aakanksha Bhatia | Shreyashi Halder | Tricia Maxwell | Beata Beigman Klebanov | Benny Longwill | Kashish Behl | Calli Shekell
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This paper examines how generative AI (GenAI) teaching simulations can be used as a formative assessment tool to gain insight into elementary preservice teachers’ (PSTs’) instructional abilities. This study investigated the teaching moves PSTs used to elicit student thinking in a GenAI simulation and their perceptions of the simulation’s