Seiji Isotani
Author directory2026
Assessing Small Language Models as Decimal-Arithmetic Tutors: A Measurement Framework
Mai Que Vuong | Shahana Ahmadli | Michelle Zhou | Minseok Kim | Shruti Mehta | Talita de Paula Cypriano de Souza | Seiji Isotani
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Mai Que Vuong | Shahana Ahmadli | Michelle Zhou | Minseok Kim | Shruti Mehta | Talita de Paula Cypriano de Souza | Seiji Isotani
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Small language models (SLMs) are increasingly proposed for educational use because they promise lower cost, offline deployment, and stronger data privacy. We report an exploratory evaluation of three sub-2B-parameter models on example-based decimal-arithmetic tutoring. Across structured interactions, all three produced fluent, confident output that masked unstable pedagogy and frequent mathematical errors even on elementary decimal addition and place-value tasks. Building on these observations, we describe an emerging interaction-based measurement framework intended to support more defensible readiness decisions about SLMs as math tutors.
Can Small Language Models Teach Math Through Socratic Dialogue?
Dizhi Zhang | Mengchen Li | Zichen Wang | Ziqiao Zeng | Shruti Mehta | Talita de Paula Cypriano de Souza | Seiji Isotani
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Dizhi Zhang | Mengchen Li | Zichen Wang | Ziqiao Zeng | Shruti Mehta | Talita de Paula Cypriano de Souza | Seiji Isotani
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study evaluates whether three locally deployed small language models (LLaMA-3.2-1B, Qwen-2.5-0.5B, and Gemma-3-1B) can implement Socratic tutoring in Grade 6–8 mathematics. Using a rubric-based protocol across 72 sessions, results show that all models struggled to sustain guided inquiry, with Gemma performing best overall.