Seiji Isotani

Author directory

2026

Small language models (SLMs) are increasingly proposed for educational use because they promise lower cost, offline deployment, and stronger data privacy. We report an exploratory evaluation of three sub-2B-parameter models on example-based decimal-arithmetic tutoring. Across structured interactions, all three produced fluent, confident output that masked unstable pedagogy and frequent mathematical errors even on elementary decimal addition and place-value tasks. Building on these observations, we describe an emerging interaction-based measurement framework intended to support more defensible readiness decisions about SLMs as math tutors.
This study evaluates whether three locally deployed small language models (LLaMA-3.2-1B, Qwen-2.5-0.5B, and Gemma-3-1B) can implement Socratic tutoring in Grade 6–8 mathematics. Using a rubric-based protocol across 72 sessions, results show that all models struggled to sustain guided inquiry, with Gemma performing best overall.