Can Small Language Models Teach Math Through Socratic Dialogue?

Dizhi Zhang, Mengchen Li, Zichen Wang, Ziqiao Zeng, Shruti Mehta, Talita de Paula Cypriano de Souza, Seiji Isotani


Abstract
This study evaluates whether three locally deployed small language models (LLaMA-3.2-1B, Qwen-2.5-0.5B, and Gemma-3-1B) can implement Socratic tutoring in Grade 6–8 mathematics. Using a rubric-based protocol across 72 sessions, results show that all models struggled to sustain guided inquiry, with Gemma performing best overall.
Anthology ID:
2026.aimecon-main.67
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
599–605
Language:
URL:
https://aclanthology.org/2026.aimecon-main.67/
DOI:
Bibkey:
Cite (ACL):
Dizhi Zhang, Mengchen Li, Zichen Wang, Ziqiao Zeng, Shruti Mehta, Talita de Paula Cypriano de Souza, and Seiji Isotani. 2026. Can Small Language Models Teach Math Through Socratic Dialogue?. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 599–605, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Can Small Language Models Teach Math Through Socratic Dialogue? (Zhang et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.67.pdf