Automated Evaluation of Mathematical Equivalence Between Personalized and Standard Word Problems

Burcu Arslan, Ikkyu Choi, Jesse R. Sparks, Reginald M. Gooch, Candace Walkington, Matthew L. Bernacki


Abstract
Generative AI enables real-time and scalable context personalization of mathematics word problems (MWPs) based on students’ self-reported interests during assessment. However, a question arises: are personalized and standard MWPs mathematically equivalent? In this paper, we present two Natural Language Processing pipelines for evaluating mathematical equivalence between personalized and standard MWPs.
Anthology ID:
2026.aimecon-sessions.32
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
291–305
Language:
URL:
https://aclanthology.org/2026.aimecon-sessions.32/
DOI:
Bibkey:
Cite (ACL):
Burcu Arslan, Ikkyu Choi, Jesse R. Sparks, Reginald M. Gooch, Candace Walkington, and Matthew L. Bernacki. 2026. Automated Evaluation of Mathematical Equivalence Between Personalized and Standard Word Problems. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 291–305, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Automated Evaluation of Mathematical Equivalence Between Personalized and Standard Word Problems (Arslan et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-sessions.32.pdf