LLM-Based Pairwise Judgment for Math Item Parameter Modeling Using Workflows and Agents

Suhwa Han, Frank Rijmen


Abstract
This study examines the feasibility of using large language models (LLMs) as judges for pairwise comparisons of item difficulty and to the extent which the resulting comparison outcomes recover banked difficulty parameters. The study in particular fine-tunes an instruction-tuned LLM to evaluate the impact of task-specific fine-tuning on the parameter prediction accuracy. The study also demonstrates an agentic approach to hyperparameter tuning of LLM-training task.
Anthology ID:
2026.aimecon-wip.32
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
248–259
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.32/
DOI:
Bibkey:
Cite (ACL):
Suhwa Han and Frank Rijmen. 2026. LLM-Based Pairwise Judgment for Math Item Parameter Modeling Using Workflows and Agents. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 248–259, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
LLM-Based Pairwise Judgment for Math Item Parameter Modeling Using Workflows and Agents (Han & Rijmen, AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.32.pdf