Benjamin Domingue
Author directory2026
Consensus without Accuracy: Investigating LLM’s Recovery of Item Difficulty Using Paired Comparisons
Michael Leon Chrzan | Benjamin Domingue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Michael Leon Chrzan | Benjamin Domingue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
This study evaluates whether large language models (LLMs) can recover item difficulty estimates through Bradley-Terry modeling from pairwise comparisons of items. Across five Item Response Warehouse datasets and four LLMs, we examine alignment between LLM pairwise-derived difficulty rankings and 1PL IRT parameters, with implications for scalable, AI-assisted item calibration.
Can LLMs Interpret Psychometric Parameters? Quantifying "Over-Knowledge" in LLM Student Simulation
Kazunori Fukuhara | Benjamin Domingue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Kazunori Fukuhara | Benjamin Domingue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
LLM simulation suffers from an “over-knowledge problem”, where models perform too well to represent struggling learners. We compare persona-, IRT-, and CDM-based prompting for student simulation and measure how well methods reflect expected behaviors and quantify this bias. Findings show LLMs follow IRT parameters, yet struggle to simulate real abilities.
2025
Dynamic Bayesian Item Response Model with Decomposition (D-BIRD): Modeling Cohort and Individual Learning Over Time
Hansol Lee | Jason B. Cho | David S. Matteson | Benjamin Domingue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Hansol Lee | Jason B. Cho | David S. Matteson | Benjamin Domingue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
We present D-BIRD, a Bayesian dynamic item response model for estimating student ability from sparse, longitudinal assessments. By decomposing ability into a cohort trend and individual trajectory, D-BIRD supports interpretable modeling of learning over time. We evaluate parameter recovery in simulation and demonstrate the model using real-world personalized learning data.