Benjamin Domingue

Author directory

2026

This study evaluates whether large language models (LLMs) can recover item difficulty estimates through Bradley-Terry modeling from pairwise comparisons of items. Across five Item Response Warehouse datasets and four LLMs, we examine alignment between LLM pairwise-derived difficulty rankings and 1PL IRT parameters, with implications for scalable, AI-assisted item calibration.
LLM simulation suffers from an “over-knowledge problem”, where models perform too well to represent struggling learners. We compare persona-, IRT-, and CDM-based prompting for student simulation and measure how well methods reflect expected behaviors and quantify this bias. Findings show LLMs follow IRT parameters, yet struggle to simulate real abilities.

2025

We present D-BIRD, a Bayesian dynamic item response model for estimating student ability from sparse, longitudinal assessments. By decomposing ability into a cohort trend and individual trajectory, D-BIRD supports interpretable modeling of learning over time. We evaluate parameter recovery in simulation and demonstrate the model using real-world personalized learning data.