Burhan Ogut
Author directory2026
When One Benchmark Hides Many Truths: Mixture-IRT and LLM Difficulty Prediction
Enis Dogan | Burhan Ogut
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Enis Dogan | Burhan Ogut
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Studies that validate LLM-predicted item difficulty conventionally benchmark predictions against a single-population estimate. Using a multimodal LLM’s pairwise judgments on a 29-item mathematics exam, mixture Rasch modeling shows the LLM tracks difficulty ordering differentially across classes, meaningful subgroup variation that the aggregate benchmark conceals entirely.