Enis Dogan

Author directory

2026

Studies that validate LLM-predicted item difficulty conventionally benchmark predictions against a single-population estimate. Using a multimodal LLM’s pairwise judgments on a 29-item mathematics exam, mixture Rasch modeling shows the LLM tracks difficulty ordering differentially across classes, meaningful subgroup variation that the aggregate benchmark conceals entirely.