2026

This study examines the feasibility of using large language models (LLMs) as judges for pairwise comparisons of item difficulty and to the extent which the resulting comparison outcomes recover banked difficulty parameters. The study in particular fine-tunes an instruction-tuned LLM to evaluate the impact of task-specific fine-tuning on the parameter prediction accuracy. The study also demonstrates an agentic approach to hyperparameter tuning of LLM-training task.
Item difficulty prediction relies on the efficient utilization of high-leverage item characteristics. Many items in the domain of mathematics include figures, charts, or other visual stimuli that are challenging to incorporate into difficulty prediction models. Multimodal large language models (LLMs) offer a way to process these visual stimuli in combination with text input, potentially enhancing the success of item difficulty prediction. In this study, we employ several open-source multimodal LLMs to predict the difficulty of math items with visual stimuli from a publicly available data set. We find that multimodal LLMs are capable of predicting math item difficulty.

2025

The study introduces novel approaches for fine-tuning pre-trained LLMs to predict item response theory parameters directly from item texts and structured item attribute variables. The proposed methods were evaluated on a dataset over 1,000 English Language Art items that are currently in the operational pool for a large scale assessment.