Ahmed H. Bediwy

Author directory

2026

Explanatory item response models let test de- velopers anticipate item difficulty from item design, but they depend on subject-matter ex- perts rating every item on every hypothesized feature—a step that is slow, costly, and the practical bottleneck limiting how many items can be modeled. We ask whether large lan- guage models can supply those ratings. Four trained experts and six LLMs independently rated 15 grade-5 mathematics items on six psy- chometric features. We evaluate the LLM rat- ings twice: against the human consensus using quadratic-weighted 𝜅and mixed-effects mod- els, and against 1,452 student responses by us- ing each source’s ratings as the design matrix of a linear logistic test model (LLTM) bench- marked against a Rasch baseline. Claude Opus 4.7 performed best in recovering the item dif- ficulty (r = 0.89), however, the human con- sensus against which agreement is measured performs worst at recovering Rasch difficulty (r = 0.33) compared to the rest of the LLMs. The reason is visible in the expert panel it- self: inter-rater Fleiss’ 𝜅is at or below chance on three of six features, so the consensus is a noisy reference rather than ground truth. We also identify two failure modes (zero-variance and perfect collinearity) that agreement statis- tics cannot detect but make a feature unusable as an LLTM covariate. We argue that agree- ment should be reported alongside criterion- referenced recovery of item parameters, not in place of it.
Fine-tuning pre-trained transformer models for constructed-response items often begins with a hyperparameter grid search using k-fold cross-validation. We study how much that search actually helps by fine-tuning two encoders—MathBERT, a smaller math-focused model, and DeBERTa-v3-large, a larger general-purpose model—to score 18 math-explanation items from a state-wide assessment. Crossing learning rate, weight decay, and label smoothing over two epochs and five folds (1,800 model-fold runs), we compare each model’s tuning gain directly against fold-to-fold noise via a gain-to-noise ratio. Gains were small relative to fold noise for both encoders: MathBERT’s ratio fell below one (0.90), and DeBERTa’s nominally higher ratio (1.42) traced to a handful of divergent fits rather than an informative search landscape. Weight decay and label smoothing were effectively inert, leaving learning rate as the only hyperparameter worth checking—though even for learning rate the best configurations offered small performance gains over a reasonable default. We accordingly recommend a lean workflow that fixes the inert hyperparameters to sensible defaults, runs a narrow learning-rate search extended modestly upward, and increases the epoch budget with early stopping, substantially reducing training compute without sacrificing accuracy.
Automated scoring of math-explanation items should handle responses that mix natural language with symbolic reasoning. The 2023 NAEP Automated Scoring Challenge showed that fine-tuning pre-trained encoders can reach human-like agreement, but its winning system. used a large, general-purpose encoder, while smaller math-specific pre-trained encoders remain a plausible and more efficient alternative. We fine-tune MathBERT (110M parameters) and DeBERTa-v3-large (183M parameters) as single-output regression scorers on 31 math-explanation items from one statewide summative assessment administration, benchmarking every model against human–human reliability and summarizing deployability with a three-tier acceptance criteria. On the 18 items fine-tuned under both encoders with identical datasets, the resulting performance between the two was practically indistinguishable: mean best QWK was 0.931 for MathBERT and 0.933 for DeBERTa-v3-large—a difference of only 0.002—and the models traded top performance at similar rates. Given similar performance, we tested a two-stage approach that used the larger general-purpose encoder only when the smaller math-focused encoder failed acceptance criteria. Extending MathBERT to all 31 items, it reached a mean QWK of 0.926 against a human–human benchmark of QWK = 0.940 and met operational acceptance criteria on 81% of items. Escalating the six items where a fine-tuned MathBERT model failed acceptance to a fine-tuned DeBERTa-v3-large model recovered four of the six, yielding acceptable automated-scoring models for 29 of 31 items. Because the smaller, math-pre-trained encoder matches the larger one at roughly 40% fewer parameters, we recommend fine-tuning a MathBERT encoder as the default and escalating to the larger, general purpose encoder like DeBERTa- v3-large only when a MathBERT model fails when deploying automated scoring models operationally at scale for math-explanation items.
One approach to automated essay scoring re- lies on neural regression models that output continuous predicted scores, yet operational scores must be reported on a discrete, ordi- nal rubric scale to match human raters’ scores. The mapping from continuous predictions to integer scores is governed by a small set of cutpoints, and the choice of cutpoint esti- mator materially changes both accuracy and the score distribution that examinees expe- rience. We introduce, formalize, and com- pare seven cutpoint methods—simple round- ing, exact-agreement optimization, balance op- timization, quadratic-weighted-kappa (QWK) maximization, maximum-likelihood Gaussian boundaries, a Bayesian ordered-prior estimator, and an ordinal cumulative link model—under a single notation. We evaluate all seven on seven prompts from the ASAP 2.0 dataset (24,728 essays), using a shared Longformer regres- sion backbone, and report agreement (quadratic weighted kappa), association (the Pearson cor- relation between integer machine and human scores), and distributional fidelity (the stan- dardized mean difference and the maximum score-point distribution gap). Our results ex- pose a consistent accuracy–calibration trade- off: QWK maximization and balance optimiza- tion tie for the highest agreement, balance opti- mization achieves the smallest score-point dis- tribution gap, the ordinal link model achieves the smallest mean bias, the generative Gaussian and Bayesian estimators recover score means but distort the distribution, and the ordinal link model is a strong all-rounder.