Gregory M. Jacobs
Author directory2026
How Much Does Hyperparameter Tuning Actually Help? An Efficiency Survey for Fine-Tuning Transformers to Score Mathematics-Explanation Items
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Fine-tuning pre-trained transformer models for constructed-response items often begins with a hyperparameter grid search using k-fold cross-validation. We study how much that search actually helps by fine-tuning two encoders—MathBERT, a smaller math-focused model, and DeBERTa-v3-large, a larger general-purpose model—to score 18 math-explanation items from a state-wide assessment. Crossing learning rate, weight decay, and label smoothing over two epochs and five folds (1,800 model-fold runs), we compare each model’s tuning gain directly against fold-to-fold noise via a gain-to-noise ratio. Gains were small relative to fold noise for both encoders: MathBERT’s ratio fell below one (0.90), and DeBERTa’s nominally higher ratio (1.42) traced to a handful of divergent fits rather than an informative search landscape. Weight decay and label smoothing were effectively inert, leaving learning rate as the only hyperparameter worth checking—though even for learning rate the best configurations offered small performance gains over a reasonable default. We accordingly recommend a lean workflow that fixes the inert hyperparameters to sensible defaults, runs a narrow learning-rate search extended modestly upward, and increases the epoch budget with early stopping, substantially reducing training compute without sacrificing accuracy.
Math-Specialized or General-Purpose? A Comparison of MathBERT and DeBERTa-v3-large for Automated Scoring of Mathematics-Explanation Items
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Automated scoring of math-explanation items should handle responses that mix natural language with symbolic reasoning. The 2023 NAEP Automated Scoring Challenge showed that fine-tuning pre-trained encoders can reach human-like agreement, but its winning system. used a large, general-purpose encoder, while smaller math-specific pre-trained encoders remain a plausible and more efficient alternative. We fine-tune MathBERT (110M parameters) and DeBERTa-v3-large (183M parameters) as single-output regression scorers on 31 math-explanation items from one statewide summative assessment administration, benchmarking every model against human–human reliability and summarizing deployability with a three-tier acceptance criteria. On the 18 items fine-tuned under both encoders with identical datasets, the resulting performance between the two was practically indistinguishable: mean best QWK was 0.931 for MathBERT and 0.933 for DeBERTa-v3-large—a difference of only 0.002—and the models traded top performance at similar rates. Given similar performance, we tested a two-stage approach that used the larger general-purpose encoder only when the smaller math-focused encoder failed acceptance criteria. Extending MathBERT to all 31 items, it reached a mean QWK of 0.926 against a human–human benchmark of QWK = 0.940 and met operational acceptance criteria on 81% of items. Escalating the six items where a fine-tuned MathBERT model failed acceptance to a fine-tuned DeBERTa-v3-large model recovered four of the six, yielding acceptable automated-scoring models for 29 of 31 items. Because the smaller, math-pre-trained encoder matches the larger one at roughly 40% fewer parameters, we recommend fine-tuning a MathBERT encoder as the default and escalating to the larger, general purpose encoder like DeBERTa- v3-large only when a MathBERT model fails when deploying automated scoring models operationally at scale for math-explanation items.