Martha Bellows
Author directory2026
Construct Validity of Small-Sample Transformer Scoring Models: A Mechanistic Interpretability Approach
Michael P. Hemenway | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Michael P. Hemenway | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Small-sample transformer scoring models ( n = 64 n=64) reach high human agreement but risk leaning on surface shortcuts like response length. Evaluating Mechanistic Interpretability strategies across 90 models, we show correlational methods suffer from seed noise, whereas interventional erasure proves small-sample models causally depend more on length. We outline an actionable audit protocol.
How Much Does Hyperparameter Tuning Actually Help? An Efficiency Survey for Fine-Tuning Transformers to Score Mathematics-Explanation Items
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Fine-tuning pre-trained transformer models for constructed-response items often begins with a hyperparameter grid search using k-fold cross-validation. We study how much that search actually helps by fine-tuning two encoders—MathBERT, a smaller math-focused model, and DeBERTa-v3-large, a larger general-purpose model—to score 18 math-explanation items from a state-wide assessment. Crossing learning rate, weight decay, and label smoothing over two epochs and five folds (1,800 model-fold runs), we compare each model’s tuning gain directly against fold-to-fold noise via a gain-to-noise ratio. Gains were small relative to fold noise for both encoders: MathBERT’s ratio fell below one (0.90), and DeBERTa’s nominally higher ratio (1.42) traced to a handful of divergent fits rather than an informative search landscape. Weight decay and label smoothing were effectively inert, leaving learning rate as the only hyperparameter worth checking—though even for learning rate the best configurations offered small performance gains over a reasonable default. We accordingly recommend a lean workflow that fixes the inert hyperparameters to sensible defaults, runs a narrow learning-rate search extended modestly upward, and increases the epoch budget with early stopping, substantially reducing training compute without sacrificing accuracy.
Math-Specialized or General-Purpose? A Comparison of MathBERT and DeBERTa-v3-large for Automated Scoring of Mathematics-Explanation Items
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Automated scoring of math-explanation items should handle responses that mix natural language with symbolic reasoning. The 2023 NAEP Automated Scoring Challenge showed that fine-tuning pre-trained encoders can reach human-like agreement, but its winning system. used a large, general-purpose encoder, while smaller math-specific pre-trained encoders remain a plausible and more efficient alternative. We fine-tune MathBERT (110M parameters) and DeBERTa-v3-large (183M parameters) as single-output regression scorers on 31 math-explanation items from one statewide summative assessment administration, benchmarking every model against human–human reliability and summarizing deployability with a three-tier acceptance criteria. On the 18 items fine-tuned under both encoders with identical datasets, the resulting performance between the two was practically indistinguishable: mean best QWK was 0.931 for MathBERT and 0.933 for DeBERTa-v3-large—a difference of only 0.002—and the models traded top performance at similar rates. Given similar performance, we tested a two-stage approach that used the larger general-purpose encoder only when the smaller math-focused encoder failed acceptance criteria. Extending MathBERT to all 31 items, it reached a mean QWK of 0.926 against a human–human benchmark of QWK = 0.940 and met operational acceptance criteria on 81% of items. Escalating the six items where a fine-tuned MathBERT model failed acceptance to a fine-tuned DeBERTa-v3-large model recovered four of the six, yielding acceptable automated-scoring models for 29 of 31 items. Because the smaller, math-pre-trained encoder matches the larger one at roughly 40% fewer parameters, we recommend fine-tuning a MathBERT encoder as the default and escalating to the larger, general purpose encoder like DeBERTa- v3-large only when a MathBERT model fails when deploying automated scoring models operationally at scale for math-explanation items.
Transformer-Aided Detection of Gaming in Constructed-Response English Language Assessments
Melanie Sharif | Scott Hellman | Martha Bellows | Sue Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Melanie Sharif | Scott Hellman | Martha Bellows | Sue Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
We compare handcrafted features, frozen transformer embeddings, and ensembles across three item types and two regimes, testing pooled versus specialist models on a gaming detection task. Ensembles perform best (ROC-AUC 0.95, 𝜅=0.75), except for rarest item type under unseen prompts. Labels reflect review detections, motivating reference-conditioned recall and blind re-review.
Multi-Agent LLM Annotation and Scoring for Training Fine-Grained K-12 Writing-Feedback Models
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Fine-grained formative writing feedback needs dense, standards-aligned labels that human annotation cannot supply at scale. We describe a multi-agent LLM pipeline producing verified silver labels, and then train small deterministic transformer scorers. On a Grade 5 pilot these reach Cohen’s 𝜅 up to 0.92 on conventions.
Hierarchical Analytic Writing Feedback with Fine-Tuned Transformers
Michael P. Hemenway | Justin O. Barber | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Michael P. Hemenway | Justin O. Barber | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
We present a production system for hierarchical analytic writing feedback that uses frontier LLMs to generate training data for fine-tuned transformer models. Applied to Grade 5 opinion essays, the system scores 5 competency dimensions and 31 binary feedback codes, achieving agreement approaching inter-rater levels from training sets of 300–500 essays.
Long-Context Transformer and State-Space Architectures for Data-Efficient Automated Essay Scoring
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
We compare long-context architectures (ModernBERT, Longformer, and causal and bidirectional Mamba) for data-efficient automated essay scoring across four training sizes on ASAP-2. Bidirectional Mamba matches the transformers. Deployment readiness depends more on label quantity (rising from 60% to 84% as training grows from 128 to 512 essays) than on backbone choice.
Seven Ways to Cut a Score: Mapping Essay-Scoring Predictions to Ordinal Ratings
Ahmed H. Bediwy | Martha Bellows | Sue Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Ahmed H. Bediwy | Martha Bellows | Sue Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
One approach to automated essay scoring re- lies on neural regression models that output continuous predicted scores, yet operational scores must be reported on a discrete, ordi- nal rubric scale to match human raters’ scores. The mapping from continuous predictions to integer scores is governed by a small set of cutpoints, and the choice of cutpoint esti- mator materially changes both accuracy and the score distribution that examinees expe- rience. We introduce, formalize, and com- pare seven cutpoint methods—simple round- ing, exact-agreement optimization, balance op- timization, quadratic-weighted-kappa (QWK) maximization, maximum-likelihood Gaussian boundaries, a Bayesian ordered-prior estimator, and an ordinal cumulative link model—under a single notation. We evaluate all seven on seven prompts from the ASAP 2.0 dataset (24,728 essays), using a shared Longformer regres- sion backbone, and report agreement (quadratic weighted kappa), association (the Pearson cor- relation between integer machine and human scores), and distributional fidelity (the stan- dardized mean difference and the maximum score-point distribution gap). Our results ex- pose a consistent accuracy–calibration trade- off: QWK maximization and balance optimiza- tion tie for the highest agreement, balance opti- mization achieves the smallest score-point dis- tribution gap, the ordinal link model achieves the smallest mean bias, the generative Gaussian and Bayesian estimators recover score means but distort the distribution, and the ordinal link model is a strong all-rounder.
Effects of Demographic Representations and Model Training Methods on Automated Scoring Engines
Yangmeng Xu | Martha Bellows | Edward W Wolfe
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Yangmeng Xu | Martha Bellows | Edward W Wolfe
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study evaluated whether automated scoring engines maintain stable performance when training data composition and training methods vary. We manipulated demographic representation (gender, English language learner, race, student with disabilities) and compared feature-based versus transformer-based models. Results showed performances were stable across subgroup-representation densities and transformer models exhibited greater stability.