Justin O. Barber
Author directoryAlso published as: Justin O Barber
2026
Multi-Agent LLM Annotation and Scoring for Training Fine-Grained K-12 Writing-Feedback Models
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Fine-grained formative writing feedback needs dense, standards-aligned labels that human annotation cannot supply at scale. We describe a multi-agent LLM pipeline producing verified silver labels, and then train small deterministic transformer scorers. On a Grade 5 pilot these reach Cohen’s 𝜅 up to 0.92 on conventions.
Hierarchical Analytic Writing Feedback with Fine-Tuned Transformers
Michael P. Hemenway | Justin O. Barber | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Michael P. Hemenway | Justin O. Barber | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
We present a production system for hierarchical analytic writing feedback that uses frontier LLMs to generate training data for fine-tuned transformer models. Applied to Grade 5 opinion essays, the system scores 5 competency dimensions and 31 binary feedback codes, achieving agreement approaching inter-rater levels from training sets of 300–500 essays.
Long-Context Transformer and State-Space Architectures for Data-Efficient Automated Essay Scoring
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
We compare long-context architectures (ModernBERT, Longformer, and causal and bidirectional Mamba) for data-efficient automated essay scoring across four training sizes on ASAP-2. Bidirectional Mamba matches the transformers. Deployment readiness depends more on label quantity (rising from 60% to 84% as training grows from 128 to 512 essays) than on backbone choice.
2025
When Does Active Learning Actually Help? Empirical Insights with Transformer-based Automated Scoring
Justin O Barber | Michael P. Hemenway | Edward Wolfe
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Justin O Barber | Michael P. Hemenway | Edward Wolfe
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Developing automated essay scoring (AES) systems typically demands extensive human annotation, incurring significant costs and requiring considerable time. Active learning (AL) methods aim to alleviate this challenge by strategically selecting the most informative essays for scoring, thereby potentially reducing annotation requirements without compromising model accuracy. This study systematically evaluates four prominent AL strategies—uncertainty sampling, BatchBALD, BADGE, and a novel GenAI-based uncertainty approach—against a random sampling baseline, using DeBERTa-based regression models across multiple assessment prompts exhibiting varying degrees of human scorer agreement. Contrary to initial expectations, we found that AL methods provided modest but meaningful improvements only for prompts characterized by poor scorer reliability (<60% agreement per score point). Notably, extensive hyperparameter optimization alone substantially reduced the annotation budget required to achieve near-optimal scoring performance, even with random sampling. Our findings underscore that while targeted AL methods can be beneficial in contexts of low scorer reliability, rigorous hyperparameter tuning remains a foundational and highly effective strategy for minimizing annotation costs in AES system development.