Justin O. Barber

Author directory

Also published as: Justin O Barber


2026

Fine-grained formative writing feedback needs dense, standards-aligned labels that human annotation cannot supply at scale. We describe a multi-agent LLM pipeline producing verified silver labels, and then train small deterministic transformer scorers. On a Grade 5 pilot these reach Cohen’s 𝜅 up to 0.92 on conventions.
We present a production system for hierarchical analytic writing feedback that uses frontier LLMs to generate training data for fine-tuned transformer models. Applied to Grade 5 opinion essays, the system scores 5 competency dimensions and 31 binary feedback codes, achieving agreement approaching inter-rater levels from training sets of 300–500 essays.
We compare long-context architectures (ModernBERT, Longformer, and causal and bidirectional Mamba) for data-efficient automated essay scoring across four training sizes on ASAP-2. Bidirectional Mamba matches the transformers. Deployment readiness depends more on label quantity (rising from 60% to 84% as training grows from 128 to 512 essays) than on backbone choice.

2025

Developing automated essay scoring (AES) systems typically demands extensive human annotation, incurring significant costs and requiring considerable time. Active learning (AL) methods aim to alleviate this challenge by strategically selecting the most informative essays for scoring, thereby potentially reducing annotation requirements without compromising model accuracy. This study systematically evaluates four prominent AL strategies—uncertainty sampling, BatchBALD, BADGE, and a novel GenAI-based uncertainty approach—against a random sampling baseline, using DeBERTa-based regression models across multiple assessment prompts exhibiting varying degrees of human scorer agreement. Contrary to initial expectations, we found that AL methods provided modest but meaningful improvements only for prompts characterized by poor scorer reliability (<60% agreement per score point). Notably, extensive hyperparameter optimization alone substantially reduced the annotation budget required to achieve near-optimal scoring performance, even with random sampling. Our findings underscore that while targeted AL methods can be beneficial in contexts of low scorer reliability, rigorous hyperparameter tuning remains a foundational and highly effective strategy for minimizing annotation costs in AES system development.