Matthew S. Johnson

Author directory

Also published as: Matthew S Johnson


2026

We present a scalable dual-agent architecture translating multi-source assessment data – integrating performance with process logs – into formative data insights. Decoupling classification from text generation, embedding expert rubrics, and optimizing latency enables rapid first-draft generation. These insights reveal underlying learning behaviors, facilitating targeted intervention without increasing teachers’ cognitive burden.
We investigate the problem of promoting individual fairness in automated scoring. Three models are trained and evaluated under different individual fairness constraints. The models show improved scoring consistency on the test set, but at the cost of slightly reduced accuracy and a tendency to regress scores towards the mean.

2025

This study uses multi-AI agents to accelerate teacher co-design efforts. It innovatively links student profiles obtained from numerical assessment data to AI agents in natural languages. The AI agents simulate human inquiry, enrich feedback and ground it in teachers’ knowledge and practice, showing significant potential for transforming assessment practice and research.

2020

The effect of noisy labels on the performance of NLP systems has been studied extensively for system training. In this paper, we focus on the effect that noisy labels have on system evaluation. Using automated scoring as an example, we demonstrate that the quality of human ratings used for system evaluation have a substantial impact on traditional performance metrics, making it impossible to compare system evaluations on labels with different quality. We propose that a new metric, PRMSE, developed within the educational measurement community, can help address this issue, and provide practical guidelines on using PRMSE.