Edward W Wolfe

Author directory

2026

This study compares automated and human scores with expert/backread scores for writing, reading, and science tasks in a K-12 assessment program. Analyses examine discrepancies by score level and reused prompts across administrations. Automated scores showed smaller discrepancies for writing but larger discrepancies for short-answer tasks; reused prompts revealed localized shifts.
This study evaluated whether automated scoring engines maintain stable performance when training data composition and training methods vary. We manipulated demographic representation (gender, English language learner, race, student with disabilities) and compared feature-based versus transformer-based models. Results showed performances were stable across subgroup-representation densities and transformer models exhibited greater stability.