Corey Palermo

Author directory

2026

This study examined whether rubric-aligned generative-AI features could augment established linguistic features in trait-based automated essay scoring. Features from both sources showed meaningful associations with human scores and only partial overlap with one another. Scoring models combining both feature sets produced modest improvements that varied across traits and evaluation metrics.

2025

In hybrid scoring systems, confidence thresholds determine which responses receive human review. This study evaluates a relative (within-batch) thresholding method against an absolute benchmark across ten items. Results show near-perfect agreement and modest distributional differences, supporting the relative method’s validity as a scalable, operationally viable approach for flagging low-confidence responses.