Corinne Huggins-Manley

Author directory

2026

Using the Facets Model for Severity and Centrality, we analyzed two human raters and 10 LLMs across four ASAP-SAS prompts. LLM centrality was task-dependent, and few-shot prompting reduced centrality inconsistently. Omitting scale-use differences altered severity estimates (r=.47), while MFRM fit diagnostics were harder to interpret in heterogeneous LLM rater pools.
Automated short-answer scoring should reflect substantive content rather than linguistic form. Using controlled rewrites of 743 ASAP-SAS responses, we compare six fine-tuned encoders and 11 LLMs. Encoder scores increased with linguistic complexity, especially for lower-scoring responses, whereas LLMs showed heterogeneous patterns, revealing model-specific construct-irrelevant scoring signals.