Construct Validity of Small-Sample Transformer Scoring Models: A Mechanistic Interpretability Approach

Michael P. Hemenway, Martha Bellows


Abstract
Small-sample transformer scoring models ( n = 64 n=64) reach high human agreement but risk leaning on surface shortcuts like response length. Evaluating Mechanistic Interpretability strategies across 90 models, we show correlational methods suffer from seed noise, whereas interventional erasure proves small-sample models causally depend more on length. We outline an actionable audit protocol.
Anthology ID:
2026.aimecon-wip.41
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
322–328
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.41/
DOI:
Bibkey:
Cite (ACL):
Michael P. Hemenway and Martha Bellows. 2026. Construct Validity of Small-Sample Transformer Scoring Models: A Mechanistic Interpretability Approach. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 322–328, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Construct Validity of Small-Sample Transformer Scoring Models: A Mechanistic Interpretability Approach (Hemenway & Bellows, AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.41.pdf