Effects of Demographic Representations and Model Training Methods on Automated Scoring Engines

Yangmeng Xu, Martha Bellows, Edward W Wolfe


Abstract
This study evaluated whether automated scoring engines maintain stable performance when training data composition and training methods vary. We manipulated demographic representation (gender, English language learner, race, student with disabilities) and compared feature-based versus transformer-based models. Results showed performances were stable across subgroup-representation densities and transformer models exhibited greater stability.
Anthology ID:
2026.aimecon-main.16
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
154–160
Language:
URL:
https://aclanthology.org/2026.aimecon-main.16/
DOI:
Bibkey:
Cite (ACL):
Yangmeng Xu, Martha Bellows, and Edward W Wolfe. 2026. Effects of Demographic Representations and Model Training Methods on Automated Scoring Engines. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 154–160, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Effects of Demographic Representations and Model Training Methods on Automated Scoring Engines (Xu et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.16.pdf