Rethinking Validity in Educational Assessment when AI Co-Produces Performance

Xiaoran Li, Tanesia Beverly


Abstract
Generative AI complicates a core assumption of assessment that observed performance reflects an individual’s own cognition. Using StudyChat, we show AI supply only moderately tracks student intent, and assignment scores are largely insensitive to either. With the disruption of validity warrant, response process validity needs reconceptualizing when AI co-produces performance.
Anthology ID:
2026.aimecon-sessions.28
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
260–266
Language:
URL:
https://aclanthology.org/2026.aimecon-sessions.28/
DOI:
Bibkey:
Cite (ACL):
Xiaoran Li and Tanesia Beverly. 2026. Rethinking Validity in Educational Assessment when AI Co-Produces Performance. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 260–266, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Rethinking Validity in Educational Assessment when AI Co-Produces Performance (Li & Beverly, AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-sessions.28.pdf