James Sharpnack
Author directory2026
S2A3: Thompson Sampling and Stochastic Exposure Control for High-Stakes CATs
James Sharpnack | Alexander Tsigler | J.R. Lockwood | Steven Nydick | Alina A. von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
James Sharpnack | Alexander Tsigler | J.R. Lockwood | Steven Nydick | Alina A. von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
We introduce S2A3, a unified Bayesian framework for high-stakes computerized adaptive testing that eliminates separate item piloting. Thompson sampling routes uncertain items to informative test-takers while soft scoring attenuates their influence on ability estimates. Stochastic Sympson-Hetter exposure control ensures bank security. Validation on the Duolingo English Test confirms rapid calibration.
2024
Detecting LLM-Assisted Cheating on Open-Ended Writing Tasks on Language Proficiency Tests
Chenhao Niu | Kevin P. Yancey | Ruidong Liu | Mirza Basim Baig | André Kenji Horie | James Sharpnack
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track
Chenhao Niu | Kevin P. Yancey | Ruidong Liu | Mirza Basim Baig | André Kenji Horie | James Sharpnack
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track
The high capability of recent Large Language Models (LLMs) has led to concerns about possible misuse as cheating assistants in open-ended writing tasks in assessments. Although various detecting methods have been proposed, most of them have not been evaluated on or optimized for real-world samples from LLM-assisted cheating, where the generated text is often copy-typed imperfectly by the test-taker. In this paper, we present a framework for training LLM-generated text detectors that can effectively detect LLM-generated samples after being copy-typed. We enhance the existing transformer-based classifier training process with contrastive learning on constructed pairwise data and self-training on unlabeled data, and evaluate the improvements on a real-world dataset from the Duolingo English Test (DET), a high-stakes online English proficiency test. Our experiments demonstrate that the improved model outperforms the original transformer-based classifier and other baselines.