Siyuan (Marco) Chen
Author directoryAlso published as: Siyuan Marco Chen
2026
Responsible AI in the Duolingo English Test: Case Studies with Automatic Item Creation and Session-Level Quality Monitoring
Siyuan Marco Chen | Xiaowan Zhang | Andrew Runge | Jacqueline Church
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Siyuan Marco Chen | Xiaowan Zhang | Andrew Runge | Jacqueline Church
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Digital-first assessments are delivered continuously and often remotely. Artificial intelligence (AI) enables digital assessments to generate content and administer tests at scale. Like any system used for high-stakes decision-making, digital assessments require responsible AI (RAI) practices to ensure fairness and validity. This paper presents two deployed systems in the Duolingo English Test (DET) lifecycle that align the DET to its RAI Standards. The Item Factory combines automated item generation with staged expert review; the Analytics for Quality Assurance in Test Taker (AQUA-TT) system applies unsupervised anomaly detection methods to continuously monitor for issues in digital test deliveries for daily individual test sessions. We present the design and performance of these systems and discuss what they imply for placing human judgment inside digital assessments.
Bayesian Consensus Calibration of Continuously Evolving IRT Item Banks
Paul A Jewsbury | Steven W Nydick | Manqian Liao | Siyuan (Marco) Chen
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Paul A Jewsbury | Steven W Nydick | Manqian Liao | Siyuan (Marco) Chen
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
AI-based item generation and NLP-based prediction of item parameters are producing item banks that are substantially larger, sparser, and more frequently updated than conventional banks. Hierarchical Bayesian item response theory (IRT) is a natural calibration framework for such banks, but the common practice of refitting the entire accumulated response history at each update is costly and can exceed available memory. We describe consensus calibration, a divide-and-conquer procedure that calibrates each time period independently and reconstructs the pooled posterior in two layers. First, the posterior draws of each period are mapped to a common metric by a robust characteristic-curve linking (Haebara) that is solved separately for each draw, which propagates the uncertainty of the linking transformation into the linked posteriors. Second, the linked item posteriors are combined as a product of Gaussian densities from which the population prior contributed by each period is removed and a single prior—obtained by consensus across the per-period population posteriors—is reinstated. The correction targets the posterior dispersion, not only its location. As evidence for consensus calibration, we compare it to a pooled single-run analysis on a large operational assessment in terms of item-parameter recovery, an uncertainty-by-exposure diagnostic, and the ability distributions.