Siyuan (Marco) Chen

Author directory

Also published as: Siyuan Marco Chen


2026

Digital-first assessments are delivered continuously and often remotely. Artificial intelligence (AI) enables digital assessments to generate content and administer tests at scale. Like any system used for high-stakes decision-making, digital assessments require responsible AI (RAI) practices to ensure fairness and validity. This paper presents two deployed systems in the Duolingo English Test (DET) lifecycle that align the DET to its RAI Standards. The Item Factory combines automated item generation with staged expert review; the Analytics for Quality Assurance in Test Taker (AQUA-TT) system applies unsupervised anomaly detection methods to continuously monitor for issues in digital test deliveries for daily individual test sessions. We present the design and performance of these systems and discuss what they imply for placing human judgment inside digital assessments.
AI-based item generation and NLP-based prediction of item parameters are producing item banks that are substantially larger, sparser, and more frequently updated than conventional banks. Hierarchical Bayesian item response theory (IRT) is a natural calibration framework for such banks, but the common practice of refitting the entire accumulated response history at each update is costly and can exceed available memory. We describe consensus calibration, a divide-and-conquer procedure that calibrates each time period independently and reconstructs the pooled posterior in two layers. First, the posterior draws of each period are mapped to a common metric by a robust characteristic-curve linking (Haebara) that is solved separately for each draw, which propagates the uncertainty of the linking transformation into the linked posteriors. Second, the linked item posteriors are combined as a product of Gaussian densities from which the population prior contributed by each period is removed and a single prior—obtained by consensus across the per-period population posteriors—is reinstated. The correction targets the posterior dispersion, not only its location. As evidence for consensus calibration, we compare it to a pooled single-run analysis on a large operational assessment in terms of item-parameter recovery, an uncertainty-by-exposure diagnostic, and the ability distributions.