Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish (Editors)
- Anthology ID:
- 2026.aimecon-sessions
- Month:
- October
- Year:
- 2026
- Address:
- Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
- Venue:
- AIME-Con
- Event:
- Artificial Intelligence in Measurement and Education Conference (AIME-Con) (2026)
- SIG:
- Publisher:
- National Council on Measurement in Education (NCME)
- URL:
- https://aclanthology.org/2026.aimecon-sessions/
- DOI:
- ISBN:
- 979-8-9983004-2-4
- PDF:
- https://aclanthology.org/2026.aimecon-sessions.pdf
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Multi-Agent LLM Annotation and Scoring for Training Fine-Grained K-12 Writing-Feedback Models
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Justin O Barber | Michael P Hemenway | Martha Bellows | Susan Lottridge
Fine-grained formative writing feedback needs dense, standards-aligned labels that human annotation cannot supply at scale. We describe a multi-agent LLM pipeline producing verified silver labels, and then train small deterministic transformer scorers. On a Grade 5 pilot these reach Cohen’s 𝜅 up to 0.92 on conventions.
LLM-Based Rubric Refinement in Multi-Agent Automated Scoring
Alexander Kwako | Cristina Everett | Harry Wang
Alexander Kwako | Cristina Everett | Harry Wang
In this study, we marry automated scoring with automated rubric revision in an iterative, mutually-reinforcing process. One LLM agent scores student responses; a second, the Error Analysis Agent, examines human–engine scoring discrepancies and suggests rubric revisions. We show that this positive feedback loop improves automated scoring performance.
Beyond Agreement: Operational Monitoring of K–12 Automated Scoring
Jing Ma | Edward W Wolfe | Anthony D Fina
Jing Ma | Edward W Wolfe | Anthony D Fina
This study compares automated and human scores with expert/backread scores for writing, reading, and science tasks in a K-12 assessment program. Analyses examine discrepancies by score level and reused prompts across administrations. Automated scores showed smaller discrepancies for writing but larger discrepancies for short-answer tasks; reused prompts revealed localized shifts.
Detecting Invalid Responses in Automated Essay Scoring with Fine-Tuned Large Language Models
YoungKoung Kim | Christopher Ormerod
YoungKoung Kim | Christopher Ormerod
This study compares generative models Gemma and Qwen with ModernBERT and Mahalanobis-LOF for invalid-response detection in automated essay scoring. Qwen maintained high recall while reducing false invalid flags on a representative test set and detected some fluent off-topic responses in a matched diagnostic set. Fluent off-topic detection remained difficult.
Compact distillation preserved clean essay-scoring agreement across three student models and six compressed teacher targets, but the best target differed by model and margins were small. Perturbation, distribution-shift, and subgroup-bootstrap analyses dominated those differences, indicating that resource-constrained educational scoring requires uncertainty-aware evaluation rather than clean held-out comparison alone.
From Efficiency to Sufficiency: A Measurement-Informed Framework for Cost-Aware Automated Essay Scoring
Yi Gui
Yi Gui
Compute cost in automated essay scoring is usually treated as engineering rather than measurement. This paper introduces the Cost-Aware Measurement Design (CAMD) framework, defines cost per defensible score (CPDS), and argues that AES systems should be compared by quality-cost sufficiency at a declared scoring volume.
Generalizability Theory for Evaluating Fine-Tuned LLMs in Automated Item Generation
Zhifei Li | Won-Chan Lee
Zhifei Li | Won-Chan Lee
Automated item generation with large language models (LLMs) is typically evaluated using aggregate accuracy metrics that conflate the quality of the source passage with noise in- troduced by the generation process itself. We address this gap by embedding a prospective i : (p×s×t) Generalizability Theory (G- theory) design into the evaluation of a QLoRA fine-tuned Qwen2.5-7B-Instruct model on the SciQ corpus. Passages (p) serve as the object of measurement; random seeds (s) and prompt templates (t) are fully crossed facets; items are nested within each (p,s,t) cell. Across 9,000 observations and five binary quality met- rics, we find that 77–81% of total variance is attributable to the passage, seed and template main effects are negligible (≤0.02%), and G- coefficients (E 𝜌2) uniformly exceed 0.97 un- der the observed design. D-study projections show that a single seed and template already achieves E 𝜌2 = 0.87, while the observed de- sign (ns = 3, nt = 3, ni = 2) reaches 0.98. Code, data, and R analysis scripts are released to support reproducible psychometric evalua- tion of future item-generation systems.
Evaluating Score Dependability in ChatGPT-Supported AP Chinese Speaking Tasks
Dan Song | Won-Chan Lee
Dan Song | Won-Chan Lee
This study applied generalizability theory to examine score variation and dependability in AP Chinese speaking tasks completed with and without ChatGPT support. Although ChatGPT-supported tasks were associated with higher scores, the NoGPT condition consistently exhibited higher dependability coefficients. Increasing the numbers of tasks and raters further improved score dependability.
From Validation to Resilience: Sustaining Measurement Quality in AI-Assisted Enemy Item Identification
Ye Ma
Ye Ma
Enemy items are item pairs that must not appear on the same test form. An automatic enemy-identification method using large language models (LLMs) has been deployed in operation. This study asks a question: how is measurement quality sustained as conditions shift after deployment? Using operational data from certification exams, two analyses examine the factors that affect the method’s resilience. Analysis 1 isolates model-version and prompt updates: changing the model with the prompt held constant reduced recall from 0.75 to 0.54, while prompt refinement with the model held constant recovered it to 0.82. It also shows that most model–reviewer disagreements are edge cases and that the human standard is itself variable, with four experts spanning 0.70 to 0.91 in recall. Analysis 2 reports a disruption case: the established method works effectively on a professional-level exam but not on a foundational-level exam. A complementary content-tag based method was added to the existing method in response, with human review surfacing the disruption and validating the fix. Responsible LLM deployment in assessment requires monitoring with labeled data, testing before deployment, and evaluation systems tied to each use case, so that validity, reliability, and fairness are sustained rather than certified once.
Semantic Variability of LLM-Generated Replies Across LLMs: Implications for Designing Conversation-Based Assessment
Jiangang Hao
Jiangang Hao
This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.
Resilience Infrastructure for Conversational AI-Based Educational Applications
Andrew Emerson | Keelan Evanini | Kevin Frome | Le An Ha | Peter Baldwin
Andrew Emerson | Keelan Evanini | Kevin Frome | Le An Ha | Peter Baldwin
Educational applications using conversational AI built on LLMs are hard to evaluate reliably because outputs vary unpredictably each turn, posing validity, reliability, fairness, and safety risks. We present a four-stage framework that combines automated and human testing to determine operational preparedness and continuously monitor deployed systems. The application of this resilience infrastructure framework is demonstrated in a case study of a conversation-based formative assessment tool for medical students to practice doctor-patient communication skills.
Responsible AI in the Duolingo English Test: Case Studies with Automatic Item Creation and Session-Level Quality Monitoring
Siyuan Marco Chen | Xiaowan Zhang | Andrew Runge | Jacqueline Church
Siyuan Marco Chen | Xiaowan Zhang | Andrew Runge | Jacqueline Church
Digital-first assessments are delivered continuously and often remotely. Artificial intelligence (AI) enables digital assessments to generate content and administer tests at scale. Like any system used for high-stakes decision-making, digital assessments require responsible AI (RAI) practices to ensure fairness and validity. This paper presents two deployed systems in the Duolingo English Test (DET) lifecycle that align the DET to its RAI Standards. The Item Factory combines automated item generation with staged expert review; the Analytics for Quality Assurance in Test Taker (AQUA-TT) system applies unsupervised anomaly detection methods to continuously monitor for issues in digital test deliveries for daily individual test sessions. We present the design and performance of these systems and discuss what they imply for placing human judgment inside digital assessments.
Developing and evaluating an AI agent to read and interpret learning maps
Dante Cisterna | Jonathan Schuster | Amber Samson
Dante Cisterna | Jonathan Schuster | Amber Samson
We describe the development and evaluation of an AI conversational agent designed to interpret learning maps. Results indicate the agent can accurately identify map components and structures, especially when different informational formats are combined. Findings highlight the tool’s feasibility to inform the development of tools that use and interpret learning maps for teachers.
Learning Progression-Guided Scientific Model Assessment Using Foundation and Supervised Vision Models
Leonora Kaldaras | Yasasvi Nalla
Leonora Kaldaras | Yasasvi Nalla
Artificial intelligence can scale assessment of student-generated scientific models, but valid educational use requires algorithms to identify features that meaningfully represent the intended learning construct. This study compares foundation and supervised computer vision approaches for evaluating learning progression (LP)-aligned evidence in approximately 1,600 high-school students’ electroscope models. Grounding DINO combined with the Segment Anything Model (SAM) was used for zero-shot detection without task-specific labeled training data, whereas a CustomCharge convolutional neural network (CNN) was trained on human-scored models. Both approaches were evaluated against expert scoring across 13 LP-aligned analytic categories. Grounding DINO+SAM achieved higher Cohen’s 𝜅 in 12 of 13 categories, with strong performance across both charge- and force-related evidence. Categories involving less frequent or more complex cross-scenario representations remained comparatively challenging. Findings demonstrate complementary strengths of foundation and supervised approaches and illustrate how LPs can provide a theoretically grounded framework for developing and validating AI assessment of scientific models.
Developing an LLM Tutor Quality Evaluation Scale
Michael Xiao | Zewei Tian | Alex Liu | Lief Esbenshade | Min Sun
Michael Xiao | Zewei Tian | Alex Liu | Lief Esbenshade | Min Sun
We introduce a unified metric set for evaluating LLM tutors by inductively coding 177 recent literature-based metrics and scoring tutoring dialogues with three LLM judges. Exploratory factor analysis identified five constructs and a strong general factor. We present the resulting 10-category, 41-item framework for confirmatory analysis and human validation.
When Language Becomes a Shortcut in Automated Short-Answer Scoring
Xiaomeng Xiong | Corinne Huggins-Manley | Jinnie Shin | Christan Grant
Xiaomeng Xiong | Corinne Huggins-Manley | Jinnie Shin | Christan Grant
Automated short-answer scoring should reflect substantive content rather than linguistic form. Using controlled rewrites of 743 ASAP-SAS responses, we compare six fine-tuned encoders and 11 LLMs. Encoder scores increased with linguistic complexity, especially for lower-scoring responses, whereas LLMs showed heterogeneous patterns, revealing model-specific construct-irrelevant scoring signals.
Disentangling Severity and Centrality: Rater Effects in LLM-Based Automated Short-Answer Scoring
Xiaomeng Xiong | Corinne Huggins-Manley | Jinnie Shin
Xiaomeng Xiong | Corinne Huggins-Manley | Jinnie Shin
Using the Facets Model for Severity and Centrality, we analyzed two human raters and 10 LLMs across four ASAP-SAS prompts. LLM centrality was task-dependent, and few-shot prompting reduced centrality inconsistently. Omitting scale-use differences altered severity estimates (r=.47), while MFRM fit diagnostics were harder to interpret in heterogeneous LLM rater pools.
Expert approval certifies that test developers would use AI-generated items, not that the scores support valid interpretations. I extend argument-based validity with a generation inference and a five-layer evidence framework, apply it diagnostically to three NAEP–Wilbur evaluations, and propose design principles and a minimal reporting standard for the field.
This study evaluated use of an AI-enabled item-generation tool for Reading item development for the National Assessment of Educational Progress (NAEP). Results were strongest when items targeted textually explicit information and weakest when passage and construct complexity increased. The tool showed potential for early-stage drafting rather than autonomous item development.
Evaluating an AI-Item Generation Tool for NAEP Mathematics
Cheryl Van Ness | Ilona Minchuk | Karen Parker
Cheryl Van Ness | Ilona Minchuk | Karen Parker
This study evaluates the effectiveness of an AI-based system (Wilbur) for generating NAEP mathematics items. Results from this investigation indicate moderate success for content-focused items but substantial challenges for items intended to assess the NAEP Mathematical Practices. Transitioning from GPT-4 to GPT-5.2 improved item quality, though human revision remained essential.
Evaluating an AI-Item Generation Tool for NAEP Science
Carolina Safar | Mitchell Price | Jeffery Ackley
Carolina Safar | Mitchell Price | Jeffery Ackley
AI-enabled systems could transform item authoring for standardized assessments. We compare framework alignment and accuracy of NAEP Science items developed by two AI tools and human authors. Items from all sources had similar acceptance rates and most accepted items would require significant revision, underscoring the importance of evaluating AI-generated content.
Evidence-Centered Design for AI-Driven Automated Item Generation in Applied Mathematics
Edith Aurora Graf | Maria Elena Oliveri | Giulia Oliveri | Emily Proctor
Edith Aurora Graf | Maria Elena Oliveri | Giulia Oliveri | Emily Proctor
We discuss how human-led Evidence-Centered Design (ECD) can inform the development of AI-supported automated item generation (AIG), as well as how it can be used as a quality-control mechanism. In this particular application, we describe how we collaborated with Claude Sonnet 5 to produce kinematics items together with interactive tools through an AIG pipeline. A preliminary qualitative analysis of a small number of generated items suggested they have identifiable strengths and weaknesses, which we discuss in depth per item.
Human–AI Collaboration in Educational Measurement: Transforming Assessment Data into Educational Action
Laura C. Egan | Judy H. Tang
Laura C. Egan | Judy H. Tang
This paper describes a framework for human–AI collaboration in educational measurement that connects four workflow components. The framework integrates educational measurement and responsible AI principles, emphasizing AI tools as support for human expertise. Illustrative applications demonstrate how human–AI collaboration can strengthen assessment processes while maintaining measurement quality and integrity.
Scaling Human-AI Collaboration: Translating Multi-Source Assessment Data into Formative Learner Profiles
Hongwen Guo | Matthew S Johnson | Luis Saldivia | Michelle Worthington | Jeremy Lee | Kadriye Ercikan
Hongwen Guo | Matthew S Johnson | Luis Saldivia | Michelle Worthington | Jeremy Lee | Kadriye Ercikan
We present a scalable dual-agent architecture translating multi-source assessment data – integrating performance with process logs – into formative data insights. Decoupling classification from text generation, embedding expert rubrics, and optimizing latency enables rapid first-draft generation. These insights reveal underlying learning behaviors, facilitating targeted intervention without increasing teachers’ cognitive burden.
Human–Machine Learning for Large-Scale Qualitative Coding in Educational Measurement
Judy H. Tang | Tom Krenzke | Jin Hui Xu | Karen Lo
Judy H. Tang | Tom Krenzke | Jin Hui Xu | Karen Lo
This paper examines a human–ML framework for large-scale qualitative data coding. Using a nationally representative sample of transcript data, the study evaluates semantic embeddings and ranked recommendations through validation and user testing. Results demonstrate improved efficiency and accuracy while maintaining human expertise, oversight, and responsibility for final coding decisions.
Rubric-Aligned Generative-AI Features as Supplementary Predictors in Feature-Based Automated Essay Scoring
Yue Huang | Duanli Yan | Corey Palermo
Yue Huang | Duanli Yan | Corey Palermo
This study examined whether rubric-aligned generative-AI features could augment established linguistic features in trait-based automated essay scoring. Features from both sources showed meaningful associations with human scores and only partial overlap with one another. Scoring models combining both feature sets produced modest improvements that varied across traits and evaluation metrics.
Using PERSUADE 2.0 discourse segments, this study examines whether effectiveness labels carry stable linguistic meaning across discourse types. Traditional NLP features and latent components show robust type-by-effectiveness interactions. Length decomposition indicates much of this signal reflects elaboration, while smaller length-robust patterns remain, cautioning against context-free diagnostic feedback in automated writing evaluation.
Rethinking Validity in Educational Assessment when AI Co-Produces Performance
Xiaoran Li | Tanesia Beverly
Xiaoran Li | Tanesia Beverly
Generative AI complicates a core assumption of assessment that observed performance reflects an individual’s own cognition. Using StudyChat, we show AI supply only moderately tracks student intent, and assignment scores are largely insensitive to either. With the disruption of validity warrant, response process validity needs reconceptualizing when AI co-produces performance.
Programmatic Tiered Prompting for LLM Generation of ELA & Math Items
David Whitecomb | Marjorie Wine | Alexander Hoffman
David Whitecomb | Marjorie Wine | Alexander Hoffman
Using the LBIDAT protocol, we evaluated items generated by three LLMs (Claude, Gemini, GPT) across six zero-shot prompting tiers for 8th-grade standards. Neither prompt tier nor model affected defect severity. ELA items were consistently poor; mathematics items passed the low bar while falling short of appropriate grade-level cognitive complexity.
The Scoring Paradox: Multi-Agent Architectures for Unbiased Feedback Optimization
Okan Bulut | Bin Tan | Elisabetta Mazzullo | Cole Walsh
Okan Bulut | Bin Tan | Elisabetta Mazzullo | Cole Walsh
Does grade awareness bias automatically generated feedback? We present a multi-agent artificial intelligence (AI) framework that contrasts score-blind and score-informed evaluators and reconciles their critiques. Across 908 feedback blocks, score exposure systematically biased the model’s critique: harsher at low grades, warmer at high (halo/horn effects). Isolating a score-blind judgment improved score-alignment without sacrificing faithfulness.
As LLMs increasingly streamline item generation, ensuring the quality of their outputs remains a critical challenge. This study examines whether a multi-agent system can serve as an automated judge to detect and revise flaws in AI-generated educational assessment items, such as ambiguity, bias, or content misalignment.
Automated Evaluation of Mathematical Equivalence Between Personalized and Standard Word Problems
Burcu Arslan | Ikkyu Choi | Jesse R. Sparks | Reginald M. Gooch | Candace Walkington | Matthew L. Bernacki
Burcu Arslan | Ikkyu Choi | Jesse R. Sparks | Reginald M. Gooch | Candace Walkington | Matthew L. Bernacki
Generative AI enables real-time and scalable context personalization of mathematics word problems (MWPs) based on students’ self-reported interests during assessment. However, a question arises: are personalized and standard MWPs mathematically equivalent? In this paper, we present two Natural Language Processing pipelines for evaluating mathematical equivalence between personalized and standard MWPs.
This study integrates assessment data with learning maps, using a neural network to recommend individualized next skills for instruction and assessment. Results indicate high-certainty, expert-validated recommendations, with empirical support for foundational assumptions. This demonstrates the potential for AI to support more personalized learning-maps-based instruction and assessment.
Score Report Conversational Agent to Guide Educators in Interpreting Assessment Results
Brian D. Gane | Dante I. Cisterna | Amy K. Clark
Brian D. Gane | Dante I. Cisterna | Amy K. Clark
We developed and evaluated a prototype conversational agent for teachers that can support their understanding and use of mastery-based score reports that describe a students’ academic performance. The agent prototype delivered grounded responses reflecting intended interpretation and uses, however some responses were inaccurate, signaling aspects that can be improved.
Automated Generation and Scoring of Maze Reading Comprehension Assessments
Hatice Kubra Karakis | Walter Leite | Logan Scott | Xinyi Tai | Akihito Kamata
Hatice Kubra Karakis | Walter Leite | Logan Scott | Xinyi Tai | Akihito Kamata
This work proposes automated generation and psychometric scoring of Maze comprehension assessments using large language models (LLMs) and a multilevel item response theory (multilevel IRT) framework. Findings show reliable ability estimates across passages and items, offering scalable, curriculum-aligned formative assessment that reduces teacher workload and supports targeted reading instruction.