Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish (Editors)
- Anthology ID:
- 2026.aimecon-wip
- Month:
- October
- Year:
- 2026
- Address:
- Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
- Venue:
- AIME-Con
- Event:
- Artificial Intelligence in Measurement and Education Conference (AIME-Con) (2026)
- SIG:
- Publisher:
- National Council on Measurement in Education (NCME)
- URL:
- https://aclanthology.org/2026.aimecon-wip/
- DOI:
- ISBN:
- 979-8-9983004-1-7
- PDF:
- https://aclanthology.org/2026.aimecon-wip.pdf
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
A Cognitive Lab Study of Student AI Chatbot Use in Online Coursework
Jinah Choi | Sonya Powers | Farzan Karimi-Malekabadi | Michelle Barrett
Jinah Choi | Sonya Powers | Farzan Karimi-Malekabadi | Michelle Barrett
This cognitive lab study examined how secondary students in Grades 7-12 interacted with an AI-enabled chatbot while completing online course activities. Using chat logs, facilitator observations, post-session reflections, and survey responses, we analyzed spontaneous and lightly seeded interactions to evaluate whether the chatbot supported its intended theory of action: helping students continue learning without simply completing work for them. Findings highlight the importance of typing fluency, answer-boundary handling, safety and privacy response patterns, and student-facing feedback mechanisms. This study illustrates how cognitive lab evidence can inform ongoing validation.
This work-in-progress study examines whether sentence embeddings can flag item pairs at risk for local dependence in multidimensional IRT. Using a 50-item personality inventory, preliminary analysis link semantic similarity to residual item-pair dependence. Future work will test robustness across larger samples, additional datasets, and alternative diagnostics.
Evaluating Multi-Phase and Granular Strategies for Evidence-Anchored Automated Scoring
Alessia Marigo | Laura Wright | Linda Malkin | Xin Xie
Alessia Marigo | Laura Wright | Linda Malkin | Xin Xie
This study evaluates few-shot large language models (LLMs) on middle-school geoscience responses (N=86 responses × 11 indicators), separating presence agreement from extract agreement. Joint prompting and annotation-like rubric guidance plus examples yield the clearest gains; lengthy rubric rewriting and clause-level parsing do not reliably improve extract alignment.
Consensus without Accuracy: Investigating LLM’s Recovery of Item Difficulty Using Paired Comparisons
Michael Leon Chrzan | Benjamin Domingue
Michael Leon Chrzan | Benjamin Domingue
This study evaluates whether large language models (LLMs) can recover item difficulty estimates through Bradley-Terry modeling from pairwise comparisons of items. Across five Item Response Warehouse datasets and four LLMs, we examine alignment between LLM pairwise-derived difficulty rankings and 1PL IRT parameters, with implications for scalable, AI-assisted item calibration.
Implementing Generative AI Assessment Guidelines in a Mexican University
Melchor Sánchez-Mendiola | Enrique Buzo-Casanova | Elibidú Ortega-Sánchez | Manuel García-Minjares | Adriana Cruz Juárez
Melchor Sánchez-Mendiola | Enrique Buzo-Casanova | Elibidú Ortega-Sánchez | Manuel García-Minjares | Adriana Cruz Juárez
Case study that examines a large public Mexican university development of GenAI assessment guidelines for upper-secondary, undergraduate, and graduate education. Using document analysis and process tracing, it identifies shared governance principles, level-specific adaptations, and a policy-to-practice pathway emphasizing validity, fairness, transparency, human oversight, responsible AI use, and implementation evaluation.
TRACE: Automated Multi-Granularity Analysis of Text Revision
Yu Tian | Katerina Christhilf | Andrew Potter | Jessica Early | Steve Graham | Danielle McNamara
Yu Tian | Katerina Christhilf | Andrew Potter | Jessica Early | Steve Graham | Danielle McNamara
This paper introduces TRACE, an automated framework for analyzing textual revisions across drafts. TRACE aligns initial and revised texts and infers interpretable revision operations, such as insert, delete, replace, and move. TRACE supports AI-powered revision analytics, writing assessment, automated feedback, and studies of human–AI writing collaboration.
We examine how strategies for representing student skill (𝜃) affect IRT-3PL item-parameter recovery from reconstructed item characteristic curves. We compare natural-language descriptors with signed decimal and scientific-notation anchors across anchor counts and placements. Our results show that dense numeric grids improve recovery, while non-uniform placements produce parameter-specific trade-offs, showing that 𝜃 representation matters.
Assessing Large Language Model Performance in Post-Certification-Examination Comment Categorization
Huaping Sun | Kristin O’Brien | Colleen Burke Kave | David Shin | Qiao Lin | Jeffrey Marc Lyness
Huaping Sun | Kristin O’Brien | Colleen Burke Kave | David Shin | Qiao Lin | Jeffrey Marc Lyness
This study evaluated an LLM for analyzing 1,406 Neurocritical Care examination comments coded for sentiment and thematic categories. Human raters showed high agreement, whereas LLM-human agreement was moderate. Thematic definitions reduced performance, but human-coded examples improved accuracy. LLMs may support preliminary coding, although human review remains necessary.
Rethinking Pilot Data: Evaluating LLM Synthetic Data for Scale Development
Margarita Olivera-Aguilar | Samuel H. Rikoon | Paul D. Bailey | Michael B. Kruse | Kamal Middlebrook
Margarita Olivera-Aguilar | Samuel H. Rikoon | Paul D. Bailey | Michael B. Kruse | Kamal Middlebrook
This study evaluated whether LLM- synthetic data can support K-12 instrument development. We generated LLM-synthetic datasets under various prompt conditions for two surveys and one assessment. Our findings indicate that while some prompt conditions successfully reproduced the overall latent structure of the instruments, recovery of item-level parameters was generally inadequate.
Hierarchical Attention Network Architecture for Automated Response Time Prediction
Wei Shuang Schneider
Wei Shuang Schneider
This paper presents a novel deep learning architecture for item response time prediction to create human readable feedback for content creators in educational assessment. The model employs a two-tier attention mechanism, mirroring the cognitive hierarchy of reading: the first tier identifies influential words nested within their sentence context, while the second tier evaluates the contribution of individual sentences to the overall item block. By leveraging this specific hierarchical structure, the model improves explainable AI (xAI) in educational measurement AI research. While architecturally simpler models can achieve comparable or superior raw predictive accuracy, the two-tier structure is designed specifically for xAI — to yield transparent, sentence- and world-level attribution that item developers can act on directly. This methodology bridges the gap between opaque, black-box predictive modeling and actionable, human interpretable feedback for instructional design and item development.
Beyond Agreement: Calibrating Automated Scoring to Human Judgment Using Control Scripts
Mark Dulhunty | Alex Codoreanu | Nathan Zoanetti
Mark Dulhunty | Alex Codoreanu | Nathan Zoanetti
An ensemble automated scoring approach was evaluated against 260,000 responses spanning 55 constructed-response reading items. Performance was then compared to 27 human markers using 477 control scripts. The ensemble achieved perfect agreement on control scripts for 39 items (71%), matching or exceeding human marker performance and supporting operational quality assurance.
Human-in-the-Loop Gemini Item Generation for Large-Scale MCQ Banks
Melchor Sánchez-Mendiola | Milton García-Lima | Elibidú Ortega-Sánchez | Manuel García-Minjares | Enrique Buzo-Casanova
Melchor Sánchez-Mendiola | Milton García-Lima | Elibidú Ortega-Sánchez | Manuel García-Minjares | Enrique Buzo-Casanova
This work-in-progress compares 2,405 Spanish-language MCQs drafted by customized Gemini Gems or faculty in a Mexican university. Expert committees reviewed all items. Gemini drafts showed lower validation-friction in three of four areas and stronger blueprint adherence, while faculty drafts offered richer contextualization. Findings support structured AI generation with human validation.
This study considers measures commonly used to validate AES scoring and their limitations for indicating comparability with human rater scoring. In data simulated to reflect validation results achieved by AES national competition winners, scoring standard differences can occur across AES and human rater scoring (especially for 4-point scales vs. 2- and 3-point scales). Equipercentile methods are described and recommended for resolving the scoring differences.
Finding Evidence of Communication Behaviors using LLMs
Tazin Afrin | Saed Rezayi | Janet Mee | Polina Harik | Le An Ha
Tazin Afrin | Saed Rezayi | Janet Mee | Polina Harik | Le An Ha
This study compares two automated scoring approaches to identify communication behaviors in physician responses to patient questions: prompt-based scoring with Large Language Models (LLM) and supervised transformer-based models. Our results show that although transformer-based models provide more consistent performance, competitive LLM performance is promising in reducing the need for extensive annotated training data.
Designing and Evaluating a Measurement-Oriented RAG Framework for ELL Writing Assessment
Daeryong Seo
Daeryong Seo
This design-focused study compared measurement-oriented ELL_RAG with conventional RAG for writing-assessment analytics. ELL_RAG showed a descriptive advantage overall and on global/population-level tasks, while local-evidence performance was nearly equivalent. Findings support task–evidence alignment rather than universal superiority and highlight tool orchestration as a continuing design challenge.
Automated Item Evaluation: Predicting Item Acceptance and Rejection using LLM-Generated Critiques
Hotaka Maeda | Yikai Lu
Hotaka Maeda | Yikai Lu
We built a near-comprehensive automated item evaluation model, predicting historical item acceptance or rejection from item text and Qwen3-generated critiques using 52,759 items from a large-scale testing program. Two DeBERTaV3 classifiers fused reached AUC .80 overall and .86 for math. Fairness-related rejections remained difficult, underscoring the need for human review.
Production-Ready Automated Item Generation in Educational Assessment: Integration with Operational Workflows
Hotaka Maeda | Kargi Chauhan | Tharunya Chandrashekar
Hotaka Maeda | Kargi Chauhan | Tharunya Chandrashekar
We present a production-ready automated item generation pipeline integrated into a large-scale assessment program’s workflows. A one-shot approach prompts LLMs from an operational source item and existing guidelines, then populates metadata, screens quality, and uploads items for human review. Generated items passed expert review, and difficulty prompting reliably shifted difficulty.
Using Generative AI to Simulate Item Responses by Skills Insight Score Band
Jianshen Chen | Chansoon Lee | Kylie Gorney | Judit Antal
Jianshen Chen | Chansoon Lee | Kylie Gorney | Judit Antal
This study evaluates GPT-5.4 Thinking simulations of reading-item responses using Skills Insight score-band descriptions and item content. Simulated and empirical item statistics and IRT parameters were compared. Difficulty recovery was strongest, with moderately strong b-parameter correlations, while a and c recovery was weaker, supporting preliminary difficulty evaluation before field testing.
Exploring AI-driven Methods for Pre-calibrating Difficulty of Mathematics Items with Visual Content
Leah Walker | Jinnie Choi
Leah Walker | Jinnie Choi
Use of AI is bringing efficiencies in replacing previously laborious field- testing with scaled item parameter pre- calibration, especially in language assessment. However, gaps in research exist in replicating this accomplishment for mathematics assessment items with visual content. This study explores the methodological options of using AI in pre-calibrating multimodal mathematics items.
Evaluating Multiple Models for Predicting Item Difficulty in a Principled Assessment Design
Alexandra Lane Perez | Christina Schneider | Sangdon Lim | Garron Gianopulos | Kang Xue
Alexandra Lane Perez | Christina Schneider | Sangdon Lim | Garron Gianopulos | Kang Xue
Item difficulty should, in theory, be predictable from features derived from RPLDs, task characteristics, and linguistic complexity. This study evaluates the performance of multiple statistical and machine learning models in estimating item difficulty. We found similar results across all four models, the item features selected explain 53%–56% of the variance across all grades.
Transformer-Aided Detection of Gaming in Constructed-Response English Language Assessments
Melanie Sharif | Scott Hellman | Martha Bellows | Sue Lottridge
Melanie Sharif | Scott Hellman | Martha Bellows | Sue Lottridge
We compare handcrafted features, frozen transformer embeddings, and ensembles across three item types and two regimes, testing pooled versus specialist models on a gaming detection task. Ensembles perform best (ROC-AUC 0.95, 𝜅=0.75), except for rarest item type under unseen prompts. Labels reflect review detections, motivating reference-conditioned recall and blind re-review.
Calibrated, Interpretable Automated Scoring with LLM Likelihoods
Thomas Christie | Markus Hauru | Anna Rafferty
Thomas Christie | Markus Hauru | Anna Rafferty
Paradigms like IRT use measurement equations to model student behavior probabilistically. We investigate an analogous LLM-based approach, the noisy channel model, to compute and compare the conditional likelihoods of student writing and produce auditable token-level inferences about student skills. We compare against other LLM-based methods on accuracy, calibration, and interpretability.
Predicting Item-to-Range Performance Level Descriptor Matches with Structured Language Model Reasoning
Kang Xue | M. Christina Schneider | Sangdon Lim | Garron Gianopulos | Alexandra Perez
Kang Xue | M. Christina Schneider | Sangdon Lim | Garron Gianopulos | Alexandra Perez
Range performance level descriptors (RPLDs) connect assessment items to claims about what students at different achievement levels know and can do. Retrospectively assigning RPLDs to a large item bank is valuable but labor intensive. This work-in-progress study evaluates whether small and large language models can assist expert item-to-RPLD matching for 524 Grade 3–5 English language arts items. We compared direct classification with structured prompts that guide a model to analyze the knowledge, skills, evidence, and cognitive demand required by an item. We also examined self-consistency voting and a rater-informed prompting. Preliminary results indicate that larger hosted models achieved the closest overall performance to humans, although some locally hosted models produced comparable results. Voting improved prediction reliability but did not consistently increase agreement with human scores, whereas training models with human scores generally improved classification accuracy. Model performance also tended to decline as grade level increased.
Using LLM Judges’ Paired Comparisons to Estimate Mathematics Item Difficulty
Lanrong Li | Mohammad A. A. Abulela | Matthew Gushta
Lanrong Li | Mohammad A. A. Abulela | Matthew Gushta
Estimating item difficulty without collecting field testing data has been a long sought-after goal in educational measurement. We prompted large language models (LLMs) to compare mathematics items in pairs to estimate their difficulty. Results showed strong correlations between difficulty based on one LLM judge’s paired comparisons and empirical item difficulty.
Math-Specialized or General-Purpose? A Comparison of MathBERT and DeBERTa-v3-large for Automated Scoring of Mathematics-Explanation Items
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Automated scoring of math-explanation items should handle responses that mix natural language with symbolic reasoning. The 2023 NAEP Automated Scoring Challenge showed that fine-tuning pre-trained encoders can reach human-like agreement, but its winning system. used a large, general-purpose encoder, while smaller math-specific pre-trained encoders remain a plausible and more efficient alternative. We fine-tune MathBERT (110M parameters) and DeBERTa-v3-large (183M parameters) as single-output regression scorers on 31 math-explanation items from one statewide summative assessment administration, benchmarking every model against human–human reliability and summarizing deployability with a three-tier acceptance criteria. On the 18 items fine-tuned under both encoders with identical datasets, the resulting performance between the two was practically indistinguishable: mean best QWK was 0.931 for MathBERT and 0.933 for DeBERTa-v3-large—a difference of only 0.002—and the models traded top performance at similar rates. Given similar performance, we tested a two-stage approach that used the larger general-purpose encoder only when the smaller math-focused encoder failed acceptance criteria. Extending MathBERT to all 31 items, it reached a mean QWK of 0.926 against a human–human benchmark of QWK = 0.940 and met operational acceptance criteria on 81% of items. Escalating the six items where a fine-tuned MathBERT model failed acceptance to a fine-tuned DeBERTa-v3-large model recovered four of the six, yielding acceptable automated-scoring models for 29 of 31 items. Because the smaller, math-pre-trained encoder matches the larger one at roughly 40% fewer parameters, we recommend fine-tuning a MathBERT encoder as the default and escalating to the larger, general purpose encoder like DeBERTa- v3-large only when a MathBERT model fails when deploying automated scoring models operationally at scale for math-explanation items.
How Much Does Hyperparameter Tuning Actually Help? An Efficiency Survey for Fine-Tuning Transformers to Score Mathematics-Explanation Items
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Gregory M. Jacobs | Ahmed H. Bediwy | Martha Bellows
Fine-tuning pre-trained transformer models for constructed-response items often begins with a hyperparameter grid search using k-fold cross-validation. We study how much that search actually helps by fine-tuning two encoders—MathBERT, a smaller math-focused model, and DeBERTa-v3-large, a larger general-purpose model—to score 18 math-explanation items from a state-wide assessment. Crossing learning rate, weight decay, and label smoothing over two epochs and five folds (1,800 model-fold runs), we compare each model’s tuning gain directly against fold-to-fold noise via a gain-to-noise ratio. Gains were small relative to fold noise for both encoders: MathBERT’s ratio fell below one (0.90), and DeBERTa’s nominally higher ratio (1.42) traced to a handful of divergent fits rather than an informative search landscape. Weight decay and label smoothing were effectively inert, leaving learning rate as the only hyperparameter worth checking—though even for learning rate the best configurations offered small performance gains over a reasonable default. We accordingly recommend a lean workflow that fixes the inert hyperparameters to sensible defaults, runs a narrow learning-rate search extended modestly upward, and increases the epoch budget with early stopping, substantially reducing training compute without sacrificing accuracy.
Responsible use of generative AI when creating reading comprehension questions: Inference matters
Zuowei Wang | Michael Flor | Jiyun Zu | Tenaha O’Reilly | Wanjing Anya Ma
Zuowei Wang | Michael Flor | Jiyun Zu | Tenaha O’Reilly | Wanjing Anya Ma
AI-generated and expert-created reading comprehension questions can show similar item statistics yet differ in the types of inferences required. This difference stemmed from AI’s failure to follow prompts during an intermediate item generation step. Evaluations of AI-generated items should document prompts and generation steps to identify and mitigate construct-relevant differences.
Fine-tuning DeBERTa-v3 to Automate Spatial Language Classification
Sam Agnoli | Qingzhou Shi | David Uttal
Sam Agnoli | Qingzhou Shi | David Uttal
To automate spatial language classification, we fine-tuned DeBERTa-v3, a transformer-based language model, using a 33,284-word dataset following a 70:15:15 train-validation-test split. The model performed well in a held-out test comparing performance to human coders (Cohen’s kappa = .88).
CDM Q-Matrix Discovery with LLMs: Fusing Domain Knowledge with Empirical Evidence
Susu Zhang | V. N. Vimal Rao
Susu Zhang | V. N. Vimal Rao
We present a method for diagnostic assessment Q-matrix specification combining LLM input with model-based empirical validation. An LLM generates an initial Q-matrix from item content, and response data guide subsequent Q refinement. The approach integrates substantive rationale with empirical evidence to support scalable and measurement theory-grounded diagnostic assessment.
Composite Scores vs. Preference Rankings: Measuring Architecture Effects in LLM Feedback
Harvey Ngoe Kolle | Carrie Demmans Epp | Amna Liaqat | Maria Cutumisu
Harvey Ngoe Kolle | Carrie Demmans Epp | Amna Liaqat | Maria Cutumisu
AI-generated feedback is often judged by quality ratings or preference rankings but rarely both. Comparing a multi-agent system, a single-agent system, and human feedback on student writing, we show that the two evaluation approaches can support different reported conclusions even when their underlying effects are nearly identical. These differences have consequences for how AI-generated feedback should be evaluated.
Identifying Communication Profiles from Physician–Patient Conversations Using Large Language Models
Reyhaneh Hosseinpourkhoshkbari | Yash Bipin Jain | Richard Golden
Reyhaneh Hosseinpourkhoshkbari | Yash Bipin Jain | Richard Golden
Large language models (LLMs) are increasingly used to score clinical communication, but validation is often limited to individual checklist items or total scores. These measures may not capture how communication behaviors occur together within a transcript. We therefore examine communication profiles: recurring combinations of behaviors that characterize different patterns of clinician communication and may support more targeted formative feedback. We analyzed 213 simulated respiratory OSCE transcripts rated on 15 binary Kalamazoo-derived communication items and compared final human ratings with GPT-4o, GPT-o3, and GPT-5.6 Sol. Raw item-level agreement was relatively high overall, but varied substantially across behaviors and was lower for several judgment-intensive items used for profile modeling. Bayesian latent-class analysis identified three stable human-derived profiles. Although the LLM-derived profiles showed broadly similar item-probability patterns, the models frequently assigned individual transcripts to different profiles than the human ratings. These findings show that agreement on individual communication skills does not necessarily translate into agreement on higher-level communication profiles. If LLMs are used to provide profile-based feedback, validation should therefore include profile-level agreement in addition to item-level performance.
LLM-Based Pairwise Judgment for Math Item Parameter Modeling Using Workflows and Agents
Suhwa Han | Frank Rijmen
Suhwa Han | Frank Rijmen
This study examines the feasibility of using large language models (LLMs) as judges for pairwise comparisons of item difficulty and to the extent which the resulting comparison outcomes recover banked difficulty parameters. The study in particular fine-tunes an instruction-tuned LLM to evaluate the impact of task-specific fine-tuning on the parameter prediction accuracy. The study also demonstrates an agentic approach to hyperparameter tuning of LLM-training task.
Beyond Volume: How Cognitive Fingerprints Predict Residual Gain in AI Tutoring
Aleena K Raj | Nina Deng
Aleena K Raj | Nina Deng
This study evaluates an AI tutoring tool in a pre-licensure exam preparation product. Using residual gain modeling, elastic-net feature selection, and Gaussian Mixture clustering on LLM-derived cognitive fingerprints, the study identifies meaningful learner profiles and shows that the cognitive depth of student–AI interaction, not volume, drives measurable learning gain.
This study evaluated automatic prompt engineering (APE) using one assignment in the PERSUADE 2.0 dataset. The APE approach achieved higher QWK (.812) than the research-informed, zero-shot baseline prompting approach (.646). Descriptive comparisons examined gender and English language learner subgroups. Findings support the benefits of APE for automated essay scoring (AES).
A Comprehensive Evaluation of GenAI-based Items in a Medical Examination
Yanlin Jiang | Marcus Walker | Andrew Dallas | Aquia Richburg | Nikole Gregg | Brittany Corrigan
Yanlin Jiang | Marcus Walker | Andrew Dallas | Aquia Richburg | Nikole Gregg | Brittany Corrigan
This study evaluated GenAI-based medical assessment items against SME-developed items. Analyses of item content and response data from a high-stakes examination showed that their content quality and psychometric characteristics are comparable. These findings empirically support the quality of GenAI-based items and their potential use in educational and professional assessment programs.
When One Benchmark Hides Many Truths: Mixture-IRT and LLM Difficulty Prediction
Enis Dogan | Burhan Ogut
Enis Dogan | Burhan Ogut
Studies that validate LLM-predicted item difficulty conventionally benchmark predictions against a single-population estimate. Using a multimodal LLM’s pairwise judgments on a 29-item mathematics exam, mixture Rasch modeling shows the LLM tracks difficulty ordering differentially across classes, meaningful subgroup variation that the aggregate benchmark conceals entirely.
LLM-As-An-Assessor: Can Open-Weight LLMs Assess Computational Thinking via Student-Designed Embodied Games?
William Lee | Sai Gattupalli | Ivon Arroyo | Brendan O’Connor
William Lee | Sai Gattupalli | Ivon Arroyo | Brendan O’Connor
We evaluate whether open-weight LLMs can assess Computational Thinking in middle-school students’ finite-state game designs against human labels. Across a rich and a sparse design, models detect some behaviors at near-human agreement but over-credit absent ones and fail at counting, loop detection, and standards mapping. Failures trace to fixable setup.
Funnel Plot Analysis of Examinee Performance in Longitudinal Assessment
Aquia Richburg | Marcus Walker
Aquia Richburg | Marcus Walker
We identify content areas where examinees underperform in a longitudinal assessment for a high-stakes medical licensure exam. Using modified funnel plots with a moving baseline to account for item difficulty, we rank topics by relative performance. Preliminary results suggest reasons beyond item difficulty, informing future analyses supporting potential educational interventions.
How Much Training Data Is Enough? Fine-Tuning LLMs for Short-Answer Scoring
Joshua A McGrane | David Torres Irribarra
Joshua A McGrane | David Torres Irribarra
Fine-tuned models approached state-of-the-art agreement using small labelled sets. Open-weight models met operational criteria on seven of ten items, with a median of 50 responses per score point among passing items. A single marking exercise may supply enough data for automated short-answer scoring.
Evaluating Machine Learning, Transformer, and LLM Approaches for Large-Scale Automated Essay Scoring
Cyrus Sai-Cheong Chan | Charles Anifowose | Su Zhang | Winnie Wing-Yee Tse | Nizam Radwan | Hyunah Kim
Cyrus Sai-Cheong Chan | Charles Anifowose | Su Zhang | Winnie Wing-Yee Tse | Nizam Radwan | Hyunah Kim
This study compared machine learning, transformer fine-tuning, prompt-based LLM, and novel ensemble approaches for automated essay scoring in a Canadian large-scale provincial assessment. Fine-tuned transformers achieved the highest reliability, followed by machine learning models. Meanwhile, prompt-based LLMs provided greater explainability, and ensemble architectures highlighted opportunities for balancing reliability and explainability.
Construct Validity of Small-Sample Transformer Scoring Models: A Mechanistic Interpretability Approach
Michael P. Hemenway | Martha Bellows
Michael P. Hemenway | Martha Bellows
Small-sample transformer scoring models ( n = 64 n=64) reach high human agreement but risk leaning on surface shortcuts like response length. Evaluating Mechanistic Interpretability strategies across 90 models, we show correlational methods suffer from seed noise, whereas interventional erasure proves small-sample models causally depend more on length. We outline an actionable audit protocol.
Applying Evidence-Centered Design to Automated Evals of AI-Powered Assessment Systems
Kristen DiCerbo | Britte Haugan Cheng | John Whitmer
Kristen DiCerbo | Britte Haugan Cheng | John Whitmer
This paper demonstrates an application of Evidence-Centered Design (ECD) as a principled approach to design automated evaluations of AI-powered assessment outputs. We demonstrate this application through Khan Academy’s “Explain Your Thinking” conversational agent for mathematics items, showing how ECD’s layered models can be translated into rigorous and interpretable systems to demonstrate validity, reliability and fairness in AI systems to diverse audiences.
This study investigates effective use of AI models as statistics tutors. Specifically, how to promote accurate student misconception identification, elicit active student cognition, and provide trustworthy responses. Initial findings from explorations of prompting strategies and indicators of sycophantic behavior with emphasis on model reasoning traces in simulated conversations are discussed.
Small language models can run offline in a classroom, but they make mathematical errors no teacher can afford to miss. This work-in-progress proposes a measurement-grounded professional development model in which upper-elementary teachers use local LLMs to generate and critically evaluate items; an illustrative pilot demonstrates feasibility, teacher-impact study planned.
Principled Approaches to Building AI Partners for Assessing and Supporting Small-Group Collaboration
Peter W Foltz | Chelsea Chandler | Mon-Lin Monica Ko
Peter W Foltz | Chelsea Chandler | Mon-Lin Monica Ko
Collaboration is complex and multifaceted, blending cognitive, social, and emotional components that resist simple measurement. We present a framework linking what is measured, where, and how evidence is warranted. This paper synthesizes key design dimensions for AI collaboration partners and demonstrates how they embed measurement science into practice.
Measuring Writing Proficiency Through Keystroke Process Data: A Cross-School Generalizability Study
Damilola Babalola | Naga Buddarapu | Caleb Scott | Piotr Mitros | Paul Deane | Collin Lynch
Damilola Babalola | Naga Buddarapu | Caleb Scott | Piotr Mitros | Paul Deane | Collin Lynch
Using a year of authentic Grade 7–8 writing from two U.S. middle schools (8,576 documents, 423 students), we test whether keystroke features correlate with English Language Arts proficiency, form interpretable latent dimensions, and vary by gender. Fluency and engagement features reliably index proficiency by Grade 8, though generalizability varies.
Self-Revising Agents as Item Writers: Benchmarking Against Professional Item Writing
Steven Tang | Zhen Li | Richard Patz
Steven Tang | Zhen Li | Richard Patz
Expert reviewers rated 156 calculus items from a self-revising multi-agent framework (Claude Sonnet 4.6 or GPT-5.4, three prompt conditions) and 26 human-written items. AI items were formative-ready nearly as often as human items (78–80% versus 88%) but summative-ready less often (26% and 10% versus 58%); a root-item reference mattered most.
Evaluating Neural Network-Based IRT in Multistage Adaptive Testing with Small Item Banks
Seong Eun Hong
Seong Eun Hong
This study investigates the performance of Neural Network-based IRT estimation in multistage adaptive language assessment with small item banks and short tests. Results showed that accuracy improved with larger samples and more training information. Moderate distribution shifts were largely accommodated under higher iteration conditions, whereas larger shifts continued to reduce estimation accuracy.
Validating LLM-Rated Item Features for Explanatory Item Response Models
Ahmed H. Bediwy | Mubarak O. Mojoyinola
Ahmed H. Bediwy | Mubarak O. Mojoyinola
Explanatory item response models let test de- velopers anticipate item difficulty from item design, but they depend on subject-matter ex- perts rating every item on every hypothesized feature—a step that is slow, costly, and the practical bottleneck limiting how many items can be modeled. We ask whether large lan- guage models can supply those ratings. Four trained experts and six LLMs independently rated 15 grade-5 mathematics items on six psy- chometric features. We evaluate the LLM rat- ings twice: against the human consensus using quadratic-weighted 𝜅and mixed-effects mod- els, and against 1,452 student responses by us- ing each source’s ratings as the design matrix of a linear logistic test model (LLTM) bench- marked against a Rasch baseline. Claude Opus 4.7 performed best in recovering the item dif- ficulty (r = 0.89), however, the human con- sensus against which agreement is measured performs worst at recovering Rasch difficulty (r = 0.33) compared to the rest of the LLMs. The reason is visible in the expert panel it- self: inter-rater Fleiss’ 𝜅is at or below chance on three of six features, so the consensus is a noisy reference rather than ground truth. We also identify two failure modes (zero-variance and perfect collinearity) that agreement statis- tics cannot detect but make a feature unusable as an LLTM covariate. We argue that agree- ment should be reported alongside criterion- referenced recovery of item parameters, not in place of it.
Efficacy of Student–AI Co-Authored Math Word Problems in an Intelligent Tutoring System
Kole Norberg | April Murphy | Steve Fancsali | Steve Ritter
Kole Norberg | April Murphy | Steve Fancsali | Steve Ritter
We evaluated student (N = 1,361) interest and performance on student-AI co-authored math word problems. Performance matched or exceeded standard problems. Students rated peer-authored problems more often, especially when authorship was disclosed. Liking predicted first-attempt accuracy when problems required greater textual engagement, supporting interest-based context personalization.
Assessing Small Language Models as Decimal-Arithmetic Tutors: A Measurement Framework
Mai Que Vuong | Shahana Ahmadli | Michelle Zhou | Minseok Kim | Shruti Mehta | Talita de Paula Cypriano de Souza | Seiji Isotani
Mai Que Vuong | Shahana Ahmadli | Michelle Zhou | Minseok Kim | Shruti Mehta | Talita de Paula Cypriano de Souza | Seiji Isotani
Small language models (SLMs) are increasingly proposed for educational use because they promise lower cost, offline deployment, and stronger data privacy. We report an exploratory evaluation of three sub-2B-parameter models on example-based decimal-arithmetic tutoring. Across structured interactions, all three produced fluent, confident output that masked unstable pedagogy and frequent mathematical errors even on elementary decimal addition and place-value tasks. Building on these observations, we describe an emerging interaction-based measurement framework intended to support more defensible readiness decisions about SLMs as math tutors.
LLM as Investigator and Judge in Pairwise Comparisons for Item Parameter Modeling
Brian Harrold | Vincent Fakiyesi | Maria D’Brot | Susan Lottridge | Justin Barber | Michael Hemenway
Brian Harrold | Vincent Fakiyesi | Maria D’Brot | Susan Lottridge | Justin Barber | Michael Hemenway
This study synthesizes feature-based and pairwise-comparison approaches to item parameter modeling by using an LLM as an investigator of item characteristics and a judge of relative item parameters. The proposed framework aims to improve recovery of item parameter rankings and predictions while providing insight into factors associated with the parameters.
Using an LLM-assisted pipeline, this study examines sociodemographic variation in English Language Learners’ hedge use across 6,500 essays. Regression analyses are expected to show that gender, race, and SES predict hedge density and category-specific variation, which subsequently predict writing scores. Beyond scaling annotation, the study evaluates whether LLM-derived metadiscourse indicators provide interpretable evidence about rhetorical development in ELL writing.
Code-editing language-model agents can change training choices beyond the hyper- parameter grids tested in automated essay scoring. Comparing these approaches requires distinguishing gains from a broader search space from evidence of a better search procedure. We compare code-editing agents with grid-restricted search, including random search, in two studies on ASAP-AES with nominally matched 12-hour search budgets. Code-editing produced the configuration with the highest test quadratic-weighted kappa (QWK) point estimate in each primary comparison. However, random search over a grid built afterwards around the Study 2 agent’s backbone, input length and head rule recovered most of its gain over BERT. Requiring a minimum validation-score improvement to accept a trial also left the accepted configuration’s validation score below the highest recorded valid validation score in every run that accepted a trial under this rule. The primary comparisons use one search run per condition, and some test folds reused essays involved in configuration selection, so these results compare selected configurations without establishing search-procedure superiority. These findings motivate reporting the available search choices and both accepted and best-scoring trials, and evaluating search procedures through repeated runs on data independent of configuration selection.
LLM Difficulty Prediction Across Three L2 Task Types and Two Examination Forms
Meng Lyu | Moses Oluoke Omopekunola | Feng Wang
Meng Lyu | Moses Oluoke Omopekunola | Feng Wang
Large language models (LLMs) have been proposed as a substitute for field trials in itemdifficulty estimation, but evidence about where they succeed is usually drawn from a single test form. We compare LLM difficulty predictions across three task types that make different language demands, grammar cloze, semantic cloze and reading comprehension, on two operational Grade 9 English examinations sat by the same cohort in China (1,199 and1,166 examinees). Three LLMs estimated difficulty for 60 MCQ items under the same standardized protocol. On Form 1, grammar cloze had the highest Pearson correlation for all three models; the Fisher-averaged correlations were .81 for grammar cloze and .34 for semantic cloze. On Form 2, reading had the highest correlation for all three models (Fisheraveraged .78), and the grammar–semantic difference was smaller and inconsistent in direction. With ten items per task type, a task-type advantage seen on one form needs testing on other forms before it is read as a property of the model.
Measuring Science Teachers’ Generative AI Literacy for Classroom Assessment: Performance Assessment
Ruiping Huang | Yue Yin
Ruiping Huang | Yue Yin
We report preliminary analyses from an ongoing study using performance-based tasks to assess science teachers’ generative AI (GenAI) literacy for classroom assessment and compare these scores with self-reported GenAI use and confidence. Preliminary findings indicate moderate correlation with use frequency but weak correlation with confidence, highlighting performance assessments’ value.
Do Synthetic Students Thrive? Validity and Fairness of Generated Assessment Data
Neba Nfonsang | Temple Lovelace
Neba Nfonsang | Temple Lovelace
Synthetic assessment data may reproduce response distributions while failing to preserve underlying construct structure. Using a retrieval-augmented generation (RAG) pipeline and a stratified majority Black and Latinx student sample, we evaluated synthetic responses using distributional and psychometric measures. Results highlight the importance of psychometric preservation alongside distributional similarity.
AI-Enabled Quality Assurance for Multiple-Choice Assessment Items
Steven James Moore | Nicholas Diana
Steven James Moore | Nicholas Diana
Generating multiple-choice questions is increasingly scalable, but establishing their quality remains difficult. We review fourteen reports on automated item-writing flaw detection, revision, psychometric screening, and benchmark auditing. High accuracy often masks weak detection of flawed items, and revision evidence is mixed. We propose evaluating quality assurance as independently validated decisions.
More than Words: Modeling What Teachers Praise and How They Praise It
Jiseung Yoo | Paiheng Xu | Jing Liu
Jiseung Yoo | Paiheng Xu | Jing Liu
Praise is one of the most frequent evaluative acts in teaching, but what a praise turn communicates depends less on its words than on how it is delivered. We propose measuring delivery as self-deviation, comparing a praise turn against two references drawn from the speaker themselves, the teacher’s other praise turns (Global Reference) and the speech immediately surrounding the turn (Local Reference).
Does an Automated Source Integration Score Respond to Revision? Extending Evidence of Construct Validity with a Feedback Experiment
Andrew Potter | Yu Tian | Kaitlin Van Houghton | Manmeet Singh | Renu Balyan | Maria Goldshtein | Laura K. Allen | Danielle S. McNamara
Andrew Potter | Yu Tian | Kaitlin Van Houghton | Manmeet Singh | Renu Balyan | Maria Goldshtein | Laura K. Allen | Danielle S. McNamara
This study examined whether an automated source integration measure responds to revision. In a randomized experiment, undergraduates drafted and revised a source-based essay. Source integration scores increased across drafts, although gains did not differ significantly between students who received source integration feedback and those who did not. Findings support use of the measure for formative evaluation and feedback.