Artificial Intelligence in Measurement and Education Conference (AIME-Con) (2026)


up

pdf (full)
bib (full)
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers

Test assembly, the process of constructing a complete test form from an item pool subject to blueprint constraints, has traditionally been treated as a static optimization problem. In AI-enabled assessment environments, however, item pools evolve continuously as newly generated items enter with uncertain psychometric parameters, and delivery is on demand. These conditions make test assembly a sequential decision-making problem under uncertainty: which form should be deployed now, given current but incomplete knowledge of item quality, to simultaneously maximize measurement precision, satisfy content-blueprint constraints, maintain pool sustainability, and accelerate calibration of uncertain new items? This paper proposes the Stochastic Constrained Hybrid (SCH) framework as a principled answer to this question. SCH recasts form-level assembly as a multi-armed bandit (MAB) problem with Fisher information as the reward, extending recent item-level approaches in computerized adaptive testing (CAT) to the form-level setting. A simulation study comparing six test assembly methods is also presented. The main contribution of this paper is a framework for incorporating items with uncertain parameters into the automatic test assembly process for linear test forms.
PsychMet is a domain-grounded chatbot that uses a GPT-4.1 conversational model with retrieval-augmented generation over a curated psychometrics corpus, with emphasis on IRT and NCME competencies. Using the RAGAS framework on a 30-question set, PsychMet achieved an Overall score of 0.539, with strengths in Answer Correctness (0.810) and Context Recall (0.671), moderate Faithfulness (0.588), and weaknesses in Answer Relevancy (0.284), Context Precision (0.425), and Context Relevancy (0.474). This pattern suggests that retrieval breadth is outpacing specificity. We outline targeted fixes — such as hybrid sparse+dense retrieval with light filtering and question-first prompting — to tighten focus without sacrificing coverage. PsychMet is accurate and transparently sourced for exploratory learning; with retrieval tightening and answer scoping, it can better support time-bound professional workflows.
How students use generative AI is a configuration of distinct behaviors, not one skill, yet scoring typically collapses it onto a single proficiency axis, conflating process with product. We develop a multidimensional Bayesian IRT model with rater effects and post-hoc Varimax-permutation identification, surfacing three substantive dimensions: prompting effort, content delegation, AI-delivered citations. Domain familiarity shifts students toward more active engagement and away from delegation.
This study predicts IRT item difficulty and discrimination for passage-based reading comprehension items using lexical, syntactic, and semantic NLP features. Modeling interactions among passages, stems, and options, Light- GBM with semantic embeddings best predicts difficulty (r = 0.594), while discrimination remains harder to recover from text alone.
This simulation study evaluates whether predictive item-parameter priors can reduce respondent requirements for 3PL IRT calibration. Across twenty seven prior-quality configurations and seven sample sizes (400–1600), six configurations were able to match a 2000- respondent baseline at 400 respondents, yielding an 80% sample-size reduction when informative priors were properly incorporated.
While automated writing evaluation (AWE) systems typically assess essay quality, we develop a process-oriented approach assessing revision alignment with feedback. Classroom deployments show that student perceptions of AWE feedback correlate with writing improvement. Furthermore, feedback designed to build revision knowledge can enhance student outcomes, shifting focus from product to process.
We examine whether LLMs generate feedback aligned with expert teachers’ practices in feedback focus and adaptivity. We present FeedType, a benchmark of annotated teacher and LLM feedback to evaluate this alignment. Results reveal that while LLMs cover most feedback types, they fail to fully replicate teachers’ feedback distributions and adaptivity.
We introduce S2A3, a unified Bayesian framework for high-stakes computerized adaptive testing that eliminates separate item piloting. Thompson sampling routes uncertain items to informative test-takers while soft scoring attenuates their influence on ability estimates. Stochastic Sympson-Hetter exposure control ensures bank security. Validation on the Duolingo English Test confirms rapid calibration.
We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice probabilities across a large training corpus of multiple-choice items containing both image and text stimuli, conditioned on a labeled set of student ability levels. By learning to reproduce the systematic error patterns of students across a discrete range of abilities, the LLM implicitly captures the underlying response probabilities encoded in the 3PL and MCM curves. This allows us to accurately approximate item difficulty on a held-out test set directly from the model’s predicted option probabilities.
In programming courses, students are often asked to explain code fragments, which is a way to assess their understanding of programming constructs and patterns. These types of problem known as “explain in plain English” are valuable in both assessment and practice contexts. The main challenge to using these types of problems at scale in both contexts is automating the scoring process, i.e., assessing whether those explanations are correct. The prevailing approach scores an explanation by its semantic similarity to an instructor’s model explanation, but this raises a measurement concern: students who reason correctly yet phrase their explanations differently from an expert may be scored as incorrect (false negatives), threatening the validity and fairness of the assessment. Given recent advances in LLM-based automated scoring, it remains unclear whether semantic similarity methods are still the most effective technique for automatically scoring free-form student responses, such as code explanations. In this paper, we present a rigorous comparison between LLMs and semantic similarity approaches for the automated scoring of student code explanations, using an open dataset Selfcode 2.0. We frame the scoring as a binary classification task and use generative AI to balance the dataset. Our results suggest that LLM-based scoring (F1 = 0.98, accuracy = 0.96) outperform semantic similarity scoring (F1 = 0.72, accuracy = 0.65). It also produces fewer false negatives and eliminate the need to produce well-formulated model explanations.
This paper describes the development and validation of an automated scoring model for open-ended chatbot-based conversational speech among English learners in the Duolingo learning app. The model strongly predicted human ratings, produced reliable scores, and showed concurrent validity with Duolingo English Test speaking items, demonstrating stealth proficiency assessment at scale.
Black-box knowledge tracing models are commonly deployed to drive personalization in online learning platforms and are typically evaluated using classification metrics such as AUC, accuracy, and F1 score. However, a model that predicts item responses using only each item’s proportion correct in the training set achieves AUC up to 0.72 and accuracy up to 0.84 on widely used benchmark datasets, despite using no information about individual students’ response histories. Inspired by the concept of marginal reliability in psychometrics, we introduce Fractional Information Gain (FIG), an information-theoretic evaluation metric for black-box predictive models of student item responses. FIG measures the fraction of a student’s response uncertainty resolved by a trained model relative to the item-only baseline. FIG equals 0 when the model adds no information beyond item base rates, and 1 when held-out responses are perfectly predicted. FIG is applicable to any model that outputs probabilities and is sensitive to calibration errors. We characterize FIG on synthetic and real data, compare it to AUC for both the item-only baseline and trained models on four benchmark datasets, and describe the operational affordances that FIG inherits from its reliability-like construction. FIG imports the conceptual benefits of score reliability into the prediction-oriented framework of ML-driven educational systems.
We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies. We estimate the factors that contribute to LLM performance, quantifying the disconnect between human and LLM difficulties in SAS and provide recommendations for developers working on SAS for schoolchildren.
Credentials often fail to capture skills developed outside formal education, limiting access to opportunity. We introduce Current Skills Validation, a multi-agent architecture that decomposes skills assessment into a set of specialized agents, supporting fine-grained adaptivity grounded in learning science. We describe its architecture, foundations, and measurement agenda.
We compare traditional machine-learning models and instruction-tuned Gemma 3 models for identifying knowledge components from student code. On 1,800 GPT-4o-annotated Dart submissions, feature-based ML is more accurate and substantially faster than LLM approaches. Joint multi-KC prompting also suffers from format failures, highlighting deployment-oriented advantages of traditional ML.
This study evaluated whether automated scoring engines maintain stable performance when training data composition and training methods vary. We manipulated demographic representation (gender, English language learner, race, student with disabilities) and compared feature-based versus transformer-based models. Results showed performances were stable across subgroup-representation densities and transformer models exhibited greater stability.
We examine model design choices for encoder-based prediction of IRT-3PL item parameters in reading comprehension tasks with passage-based multiple-choice questions (MCQs). We compare three aspects: (1) encoding strategies for passage and question texts, including optional cross-attention, (2) per-parameter vs. joint-parameter prediction with MSE and CRPS losses, and (3) random vs. passage-grouped batching. Results show that separate-then-fuse encoding is substantially more efficient than concatenated-input encoding while maintaining comparable parameter recovery. Cross-attention improves recovery for discrimination and guessing but weakens difficulty recovery. Joint-parameter prediction achieves performance close to per-parameter prediction in the strongest configurations while reducing training time. CRPS is most useful for joint prediction, and passage-grouped batching consistently improves recovery of the guessing parameter. Overall, joint prediction with CRPS and passage-grouped batching provides the most balanced performance while maintaining the efficiency of separate-then-fuse encoding without cross-attention, showing that the proposed model design can support efficient IRT parameter prediction for passage-based MCQs.
This study evaluates synthetic open-response data generation for assessment development by using real response data as ground truth. It finds that conditioning LLM generation on actual test-taker responses improves fidelity and diversity, though synthetic responses still lack the full variation of real test-takers, limiting current psychometric utility.
This study evaluates machine-learning classifiers for detecting lower justification depth in student responses from AI-supported instructional design tasks. Using 600 human-coded responses and 20 group-aware repeated train-development-test partitions, RoBERTa outperformed TF-IDF logistic regression and Complement Naive Bayes under recall-prioritized thresholding for formative assessment support.
We examine alignment at both node and taxonomy levels across four major skill taxonomies. Using semantic similarity and cluster analysis, results indicate uneven overlap: some skills converge, and some taxonomies interleave more than others. These findings could inform educational assessment by clarifying construct validity claims and score interpretations across frameworks.
Retrieval practice supports learning but requires educators to build large item banks. We compared LLM-generated and human-written retrieval practice items in an introductory psychology course to test whether LLM items match instructor-written ones in quality. LLM items exhibited overall weaker psychometric properties, suggesting that human supervision may remain necessary during item generation for retrieval practice.
In a within-subject study, students learned via search and LLMs, then wrote essays without them. Analyses show no detectable difference across stylometric and perplexity features; AI detection probability is tool-dependent; plagiarism overlap is higher after search. Results suggest asymmetric detection signals across learning pathways, with implications for interpreting academic integrity tools.
This study introduces an LLM-powered Smart Report Assistant grounded in audience analysis and assessment design principles. Using retrieval-augmented generation, the system helps teachers interpret assessment data through personalized and conversational reporting. Results from a usability study indicate promise for supporting score interpretation and data-informed instructional decision-making.
This study investigates Large Language Model (LLM) surprisal as an indicator of global linguistic naturalness in L2 Japanese automatic essay scoring. Results demonstrate that surprisal effectively distinguishes learner proficiency levels. Combining surprisal with feature-based indices achieves the highest classification accuracy, validating surprisal as a valuable metric for automated writing evaluation.
This conceptual paper presents a validity threat framework for measuring students’ GenAI use in higher education. It identifies threats related to construct definition, response processes, self-report bias, policy context, fairness, and score interpretation, offering practical guidance for stronger measurement, assessment, policy, and pedagogy in AI-integrated education.
An XGBoost prediction model was fitted to online skill practice data for a collection of 1800+ skills to construct a vertically aligned skill difficulty scale spanning kindergarten through 8th grade math. The results were validated through association with an independently derived Rasch vertical scale.
This study presents a causal approach for evalu- ating automated generation of reading passages for early readers. It consists of estimating the average treatment effect of a candidate model, then conditional average treatment effects with causal forests. We apply it to evaluate a fine- tuned model for a digital literacy platform
MCQ-Diag is a diagnostic tool for reviewing distractor quality in multiple-choice assessments through three interpretable indicators: semantic plausibility, semantic uniqueness, and lexical distinctiveness. Rather than assigning automated judgments, it presents these as evidence within an interactive review environment. Semantic plausibility and lexical distinctiveness show modest, statistically significant validity against expert ratings.
This study presents a comprehensive operational pipeline for multilingual GenAI scoring that integrates secure data preparation, prompt generation, automated scoring, post-processing, and human-in-the-loop review. Results demonstrate the pipeline’s strong adaptability across item formats, assessment languages, and assessment cycles, offering a viable solution to the persistent challenges of multilingual scoring.
Prior automated writing feedback systems have shown limited evidence of transferable writing development. This pilot longitudinal study examined whether repeated interaction with process-oriented LLM feedback improved independent writing performance over time. Results suggest positive transfer effects across novel writing tasks, supporting the potential developmental role of LLM-mediated feedback.
This study evaluates score agreement and uncertainty estimation in large language model (LLM)-based automated essay scoring. Results showed that LLMs tended to be overconfident and highly consistent in their assigned scores, and majority voting did not yield stronger agreement with human scores than deterministic scoring.
We evaluate the psychometric reliability of Bayesian Knowledge Tracing for skill mastery classification. Using large-scale intelligent tutoring system data, we compare two simulation-based approaches for estimating classification consistency as reliability. Results support both approaches, show reliability increases with response-sequence length, and highlight the necessity of adjusting for chance agreement.
This study reports on an open data science competition based on a benchmark dataset of mathematics misunderstandings, comprising over 52,000 mathematics explanations written to justify answer choices from 15 multiple-choice questions that were labeled by expert human annotators. Competitors were tasked with correctly classifying the explanations and any misunderstandings. By combining stable validation methods with efficient inference and enriched training data, top teams achieved high accuracy scores that correctly classified the explanations as correct, a misunderstanding, or neither and, if it did have a misunderstanding, what type of misunderstanding it was.
Applying Exploratory Factor Analysis to human and LLM responses on chemistry and quantitative reasoning assessments, we asked domain experts to blindly interpret emergent factors. Experts decoded human factors successfully but could not interpret most LLM factors, suggesting LLMs rely on statistically opaque mechanisms distinct from human cognitive constructs.
AI-based item generation and NLP-based prediction of item parameters are producing item banks that are substantially larger, sparser, and more frequently updated than conventional banks. Hierarchical Bayesian item response theory (IRT) is a natural calibration framework for such banks, but the common practice of refitting the entire accumulated response history at each update is costly and can exceed available memory. We describe consensus calibration, a divide-and-conquer procedure that calibrates each time period independently and reconstructs the pooled posterior in two layers. First, the posterior draws of each period are mapped to a common metric by a robust characteristic-curve linking (Haebara) that is solved separately for each draw, which propagates the uncertainty of the linking transformation into the linked posteriors. Second, the linked item posteriors are combined as a product of Gaussian densities from which the population prior contributed by each period is removed and a single prior—obtained by consensus across the per-period population posteriors—is reinstated. The correction targets the posterior dispersion, not only its location. As evidence for consensus calibration, we compare it to a pooled single-run analysis on a large operational assessment in terms of item-parameter recovery, an uncertainty-by-exposure diagnostic, and the ability distributions.
We investigate the problem of promoting individual fairness in automated scoring. Three models are trained and evaluated under different individual fairness constraints. The models show improved scoring consistency on the test set, but at the cost of slightly reduced accuracy and a tendency to regress scores towards the mean.
We develop CAP, a preference optimization-based method for aligning LLM student simulators to student ability. It constructs preference pairs from IRT ability gaps, training the simulator to generate responses that better reflect prompted ability, enabling simulation-based estimation of open-ended item difficulty.
This study identifies a selected transformer-based clustering model for operational medical assessment items that balances within blueprint-topic proximity with cluster separation. Operational analyses showed supplementary alignment with blueprint structure, flagged isolated content areas that may need additional item coverage, and identified dispersed topics for subject matter expert (SME) review.
This study presents a Generative AI (GenAI) tool designed to create educator-facing interpretive materials, referred to as use guides, tailored to specific use cases. Through a four-step workflow, system prompt design, embedded guardrails, and three-round refinement, the tool produces context-specific, data-grounded interpretations and actionable recommendations.
AI-based comparative judgment was evaluated as a tool for early item difficulty estimation. Three AI judges compared new TIMSS Grade 4 mathematics items with calibrated anchors, with rankings analyzed using a fixed-anchor Bradley-Terry model. Results show meaningful difficulty signals, supporting scalable supplementary use while highlighting anchor coverage and comparison-network design.
One approach to automated essay scoring re- lies on neural regression models that output continuous predicted scores, yet operational scores must be reported on a discrete, ordi- nal rubric scale to match human raters’ scores. The mapping from continuous predictions to integer scores is governed by a small set of cutpoints, and the choice of cutpoint esti- mator materially changes both accuracy and the score distribution that examinees expe- rience. We introduce, formalize, and com- pare seven cutpoint methods—simple round- ing, exact-agreement optimization, balance op- timization, quadratic-weighted-kappa (QWK) maximization, maximum-likelihood Gaussian boundaries, a Bayesian ordered-prior estimator, and an ordinal cumulative link model—under a single notation. We evaluate all seven on seven prompts from the ASAP 2.0 dataset (24,728 essays), using a shared Longformer regres- sion backbone, and report agreement (quadratic weighted kappa), association (the Pearson cor- relation between integer machine and human scores), and distributional fidelity (the stan- dardized mean difference and the maximum score-point distribution gap). Our results ex- pose a consistent accuracy–calibration trade- off: QWK maximization and balance optimiza- tion tie for the highest agreement, balance opti- mization achieves the smallest score-point dis- tribution gap, the ordinal link model achieves the smallest mean bias, the generative Gaussian and Bayesian estimators recover score means but distort the distribution, and the ordinal link model is a strong all-rounder.
Item difficulty prediction relies on the efficient utilization of high-leverage item characteristics. Many items in the domain of mathematics include figures, charts, or other visual stimuli that are challenging to incorporate into difficulty prediction models. Multimodal large language models (LLMs) offer a way to process these visual stimuli in combination with text input, potentially enhancing the success of item difficulty prediction. In this study, we employ several open-source multimodal LLMs to predict the difficulty of math items with visual stimuli from a publicly available data set. We find that multimodal LLMs are capable of predicting math item difficulty.
This study explores the integration of a generative language model (GLM) into an Automated Writing Evaluation (AWE) system designed to highlight both argumentative components and errors in spelling and grammar. We fine-tune an open-source GLM using parameter-efficient techniques. We evaluate the model’s capabilities in both argument analysis and error detection against established datasets. We demonstrate that a single GLM with parameter-efficient adapters can accurately identify argumentative clauses, classify their types, map relationships between them, and flag mechanical mistakes. We establish that our AWE system performs at human-level accuracy, while only requiring a fraction of the computational power of much larger models.
Prompting functions as a measurement condition that alters LLM coding behavior. We examined whether few-shot prompting improves GPT-based qualitative coding in systematic reviews. Contrary to expectations, we found that the zero-shot approach produced the highest agreement with human coders, suggesting increasing conservatism as more prompt structure was added.
Researchers claim to measure large language models. Measurement and evaluation differ epistemically: evaluation asks whether outputs are fit for use; measurement requires a pre-existing quantitative attribute. LLM variability is designed, not discovered; evaluate models, measure people. Critical AI literacy proves, in simulation, most measurable where AI is least capable.
We compare long-context architectures (ModernBERT, Longformer, and causal and bidirectional Mamba) for data-efficient automated essay scoring across four training sizes on ASAP-2. Bidirectional Mamba matches the transformers. Deployment readiness depends more on label quantity (rising from 60% to 84% as training grows from 128 to 512 essays) than on backbone choice.
LLM simulation suffers from an “over-knowledge problem”, where models perform too well to represent struggling learners. We compare persona-, IRT-, and CDM-based prompting for student simulation and measure how well methods reflect expected behaviors and quantify this bias. Findings show LLMs follow IRT parameters, yet struggle to simulate real abilities.
Student-generated mathematics metaphors reveal students’ attitudes and beliefs but are costly to code manually. We evaluate LoRA-based fine-tuning of compact open-weight LLMs for valence-intensity and thematic coding. Fine-tuning substantially improves coding performance and reliability, making these models competitive with proprietary prompt-only LLMs while supporting local, privacy-conscious deployment.
Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual sentence embeddings can replace translated English input for Linguistic-Integrated Reliability Auditing (LiRA) across 11 PIRLS constructed-response items and three embedding models. Native-language embeddings reproduced translation-based reliability estimates closely while recovering responses excluded after translation failure, with no meaningful change in reliability.
This study evaluates automated approaches for detecting concerning content in open-response SJTs used in higher education admissions. Comparing fine-tuned BERT models with zero-shot and fine-tuned LLMs, we found that fine-tuned BERT achieved the strongest performance despite not receiving the scenario context available to the LLMs
This study evaluates direct few-shot scoring and anchor based comparative judgment for LLM essay grading. Using human scored anchor essays, we compare their accuracy and stability. Although both approaches produce competitive scores, comparative judgments demonstrate better stability and consistency, suggesting it as a better alternative for grading writing assessment.
Comparing LLM-human rater rationales using semantic similarity risks conflating textual proximity with evaluative agreement. We test whether embedding-based similarity reflects qualitative coding distinctions across rationale pairs in medical education. Similarity declined as rationale length differences grew and was less effective at distinguishing whether LLMs preserved the human’s central claim.
This study provides validity evidence for automated scoring of oral reading fluency (ORF) using off-the-shelf automatic speech recognition (ASR) models. A corpus of 320 audio recordings of 63 children collected with a digital literacy platform was used to estimate words correct per minute (WCPM) with six model variants.
This work re-frames anomalous test response detection in standardized testing as binary image classification by transforming test response logs into fixed-size matrix representations and treating them as grayscale images. Deep learning architectures, including ResNet and Vision Transformer, are evaluated. Experiments demonstrate that the approach is feasible and achieves high recall.
Fine-tuning large language models for automated scoring is often formulated as either a regression or a classification task. However, scores assigned by human raters are on an ordinal scale. We summarize different deep-learning ordinal regression methods and compare their performances with those from regression and classification using a real dataset.
This study evaluated whether visual embeddings from SigLIP and DINOv2 could be used to simulate responses to Figure Matrices items. Models predicted response probabilities using item screenshots and examinee ability estimates. Results showed moderate probability recovery but limited item difficulty recovery, indicating promise for early screening but not calibration replacement.
We introduce an evidence-centered, argument-based framework for validating AI-generated assessment items. The framework organizes three claims—content comparability, construct validity, and psychometric functioning—with assumptions and evidence. Applying the framework to critical thinking assessments, we compared AI-generated and human-authored items and identified distinct mechanisms underlying the validity of AI-generated items.
We present a production system for hierarchical analytic writing feedback that uses frontier LLMs to generate training data for fine-tuned transformer models. Applied to Grade 5 opinion essays, the system scores 5 competency dimensions and 31 binary feedback codes, achieving agreement approaching inter-rater levels from training sets of 300–500 essays.
Cloze exercises offer scalable comprehension assessment, but their validity depends on which words are selected as gaps. We compared three automated methods for cloze exercises generated from summaries within an intelligent textbook platform. Conditioning masked language model predictions on the source text (contextuality-plus) produced higher quality and more source-dependent gaps.
We used confirmatory factor analysis to assess the reliability and construct representation of an LLM-based measurement instrument of language proficiency. LLMs were at least as reliable as human raters and loaded onto the same underlying factor, though analyses indicated a less than perfect alignment between LLM and human raters.
Our work examines the evolution of undergraduate students’ appraisals of genAI fairness in school curriculum across a two-year horizon. We identify and present differences in students’ fairness appraisals and rationales between this academic year and last, delving into underlying sources of change and their implications for higher education policy.
Centroid margin, an embedding-based measure of relative domain distinctiveness, was evaluated as a pre-data item-selection tool for multi-domain instruments. Across four samples, centroid-guided selection improved model fit and discriminant validity, whereas own-domain cosine similarity increased overlap among closely related domains, revealing a key limitation of absolute similarity metrics.
Artificial Intelligence use raises academic integrity concerns, yet institutions rely on detection tools of questionable reliability. Rapid version updates to chatbots, detectors, and humanization tools remain unaccounted for in detection research. Across 960 detection scores from two waves, the findings reveal detection outcomes are fundamentally reshaped by these parallel updates.
GenAI can power simulated student agents that provide opportunities for educators to engage in core teaching practices, such as leading a small group argumentation-based science discussion. To realize the potential of such simulations and support teacher reflection and learning, it is necessary to provide participants with timely feedback on their performance in the simulation. This study investigates systems for automated evaluation of and feedback on teacher performance in a simulation along the dimension of making use of student ideas to move the discussion forward. We address three research questions: (a) How well do models fine-tuned on transcripts of teacher performance in a matching human-puppeteered teaching simulation (that served as the model during the development of the GenAI one) perform in evaluating transcripts from the GenAI teaching simulation? (b) How well does a system using a few-shot LLM perform on the same task? (c) How do educators perceive the quality and usefulness of the automatically generated feedback? The findings underscore the importance of a rigorous evaluation of automated evaluation and feedback systems.
This study examines the relationship between high school students’ reliance on GenAI and potential predictors: student demographic characteristics, academic background, and learning goal orientation. The findings highlight the importance of guiding students to use GenAI appropriately, preventing over-reliance on AI, and fostering mastery goal orientations in the AI age.
This paper examines elementary educators’ instructional skills as they practice eliciting student thinking in a generative AI (GenAI) mathematics teaching simulation. Study findings indicate that educators were able to elicit some aspects of a GenAI student’s conceptual understanding and misunderstanding. Implications for using GenAI teaching simulations as a practice space and formative assessment tool to build and evaluate educators’ instructional skills are addressed.
This study evaluates whether three locally deployed small language models (LLaMA-3.2-1B, Qwen-2.5-0.5B, and Gemma-3-1B) can implement Socratic tutoring in Grade 6–8 mathematics. Using a rubric-based protocol across 72 sessions, results show that all models struggled to sustain guided inquiry, with Gemma performing best overall.
This paper evaluates an AI-based simulated administration approach for estimating item difficulty before empirical calibration. AI-generated option probabilities for synthetic examinees were converted into item difficulty estimates and compared with operational Rasch difficulty statistics. Results showed meaningful alignment, highlighting simulated administration as a potential tool for earlier item performance evidence.
Large language models are increasingly used as automated judges in education, yet their ability to score pedagogical quality in AI tutor responses to K-12 STEM student inquiries remains underexplored. This study evaluates whether two LLM-based scorers, Nemotron-3-Super-120B-A12B and GPT-OSS-120B, can approximate human judgment of single-turn Socratic style responses to STEM inquiries. Using a four-dimension rubric adapted from the CPS-R Questioning and Thinking subscale, we compare human and model ratings. Results show mixed reliability: agreement is stronger for more observable instructional features such as Cognitive Demand, but weaker for more interpretive dimensions, especially Encouraging Metacognition and Differentiation. Chance-corrected reliability is also sensitive to skewed score distributions, as shown by a base-rate effect in Differentiation. A mixed-effects analysis further reveals that the two LLM scorers differ systematically, with larger divergence on elementary-level items. We also observe prompt-adherence failures in generated tutoring responses, where some outputs briefly violate the instruction to avoid direct answers before returning to a Socratic response. Overall, the findings suggest that LLMs can assist with large-scale pedagogical evaluation, but human oversight remains necessary for nuanced instructional assessment and for maintaining Socratic response style.
To address the "learning system wall," we introduce an AI system that converts tutoring screen recordings into unified transcripts of dialogue and on-screen actions. We present a method for aligning and classifying learning processes against MATHia logs, marking an initial step toward generalizable cross-platform learner modeling.
This study measures the effectiveness of including images to predict discrimination parameters of reading items. The results suggest that providing images to language models is beneficial. By using both image and text, Lasso regressions (Tibshirani, 1996) could explain data better and random forests (Breiman, 2001) showed higher accuracy.
Item response theory forces a choice between flexible curves and tractable inference. We bridge them by representing item response functions as Bernstein polynomials. With a natural conjugate Beta prior, this gives a closed-form ability posterior as a Beta mixture. It reproduces IRT and recovers non-monotone curves with bimodal posteriors.
Regenerating LLM Q-matrices ten times per model on TIMSS 2011 items, we find two runs of the same model reassign 61–88% of student mastery profiles. Greedy decoding removes this instability; disagreement with expert judgment (63–76%) survives. Majority voting fixes neither. Report distributions, not single runs.
We present CRS, a framework representing group reasoning as contributions, relations, and derived structure. Evaluating five LLMs on student discussions, we find that reasoning is classifiable but not reliably segmentable, with identification being the bottleneck. Prompting improves labeling far more than identification; participation is recoverable, yet fine-grained structure is hard.
Geometry item creation is often slow, as every figure must be checked by hand. AI can draw figures, but they often look right yet are subtly wrong. We propose item generation where a LLM writes structured constructions and a geometry engine builds, checks, and scores diagrams automatically.
We compare the traditional three-correct-in-a-row mastery rule with an adaptive, measurement-informed mastery model using ASSISTments Skill Builder data. The adaptive model produced more consistent decisions overall, especially early in learning, while rule-based mastery varied across skills, templates, and hint use, suggesting a need for context-sensitive formative decision frameworks.
Agreement between generative AI and subject-matter expert scores is necessary but insufficient to establish equivalent score meaning. Drawing on Kane’s scoring inference, I distinguish text-proximal from expertise-dependent tasks and argue that opaque model development threatens validity when score meaning depends on domain-specific judgment that score agreement alone cannot demonstrate.

up

pdf (full)
bib (full)
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress

This cognitive lab study examined how secondary students in Grades 7-12 interacted with an AI-enabled chatbot while completing online course activities. Using chat logs, facilitator observations, post-session reflections, and survey responses, we analyzed spontaneous and lightly seeded interactions to evaluate whether the chatbot supported its intended theory of action: helping students continue learning without simply completing work for them. Findings highlight the importance of typing fluency, answer-boundary handling, safety and privacy response patterns, and student-facing feedback mechanisms. This study illustrates how cognitive lab evidence can inform ongoing validation.
This work-in-progress study examines whether sentence embeddings can flag item pairs at risk for local dependence in multidimensional IRT. Using a 50-item personality inventory, preliminary analysis link semantic similarity to residual item-pair dependence. Future work will test robustness across larger samples, additional datasets, and alternative diagnostics.
This study evaluates few-shot large language models (LLMs) on middle-school geoscience responses (N=86 responses × 11 indicators), separating presence agreement from extract agreement. Joint prompting and annotation-like rubric guidance plus examples yield the clearest gains; lengthy rubric rewriting and clause-level parsing do not reliably improve extract alignment.
This study evaluates whether large language models (LLMs) can recover item difficulty estimates through Bradley-Terry modeling from pairwise comparisons of items. Across five Item Response Warehouse datasets and four LLMs, we examine alignment between LLM pairwise-derived difficulty rankings and 1PL IRT parameters, with implications for scalable, AI-assisted item calibration.
Case study that examines a large public Mexican university development of GenAI assessment guidelines for upper-secondary, undergraduate, and graduate education. Using document analysis and process tracing, it identifies shared governance principles, level-specific adaptations, and a policy-to-practice pathway emphasizing validity, fairness, transparency, human oversight, responsible AI use, and implementation evaluation.
This paper introduces TRACE, an automated framework for analyzing textual revisions across drafts. TRACE aligns initial and revised texts and infers interpretable revision operations, such as insert, delete, replace, and move. TRACE supports AI-powered revision analytics, writing assessment, automated feedback, and studies of human–AI writing collaboration.
We examine how strategies for representing student skill (𝜃) affect IRT-3PL item-parameter recovery from reconstructed item characteristic curves. We compare natural-language descriptors with signed decimal and scientific-notation anchors across anchor counts and placements. Our results show that dense numeric grids improve recovery, while non-uniform placements produce parameter-specific trade-offs, showing that 𝜃 representation matters.
This study evaluated an LLM for analyzing 1,406 Neurocritical Care examination comments coded for sentiment and thematic categories. Human raters showed high agreement, whereas LLM-human agreement was moderate. Thematic definitions reduced performance, but human-coded examples improved accuracy. LLMs may support preliminary coding, although human review remains necessary.
This study evaluated whether LLM- synthetic data can support K-12 instrument development. We generated LLM-synthetic datasets under various prompt conditions for two surveys and one assessment. Our findings indicate that while some prompt conditions successfully reproduced the overall latent structure of the instruments, recovery of item-level parameters was generally inadequate.
This paper presents a novel deep learning architecture for item response time prediction to create human readable feedback for content creators in educational assessment. The model employs a two-tier attention mechanism, mirroring the cognitive hierarchy of reading: the first tier identifies influential words nested within their sentence context, while the second tier evaluates the contribution of individual sentences to the overall item block. By leveraging this specific hierarchical structure, the model improves explainable AI (xAI) in educational measurement AI research. While architecturally simpler models can achieve comparable or superior raw predictive accuracy, the two-tier structure is designed specifically for xAI — to yield transparent, sentence- and world-level attribution that item developers can act on directly. This methodology bridges the gap between opaque, black-box predictive modeling and actionable, human interpretable feedback for instructional design and item development.
An ensemble automated scoring approach was evaluated against 260,000 responses spanning 55 constructed-response reading items. Performance was then compared to 27 human markers using 477 control scripts. The ensemble achieved perfect agreement on control scripts for 39 items (71%), matching or exceeding human marker performance and supporting operational quality assurance.
This work-in-progress compares 2,405 Spanish-language MCQs drafted by customized Gemini Gems or faculty in a Mexican university. Expert committees reviewed all items. Gemini drafts showed lower validation-friction in three of four areas and stronger blueprint adherence, while faculty drafts offered richer contextualization. Findings support structured AI generation with human validation.
This study considers measures commonly used to validate AES scoring and their limitations for indicating comparability with human rater scoring. In data simulated to reflect validation results achieved by AES national competition winners, scoring standard differences can occur across AES and human rater scoring (especially for 4-point scales vs. 2- and 3-point scales). Equipercentile methods are described and recommended for resolving the scoring differences.
This study compares two automated scoring approaches to identify communication behaviors in physician responses to patient questions: prompt-based scoring with Large Language Models (LLM) and supervised transformer-based models. Our results show that although transformer-based models provide more consistent performance, competitive LLM performance is promising in reducing the need for extensive annotated training data.
This design-focused study compared measurement-oriented ELL_RAG with conventional RAG for writing-assessment analytics. ELL_RAG showed a descriptive advantage overall and on global/population-level tasks, while local-evidence performance was nearly equivalent. Findings support task–evidence alignment rather than universal superiority and highlight tool orchestration as a continuing design challenge.
We built a near-comprehensive automated item evaluation model, predicting historical item acceptance or rejection from item text and Qwen3-generated critiques using 52,759 items from a large-scale testing program. Two DeBERTaV3 classifiers fused reached AUC .80 overall and .86 for math. Fairness-related rejections remained difficult, underscoring the need for human review.
We present a production-ready automated item generation pipeline integrated into a large-scale assessment program’s workflows. A one-shot approach prompts LLMs from an operational source item and existing guidelines, then populates metadata, screens quality, and uploads items for human review. Generated items passed expert review, and difficulty prompting reliably shifted difficulty.
This study evaluates GPT-5.4 Thinking simulations of reading-item responses using Skills Insight score-band descriptions and item content. Simulated and empirical item statistics and IRT parameters were compared. Difficulty recovery was strongest, with moderately strong b-parameter correlations, while a and c recovery was weaker, supporting preliminary difficulty evaluation before field testing.
Use of AI is bringing efficiencies in replacing previously laborious field- testing with scaled item parameter pre- calibration, especially in language assessment. However, gaps in research exist in replicating this accomplishment for mathematics assessment items with visual content. This study explores the methodological options of using AI in pre-calibrating multimodal mathematics items.
Item difficulty should, in theory, be predictable from features derived from RPLDs, task characteristics, and linguistic complexity. This study evaluates the performance of multiple statistical and machine learning models in estimating item difficulty. We found similar results across all four models, the item features selected explain 53%–56% of the variance across all grades.
We compare handcrafted features, frozen transformer embeddings, and ensembles across three item types and two regimes, testing pooled versus specialist models on a gaming detection task. Ensembles perform best (ROC-AUC 0.95, 𝜅=0.75), except for rarest item type under unseen prompts. Labels reflect review detections, motivating reference-conditioned recall and blind re-review.
Paradigms like IRT use measurement equations to model student behavior probabilistically. We investigate an analogous LLM-based approach, the noisy channel model, to compute and compare the conditional likelihoods of student writing and produce auditable token-level inferences about student skills. We compare against other LLM-based methods on accuracy, calibration, and interpretability.
Range performance level descriptors (RPLDs) connect assessment items to claims about what students at different achievement levels know and can do. Retrospectively assigning RPLDs to a large item bank is valuable but labor intensive. This work-in-progress study evaluates whether small and large language models can assist expert item-to-RPLD matching for 524 Grade 3–5 English language arts items. We compared direct classification with structured prompts that guide a model to analyze the knowledge, skills, evidence, and cognitive demand required by an item. We also examined self-consistency voting and a rater-informed prompting. Preliminary results indicate that larger hosted models achieved the closest overall performance to humans, although some locally hosted models produced comparable results. Voting improved prediction reliability but did not consistently increase agreement with human scores, whereas training models with human scores generally improved classification accuracy. Model performance also tended to decline as grade level increased.
Estimating item difficulty without collecting field testing data has been a long sought-after goal in educational measurement. We prompted large language models (LLMs) to compare mathematics items in pairs to estimate their difficulty. Results showed strong correlations between difficulty based on one LLM judge’s paired comparisons and empirical item difficulty.
Automated scoring of math-explanation items should handle responses that mix natural language with symbolic reasoning. The 2023 NAEP Automated Scoring Challenge showed that fine-tuning pre-trained encoders can reach human-like agreement, but its winning system. used a large, general-purpose encoder, while smaller math-specific pre-trained encoders remain a plausible and more efficient alternative. We fine-tune MathBERT (110M parameters) and DeBERTa-v3-large (183M parameters) as single-output regression scorers on 31 math-explanation items from one statewide summative assessment administration, benchmarking every model against human–human reliability and summarizing deployability with a three-tier acceptance criteria. On the 18 items fine-tuned under both encoders with identical datasets, the resulting performance between the two was practically indistinguishable: mean best QWK was 0.931 for MathBERT and 0.933 for DeBERTa-v3-large—a difference of only 0.002—and the models traded top performance at similar rates. Given similar performance, we tested a two-stage approach that used the larger general-purpose encoder only when the smaller math-focused encoder failed acceptance criteria. Extending MathBERT to all 31 items, it reached a mean QWK of 0.926 against a human–human benchmark of QWK = 0.940 and met operational acceptance criteria on 81% of items. Escalating the six items where a fine-tuned MathBERT model failed acceptance to a fine-tuned DeBERTa-v3-large model recovered four of the six, yielding acceptable automated-scoring models for 29 of 31 items. Because the smaller, math-pre-trained encoder matches the larger one at roughly 40% fewer parameters, we recommend fine-tuning a MathBERT encoder as the default and escalating to the larger, general purpose encoder like DeBERTa- v3-large only when a MathBERT model fails when deploying automated scoring models operationally at scale for math-explanation items.
Fine-tuning pre-trained transformer models for constructed-response items often begins with a hyperparameter grid search using k-fold cross-validation. We study how much that search actually helps by fine-tuning two encoders—MathBERT, a smaller math-focused model, and DeBERTa-v3-large, a larger general-purpose model—to score 18 math-explanation items from a state-wide assessment. Crossing learning rate, weight decay, and label smoothing over two epochs and five folds (1,800 model-fold runs), we compare each model’s tuning gain directly against fold-to-fold noise via a gain-to-noise ratio. Gains were small relative to fold noise for both encoders: MathBERT’s ratio fell below one (0.90), and DeBERTa’s nominally higher ratio (1.42) traced to a handful of divergent fits rather than an informative search landscape. Weight decay and label smoothing were effectively inert, leaving learning rate as the only hyperparameter worth checking—though even for learning rate the best configurations offered small performance gains over a reasonable default. We accordingly recommend a lean workflow that fixes the inert hyperparameters to sensible defaults, runs a narrow learning-rate search extended modestly upward, and increases the epoch budget with early stopping, substantially reducing training compute without sacrificing accuracy.
AI-generated and expert-created reading comprehension questions can show similar item statistics yet differ in the types of inferences required. This difference stemmed from AI’s failure to follow prompts during an intermediate item generation step. Evaluations of AI-generated items should document prompts and generation steps to identify and mitigate construct-relevant differences.
To automate spatial language classification, we fine-tuned DeBERTa-v3, a transformer-based language model, using a 33,284-word dataset following a 70:15:15 train-validation-test split. The model performed well in a held-out test comparing performance to human coders (Cohen’s kappa = .88).
We present a method for diagnostic assessment Q-matrix specification combining LLM input with model-based empirical validation. An LLM generates an initial Q-matrix from item content, and response data guide subsequent Q refinement. The approach integrates substantive rationale with empirical evidence to support scalable and measurement theory-grounded diagnostic assessment.
AI-generated feedback is often judged by quality ratings or preference rankings but rarely both. Comparing a multi-agent system, a single-agent system, and human feedback on student writing, we show that the two evaluation approaches can support different reported conclusions even when their underlying effects are nearly identical. These differences have consequences for how AI-generated feedback should be evaluated.
Large language models (LLMs) are increasingly used to score clinical communication, but validation is often limited to individual checklist items or total scores. These measures may not capture how communication behaviors occur together within a transcript. We therefore examine communication profiles: recurring combinations of behaviors that characterize different patterns of clinician communication and may support more targeted formative feedback. We analyzed 213 simulated respiratory OSCE transcripts rated on 15 binary Kalamazoo-derived communication items and compared final human ratings with GPT-4o, GPT-o3, and GPT-5.6 Sol. Raw item-level agreement was relatively high overall, but varied substantially across behaviors and was lower for several judgment-intensive items used for profile modeling. Bayesian latent-class analysis identified three stable human-derived profiles. Although the LLM-derived profiles showed broadly similar item-probability patterns, the models frequently assigned individual transcripts to different profiles than the human ratings. These findings show that agreement on individual communication skills does not necessarily translate into agreement on higher-level communication profiles. If LLMs are used to provide profile-based feedback, validation should therefore include profile-level agreement in addition to item-level performance.
This study examines the feasibility of using large language models (LLMs) as judges for pairwise comparisons of item difficulty and to the extent which the resulting comparison outcomes recover banked difficulty parameters. The study in particular fine-tunes an instruction-tuned LLM to evaluate the impact of task-specific fine-tuning on the parameter prediction accuracy. The study also demonstrates an agentic approach to hyperparameter tuning of LLM-training task.
This study evaluates an AI tutoring tool in a pre-licensure exam preparation product. Using residual gain modeling, elastic-net feature selection, and Gaussian Mixture clustering on LLM-derived cognitive fingerprints, the study identifies meaningful learner profiles and shows that the cognitive depth of student–AI interaction, not volume, drives measurable learning gain.
This study evaluated automatic prompt engineering (APE) using one assignment in the PERSUADE 2.0 dataset. The APE approach achieved higher QWK (.812) than the research-informed, zero-shot baseline prompting approach (.646). Descriptive comparisons examined gender and English language learner subgroups. Findings support the benefits of APE for automated essay scoring (AES).
This study evaluated GenAI-based medical assessment items against SME-developed items. Analyses of item content and response data from a high-stakes examination showed that their content quality and psychometric characteristics are comparable. These findings empirically support the quality of GenAI-based items and their potential use in educational and professional assessment programs.
Studies that validate LLM-predicted item difficulty conventionally benchmark predictions against a single-population estimate. Using a multimodal LLM’s pairwise judgments on a 29-item mathematics exam, mixture Rasch modeling shows the LLM tracks difficulty ordering differentially across classes, meaningful subgroup variation that the aggregate benchmark conceals entirely.
We evaluate whether open-weight LLMs can assess Computational Thinking in middle-school students’ finite-state game designs against human labels. Across a rich and a sparse design, models detect some behaviors at near-human agreement but over-credit absent ones and fail at counting, loop detection, and standards mapping. Failures trace to fixable setup.
We identify content areas where examinees underperform in a longitudinal assessment for a high-stakes medical licensure exam. Using modified funnel plots with a moving baseline to account for item difficulty, we rank topics by relative performance. Preliminary results suggest reasons beyond item difficulty, informing future analyses supporting potential educational interventions.
Fine-tuned models approached state-of-the-art agreement using small labelled sets. Open-weight models met operational criteria on seven of ten items, with a median of 50 responses per score point among passing items. A single marking exercise may supply enough data for automated short-answer scoring.
This study compared machine learning, transformer fine-tuning, prompt-based LLM, and novel ensemble approaches for automated essay scoring in a Canadian large-scale provincial assessment. Fine-tuned transformers achieved the highest reliability, followed by machine learning models. Meanwhile, prompt-based LLMs provided greater explainability, and ensemble architectures highlighted opportunities for balancing reliability and explainability.
Small-sample transformer scoring models ( n = 64 n=64) reach high human agreement but risk leaning on surface shortcuts like response length. Evaluating Mechanistic Interpretability strategies across 90 models, we show correlational methods suffer from seed noise, whereas interventional erasure proves small-sample models causally depend more on length. We outline an actionable audit protocol.
This paper demonstrates an application of Evidence-Centered Design (ECD) as a principled approach to design automated evaluations of AI-powered assessment outputs. We demonstrate this application through Khan Academy’s “Explain Your Thinking” conversational agent for mathematics items, showing how ECD’s layered models can be translated into rigorous and interpretable systems to demonstrate validity, reliability and fairness in AI systems to diverse audiences.
This study investigates effective use of AI models as statistics tutors. Specifically, how to promote accurate student misconception identification, elicit active student cognition, and provide trustworthy responses. Initial findings from explorations of prompting strategies and indicators of sycophantic behavior with emphasis on model reasoning traces in simulated conversations are discussed.
Small language models can run offline in a classroom, but they make mathematical errors no teacher can afford to miss. This work-in-progress proposes a measurement-grounded professional development model in which upper-elementary teachers use local LLMs to generate and critically evaluate items; an illustrative pilot demonstrates feasibility, teacher-impact study planned.
Collaboration is complex and multifaceted, blending cognitive, social, and emotional components that resist simple measurement. We present a framework linking what is measured, where, and how evidence is warranted. This paper synthesizes key design dimensions for AI collaboration partners and demonstrates how they embed measurement science into practice.
Using a year of authentic Grade 7–8 writing from two U.S. middle schools (8,576 documents, 423 students), we test whether keystroke features correlate with English Language Arts proficiency, form interpretable latent dimensions, and vary by gender. Fluency and engagement features reliably index proficiency by Grade 8, though generalizability varies.
Expert reviewers rated 156 calculus items from a self-revising multi-agent framework (Claude Sonnet 4.6 or GPT-5.4, three prompt conditions) and 26 human-written items. AI items were formative-ready nearly as often as human items (78–80% versus 88%) but summative-ready less often (26% and 10% versus 58%); a root-item reference mattered most.
This study investigates the performance of Neural Network-based IRT estimation in multistage adaptive language assessment with small item banks and short tests. Results showed that accuracy improved with larger samples and more training information. Moderate distribution shifts were largely accommodated under higher iteration conditions, whereas larger shifts continued to reduce estimation accuracy.
Explanatory item response models let test de- velopers anticipate item difficulty from item design, but they depend on subject-matter ex- perts rating every item on every hypothesized feature—a step that is slow, costly, and the practical bottleneck limiting how many items can be modeled. We ask whether large lan- guage models can supply those ratings. Four trained experts and six LLMs independently rated 15 grade-5 mathematics items on six psy- chometric features. We evaluate the LLM rat- ings twice: against the human consensus using quadratic-weighted 𝜅and mixed-effects mod- els, and against 1,452 student responses by us- ing each source’s ratings as the design matrix of a linear logistic test model (LLTM) bench- marked against a Rasch baseline. Claude Opus 4.7 performed best in recovering the item dif- ficulty (r = 0.89), however, the human con- sensus against which agreement is measured performs worst at recovering Rasch difficulty (r = 0.33) compared to the rest of the LLMs. The reason is visible in the expert panel it- self: inter-rater Fleiss’ 𝜅is at or below chance on three of six features, so the consensus is a noisy reference rather than ground truth. We also identify two failure modes (zero-variance and perfect collinearity) that agreement statis- tics cannot detect but make a feature unusable as an LLTM covariate. We argue that agree- ment should be reported alongside criterion- referenced recovery of item parameters, not in place of it.
We evaluated student (N = 1,361) interest and performance on student-AI co-authored math word problems. Performance matched or exceeded standard problems. Students rated peer-authored problems more often, especially when authorship was disclosed. Liking predicted first-attempt accuracy when problems required greater textual engagement, supporting interest-based context personalization.
Small language models (SLMs) are increasingly proposed for educational use because they promise lower cost, offline deployment, and stronger data privacy. We report an exploratory evaluation of three sub-2B-parameter models on example-based decimal-arithmetic tutoring. Across structured interactions, all three produced fluent, confident output that masked unstable pedagogy and frequent mathematical errors even on elementary decimal addition and place-value tasks. Building on these observations, we describe an emerging interaction-based measurement framework intended to support more defensible readiness decisions about SLMs as math tutors.
This study synthesizes feature-based and pairwise-comparison approaches to item parameter modeling by using an LLM as an investigator of item characteristics and a judge of relative item parameters. The proposed framework aims to improve recovery of item parameter rankings and predictions while providing insight into factors associated with the parameters.
Using an LLM-assisted pipeline, this study examines sociodemographic variation in English Language Learners’ hedge use across 6,500 essays. Regression analyses are expected to show that gender, race, and SES predict hedge density and category-specific variation, which subsequently predict writing scores. Beyond scaling annotation, the study evaluates whether LLM-derived metadiscourse indicators provide interpretable evidence about rhetorical development in ELL writing.
Code-editing language-model agents can change training choices beyond the hyper- parameter grids tested in automated essay scoring. Comparing these approaches requires distinguishing gains from a broader search space from evidence of a better search procedure. We compare code-editing agents with grid-restricted search, including random search, in two studies on ASAP-AES with nominally matched 12-hour search budgets. Code-editing produced the configuration with the highest test quadratic-weighted kappa (QWK) point estimate in each primary comparison. However, random search over a grid built afterwards around the Study 2 agent’s backbone, input length and head rule recovered most of its gain over BERT. Requiring a minimum validation-score improvement to accept a trial also left the accepted configuration’s validation score below the highest recorded valid validation score in every run that accepted a trial under this rule. The primary comparisons use one search run per condition, and some test folds reused essays involved in configuration selection, so these results compare selected configurations without establishing search-procedure superiority. These findings motivate reporting the available search choices and both accepted and best-scoring trials, and evaluating search procedures through repeated runs on data independent of configuration selection.
Large language models (LLMs) have been proposed as a substitute for field trials in itemdifficulty estimation, but evidence about where they succeed is usually drawn from a single test form. We compare LLM difficulty predictions across three task types that make different language demands, grammar cloze, semantic cloze and reading comprehension, on two operational Grade 9 English examinations sat by the same cohort in China (1,199 and1,166 examinees). Three LLMs estimated difficulty for 60 MCQ items under the same standardized protocol. On Form 1, grammar cloze had the highest Pearson correlation for all three models; the Fisher-averaged correlations were .81 for grammar cloze and .34 for semantic cloze. On Form 2, reading had the highest correlation for all three models (Fisheraveraged .78), and the grammar–semantic difference was smaller and inconsistent in direction. With ten items per task type, a task-type advantage seen on one form needs testing on other forms before it is read as a property of the model.
We report preliminary analyses from an ongoing study using performance-based tasks to assess science teachers’ generative AI (GenAI) literacy for classroom assessment and compare these scores with self-reported GenAI use and confidence. Preliminary findings indicate moderate correlation with use frequency but weak correlation with confidence, highlighting performance assessments’ value.
Synthetic assessment data may reproduce response distributions while failing to preserve underlying construct structure. Using a retrieval-augmented generation (RAG) pipeline and a stratified majority Black and Latinx student sample, we evaluated synthetic responses using distributional and psychometric measures. Results highlight the importance of psychometric preservation alongside distributional similarity.
Generating multiple-choice questions is increasingly scalable, but establishing their quality remains difficult. We review fourteen reports on automated item-writing flaw detection, revision, psychometric screening, and benchmark auditing. High accuracy often masks weak detection of flawed items, and revision evidence is mixed. We propose evaluating quality assurance as independently validated decisions.
Praise is one of the most frequent evaluative acts in teaching, but what a praise turn communicates depends less on its words than on how it is delivered. We propose measuring delivery as self-deviation, comparing a praise turn against two references drawn from the speaker themselves, the teacher’s other praise turns (Global Reference) and the speech immediately surrounding the turn (Local Reference).
This study examined whether an automated source integration measure responds to revision. In a randomized experiment, undergraduates drafted and revised a source-based essay. Source integration scores increased across drafts, although gains did not differ significantly between students who received source integration feedback and those who did not. Findings support use of the measure for formative evaluation and feedback.

up

pdf (full)
bib (full)
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers

Fine-grained formative writing feedback needs dense, standards-aligned labels that human annotation cannot supply at scale. We describe a multi-agent LLM pipeline producing verified silver labels, and then train small deterministic transformer scorers. On a Grade 5 pilot these reach Cohen’s 𝜅 up to 0.92 on conventions.
In this study, we marry automated scoring with automated rubric revision in an iterative, mutually-reinforcing process. One LLM agent scores student responses; a second, the Error Analysis Agent, examines human–engine scoring discrepancies and suggests rubric revisions. We show that this positive feedback loop improves automated scoring performance.
This study compares automated and human scores with expert/backread scores for writing, reading, and science tasks in a K-12 assessment program. Analyses examine discrepancies by score level and reused prompts across administrations. Automated scores showed smaller discrepancies for writing but larger discrepancies for short-answer tasks; reused prompts revealed localized shifts.
This study compares generative models Gemma and Qwen with ModernBERT and Mahalanobis-LOF for invalid-response detection in automated essay scoring. Qwen maintained high recall while reducing false invalid flags on a representative test set and detected some fluent off-topic responses in a matched diagnostic set. Fluent off-topic detection remained difficult.
Compact distillation preserved clean essay-scoring agreement across three student models and six compressed teacher targets, but the best target differed by model and margins were small. Perturbation, distribution-shift, and subgroup-bootstrap analyses dominated those differences, indicating that resource-constrained educational scoring requires uncertainty-aware evaluation rather than clean held-out comparison alone.
Compute cost in automated essay scoring is usually treated as engineering rather than measurement. This paper introduces the Cost-Aware Measurement Design (CAMD) framework, defines cost per defensible score (CPDS), and argues that AES systems should be compared by quality-cost sufficiency at a declared scoring volume.
Automated item generation with large language models (LLMs) is typically evaluated using aggregate accuracy metrics that conflate the quality of the source passage with noise in- troduced by the generation process itself. We address this gap by embedding a prospective i : (p×s×t) Generalizability Theory (G- theory) design into the evaluation of a QLoRA fine-tuned Qwen2.5-7B-Instruct model on the SciQ corpus. Passages (p) serve as the object of measurement; random seeds (s) and prompt templates (t) are fully crossed facets; items are nested within each (p,s,t) cell. Across 9,000 observations and five binary quality met- rics, we find that 77–81% of total variance is attributable to the passage, seed and template main effects are negligible (≤0.02%), and G- coefficients (E 𝜌2) uniformly exceed 0.97 un- der the observed design. D-study projections show that a single seed and template already achieves E 𝜌2 = 0.87, while the observed de- sign (ns = 3, nt = 3, ni = 2) reaches 0.98. Code, data, and R analysis scripts are released to support reproducible psychometric evalua- tion of future item-generation systems.
This study applied generalizability theory to examine score variation and dependability in AP Chinese speaking tasks completed with and without ChatGPT support. Although ChatGPT-supported tasks were associated with higher scores, the NoGPT condition consistently exhibited higher dependability coefficients. Increasing the numbers of tasks and raters further improved score dependability.
Enemy items are item pairs that must not appear on the same test form. An automatic enemy-identification method using large language models (LLMs) has been deployed in operation. This study asks a question: how is measurement quality sustained as conditions shift after deployment? Using operational data from certification exams, two analyses examine the factors that affect the method’s resilience. Analysis 1 isolates model-version and prompt updates: changing the model with the prompt held constant reduced recall from 0.75 to 0.54, while prompt refinement with the model held constant recovered it to 0.82. It also shows that most model–reviewer disagreements are edge cases and that the human standard is itself variable, with four experts spanning 0.70 to 0.91 in recall. Analysis 2 reports a disruption case: the established method works effectively on a professional-level exam but not on a foundational-level exam. A complementary content-tag based method was added to the existing method in response, with human review surfacing the disruption and validating the fix. Responsible LLM deployment in assessment requires monitoring with labeled data, testing before deployment, and evaluation systems tied to each use case, so that validity, reliability, and fairness are sustained rather than certified once.
This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.
Educational applications using conversational AI built on LLMs are hard to evaluate reliably because outputs vary unpredictably each turn, posing validity, reliability, fairness, and safety risks. We present a four-stage framework that combines automated and human testing to determine operational preparedness and continuously monitor deployed systems. The application of this resilience infrastructure framework is demonstrated in a case study of a conversation-based formative assessment tool for medical students to practice doctor-patient communication skills.
Digital-first assessments are delivered continuously and often remotely. Artificial intelligence (AI) enables digital assessments to generate content and administer tests at scale. Like any system used for high-stakes decision-making, digital assessments require responsible AI (RAI) practices to ensure fairness and validity. This paper presents two deployed systems in the Duolingo English Test (DET) lifecycle that align the DET to its RAI Standards. The Item Factory combines automated item generation with staged expert review; the Analytics for Quality Assurance in Test Taker (AQUA-TT) system applies unsupervised anomaly detection methods to continuously monitor for issues in digital test deliveries for daily individual test sessions. We present the design and performance of these systems and discuss what they imply for placing human judgment inside digital assessments.
We describe the development and evaluation of an AI conversational agent designed to interpret learning maps. Results indicate the agent can accurately identify map components and structures, especially when different informational formats are combined. Findings highlight the tool’s feasibility to inform the development of tools that use and interpret learning maps for teachers.
Artificial intelligence can scale assessment of student-generated scientific models, but valid educational use requires algorithms to identify features that meaningfully represent the intended learning construct. This study compares foundation and supervised computer vision approaches for evaluating learning progression (LP)-aligned evidence in approximately 1,600 high-school students’ electroscope models. Grounding DINO combined with the Segment Anything Model (SAM) was used for zero-shot detection without task-specific labeled training data, whereas a CustomCharge convolutional neural network (CNN) was trained on human-scored models. Both approaches were evaluated against expert scoring across 13 LP-aligned analytic categories. Grounding DINO+SAM achieved higher Cohen’s 𝜅 in 12 of 13 categories, with strong performance across both charge- and force-related evidence. Categories involving less frequent or more complex cross-scenario representations remained comparatively challenging. Findings demonstrate complementary strengths of foundation and supervised approaches and illustrate how LPs can provide a theoretically grounded framework for developing and validating AI assessment of scientific models.
We introduce a unified metric set for evaluating LLM tutors by inductively coding 177 recent literature-based metrics and scoring tutoring dialogues with three LLM judges. Exploratory factor analysis identified five constructs and a strong general factor. We present the resulting 10-category, 41-item framework for confirmatory analysis and human validation.
Automated short-answer scoring should reflect substantive content rather than linguistic form. Using controlled rewrites of 743 ASAP-SAS responses, we compare six fine-tuned encoders and 11 LLMs. Encoder scores increased with linguistic complexity, especially for lower-scoring responses, whereas LLMs showed heterogeneous patterns, revealing model-specific construct-irrelevant scoring signals.
Using the Facets Model for Severity and Centrality, we analyzed two human raters and 10 LLMs across four ASAP-SAS prompts. LLM centrality was task-dependent, and few-shot prompting reduced centrality inconsistently. Omitting scale-use differences altered severity estimates (r=.47), while MFRM fit diagnostics were harder to interpret in heterogeneous LLM rater pools.
Expert approval certifies that test developers would use AI-generated items, not that the scores support valid interpretations. I extend argument-based validity with a generation inference and a five-layer evidence framework, apply it diagnostically to three NAEP–Wilbur evaluations, and propose design principles and a minimal reporting standard for the field.
This study evaluated use of an AI-enabled item-generation tool for Reading item development for the National Assessment of Educational Progress (NAEP). Results were strongest when items targeted textually explicit information and weakest when passage and construct complexity increased. The tool showed potential for early-stage drafting rather than autonomous item development.
This study evaluates the effectiveness of an AI-based system (Wilbur) for generating NAEP mathematics items. Results from this investigation indicate moderate success for content-focused items but substantial challenges for items intended to assess the NAEP Mathematical Practices. Transitioning from GPT-4 to GPT-5.2 improved item quality, though human revision remained essential.
AI-enabled systems could transform item authoring for standardized assessments. We compare framework alignment and accuracy of NAEP Science items developed by two AI tools and human authors. Items from all sources had similar acceptance rates and most accepted items would require significant revision, underscoring the importance of evaluating AI-generated content.
We discuss how human-led Evidence-Centered Design (ECD) can inform the development of AI-supported automated item generation (AIG), as well as how it can be used as a quality-control mechanism. In this particular application, we describe how we collaborated with Claude Sonnet 5 to produce kinematics items together with interactive tools through an AIG pipeline. A preliminary qualitative analysis of a small number of generated items suggested they have identifiable strengths and weaknesses, which we discuss in depth per item.
This paper describes a framework for human–AI collaboration in educational measurement that connects four workflow components. The framework integrates educational measurement and responsible AI principles, emphasizing AI tools as support for human expertise. Illustrative applications demonstrate how human–AI collaboration can strengthen assessment processes while maintaining measurement quality and integrity.
We present a scalable dual-agent architecture translating multi-source assessment data – integrating performance with process logs – into formative data insights. Decoupling classification from text generation, embedding expert rubrics, and optimizing latency enables rapid first-draft generation. These insights reveal underlying learning behaviors, facilitating targeted intervention without increasing teachers’ cognitive burden.
This paper examines a human–ML framework for large-scale qualitative data coding. Using a nationally representative sample of transcript data, the study evaluates semantic embeddings and ranked recommendations through validation and user testing. Results demonstrate improved efficiency and accuracy while maintaining human expertise, oversight, and responsibility for final coding decisions.
This study examined whether rubric-aligned generative-AI features could augment established linguistic features in trait-based automated essay scoring. Features from both sources showed meaningful associations with human scores and only partial overlap with one another. Scoring models combining both feature sets produced modest improvements that varied across traits and evaluation metrics.
Using PERSUADE 2.0 discourse segments, this study examines whether effectiveness labels carry stable linguistic meaning across discourse types. Traditional NLP features and latent components show robust type-by-effectiveness interactions. Length decomposition indicates much of this signal reflects elaboration, while smaller length-robust patterns remain, cautioning against context-free diagnostic feedback in automated writing evaluation.
Generative AI complicates a core assumption of assessment that observed performance reflects an individual’s own cognition. Using StudyChat, we show AI supply only moderately tracks student intent, and assignment scores are largely insensitive to either. With the disruption of validity warrant, response process validity needs reconceptualizing when AI co-produces performance.
Using the LBIDAT protocol, we evaluated items generated by three LLMs (Claude, Gemini, GPT) across six zero-shot prompting tiers for 8th-grade standards. Neither prompt tier nor model affected defect severity. ELA items were consistently poor; mathematics items passed the low bar while falling short of appropriate grade-level cognitive complexity.
Does grade awareness bias automatically generated feedback? We present a multi-agent artificial intelligence (AI) framework that contrasts score-blind and score-informed evaluators and reconciles their critiques. Across 908 feedback blocks, score exposure systematically biased the model’s critique: harsher at low grades, warmer at high (halo/horn effects). Isolating a score-blind judgment improved score-alignment without sacrificing faithfulness.
As LLMs increasingly streamline item generation, ensuring the quality of their outputs remains a critical challenge. This study examines whether a multi-agent system can serve as an automated judge to detect and revise flaws in AI-generated educational assessment items, such as ambiguity, bias, or content misalignment.
Generative AI enables real-time and scalable context personalization of mathematics word problems (MWPs) based on students’ self-reported interests during assessment. However, a question arises: are personalized and standard MWPs mathematically equivalent? In this paper, we present two Natural Language Processing pipelines for evaluating mathematical equivalence between personalized and standard MWPs.
This study integrates assessment data with learning maps, using a neural network to recommend individualized next skills for instruction and assessment. Results indicate high-certainty, expert-validated recommendations, with empirical support for foundational assumptions. This demonstrates the potential for AI to support more personalized learning-maps-based instruction and assessment.
We developed and evaluated a prototype conversational agent for teachers that can support their understanding and use of mastery-based score reports that describe a students’ academic performance. The agent prototype delivered grounded responses reflecting intended interpretation and uses, however some responses were inaccurate, signaling aspects that can be improved.
This work proposes automated generation and psychometric scoring of Maze comprehension assessments using large language models (LLMs) and a multilevel item response theory (multilevel IRT) framework. Findings show reliable ability estimates across passages and items, offering scalable, curriculum-aligned formative assessment that reduces teacher workload and supports targeted reading instruction.