Christopher Ormerod
Author directory2026
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Detecting Invalid Responses in Automated Essay Scoring with Fine-Tuned Large Language Models
YoungKoung Kim | Christopher Ormerod
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
YoungKoung Kim | Christopher Ormerod
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
This study compares generative models Gemma and Qwen with ModernBERT and Mahalanobis-LOF for invalid-response detection in automated essay scoring. Qwen maintained high recall while reducing false invalid flags on a representative test set and detected some fluent off-topic responses in a matched diagnostic set. Fluent off-topic detection remained difficult.
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Generative Language Models for Argumentation Annotation
Kai North | Christopher Ormerod
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Kai North | Christopher Ormerod
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study explores the integration of a generative language model (GLM) into an Automated Writing Evaluation (AWE) system designed to highlight both argumentative components and errors in spelling and grammar. We fine-tune an open-source GLM using parameter-efficient techniques. We evaluate the model’s capabilities in both argument analysis and error detection against established datasets. We demonstrate that a single GLM with parameter-efficient adapters can accurately identify argumentative clauses, classify their types, map relationships between them, and flag mechanical mistakes. We establish that our AWE system performs at human-level accuracy, while only requiring a fraction of the computational power of much larger models.
Math Item Difficulty Prediction with Multimodal Input
Kai North | Christopher Ormerod | Alexander Kwako | Suhwa Han
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Kai North | Christopher Ormerod | Alexander Kwako | Suhwa Han
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Item difficulty prediction relies on the efficient utilization of high-leverage item characteristics. Many items in the domain of mathematics include figures, charts, or other visual stimuli that are challenging to incorporate into difficulty prediction models. Multimodal large language models (LLMs) offer a way to process these visual stimuli in combination with text input, potentially enhancing the success of item difficulty prediction. In this study, we employ several open-source multimodal LLMs to predict the difficulty of math items with visual stimuli from a publicly available data set. We find that multimodal LLMs are capable of predicting math item difficulty.
CAP: Cross-Ability Preferences to Align Student Simulators for Item Difficulty Prediction
Nigel Fernandez | Alexander Scarlatos | Christopher Ormerod | Andrew Lan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Nigel Fernandez | Alexander Scarlatos | Christopher Ormerod | Andrew Lan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
We develop CAP, a preference optimization-based method for aligning LLM student simulators to student ability. It constructs preference pairs from IRT ability gaps, training the simulator to generate responses that better reflect prompted ability, enabling simulation-based estimation of open-ended item difficulty.
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting-Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
2025
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
Alexander Scarlatos | Nigel Fernandez | Christopher Ormerod | Susan Lottridge | Andrew Lan
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Alexander Scarlatos | Nigel Fernandez | Christopher Ormerod | Susan Lottridge | Andrew Lan
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Item (question) difficulties play a crucial role in educational assessments, enabling accurate and efficient assessment of student abilities and personalization to maximize learning outcomes. Traditionally, estimating item difficulties can be costly, requiring real students to respond to items, followed by fitting an item response theory (IRT) model to get difficulty estimates. This approach cannot be applied to the cold-start setting for previously unseen items either. In this work, we present SMART (Simulated Students Aligned with IRT), a novel method for aligning simulated students with instructed ability, which can then be used in simulations to predict the difficulty of open-ended items. We achieve this alignment using direct preference optimization (DPO), where we form preference pairs based on how likely responses are under a ground-truth IRT model. We perform a simulation by generating thousands of responses, evaluating them with a large language model (LLM)-based scoring model, and fit the resulting data to an IRT model to obtain item difficulty estimates. Through extensive experiments on two real-world student response datasets, we show that SMART outperforms other item difficulty prediction methods by leveraging its improved ability alignment.
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Joshua Wilson | Christopher Ormerod | Magdalen Beiting Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Joshua Wilson | Christopher Ormerod | Magdalen Beiting Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
Christopher Ormerod
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Christopher Ormerod
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
This study illustrates how incorporating feedback-oriented annotations into the scoring pipeline can enhance the accuracy of automated essay scoring (AES). This approach is demonstrated with the Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements (PERSUADE) corpus. We integrate two types of feedback-driven annotations: those that identify spelling and grammatical errors, and those that highlight argumentative components. To illustrate how this method could be applied in real-world scenarios, we employ two LLMs to generate annotations – a generative language model used for spell correction and an encoder-based token-classifier trained to identify and mark argumentative elements. By incorporating annotations into the scoring process, we demonstrate improvements in performance using encoder-based large language models fine-tuned as classifiers.
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Long context Automated Essay Scoring with Language Models
Christopher Ormerod | Gitit Kehat
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Christopher Ormerod | Gitit Kehat
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
In this study, we evaluate several models that incorporate architectural modifications to overcome the length limitations of the standard transformer architecture using the Kaggle ASAP 2.0 dataset. The models considered in this study include fine-tuned versions of XLNet, Longformer, ModernBERT, Mamba, and Llama models.
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Joshua Wilson | Christopher Ormerod | Magdalen Beiting Parrish
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
2024
Can Language Models Guess Your Identity? Analyzing Demographic Biases in AI Essay Scoring
Alexander Kwako | Christopher Ormerod
Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024)
Alexander Kwako | Christopher Ormerod
Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024)
Large language models (LLMs) are increasingly used for automated scoring of student essays. However, these models may perpetuate societal biases if not carefully monitored. This study analyzes potential biases in an LLM (XLNet) trained to score persuasive student essays, based on data from the PERSUADE corpus. XLNet achieved strong performance based on quadratic weighted kappa, standardized mean difference, and exact agreement with human scores. Using available metadata, we performed analyses of scoring differences across gender, race/ethnicity, English language learning status, socioeconomic status, and disability status. Automated scores exhibited small magnifications of marginal differences in human scoring, favoring female students over males and White students over Black students. To further probe potential biases, we found that separate XLNet classifiers and XLNet hidden states weakly predicted demographic membership. Overall, results reinforce the need for continued fairness analyses as use of LLMs expands in education.