Anna N. Rafferty

Author directory

Also published as: Anna Rafferty


2026

Paradigms like IRT use measurement equations to model student behavior probabilistically. We investigate an analogous LLM-based approach, the noisy channel model, to compute and compare the conditional likelihoods of student writing and produce auditable token-level inferences about student skills. We compare against other LLM-based methods on accuracy, calibration, and interpretability.
Black-box knowledge tracing models are commonly deployed to drive personalization in online learning platforms and are typically evaluated using classification metrics such as AUC, accuracy, and F1 score. However, a model that predicts item responses using only each item’s proportion correct in the training set achieves AUC up to 0.72 and accuracy up to 0.84 on widely used benchmark datasets, despite using no information about individual students’ response histories. Inspired by the concept of marginal reliability in psychometrics, we introduce Fractional Information Gain (FIG), an information-theoretic evaluation metric for black-box predictive models of student item responses. FIG measures the fraction of a student’s response uncertainty resolved by a trained model relative to the item-only baseline. FIG equals 0 when the model adds no information beyond item base rates, and 1 when held-out responses are perfectly predicted. FIG is applicable to any model that outputs probabilities and is sensitive to calibration errors. We characterize FIG on synthetic and real data, compare it to AUC for both the item-only baseline and trained models on four benchmark datasets, and describe the operational affordances that FIG inherits from its reliability-like construction. FIG imports the conceptual benefits of score reliability into the prediction-oriented framework of ML-driven educational systems.

2020

We train neural machine translation (NMT) models from English to six target languages, using NMT encoder representations to predict ancestor constituent labels of source language words. We find that NMT encoders learn similar source syntax regardless of NMT target language, relying on explicit morphosyntactic cues to extract syntactic features from source sentences. Furthermore, the NMT encoders outperform RNNs trained directly on several of the constituent label prediction tasks, suggesting that NMT encoder representations can be used effectively for natural language tasks involving syntax. However, both the NMT encoders and the directly-trained RNNs learn substantially different syntactic information from a probabilistic context-free grammar (PCFG) parser. Despite lower overall accuracy scores, the PCFG often performs well on sentences for which the RNN-based models perform poorly, suggesting that RNN architectures are constrained in the types of syntax they can learn.

2011

2009

2008