Evelyn Johnson

Author directory

2026

This study evaluated whether visual embeddings from SigLIP and DINOv2 could be used to simulate responses to Figure Matrices items. Models predicted response probabilities using item screenshots and examinee ability estimates. Results showed moderate probability recovery but limited item difficulty recovery, indicating promise for early screening but not calibration replacement.

2025

This work-in-progress study compares the accuracy of machine learning and large language models to predict student responses to field-test items on a social-emotional learning assessment. We evaluate how well each method replicates actual responses and examine the item parameters generated by synthetic data to those derived from actual student data.
This study explores the use of large language models to simulate human responses to Likert-scale items. A DeBERTa-base model fine-tuned with item text and examinee ability emulates a graded response model (GRM). High alignment with GRM probabilities and reasonable threshold recovery support LLMs as scalable tools for early-stage item evaluation.