Kang Xue
Author directory2026
Predicting Item-to-Range Performance Level Descriptor Matches with Structured Language Model Reasoning
Kang Xue | M. Christina Schneider | Sangdon Lim | Garron Gianopulos | Alexandra Perez
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Kang Xue | M. Christina Schneider | Sangdon Lim | Garron Gianopulos | Alexandra Perez
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Range performance level descriptors (RPLDs) connect assessment items to claims about what students at different achievement levels know and can do. Retrospectively assigning RPLDs to a large item bank is valuable but labor intensive. This work-in-progress study evaluates whether small and large language models can assist expert item-to-RPLD matching for 524 Grade 3–5 English language arts items. We compared direct classification with structured prompts that guide a model to analyze the knowledge, skills, evidence, and cognitive demand required by an item. We also examined self-consistency voting and a rater-informed prompting. Preliminary results indicate that larger hosted models achieved the closest overall performance to humans, although some locally hosted models produced comparable results. Voting improved prediction reliability but did not consistently increase agreement with human scores, whereas training models with human scores generally improved classification accuracy. Model performance also tended to decline as grade level increased.
Evaluating Multiple Models for Predicting Item Difficulty in a Principled Assessment Design
Alexandra Lane Perez | Christina Schneider | Sangdon Lim | Garron Gianopulos | Kang Xue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Alexandra Lane Perez | Christina Schneider | Sangdon Lim | Garron Gianopulos | Kang Xue
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Item difficulty should, in theory, be predictable from features derived from RPLDs, task characteristics, and linguistic complexity. This study evaluates the performance of multiple statistical and machine learning models in estimating item difficulty. We found similar results across all four models, the item features selected explain 53%–56% of the variance across all grades.
2020
Predicting the Difficulty and Response Time of Multiple Choice Questions Using Transfer Learning
Kang Xue | Victoria Yaneva | Christopher Runyon | Peter Baldwin
Proceedings of the Fifteenth Workshop on Innovative Use of NLP for Building Educational Applications
Kang Xue | Victoria Yaneva | Christopher Runyon | Peter Baldwin
Proceedings of the Fifteenth Workshop on Innovative Use of NLP for Building Educational Applications
This paper investigates whether transfer learning can improve the prediction of the difficulty and response time parameters for 18,000 multiple-choice questions from a high-stakes medical exam. The type the signal that best predicts difficulty and response time is also explored, both in terms of representation abstraction and item component used as input (e.g., whole item, answer options only, etc.). The results indicate that, for our sample, transfer learning can improve the prediction of item difficulty when response time is used as an auxiliary task but not the other way around. In addition, difficulty was best predicted using signal from the item stem (the description of the clinical case), while all parts of the item were important for predicting the response time.