Yikai Lu
Author directory2026
Automated Item Evaluation: Predicting Item Acceptance and Rejection using LLM-Generated Critiques
Hotaka Maeda | Yikai Lu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Hotaka Maeda | Yikai Lu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
We built a near-comprehensive automated item evaluation model, predicting historical item acceptance or rejection from item text and Qwen3-generated critiques using 52,759 items from a large-scale testing program. Two DeBERTaV3 classifiers fused reached AUC .80 overall and .86 for math. Fairness-related rejections remained difficult, underscoring the need for human review.