Hotaka Maeda
Author directory2026
Production-Ready Automated Item Generation in Educational Assessment: Integration with Operational Workflows
Hotaka Maeda | Kargi Chauhan | Tharunya Chandrashekar
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Hotaka Maeda | Kargi Chauhan | Tharunya Chandrashekar
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
We present a production-ready automated item generation pipeline integrated into a large-scale assessment program’s workflows. A one-shot approach prompts LLMs from an operational source item and existing guidelines, then populates metadata, screens quality, and uploads items for human review. Generated items passed expert review, and difficulty prompting reliably shifted difficulty.
Automated Item Evaluation: Predicting Item Acceptance and Rejection using LLM-Generated Critiques
Hotaka Maeda | Yikai Lu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Hotaka Maeda | Yikai Lu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
We built a near-comprehensive automated item evaluation model, predicting historical item acceptance or rejection from item text and Qwen3-generated critiques using 52,759 items from a large-scale testing program. Two DeBERTaV3 classifiers fused reached AUC .80 overall and .86 for math. Fairness-related rejections remained difficult, underscoring the need for human review.