Richard Patz
Author directory2026
Self-Revising Agents as Item Writers: Benchmarking Against Professional Item Writing
Steven Tang | Zhen Li | Richard Patz
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Steven Tang | Zhen Li | Richard Patz
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Expert reviewers rated 156 calculus items from a self-revising multi-agent framework (Claude Sonnet 4.6 or GPT-5.4, three prompt conditions) and 26 human-written items. AI items were formative-ready nearly as often as human items (78–80% versus 88%) but summative-ready less often (26% and 10% versus 58%); a root-item reference mattered most.