Marcus Walker
Author directory2026
Funnel Plot Analysis of Examinee Performance in Longitudinal Assessment
Aquia Richburg | Marcus Walker
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Aquia Richburg | Marcus Walker
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
We identify content areas where examinees underperform in a longitudinal assessment for a high-stakes medical licensure exam. Using modified funnel plots with a moving baseline to account for item difficulty, we rank topics by relative performance. Preliminary results suggest reasons beyond item difficulty, informing future analyses supporting potential educational interventions.
A Comprehensive Evaluation of GenAI-based Items in a Medical Examination
Yanlin Jiang | Marcus Walker | Andrew Dallas | Aquia Richburg | Nikole Gregg | Brittany Corrigan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Yanlin Jiang | Marcus Walker | Andrew Dallas | Aquia Richburg | Nikole Gregg | Brittany Corrigan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
This study evaluated GenAI-based medical assessment items against SME-developed items. Analyses of item content and response data from a high-stakes examination showed that their content quality and psychometric characteristics are comparable. These findings empirically support the quality of GenAI-based items and their potential use in educational and professional assessment programs.
Auxiliary Information for Semantic Clustering of Assessment Items
Josiah Hunsberger | Aquia Richburg | Marcus Walker
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Josiah Hunsberger | Aquia Richburg | Marcus Walker
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study identifies a selected transformer-based clustering model for operational medical assessment items that balances within blueprint-topic proximity with cluster separation. Operational analyses showed supplementary alignment with blueprint structure, flagged isolated content areas that may need additional item coverage, and identified dispersed topics for subject matter expert (SME) review.