Marcus Walker

Author directory

2026

We identify content areas where examinees underperform in a longitudinal assessment for a high-stakes medical licensure exam. Using modified funnel plots with a moving baseline to account for item difficulty, we rank topics by relative performance. Preliminary results suggest reasons beyond item difficulty, informing future analyses supporting potential educational interventions.
This study evaluated GenAI-based medical assessment items against SME-developed items. Analyses of item content and response data from a high-stakes examination showed that their content quality and psychometric characteristics are comparable. These findings empirically support the quality of GenAI-based items and their potential use in educational and professional assessment programs.
This study identifies a selected transformer-based clustering model for operational medical assessment items that balances within blueprint-topic proximity with cluster separation. Operational analyses showed supplementary alignment with blueprint structure, flagged isolated content areas that may need additional item coverage, and identified dispersed topics for subject matter expert (SME) review.