Rodica Ivan
Author directory2026
To Flag or Not to Flag? Detecting Concerning Content in Situational Judgment
Susha Suresh | Cole Walsh | Rodica Ivan | Colleen Robb
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Susha Suresh | Cole Walsh | Rodica Ivan | Colleen Robb
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study evaluates automated approaches for detecting concerning content in open-response SJTs used in higher education admissions. Comparing fine-tuned BERT models with zero-shot and fine-tuned LLMs, we found that fine-tuned BERT achieved the strongest performance despite not receiving the scenario context available to the LLMs
It Quacks Like a Response: Evaluating Synthetically-Generated Responses in Assessment Research
Cole Walsh | Rodica Ivan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Cole Walsh | Rodica Ivan
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study evaluates synthetic open-response data generation for assessment development by using real response data as ground truth. It finds that conditioning LLM generation on actual test-taker responses improves fidelity and diversity, though synthetic responses still lack the full variation of real test-takers, limiting current psychometric utility.
2025
Using LLMs to identify features of personal and professional skills in an open-response situational judgment test
Cole Walsh | Rodica Ivan | Muhammad Zafar Iqbal | Colleen Robb
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Cole Walsh | Rodica Ivan | Muhammad Zafar Iqbal | Colleen Robb
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Current methods for assessing personal and professional skills lack scalability due to reliance on human raters, while NLP-based systems for assessing these skills fail to demonstrate construct validity. This study introduces a new method utilizing LLMs to extract construct-relevant features from responses to an assessment of personal and professional skills.