Rodica Ivan

Author directory

2026

This study evaluates automated approaches for detecting concerning content in open-response SJTs used in higher education admissions. Comparing fine-tuned BERT models with zero-shot and fine-tuned LLMs, we found that fine-tuned BERT achieved the strongest performance despite not receiving the scenario context available to the LLMs
This study evaluates synthetic open-response data generation for assessment development by using real response data as ground truth. It finds that conditioning LLM generation on actual test-taker responses improves fidelity and diversity, though synthetic responses still lack the full variation of real test-takers, limiting current psychometric utility.

2025

Current methods for assessing personal and professional skills lack scalability due to reliance on human raters, while NLP-based systems for assessing these skills fail to demonstrate construct validity. This study introduces a new method utilizing LLMs to extract construct-relevant features from responses to an assessment of personal and professional skills.