To Flag or Not to Flag? Detecting Concerning Content in Situational Judgment

Susha Suresh, Cole Walsh, Rodica Ivan, Colleen Robb


Abstract
This study evaluates automated approaches for detecting concerning content in open-response SJTs used in higher education admissions. Comparing fine-tuned BERT models with zero-shot and fine-tuned LLMs, we found that fine-tuned BERT achieved the strongest performance despite not receiving the scenario context available to the LLMs
Anthology ID:
2026.aimecon-main.50
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
450–457
Language:
URL:
https://aclanthology.org/2026.aimecon-main.50/
DOI:
Bibkey:
Cite (ACL):
Susha Suresh, Cole Walsh, Rodica Ivan, and Colleen Robb. 2026. To Flag or Not to Flag? Detecting Concerning Content in Situational Judgment. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 450–457, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
To Flag or Not to Flag? Detecting Concerning Content in Situational Judgment (Suresh et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.50.pdf