Jiyun Zu
Author directory2026
Responsible use of generative AI when creating reading comprehension questions: Inference matters
Zuowei Wang | Michael Flor | Jiyun Zu | Tenaha O’Reilly | Wanjing Anya Ma
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Zuowei Wang | Michael Flor | Jiyun Zu | Tenaha O’Reilly | Wanjing Anya Ma
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
AI-generated and expert-created reading comprehension questions can show similar item statistics yet differ in the types of inferences required. This difference stemmed from AI’s failure to follow prompts during an intermediate item generation step. Evaluations of AI-generated items should document prompts and generation steps to identify and mitigate construct-relevant differences.
Fine-tuning Large Language Models for Automated Scoring: Classification, Regression, vs. Ordinal Regression
Jiyun Zu | Akshay Badola
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Jiyun Zu | Akshay Badola
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Fine-tuning large language models for automated scoring is often formulated as either a regression or a classification task. However, scores assigned by human raters are on an ordinal scale. We summarize different deep-learning ordinal regression methods and compare their performances with those from regression and classification using a real dataset.
2025
Exploring the Interpretability of AI-Generated Response Detection with Probing
Ikkyu Choi | Jiyun Zu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Ikkyu Choi | Jiyun Zu
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Multiple strategies for AI-generated response detection have been proposed, with many high-performing ones built on language models. However, the decision-making processes of these detectors remain largely opaque. We addressed this knowledge gap by fine-tuning a language model for the detection task and applying probing techniques using adversarial examples. Our adversarial probing analysis revealed that the fine-tuned model relied heavily on a narrow set of lexical cues in making the classification decision. These findings underscore the importance of interpretability in AI-generated response detectors and highlight the value of adversarial probing as a tool for exploring model interpretability.
Effects of Generation Model on Detecting AI-generated Essays in a Writing Test
Jiyun Zu | Michael Fauss | Chen Li
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Jiyun Zu | Michael Fauss | Chen Li
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Various detectors have been developed to detect AI-generated essays using labeled datasets of human-written and AI-generated essays, with many reporting high detection accuracy. In real-world settings, essays may be generated by models different from those used to train the detectors. This study examined the effects of generation model on detector performance. We focused on two generation models – GPT-3.5 and GPT-4 – and used writing items from a standardized English proficiency test. Eight detectors were built and evaluated. Six were trained on three training sets (human-written essays combined with either GPT-3.5-generated essays, or GPT-4-generated essays, or both) using two training approaches (feature-based machine learning and fine-tuning RoBERTa), and the remaining two were ensemble detectors. Results showed that a) fine-tuned detectors outperformed feature-based machine learning detectors on all studied metrics; b) detectors trained with essays generated from only one model were more likely to misclassify essays generated by the other model as human-written essays (false negatives), but did not misclassify more human-written essays as AI-generated (false positives); c) the ensemble fine-tuned RoBERTa detector had fewer false positives, but slightly more false negatives than detectors trained with essays generated by both models.