Ummugul Bezirhan
Author directory2026
Lost Without Translation? Multilingual Sentence Embeddings for Linguistic-integrated Reliability Auditing
Ummugul Bezirhan | Ji Yoon Jung | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Ummugul Bezirhan | Ji Yoon Jung | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual sentence embeddings can replace translated English input for Linguistic-Integrated Reliability Auditing (LiRA) across 11 PIRLS constructed-response items and three embedding models. Native-language embeddings reproduced translation-based reliability estimates closely while recovering responses excluded after translation failure, with no meaningful change in reliability.
Anchored Bradley-Terry Calibration Using LLM Comparative Judgments
Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
AI-based comparative judgment was evaluated as a tool for early item difficulty estimation. Three AI judges compared new TIMSS Grade 4 mathematics items with calibrated anchors, with rankings analyzed using a fixed-anchor Bradley-Terry model. Results show meaningful difficulty signals, supporting scalable supplementary use while highlighting anchor coverage and comparison-network design.
Comprehensive Pipeline for Multilingual GenAI Scoring
Ji Yoon Jung | Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Ji Yoon Jung | Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study presents a comprehensive operational pipeline for multilingual GenAI scoring that integrates secure data preparation, prompt generation, automated scoring, post-processing, and human-in-the-loop review. Results demonstrate the pipeline’s strong adaptability across item formats, assessment languages, and assessment cycles, offering a viable solution to the persistent challenges of multilingual scoring.
2025
Optimizing Reliability Scoring for ILSAs
Ji Yoon Jung | Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Ji Yoon Jung | Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study proposes an innovative method for evaluating cross-country scoring reliability (CCSR) in multilingual assessments, using hyperparameter optimization and a similarity-based weighted majority scoring within a single human scoring framework. Results show that this approach provides a cost-effective and comprehensive assessment of CCSR without the need for additional raters.
AI-Based Classification of TIMSS Items for Framework Alignment
Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Large-scale assessments rely on expert panels to verify that test items align with prescribed frameworks, a labor-intensive process. This study evaluates the use of GPT-4o to classify TIMSS items to content domain, cognitive domain, and difficulty categories. Findings highlight the potential of language models to support scalable, framework-aligned item verification.
Input Optimization for Automated Scoring in Reading Assessment
Ji Yoon Jung | Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Ji Yoon Jung | Ummugul Bezirhan | Matthias von Davier
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
This study examines input optimization for enhanced efficiency in automated scoring (AS) of reading assessments, which typically involve lengthy passages and complex scoring guides. We propose optimizing input size using question-specific summaries and simplified scoring guides. Findings indicate that input optimization via compression is achievable while maintaining AS performance.