Mohammad Hassany
Author directory2026
Automated Knowledge Component Generation and Interpretable Knowledge Tracing in Coding Problems
Zhangqi Duan | Nigel Fernandez | Arun Balajiee Lekshmi Narayanan | Mohammad Hassany | Rafaella Sampaio de Alencar | Peter Brusilovsky | Bita Akram | Andrew Lan
Findings of the Association for Computational Linguistics: ACL 2026
Zhangqi Duan | Nigel Fernandez | Arun Balajiee Lekshmi Narayanan | Mohammad Hassany | Rafaella Sampaio de Alencar | Peter Brusilovsky | Bita Akram | Andrew Lan
Findings of the Association for Computational Linguistics: ACL 2026
Knowledge components (KCs) are key to assessing student knowledge levels on fine-grained skills and driving personalization and feedback. However, crafting KCs and tagging them for problems, traditionally performed by human domain experts, is highly labor-intensive. Prior work has studied automated KC generation only for multiple-choice questions but not open-ended ones. We bridge this gap and present an automated, large language model (LLM)-based pipeline for KC generation and tagging for open-ended programming problems. We also develop an LLM-based knowledge tracing (KT) framework to leverage these LLM-generated KCs. We conduct extensive quantitative and qualitative evaluations on two real-world student code submission datasets. Results show that our KT method outperforms existing ones and LLM-generated KCs outperform human-written KCs on future student response prediction. We also investigate how these KCs enable us to analyze student learning curves and conduct human evaluation with course instructors to further verify the quality of KC-problem tagging.
Can LLMs Replace Semantic Similarity for Scoring Student Code Explanations?
Arun Balajiee Lekshmi Narayanan | Mohammad Hassany | Peter Brusilovsky
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Arun Balajiee Lekshmi Narayanan | Mohammad Hassany | Peter Brusilovsky
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
In programming courses, students are often asked to explain code fragments, which is a way to assess their understanding of programming constructs and patterns. These types of problem known as “explain in plain English” are valuable in both assessment and practice contexts. The main challenge to using these types of problems at scale in both contexts is automating the scoring process, i.e., assessing whether those explanations are correct. The prevailing approach scores an explanation by its semantic similarity to an instructor’s model explanation, but this raises a measurement concern: students who reason correctly yet phrase their explanations differently from an expert may be scored as incorrect (false negatives), threatening the validity and fairness of the assessment. Given recent advances in LLM-based automated scoring, it remains unclear whether semantic similarity methods are still the most effective technique for automatically scoring free-form student responses, such as code explanations. In this paper, we present a rigorous comparison between LLMs and semantic similarity approaches for the automated scoring of student code explanations, using an open dataset Selfcode 2.0. We frame the scoring as a binary classification task and use generative AI to balance the dataset. Our results suggest that LLM-based scoring (F1 = 0.98, accuracy = 0.96) outperform semantic similarity scoring (F1 = 0.72, accuracy = 0.65). It also produces fewer false negatives and eliminate the need to produce well-formulated model explanations.