@inproceedings{lekshmi-narayanan-etal-2026-llms,
title = "Can {LLM}s Replace Semantic Similarity for Scoring Student Code Explanations?",
author = "Lekshmi Narayanan, Arun Balajiee and
Hassany, Mohammad and
Brusilovsky, Peter",
editor = "Wilson, Joshua and
Ormerod, Christopher and
Beiting-Parrish, Magdalen",
booktitle = "Proceedings of the Artificial Intelligence in Measurement and Education Conference ({AIME}-Con): Full Papers",
month = oct,
year = "2026",
address = "Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States",
publisher = "National Council on Measurement in Education (NCME)",
url = "https://aclanthology.org/2026.aimecon-main.10/",
pages = "92--99",
ISBN = "979-8-9983004-0-0",
abstract = "In programming courses, students are often asked to explain code fragments, which is a way to assess their understanding of programming constructs and patterns. These types of problem known as ``explain in plain English'' are valuable in both assessment and practice contexts. The main challenge to using these types of problems at scale in both contexts is automating the scoring process, i.e., assessing whether those explanations are correct. The prevailing approach scores an explanation by its semantic similarity to an instructor{'}s model explanation, but this raises a measurement concern: students who reason correctly yet phrase their explanations differently from an expert may be scored as incorrect (false negatives), threatening the validity and fairness of the assessment. Given recent advances in LLM-based automated scoring, it remains unclear whether semantic similarity methods are still the most effective technique for automatically scoring free-form student responses, such as code explanations. In this paper, we present a rigorous comparison between LLMs and semantic similarity approaches for the automated scoring of student code explanations, using an open dataset Selfcode 2.0. We frame the scoring as a binary classification task and use generative AI to balance the dataset. Our results suggest that LLM-based scoring (F1 = 0.98, accuracy = 0.96) outperform semantic similarity scoring (F1 = 0.72, accuracy = 0.65). It also produces fewer false negatives and eliminate the need to produce well-formulated model explanations."
}<?xml version="1.0" encoding="UTF-8"?>
<modsCollection xmlns="http://www.loc.gov/mods/v3">
<mods ID="lekshmi-narayanan-etal-2026-llms">
<titleInfo>
<title>Can LLMs Replace Semantic Similarity for Scoring Student Code Explanations?</title>
</titleInfo>
<name type="personal">
<namePart type="given">Arun</namePart>
<namePart type="given">Balajiee</namePart>
<namePart type="family">Lekshmi Narayanan</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Mohammad</namePart>
<namePart type="family">Hassany</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Peter</namePart>
<namePart type="family">Brusilovsky</namePart>
<role>
<roleTerm authority="marcrelator" type="text">author</roleTerm>
</role>
</name>
<originInfo>
<dateIssued>2026-10</dateIssued>
</originInfo>
<typeOfResource>text</typeOfResource>
<relatedItem type="host">
<titleInfo>
<title>Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers</title>
</titleInfo>
<name type="personal">
<namePart type="given">Joshua</namePart>
<namePart type="family">Wilson</namePart>
<role>
<roleTerm authority="marcrelator" type="text">editor</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Christopher</namePart>
<namePart type="family">Ormerod</namePart>
<role>
<roleTerm authority="marcrelator" type="text">editor</roleTerm>
</role>
</name>
<name type="personal">
<namePart type="given">Magdalen</namePart>
<namePart type="family">Beiting-Parrish</namePart>
<role>
<roleTerm authority="marcrelator" type="text">editor</roleTerm>
</role>
</name>
<originInfo>
<publisher>National Council on Measurement in Education (NCME)</publisher>
<place>
<placeTerm type="text">Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States</placeTerm>
</place>
</originInfo>
<genre authority="marcgt">conference publication</genre>
<identifier type="isbn">979-8-9983004-0-0</identifier>
</relatedItem>
<abstract>In programming courses, students are often asked to explain code fragments, which is a way to assess their understanding of programming constructs and patterns. These types of problem known as “explain in plain English” are valuable in both assessment and practice contexts. The main challenge to using these types of problems at scale in both contexts is automating the scoring process, i.e., assessing whether those explanations are correct. The prevailing approach scores an explanation by its semantic similarity to an instructor’s model explanation, but this raises a measurement concern: students who reason correctly yet phrase their explanations differently from an expert may be scored as incorrect (false negatives), threatening the validity and fairness of the assessment. Given recent advances in LLM-based automated scoring, it remains unclear whether semantic similarity methods are still the most effective technique for automatically scoring free-form student responses, such as code explanations. In this paper, we present a rigorous comparison between LLMs and semantic similarity approaches for the automated scoring of student code explanations, using an open dataset Selfcode 2.0. We frame the scoring as a binary classification task and use generative AI to balance the dataset. Our results suggest that LLM-based scoring (F1 = 0.98, accuracy = 0.96) outperform semantic similarity scoring (F1 = 0.72, accuracy = 0.65). It also produces fewer false negatives and eliminate the need to produce well-formulated model explanations.</abstract>
<identifier type="citekey">lekshmi-narayanan-etal-2026-llms</identifier>
<location>
<url>https://aclanthology.org/2026.aimecon-main.10/</url>
</location>
<part>
<date>2026-10</date>
<extent unit="page">
<start>92</start>
<end>99</end>
</extent>
</part>
</mods>
</modsCollection>
%0 Conference Proceedings
%T Can LLMs Replace Semantic Similarity for Scoring Student Code Explanations?
%A Lekshmi Narayanan, Arun Balajiee
%A Hassany, Mohammad
%A Brusilovsky, Peter
%Y Wilson, Joshua
%Y Ormerod, Christopher
%Y Beiting-Parrish, Magdalen
%S Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
%D 2026
%8 October
%I National Council on Measurement in Education (NCME)
%C Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
%@ 979-8-9983004-0-0
%F lekshmi-narayanan-etal-2026-llms
%X In programming courses, students are often asked to explain code fragments, which is a way to assess their understanding of programming constructs and patterns. These types of problem known as “explain in plain English” are valuable in both assessment and practice contexts. The main challenge to using these types of problems at scale in both contexts is automating the scoring process, i.e., assessing whether those explanations are correct. The prevailing approach scores an explanation by its semantic similarity to an instructor’s model explanation, but this raises a measurement concern: students who reason correctly yet phrase their explanations differently from an expert may be scored as incorrect (false negatives), threatening the validity and fairness of the assessment. Given recent advances in LLM-based automated scoring, it remains unclear whether semantic similarity methods are still the most effective technique for automatically scoring free-form student responses, such as code explanations. In this paper, we present a rigorous comparison between LLMs and semantic similarity approaches for the automated scoring of student code explanations, using an open dataset Selfcode 2.0. We frame the scoring as a binary classification task and use generative AI to balance the dataset. Our results suggest that LLM-based scoring (F1 = 0.98, accuracy = 0.96) outperform semantic similarity scoring (F1 = 0.72, accuracy = 0.65). It also produces fewer false negatives and eliminate the need to produce well-formulated model explanations.
%U https://aclanthology.org/2026.aimecon-main.10/
%P 92-99
Markdown (Informal)
[Can LLMs Replace Semantic Similarity for Scoring Student Code Explanations?](https://aclanthology.org/2026.aimecon-main.10/) (Lekshmi Narayanan et al., AIME-Con 2026)
ACL
- Arun Balajiee Lekshmi Narayanan, Mohammad Hassany, and Peter Brusilovsky. 2026. Can LLMs Replace Semantic Similarity for Scoring Student Code Explanations?. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 92–99, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).