Don’t Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

Jiwon Moon; Yerin Hwang; Dongryeol Lee; Taegwan Kang; Yongil Kim; Kyomin Jung

Don’t Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, Kyomin Jung

Abstract

With the growing use of large language models (LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While this offers scalability and flexibility, it also raises a critical, unresolved question: Can LLM judges fairly and robustly evaluate semantically equivalent code with superficial variations? Functionally correct code often exhibits variations—such as differences in variable names, comments, or formatting—that should not influence its correctness. Yet, whether LLM judges can reliably handle these variations remains unclear. We present the first comprehensive study of this issue, defining six types of potential bias in code evaluation and revealing their systematic impact on LLM judges. Across five programming languages and multiple LLMs, we empirically demonstrate that all tested LLM judges are susceptible to both positive and negative biases, resulting in inflated or unfairly low scores. Moreover, we observe that LLM judges remain vulnerable to these biases even when prompted to generate test cases before scoring, highlighting the need for more robust code evaluation method.

Anthology ID:: 2026.findings-eacl.70
Volume:: Findings of the Association for Computational Linguistics: EACL 2026
Month:: March
Year:: 2026
Address:: Rabat, Morocco
Editors:: Vera Demberg, Kentaro Inui, Lluís Marquez
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1364–1389
Language:
URL:: https://aclanthology.org/2026.findings-eacl.70/
DOI:
Bibkey:
Cite (ACL):: Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, and Kyomin Jung. 2026. Don’t Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation. In Findings of the Association for Computational Linguistics: EACL 2026, pages 1364–1389, Rabat, Morocco. Association for Computational Linguistics.
Cite (Informal):: Don’t Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation (Moon et al., Findings 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.findings-eacl.70.pdf
Checklist:: 2026.findings-eacl.70.checklist.pdf

PDF Cite Search Checklist Fix data