Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generation

Shiwen Zhang; Lingxiang Wang; Hainan Zhang; Ziwei Wang; Sijia Wen; Zhiming Zheng

doi:10.18653/v1/2025.findings-emnlp.281

Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generation

Shiwen Zhang, Lingxiang Wang, Hainan Zhang, Ziwei Wang, Sijia Wen, Zhiming Zheng

Abstract

In competitive programming task, problem statements are often embedded within elaborate narrative backgrounds, requiring deep understanding of the underlying solutions to successfully complete the tasks. Current code generation models primarily focus on token-level semantic modeling, highly susceptible to distractions from irrelevant narrative statements. Inspired by RAG, retrieving reference code with similar solutions may help enhance model performance on difficult problems. However, existing retrieval models also emphasize surface-level semantic similarity, neglecting the deeper solution-level logical similarities that are critical in competitive programming. Therefore, designing ranking models capable of accurately identifying and retrieving problems and corresponding codes remains an urgent research problem in competitive code generation. In this paper, we propose SolveRank, a solution-aware ranking model empowered by synthetic data for competitive programming tasks. Specifically, we leverage the DeepSeek-R1 model to generate logically equivalent but differently phrased new problems, verified by GPT-4o for solution consistency. Then, we train SolveRank with these as positive samples and BM25/random-retrieved problems as negatives. During inference, SolveRank retrieves relevant problems and corresponding code from the corpus to assist a downstream code generator. Experiments on the xCodeEval dataset demonstrate that SolveRank outperforms SOTA ranking methods in precision and recall metrics, and boosts code generation performance for difficult problems.

Anthology ID:: 2025.findings-emnlp.281
Original:: 2025.findings-emnlp.281v1
Version 2:: 2025.findings-emnlp.281v2
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 5237–5246
Language:
URL:: https://aclanthology.org/2025.findings-emnlp.281/
DOI:: 10.18653/v1/2025.findings-emnlp.281
Bibkey:
Cite (ACL):: Shiwen Zhang, Lingxiang Wang, Hainan Zhang, Ziwei Wang, Sijia Wen, and Zhiming Zheng. 2025. Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 5237–5246, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generation (Zhang et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-emnlp.281.pdf
Checklist:: 2025.findings-emnlp.281.checklist.pdf

PDF (v2) PDF (v1) Cite Search Checklist Fix data