OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Siming Huang; Tianhao Cheng; Jason Klein Liu; Weidi Xu; Jiaran Hao; Liuyihan Song; Yang Xu; Jian Yang; Jiaheng Liu; Chenchen Zhang; Linzheng Chai; Ruifeng Yuan; Xianzhen Luo; Qiufeng Wang; Yuantao Fan; Qingfu Zhu; Zhaoxiang Zhang; Yang Gao (扬 高); Jie Fu; Qian Liu; Houyi Li; Ge Zhang; Yuan Qi; Xu Yinghui; Wei Chu; Zili Wang

doi:10.18653/v1/2025.acl-long.1591

OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, Zili Wang

Abstract

Code LLMs have been widely used in various domains, including code generation, logical reasoning, and agent systems. However, open-access code LLMs mostly only release weights, lacking key features such as reproducible data pipelines and transparent training protocols, which are crucial for advancing deeper, more reliable investigations. To address the gap, we introduce OpenCoder, a top-tier code LLM that not only achieves performance comparable to leading models but also serves as an “open cookbook” for the research community. Unlike most prior efforts, we release not only model weights and inference code, but also the reproducible training data, complete data processing pipeline, rigorous experimental ablation results, and detailed training protocols for open scientific research. Our work identifies the key ingredients for building a top-tier code LLM: optimized heuristic rules for data cleaning and deduplication, effective recall of code-related text corpus, and high-quality synthetic data for both annealing and supervised fine-tuning stages. By offering this level of openness, we aim to broaden access to all aspects of a top-tier code LLM, with OpenCoder serving as both a powerful model and an open foundation to accelerate research and enable reproducible advancements in code intelligence. The released resource is available at https://opencoder-llm.github.io.

Anthology ID:: 2025.acl-long.1591
Volume:: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2025
Address:: Vienna, Austria
Editors:: Wanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 33167–33193
Language:
URL:: https://aclanthology.org/2025.acl-long.1591/
DOI:: 10.18653/v1/2025.acl-long.1591
Bibkey:
Cite (ACL):: Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, and Zili Wang. 2025. OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 33167–33193, Vienna, Austria. Association for Computational Linguistics.
Cite (Informal):: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models (Huang et al., ACL 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.acl-long.1591.pdf

PDF Cite Search Fix data