Protein Large Language Models: A Comprehensive Survey

Yijia Xiao; Wanjia Zhao; Junkai Zhang; Yiqiao Jin; Han Zhang; Zhicheng Ren; Renliang Sun; Haixin Wang; Guancheng Wan; Pan Lu; Xiao Luo; Yu Zhang; James Zou; Yizhou Sun; Wei Wang (王巍)

Protein Large Language Models: A Comprehensive Survey

Yijia Xiao, Wanjia Zhao, Junkai Zhang, Yiqiao Jin, Han Zhang, Zhicheng Ren, Renliang Sun, Haixin Wang, Guancheng Wan, Pan Lu, Xiao Luo, Yu Zhang, James Zou, Yizhou Sun, Wei Wang

Abstract

Protein-specific large language models (ProteinLLMs) are revolutionizing protein science by enabling more efficient protein structure prediction, function annotation, and design. While existing surveys focus on specific aspects or applications, this work provides the first comprehensive overview of ProteinLLMs, covering their architectures, training datasets, evaluation metrics, and diverse applications. Through a systematic analysis of over 100 articles, we propose a structured taxonomy of state-of-the-art ProteinLLMs, analyze how they leverage large-scale protein sequence data for improved accuracy, and explore their potential in advancing protein engineering and biomedical research. Additionally, we discuss key challenges and future directions, positioning ProteinLLMs as essential tools for scientific discovery in protein science. Resources are maintained at https://github.com/Yijia-Xiao/Protein-LLM-Survey.

Anthology ID:: 2025.findings-emnlp.1255
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 23080–23103
Language:
URL:: https://aclanthology.org/2025.findings-emnlp.1255/
DOI:
Bibkey:
Cite (ACL):: Yijia Xiao, Wanjia Zhao, Junkai Zhang, Yiqiao Jin, Han Zhang, Zhicheng Ren, Renliang Sun, Haixin Wang, Guancheng Wan, Pan Lu, Xiao Luo, Yu Zhang, James Zou, Yizhou Sun, and Wei Wang. 2025. Protein Large Language Models: A Comprehensive Survey. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 23080–23103, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Protein Large Language Models: A Comprehensive Survey (Xiao et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-emnlp.1255.pdf
Checklist:: 2025.findings-emnlp.1255.checklist.pdf

PDF Cite Search Checklist Fix data