Evidence map›Paper›PMID 42467985›Full record

ReviewBriefings in bioinformatics2026

Decoding viral protein sequences by large language models.

Tianyi Fei, Siqi Li, Ziyue Yang, Kaitao Zhou, Yixue Li, Tao Zeng

Abstract readReview
In one paragraph

Review in Briefings in bioinformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Tianyi FeiGMU-GIBH Joint School of Life Sciences, Guangdong Provincial Key Laboratory of Protein Modification and Disease, The Guangdong-Hong Kong-Macao Joint Laboratory for Cell Fate Regulation and Diseases, Guangzhou Medical University, Guangzhou 511436, Guangdong Province, China.
Siqi LiGMU-GIBH Joint School of Life Sciences, Guangdong Provincial Key Laboratory of Protein Modification and Disease, The Guangdong-Hong Kong-Macao Joint Laboratory for Cell Fate Regulation and Diseases, Guangzhou Medical University, Guangzhou 511436, Guangdong Province, China.
Ziyue YangGuangzhou National Laboratory, No. 9 XingDaoHuanBei Road, Guangzhou International Bio Island, Guangzhou 510005, Guangdong Province, China.
Kaitao ZhouGuangzhou National Laboratory, No. 9 XingDaoHuanBei Road, Guangzhou International Bio Island, Guangzhou 510005, Guangdong Province, China.
Yixue LiGMU-GIBH Joint School of Life Sciences, Guangdong Provincial Key Laboratory of Protein Modification and Disease, The Guangdong-Hong Kong-Macao Joint Laboratory for Cell Fate Regulation and Diseases, Guangzhou Medical University, Guangzhou 511436, Guangdong Province, China.ORCID 0000-0002-1198-7176
Tao ZengGMU-GIBH Joint School of Life Sciences, Guangdong Provincial Key Laboratory of Protein Modification and Disease, The Guangdong-Hong Kong-Macao Joint Laboratory for Cell Fate Regulation and Diseases, Guangzhou Medical University, Guangzhou 511436, Guangdong Province, China.ORCID 0000-0002-0295-3994

Funding

Major Project of Guangzhou National Laboratory GZNL2025C01013National Key Research and Development Program of China 2023YFF1204700National Natural Science Foundation of China 12371485National Natural Science Foundation of China 92574105Natural Science Foundation of Guangdong Province of China 2025A1515011988Prevention and Control of Emerging and Major Infectious Diseases-National Science and Technology Major Project 2025ZD01901900
6 · The paper itself

Abstract

Large language models (LLMs) for biological sequences are transforming computational biology, enabling a nuanced understanding of protein and nucleotide sequence data. Recent models, including ESM2, ESM3, AlphaGenome, Evo-1, and Evo-2, adapt natural language processing principles to the biological domain by learning high-dimensional hidden representations that capture evolutionary constraints, structural patterns, and functional motifs. This mini-review summarizes recent developments in devising and applying such models, emphasizing viral protein analysis. We highlight studies that have leveraged sequence-based LLMs in the protein domain (i.e. protein language models, or PLMs) for important application tasks such as viral protein annotation, variant effect prediction, and immune escape characterization. Additionally, we present a benchmark evaluation of these state-of-the-art protein language models to evaluate their core ability to capture evolutionary relationships between viral protein sequences. By discussing the opportunities and challenges of PLMs, the review outlines a road map for the potential application of LLMs in empowering virology research and pathogen surveillance.

Indexed as

Computational BiologyLarge Language ModelsViral ProteinsAmino Acid SequenceHumansViral Proteinslarge language modelsprotein language models

Identifiers

PMID42467985
PMCPMC13379068

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.