Evidence map›Paper›PMID 38259343›Full record

ArticleArXiv2025

Advancing bioinformatics with large language models: components, applications and perspectives.

Jiajia Liu, Mengyuan Yang, Yankai Yu, Haixia Xu, Tiangang Wang, Kang Li, Xiaobo Zhou

Abstract readPreprint
In one paragraph

Article in ArXiv, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

7 authors.

Jiajia LiuCenter for Computational Systems Medicine, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, Texas, 77030, USA.
Mengyuan YangDepartment of Cell Biology and Genetics, School of Basic Medical Sciences, Xi'an Jiaotong University Health Science Center, Xi'an, China.
Yankai YuSchool of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu, Sichuan 611756, China.
Haixia XuCenter for Computational Systems Medicine, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, Texas, 77030, USA.
Tiangang WangCenter for Computational Systems Medicine, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, Texas, 77030, USA.
Kang LiWest China Biomedical Big Data Center, West China Hospital, Sichuan University, Chengdu, Sichuan 610041, China.
Xiaobo ZhouCenter for Computational Systems Medicine, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, Texas, 77030, USA.

Funding

Systems Modeling Guided Bone regenerationU01AR069395 · NIAMS · WAKE FOREST UNIVERSITY HEALTH SCIENCES · PI YANG, YUNZHI, ZHOU, XIAOBO · 2016 to 2021
$3.4M
Multiscale Resolution and Deep Network Approaches for Deconvolving Different Cell Types in Bulk Tumor using Single-cell Sequencing Data (scDEC)R01CA241930 · NCI · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI ZHOU, XIAOBO · 2019 to 2023
$2.7M
Integrative approach to studying LncRNA functionsR01GM123037 · NIGMS · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI ZHOU, XIAOBO · 2017 to 2020
$1.5M
Optimizing mRNA sequences with deep neural networksR01LM014156 · NLM · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI Xiaobo Zhou · 2024 to 2026
$1.1M
Developing mRNAdesigner tool package for optimization of mRNA sequenceR01GM153822 · NIGMS · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI ZHOU, XIAOBO · 2024 to 2025
$624k
NCI NIH HHS R01 CA241930NIAMS NIH HHS U01 AR069395NIGMS NIH HHS R01 GM123037NIGMS NIH HHS R01 GM153822NLM NIH HHS R01 LM014156
6 · The paper itself

Abstract

Large language models (LLMs) are a class of artificial intelligence models based on deep learning, which have great performance in various tasks, especially in natural language processing (NLP). Large language models typically consist of artificial neural networks with numerous parameters, trained on large amounts of unlabeled input using self-supervised or semi-supervised learning. However, their potential for solving bioinformatics problems may even exceed their proficiency in modeling human language. In this review, we will provide a comprehensive overview of the essential components of large language models (LLMs) in bioinformatics, spanning genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. Key aspects covered include tokenization methods for diverse data types, the architecture of transformer models, the core attention mechanism, and the pre-training processes underlying these models. Additionally, we will introduce currently available foundation models and highlight their downstream applications across various bioinformatics domains. Finally, drawing from our experience, we will offer practical guidance for both LLM users and developers, emphasizing strategies to optimize their use and foster further innovation in the field.

Identifiers

PMID38259343
PMCPMC10802675

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.