Evidence map›Paper›PMID 40671295›Full record

ArticleProtein science : a publication of the Protein Society2025

A large language model for predicting neurotoxic peptides and neurotoxins.

Anand Singh Rathore, Saloni Jain, Shubham Choudhury, Gajendra P S Raghava

Abstract read
In one paragraph

Article in Protein science : a publication of the Protein Society, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. A large language model for predicting neurotoxic peptides and neurotoxins.Protein science : a publication of the Protein Society · 2025
    Article
  5. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Anand Singh RathoreDepartment of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India.ORCID 0009-0004-8907-7174
Saloni JainDepartment of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India.ORCID 0009-0009-3181-1037
Shubham ChoudhuryDepartment of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India.ORCID 0000-0002-4509-4683
Gajendra P S RaghavaDepartment of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India.ORCID 0000-0002-8902-2876

Funding

Council of Scientific and Industrial Research (CSIR)Department of Biotechnology (DBT), Ministry of Science and Technology, India BT/PR40158/BTIS/137/24/2021Department of Science and Technology (DST-INSPIRE)University Grants Commission (UGC)
6 · The paper itself

Abstract

The accurate prediction of neurotoxicity in peptides and proteins is essential for the safety evaluation of therapeutic proteins and genetically modified (GM) organisms. Existing tools, including our earlier method NTxPred, typically use a single predictive model for both neurotoxic peptides and proteins, despite their structural and functional differences. This lack of specialization may lead to suboptimal performance and limited generalizability. To address this, we developed NTxPred2, distinct, specialized models for predicting neurotoxic peptides and neurotoxins (proteins). Our curated datasets include 877 neurotoxic and 877 non-toxic peptides, and 775 neurotoxic and 775 non-toxic proteins. Certain residues, like cysteine, are prevalent in both but in different magnitudes. Using composition and binary profiles, our machine-learning models achieved an area under the curve (AUC) of 0.97 for peptides and 0.85 for proteins, improving to 0.89 with evolutionary information. Models using protein embeddings reached 0.96 AUC for peptides and 0.94 for proteins, while protein language models achieved 0.98 (esm2-t30) and 0.91 (esm2-t6). All models were validated via five-fold cross-validation, and the final models were evaluated on an independent dataset. We further assessed protein models on the peptide dataset and vice versa, highlighting the necessity of separate models. The proposed models outperform existing methods on independent datasets that are not used for training. Our neurotoxicity prediction models will aid in the safety assessment of GM foods and therapeutic proteins by minimizing the need for animal testing. To support the scientific community, we developed a standalone software and web server NTxPred2 for predicting and scanning neurotoxins (https://webs.iiitd.edu.in/raghava/ntxpred2/, https://github.com/raghavagps/ntxpred2/).

Indexed as

Machine LearningNeurotoxinsPeptidesDatabases, ProteinHumansLarge Language ModelsSoftwareNeurotoxinsPeptidesembeddingsmachine‐learning techniquesneurotoxic peptidesneurotoxinsprotein language modelsWHO guidelines

Identifiers

PMID40671295
PMCPMC12267675

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.