Evidence map›Paper›PMID 40924541›Full record

ArticleBioinformatics (Oxford, England)2025

SPACE: STRING proteins as complementary embeddings.

Dewei Hu, Damian Szklarczyk, Christian von Mering, Lars Juhl Jensen

Erratum issuedAbstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Molecular maps of diseases from omics data and network embeddings.NPJ systems biology and applications · 2026
    Article
  2. Article
  3. Protein Language Models in Virology: A Review of Advances and Applications.Methods in molecular biology (Clifton, N.J.) · 2026
    Review
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

4 authors.

Dewei HuNovo Nordisk Foundation Center for Protein Research, Department of Cellular and Molecular Medicine, Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen 2200, Denmark.ORCID 0009-0005-5823-1498
Damian SzklarczykDepartment of Molecular Life Sciences, University of Zurich, Zurich 8057, Switzerland.
Christian von MeringDepartment of Molecular Life Sciences, University of Zurich, Zurich 8057, Switzerland.ORCID 0000-0001-7734-9102
Lars Juhl JensenNovo Nordisk Foundation Center for Protein Research, Department of Cellular and Molecular Medicine, Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen 2200, Denmark.

Funding

Novo Nordisk FoundationNovo Nordisk Foundation NNF14CC0001Novo Nordisk Foundation NNF20SA0035590Swiss Institute of Bioinformatics
6 · The paper itself

Abstract

motivationRepresentation learning has revolutionized sequence-based prediction of protein function and subcellular localization. Protein networks are an important source of information complementary to sequences, but the use of protein networks has proven to be challenging in the context of machine learning, especially in a cross-species setting.

resultsWe leveraged the STRING database of protein networks and orthology relations for 1322 eukaryotes to generate network-based cross-species protein embeddings. We did this by first creating species-specific network embeddings and subsequently aligning them based on orthology relations to facilitate direct cross-species comparisons. We show that these aligned network embeddings ensure consistency across species without sacrificing quality compared to species-specific network embeddings. We also show that the aligned network embeddings are complementary to sequence embedding techniques, despite the use of sequence-based orthology relations in the alignment process. Finally, we validated the embeddings by using them for two well-established tasks: subcellular localization prediction and protein function prediction. Training logistic regression classifiers on aligned network embeddings and sequence embeddings improved the accuracy over using sequence alone, reaching performance numbers close to state-of-the-art deep-learning methods. AVAILABILITY AND IMPLEMENTATION: The source code and scripts for generating the network-based cross-species protein embeddings are available at https://github.com/deweihu96/SPACE. Precomputed network embeddings and sequence embeddings for all eukaryotic proteins are included in STRING version 12.0 (https://string-db.org/cgi/download).

Indexed as

Computational BiologyProteinsDatabases, ProteinDeep LearningHumansMachine LearningSequence Analysis, ProteinProteins

Identifiers

PMID40924541
PMCPMC12453690

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.