Evidence map›Paper›PMID 40033432›Full record

ArticleFluids and barriers of the CNS2025

Applying machine learning to high-dimensional proteomics datasets for the identification of Alzheimer's disease biomarkers.

Christoffer Ivarsson Orrelid, Oscar Rosberg, Sophia Weiner, Fredrik D Johansson, Johan Gobom, Henrik Zetterberg, Newton Mwai, Lena Stempfle

Abstract read
In one paragraph

Article in Fluids and barriers of the CNS, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 9 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
9citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

9 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Review
  4. Review
  5. Article
  6. Article
  7. Advancing global dementia research through equity and inclusion.Alzheimer's & dementia : the journal of the Alzheimer's Association · 2026
    Article
  8. Article
  9. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Christoffer Ivarsson Orrelid *Computer Science and Engineering, Chalmers University of Technology and University of Gothenburg, Rännvägen 6b, 41296, Gothenburg, Västra Götalandsregionen, Sweden. christoffer.orrelid@gmail.com.
Oscar Rosberg *Computer Science and Engineering, Chalmers University of Technology and University of Gothenburg, Rännvägen 6b, 41296, Gothenburg, Västra Götalandsregionen, Sweden.
Sophia WeinerDepartment of Psychiatry and Neurochemistry, The Sahlgrenska Academy at the University of Gothenburg, Wallinsgatan 6, 43141, Möndal, Västra Götalandsregionen, Sweden.
Fredrik D JohanssonComputer Science and Engineering, Chalmers University of Technology and University of Gothenburg, Rännvägen 6b, 41296, Gothenburg, Västra Götalandsregionen, Sweden.
Johan GobomDepartment of Psychiatry and Neurochemistry, The Sahlgrenska Academy at the University of Gothenburg, Wallinsgatan 6, 43141, Möndal, Västra Götalandsregionen, Sweden.
Henrik ZetterbergDepartment of Psychiatry and Neurochemistry, The Sahlgrenska Academy at the University of Gothenburg, Wallinsgatan 6, 43141, Möndal, Västra Götalandsregionen, Sweden.
Newton MwaiComputer Science and Engineering, Chalmers University of Technology and University of Gothenburg, Rännvägen 6b, 41296, Gothenburg, Västra Götalandsregionen, Sweden.
Lena StempfleComputer Science and Engineering, Chalmers University of Technology and University of Gothenburg, Rännvägen 6b, 41296, Gothenburg, Västra Götalandsregionen, Sweden.

Funding

AD Strategic Fund and the Alzheimer's Association #ADSF-21-831376-C, #ADSF-21-831381-C, #ADSF-21-831377-C, and #ADSF-24-1284328-CAlzheimer's Drug Discovery Foundation 201809-2016862European Partnership on Metrology, co-financed from the European Union's Horizon Europe Research and Innovation Programme and by the Participating States NEuroBioStand, #22HLT07European Union Joint Programme- Neurodegenerative Disease Research JPND2021-00694European Union's Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie 860197 (MIRI ADE)European Union's Horizon Europe research and innovation programme 101053962Hjärnfonden FO2022-0270National Institute for Health and Care Research Univer sity College London Hospitals Biomedical Research Centre, and the UK Dementia Research Institute at UCL UKDRI-1003Swedish Research Council #2023-00356; #2022-01018 and #2019-02397Swedish State Support for Clinical Research ALFGBG-71320
6 · The paper itself

Abstract

purposeThis study explores the application of machine learning to high-dimensional proteomics datasets for identifying Alzheimer's disease (AD) biomarkers. AD, a neurodegenerative disorder affecting millions worldwide, necessitates early and accurate diagnosis for effective management.

methodsWe leverage Tandem Mass Tag (TMT) proteomics data from the cerebrospinal fluid (CSF) samples from the frontal cortex of patients with idiopathic normal pressure hydrocephalus (iNPH), a condition often comorbid with AD, with rare access to both lumbar and ventricular samples. Our methodology includes extensive data preprocessing to address batch effects and missing values, followed by the use of the Synthetic Minority Over-sampling Technique (SMOTE) for data augmentation to overcome the small sample size. We apply linear, and non-linear machine learning models, and ensemble methods, to compare iNPH patients with and without biomarker evidence of AD pathology (

resultsWe present a machine learning workflow for working with high-dimensional TMT proteomics data that addresses their inherent data characteristics. Our results demonstrate that batch effect correction has no or minor impact on the models' performance and robust feature selection is critical for model stability and performance, especially in the high-dimensional proteomics data setting for AD diagnostics. The results further indicated that removing features with missing values produced stronger models than imputing them, and the batch effect had minimal impact on the models Our best-performing disease-progression detection model, a random forest, achieves an AUC of 0.84 (± 0.03).

conclusionWe identify several novel protein biomarkers candidates, such as FABP3 and GOT1, with potential diagnostic value for AD pathology detection, suggesting the necessity of different biomarkers for AD diagnoses for patients with iNPH, and considering different biomarkers for ventricular and lumbar CSF samples. This work underscores the importance of a meticulous machine learning process in enhancing biomarker discovery. Our study also provides insights in translating biomarkers from other central nervous system diseases like iNPH, and both ventricular and lumbar CSF samples for biomarker discovery, providing a foundation for future research and clinical applications.

Indexed as

Alzheimer DiseaseHydrocephalus, Normal PressureMachine LearningProteomicsAgedAged, 80 and overBiomarkersFemaleHumansMaleBiomarkersAlzheimer’s diseaseBiomarkersFeature selectionHigh-dimensional dataMachine learningMass spectrometryProteomics

Identifiers

PMID40033432
PMCPMC11874791

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.