ArticleFluids and barriers of the CNS2025
Applying machine learning to high-dimensional proteomics datasets for the identification of Alzheimer's disease biomarkers.
Article in Fluids and barriers of the CNS, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 9 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
9 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Decoding the CSF Proteomic Signature of Idiopathic Normal Pressure Hydrocephalus: A Systematic Review.Molecules (Basel, Switzerland) · 2026Pooled it
- Progress and trends on machine learning in proteomics during 1997-2024: a bibliometric analysis.Frontiers in medicine · 2025Pooled it
- Artificial intelligence and extracellular vesicles in oncology: towards tumor diagnosis, prediction, and therapy.Drug delivery · 2026Review
- Functional Foods and Micro- and Nanoplastics: Advances in Precision Nutritional Medicine for Oral-Gut-Brain Axis Health.Antioxidants (Basel, Switzerland) · 2026Review
- Discovery-Driven Plasma Proteomics Identifies a Multi-Protein Signature for Amyloid PET Positivity: A Machine Learning Analysis of the Bio-Hermes Cohort.International journal of molecular sciences · 2026Article
- Multi-omics analysis identifies oxidative stress-related biomarkers and therapeutic targets linking periodontitis and ulcerative colitis via the oral-gut axis.Frontiers in immunology · 2026Article
- Advancing global dementia research through equity and inclusion.Alzheimer's & dementia : the journal of the Alzheimer's Association · 2026Article
- A nine-gene signature with potential targets for predicting the prognosis of patients with esophageal cancer.Translational cancer research · 2025Article
- Multi-modal machine learning and gut microbiome pathway analysis for Alzheimer's risk prediction.Alzheimer's & dementia (Amsterdam, Netherlands)Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
Abstract
purposeThis study explores the application of machine learning to high-dimensional proteomics datasets for identifying Alzheimer's disease (AD) biomarkers. AD, a neurodegenerative disorder affecting millions worldwide, necessitates early and accurate diagnosis for effective management.
methodsWe leverage Tandem Mass Tag (TMT) proteomics data from the cerebrospinal fluid (CSF) samples from the frontal cortex of patients with idiopathic normal pressure hydrocephalus (iNPH), a condition often comorbid with AD, with rare access to both lumbar and ventricular samples. Our methodology includes extensive data preprocessing to address batch effects and missing values, followed by the use of the Synthetic Minority Over-sampling Technique (SMOTE) for data augmentation to overcome the small sample size. We apply linear, and non-linear machine learning models, and ensemble methods, to compare iNPH patients with and without biomarker evidence of AD pathology (
resultsWe present a machine learning workflow for working with high-dimensional TMT proteomics data that addresses their inherent data characteristics. Our results demonstrate that batch effect correction has no or minor impact on the models' performance and robust feature selection is critical for model stability and performance, especially in the high-dimensional proteomics data setting for AD diagnostics. The results further indicated that removing features with missing values produced stronger models than imputing them, and the batch effect had minimal impact on the models Our best-performing disease-progression detection model, a random forest, achieves an AUC of 0.84 (± 0.03).
conclusionWe identify several novel protein biomarkers candidates, such as FABP3 and GOT1, with potential diagnostic value for AD pathology detection, suggesting the necessity of different biomarkers for AD diagnoses for patients with iNPH, and considering different biomarkers for ventricular and lumbar CSF samples. This work underscores the importance of a meticulous machine learning process in enhancing biomarker discovery. Our study also provides insights in translating biomarkers from other central nervous system diseases like iNPH, and both ventricular and lumbar CSF samples for biomarker discovery, providing a foundation for future research and clinical applications.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.