ArticleRespiratory research2026
Integration of blood protein-metabolic profiles via machine learning to enable the accurate early detection of non-small cell lung cancer.
Article in Respiratory research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
11 authors.
Funding
Abstract
backgroundNon-small cell lung cancer (NSCLC) poses a major threat to human health due to its high morbidity and mortality. Early accurate diagnosis and differential diagnosis from other respiratory diseases, are pivotal to improving patient prognosis. This study aimed to construct an NSCLC diagnostic model based on multidimensional datasets by leveraging machine learning (ML) driven feature selection and classification algorithms, clarifying the diagnostic value of blood protein metabolic profiles, and provide a novel non-invasive diagnostic scheme for clinical practice.
methodsA total of 144 lung cancer patients (LC), 132 healthy controls (HCs), and 130 patients with other respiratory diseases (ORDs) were recruited from three medical centers. A panel of 26 serum protein metabolism indicators and demographic variables was included as candidate features. Five feature-selection algorithms were used to identify core variables from the high-dimensional data. Ten ML classifiers were subsequently constructed, and their performance was comprehensively evaluated using metrics such as the area under the receiver operating characteristic curve (AUC), accuracy, positive precision, negative precision, positive recall, negative recall, F1 score, and Cohen’s kappa. Moreover, SHapley Additive exPlanations (SHAP) analysis was performed to decipher key predictive factors.
resultsSignificant differences were observed in 21 serum protein metabolic indicators across the LC, HC, and ORD cohorts (p < 0.05). Among the tested algorithms, LightGBM emerged as the superior model for early detection of NSCLC, outperforming logistic regression, SVM, and ANN. It yielded the highest performance metrics, including an AUC of 0.867, positive recall of 0.977, and F1 score of 0.807. Notably, the model minimized both missed and false-positive diagnoses-a crucial factor for clinical utility-while demonstrating robust generalizability (AUC > 0.840 across all datasets). Furthermore, SHAP analysis identified age as the top predictor of NSCLC, followed by Glu, His, ApoB, FN, ApoA2, and Arg.
conclusionsIn this study, a stable, high-performance LightGBM model for NSCLC diagnosis was developed, and key clinical metrics and probability-based outputs were optimized. This model is expected to enhance early detection and differential diagnosis, thereby helping reduce lung cancer morbidity and mortality.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.