ArticleBMC medical informatics and decision making2026
Decoding cardiovascular risk in Chinese middle-aged and elderly adults: a 9-year prospective study integrating machine learning with explainable AI based on CHARLS cohort.
Article in BMC medical informatics and decision making, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
Abstract
backgroundCardiovascular disease constitutes the most formidable public health challenge in China, accounting for 48.98% and 47.35% of mortality in rural and urban populations, respectively, affecting approximately 330 million individuals. Existing risk stratification models predominantly derive from Western populations, with the Framingham Risk Equation systematically overestimating cardiovascular risk by 276% in Chinese men and 102% in Chinese women, underscoring the critical imperative for population-specific predictive instruments. Although machine learning methodologies demonstrate considerable promise in cardiovascular risk prognostication, their inherent "black-box" characteristics substantially impede clinical translational implementation.
objectiveLeveraging longitudinal cohort data from the China Health and Retirement Longitudinal Study (CHARLS) and integrating machine learning with explainable artificial intelligence techniques, we sought to develop and validate a cardiovascular disease long-term risk prediction model tailored to the Chinese middle-aged and elderly population, achieving optimal synthesis of predictive accuracy and clinical interpretability through quantitative risk factor contribution analysis.
methodsWe incorporated four waves of CHARLS surveillance data spanning 2011-2020, with 8,080 participants aged ≥ 45 years completing 9-year follow-up after rigorous inclusion criteria application. Recursive feature elimination was employed to identify optimal predictors from 90 candidate variables. We systematically evaluated 12 machine learning algorithms encompassing linear, non-linear, ensemble learning, and deep learning methodologies, utilizing stratified random 7:3 partitioning for training and validation cohorts. SHAP (SHapley Additive exPlanations) methodology facilitated comprehensive global and local interpretability analyses, with decision curve analysis assessing clinical net benefit.
resultsAmong 5,699 training cohort participants, 1,248 (21.9%) experienced cardiovascular events during follow-up. Recursive feature elimination identified 18 pivotal predictive factors spanning lipid metabolism, anthropometric parameters, renal function, and glucose homeostasis domains. The gradient boosting machine demonstrated superior comprehensive performance, achieving validation cohort AUC of 0.798 (95% CI: 0.776-0.820), specificity of 98%, and positive predictive value of 78%. SHAP analysis revealed waist circumference, triglycerides, and hypertension history as the three predominant predictive factors, with mean absolute SHAP values significantly exceeding other variables. Individual risk attribution analysis demonstrated substantial heterogeneity: extremely high-risk specimens (predicted probability 0.991) exhibited synergistic multi-factorial risk amplification, with standardized waist circumference contributing + 0.0778 SHAP value and triglycerides (477 mg/dL) contributing + 0.0729; conversely, low-risk specimens (predicted probability - 0.0393) demonstrated triglycerides (45.1 mg/dL) providing the maximal singular protective contribution of -0.166. Decision curve analysis confirmed positive net benefit across the 0-0.95 threshold probability spectrum, systematically surpassing conventional strategies.
conclusionsThe gradient boosting machine model achieved superior discrimination (AUC 0.798, 95% CI 0.785-0.825) compared to Framingham (0.638) and China-PAR (0.654) scores for 9-year cardiovascular disease prediction in Chinese adults aged ≥ 45 years. Waist circumference, triglycerides, and hypertension emerged as principal predictive features, though SHAP-derived importance reflects statistical contribution rather than causal effects. Decision curve analysis demonstrated clinical utility across threshold probabilities 0.05-0.95, enabling flexible deployment from population screening (98.3% sensitivity) to targeted intervention (98.7% specificity). External validation in independent cohorts is essential to establish generalizability before clinical implementation. CLINICAL TRIAL NUMBER: Not applicable.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.