ArticleEuropean journal of medical research2026
Machine learning and SHAP values for predicting coronary artery disease risk in Xinjiang, China.
Article in European journal of medical research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
Abstract
backgroundAccurate individual risk assessment is crucial for guiding and improving the prevention of atherosclerotic cardiovascular disease (ASCVD). Existing prediction models are primarily derived from Western Caucasian and Chinese Han populations. Our objective is to develop and validate an interpretable machine learning (ML) model based on biomarkers for predicting coronary artery disease (CAD) risk among multi-ethnic patients in Xinjiang, China.
methodsThis retrospective cohort study enrolled patients who underwent coronary angiography or coronary computed tomography angiography (CCTA) at the First Affiliated Hospital of Xinjiang Medical University. The cohort was divided into training, validation, and test sets. Feature selection was performed using logistic regression and LASSO, followed by prediction model development with six machine learning algorithms (XGBoost, RF, MLP, SVM, KNN, AdaBoost). Predictive performance was evaluated using the area under the receiver operating characteristic curve (AUROC) as the primary metric to identify the optimal algorithm. The selected algorithm was further validated on both the validation and testing sets. Shapley Additive Explanations (SHAP) were applied to quantify each feature's contribution to CAD risk prediction, generating individualized risk explanations. Furthermore, the model's calibration was assessed using calibration curves and the Brier score. Its clinical utility was evaluated through decision curve analysis and was benchmarked against the established SCORE2 Asia Pacific risk model.
resultsThis study enrolled 7655 male and 4461 female participants, divided into training, validation, and test sets in a 7:1.5:1.5 ratio. XGBoost demonstrated optimal performance in both cohorts: the male model achieved AUROCs of 0.845 (95% CI: 0.834-0.855), 0.814 (0.789-0.839), and 0.826 (0.802-0.850) in the training, validation, and test sets, respectively, while the female model attained values of 0.817 (0.802-0.832), 0.759 (0.721-0.796), and 0.786 (0.751-0.821). Male CAD risk was significantly associated with advanced age, multiple abnormal clinical indicators (elevated creatinine, total cholesterol, lipoprotein(a), etc., and decreased HDL-C), hypertension, and diabetes, with higher risk observed in Kazakh and Hui ethnicities, whereas higher education and married status served as protective factors. In females, hypertension was the strongest predictor, while elevated uric acid, systolic blood pressure, fasting blood glucose, along with histories of hypertension and diabetes increased risk; married status and higher education similarly exhibited protective effects. The prediction model demonstrated favorable clinical utility and accuracy in both cohorts, with calibration significantly enhancing predictive performance. Compared to the SCORE2 Asia Pacific risk model, our model exhibited superior discriminatory ability (male: 0.826 vs. 0.662; female: 0.786 vs. 0.720) and improved calibration.
conclusionsMachine learning models can provide personalized and highly accurate predictions of CAD risk. The interpretability of these models facilitates the identification of modifiable risk factors in individual patients, offering valuable insights to enhance primary prevention and management of cardiovascular disease in the Xinjiang region.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.