Evidence map›Paper›PMID 41934166›Full record

ArticleJournal of diabetes research2026

Machine Learning-Based Prediction Model Construction for Type 2 Diabetes Mellitus: A Comparison of Algorithms and Multilevel Risk Factor Analysis.

Qian Xu, Ruicong Yu, Huixin Qiu, Yannan Jiang, Jake Ball, Cuirong Xu, Jing Sun

Abstract readComparative Study
In one paragraph

Article in Journal of diabetes research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Qian XuOperating Room, Zhongda Hospital Southeast University, Nanjing, Jiangsu Province, China, cis.seu.edu.cn.
Ruicong YuSchool of Medicine, Southeast University, Nanjing, Jiangsu Province, China, seu.edu.bd.ORCID https://orcid.org/0009-0006-0183-7525
Huixin QiuSchool of Medicine, Southeast University, Nanjing, Jiangsu Province, China, seu.edu.bd.
Yannan JiangData Science Institute, University of Technology Sydney, Sydney, New South Wales, Australia, uts.edu.au.
Jake BallData Science Institute, University of Technology Sydney, Sydney, New South Wales, Australia, uts.edu.au.
Cuirong XuDepartment of Nursing, Zhongda Hospital Southeast University, Nanjing, Jiangsu Province, China, cis.seu.edu.cn.ORCID https://orcid.org/0000-0002-8979-0533
Jing SunData Science Institute, University of Technology Sydney, Sydney, New South Wales, Australia, uts.edu.au.ORCID https://orcid.org/0000-0002-0097-2438

Funding

Jiangsu Commission of Health ZD2022057
6 · The paper itself

Abstract

backgroundAgainst the backdrop of the global high incidence of Type 2 diabetes mellitus (T2DM), existing prediction models are largely confined to single-dimensional risk factors, suffering from a core limitation of lacking multilevel integrated analysis. Given the severe impact of T2DM on individual health and healthcare systems, the construction of a comprehensive and accurate prediction model is of great significance.

objectiveThis study is aimed at constructing a T2DM prediction model, identifying multilevel risk factors, and enabling early screening, so as to help clinicians identify high-risk individuals and provide targets for public health interventions.

methodsData from the National Health and Nutrition Examination Survey (NHANES) 2021-2023 were used, including 6337 participants aged 18 years and older. Missing values were handled using Monte Carlo multiple imputation, collinearity was reduced via principal component analysis (PCA), and feature selection was performed using random forest (RF) and recursive feature elimination (RFE). The adaptive synthetic sampling (ADASYN) method was applied to address class imbalance. The performance of seven machine learning models, including decision tree, random forest, extreme gradient boosting (XGBoost), and adaptive boosting (AdaBoost), was compared.

resultsThe AdaBoost model exhibited the optimal performance, with an area under the curve (AUC) of 0.85 (95% confidence interval: 0.85-0.86), an accuracy of 0.71 (95% confidence interval: 0.70-0.72), and an F1 score of 0.71; its performance was further improved after parameter optimization. A total of 24 key risk factors were identified, including 19 at the individual trait level, 3 at the individual behavior level, and 2 related to working and living conditions.

conclusionsMachine learning models integrating multidimensional risk factors based on the health ecology framework can more accurately predict T2DM risk, providing a scientific basis for multilevel interventions. The innovation of this study lies in the first integration of the health ecology model with machine learning technology to systematically identify cross-level risk factors. Compared with traditional models, it is more comprehensive, breaks through the limitations of previous studies, and provides a new and effective tool for the precise prevention of T2DM and public health interventions.

Indexed as

Diabetes Mellitus, Type 2Machine LearningAdultAlgorithmsBoosting Machine Learning AlgorithmsClassification AlgorithmsFemaleHumansMaleMiddle AgedNutrition SurveysPrediction AlgorithmsPredictive Learning ModelsRandom ForestRisk Factorshealth ecologymachine learningprediction modelType 2 diabetes mellitus

Identifiers

PMID41934166
PMCPMC13051855

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.