ArticleKidney diseases (Basel, Switzerland)
Predicting Chronic Kidney Disease in Type 2 Diabetes Using Natural Language Processing on Healthcare Data.
Article in Kidney diseases (Basel, Switzerland). The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
3 citing papers in PubMed.
- Real-World Insights into Stage I-III Non-Small Cell Lung Cancer in Spain in the Pre-Immunotherapy Era Using AI Techniques: The IntellyLUNG Study.Life (Basel, Switzerland) · 2026Article
- Natural Language Processing of Unstructured Healthcare Data for Predicting Heart Failure in Individuals with Type 2 Diabetes.Journal of clinical medicine · 2026Article
- Shaping the Future of Evidence Generation: Real-World Data to Drive Healthcare Transformation and Patient-Centered Decisions.Pharmaceutical medicine · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
18 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Introduction: Persons with type 2 diabetes mellitus (T2DM) attending hospitals frequently experience major complications. We assessed the potential use of unstructured free-text data extracted from electronic health records (EHRs) using natural language processing (NLP) and machine learning (ML) to develop a predictive model for chronic kidney disease (CKD) in T2DM. Methods: This multicenter retrospective study included data from eight Spanish hospitals (2013-2018), extracted using NLP and ML techniques (EHRead®) based on SNOMED CT terminology. From a cohort of individuals with T2DM, we identified those with and without CKD at inclusion. Among individuals without CKD, we trained and validated a 2-year predictive model for CKD development. The model showing the best balance between performance and clinical interpretability was selected for integration into a web-based tool to support early detection and risk stratification. Results: Of 588,786 individuals with T2DM, 316,597 were included for model development (training: 291,429 [92.1%]; validation: 25,168 [7.9%]; CKD incidence: 15.4% and 18.4%, respectively). A high proportion of missing data was observed in key clinical variables. Among models evaluated, logistic regression achieved the best performance (receiver operating characteristic area under the curve 0.72) using 27 predictors. Both a reduced 10-predictor model and a clinically refined 8-predictor model showed comparable performance to the full model in training and validation cohorts. The clinically refined model was selected for implementation in the web-based tool. Conclusion: Unstructured EHR data enabled the development of a predictive model for 2-year CKD risk in persons with T2DM. Improving EHR data completeness remains essential to enhance future predictive modeling.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.