Evidence map›Paper›PMID 41491227›Full record

ArticleScientific reports2026

Mapping neighbourhood-level drivers of type 2 diabetes for precision public health using predictive and causal machine learning.

Mohammad Noaeen, Amirhosein Rostami, Ibrahim Ghanem, Olli Saarela, Karim Keshavjee, Jeffrey R Brook, Zahra Shakeri

Abstract read
In one paragraph

Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Mohammad Noaeen *Dalla Lana School of Public Health, University of Toronto, Toronto, Canada.
Amirhosein Rostami *Institute of Health Policy, Management and Evaluation, University of Toronto, Toronto, Canada.
Ibrahim Ghanem *Department of Geography, Geomatics and Environment, University of Toronto Mississauga, Mississauga, Canada.
Olli SaarelaDalla Lana School of Public Health, University of Toronto, Toronto, Canada.
Karim KeshavjeeInstitute of Health Policy, Management and Evaluation, University of Toronto, Toronto, Canada.
Jeffrey R BrookDalla Lana School of Public Health, University of Toronto, Toronto, Canada.
Zahra ShakeriDalla Lana School of Public Health, University of Toronto, Toronto, Canada. zahra.shakeri@utoronto.ca.

Funding

Data Science Institute, University of Toronto DSI-PDFY3R1P31Natural Sciences and Engineering Research Council of Canada (NSERC) RGPIN-2025-07037
6 · The paper itself

Abstract

Type 2 diabetes has become an urban epidemic influenced by neighbourhood environments. However, conventional risk models focusing solely on individual factors fail to account for these neighbourhood influences and often require detailed patient data that may not be available. To address this gap, we developed an integrated approach combining machine learning and causal inference to map type 2 diabetes risk at the neighbourhood level. Using demographic, health, and socioeconomic data from 1,149 Census Tracts (CTs; the neighbourhood unit in this study) in a large metropolitan region, we trained seven machine learning models to identify neighbourhoods with high diabetes prevalence. Although neighbourhood-level diabetes data were available for this study area, our model's high predictive accuracy on external validation data (area under the curve (AUC) = 0.95), particularly from a distinct geographical region, suggests potential utility for predicting diabetes risk in other Canadian regions or elsewhere where such data are unavailable, provided comparable covariates are available and the model is locally retrained and validated using spatially aware procedures. The top models achieved high recall ([Formula: see text]) and AUC up to 0.96 on test data, indicating accurate identification of high-risk neighbourhoods with few missed high-risk areas. Survey-derived neighbourhood health indicators, including obesity rate, physical inactivity, and median age were strong predictors of diabetes prevalence. We then applied a Causal Forest approach to estimate conditional average treatment effects (CATE, τ) for selected potentially modifiable factors and summarized the results with the mean [Formula: see text]. Higher work stress ([Formula: see text]) and daily smoking ([Formula: see text]) were moderately associated with increased risk, whereas better mental health ([Formula: see text]) was protective, highlighting mental health as a priority for further evaluation, especially in neighbourhoods predicted to have high diabetes prevalence. These findings could help identify modifiable neighbourhood-level factors for local prevention efforts and inform equity-oriented planning in diverse urban populations. Prospective or quasi-experimental studies are needed to evaluate intervention effects. Our integrated machine-learning and causal framework lays the groundwork for precision public health, suggesting that modifiable neighbourhood factors may indicate diabetes risk when patient-level data are scarce. Furthermore, the pipeline is conceptually adaptable to other chronic diseases influenced by social and environmental determinants and may inform targeted prevention beyond type 2 diabetes, contingent on disease-specific feature sets and external validation.

Indexed as

Diabetes Mellitus, Type 2Machine LearningPublic HealthResidence CharacteristicsAdultAgedCanadaFemaleHumansMaleMiddle AgedPrevalenceRisk FactorsSocioeconomic Factors

Identifiers

PMID41491227
PMCPMC12858814

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.