Evidence mapPaperPMID 39908080Full record

ArticleJMIR medical informatics2025

Identification of Clusters in a Population With Obesity Using Machine Learning: Secondary Analysis of The Maastricht Study.

Maik Jm Beuken, Melanie Kleynen, Susy Braun, Kees Van Berkel, Carla van der Kallen, Annemarie Koster, Hans Bosma, Tos Tjm Berendschot, Alfons Jhm Houben, Nicole Dukers-Muijrers and 4 more

Abstract read
In one paragraph

Article in JMIR medical informatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

14 authors.

Maik Jm BeukenFaculty of Financial Management, Research Center for Statistics & Data Science, Zuyd University of Applied Sciences, Sittard, Netherlands.ORCID https://orcid.org/0000-0002-8356-2716
Melanie KleynenFaculty of Health, School of Physiotherapy, Research Center for Nutrition, Lifestyle and Exercise, Zuyd University of Applied Sciences, Heerlen, Netherlands.ORCID https://orcid.org/0000-0002-6543-6994
Susy BraunFaculty of Health, School of Physiotherapy, Research Center for Nutrition, Lifestyle and Exercise, Zuyd University of Applied Sciences, Heerlen, Netherlands.ORCID https://orcid.org/0000-0002-3037-3428
Kees Van BerkelFaculty of Financial Management, Research Center for Statistics & Data Science, Zuyd University of Applied Sciences, Sittard, Netherlands.ORCID https://orcid.org/0000-0003-0139-9298
Carla van der KallenDepartment of Internal Medicine, Maastricht University Medical Center+, Maastricht, Netherlands.ORCID https://orcid.org/0000-0003-1468-8793
Annemarie KosterDepartment of Social Medicine, Maastricht University, Maastricht, Netherlands.ORCID https://orcid.org/0000-0003-1583-7391
Hans BosmaDepartment of Social Medicine, Maastricht University, Maastricht, Netherlands.ORCID https://orcid.org/0000-0003-4333-4564
Tos Tjm BerendschotUniversity Eye Clinic Maastricht, Maastricht University, Maastricht, Netherlands.ORCID https://orcid.org/0000-0002-8101-939X
Alfons Jhm HoubenDepartment of Internal Medicine, Maastricht University Medical Center+, Maastricht, Netherlands.ORCID https://orcid.org/0000-0002-1747-8452
Nicole Dukers-MuijrersDepartment of Health Promotion, Care and Public Health Research Institute, Maastricht University, Maastricht, Netherlands.ORCID https://orcid.org/0000-0003-4896-758X
Joop P van den BerghDepartment of Internal Medicine, Maastricht University Medical Center+, Maastricht, Netherlands.ORCID https://orcid.org/0000-0003-3984-2232
Abraham A KroonDepartment of Internal Medicine, Maastricht University Medical Center+, Maastricht, Netherlands.ORCID https://orcid.org/0000-0001-7750-8249
Maastricht Study ManagementSee Acknowledgments, .
Iris M KaneraFaculty of Health, School of Physiotherapy, Research Center for Nutrition, Lifestyle and Exercise, Zuyd University of Applied Sciences, Heerlen, Netherlands.ORCID https://orcid.org/0000-0001-6863-2096

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundModern lifestyle risk factors, like physical inactivity and poor nutrition, contribute to rising rates of obesity and chronic diseases like type 2 diabetes and heart disease. Particularly personalized interventions have been shown to be effective for long-term behavior change. Machine learning can be used to uncover insights without predefined hypotheses, revealing complex relationships and distinct population clusters. New data-driven approaches, such as the factor probabilistic distance clustering algorithm, provide opportunities to identify potentially meaningful clusters within large and complex datasets.

objectiveThis study aimed to identify potential clusters and relevant variables among individuals with obesity using a data-driven and hypothesis-free machine learning approach.

methodsWe used cross-sectional data from individuals with abdominal obesity from The Maastricht Study. Data (2971 variables) included demographics, lifestyle, biomedical aspects, advanced phenotyping, and social factors (cohort 2010). The factor probabilistic distance clustering algorithm was applied in order to detect clusters within this high-dimensional data. To identify a subset of distinct, minimally redundant, predictive variables, we used the statistically equivalent signature algorithm. To describe the clusters, we applied measures of central tendency and variability, and we assessed the distinctiveness of the clusters through the emerged variables using the F test for continuous variables and the chi-square test for categorical variables at a confidence level of α=.001.

resultsWe identified 3 distinct clusters (including 4128/9188, 44.93% of all data points) among individuals with obesity (n=4128). The most significant continuous variable for distinguishing cluster 1 (n=1458) from clusters 2 and 3 combined (n=2670) was the lower energy intake (mean 1684, SD 393 kcal/day vs mean 2358, SD 635 kcal/day; P<.001). The most significant categorical variable was occupation (P<.001). A significantly higher proportion (1236/1458, 84.77%) in cluster 1 did not work compared to clusters 2 and 3 combined (1486/2670, 55.66%; P<.001). For cluster 2 (n=1521), the most significant continuous variable was a higher energy intake (mean 2755, SD 506.2 kcal/day vs mean 1749, SD 375 kcal/day; P<.001). The most significant categorical variable was sex (P<.001). A significantly higher proportion (997/1521, 65.55%) in cluster 2 were male compared to the other 2 clusters (885/2607, 33.95%; P<.001). For cluster 3 (n=1149), the most significant continuous variable was overall higher cognitive functioning (mean 0.2349, SD 0.5702 vs mean -0.3088, SD 0.7212; P<.001), and educational level was the most significant categorical variable (P<.001). A significantly higher proportion (475/1149, 41.34%) in cluster 3 received higher vocational or university education in comparison to clusters 1 and 2 combined (729/2979, 24.47%; P<.001).

conclusionsThis study demonstrates that a hypothesis-free and fully data-driven approach can be used to identify distinguishable participant clusters in large and complex datasets and find relevant variables that differ within populations with obesity.

Indexed as

Machine LearningObesityAdultAlgorithmsCluster AnalysisCross-Sectional StudiesFemaleHumansLife StyleMaleMiddle AgedRisk Factorschronic diseasecluster analysisdiabetesfactor probabilistic distance clusteringFPDC algorithmheart diseasehypothesis freelong-term behavior changeMaastricht Studyobesityparticipant clustersphysical activityphysical inactivitypoor nutritionrisk factorSES feature selectionstatistically equivalent signaturetype 2 diabetesunsupervised machine learning

Identifiers

PMID39908080
PMCPMC11840370

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.