Evidence mapPaperPMID 41345701Full record

ArticleBMC sports science, medicine & rehabilitation2025

Comparative evaluation of ChatGPT versions in training program design: scientific approach, accuracy, and practical applicability.

Ayça Genç, Gökhan Recep Aydın, Murat Kasap, Ali Özkan, Alpamys Rakhymzhanov, Hasan Hüseyin Gürcan, Sevim Güllü, Bilal Demirhan

Abstract read
In one paragraph

Article in BMC sports science, medicine & rehabilitation, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Ayça GençBartın University, Faculty of Sport Science, Bartin, Türkiye. aycagenc@bartin.edu.tr.ORCID http://orcid.org/0000-0003-2498-7092
Gökhan Recep AydınBartın University, Faculty of Sport Science, Bartin, Türkiye.ORCID http://orcid.org/0000-0001-8755-226X
Murat KasapBartın University, Faculty of Sport Science, Bartin, Türkiye.ORCID http://orcid.org/0000-0003-4740-7118
Ali ÖzkanYozgat Bozok University, Faculty of Sport Science, Yozgat, Türkiye.ORCID http://orcid.org/0000-0002-2859-2824
Alpamys RakhymzhanovKhoja Akhmet Yassawi International Kazakh-Turkish University, Turkekistan, Kazakhstan.ORCID http://orcid.org/0000-0001-7261-8641
Hasan Hüseyin GürcanBartın University, Faculty of Sport Science, Bartin, Türkiye.ORCID http://orcid.org/0000-0002-2319-0981
Sevim GüllüIstanbul University-Cerrahpaşa, Faculty of Sport Science, Istanbul, Türkiye.ORCID http://orcid.org/0000-0002-8027-8891
Bilal DemirhanKyrgyz-Turkish Manas University, Bishkek, Kyrgyzstan.ORCID http://orcid.org/0000-0002-3063-9863

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

objectiveThis study aimed to comparatively evaluate the scientific approach, accuracy, and practical applicability of different versions of ChatGPT in generating training programs.

methodAdopting a mixed-methods design, the study employed seven distinct instruction sets, each developed with input from seven experts possessing a minimum of 10 years of professional experience (Certified Strength and Conditioning Specialists (CSCS), academicians holding PhDs in Sports Sciences with research focus on exercise physiology and training methodology) in their respective fields. Using these instruction sets, three versions of ChatGPT (ChatGPT-3.5, ChatGPT-4o, and ChatGPT-4.1) were tasked with generating 12-week resistance training programs for a hypothetical, healthy young adult male with a moderate training background. Using a rubric scoring scale, the generated programs were systematically evaluated and scored separately based on the following criteria: compliance with the initial program request, inclusion of literature references, adherence to exercise variety and progressive loading principles, individualization of progression and justification, program modifications, inclusion of warm-up and cool-down components, injury risk considerations, presence of incorrect recommendations and practical applicability, and accessibility.To determine whether differences in mean scores were statistically significant, the Friedman non-parametric test was applied. When significant differences were identified, pairwise comparisons were conducted using the Wilcoxon signed-rank test to determine which groups accounted for these differences. In addition, qualitative data analysis was performed to explore expert evaluations in depth, employing both content analysis and descriptive analysis techniques.

resultsStatistically significant differences were identified among the ChatGPT versions examined in this study: between ChatGPT-4o and ChatGPT-3.5 (p = .018), ChatGPT-4.1 and ChatGPT-3.5 (p = .018), and ChatGPT-4.1 and ChatGPT-4o (p = .018). Expert content analysis further indicated that ChatGPT-4.1 produced responses that were more detailed, internally consistent, and better supported by scientific literature compared to the other versions.

conclusionAlthough the ChatGPT versions examined in this study exhibited certain limitations, they demonstrated the potential to deliver structured exercise programs aligned with established training principles and relevant scientific literature. Nonetheless, to ensure that AI-assisted training plans provide safe, evidence-based, and individualized content, the involvement of qualified human expertise remains essential.

trial registrationNo official trial registration number was assigned.

Indexed as

AI-Assisted coachingArtificial intelligenceChatGPTExercise program designResistance training

Identifiers

PMID41345701
PMCPMC12797729

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.