Evidence mapPaperPMID 40206348Full record

ArticleComputational and structural biotechnology journal2025

Artificial intelligence-large language models (AI-LLMs) for reliable and accurate cardiotocography (CTG) interpretation in obstetric practice.

Khanisyah Erza Gumilar, Manggala Pasca Wardhana, Muhammad Ilham Aldika Akbar, Agung Sunarko Putra, Dharma Putra Perjuangan Banjarnahor, Ryan Saktika Mulyana, Ita Fatati, Zih-Ying Yu, Yu-Cheng Hsu, Erry Gumilar Dachlan and 3 more

Abstract read
In one paragraph

Article in Computational and structural biotechnology journal, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Journal of clinical medicine · 2025
    Review
  5. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Khanisyah Erza GumilarGraduate Institute of Biomedical Science, China Medical University, Taichung, Taiwan.
Manggala Pasca WardhanaDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Muhammad Ilham Aldika AkbarDepartment of Obstetrics and Gynecology, Universitas Airlangga Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Agung Sunarko PutraDepartment of Obstetrics and Gynecology, Dr. Ramelan Naval Hospital, Surabaya, Indonesia.
Dharma Putra Perjuangan BanjarnahorDepartment of Obstetrics and Gynecology, Dr. Mohamad Soewandhie Hospital, Surabaya, Indonesia.
Ryan Saktika MulyanaDepartment of Obstetrics and Gynecology, Department of Obstetrics and Gynecology, Udayana University Hospital, Denpasar, Indonesia.
Ita FatatiDepartment of Obstetrics and Gynecology, Bandung Kiwari General Hospital, Bandung, Indonesia.
Zih-Ying YuDepartment of Public Health, China Medical University, Taichung, Taiwan.
Yu-Cheng HsuDepartment of Public Health, China Medical University, Taichung, Taiwan.
Erry Gumilar DachlanDepartment of Obstetrics and Gynecology, Universitas Airlangga Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Chien-Hsing LuDepartment of Obstetrics and Gynecology, Taichung Veteran General Hospital, Taichung, Taiwan.
Li-Na LiaoDepartment of Public Health, China Medical University, Taichung, Taiwan.
Ming TanGraduate Institute of Biomedical Science, China Medical University, Taichung, Taiwan.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Accurate cardiotocography (CTG) interpretation is vital for the monitoring of fetal well-being during pregnancy and labor. Advanced artificial intelligence (AI) tools such as AI-large language models (AI-LLMs) may enhance the accuracy of CTG interpretation, but their potential has not been extensively evaluated. Objective: This study aimed to assess the performance of three AI-LLMs (ChatGPT-4o, Gemini Advanced, and Copilot) in CTG image interpretation, compare their results to those of junior (JHDs) and senior human doctors (SHDs), and evaluate their reliability in clinical decision-making. Study design: Seven CTG images were interpreted by the three AI-LLMs, five SHDs, and five JHDs, with the evaluations scored by five blinded maternal-fetal medicine experts using a Likert scale for five parameters (relevance, clarity, depth, focus, and coherence). The homogeneity of the expert ratings and group performances were statistically compared. Results: ChatGPT-4o scored 77.86, outperforming the Gemini Advanced (57.14), Copilot (47.29), and JHDs (61.57). Its performance closely approached that of the SHDs (80.43), with no statistically significant difference between the two (p > 0.05). ChatGPT-4o excelled in the depth parameter and was only marginally inferior to the SHDs regarding the other parameters. Conclusion: ChatGPT-4o demonstrated superior performance among the AI-LLMs, surpassed JHDs in CTG interpretation, and closely matched the performance level of SHDs. AI-LLMs, particularly ChatGPT-4o, are promising tools for assisting obstetricians, improving diagnostic accuracy, and enhancing obstetric patient care.

Indexed as

Artificial intelligence-large language models (AI-LLMs)Cardiotocography (CTG)ChatGPTCopilotFetal monitoringGeminiObstetrics

Identifiers

PMID40206348
PMCPMC11981782

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.