Evidence map›Paper›PMID 39610903›Full record

ArticleComputational and structural biotechnology journal2024

Assessment of Large Language Models (LLMs) in decision-making support for gynecologic oncology.

Khanisyah Erza Gumilar, Birama R Indraprasta, Ach Salman Faridzi, Bagus M Wibowo, Aditya Herlambang, Eccita Rahestyningtyas, Budi Irawan, Zulkarnain Tambunan, Ahmad Fadhli Bustomi, Bagus Ngurah Brahmantara and 15 more

Abstract read
In one paragraph

Article in Computational and structural biotechnology journal, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.

0numbers the graph read from it
0cells of the map it votes in
14citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

14 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

25 authors.

Khanisyah Erza GumilarGraduate Institute of Biomedical Science, China Medical University, Taichung, Taiwan.
Birama R IndraprastaDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Ach Salman FaridziDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Bagus M WibowoDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Aditya HerlambangDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Eccita RahestyningtyasDepartment of Obstetrics and Gynecology, Hospital of Universitas Airlangga - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Budi IrawanDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Zulkarnain TambunanDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Ahmad Fadhli BustomiDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Bagus Ngurah BrahmantaraDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Zih-Ying YuDepartment of Public Health, China Medical University, Taichung, Taiwan.
Yu-Cheng HsuDepartment of Public Health, China Medical University, Taichung, Taiwan.
Herlangga PramudityaDepartment of Obstetrics and Gynecology, Dr. Ramelan Naval Hospital, Surabaya, Indonesia.
Very Great E PutraDepartment of Obstetrics and Gynecology, Dr. Kariadi Central General Hospital, Semarang, Indonesia.
Hari NugrohoDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Pungky MulawardhanaDepartment of Obstetrics and Gynecology, Hospital of Universitas Airlangga - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Brahmana A TjokroprawiroDepartment of Obstetrics and Gynecology, Dr. Soetomo General Hospital - Faculty of Medicine, Universitas Airlangga, Surabaya, Indonesia.
Tri HediantoFaculty of Medicine and Health, Institut Teknologi Sepuluh Nopember, Surabaya, Indonesia.
Ibrahim H IbrahimGraduate Institute of Biomedical Science, China Medical University, Taichung, Taiwan.
Jingshan HuangSchool of Computing, College of Medicine, University of South Alabama, Mobile, AL, USA.
Dongqi LiSchool of Information and Computer Sciences, School of Social and Behavioral Sciences, University of California, Irvine, CA, USA.
Chien-Hsing LuDepartment of Obstetrics and Gynecology, Taichung Veteran General Hospital, Taichung, Taiwan.
Jer-Yen YangGraduate Institute of Biomedical Science, China Medical University, Taichung, Taiwan.
Li-Na LiaoDepartment of Public Health, China Medical University, Taichung, Taiwan.
Ming TanGraduate Institute of Biomedical Science, China Medical University, Taichung, Taiwan.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: This study investigated the ability of Large Language Models (LLMs) to provide accurate and consistent answers by focusing on their performance in complex gynecologic cancer cases. Background: LLMs are advancing rapidly and require a thorough evaluation to ensure that they can be safely and effectively used in clinical decision-making. Such evaluations are essential for confirming LLM reliability and accuracy in supporting medical professionals in casework. Study design: We assessed three prominent LLMs-ChatGPT-4 (CG-4), Gemini Advanced (GemAdv), and Copilot-evaluating their accuracy, consistency, and overall performance. Fifteen clinical vignettes of varying difficulty and five open-ended questions based on real patient cases were used. The responses were coded, randomized, and evaluated blindly by six expert gynecologic oncologists using a 5-point Likert scale for relevance, clarity, depth, focus, and coherence. Results: GemAdv demonstrated superior accuracy (81.87 %) compared to both CG-4 (61.60 %) and Copilot (70.67 %) across all difficulty levels. GemAdv consistently provided correct answers more frequently (>60 % every day during the testing period). Although CG-4 showed a slight advantage in adhering to the National Comprehensive Cancer Network (NCCN) treatment guidelines, GemAdv excelled in the depth and focus of the answers provided, which are crucial aspects of clinical decision-making. Conclusion: LLMs, especially GemAdv, show potential in supporting clinical practice by providing accurate, consistent, and relevant information for gynecologic cancer. However, further refinement is needed for more complex scenarios. This study highlights the promise of LLMs in gynecologic oncology, emphasizing the need for ongoing development and rigorous evaluation to maximize their clinical utility and reliability.

Indexed as

AccuracyArtificial intelligenceConsistencyGynecologic cancerLarge Language Models

Identifiers

PMID39610903
PMCPMC11603009

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.