ArticleComputational and structural biotechnology journal2024
Assessment of Large Language Models (LLMs) in decision-making support for gynecologic oncology.
Article in Computational and structural biotechnology journal, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
14 citing papers in PubMed.
- Social Status and Clinical Resource Allocation by a Large Language Model: An Evaluation of 30,618 Decisions.Journal of personalized medicine · 2026Article
- The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.Journal of medical Internet research · 2026Article
- Leveraging artificial intelligence to streamline documentation and support patient-centered gynecologic oncology outpatient visits.Gynecologic oncology reports · 2026Article
- Feasibility and Concordance of a Large Language Model (ChatGPT-5) as a Clinical Decision Support Tool in Gynecologic Oncology Tumor Boards: A Blinded, Multi-Observer Study.Journal of clinical medicine · 2026Article
- Article
- Association between Urinary-Extracellular-Vesicle-Enriched Proteome Dynamics and Oncological Outcomes Following Concurrent Chemoradiation in Locally Advanced Cervical Cancer.Computational and structural biotechnology journal · 2026Article
- Can large language models serve as consultants for forensic cause of death analysis? A multidimensional evaluation.Frontiers in artificial intelligence · 2026Article
- Evaluating LLMs in non-metastatic melanoma care: a comparative analysis.Frontiers in medicine · 2026Article
- Exploring transformer models: Fine-tuning VS inference on relation extraction from biomedical texts.Computational and structural biotechnology journal · 2026Article
- Construction and evaluation of the knowledge graph and large model question-answering system for Jin San Zhen therapy: a tool study for primary care and general practice.Frontiers in medicine · 2026Article
- Comparative analysis of ChatGPT 3.5 and ChatGPT 4 obstetric and gynecological knowledge.Scientific reports · 2025Article
- Performance of ChatGPT in Pediatric Audiology as Rated by Students and Experts.Journal of clinical medicine · 2025Article
- Evaluating ChatGPT, Gemini and other Large Language Models (LLMs) in orthopaedic diagnostics: A prospective clinical study.Computational and structural biotechnology journal · 2025Article
- Artificial intelligence-large language models (AI-LLMs) for reliable and accurate cardiotocography (CTG) interpretation in obstetric practice.Computational and structural biotechnology journal · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
25 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Objective: This study investigated the ability of Large Language Models (LLMs) to provide accurate and consistent answers by focusing on their performance in complex gynecologic cancer cases. Background: LLMs are advancing rapidly and require a thorough evaluation to ensure that they can be safely and effectively used in clinical decision-making. Such evaluations are essential for confirming LLM reliability and accuracy in supporting medical professionals in casework. Study design: We assessed three prominent LLMs-ChatGPT-4 (CG-4), Gemini Advanced (GemAdv), and Copilot-evaluating their accuracy, consistency, and overall performance. Fifteen clinical vignettes of varying difficulty and five open-ended questions based on real patient cases were used. The responses were coded, randomized, and evaluated blindly by six expert gynecologic oncologists using a 5-point Likert scale for relevance, clarity, depth, focus, and coherence. Results: GemAdv demonstrated superior accuracy (81.87 %) compared to both CG-4 (61.60 %) and Copilot (70.67 %) across all difficulty levels. GemAdv consistently provided correct answers more frequently (>60 % every day during the testing period). Although CG-4 showed a slight advantage in adhering to the National Comprehensive Cancer Network (NCCN) treatment guidelines, GemAdv excelled in the depth and focus of the answers provided, which are crucial aspects of clinical decision-making. Conclusion: LLMs, especially GemAdv, show potential in supporting clinical practice by providing accurate, consistent, and relevant information for gynecologic cancer. However, further refinement is needed for more complex scenarios. This study highlights the promise of LLMs in gynecologic oncology, emphasizing the need for ongoing development and rigorous evaluation to maximize their clinical utility and reliability.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.