Evidence mapPaperPMID 42258573Full record

ArticleJMIR formative research2026

Supervised Fine-Tuning of Large Language Models With Chain-of-Thought Reasoning for Pediatric Heart Disease Detection in Unstructured Echocardiogram Reports: Algorithm Development and Validation.

Haoming Shi, Justin B Long, Michael C Fiedorek, Hannah D Kilday, Henry P Foote, Christoph P Hornik, Aditya Nagori, Yifan Xiang, Rishikesan Kamaleswaran

Abstract read
In one paragraph

Article in JMIR formative research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Haoming ShiDepartment of Biomedical Engineering, Duke University, 1427 Fitzpatrick Center Box 90281, Durham, NC, 27708, United States, 1 9196695131.ORCID http://orcid.org/0000-0002-0228-6116
Justin B LongDepartment of Anesthesiology, Emory University School of Medicine, Atlanta, GA, United States.ORCID http://orcid.org/0000-0002-7660-6063
Michael C FiedorekDepartment of Anesthesiology, Emory University School of Medicine, Atlanta, GA, United States.ORCID http://orcid.org/0009-0002-0944-4161
Hannah D KildayDepartment of Anesthesiology, Emory University School of Medicine, Atlanta, GA, United States.ORCID http://orcid.org/0009-0003-9198-0787
Henry P FooteDepartment of Pediatrics, Duke University School of Medicine, Durham, NC, United States.ORCID http://orcid.org/0009-0008-5973-2579
Christoph P HornikDepartment of Pediatrics, Duke University School of Medicine, Durham, NC, United States.ORCID http://orcid.org/0000-0001-7056-8759
Aditya NagoriDepartment of Surgery, Duke University School of Medicine, Durham, NC, United States.ORCID http://orcid.org/0000-0002-6389-2179
Yifan XiangDepartment of Biomedical Engineering, Duke University, 1427 Fitzpatrick Center Box 90281, Durham, NC, 27708, United States, 1 9196695131.ORCID http://orcid.org/0009-0001-2931-6100
Rishikesan KamaleswaranDepartment of Biomedical Engineering, Duke University, 1427 Fitzpatrick Center Box 90281, Durham, NC, 27708, United States, 1 9196695131.ORCID http://orcid.org/0000-0001-8366-4811

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Pediatric heart disease (PHD), including congenital heart defects, is often incompletely captured in electronic health records, particularly when clinical significance must be inferred from unstructured echocardiogram reports. Automated methods capable of extracting clinically meaningful PHD from narrative reports could improve clinical decision support and research applications. Objective: The aim of the study is to evaluate the feasibility of using supervised fine-tuning of large language models (LLMs), with and without chain-of-thought (CoT) reasoning, to characterize patients with clinically significant or historical PHD from unstructured echocardiogram reports. Methods: We developed a PHD detection algorithm using fine-tuned open-source LLMs, including LLaMA (Meta) and Qwen (Alibaba), to analyze 9749 echocardiogram reports. A subset of 712 reports was adjudicated by 2 pediatric cardiac anesthesiologists, classifying 506 (71.1%) as clinically significant PHD and 206 (28.9%) as not significant. While DeepSeek R1 has shown improved performance with CoT reasoning, its application in medical contexts is underexplored. We incorporated R1-generated CoT into model prompts and fine-tuned backbone LLMs. Results: The fine-tuned Qwen-7B-10k-overthink-CoT achieved the highest accuracy (92.4%), outperforming Qwen-7B-without-CoT (90%), LLaMA-3B-without-CoT (87.9%), Qwen-3B-without-CoT (85.6%), Qwen-3B-10k-overthink-CoT (68.5%), and LLaMA-3B-10k-overthink-CoT (46.2%). In a second dataset, an external validation was performed (n=113; 64 positive, 49 negative), Qwen-7B-10k-overthink-CoT sustained a strong, balanced performance (82.7%), followed by Qwen-7B-without-CoT (88.4%), LLaMA-3B-without-CoT (86.8%), Qwen-3B-without-CoT (84.5%), Qwen-3B-10k-overthink-CoT (58.9%), and LLaMA-3B-10k-overthink-CoT (46.2%). The fine-tuned Qwen-7B model with overthinking CoT (10,000 tokens) achieved the highest internal accuracy (92.4%), with balanced sensitivity and specificity. Across repeated runs, CoT-enhanced models demonstrated improved classification consistency compared to non-CoT models (Qwen-7B-without-CoT: 90%, LLaMA-3B-without-CoT: 87.9%, Qwen-3B-without-CoT: 85.6%). In external validation (n=113), non-CoT variants achieved higher accuracy (up to 88.4%), whereas the Qwen-7B CoT model demonstrated more balanced class performance (accuracy=82.7%). Conclusions: Supervised fine-tuning of LLMs with CoT offers an effective approach for automated PHD detection within unstructured data in the electronic medical record. While CoT-enhanced models demonstrated improved internal performance and more balanced classification, they did not consistently achieve higher accuracy in external validation, highlighting trade-offs between accuracy and class balance. These findings highlight the promise of LLM-based approaches for clinical text phenotyping while underscoring the need for larger, multicenter validation and careful calibration for real-world deployment. Continued validation and integration into the electronic medical record are essential for real-world, artificial intelligence-driven clinical decision support.

Indexed as

AlgorithmsEchocardiographyHeart DiseasesLarge Language ModelsChildElectronic Health RecordsHeart Defects, CongenitalHumansMalePediatricscongenital heart defectsechocardiogramlarge language modelmachine learningnatural language processing

Identifiers

PMID42258573
PMCPMC13245642

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.