Evidence map›Paper›PMID 42701163›Full record

ArticleInsights into imaging2026

From zero-shot to fine-tuning: optimize large language models for error detection of ultrasound reports.

Min Lai, Jin Zhang, Yaer Lv, Chunjie Hou, Ceng Wang, Fengzhi Li, Songcheng Xu, Meixia Du, Min Wei, Dong Xu and 1 more

Abstract read
In one paragraph

Article in Insights into imaging, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Min Lai *Cancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Jin Zhang *Department of Ultrasound Medicine, Hangzhou Women's Hospital (Hangzhou Maternity and Child Health Care Hospital), Hangzhou, China.
Yaer LvCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Chunjie HouCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Ceng WangCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Fengzhi LiCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Songcheng XuCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Meixia DuCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Min WeiCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
Dong XuDepartment of Diagnostic Ultrasound Imaging & Interventional Therapy, Zhejiang Cancer Hospital, Hangzhou, China. xudong@zjcc.org.cn.
Litao SunCancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China. litaosun1971@sina.com.ORCID http://orcid.org/0000-0002-4724-0971

Funding

Key Research and Development Project of Vanguard and Leading Goose in Zhejiang Province Key Research and Development Project of Vanguard and Leading Goose in Zhejiang ProvinceMajor Science and Technology Program for Health in Zhejiang Province WKJ-ZJ-2546Science and Technology Program of Zhejiang Province for Traditional Chinese Medicine 2025ZL028
6 · The paper itself

Abstract

backgroundHigh workload and inconsistent quality of ultrasound report writing may lead to diagnostic errors. This study aims to ascertain whether fine-tuned open-source large language models (LLMs) can achieve promising performance for automated quality control of Chinese ultrasound reports, when compared to proprietary LLMs. MATERIALS AND

methodsThis retrospective, multi-center study included a multi-subspecialty dataset of 1800 Chinese ultrasound reports, comprising 1500 quality-controlled reports injected artificially with six predefined error types and 300 reports with naturally occurring errors. Nine proprietary LLMs (under zero-shot and few-shot paradigms) and seven open-source LLMs (under fine-tuning) were evaluated, with performance compared against that of radiologists of varying seniority. Performance was measured by detection accuracy, Macro-F1 score, precision, recall, and mean absolute error across six categories.

resultsFine-tuned open-source LLMs, notably Qwen3-14B, achieved a detection accuracy of 0.931 and a Macro-F1 of 0.739, approaching the performance of senior radiologists. Some fine-tuned open-source LLMs maintained performance despite smaller parameter sizes and outperformed most proprietary LLMs with vastly larger parameter counts. The fine-tuned Qwen3-14B demonstrated superior recognition capability for semantic errors such as redundancy, spelling, orientation, and unit or value errors.

conclusionThis study demonstrates that task-specific fine-tuning enables open-source LLMs to rival proprietary LLMs and expert radiologists in Chinese ultrasound report error detection, offering a locally deployable and privacy-compliant alternative for AI-assisted clinical quality control workflows. KEY POINTS: Question Can task-specific fine-tuning improve error detection by open-source LLMs in Chinese ultrasound reports and provide an effective approach to automated report quality control? Findings Task-specific fine-tuning enabled open-source LLMs to achieve Macro-F1 scores up to 0.739, approaching that of experienced radiologists (0.764) and outperforming most proprietary LLMs. Critical relevance statement Task-specific fine-tuning allows open-source LLMs to become feasible assistants in ultrasound report quality control workflows, offering a locally deployable, privacy-preserving, and regulation-compliant solution for enhancing reporting consistency and reducing diagnostic errors in high-volume ultrasound examination procedures.

Indexed as

Diagnostic errorsLarge language modelsOpen sourceUltrasonic diagnosis

Identifiers

PMID42701163
PMCPMC13546399

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.