ArticleInsights into imaging2024
Utilizing a domain-specific large language model for LI-RADS v2018 categorization of free-text MRI reports: a feasibility study.
Article in Insights into imaging, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 12 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
12 citing papers in PubMed.
- Application of artificial intelligence in paediatric oncology imaging.Pediatric radiology · 2026Review
- Transforming adnexal mass assessment: how large language model improve ovarian-adnexal reporting and data system interpretation and sonographer performance.Abdominal radiology (New York) · 2026Article
- Leveraging Large Language Models for Accurate AO Fracture Classification from CT Text Reports.Journal of imaging informatics in medicine · 2026Article
- Natural language processing and LLMs in liver imaging: a practical review of clinical applications.Abdominal radiology (New York) · 2026Review
- Applying large language model for automated quality scoring of radiology requisitions using a standardized criteria.European radiology · 2026Article
- Evaluation of large language models in assigning PI-RADS v2.1 categories for prostate MRI reports.BMC urology · 2026Article
- Large language models for clinical decision support in gastroenterology and hepatology.Nature reviews. Gastroenterology & hepatology · 2025Review
- Current State of Evidence for Use of MRI in LI-RADS.Journal of magnetic resonance imaging : JMRI · 2025Review
- Article
- Areas of research focus and trends in the research on the application of AIGC in healthcare.Journal of health, population, and nutrition · 2025Article
- Artificial Intelligence-Empowered Radiology-Current Status and Critical Review.Diagnostics (Basel, Switzerland) · 2025Review
- Leveraging large language models for accurate classification of liver lesions from MRI reports.Computational and structural biotechnology journal · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
11 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
objectiveTo develop a domain-specific large language model (LLM) for LI-RADS v2018 categorization of hepatic observations based on free-text descriptions extracted from MRI reports. MATERIAL AND
methodsThis retrospective study included 291 small liver observations, divided into training (n = 141), validation (n = 30), and test (n = 120) datasets. Of these, 120 were fictitious, and 171 were extracted from 175 MRI reports from a single institution. The algorithm's performance was compared to two independent radiologists and one hepatologist in a human replacement scenario, and considering two combined strategies (double reading with arbitration and triage). Agreement on LI-RADS category and dichotomic malignancy (LR-4, LR-5, and LR-M) were estimated using linear-weighted κ statistics and Cohen's κ, respectively. Sensitivity and specificity for LR-5 were calculated. The consensus agreement of three other radiologists served as the ground truth.
resultsThe model showed moderate agreement against the ground truth for both LI-RADS categorization (κ = 0.54 [95% CI: 0.42-0.65]) and the dichotomized approach (κ = 0.58 [95% CI: 0.42-0.73]). Sensitivity and specificity for LR-5 were 0.76 (95% CI: 0.69-0.86) and 0.96 (95% CI: 0.91-1.00), respectively. When the chatbot was used as a triage tool, performance improved for LI-RADS categorization (κ = 0.86/0.87 for the two independent radiologists and κ = 0.76 for the hepatologist), dichotomized malignancy (κ = 0.94/0.91 and κ = 0.87) and LR-5 identification (1.00/0.98 and 0.85 sensitivity, 0.96/0.92 and 0.92 specificity), with no statistical significance compared to the human readers' individual performance. Through this strategy, the workload decreased by 45%.
conclusionLI-RADS v2018 categorization from unlabelled MRI reports is feasible using our LLM, and it enhances the efficiency of data curation. CRITICAL RELEVANCE STATEMENT: Our proof-of-concept study provides novel insights into the potential applications of LLMs, offering a real-world example of how these tools could be integrated into a local workflow to optimize data curation for research purposes. KEY POINTS: Automatic LI-RADS categorization from free-text reports would be beneficial to workflow and data mining. LiverAI, a GPT-4-based model, supported various strategies improving data curation efficiency by up to 60%. LLMs can integrate into workflows, significantly reducing radiologists' workload.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.