Evidence map›Paper›PMID 42390385›Full record

ArticleJournal of medical Internet research2026

Extracting Medical Information From Unstructured Clinical Text Using Large Language Models to Enhance Health Care Interoperability: Proof-of-Concept Study.

Bahadır Eryılmaz, Kamyar Arzideh, Mikel Bahn, Hendrik Damm, Sina Warmer, Henning Schäfer, Ahmad Idrissi-Yaghir, Tabea M G Pakull, Lea Jessica Albrecht, Jens Kleesiek and 7 more

Abstract read
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

17 authors.

Bahadır Eryılmaz *University Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0009-0002-8743-4751
Kamyar Arzideh *University Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0009-0005-6074-804X
Mikel BahnUniversity Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0009-0002-0866-4023
Hendrik DammDepartment of Computer Science, University of Applied Sciences and Arts Dortmund, Dortmund, NRW, Germany.ORCID http://orcid.org/0000-0002-7464-4293
Sina WarmerUniversity Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0009-0002-2262-2655
Henning SchäferUniversity Hospital Essen, Institute for Transfusion Medicine, Essen, NRW, Germany.ORCID http://orcid.org/0000-0002-4123-0406
Ahmad Idrissi-YaghirUniversity Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0000-0003-1507-9690
Tabea M G PakullDepartment of Computer Science, University of Applied Sciences and Arts Dortmund, Dortmund, NRW, Germany.ORCID http://orcid.org/0009-0009-9802-7167
Lea Jessica AlbrechtDepartment of Dermatology, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0002-0307-6006
Jens KleesiekUniversity Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0000-0001-8686-0682
Georg LoddeDepartment of Dermatology, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0002-6930-2694
Christoph M FriedrichDepartment of Computer Science, University of Applied Sciences and Arts Dortmund, Dortmund, NRW, Germany.ORCID http://orcid.org/0000-0001-7906-0038
Elisabeth LivingstoneDepartment of Dermatology, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0001-8279-9239
Dirk SchadendorfDepartment of Dermatology, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0003-3524-7858
Katarzyna BorysUniversity Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0000-0001-6987-6041
Felix Nensa *University Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0000-0002-5811-7100
René Hosch *University Hospital Essen, Institute for Artificial Intelligence in Medicine (IKIM), Girardetstraße 2, Essen, NRW, 45131, Germany.ORCID http://orcid.org/0000-0003-1760-2342

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Unstructured clinical text remains a major barrier to interoperable data reuse and large-scale secondary analysis in health care. Large language models (LLMs) have the potential to automate the extraction of structured clinical information; however, their application is limited by the scarcity of high-quality annotated training data. Objective: To address these limitations, this study aims to develop and validate a scalable, privacy-preserving framework that uses synthetic data generated from structured Fast Healthcare Interoperability Resources (FHIR) to fine-tune open-source LLMs for the effective extraction of interoperable clinical information from unstructured text. Methods: We evaluated an LLM-based framework for extracting structured clinical information from cancer-related discharge letters and mapping it to representations compatible with FHIR. To enable large-scale supervised training, we developed a random sample generator that creates synthetic discharge letters using Qwen3-235B by randomly sampling and aggregating structured FHIR data from 41,175 patients with cancer. The resulting synthetic discharge letters (n=75,000) were paired with their originating structured data, forming a large-scale dataset for fine-tuning MedGemma 27B, a 27-billion-parameter medical language model. Evaluation was conducted on the synthetic test dataset (n=7500), real-world discharge letters (n=30), which were evaluated by physicians and a medical student, and a comparative one-shot approach using open-source models (Qwen3, LLaMA, and GPT-OSS). Results: The fine-tuned model achieved high extraction performance across multiple clinical entities on the synthetic test set, with F1-scores of 0.84 for full International Classification of Diseases diagnosis codes, 0.99 for tumor-related information, 0.99 for laboratory values, 0.99 for medication names and dosages, and 0.94 for Anatomical Therapeutic Chemical medication codes. The extraction of procedure-related information was more challenging, with F1-scores of 0.63 for OPS codes and 0.90 for procedure descriptions. The fine-tuned model consistently outperformed general-purpose LLMs in a one-shot comparison across nearly all extraction categories. When evaluated by physicians on real-world discharge letters, the model achieved case-level correctness rates of 78.9% for International Classification of Diseases diagnoses, 86.1% for tumor-related information, 93.0% for medications, and 61.3% for procedures. Conclusions: These results demonstrate that synthetic text generation from structured clinical data enables the effective and scalable training of LLMs for extracting interoperable, multientity clinical information from unstructured documentation.

Indexed as

Electronic Health RecordsHealth Information InteroperabilityLarge Language ModelsHumansNeoplasmsProof of Concept StudyAI in health careartificial intelligenceentity extractionFast Healthcare Interoperability ResourcesFHIRgenerative AIinteroperabilitylarge language models

Identifiers

PMID42390385
PMCPMC13325620

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.