ArticleAMIA ... Annual Symposium proceedings. AMIA Symposium2024
Large Language Models Struggle in Token-Level Clinical Named Entity Recognition.
Article in AMIA ... Annual Symposium proceedings. AMIA Symposium, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
14 citing papers in PubMed.
- Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative Evaluation Study.JMIR medical informatics · 2026Article
- Automated identification of incidentalomas requiring follow-up: A multi-anatomy evaluation of LLM-based and supervised approaches.Journal of biomedical informatics · 2026Article
- Auditing frontier general-purpose large language models in biomedical tasks: reasoning gains, extraction limits, and benchmark reliability.Research square · 2026Article
- Leveraging large language models for rare disease named entity recognition.PLOS digital health · 2026Article
- Evidence-based AI: from trailblazer to trustblazer?Frontiers in artificial intelligence · 2026Article
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset.Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing · 2026Article
- Span-based annotation framework for LLM-based clinical named entity recognition: development and validation using Korean emergency department notes.JAMIA open · 2025Article
- Quantifying Global Foreign Affairs with a Multimodal Dataset of Diplomatic Websites.Scientific data · 2025Article
- Exploring multimodal large language models on transthoracic Echocardiogram (TTE) tasks for cardiovascular decision support.Journal of biomedical informatics · 2025Article
- Clinical Information Extraction From Notes of Veterans With Lymphoid Malignancies: Natural Language Processing Study.JMIR medical informatics · 2025Article
- Understanding Cancer Survivorship Care Needs Using Amazon Reviews: Content Analysis, Algorithm Development, and Validation Study.JMIR cancer · 2025Article
- Developing an ICD-10 Coding Assistant: Pilot Study Using RoBERTa and GPT-4 for Term Extraction and Description-Based Code Selection.JMIR formative research · 2025Article
- Article
- Physicians' Knowledge, Perceptions and Use of Large Language Models in Clinical Practice:Sultan Qaboos University medical journal · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
Abstract
Large Language Models (LLMs) have revolutionized various sectors, including healthcare where they are employed in diverse applications. Their utility is particularly significant in the context of rare diseases, where data scarcity, complexity, and specificity pose considerable challenges. In the clinical domain, Named Entity Recognition (NER) stands out as an essential task and it plays a crucial role in extracting relevant information from clinical texts. Despite the promise of LLMs, current research mostly concentrates on document-level NER, identifying entities in a more general context across entire documents, without extracting their precise location. Additionally, efforts have been directed towards adapting ChatGPTfor token-level NER. However, there is a significant research gap when it comes to employing token-level NER for clinical texts, especially with the use of local open-source LLMs. This study aims to bridge this gap by investigating the effectiveness of both proprietary and local LLMs in token-level clinical NER. Essentially, we delve into the capabilities of these models through a series of experiments involving zero-shot prompting, few-shot prompting, retrieval-augmented generation (RAG), and instruction-fine-tuning. Our exploration reveals the inherent challenges LLMs face in token-level NER, particularly in the context of rare diseases, and suggests possible improvements for their application in healthcare. This research contributes to narrowing a significant gap in healthcare informatics and offers insights that could lead to a more refined application of LLMs in the healthcare sector.
Indexed as
Identifiers
40417588PMC12099373What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.