ArticleJAMIA open2026
Longitudinal symptom severity tracking in vagus nerve stimulation patients: a 2-stage LLM-based pipeline with explainable AI.
Article in JAMIA open, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
13 authors.
Funding
Abstract
Objectives: To produce the first structured longitudinal dataset of symptom severity trajectories in vagus nerve stimulation (VNS) patients using a 2-stage large language model pipeline for automated symptom severity extraction from clinical notes, as part of the NIH-funded research evaluating vagal excitations and anatomical linkages (U54AT012307). Materials and methods: We analyzed 1427 annotated clinical notes from 95 patients across 5 corpora. LLaMA models (3.2-1B, 3.2-3B, and 3.3-70B) were fine-tuned to classify 13 symptoms as present, absent, or negated (stage 1), then assessed severity (indeterminate, mild, moderate, or severe; stage 2). Longitudinal trajectories were tracked quarterly in 23 VNS patients against a 10-patient surgical control cohort (NSQIP). Results: Severity F1-scores ranged from 0.76 to 0.96, and symptom classification reached 0.99 for high-prevalence symptoms, with the validation corpus (VNS vs NSQIP) as the dominant factor (29.4% of variance; no significant source-corpus effect), supporting cross-domain generalizability. Macro F1-scores were lower than weighted F1-scores for smaller models and NSQIP; suicidal-ideation sensitivity was lower on NSQIP than VNS (0.80 vs 1.00, largest model), limiting generalization for psychiatric risk detection. Equity analyses were exploratory given small subgroups (gender precision 0.95). The longitudinal dataset revealed a median time to clinically meaningful seizure improvement of 0.5 quarters, with 50% responders, 25% non-responders, and 25% worsened. Post-surgical decline was absent in the unmatched NSQIP control cohort, consistent with, but not confirming, VNS specificity. Discussion: To our knowledge, this is the first system producing structured symptom trajectories from free-text VNS documentation, supporting individualized over population-level monitoring. Conclusion: This dataset supports future treatment-response prediction, personalized device programming, and a reproducible clinical natural langauge processing benchmark for neuromodulation research.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.