Evidence mapPaperPMID 40410898Full record

ArticleJournal of biomedical semantics2025

Unveiling differential adverse event profiles in vaccines via LLM text embeddings and ontology semantic analysis.

Zhigang Wang, Xingxian Li, Jie Zheng, Yongqun He

Abstract read
In one paragraph

Article in Journal of biomedical semantics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed, 1 pooled it
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. VO: The Vaccine Ontology.Scientific data · 2026
    Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Zhigang WangDepartment of Biomedical Engineering, School of Basic Medicine, Institute of Basic Medical Sciences Chinese Academy of Medical Sciences, Peking Union Medical College, Beijing, China. wangzg@pumc.edu.cn.
Xingxian LiUniversity of Michigan Medical School, Ann Arbor, MI, USA.
Jie ZhengUniversity of Michigan Medical School, Ann Arbor, MI, USA.
Yongqun HeUniversity of Michigan Medical School, Ann Arbor, MI, USA. yongqunh@med.umich.edu.

Funding

VIOLIN 2.0: Vaccine Information and Ontology LInked kNowledgebaseU24AI171008 · UNIVERSITY OF MICHIGAN AT ANN ARBOR · 2025 to 2025
$726k
NIAID NIH HHS U24 AI171008NIH grant U24AI171008Peking Union Medical College 2023jcjg0107
6 · The paper itself

Abstract

backgroundVaccines are crucial for preventing infectious diseases; however, they may also be associated with adverse events (AEs). Conventional analysis of vaccine AEs relies on manual review and assignment of AEs to terms in terminology or ontology, which is a time-consuming process and constrained in scope. This study explores the potential of using Large Language Models (LLMs) and LLM text embeddings for efficient and comprehensive vaccine AE analysis.

resultsWe used Llama-3 LLM to extract AE information from FDA-approved vaccine package inserts for 111 licensed vaccines, including 15 influenza vaccines. Text embeddings were then generated for each vaccine's AEs using the nomic-embed-text and mxbai-embed-large models. Llama-3 achieved over 80% accuracy in extracting AE text from vaccine package inserts. To further evaluate the performance of text embedding, the vaccines were clustered using two clustering methods: (1) LLM text embedding-based clustering and (2) ontology-based semantic similarity analysis. The ontology-based method mapped AEs to the Human Phenotype Ontology (HPO) and Ontology of Adverse Events (OAE), with semantic similarity analyzed using Lin's method. Text embeddings were generated for each vaccine's AE description using the LLM nomic-embed-text and mxbai-embed-large models. Compared to the semantic similarity analysis, the LLM approach was able to capture more differential AE profiles. Furthermore, LLM-derived text embeddings were used to develop a Lasso logistic regression model to predict whether a vaccine is "Live" or "Non-Live". The term "Non-Live" refers to all vaccines that do not contain live organisms, including inactivated and mRNA vaccines. A comparative analysis showed that, despite similar clustering patterns, the nomic-embed-text model outperformed the other. It achieved 80.00% sensitivity, 83.06% specificity, and 81.89% accuracy in a 10-fold cross-validation. Many AE patterns, with examples demonstrated, were identified from our analysis with AE LLM embeddings.

conclusionThis study demonstrates the effectiveness of LLMs for automated AE extraction and analysis, and LLM text embeddings capture latent information about AEs, enabling more comprehensive knowledge discovery. Our findings suggest that LLMs demonstrate substantial potential for improving vaccine safety and public health research.

Indexed as

Biological OntologiesDrug-Related Side Effects and Adverse ReactionsNatural Language ProcessingSemanticsVaccinesHumansVaccinesAdverse eventFDA package insertsHuman phenotype ontology (HPO)Large language models (LLM)LIama-3 modelOntology of adverse events (OAE)VaccineVaccine ontology (VO)

Identifiers

PMID40410898
PMCPMC12102970

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.