Evidence map›Paper›PMID 40179876›Full record

ArticleMed (New York, N.Y.)2025

OnSIDES database: Extracting adverse drug events from drug labels using natural language processing models.

Yutaro Tanaka, Hsin Yi Chen, Pietro Belloni, Undina Gisladottir, Jenna Kefeli, Jason Patterson, Apoorva Srinivasan, Michael Zietz, Gaurav Sirdeshmukh, Jacob Berkowitz and 2 more

Abstract read
In one paragraph

Article in Med (New York, N.Y.), 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Review
  9. Article
  10. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Yutaro TanakaDepartment of Biomedical Informatics, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA; Department of Applied Physics and Applied Mathematics, Fu Foundation School of Engineering and Applied Sciences, Columbia University, New York, NY 10027, USA.
Hsin Yi ChenDepartment of Biomedical Informatics, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA.
Pietro BelloniDepartment of Statistical Sciences, University of Padova, Padova, Italy.
Undina GisladottirDepartment of Biomedical Informatics, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA.
Jenna KefeliDepartment of Systems Biology, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA.
Jason PattersonDepartment of Biomedical Informatics, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA.
Apoorva SrinivasanDepartment of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA; Cedars-Sinai Cancer, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA.
Michael ZietzDepartment of Biomedical Informatics, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA; Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA; Cedars-Sinai Cancer, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA.
Gaurav SirdeshmukhDepartment of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA; Cedars-Sinai Cancer, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA.
Jacob BerkowitzDepartment of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA; Cedars-Sinai Cancer, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA.
Kathleen LaRow BrownDepartment of Biomedical Informatics, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA.
Nicholas P TatonettiDepartment of Biomedical Informatics, Columbia University Irving Medical Center, Columbia University, New York, NY 10032, USA; Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA; Cedars-Sinai Cancer, Cedars-Sinai Medical Center, Los Angeles, CA 90069, USA; Herbert Irving Comprehensive Cancer Center, New York Presbyterian/Columbia University Irving Medical Center, New York, NY 10032, USA. Electronic address: nicholas.tatonetti@cshs.org.

Funding

Training in Biomedical Informatics at Columbia UniversityT15LM007079 · NLM · COLUMBIA UNIV NEW YORK MORNINGSIDE · PI NOEMIE ELHADAD, GEORGE M HRIPCSAK · 1992 to 2026
$28.9M
Medical Scientist Training ProgramT32GM145440 · NIGMS · COLUMBIA UNIVERSITY HEALTH SCIENCES · PI STEVEN L REINER · 2022 to 2026
$7.3M
Precision Pharmacology and Pharmacovigilance: Leveraging AI to address drug safety knowledge gapsR35GM131905 · NIGMS · COLUMBIA UNIVERSITY HEALTH SCIENCES · PI Nicholas P Tatonetti · 2019 to 2026
$3.3M
NIGMS NIH HHS R35 GM131905NIGMS NIH HHS T32 GM145440NLM NIH HHS T15 LM007079
6 · The paper itself

Abstract

backgroundAdverse drug events (ADEs) are the fourth leading cause of death in the US and cost billions of dollars annually in increased healthcare costs. However, few machine-readable databases of ADEs exist, limiting our capacity to study drug safety on a broader, systematic scale. Recent advances in natural language processing methods, such as BERT models, present an opportunity to accurately extract relevant information from unstructured biomedical text.

methodsWe fine-tune a PubMedBERT model to extract ADE terms from text in FDA Structured Product Labels for prescription drugs. Here, we present OnSIDES (on-label side effects resource), a compiled, machine-friendly database of drug-ADE pairs generated with this method. We further utilize this method to extract pediatric-specific ADEs, serious ADEs from labels' "Boxed Warnings" section, and ADEs from drug labels of other major nations-the UK, the European Union, and Japan-to build a complementary OnSIDES-INTL database. To present OnSIDES' potential applications, we leverage the database to predict novel drug targets and indications, analyze enrichment of ADEs across drug classes, and predict novel ADEs from chemical compound structures.

findingsWe achieve an F1 score of 0.90, AUROC of 0.92, and AUPR of 0.95 at extracting ADEs from the labels' "Adverse Reactions" section. OnSIDES contains over 3.6 million drug-ADE pairs for 3,233 unique drug ingredient combinations extracted from 47,211 labels.

conclusionsOnSIDES can be used as a comprehensive resource to study and enhance drug safety.

fundingR35GM131905 to N.P.T.; T32GM145440 to H.Y.C.; and T15LM007079 to U.G., M.Z., and K.L.B.

Indexed as

Databases, FactualDrug LabelingDrug-Related Side Effects and Adverse ReactionsNatural Language ProcessingHumansUnited StatesUnited States Food and Drug Administrationadverse drug eventsBERTdrug labelsdrug safetyFoundational researchnatural language processingpharmacovigilance

Identifiers

PMID40179876
PMCPMC12256195

What Socratic holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.