Evidence map›Paper›PMID 42504271›Full record

ArticleTherapeutic advances in drug safety2026

Development and prospective validation of a machine learning model for risk stratification of drug-induced liver injury using real-world clinical data.

Ngan Thi Tran, Tung My Pham, Mai Thi Quynh Ngo, Anh Van Tran, Thu Thi Kim Ninh, Long Duc Nguyen, Dung Van Hoang, Phuong Thi Thu Nguyen

Abstract read
In one paragraph

Article in Therapeutic advances in drug safety, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Ngan Thi TranFaculty of Pharmacy, Hai Phong University of Medicine and Pharmacy, Hai Phong, Vietnam.
Tung My PhamFaculty of Pharmacy, Hai Phong University of Medicine and Pharmacy, Hai Phong, Vietnam.
Mai Thi Quynh NgoFaculty of Pharmacy, Hai Phong University of Medicine and Pharmacy, Hai Phong, Vietnam.
Anh Van TranFaculty of Pharmacy, Hai Phong University of Medicine and Pharmacy, Hai Phong, Vietnam.
Thu Thi Kim NinhFaculty of Pharmacy, Hai Phong University of Medicine and Pharmacy, Hai Phong, Vietnam.
Long Duc NguyenDepartment of Pharmacy, Hai Phong International Hospital, Hai Phong, Vietnam.
Dung Van HoangDepartment of Internal Medicine, Hai Phong International Hospital, Hai Phong, Vietnam.
Phuong Thi Thu NguyenFaculty of Pharmacy, Hai Phong University of Medicine and Pharmacy, Hai Phong 180000, Vietnam.ORCID https://orcid.org/0000-0003-0523-0852

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Drug-induced liver injury (DILI) is difficult to diagnose and manage in routine care because it lacks pathognomonic biomarkers and is often recognized only after clinically meaningful injury has occurred. Existing computational approaches are largely drug-centric and do not routinely incorporate patient-level clinical data available in electronic health records (EHRs). Objectives: To develop and temporally validate a machine learning model for episode-level risk stratification of Roussel Uclaf Causality Assessment Method (RUCAM)-defined DILI using routinely available baseline clinical data. Design: This was an observational cohort study conducted at Hai Phong International Hospital using linked EHR, laboratory, and pharmacy data. The final labeled cohort was partitioned chronologically at the patient level into a retrospective development cohort (2019-2023) and a temporally subsequent prospective validation cohort (2024-2025). Methods: Eligible drug-exposure episodes with complete baseline liver biochemistry and key exposure covariates were included. Analysis-ready episodes were monitored for biochemical liver injury triggers, and trigger-positive episodes underwent clinical review and RUCAM adjudication. DILI was defined as RUCAM ⩾6. Predictors were limited to baseline demographics, comorbidities, laboratory values, drug-exposure features, and FDA DILIrank 2.0 metadata. Candidate models included logistic regression, elastic-net logistic regression, random forest, ExtraTrees, XGBoost, and LightGBM. Results: Among 5095 eligible episodes from 3579 patients, 2786 episodes from 2712 patients were analysis-ready after exclusions. The final labeled cohort comprised 274 DILI-positive and 2512 non-DILI episodes. The prospective validation cohort included 828 episodes, of which 108 (13.0%) were DILI-positive. Tree-based ensemble models outperformed regression-based models. Logistic regression achieved an area under the receiver operating characteristic curve (AUROC) of 0.777 and an area under the precision-recall curve (PR-AUC) of 0.349, whereas the final LightGBM model achieved an AUROC of 0.965 (95% CI 0.942-0.983), a PR-AUC of 0.903 (95% CI 0.856-0.943), and a Brier score of 0.034. Conclusion: A prospectively validated machine learning model using routinely collected baseline clinical data showed excellent performance for DILI risk stratification and may strengthen hospital pharmacovigilance.

Indexed as

drug-induced liver injuryelectronic health recordsmachine learningpharmacovigilanceRUCAMtemporal validation

Identifiers

PMID42504271
PMCPMC13401651

What Socratic holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.