Evidence map›Paper›PMID 42414971›Full record

ArticleBMC public health2026

Development and validation of a machine learning model for evidence of prior or current HCV infection: a demographic screening approach for the US population.

Dorian G Ding, Taoyi Chen, Yu Sheng, Jeffrey S H Lin, Ye Yuan

Abstract readValidation Study
In one paragraph

Article in BMC public health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Dorian G DingCollege of Arts and Sciences, Emory University, Atlanta, USA.
Taoyi ChenSchool of Mathematics and Statistics, Nantong University, Nantong, China.
Yu ShengSchool of Biomedical Engineering, Tsinghua University, Beijing, China.
Jeffrey S H LinDepartment of Cellular and Physiological Sciences, Life Sciences Institute, University of British Columbia, Vancouver, Canada.
Ye YuanState Key Laboratory for Diagnosis and Treatment of Infectious Diseases, The First Affiliated Hospital, Zhejiang University School of Medicine, 15th Floor, Building 6, No. 79 Qingchun Road, Hangzhou, 310003, Zhejiang, China. y.yuan@zju.edu.cn.ORCID https://orcid.org/0000-0003-2661-0635

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundHepatitis C virus (HCV) remains underdiagnosed in the United States despite recommendations for universal screening. A simple approach based on readily available demographic information may help target screening in settings where screening implementation continues to be incomplete.

methodsWe analyzed 10 NHANES cycles (1999-2014 and 2017-2023) and defined HCV exposure as a positive HCV antibody or RNA result. Using sex, birth year, race/ethnicity, birthplace, and income-to-poverty ratio, we trained and compared logistic regression (LR) and machine learning models in training and cycle-based internal validation cohorts (48,434 and 20,762 participants, respectively). Model performance was evaluated based on sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and the area under the receiver operating characteristic curve (AUROC). A web-based calculator was developed to estimate HCV-marker positivity risk and support prioritization of confirmatory HCV testing.

results69,196 participants were included, with 967 showing evidence of HCV exposure. Weighted HCV prevalence remained relatively stable across cycles, ranging from 1.22% to 1.93%. The prevalence did not change significantly after the pandemic. Earlier birth year, male sex, non-Hispanic Black race, US-born status, and lower income-to-poverty ratio were independently associated with HCV exposure. XGBoost performed better than LR in the validation cohort (AUROC 0.859 vs. 0.762, p < 0.001), but PPV remained low (5.1%). In the RNA-positive sensitivity analysis, XGBoost again showed the best validation performance (AUROC 0.867). Predicted risk showed clear risk stratification: observed HCV prevalence increased from 0.05% in the lowest-risk decile to 7.80% in the highest, with the top decile containing 57.2% of participants with HCV exposure and the top three deciles containing 86.2%.

conclusionsFive demographic variables were sufficient to build a supportive HCV-exposure risk stratification model in a nationally representative US sample. Most HCV-exposed individuals were concentrated in the highest predicted-risk groups, suggesting that this approach could help prioritize outreach and confirmatory testing where universal screening uptake remains incomplete. Because no laboratory data are required, it may also be practical in data-limited settings.

Indexed as

Hepatitis CMachine LearningMass ScreeningAdultDemographyFemaleHumansMaleMiddle AgedNutrition SurveysPredictive Learning ModelsPrevalenceSensitivity and SpecificityUnited StatesDemographicsHepatitis CMachine learningPrevalenceScreening

Identifiers

PMID42414971
PMCPMC13587525

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.