Evidence map›Paper›PMID 41416738›Full record

ArticleNicotine & tobacco research : official journal of the Society for Research on Nicotine and Tobacco2026

Identifying Key Predictors of Smoking Cessation Success: Text-Based Feature Selection Using a Large Language Model.

Thuy T T Le, Jiongxuan Yang, Zimo Zhao, Kaidi Zhang, Wenjun Li, Yan Hu

Abstract read
In one paragraph

Article in Nicotine & tobacco research : official journal of the Society for Research on Nicotine and Tobacco, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

6 authors.

Thuy T T LeDepartment of Health Management and Policy, University of Michigan School of Public Health, Ann Arbor, MI, 48109, United States.ORCID 0000-0002-3106-4045
Jiongxuan YangDepartment of Biostatistics, University of Michigan School of Public Health, Ann Arbor, MI, 48109, United States.
Zimo ZhaoSchool of Data Science, The Chinese University of Hong Kong Shenzhen, Shenzhen, Guangdong, 518172, China.
Kaidi ZhangSchool of Data Science, The Chinese University of Hong Kong Shenzhen, Shenzhen, Guangdong, 518172, China.
Wenjun LiDepartment of Public Health and Center for Health Statistics, University of Massachusetts Lowell, Lowell, MA, 01854, United States.ORCID 0000-0001-5335-7386
Yan HuSchool of Data Science, The Chinese University of Hong Kong Shenzhen, Shenzhen, Guangdong, 518172, China.

Funding

Research Project 3: Modeling the Impact of Tobacco Control Policies on Polytobacco Use and Associated Health DisparitiesU54CA229974 · NCI · UNIVERSITY OF MICHIGAN AT ANN ARBOR · PI David Mendez Emilien · 2018 to 2026
$39.2M
National Cancer Institute of the National Institutes of Health and the Food and Drug Administration Center for Tobacco Products 2U54CA229974NCI NIH HHS U54 CA229974
6 · The paper itself

Abstract

introductionThe most effective way to reduce mortality and morbidity among current smokers is to quit smoking. Although about half of smokers attempted to quit, only one-tenth succeeded in 2022. Understanding key predictors of smoking cessation success would inform smoking cessation interventions and increase quitting rates.

methodsWe analyzed data from waves 5 and 6 of the Population Assessment of Tobacco and Health (PATH) study (December 2018 to November 2021). Using OpenAI's GPT-4.1, we identified the top 45 variables from wave 5 that are highly predictive of 12-month smoking abstinence in wave 6, based on descriptions of survey variables. We then validated the predictive power of the GPT-4.1-selected variables by comparing the performance of eXtreme Gradient Boosting (XGBoost) trained on different sets of variables. Finally, we derived insights into the top 10 variables, ranked according to their SHapley Additive exPlanations values.

resultsThe performance of XGBoost trained with all possible wave 5 variables and the 45 selected variables was almost identical (AUC:0.749 vs AUC:0.752). The top 10 variables included past 30-day smoking frequency, minutes from waking up to smoking first cigarette, important people's views on tobacco use, prevalence of tobacco use among close associates, daily electronic nicotine product use, emotional dependence, and health harm concerns.

conclusionsThe high predictive performance of XGBoost, when trained on the selected variables, underscores the efficiency and efficacy of GPT-4.1-based feature selection. The top 10 variables include various risk factors that have been previously reported in the literature for their influence on smoking behavior. IMPLICATIONS: Our findings do not establish causal relationships between the selected predictors and 12-month smoking abstinence. However, identifying these key predictors provides valuable insights into the factors highly associated with smoking cessation success. This study demonstrates the ability of OpenAI's GPT-4.1 to perform feature selection using only the textual descriptions of variables. The efficient and successful application of GPT-4.1 for variable selection highlights the potential of integrating artificial intelligence tools into tobacco research to guide resource-efficient and targeted intervention strategies.

Indexed as

LanguageSmoking CessationAdultFemaleHumansLarge Language ModelsMaleMiddle Aged

Identifiers

PMID41416738
PMCPMC12747148

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.