Evidence map›Paper›PMID 40134877›Full record

ArticlePeerJ. Computer science2025

Generative artificial intelligence and machine learning methods to screen social media content.

Kellen Sharp, Rachel R Ouellette, Rujula Singh Rajendra Singh, Elise E DeVito, Neil Kamdar, Amanda de la Noval, Dhiraj Murthy, Grace Kong

Abstract read
In one paragraph

Article in PeerJ. Computer science, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Kellen Sharp *Department of Radio-Television-Film, University of Texas at Austin, Austin, Texas, United States.ORCID 0009-0005-5519-9787
Rachel R Ouellette *Department of Psychiatry, Yale School of Medicine, New Haven, Connecticut, United States.ORCID 0000-0001-6982-0830
Rujula Singh Rajendra SinghDepartment of Computer Science, University of Texas at Austin, Austin, Texas, United States.
Elise E DeVitoDepartment of Psychiatry, Yale School of Medicine, New Haven, Connecticut, United States.
Neil KamdarDepartment of Computer Science, University of Texas at Austin, Austin, Texas, United States.
Amanda de la NovalDepartment of Psychiatry, Yale School of Medicine, New Haven, Connecticut, United States.ORCID 0009-0000-0808-0488
Dhiraj MurthySchool of Journalism and Media, University of Texas at Austin, Austin, Texas, United States.ORCID 0000-0001-9734-1124
Grace KongDepartment of Psychiatry, Yale School of Medicine, New Haven, Connecticut, United States.

Funding

Yale Center for the Study of Tobacco Product Use and Addiction (YCSTP) Project 2: Addictive Threshold of Nicotine and the Impact of SweetenersU54DA036151 · NIDA · YALE UNIVERSITY · PI SUCHITRA KRISHNAN-SARIN · 2018 to 2026
$41.6M
Research Training Program in Substance Use PreventionT32DA019426 · NIDA · YALE UNIVERSITY · PI TAMI P SULLIVAN · 2005 to 2026
$7.5M
Modified Use of E-cigarettes and Marketing on YouTubeR01DA049878 · NIDA · YALE UNIVERSITY · PI KONG, GRACE · 2020 to 2022
$1.4M
NIDA NIH HHS R01 DA049878NIDA NIH HHS T32 DA019426NIDA NIH HHS U54 DA036151
6 · The paper itself

Abstract

Background: Social media research is confronted by the expansive and constantly evolving nature of social media data. Hashtags and keywords are frequently used to identify content related to a specific topic, but these search strategies often result in large numbers of irrelevant results. Therefore, methods are needed to quickly screen social media content based on a specific research question. The primary objective of this article is to present generative artificial intelligence (AI; Methods: We searched TikTok for pregnancy and vaping content using 70 hashtag pairs related to "pregnancy" and "vaping" ( Results: Our results indicated ChatGPT-4 classified 44.86% of the videos as exclusively related to pregnancy, 36.91% to vaping, and 8.91% as containing both topics. A human reviewer confirmed for vaping and pregnancy content in 45.38% of the TikTok posts identified by ChatGPT as containing relevant content. Human review of 10% of the posts screened out by ChatGPT identified a 99.06% agreement rate for excluded posts. Conclusions: ChatGPT has mixed capacity to screen social media content that has been converted into text data using machine learning techniques such as object detection. ChatGPT's sensitivity was found to be lower than a human coder in the current case example but has demonstrated power for screening out irrelevant content and can be used as an initial pass at screening content. Future studies should explore ways to enhance ChatGPT's sensitivity.

Indexed as

ChatGPTComputer visione-cigaretteENDSGenerative AIMachine learningPregnancySocial mediaTikTokVaping

Identifiers

PMID40134877
PMCPMC11935761

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.