Evidence map›Paper›PMID 40986859›Full record

ArticleJMIR cancer2025

Understanding Cancer Survivorship Care Needs Using Amazon Reviews: Content Analysis, Algorithm Development, and Validation Study.

Liwei Wang, Qiuhao Lu, Rui Li, Taylor B Harrison, Heling Jia, Ming Huang, Heidi Dowst, Rui Zhang, Hoda Badr, Jungwei W Fan and 1 more

Abstract readValidation Study
In one paragraph

Article in JMIR cancer, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Liwei WangDepartment of Clinical and Health Informatics, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, 7000 Fannin Street, Suite 600, Houston, TX, 77030, United States, 1 713-500-3900.ORCID 0000-0001-9970-8604
Qiuhao LuDepartment of Health Data Science and Artificial Intelligence, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX, United States.ORCID 0000-0003-2368-8410
Rui LiDepartment of Health Data Science and Artificial Intelligence, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX, United States.ORCID 0009-0009-7292-4011
Taylor B HarrisonDepartment of Artificial Intelligence and Informatics, Mayo Clinic, Rochester, MN, United States.ORCID 0009-0002-9997-291X
Heling JiaDepartment of Artificial Intelligence and Informatics, Mayo Clinic, Rochester, MN, United States.ORCID 0009-0009-5906-6577
Ming HuangDepartment of Health Data Science and Artificial Intelligence, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX, United States.ORCID 0000-0001-7367-3626
Heidi DowstDepartment of Clinical and Health Informatics, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, 7000 Fannin Street, Suite 600, Houston, TX, 77030, United States, 1 713-500-3900.ORCID 0000-0001-6824-8553
Rui ZhangDepartment of Surgery, University of Minnesota, Twin Cities, MN, United States.ORCID 0000-0001-8258-3585
Hoda BadrDepartment of Medical Oncology, Thomas Jefferson University, Philadelphia, PA, United States.ORCID 0000-0002-4549-9111
Jungwei W FanDepartment of Artificial Intelligence and Informatics, Mayo Clinic, Rochester, MN, United States.ORCID 0000-0001-6349-3752
Hongfang LiuDepartment of Health Data Science and Artificial Intelligence, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX, United States.ORCID 0000-0003-2570-3741

Funding

Semi-structured Information Retrieval in Clinical Text for Cohort IdentificationR01LM011934 · NLM · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI HERSH, WILLIAM R, LIU, HONGFANG · 2014 to 2025
$5.1M
A Translational Informatics Framework to Mine Efficacy and Safety of Dietary SupplementsR01AT009457 · NCCIH · UNIVERSITY OF MINNESOTA · PI RUI ZHANG · 2017 to 2026
$4.1M
Predictive modeling of Alzheimer's Disease Related Dementias (ADRD) in the elderly population empowered by knowledge-driven data miningR01HG012748 · NHGRI · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI HONGFANG LIU · 2023 to 2026
$3.1M
Facilitate Observational Studies of Alzheimer's Disease and Alzheimer's Disease-Related Dementias Using Ontology and Natural Language ProcessingR01AG072799 · NIA · YALE UNIVERSITY · PI HONGFANG LIU, Cui Tao · 2025 to 2026
$1.7M
NCCIH NIH HHS R01 AT009457NHGRI NIH HHS R01 HG012748NIA NIH HHS R01 AG072799NLM NIH HHS R01 LM011934
6 · The paper itself

Abstract

Background: Complementary therapies are being increasingly used by cancer survivors. As a channel for customers to share their feelings, outcomes, and perceived knowledge about the products purchased from e-commerce platforms, Amazon consumer reviews are a valuable real-world data source for understanding cancer survivorship care needs. Objective: In this study, we aimed to highlight the potential of using Amazon consumer reviews as a novel source for identifying cancer survivorship care needs, particularly related to symptom self-management. Specifically, we present a publicly available, manually annotated corpus derived from Amazon reviews of health-related products and develop baseline natural language processing models using deep learning and large language model (LLM) to demonstrate the usability of this dataset. Methods: We preprocessed the Amazon review dataset to identify sentences with cancer mentions through a rule-based method and conducted content analysis including text feature analysis, sentiment analysis, topic modeling, cancer type, and symptom association analysis. We then designed an annotation guideline, targeting survivorship-relevant constructs. A total of 159 reviews were annotated, and baseline models were developed based on deep learning and large language model (LLM) for named entity recognition and text classification tasks. Results: A total of 4703 sentences containing positive cancer mentions were identified, drawn from 3349 reviews associated with 2589 distinct products. The identified topics through topic modeling revealed meaningful insights into cancer symptom management and survivorship experiences. Examples included discussions of green tea use during chemotherapy, cancer prevention strategies, and product recommendations for breast cancer. Top 15 symptoms in reviews were also identified, with pain being the most frequent symptom, followed by inflammation, fatigue, etc. The annotation labels were designed to capture cancer types, indicated symptoms, and symptom management outcomes. The resulting annotation corpus contains 2067 labels from 159 Amazon reviews. It is publicly accessible, together with the annotation guideline through the Open Health Natural Language Processing (OHNLP) GitHub. Our baseline model, Bert-base-cased, achieved the highest weighted average F1-score, that is, 66.92%, for named entity recognition, and LLM gpt4-1106-preview-chat achieved the highest F1-score for text classification tasks, that is, 66.67% for "Harmful outcome," 88.46% for "Favorable outcome" and 73.33% for "Ambiguous outcome." Conclusions: Our results demonstrate the potential of Amazon consumer reviews as a novel data source for identifying persistent symptoms, concerns, and self-management strategies among cancer survivors. This corpus, along with the baseline natural language processing models developed for named entity recognition and text classification, lays the groundwork for future methodological advancements in cancer survivorship research. Importantly, insights from this study could be evaluated against established clinical guidelines for symptom management in cancer survivorship care. By revealing the feasibility of using consumer-generated data for mining survivorship-related experiences, this study offers a promising foundation for future research and argumentation analysis aimed at improving long-term outcomes and support for cancer survivors.

Indexed as

Cancer SurvivorsComplementary TherapiesHealth Services Needs and DemandNeoplasmsSurvivorshipAlgorithmsDeep LearningHumansNatural Language Processingannotationbaseline modelscancer researchcancer survivorship caredeep learninglarge language modelnatural language processingreal-world data

Identifiers

PMID40986859
PMCPMC12456872

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.