Evidence map›Paper›PMID 40939164›Full record

ArticleJMIR medical informatics2025

Extracting Symptoms of Complex Conditions From Online Discourse (Subreddit to Symptomatology): Lexicon-Based Approach.

Bushra Hossain, Sarah M Preum, Md Fazle Rabbi, Rifat Ara, Mohammed Eunus Ali

Abstract read
In one paragraph

Article in JMIR medical informatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Bushra HossainDepartment of Computer Science and Engineering, Bangladesh University of Engineering and Technology, Dhaka, Bangladesh.ORCID https://orcid.org/0009-0006-0941-0374
Sarah M PreumDepartment of Computer Science, Dartmouth College, Hanover, NH, United States.ORCID https://orcid.org/0000-0002-7771-8323
Md Fazle RabbiDepartment of Computer Science and Engineering, Bangladesh University of Professionals, Dhaka, Bangladesh.ORCID https://orcid.org/0009-0004-1628-4519
Rifat AraDepartment of Obstetrics & Gynecology, Bangladesh College of Physicians and Surgeons, Dhaka, Bangladesh.ORCID https://orcid.org/0009-0007-4191-5492
Mohammed Eunus AliFaculty of Information Technology, Monash University, Clayton, Australia.ORCID https://orcid.org/0000-0002-0384-7616

Funding

Treatment Development & Evaluation CoreP30DA029926 · NIDA · DARTMOUTH COLLEGE · PI Lisa A. Marsch · 2011 to 2026
$21.5M
NIDA NIH HHS P30 DA029926
6 · The paper itself

Abstract

backgroundMillions of people affected with complex medical conditions with diverse symptoms often turn to online discourse to share their experiences. While some studies have explored natural language processing methods and medical information extraction tools, these typically focus on generic symptoms in clinical notes and struggle to identify patient-reported, disease-specific, subtle symptoms from online health discourse.

objectiveWe aimed to extract patient-reported, disease-specific symptoms shared on social media reflecting the lived experiences of thousands of affected individuals and explore the characteristics, prevalence, and occurrence patterns of the symptoms.

methodsWe propose a lexicon-based symptom extraction (LSE) method to identify a diverse list of disease-specific, patient-reported symptoms. We initially used a large language model to accelerate the extraction of symptom-related key phrases that formed the lexicon. We evaluated the effectiveness of lexicon extraction against human annotation using a Jaccard index score. We then leveraged BERT-Base, BioBERT, and Phrase-BERT-based embeddings to learn representations of these symptom-related key phrases and cluster similar symptoms using k-means and hierarchical density-based spatial clustering of applications with noise (HDBSCAN). Among the different options explored in our experiments, BioBERT-based k-means clustering was found to be the most effective. Finally, we applied symptom normalization to eliminate duplicate and redundant entries in the comprehensive symptom list.

resultsIn a real-world polycystic ovary syndrome (PCOS) subreddit dataset, we found that LSE significantly outperformed state-of-the-art baselines, achieving at least 41% and 20% higher F

conclusionsThe comprehensive patient-reported, disease-specific symptom list can help patients and health practitioners resolve uncertainties surrounding the disease, eliminating the variability of PCOS symptoms prevailing in the community. Analyzing PCOS symptomatology across varied dimensions provides valuable insights for public health research.

Indexed as

Data MiningNatural Language ProcessingSocial MediaFemaleHumanscomplex medical conditiondisease-specific symptomshealth information extractionlarge language modelsnatural language processingonline discoursepolycystic ovary syndromesymptomatology

Identifiers

PMID40939164
PMCPMC12475878

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.