ArticlePain2025
Racial, ethnic, and sex bias in large language model opioid recommendations for pain management.
Article in Pain, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 12 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
12 citing papers in PubMed.
- Large language models could be applied in personalized out-of-hospital management for breast cancer: a prospective randomized single blind study.Scientific reports · 2025Trial
- Large Language Model Performance and Clinical Reasoning Tasks.JAMA network open · 2026Article
- Clinical Use of Non-Certified Generative AI in Healthcare: Governing the Regulatory Grey Zone from Convenience to Legal Accountability.Journal of medical systems · 2026Article
- Designing Patient-Centered Communication Aids in Pediatric Surgery Using Large Language Models.Journal of pediatric surgery · 2025Article
- The Digital Standardized Patient: An Artificial Intelligence Coach for Cultural Dexterity in Surgical Care.Journal of the American College of Surgeons · 2025Article
- Development and Evaluation of an Artificial Intelligence-Powered Surgical Oral Examination Simulator: A Pilot Study.Mayo Clinic proceedings. Digital health · 2025Article
- A Future of Self-Directed Patient Internet Research: Large Language Model-Based Tools Versus Standard Search Engines.Annals of biomedical engineering · 2025Article
- Implementing large language models in healthcare while balancing control, collaboration, costs and security.NPJ digital medicine · 2025Article
- Article
- Diagnostic Accuracy of a Custom Large Language Model on Rare Pediatric Disease Case Reports.American journal of medical genetics. Part A · 2025Article
- Pilot Study of Large Language Models as an Age-Appropriate Explanatory Tool for Chronic Pediatric Conditions.medRxiv : the preprint server for health sciences · 2024Article
- Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
Abstract
abstractUnderstanding how large language model (LLM) recommendations vary with patient race/ethnicity provides insight into how LLMs may counter or compound bias in opioid prescription. Forty real-world patient cases were sourced from the MIMIC-IV Note dataset with chief complaints of abdominal pain, back pain, headache, or musculoskeletal pain and amended to include all combinations of race/ethnicity and sex. Large language models were instructed to provide a subjective pain rating and comprehensive pain management recommendation. Univariate analyses were performed to evaluate the association between racial/ethnic group or sex and the specified outcome measures-subjective pain rating, opioid name, order, and dosage recommendations-suggested by 2 LLMs (GPT-4 and Gemini). Four hundred eighty real-world patient cases were provided to each LLM, and responses included pharmacologic and nonpharmacologic interventions. Tramadol was the most recommended weak opioid in 55.4% of cases, while oxycodone was the most frequently recommended strong opioid in 33.2% of cases. Relative to GPT-4, Gemini was more likely to rate a patient's pain as "severe" (OR: 0.57 95% CI: [0.54, 0.60]; P < 0.001), recommend strong opioids (OR: 2.05 95% CI: [1.59, 2.66]; P < 0.001), and recommend opioids later (OR: 1.41 95% CI: [1.22, 1.62]; P < 0.001). Race/ethnicity and sex did not influence LLM recommendations. This study suggests that LLMs do not preferentially recommend opioid treatment for one group over another. Given that prior research shows race-based disparities in pain perception and treatment by healthcare providers, LLMs may offer physicians a helpful tool to guide their pain management and ensure equitable treatment across patient groups.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.