Evidence map›Paper›PMID 41315569›Full record

ArticleScientific reports2025

Large language models versus classical machine learning performance in COVID-19 mortality prediction using high-dimensional tabular data.

Mohammadreza Ghaffarzadeh-Esfahani, Mahdi Ghaffarzadeh-Esfahani, Aryan Salahi-Niri, Hossein Toreyhi, Zahra Atf, Amirali Mohsenzadeh-Kermani, Mahshad Sarikhani, Zohreh Tajabadi, Fatemeh Shojaeian, Mohammad Hassan Bagheri and 32 more

Abstract read
In one paragraph

Article in Scientific reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

42 authors.

Mohammadreza Ghaffarzadeh-EsfahaniResearch Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Mahdi Ghaffarzadeh-EsfahaniFaculty of Medicine, Isfahan University of Medical Sciences, Isfahan, Iran.
Aryan Salahi-NiriResearch Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Hossein ToreyhiResearch Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Zahra AtfFaculty of Business and Information Technology, Ontario Tech University, Oshawa, Canada.
Amirali Mohsenzadeh-KermaniFaculty of Medicine, Isfahan University of Medical Sciences, Isfahan, Iran.
Mahshad SarikhaniSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Zohreh TajabadiDigestive Disease Research Institute, Tehran University of Medical Sciences, Tehran, Iran.
Fatemeh ShojaeianDepartment of Surgery, The Johns Hopkins University, Baltimore, MD, USA.
Mohammad Hassan BagheriFaculty of Medicine, Isfahan University of Medical Sciences, Isfahan, Iran.
Aydin FeyziStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Mohamadamin Tarighat-PaymaSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Narges GazmehStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Fateme HeydariSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Hossein AfsharStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Amirreza AllahgholipourStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Farid AlimardaniStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Ameneh SalehiSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Naghmeh AsadimaneshSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Mohammad Amin KhalafiSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Hadis ShabanipourStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Ali MoradiStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Sajjad Hossein ZadehStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Omid YazdaniSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Romina EsbatiSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Moozhan MalekiStudent Research Committee, School of Nursing and Midwifery, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Danial Samiei NasrSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Amirali SoheiliSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Hossein MajlesiSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Saba ShahsavanSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Alireza SoheilipourSchool of Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Nooshin GoudarziResearch Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Erfan TaherifardMPH department, Shiraz University of Medical Sciences, Shiraz, Iran.
Hamidreza HatamabadiDepartment of Emergency Medicine, School of Medicine, Safety Promotion and Injury Prevention Research Center, Imam Hossein Hospital, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Jamil S SamaanKarsh Division of Gastroenterology and Hepatology, Cedars-Sinai Medical Center, 8700 Beverly Blvd, Los Angeles, CA, 90048, USA.
Thomas SavageDepartment of Medicine, Stanford University, Stanford, CA, USA.
Ankit SakhujaDivision of Data Driven and Digital Health (D3M), The Charles Bronfman Institute for Personalized Medicine, Icahn School of Medicine at Mount Sinai, New York, NY, USA.
Ali SoroushDivision of Data Driven and Digital Health (D3M), The Charles Bronfman Institute for Personalized Medicine, Icahn School of Medicine at Mount Sinai, New York, NY, USA.
Girish NadkarniDivision of Data Driven and Digital Health (D3M), The Charles Bronfman Institute for Personalized Medicine, Icahn School of Medicine at Mount Sinai, New York, NY, USA.
Ilad Alavi DarazamInfectious Diseases and Tropical Medicine Research Center, Shahid Beheshti University of Medical Sciences, Tehran, Iran. ilad13@yahoo.com.
Mohamad Amin PourhoseingholiResearch Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran. aminphg@gmail.com.
Seyed Amir Ahmad Safavi-NainiResearch Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran. sdamirsa@ymail.com.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

This study compared the performance of classical feature-based machine learning models (CMLs) and large language models (LLMs) in predicting COVID-19 mortality using high-dimensional tabular data from 9,134 patients across four hospitals. Seven CML models, including XGBoost and random forest (RF), were evaluated alongside eight LLMs, such as GPT-4 and Mistral-7b, which performed zero-shot classification on text-converted structured data. Additionally, Mistral-7b was fine-tuned using the QLoRA approach. XGBoost and RF demonstrated superior performance among CMLs, achieving F1 scores of 0.87 and 0.83 for internal and external validation, respectively. GPT-4 led the LLM category with an F1 score of 0.43, while fine-tuning Mistral-7b significantly improved its recall from 1% to 79%, yielding a stable F1 score of 0.74 during external validation. Although LLMs showed moderate performance in zero-shot classification, fine-tuning substantially enhanced their effectiveness, potentially bridging the gap with CML models. However, CMLs still outperformed LLMs in handling high-dimensional tabular data tasks. This study highlights the potential of both CMLs and fine-tuned LLMs in medical predictive modeling, while emphasizing the current superiority of CMLs for structured data analysis.

Indexed as

COVID-19LanguageMachine LearningFemaleHumansLarge Language ModelsMaleSARS-CoV-2COVID-19 mortalityFine-tuningLarge language modelsMachine learningStructured dataZero-shot classification

Identifiers

PMID41315569
PMCPMC12663554

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.