Evidence map›Paper›PMID 41239388›Full record

ArticleBMC medical informatics and decision making2025

Can general purpose large language models assist pediatricians in predicting infants with serious bacterial infection?

Ivan Šimunović, Klara Rezić, Nikola Franić, Gabrijel Boduljak, Marijan Batinić, Ivana Jukić, Ivana Jelovina, Jela Biočić, Zenon Pogorelić, Joško Markić

Abstract read
In one paragraph

Article in BMC medical informatics and decision making, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Ivan ŠimunovićSchool of Medicine, University of Split, Split, 21000, Croatia.
Klara RezićSchool of Medicine, University of Split, Split, 21000, Croatia.
Nikola FranićFaculty of Electrical Engineering, Mechanical Engineering and Naval Architecture, University of Split, Split, 21000, Croatia.
Gabrijel BoduljakDepartment of Computer Science, University of Oxford, Oxford, UK.
Marijan BatinićDepartment of Pediatrics, University Hospital of Split, Split, 21000, Croatia.
Ivana JukićDepartment of Pediatrics, University Hospital of Split, Split, 21000, Croatia.
Ivana JelovinaDepartment of Pediatrics, University Hospital of Split, Split, 21000, Croatia.
Jela BiočićSchool of Medicine, University of Split, Split, 21000, Croatia.
Zenon PogorelićSchool of Medicine, University of Split, Split, 21000, Croatia. zenon.pogorelic@mefst.hr.
Joško MarkićSchool of Medicine, University of Split, Split, 21000, Croatia. jmarkic@mefst.hr.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundSerious Bacterial Infection (SBI) in neonates and young infants often exhibit nonspecific symptoms and clinical signs in the early stages of illness, making early diagnosis challenging. Timely recognition and appropriate treatment are essential to prevent adverse outcomes. While several clinical algorithms are widely used for SBI risk stratification, these tools have limitations, particularly low positive predictive value. This study evaluates the diagnostic accuracy of general-purpose large language models (LLMs) in detecting SBI in neonates and infants under 90 days of age admitted to the emergency department. Our objective is to improve diagnostic precision, reduce unnecessary interventions, and enhance patient outcomes. LLM performance was compared against traditional machine learning models, state-of-the-art rule-based methods, and an ensemble of physicians to assess their potential as clinical decision-support tools in scenarios of diagnostic uncertainty.

resultsOn a dataset of 742 patients, LLMs demonstrated diagnostic accuracy comparable to traditional machine learning models and state-of-the-art rule-based methods. The optimized CatBoost (class-weighted) model achieved the best overall performance, with a PPV of 0.70, NPV of 0.90, sensitivity of 0.54, specificity of 0.95, F1-score of 0.60, and MCC of 0.54, outperforming the baseline CatBoost model and achieving results on par with large language models (LLMs) and physicians. When optimally prompted, LLMs performed on par with ensembles of experienced clinicians. Additionally, LLMs exhibited effective medical reasoning and provided credible diagnostic predictions, particularly valuable in cases of clinician uncertainty. The models achieved balanced performance across multiple evaluation metrics, including PPV, NPV, sensitivity, specificity, F1-score, and Matthew’s correlation coefficient (MCC). ChatGPT-4o achieved a sensitivity of 0.65 and specificity of 0.83, with an MCC of 0.41. Claude Sonnet 3.5 reached a sensitivity of 0.60 and specificity of 0.86, MCC 0.42 and Google Gemini 2.0 Flash had lower sensitivity (0.43) but the highest specificity (0.94), with an MCC of 0.43. In comparison, the best-performing individual pediatrician achieved a higher sensitivity (0.74) but lower specificity (0.68), with an MCC of 0.33, while the pediatricians’ majority vote yielded sensitivity of 0.69, specificity of 0.81, and MCC of 0.43 — comparable to the top-performing LLMs.

conclusionsThese Artificial intelligence tools offer a promising direction for SBI risk prediction, achieving performance comparable to that of experienced pediatric specialists, while maintaining simplicity of use/data-preprocessing for potential real-world applications.

Indexed as

Bacterial InfectionsLarge Language ModelsPediatriciansBoosting Machine Learning AlgorithmsHumansInfantInfant, NewbornMachine LearningPredictive Learning ModelsDiagnosticsInfectologyLarge language modelsMachine learningPediatricsPredictionSerious bacterial infection

Identifiers

PMID41239388
PMCPMC12619361

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.