Evidence map›Paper›PMID 41998703›Full record

ArticleBMC medical informatics and decision making2026

A leakage-controlled and SHAP driven machine learning framework for paediatric respiratory disease classification using Indian hospital EHR data.

Anusha Prashanth Shetty, Surendra Shetty, Pavan Hegde, Shwetha Shetty, Nagaraja Shetty

Abstract read
In one paragraph

Article in BMC medical informatics and decision making, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Anusha Prashanth ShettyDepartment of Master of Computer Applications, Nitte (Deemed to be University), NMAM Institute of Technology (NMAMIT), Nitte, Udupi, Karnataka, 574110, India.
Surendra ShettyDepartment of Master of Computer Applications, Nitte (Deemed to be University), NMAM Institute of Technology (NMAMIT), Nitte, Udupi, Karnataka, 574110, India. hsshetty@nitte.edu.in.
Pavan HegdeDepartment of Paediatrics, Father Muller Medical College and Hospital, Kankanady, Mangalore, Karnataka, 575002, India.
Shwetha ShettyDepartment of Master of Computer Applications, Nitte (Deemed to be University), NMAM Institute of Technology (NMAMIT), Nitte, Udupi, Karnataka, 574110, India.
Nagaraja ShettyManipal Institute of Technology, Manipal Academy of Higher Education, Manipal, India. nagaraj.shetty@manipal.edu.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundClassification of paediatric respiratory diseases is important for timely diagnosis and decision-making. Real-world clinical data is often compounded by challenges of data leaking, class imbalance, and high dimensional data modalities, which might compromise the reliability, methodological rigour, and interpretability of the model.

methodsThis research work introduces a leakage-controlled and explainable machine learning model for multiclass classification of respiratory diseases in children namely Bronchitis, Bronchopneumonia, Upper Respiratory Tract Infection (URTI), and Wheeze using both structured numerical and unstructured free-text information obtained from practical Electronic Health Records (EHRs). The dataset includes 1,121 paediatric cases sourced from an Indian hospital. A multi-step pre-processing pipeline was applied, including data quality filtering prior to partitioning train-only imputation, and a two-stage data leakage prevention mechanism. Numerical variables were standardized, and clinical free-text narratives were vectorized using Term Frequency–Inverse Document Frequency (TF–IDF) encoding. Class imbalance was handled using the Synthetic Minority Oversampling Technique (SMOTE) applied strictly within cross-validation folds to avoid contamination. SHapley Additive exPlanations (SHAP) were used on a Random Forest baseline to audit for diagnostic leakage and guide feature selection. The top 100 SHAP-ranked features were then used to train Logistic Regression, Random Forest, XGBoost and Stacking Ensemble models. SHAP ranking was performed exclusively on the training partition following train-test splitting, ensuring that feature selection was not influenced by test set information.

resultsSHAP analysis confirmed that predictions were driven by clinically relevant features such as age, breathing difficulty days, Peripheral Oxygen Saturation (SpO₂), respiratory rate (RR), and serum bicarbonate. “Respiratory System” feature was identified as leakage-prone and was excluded from final model training. Random Forest achieved the highest hold-out test accuracy of 0.8578 (95% Confidence Interval (CI): 0.8089–0.9022), with XGBoost achieving the highest micro-averaged Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.9706 and Area Under the Precision-Recall Curve (AUPRC) of 0.9285, computed using a one-vs-rest strategy. The narrow Cross Validation (CV) test gap is consistent with limited overfitting within this single-centre internal validation setting.

conclusionThe proposed pipeline supports transparent and leakage-controlled validation within a single-centre internal setting, contributing to methodological rigour in paediatric EHR-based Machine Learning (ML research). External validation across diverse clinical settings would be required before consideration of clinical deployment. The study also contributes to global health goals by supporting the development of equitable and reliable diagnostic technologies aligned with Sustainable Development Goals, particularly SDG 3 (Good Health and Well-being) and SDG 9 (Industry, Innovation, and Infrastructure).

Indexed as

Electronic Health RecordsMachine LearningRespiratory Tract DiseasesBoosting Machine Learning AlgorithmsChildChild, PreschoolClassification AlgorithmsHumansIndiaRandom ForestClass imbalanceClinical machine learningClinical text analysisData leakage preventionElectronic health recordsExplainable AIFeature selectionPaediatric diagnosticsSDG 3SDG 9SHAP analysisSMOTETF-IDFXGBoost

Identifiers

PMID41998703
PMCPMC13227820

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.