Evidence map›Paper›PMID 40691682›Full record

ArticleScientific reports2025

Multiclass classification of thalassemia types using complete blood count and HPLC data with machine learning.

Muhammad Umar Nasir, Muhammad Zubair, Muhammad Tahir Naseem, Tariq Shahzad, Ahmed Saeed, Khan Muhammad Adnan, Amir H Gandomi

Abstract read
In one paragraph

Article in Scientific reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Muhammad Umar Nasir *Faculty of Computing, Riphah International University, Islamabad, Pakistan.
Muhammad Zubair *Faculty of Computing, Riphah International University, Islamabad, Pakistan.
Muhammad Tahir Naseem *Department of Electronic Engineering, Yeungnam University, Gyeongsan, 38541, Republic of Korea.
Tariq ShahzadDepartment of Computer Engineering, COMSATS University Islamabad, Sahiwal Campus, Sahiwal, 57000, Pakistan.
Ahmed SaeedDivision of Computing Science and Mathematics, University of Stirling, Stirling, FK9 4LA, Scotland, UK.
Khan Muhammad AdnanDepartment of Software, Faculty of AI and Software, Gachon University, Seongnam-si, 13120, Republic of Korea. adnan@gachon.ac.kr.
Amir H GandomiFaculty of Engineering and IT, University of Technology Sydney, Sydney, NSW, 2007, Australia. gandomi@uts.edu.au.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Mild to severe anemia is caused by thalassemia, a common genetic disorder affecting over 100 countries worldwide, that results from the abnormality of one or several of the four globin genes. This leads to chronic hemolytic anemia and disrupted synthesis of hemoglobin chains, iron overload, and poor erythropoiesis. Although the diagnosis of thalassemia has improved globally along with the treatment and transfusion support, it is still a major problem in diagnosing in high-prevalence areas like Pakistan. This work aims to assess the performance of numerous combinations of machine learning methods to detect alpha and beta-thalassemia in their minor and major types. These results are obtained from CBC and HPLC analysis. The analyzed models are K-nearest Neighbor (KNN), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost). The study aims to examine the effectiveness of the developed models in discriminating thalassemia variants, especially in the light of Pakistani patients' data. The study found that XGBoost achieved the highest performance on both the CBC and HPLC datasets, with training accuracies of roughly 99.5% for CBC and 99.3% for HPLC. The test accuracy across both datasets was consistently high and thus the best model for detecting thalassemia in this research study. The imported SVM model, slightly less accurate than XGBoost, still has strong performance, particularly on the HPLC data where the cumulative testing accuracy of the model stood at 99.4%. As can be seen from the results, XGBoost specifically shows a very high accuracy of above 99% in the detection of thalassemia types using CBC and HPLC data for Pakistani patients. To the author's knowledge, this research is the first to predict alpha and beta-thalassemia in its major and minor forms using these diagnostic reports. These models indicate that they can offer significant support in detecting thalassemia in resource-constrained settings such as Pakistan. If deep learning is incorporated, even greater accuracy could be achieved.

Indexed as

alpha-Thalassemiabeta-ThalassemiaMachine LearningThalassemiaBlood Cell CountChromatography, High Pressure LiquidHumansPakistanSupport Vector MachineAlpha majorAlpha minorAlpha thalassemiaBeta majorBeta minorBeta thalassemiaComplete blood count (CBC)High-performance liquid chromatography (HPLC)KNNSVMXGBoost

Identifiers

PMID40691682
PMCPMC12279977

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.