Evidence map›Paper›PMID 40053773›Full record

SynthesisJournal of medical Internet research2025

Diagnostic Performance of Artificial Intelligence-Based Methods for Tuberculosis Detection: Systematic Review.

Seng Hansun, Ahmadreza Argha, Ivan Bakhshayeshi, Arya Wicaksana, Hamid Alinejad-Rokny, Greg J Fox, Siaw-Teng Liaw, Branko G Celler, Guy B Marks

Abstract readSystematic Review
In one paragraph

Synthesis in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
15citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

15 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Review
  4. Review
  5. Article
  6. Article
  7. Review
  8. Article
  9. Review
  10. Article
  11. Review
  12. Article
  13. Review
  14. Article
  15. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Seng HansunSchool of Clinical Medicine, South West Sydney, UNSW Medicine & Health, UNSW Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0001-6619-9751
Ahmadreza ArghaGraduate School of Biomedical Engineering, UNSW Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0002-8276-9774
Ivan BakhshayeshiGraduate School of Biomedical Engineering, UNSW Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0001-5797-586X
Arya WicaksanaInformatics Department, Universitas Multimedia Nusantara, Tangerang, Indonesia.ORCID https://orcid.org/0000-0002-0888-036X
Hamid Alinejad-RoknyTyree Institute of Health Engineering, UNSW Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0002-2189-9153
Greg J FoxNHMRC Clinical Trials Centre, Faculty of Medicine and Health, University of Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0002-4085-1411
Siaw-Teng LiawSchool of Population Health and School of Clinical Medicine, UNSW Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0001-5989-3614
Branko G CellerBiomedical Systems Research Laboratory, School of Electrical Engineering and Telecommunications, UNSW Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0003-3790-2895
Guy B MarksSchool of Clinical Medicine, South West Sydney, UNSW Medicine & Health, UNSW Sydney, Sydney, Australia.ORCID https://orcid.org/0000-0002-8976-8053

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundTuberculosis (TB) remains a significant health concern, contributing to the highest mortality among infectious diseases worldwide. However, none of the various TB diagnostic tools introduced is deemed sufficient on its own for the diagnostic pathway, so various artificial intelligence (AI)-based methods have been developed to address this issue.

objectiveWe aimed to provide a comprehensive evaluation of AI-based algorithms for TB detection across various data modalities.

methodsFollowing PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analysis) 2020 guidelines, we conducted a systematic review to synthesize current knowledge on this topic. Our search across 3 major databases (Scopus, PubMed, Association for Computing Machinery [ACM] Digital Library) yielded 1146 records, of which we included 152 (13.3%) studies in our analysis. QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies version 2) was performed for the risk-of-bias assessment of all included studies.

resultsRadiographic biomarkers (n=129, 84.9%) and deep learning (DL; n=122, 80.3%) approaches were predominantly used, with convolutional neural networks (CNNs) using Visual Geometry Group (VGG)-16 (n=37, 24.3%), ResNet-50 (n=33, 21.7%), and DenseNet-121 (n=19, 12.5%) architectures being the most common DL approach. The majority of studies focused on model development (n=143, 94.1%) and used a single modality approach (n=141, 92.8%). AI methods demonstrated good performance in all studies: mean accuracy=91.93% (SD 8.10%, 95% CI 90.52%-93.33%; median 93.59%, IQR 88.33%-98.32%), mean area under the curve (AUC)=93.48% (SD 7.51%, 95% CI 91.90%-95.06%; median 95.28%, IQR 91%-99%), mean sensitivity=92.77% (SD 7.48%, 95% CI 91.38%-94.15%; median 94.05% IQR 89%-98.87%), and mean specificity=92.39% (SD 9.4%, 95% CI 90.30%-94.49%; median 95.38%, IQR 89.42%-99.19%). AI performance across different biomarker types showed mean accuracies of 92.45% (SD 7.83%), 89.03% (SD 8.49%), and 84.21% (SD 0%); mean AUCs of 94.47% (SD 7.32%), 88.45% (SD 8.33%), and 88.61% (SD 5.9%); mean sensitivities of 93.8% (SD 6.27%), 88.41% (SD 10.24%), and 93% (SD 0%); and mean specificities of 94.2% (SD 6.63%), 85.89% (SD 14.66%), and 95% (SD 0%) for radiographic, molecular/biochemical, and physiological types, respectively. AI performance across various reference standards showed mean accuracies of 91.44% (SD 7.3%), 93.16% (SD 6.44%), and 88.98% (SD 9.77%); mean AUCs of 90.95% (SD 7.58%), 94.89% (SD 5.18%), and 92.61% (SD 6.01%); mean sensitivities of 91.76% (SD 7.02%), 93.73% (SD 6.67%), and 91.34% (SD 7.71%); and mean specificities of 86.56% (SD 12.8%), 93.69% (SD 8.45%), and 92.7% (SD 6.54%) for bacteriological, human reader, and combined reference standards, respectively. The transfer learning (TL) approach showed increasing popularity (n=89, 58.6%). Notably, only 1 (0.7%) study conducted domain-shift analysis for TB detection.

conclusionsFindings from this review underscore the considerable promise of AI-based methods in the realm of TB detection. Future research endeavors should prioritize conducting domain-shift analyses to better simulate real-world scenarios in TB detection.

trial registrationPROSPERO CRD42023453611; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023453611.

Indexed as

Artificial IntelligenceTuberculosisAlgorithmsDeep LearningHumansNeural Networks, ComputerAIartificial intelligencedeep learningdiagnostic performancemachine learningPreferred Reporting Items for Systematic Reviews and Meta-AnalysisPRISMAQUADAS-2Quality Assessment of Diagnostic Accuracy Studies version 2systematic literature reviewtuberculosis detection

Identifiers

PMID40053773
PMCPMC11928776

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.