Evidence map›Paper›PMID 39144030›Full record

ArticleQuantitative imaging in medicine and surgery2024

Independent evaluation of the accuracy of 5 artificial intelligence software for detecting lung nodules on chest X-rays.

Kirill Arzamasov, Yuriy Vasilev, Maria Zelenova, Lev Pestrenin, Yulia Busygina, Tatiana Bobrovskaya, Sergey Chetverikov, David Shikhmuradov, Andrey Pankratov, Yury Kirpichev and 3 more

Abstract read
In one paragraph

Article in Quantitative imaging in medicine and surgery, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Review
  5. Artificial intelligence in automated detection of lung nodules: a narrative review.International journal of physiology, pathophysiology and pharmacology · 2025
    Review
  6. Article
  7. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Kirill ArzamasovState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0001-7786-0349
Yuriy VasilevState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-5283-5961
Maria ZelenovaState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0001-7458-5396
Lev PestreninState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-1786-4329
Yulia BusyginaState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-4775-258X
Tatiana BobrovskayaState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-2746-7554
Sergey ChetverikovState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-3097-8881
David ShikhmuradovState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0003-1597-5786
Andrey PankratovState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.
Yury KirpichevState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-9583-5187
Valentin SinitsynState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-5649-2193
Irina SonFederal State Budgetary Educational Institution of Further Professional Education "Russian Medical Academy of Continuous Professional Education" of the Ministry of Healthcare of the Russian Federation, Moscow, Russian Federation.ORCID https://orcid.org/0000-0001-9309-2853
Olga OmelyanskayaState Budget-Funded Health Care Institution of the City of Moscow "Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department", Moscow, Russian Federation.ORCID https://orcid.org/0000-0002-0245-4431

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: The integration of artificial intelligence (AI) into medicine is growing, with some experts predicting its standalone use soon. However, skepticism remains due to limited positive outcomes from independent validations. This research evaluates AI software's effectiveness in analyzing chest X-rays (CXR) to identify lung nodules, a possible lung cancer indicator. Methods: This retrospective study analyzed 7,670,212 record pairs from radiological exams conducted between 2020 and 2022 during the Moscow Computer Vision Experiment, focusing on CXR and computed tomography (CT) scans. All images were acquired during clinical routine. The final dataset comprised 100 CXR images (50 with lung nodules, 50 without), selected consecutively and based on inclusion and exclusion criteria, to evaluate the performance of all five AI-based solutions, participating in the Moscow Computer Vision Experiment and analyzing CXR. The evaluation was performed in 3 stages. In the first stage, the probability of a nodule in the lung obtained from AI services was compared with the Ground Truth (1-there is a nodule, 0-there is no nodule). In the second stage, 3 radiologists evaluated the segmentation of nodules performed by the AI services (1-nodule correctly segmented, 0-nodule incorrectly segmented or not segmented at all). In the third stage, the same radiologists additionally evaluated the classification of the nodules (1-nodule correctly segmented and classified, 0-all other cases). The results obtained in stages 2 and 3 were compared with Ground Truth, which was common to all three stages. For each stage, diagnostic accuracy metrics were calculated for each AI service. Results: Three software solutions (Celsus, Lunit INSIGHT CXR, and qXR) demonstrated diagnostic metrics that matched or surpassed the vendor specifications, and achieved the highest area under the receiver operating characteristic curve (AUC) of 0.956 [95% confidence interval (CI): 0.918 to 0.994]. However, when evaluated by three radiologists for accurate nodule segmentation and classification, all solutions performed below the vendor-declared metrics, with the highest AUC reaching 0.812 (95% CI: 0.744 to 0.879). Meanwhile, all AI services demonstrated 100% specificity at stages 2 and 3 of the study. Conclusions: To ensure the reliability and applicability of AI-based software, it is crucial to validate performance metrics using high-quality datasets and engage radiologists in the evaluation process. Developers are recommended to improve the accuracy of the underlying models before allowing the standalone use of the software for lung nodule detection. The dataset created during the study may be accessed at https://mosmed.ai/datasets/mosmeddatargogksnalichiemiotsutstviemlegochnihuzlovtipvii/.

Indexed as

artificial intelligence (AI)Chest X-ray (CXR)computer visionlung nodulesradiology

Identifiers

PMID39144030
PMCPMC11320553

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.