Evidence map›Paper›PMID 39198744›Full record

SynthesisBMC medical research methodology2024

Identify the most appropriate imputation method for handling missing values in clinical structured datasets: a systematic review.

Marziyeh Afkanpour, Elham Hosseinzadeh, Hamed Tabesh

Abstract readSystematic Review
In one paragraph

Synthesis in BMC medical research methodology, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 35 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
35citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

35 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Observational
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Review
  14. Review
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Marziyeh AfkanpourDepartment of Medical Informatics, Faculty of Medicine, Mashhad University of Medical Sciences, Mashhad, Iran.ORCID 0000-0001-7365-0782
Elham HosseinzadehDepartment of Medical Informatics, Faculty of Medicine, Mashhad University of Medical Sciences, Mashhad, Iran.
Hamed TabeshDepartment of Medical Informatics, Faculty of Medicine, Mashhad University of Medical Sciences, Mashhad, Iran. tabesh79@gmail.com.ORCID 0000-0003-3081-0488

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

BACKGROUND AND

objectivesComprehending the research dataset is crucial for obtaining reliable and valid outcomes. Health analysts must have a deep comprehension of the data being analyzed. This comprehension allows them to suggest practical solutions for handling missing data, in a clinical data source. Accurate handling of missing values is critical for producing precise estimates and making informed decisions, especially in crucial areas like clinical research. With data's increasing diversity and complexity, numerous scholars have developed a range of imputation techniques. To address this, we conducted a systematic review to introduce various imputation techniques based on tabular dataset characteristics, including the mechanism, pattern, and ratio of missingness, to identify the most appropriate imputation methods in the healthcare field. MATERIALS AND

methodsWe searched four information databases namely PubMed, Web of Science, Scopus, and IEEE Xplore, for articles published up to September 20, 2023, that discussed imputation methods for addressing missing values in a clinically structured dataset. Our investigation of selected articles focused on four key aspects: the mechanism, pattern, ratio of missingness, and various imputation strategies. By synthesizing insights from these perspectives, we constructed an evidence map to recommend suitable imputation methods for handling missing values in a tabular dataset.

resultsOut of 2955 articles, 58 were included in the analysis. The findings from the development of the evidence map, based on the structure of the missing values and the types of imputation methods used in the extracted items from these studies, revealed that 45% of the studies employed conventional statistical methods, 31% utilized machine learning and deep learning methods, and 24% applied hybrid imputation techniques for handling missing values.

conclusionConsidering the structure and characteristics of missing values in a clinical dataset is essential for choosing the most appropriate data imputation technique, especially within conventional statistical methods. Accurately estimating missing values to reflect reality enhances the likelihood of obtaining high-quality and reusable data, contributing significantly to precise medical decision-making processes. Performing this review study creates a guideline for choosing the most appropriate imputation methods in data preprocessing stages to perform analytical processes on structured clinical datasets.

Indexed as

Biomedical ResearchData Interpretation, StatisticalDatasets as TopicHumansClinical datasetImputation methodsMechanism of missingnessMissing ratioMissing valuesPattern of missingnessSimulation study

Identifiers

PMID39198744
PMCPMC11351057

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.