ArticleJournal of biomedical informatics2023
A method for comparing multiple imputation techniques: A case study on the U.S. national COVID cohort collaborative.
Article in Journal of biomedical informatics, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
10 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Identify the most appropriate imputation method for handling missing values in clinical structured datasets: a systematic review.BMC medical research methodology · 2024Pooled it
- A novel average treatment effect estimation approach for integrated incomplete datasets, with application in childhood anemia.BMC medical research methodology · 2026Article
- Evaluating the effectiveness and safety of Xiyanping injection for severe pneumonia: a target trial emulation protocol using real-world data.Frontiers in pharmacology · 2026Article
- Increasing the Utility of Real-World Data to Inform Public Health Decision Making Through a US-based Private-Public Partnership: 10 Lessons Learned from a Principled Approach to Rapid Pandemic RWE Generation.Therapeutic innovation & regulatory science · 2025Article
- Conceptual framework as a guide to choose appropriate imputation method for missing values in a clinical structured dataset.BMC medical research methodology · 2025Article
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests.International journal of medical informatics · 2024Article
- Thrombosis risk prediction in lymphoma patients: A multi-institutional, retrospective model development and validation study.American journal of hematology · 2024Article
- Association of post-COVID phenotypic manifestations with new-onset psychiatric disease.Translational psychiatry · 2024Article
- Uncovering COVID-19 transmission tree: identifying traced and untraced infections in an infection network.Frontiers in public health · 2024Article
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests.medRxiv : the preprint server for health sciences · 2023Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
24 authors.
Funding
Abstract
Healthcare datasets obtained from Electronic Health Records have proven to be extremely useful for assessing associations between patients' predictors and outcomes of interest. However, these datasets often suffer from missing values in a high proportion of cases, whose removal may introduce severe bias. Several multiple imputation algorithms have been proposed to attempt to recover the missing information under an assumed missingness mechanism. Each algorithm presents strengths and weaknesses, and there is currently no consensus on which multiple imputation algorithm works best in a given scenario. Furthermore, the selection of each algorithm's parameters and data-related modeling choices are also both crucial and challenging. In this paper we propose a novel framework to numerically evaluate strategies for handling missing data in the context of statistical analysis, with a particular focus on multiple imputation techniques. We demonstrate the feasibility of our approach on a large cohort of type-2 diabetes patients provided by the National COVID Cohort Collaborative (N3C) Enclave, where we explored the influence of various patient characteristics on outcomes related to COVID-19. Our analysis included classic multiple imputation techniques as well as simple complete-case Inverse Probability Weighted models. Extensive experiments show that our approach can effectively highlight the most promising and performant missing-data handling strategy for our case study. Moreover, our methodology allowed a better understanding of the behavior of the different models and of how it changed as we modified their parameters. Our method is general and can be applied to different research fields and on datasets containing heterogeneous types.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.