ArticleBMC biology2021
Current cancer driver variant predictors learn to recognize driver genes instead of functional variants.
Article in BMC biology, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 13 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
13 citing papers in PubMed.
- PICDGI: A framework for predicting cancer driver genes through dynamic gene-gene interaction modeling of single-cell data.PLoS computational biology · 2026Article
- Classification of driver and passenger mutations in different cancer types using deep neural networks.Bioinformatics advances · 2026Article
- Article
- Explainable deep learning for stratified medicine in inflammatory bowel disease.Genome biology · 2025Article
- The specification game: rethinking the evaluation of drug response prediction for precision oncology.Journal of cheminformatics · 2025Article
- Combining evolution and protein language models for an interpretable cancer driver mutation prediction with D2Deep.Briefings in bioinformatics · 2024Article
- AI-derived comparative assessment of the performance of pathogenicity prediction tools on missense variants of breast cancer genes.Human genomics · 2024Article
- Discovering predisposing genes for hereditary breast cancer using deep learning.Briefings in bioinformatics · 2024Article
- Metabolic Interplay in the Tumor Microenvironment: Implications for Immune Function and Anticancer Response.Current issues in molecular biology · 2023Review
- VIPpred: a novel model for predicting variant impact on phosphorylation events driving carcinogenesis.Briefings in bioinformatics · 2023Article
- Cancer driver mutations: predictions and reality.Trends in molecular medicine · 2023Review
- Predicting functional consequences of mutations using molecular interaction network features.Human genetics · 2022Article
- HPMPdb: A machine learning-ready database of protein molecular phenotypes associated to human missense variants.Current research in structural biology · 2022Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundIdentifying variants that drive tumor progression (driver variants) and distinguishing these from variants that are a byproduct of the uncontrolled cell growth in cancer (passenger variants) is a crucial step for understanding tumorigenesis and precision oncology. Various bioinformatics methods have attempted to solve this complex task.
resultsIn this study, we investigate the assumptions on which these methods are based, showing that the different definitions of driver and passenger variants influence the difficulty of the prediction task. More importantly, we prove that the data sets have a construction bias which prevents the machine learning (ML) methods to actually learn variant-level functional effects, despite their excellent performance. This effect results from the fact that in these data sets, the driver variants map to a few driver genes, while the passenger variants spread across thousands of genes, and thus just learning to recognize driver genes provides almost perfect predictions.
conclusionsTo mitigate this issue, we propose a novel data set that minimizes this bias by ensuring that all genes covered by the data contain both driver and passenger variants. As a result, we show that the tested predictors experience a significant drop in performance, which should not be considered as poorer modeling, but rather as correcting unwarranted optimism. Finally, we propose a weighting procedure to completely eliminate the gene effects on such predictions, thus precisely evaluating the ability of predictors to model the functional effects of single variants, and we show that indeed this task is still open.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.