ArticleJournal of chemical information and modeling2026
An End-User Audit of Reproducibility, Data Leakage, and Overfitting of the Top-Ranked ADMET Prediction Models in TDC Leaderboards.
Article in Journal of chemical information and modeling, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Public leaderboards such as the Therapeutics Data Commons (TDC) ADMET benchmark are widely treated as a ranking of state-of-the-art models. However, a high leaderboard position is only meaningful if the corresponding model can actually be reproduced and deployed by an independent researcher. In this work, we audit whether the top-ranked TDC ADMET models meet that bar. We assessed the top-ranked models of all 22 TDC ADMET leaderboards from the perspective of an end user with access only to the publicly released artifacts of each model─its publication, code repository, and installation instructions. For every end point, the top three models were screened with a unified protocol including an execution environment reproducibility check, a data-leakage assessment, verification of the hyperparameter-optimization procedure, and a reevaluation against the current leaderboard. Only three models (CaliciBoost, MapLight, and MapLight + GNN) passed all stages and reproduced their reported performance. The remaining models failed because of unavailable code, nonreproducible environments, runtime incompatibilities, or methodological flaws. We traced direct or indirect data leakage in the MiniMol, GradientBoost, and XGBoost models, and used deliberately overfitted variants of our own Mol2Vec-based models to show that tuning on the public test set─whether accidental or intentional─can substantially inflate both metrics and leaderboard rank. These results indicate that current TDC leaderboard positions cannot be read as a direct measure of model quality and practical applicability and emphasize the urgent need for better public ADMET benchmarks based on the hidden test sets, strict data set versioning and model submission with standardized inference environments.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.