Evidence map›Paper›PMID 42392971›Full record

ArticleJournal of chemical information and modeling2026

An End-User Audit of Reproducibility, Data Leakage, and Overfitting of the Top-Ranked ADMET Prediction Models in TDC Leaderboards.

Ihor Koleiev, Roman Stratiichuk, Nazar Shevchuk, Mykola Melnychenko, Alex Nyporko, Daniil Todoryshyn, Vladyslav Husak, Sergii Starosyla, Semen Yesylevskyy, Alan Nafiiev

Abstract read
In one paragraph

Article in Journal of chemical information and modeling, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Ihor KoleievReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.
Roman StratiichukReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.
Nazar ShevchukReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.
Mykola MelnychenkoReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.
Alex NyporkoTaras Shevchenko Kyiv National University, 64 Volodymyrska Str., Kyiv01601, Ukraine.
Daniil TodoryshynReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.
Vladyslav HusakReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.
Sergii StarosylaReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.ORCID 0000-0002-5103-0635
Semen YesylevskyyReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.ORCID 0000-0002-6748-8931
Alan NafiievReceptor.AI Inc., 20-22 Wenlock Road, LondonN1 7GU, U.K.ORCID 0009-0004-8604-377X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Public leaderboards such as the Therapeutics Data Commons (TDC) ADMET benchmark are widely treated as a ranking of state-of-the-art models. However, a high leaderboard position is only meaningful if the corresponding model can actually be reproduced and deployed by an independent researcher. In this work, we audit whether the top-ranked TDC ADMET models meet that bar. We assessed the top-ranked models of all 22 TDC ADMET leaderboards from the perspective of an end user with access only to the publicly released artifacts of each model─its publication, code repository, and installation instructions. For every end point, the top three models were screened with a unified protocol including an execution environment reproducibility check, a data-leakage assessment, verification of the hyperparameter-optimization procedure, and a reevaluation against the current leaderboard. Only three models (CaliciBoost, MapLight, and MapLight + GNN) passed all stages and reproduced their reported performance. The remaining models failed because of unavailable code, nonreproducible environments, runtime incompatibilities, or methodological flaws. We traced direct or indirect data leakage in the MiniMol, GradientBoost, and XGBoost models, and used deliberately overfitted variants of our own Mol2Vec-based models to show that tuning on the public test set─whether accidental or intentional─can substantially inflate both metrics and leaderboard rank. These results indicate that current TDC leaderboard positions cannot be read as a direct measure of model quality and practical applicability and emphasize the urgent need for better public ADMET benchmarks based on the hidden test sets, strict data set versioning and model submission with standardized inference environments.

Indexed as

Pharmaceutical PreparationsPrediction AlgorithmsReproducibility of ResultsPharmaceutical Preparations

Identifiers

PMID42392971
PMCPMC13417885

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.