Evidence map›Paper›PMID 41361882›Full record

ArticleBMC medicine2025

Dramatic increases in redundant publications in the Generative AI era.

Danny Maupin, Tulsi Suchak, Adrian Barnett, Matt Spick

Abstract read
In one paragraph

Article in BMC medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Danny Maupin *School of Health Sciences, Faculty of Health and Medical Sciences, University of Surrey, Guildford, Surrey, GU2 7XH, UK.
Tulsi Suchak *School of Health Sciences, Faculty of Health and Medical Sciences, University of Surrey, Guildford, Surrey, GU2 7XH, UK.
Adrian BarnettSchool of Public Health and Social Work, Queensland University of Technology, Kelvin Grove, Australia.ORCID http://orcid.org/0000-0001-6339-0374
Matt SpickSchool of Health Sciences, Faculty of Health and Medical Sciences, University of Surrey, Guildford, Surrey, GU2 7XH, UK. matt.spick@surrey.ac.uk.ORCID http://orcid.org/0000-0002-9417-6511

Funding

UK Research and Innovation 1095UK Research and Innovation 2604
6 · The paper itself

Abstract

backgroundRedundant publication, the practice of submitting the same or substantially overlapping manuscripts multiple times, distorts the scientific record and wastes resources. Since 2022, publications using large open-science data resources have increased substantially, raising concerns that Generative AI (GenAI) may be facilitating the production of formulaic, redundant manuscripts. In this work, we aim to quantify the extent of redundant publication from a single, large health dataset and to investigate whether GenAI can create submissions that evade standard integrity checks.

methodsWe conducted a systematic search for the years 2021 to 2025 (year to end-July) to identify redundant publications using the US Centers for Disease Control and Prevention National Health and Nutrition Examination Survey (NHANES) dataset. Redundancy was defined as publications analysing the same exposures associated with the same outcomes in the same national population. To test whether GenAI could facilitate creating these papers, we prompted large language models to write three synthetic manuscripts using redundant publications from our dataset as input, instructing them to maximise syntactic differences and evade plagiarism detectors. These three synthetic manuscripts were then tested using a leading plagiarism detection platform to assess their similarity scores.

resultsOur search identified 411 redundant publications across 156 unique exposure-outcome pairings; for example, the association between oxidative balance score and chronic kidney disease using NHANES data was published six times in 1 year. In many instances, redundant articles appeared within the same journals. The three synthetic manuscripts created by GenAI to evade detection yielded overall similarity scores of 26%, 19%, and 14%, with individual similarity contributions below the typical 5% warning thresholds used by plagiarism detectors.

conclusionsThe rapid growth in redundant publications (a 17-fold increase between 2022 and 2024) suggests a systemic failure of editorial checks. These papers distort meta-analyses and scientometric studies, waste scarce peer review resources, and pose a significant threat to the integrity of the scientific record. Current checks for redundant publications and plagiarism are no longer fit for purpose in the GenAI era; greater co-operation between publishers and modified guidelines will be needed to address new innovations in paper mill production.

Indexed as

Artificial IntelligencePublicationsPublishingHumansPlagiarismUnited StatesCOPEGenerative AIIntegrityNHANESPaper millsPlagiarismRedundant publication

Identifiers

PMID41361882
PMCPMC12801743

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.