Evidence map›Paper›PMID 42393429›Full record

ArticleLangenbeck's archives of surgery2026

The role of large language models in the writing of surgical reviews: fact or fantasy?

Daphne E DeTemple, Simon Störzer, Sahar Arbabzadah, Anna Riddermann, Felix Gronau, Kai Timrott, Hüseyin Bektas, Moritz Kleine

Abstract read
In one paragraph

Article in Langenbeck's archives of surgery, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Daphne E DeTempleDepartment for General-, Visceral- and Transplant Surgery, Hannover Medical School, Hannover, Germany. ddetemple@ukaachen.de.ORCID http://orcid.org/0000-0002-8100-9394
Simon StörzerDepartment for General-, Visceral- and Transplant Surgery, Hannover Medical School, Hannover, Germany.
Sahar ArbabzadahDepartment for General-, Visceral- and Transplant Surgery, Hannover Medical School, Hannover, Germany.
Anna RiddermannDepartment for General-, Visceral- and Transplant Surgery, Hannover Medical School, Hannover, Germany.
Felix GronauDepartment for General-, Visceral- and Transplant Surgery, Hannover Medical School, Hannover, Germany.ORCID http://orcid.org/0000-0002-1803-2866
Kai TimrottClinic for General- and Visceral Surgery, Agaplesion Ev. Klinikum Schaumburg, Obernkirchen, Germany.ORCID http://orcid.org/0000-0003-1192-3102
Hüseyin BektasClinic for General-, Visceral- and Oncological Surgery, Klinikum Bremen Mitte, Bremen, Germany.ORCID http://orcid.org/0000-0002-5585-8907
Moritz KleineClinic for General-, Visceral- and Oncological Surgery and Coloproctology, Vinzenzkrankenhaus Hannover, Hannover, Germany.ORCID http://orcid.org/0000-0002-4515-3209

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAmidst the current enthusiasm concerning artificial intelligence and its possible application in the composition of different kinds of scientific and non-scientific written documents, we evaluated the usage of artificial intelligence for writing surgical short reviews.

methodsIn order to assess the formal and content quality of AI-generated texts compared to human written texts, ten AI-based text generators (five chatbots and five content creators) and four surgeons in training received the same prompt for a short scientific article on a liver surgery theme. All texts were anonymized and subsequently evaluated by three experienced liver surgeons based on a pre-defined scoring scheme, as well as for quality of references and readability according to readability indices. Furthermore, all texts were tested for plagiarism using PlagScan.

resultsOverall percentage of correct assessment for AI/non-AI generation by experienced surgeons lay at 78.57%. Human written text had a mean word count of 1054 versus 874 in AI-generated text, with a higher mean Flesh Reading Ease Score (FRE, 26.2 ± 5.1 versus 17.7 ± 6.1). References were PubMed-listed in 100% for human versus 46% for AI-generated text, with only one non-human text reaching 100% formally correct citation of references. PlagScan found 6.4%±1.3 mean resemblance to existing texts for human versus 7.6%±4.5 for AI-generated text. DISCUSSION: Overall, AI could already mislead experienced scientific surgeons in 26.7% of cases into believing it to be human. However, formal requirements, especially considering referencing, are still in great need of improvement with only one of AI-generated articles fulfilling our quality requirements.

Indexed as

Artificial IntelligenceLarge Language ModelsWritingGenerative Artificial IntelligenceHumansAI-text generationHepatobiliary surgeryLarge language modelReadabilityScientific writing

Identifiers

PMID42393429
PMCPMC13331823

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.