Evidence mapPaperPMID 41559613Full record

ArticleBMC medical research methodology2026

Human-AI collaboration enhances the performance of large language models in risk of bias assessment.

Yingyin Li, Fengchun Yang, Meng Wu, Jiao Li

Abstract read
In one paragraph

Article in BMC medical research methodology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Yingyin LiInstitute of Medical Information, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.
Fengchun YangDepartment of Infectious Diseases, Tongji Hospital, Tongji Medical College and State Key Laboratory for Diagnosis and Treatment of Severe Zoonostic Infectious Disease, Huazhong University of Science and Technology, Wuhan, China.
Meng WuInstitute of Medical Information, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.
Jiao LiInstitute of Medical Information, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China. jiao.li@pumc.edu.cn.

Funding

the Chinese Academy of Medical Sciences (CAMS) Innovation Fund for Medical Sciences 2021-I2M-1-056the National Science and Technology Major Project 2023ZD0509701
6 · The paper itself

Abstract

backgroundRisk of bias (RoB) assessment is essential in systematic reviews and clinical guideline development. Current manual processes are complex, inefficient, and inconsistent. Large language models (LLMs) have potential to assist RoB assessment, but their standalone performance is limited. Efficient integration of LLMs with human judgment remains a challenge.

methodsA high-quality dataset of medical literature was developed, encompassing seven bias domains as defined by the Cochrane RoB 1.0 tool. Using structured prompt engineering, two LLMs, DeepSeek V3 and Qwen-plus, were compared on three-class bias risk classification tasks. Three Human-AI collaboration modes were developed: (M1) Evidence Extraction Mode—human judgment based solely on LLM-extracted evidence; (M2) Reasoning Support Mode—combining LLM extracted evidence with reasoning explanations; (M3) Disagreement Trigger Mode—human intervention triggered by model disagreement. Accuracy and human intervention rates were used to evaluate performance and efficiency.

resultsWe randomly sampled 300 instances per domain from the dataset and evaluated both LLMs using the same structured prompting approach. DeepSeek V3 outperformed Qwen-plus in accuracy in four of seven bias domains, demonstrating superior overall judgment. Model accuracy ranged from 0.403 to 0.777 across domains. In Human-AI collaboration, 60 samples per domain were evaluated with human involvement. M1 showed relatively limited performance in both accuracy and intervention rate; M2 achieved the highest accuracy in most tasks; M3 markedly reduced intervention rates. Accuracy under M2 and M3 ranged from 0.633 to 0.900. Incorporating LLM reasoning improved consistency in human judgments. Disagreement Trigger Mode exhibited high cost-effectiveness in structured and moderate-reasoning tasks, enhancing assessment efficiency, while Reasoning Support Mode was more stable and practical for open-ended and highly subjective tasks.

conclusionsLLMs (like DeepSeek V3 and Qwen-plus) cannot yet replace human RoB assessment but well-designed Human-AI collaboration can improve accuracy and reduce manual workload.

Indexed as

Artificial IntelligenceLarge Language ModelsBiasHumansRisk AssessmentHuman-AI collaborationLarge language modelsRisk of biasSystematic reviews

Identifiers

PMID41559613
PMCPMC12903640

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.