Evidence map›Paper›PMID 42582480›Full record

ArticleFrontiers in artificial intelligence2026

Can large language models serve as consultants for forensic cause of death analysis? A multidimensional evaluation.

Enhao Fu, Haojie Qin, Zhiling Tian, Hewen Dong, Donghua Zou, Xiaotian Yu, Ningguo Liu

Abstract read
In one paragraph

Article in Frontiers in artificial intelligence, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Enhao FuAcademy of Forensic Science, Shanghai Key Laboratory of Forensic Medicine, Shanghai Forensic Service Platform, Key Laboratory of Forensic Medicine, Ministry of Justice, Shanghai, China.
Haojie QinSchool of Basic Medicine and Forensic Medicine, Institute of Medical Aspects of Specific Environments, Henan University of Science and Technology, Forensic Medicine Identification Center, Luoyang, Henan, China.
Zhiling TianAcademy of Forensic Science, Shanghai Key Laboratory of Forensic Medicine, Shanghai Forensic Service Platform, Key Laboratory of Forensic Medicine, Ministry of Justice, Shanghai, China.
Hewen DongAcademy of Forensic Science, Shanghai Key Laboratory of Forensic Medicine, Shanghai Forensic Service Platform, Key Laboratory of Forensic Medicine, Ministry of Justice, Shanghai, China.
Donghua ZouAcademy of Forensic Science, Shanghai Key Laboratory of Forensic Medicine, Shanghai Forensic Service Platform, Key Laboratory of Forensic Medicine, Ministry of Justice, Shanghai, China.
Xiaotian YuAcademy of Forensic Science, Shanghai Key Laboratory of Forensic Medicine, Shanghai Forensic Service Platform, Key Laboratory of Forensic Medicine, Ministry of Justice, Shanghai, China.
Ningguo LiuAcademy of Forensic Science, Shanghai Key Laboratory of Forensic Medicine, Shanghai Forensic Service Platform, Key Laboratory of Forensic Medicine, Ministry of Justice, Shanghai, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Introduction: Large language models (LLMs) have been proposed as decision support tools in medicine, yet their role in forensic cause of death analysis remains unexplored. Methods: In this study, we used 118 real-world cases spanning diverse categories of death to systematically evaluate the performance of four representative LLMs (GPT-4o, OpenAI o3, Gemini-2.5pro, and DeepSeek-R1) in forensic cause of death analysis. Two senior forensic pathologists independently evaluated each model's decision-making capabilities regarding inference quality and conclusion accuracy. These metrics were assessed using an expert scoring system with a 5-point Likert scale, with original analytical statements and legally valid expert opinions serving as objective gold standards. In a sub-study, we examined the application potential of the locally deployed open-source model DeepSeek-R1:32b. Additionally, a targeted retrospective analysis was conducted to quantify the incidence and typologies of AI hallucinations. Results: DeepSeek-R1 demonstrated a statistically significant advantage in inference quality scores over GPT-4o ( Discussion: LLMs can provide limited auxiliary value in cause of death analysis but should not replace the final judgment of forensic experts. LLMs still require expert oversight to ensure evidence integrity and mitigate risks such as hallucination. Open source LLMs can further mitigate data privacy concerns and provide practical support for cause of death analysis.

Indexed as

cause of deathdecision supportforensic scienceLLMlocal deployment

Identifiers

PMID42582480
PMCPMC13457641

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.