Evidence map›Paper›PMID 40456109›Full record

ArticleJournal of medical Internet research2025

Enhancing the Accuracy of Human Phenotype Ontology Identification: Comparative Evaluation of Multimodal Large Language Models.

Wei Zhong, Mingyue Sun, Shun Yao, YiFan Liu, Dingchuan Peng, Yan Liu, Kai Yang, HuiMin Gao, HuiHui Yan, WenJing Hao and 2 more

Abstract readComparative Study
In one paragraph

Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
  4. Article
  5. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Wei ZhongDepartment of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0000-0001-9823-9500
Mingyue SunDepartment of Reproductive Medicine, Shijiazhuang People's Hospital, Hebei Province, Shijiazhuang, China.ORCID http://orcid.org/0009-0003-5012-4538
Shun YaoDepartment of Gynecology and Obstetrics, Yijishan Hospital of Wannan Medical College, Anhui province, Wuhu, China.ORCID http://orcid.org/0009-0004-6235-2201
YiFan LiuDepartment of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0009-0008-7339-4756
Dingchuan PengSchool of Medicine, South China University of Technology, Guangdong Province, Guangzhou, China.ORCID http://orcid.org/0009-0000-9588-6809
Yan LiuDepartment of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0000-0003-1698-5783
Kai YangDepartment of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0000-0002-7457-3106
HuiMin GaoDepartment of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0009-0004-8874-6022
HuiHui YanDepartment of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0009-0008-2979-9895
WenJing HaoDepartment of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0009-0006-8537-0036
YouSheng Yan *Department of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0000-0002-0405-1302
ChengHong Yin *Department of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.ORCID http://orcid.org/0000-0002-2503-3285

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Identifying Human Phenotype Ontology (HPO) terms is crucial for diagnosing and managing rare diseases. However, clinicians, especially junior physicians, often face challenges due to the complexity of describing patient phenotypes accurately. Traditional manual search methods using HPO databases are time-consuming and prone to errors. Objective: The aim of the study is to investigate whether the use of multimodal large language models (MLLMs) can improve the accuracy of junior physicians in identifying HPO terms from patient images related to rare diseases. Methods: In total, 20 junior physicians from 10 specialties participated. Each physician evaluated 27 patient images sourced from publicly available literature, with phenotypes relevant to rare diseases listed in the Chinese Rare Disease Catalogue. The study was divided into 2 groups: the manual search group relied on the Chinese Human Phenotype Ontology website, while the MLLM-assisted group used an electronic questionnaire that included HPO terms preidentified by ChatGPT-4o as prompts, followed by a search using the Chinese Human Phenotype Ontology. The primary outcome was the accuracy of HPO identification, defined as the proportion of correctly identified HPO terms compared to a standard set determined by an expert panel. Additionally, the accuracy of outputs from ChatGPT-4o and 2 open-source MLLMs (Llama3.2:11b and Llama3.2:90b) was evaluated using the same criteria, with hallucinations for each model documented separately. Furthermore, participating physicians completed an additional electronic questionnaire regarding their rare disease background to identify factors affecting their ability to accurately describe patient images using standardized HPO terms. Results: A total of 270 descriptions were evaluated per group. The MLLM-assisted group achieved a significantly higher accuracy rate of 67.4% (182/270) compared to 20.4% (55/270) in the manual group (relative risk 3.31, 95% CI 2.58-4.25; P<.001). The MLLM-assisted group demonstrated consistent performance across departments, whereas the manual group exhibited greater variability. Among standalone MLLMs, ChatGPT-4o achieved an accuracy of 48% (13/27), while the open-source models Llama3.2:11b and Llama3.2:90b achieved 15% (4/27) and 18% (5/27), respectively. However, MLLMs exhibited a high hallucination rate, frequently generating HPO terms with incorrect IDs or entirely fabricated content. Specifically, ChatGPT-4o, Llama3.2:11b, and Llama3.2:90b generated incorrect IDs in 57.3% (67/117), 98% (62/63), and 82% (46/56) of cases, respectively, and fabricated terms in 34.2% (40/117), 41% (26/63), and 32% (18/56) of cases, respectively. Additionally, a survey on the rare disease knowledge of junior physicians suggests that participation in rare disease and genetic disease training may enhance the performance of some physicians. Conclusions: The integration of MLLMs into clinical workflows significantly enhances the accuracy of HPO identification by junior physicians, offering promising potential to improve the diagnosis of rare diseases and standardize phenotype descriptions in medical research. However, the notable hallucination rate observed in MLLMs underscores the necessity for further refinement and rigorous validation before widespread adoption in clinical practice.

Indexed as

Biological OntologiesLanguagePhenotypeHumansLarge Language ModelsChatGPThuman phenotype ontologylarge language modelmultimodal large language modelsopen-source LLMsrare diseases

Identifiers

PMID40456109
PMCPMC12148245

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.