Evidence map›Paper›PMID 42060907›Full record

ArticleJMIR formative research2026

Korean Medical Consultation With Open-Weight Large Language Models: Pilot Comparative Evaluation of Retrieval-Augmented Generation With Metadata Filtering.

Saeyoun Choi, Donghyun Kim, Ji-Hwan Jeon, Minji Kim, Dong Hun Lee, DaeHwan Ahn, Eu Sun Lee, Yoon Ji Kim, Hyun Youk

Abstract readComparative StudyCase Reports
In one paragraph

Article in JMIR formative research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Saeyoun ChoiMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0009-0008-9730-6298
Donghyun KimDepartment of Industrial and Systems Engineering, Dongguk University, Seoul, Republic of Korea.ORCID http://orcid.org/0000-0003-0841-3830
Ji-Hwan JeonMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0009-0006-9165-8175
Minji KimMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0009-0006-8492-4084
Dong Hun LeeMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0009-0002-2234-9089
DaeHwan AhnMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0009-0005-1409-3361
Eu Sun LeeMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0000-0001-6424-9946
Yoon Ji KimMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0009-0008-3767-3889
Hyun YoukMAIN Corp, 1 Gangwon-daehak-gil, Room 1201 (Bodeum-gwan), Chuncheon City, Gangwon Province, 24341, Republic of Korea, 82 10-9840-2120.ORCID http://orcid.org/0000-0002-4631-1504

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: This study develops an open-source large language model-based chatbot tailored for Korean health consultations. The chatbot was implemented using the retrieval-augmented generation (RAG) technique alongside metadata filtering to enhance its performance. Objective: This study aims to analyze and compare the performance of a RAG-based chatbot with other leading language models in the context of Korean health consultations. Methods: A 10.4 GB Korean medical document corpus (487,277 segments) was constructed from official websites of major Korean hospitals, public health sources, and medical textbooks. This study quantitatively compared 5 open-source large language models (Qwen3:4B, Mistral:7B, Llama-3.1:8B, Gpt-Oss:20B, and Gemma3:27B) in 3 configurations: baseline (model only), RAG-only, and RAG with metadata filtering. The RAG system used a specialized Korean embedding model (upskyy/bge-m3-korean) and an Elasticsearch store. Performance was assessed by an emergency medicine specialist using a validation set of 226 questions across 7 common diseases and scoring responses based on accuracy, safety, and helpfulness. Results: The application of RAG alone failed to yield statistically significant performance improvements and, in some cases (Llama 3.1: 8B and Gemma 3: 27B), resulted in decreased scores. However, the combination of RAG with metadata filtering yielded statistically significant (P<.05) performance increases in most models. Notably, the average score for Mistral:7B increased from 3.79, SD 0.08, to 4.10, SD 0.10, and Gpt-Oss:20B increased from 4.43, SD 0.05, to 4.51, SD 0.04, with the latter achieving the highest safety score (4.61, SD 0.03). The Gemma3:27B model, which possessed a high baseline performance (4.42, SD 0.03), was an exception, exhibiting no significant improvement (P=.14) even with filtering. Conclusions: The effectiveness of RAG for specialized domains such as Korean medical consultation is highly dependent on a metadata filtering process that controls the quality of retrieved information; simple information augmentation is insufficient. Furthermore, the benefit of RAG is limited when a model's intrinsic knowledge (eg, Gemma3:27B) already meets or exceeds the quality of the external knowledge base. This finding indicates that performance enhancement strategies must account for both the retrieval mechanism's quality and the model's preexisting capabilities.

Indexed as

Information Storage and RetrievalMetadataReferral and ConsultationHumansLarge Language ModelsPilot ProjectsRepublic of Koreahealth chatbotKorean health carelarge language modelLLMmetadata filteringRAGretrieval-augmented generation

Identifiers

PMID42060907
PMCPMC13132483

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.