Evidence map›Paper›PMID 41539675›Full record

ArticleJMIR medical informatics2026

Prompting and Fine-Tuning Large Language Models for Parkinson Disease Diagnosis: Comparative Evaluation Study Using the PPMI Structured Dataset.

Hyun-Ji Shin, Young Jin Jeong, Sungmin Jun, Do-Young Kang

Abstract readComparative Study
In one paragraph

Article in JMIR medical informatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Hyun-Ji ShinDepartment of Data Sciences Convergence, Graduate School, Dong-A University, Busan, Republic of Korea.ORCID 0000-0001-5239-8979
Young Jin JeongDepartment of Data Sciences Convergence, Graduate School, Dong-A University, Busan, Republic of Korea.ORCID 0000-0001-7611-8185
Sungmin JunInstitute of Convergence Bio-Health, Dong-A University, Busan, Republic of Korea.ORCID 0000-0003-0838-9236
Do-Young KangDepartment of Data Sciences Convergence, Graduate School, Dong-A University, Busan, Republic of Korea.ORCID 0000-0003-1688-0818

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundParkinson disease (PD) presents diagnostic challenges due to its heterogeneous motor and nonmotor manifestations. Traditional machine learning (ML) approaches have been evaluated on structured clinical variables. However, the diagnostic utility of large language models (LLMs) using natural language representations of structured clinical data remains underexplored.

objectiveThis study aimed to evaluate the diagnostic classification performance of multiple LLMs using natural language prompts derived from structured clinical data and to compare their performance with traditional ML baselines.

methodsWe reformatted structured clinical variables from the Parkinson's Progression Markers Initiative (PPMI) dataset into natural language prompts and used them as inputs for several LLMs. Variables with high multicollinearity were removed, and the top 10 features were selected using Shapley additive explanations (SHAP)-based feature ranking. LLM performance was examined across few-shot prompting, dual-output prompting that additionally generated post hoc explanatory text as an exploratory component, and supervised fine-tuning. Logistic regression (LR) and support vector machine (SVM) classifiers served as ML baselines. Model performance was evaluated using F

resultsOn the test set of 122 participants, LR and SVM trained on the 10 SHAP-selected clinical variables each achieved a macro-averaged F

conclusionsThis study provides an exploratory benchmark of how modern LLMs process structured clinical variables in natural language form. While several models achieved diagnostic performance comparable to LR across both the test and temporal validation datasets, their outputs were sensitive to prompting formats, model choice, and class distributions. Occasional variability across repeated output generations reflected the stochastic nature of LLMs, and lightweight models required supervised fine-tuning for stable generalization. These findings highlight the capabilities and limitations of current LLMs in handling tabular clinical information and underscore the need for cautious application and further investigation.

Indexed as

Large Language ModelsParkinson DiseaseHumansMachine LearningPredictive Learning ModelsSupport Vector MachineClaudediagnostic classificationfine-tuningGeminiGPTlarge language modelsLLaMAParkinson diseaseParkinson’s Progression Markers InitiativePPMIprompt engineering

Identifiers

PMID41539675
PMCPMC12856398

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.