ArticleBMC oral health2026
Prompt sensitivity of large language models in orthodontic patient counseling: a scenario-based experimental study.
Article in BMC oral health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundLarge language models (LLMs) are increasingly used as accessible sources of health information, including orthodontic patient counseling. While previous studies have evaluated the accuracy and reliability of AI-generated responses, the effect of prompt formulation on the safety and quality of orthodontic advice remains unclear. Understanding prompt sensitivity is essential for assessing the real-world reliability of conversational AI systems, as patient queries are typically expressed in diverse linguistic forms.
methodsThis in silico experimental study evaluated prompt sensitivity using 24 standardized orthodontic clinical scenarios. Each scenario was queried using four prompt formulations (brief layperson, detailed patient, professional clinical, and anxiety-driven), resulting in 96 prompts. These were submitted to four LLMs (ChatGPT, Gemini, Copilot, and Claude), generating 384 responses. Responses were independently evaluated by two orthodontic experts using a predefined expert scoring rubric across accuracy, safety, completeness, and clarity. Consensus scores were analyzed using the Friedman test for prompt effects and the Kruskal-Wallis test for model comparisons. Unsafe response rates and prompt robustness indices were also calculated.
resultsSafety scores did not differ significantly across prompt formulations (χ²(3) = 3.40, p = 0.334) or between models (H = 0.17, p = 0.982). A total of 5 of 384 responses (1.3%) were classified as unsafe. Prompt robustness analysis demonstrated low variability (mean prompt robustness index = 0.15). Response length differed significantly across prompt types (p = 0.0029), whereas response time did not (p = 0.998). A significant difference in clarity scores was observed across models (p = 0.029), with post hoc analysis indicating higher clarity scores for ChatGPT than Claude (adjusted p = 0.020).
conclusionsLLMs demonstrated consistent and clinically safe performance in orthodontic patient counseling, with minimal sensitivity to prompt formulation. While prompt wording influenced response length, it did not affect clinical reliability. Differences between models were primarily related to clarity rather than content. LLMs may provide stable informational support across diverse patient queries; however, their outputs should remain adjunctive to professional care.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.