ArticleJMIR formative research2026
Structured Large Language Model Workflows for Motivational Interviewing in Health Behavior Change: Proof-of-Concept Study.
Article in JMIR formative research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
7 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Motivational interviewing (MI) is an effective approach for supporting health behaviorchange, but face-to-face delivery is resource-intensive and difficult to scale. Rule-based conversational agents (CAs) can improve access; however, their scripted interactions and limited language flexibility constrain MI delivery. While large language models (LLMs) are increasingly being used for MI coaching, their conversational fidelity and quality compared with human coaches and rule-based CAs remain understudied. Objective: This study aimed to describe the development of an LLM-based CA, Artificially Intelligent Motivational Interviewing (Aimi), orchestrated with structured workflows, and to evaluate its feasibility, conversational fidelity, and user perceptions during MI coaching interactions. Methods: We developed Aimi using structured LLM workflows designed to enhance MI fidelity. We conducted a within-participants study, where 18 adults interacted with (1) Aimi, (2) a novice MI-trained human coach, and (3) a rule-based CA during live text-based role-play coaching sessions. Transcripts were independently evaluated by an MI expert using the Motivational Interviewing Skill Code, Version 2.0 (MISC-2), to assess MI competency and fidelity. Participants completed a user experience questionnaire to provide general feedback and to assess session alliance, dialogue relevance, empathy, engagement, linguistic quality, and perceived motivation to change. Feedback from users was thematically summarized and categorized under strengths and weaknesses for each approach. Results: Aimi achieved fidelity scores comparable to those of the novice human coach and higher than those of the rule-based CA on summary metrics, including higher reflection-to-question ratios (median 0.84, IQR 0.62-0.92 vs 0.62, IQR 0.42-0.74 vs 0.25, IQR 0.17-0.38), more complex reflections (median 66.67%, IQR 46.97%-76.92% vs 50%, IQR 34.38%-61.88% vs 0.00%, IQR 0%-50%), and greater elicitation of client change talk (median 90.83%, IQR 85.89%-100% vs 73.21%, IQR 63.10%-83.19% vs 66.67%, IQR 57.86%-81.94%). User experience ratings showed no significant differences across conditions. User feedback revealed distinct strengths and limitations across the coaching interactions. Participants described Aimi's interactions as personalized, fluid, and adaptive, though sometimes overly reflective and lengthy. The novice human coach was viewed as empathetic and supportive but slow to respond, whereas the rule-based coach was viewed as efficient and structured yet limited in depth and personalization. Conclusions: This study demonstrates the technical feasibility of structured LLM-workflows for MI coaching and their capacity to maintain conversational fidelity comparable to that of a novice MI-trained human coach. Given the role-play paradigm, single-rater coding, and small convenience sample, these comparative findings should be interpreted as exploratory. Our findings serve as a foundational baseline for the development of scalable behavior change interventions in clinical settings.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.