Study
A GPT-4 chatbot held one alcohol counseling session with 45 young adults. The study did not measure whether they drank less
Source date · Reviewed
A Stanford GPT-4 chatbot gave one alcohol counseling session to 45 young adults. Reviewers found no unsafe replies. Drinking afterward was not measured.
Study card
- Who was studied
- Adults aged 18 to 25 living in the U.S. who reported drinking 10 or more standard drinks a week, recruited through the online research platform Prolific
- How many
- 45 (8 in Phase I, 37 in Phase II)
- Design
- Single-arm pilot in two phases. The researchers rewrote the chatbot's instructions between phases. No comparison group.
- What was tested
- One text-based counseling session with MICA, a chatbot built on GPT-4 and run on Stanford's secure servers, prompted to use motivational interviewing
- Compared with
- None. Phase II was compared with Phase I after the prompts were changed.
- Main outcome
- Safety of the chatbot's replies, participants' ratings of how closely it followed motivational interviewing, and usability
- Result
- No unsafe or inappropriate replies were found. Ratings of the chatbot's relational skill rose from 67.2% in Phase I to 82.6% in Phase II (p = 0.03). Usability scored 85.4 and 80.9 out of 100 (p = 0.45).
- Limitations
- Small sample, no comparison group, one session, and no follow-up on drinking. Participants came from one online platform, and most were male and White and had at least some college.
- Funding and conflicts
- Supported by Stanford University School of Medicine and the National Institute on Alcohol Abuse and Alcoholism (grant 1R01AA030986), which the first author declared as financial support. The paper states the NIAAA had no role in the study. The other authors declared no competing interests.
- What this does NOT tell us
- Whether talking to the chatbot changes how much anyone drinks, or how it would respond to a person in crisis, since the paper does not report testing that.
The short version
Researchers at Stanford built a chatbot on GPT-4 to deliver brief alcohol counseling in the style of motivational interviewing, and 45 young adults who drink heavily each tried it once. The authors, several of them clinicians, read every transcript and found no unsafe replies, and participants rated it easy to use. The pilot tested safety and acceptability, not effect. It did not measure drinking afterward.
What they did
The chatbot, MICA, ran on a secure version of GPT-4 hosted on Stanford's servers, and the authors say the setup met HIPAA standards and passed an institutional cybersecurity review. It was instructed to counsel in the style of motivational interviewing, a conversational approach that helps a person talk through their own reasons to change.
Participants were 18 to 25, lived in the U.S., and reported drinking at least 10 standard drinks a week. They were recruited through Prolific, an online platform that pays people to take part in research. Each had one text session, with no time limit.
Testing ran in two phases. Eight people used the first version. The researchers then rewrote the chatbot's instructions, and 37 people used the second. All seven authors, who include an adolescent psychiatrist, a clinical psychologist trained in motivational interviewing, an addiction psychologist and an emergency physician, read the transcripts for replies that were medically inappropriate, encouraged self-harm, played down serious disclosures, or were discriminatory or off in tone.
What it found
No unsafe or inappropriate replies were found in either phase.
Participants rated how closely the chatbot followed motivational interviewing on a standard two-part questionnaire. The relational part, which covers things like collaboration and empathy, rose from 67.2% in Phase I to 82.6% in Phase II (p = 0.03). The technical part rose from 69.6% to 81.3%, a difference that was not statistically significant (p = 0.13). The authors set 80% as their benchmark, and Phase II cleared it on both.
Usability scored 85.4 out of 100 in Phase I and 80.9 in Phase II (p = 0.45), above the authors' benchmark of 68.
The authors also counted "change talk," statements in which a participant voices a desire, reason, need or plan to change their drinking, against "sustain talk," statements in favor of drinking as before. On average, change talk made up 65.2% of these statements per session in Phase I and 75.8% in Phase II, a difference that was not statistically significant (p = 0.10). One author coded all of them by hand.
Sessions averaged 6.3 minutes in Phase I and 10.6 minutes in Phase II. Some participants said the chatbot felt supportive. Others found it formulaic, for example opening each reply by restating what they had just said.
What it does not show
It does not show that the chatbot helps anyone drink less. There was no comparison group and no follow-up on drinking, and the authors say a randomized trial with longer follow-up is needed. Change talk is a stand-in measure, not an outcome.
The safety finding is narrower than it sounds. Seven reviewers found no bad replies in 45 conversations with people recruited for a study. The paper does not report testing the chatbot with someone in crisis, someone at risk of alcohol withdrawal, or someone asking a medical question it should not answer.
The sample was small and not representative. Across both phases, 71.1% were male, 66.7% were White and 82.2% had at least some college. The authors note that people with other mental health conditions or limited digital skills may need more support than a chatbot can give.
The study was supported by Stanford and by a federal grant from the National Institute on Alcohol Abuse and Alcoholism. No company funded it, and the authors reported no commercial conflicts. The version read for this entry was the authors' accepted manuscript on PubMed Central.
Why it matters
Two earlier entries on this site tested general-purpose chatbots with lists of questions (dangerous advice on recovery questions, ChatGPT-4 on alcohol). This one built a narrow tool for a single task, kept it on secure servers, and had clinicians read every conversation. That is a careful way to start. It is still only a start.
For someone in recovery or cutting back, the takeaway is modest: in a small test, a purpose-built chatbot held a counseling-style conversation without saying anything harmful. Whether it changes drinking is unknown. The authors describe tools like this as one piece of a wider system of care, not a standalone answer.
For clinicians and programs weighing a chatbot, the useful questions are the ones this pilot could not yet answer: what was it compared against, how did it handle a person in danger, and did anyone's drinking change?
If you have been drinking heavily for a long time, stopping suddenly can bring on withdrawal that is painful and can be life-threatening. The National Institute on Alcohol Abuse and Alcoholism advises getting medical help to plan a safe way to stop. Seizures, severe confusion, fever, or seeing or feeling things that are not there after you stop can be an emergency: call 911 or go to an emergency room.
Sources
- Suffoletto B, Clark DB, Lee C, Mason M, Schultz J, Szeto I, Walker D. Development and preliminary testing of a secure large language model-based chatbot for brief alcohol counseling in young adults. Drug and Alcohol Dependence. 2025 Jul 1;272:112697. Epub 2025 Apr 28. doi:10.1016/j.drugalcdep.2025.112697; PMID 40334327; PMCID PMC12207782 https://pubmed.ncbi.nlm.nih.gov/40334327/
- National Institute on Alcohol Abuse and Alcoholism. To Cut Down or to Quit. Rethinking Drinking. Accessed September 29, 2026. https://rethinkingdrinking.niaaa.nih.gov/thinking-about-change/cut-down-or-quit
- MedlinePlus Medical Encyclopedia. Alcohol withdrawal. U.S. National Library of Medicine. Review date January 1, 2025. https://medlineplus.gov/ency/article/000764.htm
Published by ZSKFL Management.