Study · Chatbots and AI companions in recovery
Chatbots answered real recovery questions well overall, and some answers were dangerous
Source date · Reviewed
Clinicians rated ChatGPT and LLaMA-2 answers to recovery-forum questions as high quality, but some gave wrong helplines or endorsed home detox.
Study card
- Who was studied
- Real questions about alcohol, marijuana, and opioids from three Reddit recovery forums, posted between January and March 2023
- How many
- 75 questions, 150 answers, rated by 7 clinicians
- Design
- Evaluation of chatbot answers. Clinicians rated them without being told the answers came from AI.
- What was tested
- Answers from ChatGPT (running GPT-4) and LLaMA-2
- Main outcome
- Clinicians' ratings of the answers
- Result
- Clinicians rated the answers high quality overall. The researchers also found dangerous misinformation mixed in, including ignoring signs of suicidal thinking, incorrect emergency helpline information, and endorsing detox at home. Advice shifted with how a question was worded.
- Limitations
- The fact-check was a search for examples, not a systematic count. Two chatbots at one point in time. The researchers had to work around the chatbots' safety settings.
- Funding and conflicts
- Supported in part by the NIH Intramural Research Program at the National Institute on Drug Abuse. The authors declared no competing interests.
- What this does NOT tell us
- How often the chatbots got things dangerously wrong, or what happened to anyone who received these answers.
The short version
Researchers took questions people had posted in online recovery forums and put them to two AI chatbots. Clinicians who treat substance use disorders judged the answers high quality overall. Mixed in were some that could hurt someone, including wrong emergency helplines and advice that endorsed detoxing at home.
What they did
The team collected real questions about substance use and recovery from online recovery forums. They asked two generative AI systems, OpenAI's ChatGPT (running GPT-4) and Meta's LLaMA-2, to answer them. Clinicians who work with substance use disorders then rated the responses.
The full paper reports 75 questions, 25 each about alcohol, marijuana, and opioids, taken from three Reddit recovery forums and posted between January and March 2023. Seven clinicians at a substance use treatment research facility rated the 150 answers, three ratings per answer, without being told the answers came from AI. The study appeared in the peer-reviewed journal Psychiatry Research. It was supported in part by the NIH Intramural Research Program at the National Institute on Drug Abuse, and the authors declared no competing interests.
What it found
Overall, clinicians rated the AI answers as high quality.
The researchers also found dangerous misinformation. Examples named in the study include:
- ignoring signs of suicidal thinking in a question
- giving incorrect emergency helpline information
- endorsing detox at home
The advice also shifted depending on how a question was worded. Two people asking about the same situation in different words could get different guidance.
The authors concluded that these tools need more safeguards and clinical validation before they are used in real care.
What it does not show
It does not tell us how often the chatbots got things dangerously wrong. The authors say their fact-check was a search for examples, not a systematic count, so they did not report an overall rate. In one test, they asked GPT-4 whether it is safe to detox at home when quitting long-term heroin use, worded 100 different ways. It said yes 23 times.
It tested two specific chatbots at one point in time. It does not test later versions of these systems, which may behave differently.
The researchers had to work around the chatbots' safety settings. At default settings, both systems often declined these questions, so the team used a prompt saying the request was for research and no one was at risk, and told LLaMA-2 not to avoid giving medical advice. Answers to an ordinary user may differ.
It does not show what happened to anyone who received these answers. It is an evaluation of the text, not a study of patient outcomes.
Why it matters
The danger is in the combination. The authors call it "a risky mix." An answer can look high quality on first inspection and still contain inaccurate and, in the authors' words, "potentially deadly medical advice." A wrong helpline or an endorsement of home detox inside an otherwise good answer is easy to miss. In this study, clinicians did not reliably catch it. They were asked to rate the answers, not to fact-check them, and the full paper reports that only one clinician flagged the answer to a question about detoxing at home from long-term heroin use. A person reading alone may not catch it either.
If you use a chatbot for recovery questions, do not rely on any emergency number it gives you. In an emergency, call 911. For a suicidal crisis, call or text 988. Check any medical step against a clinician or an official source before acting on it. The wording finding matters too. Rephrasing a question can change the answer, so a single answer is not a settled one.
Clinicians can put this to use at the next visit: ask patients whether they use chatbots for recovery questions, and what those tools have told them.
Sources
- Giorgi S, Isman K, Liu T, Fried Z, Sedoc J, Curtis B. Evaluating generative AI responses to real-world drug-related questions. Psychiatry Research. 2024 Sep;339:116058. Epub 2024 Jun 26. doi:10.1016/j.psychres.2024.116058; PMID 39059040; PMCID PMC11705880 https://pubmed.ncbi.nlm.nih.gov/39059040/
Published by ZSKFL Management.