Study
Researchers used AI to read 8 years of an opioid recovery forum. It described the past well but predicted setbacks poorly
Source date · Reviewed
Language models sorted 32,810 r/OpiatesRecovery posts by recovery stage and replies by support type. Predicting who would progress or slip back stayed hard.
Study card
- Who was studied
- Public posts and comments in r/OpiatesRecovery, a Reddit community for opioid recovery, from January 2014 to May 2022
- How many
- 32,810 posts and 324,224 comments. Recovery paths for 2,936 users who posted at least twice (8,070 pairs of back-to-back posts).
- Design
- Observational study of public online posts. Language models trained on human-labeled examples assigned recovery stages to posts and support types to comments.
- What was tested
- Fine-tuned language models (BERT and RoBERTa) and two open-source large language models (Llama-3.1-8B and Qwen2.5-14B) used without extra training
- Compared with
- For prediction: a baseline that always guesses no change in recovery stage
- Main outcome
- Differences in language and support across recovery stages, and how well models could predict a user's next stage change from a current post
- Result
- Posts from people still using had more negative, painful, and passive language and drew more advice and facts but less encouragement and sympathy. The best model predicted the next stage change with a weighted F1 score of 0.59, against 0.44 for always guessing no change. The large language models scored 0.11 to 0.43.
- Limitations
- One online community. Recovery stages were self-described, not clinically verified, and labeled by software. Associations only. Effect sizes were small. Time between posts ranged from the same day to more than seven years.
- Funding and conflicts
- The authors declared no funding and no conflicts of interest.
- What this does NOT tell us
- Whether advice or encouragement from other members caused anyone's recovery to move forward or back, or how people who never post, or who get support offline, are doing.
The short version
Researchers used AI language models to read eight years of posts in r/OpiatesRecovery, which the authors describe as the largest opioid recovery community on Reddit. The models sorted each post by the writer's stage of recovery and each reply by the kind of support it offered.
People still using wrote in a darker, more passive voice and drew more advice and less encouragement. Predicting who would move forward or slip back next was hard: the best model did modestly, and two open-source large language models did worse than guessing "no change" every time.
What they did
The team collected 32,810 posts and 324,224 comments from the forum, covering January 2014 to May 2022.
Each post was labeled with one of four recovery stages: still using, early recovery (abstinent less than a month), sustained recovery (one month to five years), and stable recovery (more than five years). Posts that did not reveal a stage were set aside. Two researchers hand-labeled a sample first and agreed on 74.3% of posts, which the authors call moderate reliability.
Each comment was labeled for 11 kinds of support: five informational, such as advice, facts, referrals, personal stories, and opinions, and six emotional, such as encouragement and sympathy. Software trained on labeled examples, including outside labeled data sets, sorted the rest.
To follow people over time, the team took the 2,936 users who posted at least twice and paired each post with that person's next one. Of 8,070 pairs, 24% showed forward movement, 55.1% no change, and 20.9% a step back.
What it found
Language. Posts from people still using had fewer joyful, positive, and trusting words and more negative, pain-related, fearful, and passive ones than posts from people further along.
Support. Posts from people still using drew more advice, facts, referrals, and opinions, and less encouragement, sympathy, and emotional reaction, than posts from people in recovery. The authors suggest this may reflect stigma toward people who are still using, even inside a recovery community, and note the study cannot show a cause.
Stage changes. When a person's next post showed forward movement, the earlier post had drawn more advice, facts, and opinions than posts followed by no change. Posts followed by a step back also drew somewhat more facts than posts followed by no change. Posts followed by no change drew more encouragement. The authors read this as change in either direction going with more practical replies. The differences were statistically significant but small.
Prediction. The best model, a fine-tuned RoBERTa that also learned from the comments, predicted the next stage change with a weighted F1 score of 0.59. F1 runs from 0 to 1 and balances false alarms against misses. Always guessing "no change" scored 0.44. The model was weakest at spotting setbacks, with an F1 of 0.33 for that group. Two open-source large language models, prompted without extra training, scored 0.11 to 0.32 (Llama) and 0.35 to 0.43 (Qwen). Neither beat the always-no-change guess, and both scored near zero at spotting setbacks in almost every setting.
What it does not show
This is a study of what people wrote, not of what happened to them. Recovery stages came from posts, not a clinician, and were assigned by software that was right most of the time but not always. Early recovery was the hardest label for it.
The link between practical support and forward progress is an association. People about to make progress may have written posts that invited advice, or had help offline that the study cannot see. The authors say directly that the study does not show that forum support influences recovery.
This is one community on one platform, and people who post there, and keep posting, are a self-selected group. The time between a person's posts ranged from the same day to more than seven years, which makes "next stage" a loose idea.
The large language models were tested only with prompts, not trained for the task, so the result says little about how a purpose-built model would do.
The posts were public and collected through Reddit's API. An institutional review board at the University of Kentucky determined the project was not human-subjects research. The authors say they did not report usernames, post titles, links, or verbatim quotes, and they are not releasing the data publicly. They declared no funding and no conflicts of interest, and disclosed using ChatGPT to polish their writing, not for analysis.
Why it matters
If you post in an online recovery community, your public posts can become research data without anyone asking you. These researchers took steps to keep people from being identified; nothing here says others do the same.
People who were still using got more instructions and less warmth. That may be what they asked for, or it may reflect the stigma the authors suggest. Either way, it is a pattern a community can notice.
The prediction result matters most for anyone building or buying recovery tools. Some apps and services may claim they can tell from what you write that you are headed for a setback. In this study, trained models were weak at exactly that, and off-the-shelf language models failed at it. A claim like that should come with evidence from real people and real outcomes.
Sources
- Yu X, Chen HY, Chi Y. Language, Social Support, and Recovery-Stage Transitions in Opioid Use Disorder on Reddit: Computational Analysis. Journal of Medical Internet Research. 2026 Sep 17;28:e91054. doi:10.2196/91054; PMID 42752349; PMCID PMC13585015 https://pubmed.ncbi.nlm.nih.gov/42752349/
Published by ZSKFL Management.