SUDrecovery.ai
Help now

Study

An AI model read VA notes for substance use disorder severity. It beat simple text rules in most categories, but exact matches were far from perfect

Source date · Reviewed

An open-source AI model read 520 VA notes for substance use disorder severity. It beat text rules in 7 of 11 categories; exact-match scores ran 40 to 67.

Study card

Who was studied
Clinical notes from Department of Veterans Affairs patients, drawn from visits between October 2015 and November 2023 and selected for substance use diagnosis codes or SUD clinic visits
How many
577 notes from 574 patients: 57 used to write the prompts, 520 used for testing
Design
Methods study comparing a language model with hand-written text rules against one expert's labels. Run as a VA quality improvement project.
What was tested
Flan-T5-XXL, an open-source language model, given written instructions (zero-shot prompts) to pull out each substance use disorder diagnosis with its severity or remission wording, plus filtering steps to cut made-up answers
Compared with
Rule-based text matching (regular expressions)
Main outcome
Agreement with a licensed clinical psychologist's labels across 11 substance categories, scored as F1 (0 to 100) for exact and partial matches
Result
On notes that contained a diagnosis, exact-match F1 ranged from 40.00 to 66.67 by category (alcohol 61.80, opioids 58.90); partial-match F1 ranged from 59.05 to 83.05. The model beat the text rules on exact match in 7 of 11 categories.
Limitations
One health system, one annotator, a small test set with few notes containing each diagnosis, and one model family. The authors say the models are not expected to work outside VA notes.
Funding and conflicts
Supported by the VA Office of Mental Health and Office of Suicide Prevention, using VA-funded computing at Oak Ridge National Laboratory, which is also supported by the Department of Energy. The authors declared no competing interests.
What this does NOT tell us
Whether using this tool changes care or outcomes for any patient, or how accurate it would be on notes from another health system.

The short version

The diagnosis codes used for billing, such as ICD-10, lack detail for some substance use diagnoses, the authors write, so clinicians often record the fuller diagnosis, including how severe it is, in the text of the note. Researchers at Oak Ridge National Laboratory and the VA tested whether an open-source AI language model could find those diagnoses and their severity in VA clinical notes. It beat simple text rules in most substance categories, but its exact-match scores ran from 40 to 67 out of 100, and it made errors the study's clinical expert flagged, such as reading an old problem list as a current diagnosis.

What they did

The team took 577 clinical notes from 574 VA patients, from visits between October 2015 and November 2023. Notes were picked because they carried a substance use diagnosis code or came from a substance use clinic visit. They spanned 79 note types, from mental health and psychiatry to nursing, suicide prevention and administrative notes.

A licensed clinical psychologist read each note and marked every substance use disorder diagnosis and its severity or remission wording, such as "mild," "severe," or "in sustained remission," across 11 substance categories: alcohol, opioids, cannabis, sedatives, cocaine, amphetamines, caffeine, hallucinogens, nicotine, inhalants and other substances.

The researchers used 57 notes to write and refine instructions for Flan-T5, an open-source model from Google, without training it further. They ran all five sizes of the model on the other 520 notes and report the largest, with 11 billion parameters, which did best. They added filtering steps to throw out answers that did not match text actually in the note, and compared the results with hand-written text-matching rules. The paper says closed models such as GPT-4 were not an option because of privacy concerns and data-use restrictions.

What it found

Scores are F1, a measure from 0 to 100 that balances catching what is there against adding what is not.

On notes that contained a diagnosis, exact-match F1 ranged from 40.00 (other substances) to 66.67 (nicotine and inhalants), with alcohol at 61.80 and opioids at 58.90. When partial overlap counted, the range was 59.05 to 83.05. On notes with no diagnosis for a category, the model was usually right to find nothing, with exact-match F1 from 92.44 to 100.

The model beat the text rules on exact match in 7 of the 11 categories; the rules did better in the other 4. The authors say the model tended to add extra words rather than miss the diagnosis, and that telling it to copy text exactly made it drop important details instead.

Reviewing 55 errors with the psychologist turned up real mistakes. The model sometimes read an old problem list as a current diagnosis. It sometimes attached the wrong severity: in one example, a note listing unspecified alcohol use disorder and severe cocaine use disorder came back as severe alcohol use disorder. It also labeled a stimulant disorder as amphetamine when the expert could tell from context that the drug was cocaine.

The model took about 34 minutes of GPU time per substance category for the 520 notes. The text rules took about 0.02 minutes, though the authors note that writing and maintaining rules takes expert time that is hard to measure.

What it does not show

This is a proof of concept, not a working clinical tool. The study used existing notes and did not put the tool into anyone's care, and it did not test whether better-organized severity data changes treatment or outcomes.

The ground truth came from one psychologist. The test set was small, and the share of notes containing a given diagnosis ranged from 3% to 33%. The VA would not allow exact counts for some categories to be published.

The authors say the models were built for VA decision-support systems and are not expected to be valid outside VA records.

The work was done as VA quality improvement, which the VA treats as non-research; Oak Ridge National Laboratory's review board also approved it. It was funded by the VA and run on VA-funded computing, and the authors declared no competing interests.

Why it matters

Much of the detail about a patient's substance use, such as severity, lives in free text rather than billing codes, which is why health systems want software that can read notes. This study suggests the idea is workable, and it shows exactly where it breaks: reading history as current, and pairing the wrong severity with the wrong substance. Those mistakes matter when a record follows someone into their next visit.

The team also made a privacy choice worth noticing. Because the notes are sensitive, it ran an open model on computers it controlled rather than sending records to a commercial service. Outside the VA, records from federally assisted substance use treatment programs carry extra federal protection under 42 CFR Part 2, covered in our entry on the rule. Part 2 does not apply to VA health records; VA substance use records are protected under a separate federal law, 38 U.S.C. 7332. The paper discusses neither law; connecting them is this site's framing.

For patients: what a clinician writes in a note can be read by software later. It is reasonable to ask a program or health system whether AI tools read your records, and for what.

Sources

  1. Mahbub M, Dams GM, Srinivasan S, Rizy C, Danciu I, Trafton J, Knight K. Decoding substance use disorder severity from clinical notes using a large language model. npj Mental Health Research. 2025 Feb 7;4(1):5. doi:10.1038/s44184-024-00114-6; PMID 39915681; PMCID PMC11802718 https://pubmed.ncbi.nlm.nih.gov/39915681/
  2. 42 CFR 2.12(c)(1), Applicability: exceptions, Department of Veterans Affairs. Electronic Code of Federal Regulations, current as of September 29, 2026. https://www.ecfr.gov/current/title-42/chapter-I/subchapter-A/part-2/subpart-B/section-2.12

Published by ZSKFL Management.

Help now

Someone won't wake up or isn't breathing well

Call 911 and give naloxone (Narcan) now if you have it. If they don't wake up in 2 to 3 minutes, give another dose. Keep giving a dose every 2 to 3 minutes until they wake up. Stay with them until help arrives.

Call 911

Thinking about suicide or hurting yourself

I want to find treatment

SAMHSA National Helpline, free and confidential, 24/7.

Call 1-800-662-4357

No one monitors this site. These numbers connect you to real people.