PHQ‑9 scores run 0 to 27 and GAD‑7 scores run 0 to 21, and on both instruments a score of 10 or above is the standard cutoff for a probable diagnosis. To bill CPT 96127 for administering either one, a progress note has to do four things: name the instrument, record the score, interpret it, and show that the score changed clinical decision‑making. Most notes do the first two and stop.
This guide covers how both instruments are scored, what counts as a clinically meaningful change, what a payer expects to find in the note, and how to capture all of it without adding a second documentation task to your day. Clinical figures verified against the original validation studies on 27 August 2026.
PHQ-9 and GAD-7 at a Glance
Both instruments come from the same research group and share a structure: every item is scored 0 to 3, and the item scores are summed. Neither has reverse‑keyed items, so a higher total always means more symptom burden.
PHQ-9 | GAD-7 | |
|---|---|---|
Measures | Depression severity | Generalized anxiety severity |
Items | 9 | 7 |
Score range | 0–27 | 0–21 |
Standard cutoff | 10 or above | 10 or above |
Sensitivity at cutoff | 88% | 89% |
Specificity at cutoff | 88% | 82% |
Validation study | Kroenke, Spitzer & Williams (2001) | Spitzer, Kroenke, Williams & Löwe (2006) |
Severity Bands
These bands are what your interpretation line should reference. Using the published band language also makes the note legible to anyone reviewing it later, including a payer.
Severity | PHQ-9 score | GAD-7 score |
|---|---|---|
Minimal or none | 0–4 | 0–4 |
Mild | 5–9 | 5–9 |
Moderate | 10–14 | 10–14 |
Moderately severe | 15–19 | — |
Severe | 20–27 | 15–21 |
Note that the two scales do not map onto each other. The GAD‑7 has four bands and tops out at 21; the PHQ‑9 has five and tops out at 27. A PHQ‑9 of 16 and a GAD‑7 of 16 are not the same clinical picture, and writing them as though they were is a common error in combined mental‑health notes.

What Counts as a Meaningful Change
This is the part that turns a score into a treatment decision, and it is where most documentation is weakest. A score on its own says where the patient is. A change says whether what you are doing is working.
- PHQ-9: a change of about 5 points is the commonly used threshold for a clinically meaningful shift.
- GAD-7: a change of about 4 points is the commonly used threshold.
Both of those are rules of thumb, and it is worth knowing why they are approximate. More recent work using the ED50 method found that the minimal clinically important difference scales with baseline severity: a patient starting at a very high score needs a much larger absolute drop to feel meaningfully better than a patient starting in the mild range. Treating 5 points as a universal constant will overstate improvement in your most severe patients and understate it in your mildest.
What a Payer Expects to Find in the Note
CPT 96127 covers a brief emotional or behavioral assessment using a standardized instrument, and the PHQ‑9 and GAD‑7 are the two most commonly billed under it. Administering the instrument is not what makes it billable — documenting it properly is. Four elements have to be present:
The fourth element is the one that gets dropped, and it is the one that most often decides an audit. A score recorded and never referred to again reads as a box being ticked rather than an assessment being used.
On limits: Medicare applies a Medically Unlikely Edit of 3 units per date of service, and the 2026 national average reimbursement is roughly $5 per unit, varying by locality. Commercial payers set their own unit caps and annual frequency limits, and several differ from Medicare. Confirm the policy for each payer you bill rather than assuming the Medicare rule applies.

What a Complete Entry Looks Like
Short, and it contains all four elements. The point is that this is two sentences, not a paragraph — the barrier to doing it properly is habit, not time:
Every billable element is there. The instrument is named, the score is recorded, the interpretation includes both the band and the delta, and the final sentence shows the score driving the plan.
The Ninth Item Is a Safety Item
The PHQ‑9's ninth item asks about thoughts of self‑harm or being better off dead. It is scored the same 0 to 3 as every other item and folds into the total, which creates a documentation trap: a patient can endorse that item and still land in a low or moderate total.
Treat a non‑zero response on that item as a separate finding from the total score, and document the risk response you took — the assessment performed, the safety plan discussed, the follow‑up interval agreed. A total score alone does not evidence that you addressed it, and the GAD‑7 has no equivalent item, so anxiety screening does not cover this ground.
How Often to Administer
There is no single mandated interval, and cadence should follow the clinical question rather than a calendar. In practice most measurement‑based care protocols land on:
- At intake, to establish a baseline — without one, no later score can show change.
- Every 2 to 4 weeks during active treatment changes, which is roughly when a medication or therapy shift becomes measurable.
- Every 1 to 3 months during maintenance, to catch drift before it becomes relapse.
- At any point the clinical picture changes materially, regardless of when the last one was administered.
The failure mode is administering it often and never comparing. Ten scores in a chart with no delta ever written is ten data points and zero measurement‑based care.
Making It Repeatable
The reason outcome measures get documented inconsistently is almost never that clinicians do not know the bands. It is that the note has no fixed place for them, so the score lands wherever there is room — sometimes in Subjective, sometimes in Assessment, sometimes in a sentence that never mentions the instrument by name.
The fix is structural. A dedicated section in your note template, with the four billable elements as its prompts, means the same information gets captured the same way at every visit regardless of how busy the day is. In Twofold, custom templates let you define exactly that: a named section that asks for the instrument, the score, the interpretation with the delta, and the resulting decision.
That does not make the clinical judgement for you, and it should not. What it removes is the variability — the reason a chart review turns up three different formats for the same measure across six months.
Common Failure Modes
What goes in the note | Why it fails |
|---|---|
"Depression screen completed" | Instrument not named — not billable under CPT 96127 |
"PHQ-9 in the moderate range" | No score recorded; the band cannot be verified |
"PHQ-9 14" | No interpretation and no change from baseline; a data point, not an assessment |
Score recorded, plan unchanged and unexplained | No link to clinical decision-making — the element audits catch |
Item 9 endorsed, only the total documented | Risk response not evidenced; the clinical exposure, not just a billing one |
Scores at every visit, no delta ever written | Measurement without measurement-based care |
The Short Version
Score both instruments 0 to 3 per item, read the total against the published bands, and write the change alongside the number. Name the instrument, record the score, interpret it, and show what you did about it. Treat a non‑zero item 9 as its own finding with its own documented response. Then put all of that in a fixed section of your template so it happens the same way every time.

