Loading…
Founder & DClinPsy Trainee · 23 June 2026

The PHQ-9 and GAD-7 are the two most widely used routine outcome measures in UK psychological therapy — the backbone of session-by-session monitoring in NHS Talking Therapies (formerly IAPT) and a fixture of private practice too. They are quick, free, validated, and sensitive to change. But a score is only as useful as your ability to interpret it — and the metrics that services actually report (recovery, reliable improvement) are easy to muddle.
This guide covers how to score each measure, what the bands mean, the clinical cut-offs, and — the part that trips people up — how the NHS Talking Therapies outcome metrics are defined. It's written for therapists; the measures themselves are screening and monitoring tools, not diagnostic instruments.
The Patient Health Questionnaire-9 (PHQ-9) asks about the nine symptoms of major depression from the DSM, over the last two weeks. Each item is rated:
0 — Not at all
1 — Several days
2 — More than half the days
3 — Nearly every day
Sum the nine items for a total from 0 to 27. The conventional severity bands are:
0–4 — None / minimal
5–9 — Mild
10–14 — Moderate
15–19 — Moderately severe
20–27 — Severe
A score of 10 or above is the commonly used cut-off suggesting clinically significant depression — at this threshold the PHQ-9 has good sensitivity and specificity for major depression. Treat that as a flag for fuller assessment, not a diagnosis (more on that below).
One item needs special attention. PHQ-9 item 9 asks about "thoughts that you would be better off dead, or of hurting yourself in some way." Any response above 0 warrants a direct risk follow-up regardless of the total score — a client can sit in a "mild" total band and still endorse item 9. Make checking item 9 a habit, every administration.
The Generalised Anxiety Disorder-7 (GAD-7) uses the same response scale (0–3) over the last two weeks, across seven items, for a total from 0 to 21. The severity bands:
0–4 — Minimal
5–9 — Mild
10–14 — Moderate
15–21 — Severe
A score of 10 or above is the standard clinical cut-off for probable GAD, with good sensitivity and specificity. Although designed for GAD, it also performs reasonably as a screen for panic, social anxiety, and PTSD — useful, but a reason to follow a high score with proper assessment rather than assuming the diagnosis.
Anxiety and depression are highly comorbid, so the PHQ-9 and GAD-7 are typically administered as a pair — together they give a fuller picture than either alone, and in NHS Talking Therapies they're the two measures on which the headline outcomes are calculated. Administering both, every session, is what makes genuine outcome tracking possible.
This is where interpretation gets technical — and where a lot of therapists are fuzzy. NHS Talking Therapies (formerly IAPT) reports outcomes using a few specific, defined terms. Getting them right matters, because they mean precise things:
Caseness. A client is "a case" at the start of treatment if they score at or above the clinical threshold on either measure. The thresholds used in the dataset are PHQ-9 ≥ 10 and GAD-7 ≥ 8. (Note the asymmetry: the service "caseness" threshold for the GAD-7 is 8, slightly lower than the 10 clinical cut-off quoted above — a common source of confusion. Check the current manual for the figures your service uses.)
Recovery. A client who started as a case moves to below caseness on both measures by the end of treatment. (If they were never a case to begin with, they can't "recover" in this technical sense.)
Reliable improvement. The score drops by more than the measure's measurement error — i.e., a real change, not noise. The reliable-change thresholds commonly applied are a fall of ≥ 6 points on the PHQ-9 and ≥ 4 points on the GAD-7.
Reliable recovery. The gold-standard outcome: the client has both recovered and reliably improved.
Reliable deterioration. The mirror image — a score that rises by more than the measurement error, signalling the client is reliably worse. Worth catching early.
Verify the exact figures for your setting. Caseness thresholds and reliable-change values are defined in the NHS Talking Therapies for Anxiety and Depression Manual and are periodically updated. Use this section to understand the logic of the metrics; confirm the precise numbers against the current manual before reporting against them.
The practical upshot: a single end-of-therapy score tells you severity; two scores, start and end, tell you whether the person reliably improved and recovered — which is the outcome that actually matters.
It's worth saying plainly, because the numbers can feel deceptively definitive: the PHQ-9 and GAD-7 are screening and monitoring tools, not diagnostic instruments. A high score indicates a likely problem worth assessing properly; it does not, by itself, diagnose major depression or GAD. A low score doesn't rule out a clinically significant problem, and neither measure captures risk, function, or context on its own. They earn their place by being quick and repeatable — ideal for tracking change — not by being definitive.
A score of 10 or above is the commonly used cut-off suggesting clinically significant depression, where the PHQ-9 has good sensitivity and specificity. Treat it as a flag for fuller assessment, not a diagnosis — the PHQ-9 is a screening and monitoring tool, not a diagnostic instrument. And check item 9 (thoughts of being better off dead or of self-harm) on every administration, regardless of the total.
No. A GAD-7 of 10 or above is the standard clinical cut-off for probable GAD, but a high score is a reason to follow up with proper assessment, not to assume the diagnosis. Like the PHQ-9, it is a screening and monitoring measure.
Recovery means a client who started as a case moves below the caseness threshold on both measures by the end of treatment. Reliable improvement means the score dropped by more than the measure's measurement error — a real change, not noise. Reliable recovery — both at once — is the gold-standard outcome. The exact thresholds are defined in the national manual and updated periodically, so confirm the current figures before reporting against them.
A measure administered once is a snapshot. Administered every session and plotted over time, it becomes one of the most useful things in the room:
It makes progress visible to a client who feels stuck — a falling line is powerful evidence that things are moving.
It flags non-response early — if scores aren't shifting after several sessions, that's a prompt to revisit the formulation, not to keep going regardless.
It anchors shared decisions about stepping up, continuing, or ending.
It catches reliable deterioration before it becomes a crisis.
The friction is in the admin — scoring by hand, transcribing totals, plotting trends. That's exactly the busywork worth automating: in Formulate, you can assign the PHQ-9 and GAD-7 as homework, have them scored automatically, and see the trend line plotted session by session — so the measurement supports the therapy instead of competing with it for your time.
Track PHQ-9, GAD-7 and the full validated measures suite — auto-scored, plotted, and assignable as homework. Start free →
This article is for qualified therapists and trainees and is educational, not clinical advice. Scoring thresholds and NHS Talking Therapies outcome definitions should be confirmed against the current national manual. The PHQ-9 and GAD-7 are screening measures and do not replace clinical assessment; always follow up endorsement of PHQ-9 item 9 with an appropriate risk assessment.
Formulate provides educational CBT resources for use with a qualified therapist. They are not a substitute for professional assessment, diagnosis, or crisis care.
If you need urgent help now:
Founder & DClinPsy Trainee
Tarun Vermani is the founder of Formulate and a trainee clinical psychologist (DClinPsy). He writes about CBT formulation, outcome measurement, and the tools that help clients engage with therapy between sessions. These articles are educational, written for qualified therapists and trainees, and are intended to support — not replace — clinical training, supervision and judgement.
All articles by Tarun Vermani →