Designing Outcome Dashboards That Commissioners Can Audit: Avoiding Vanity Metrics in Community Mental Health

Dashboards are now a routine feature of contract management, mobilisation and tender evaluation. The risk is that services build dashboards that look sophisticated but cannot be audited: counts of contacts, satisfaction snapshots, or composite scores with no clear cohort rules and no evidence trail back to case files. A defensible dashboard does the opposite. It uses a small set of outcome measures aligned to pathway intent, defines cohorts clearly, and embeds reconciliation checks so the numbers can be tested. This article connects dashboard design to mental health outcomes and recovery and mental health service models and pathways, because dashboards only carry weight when they reflect real delivery and pathway reality.

Why dashboards fail: the common audit gaps

Commissioners tend to probe the same weak points:

  • Vanity metrics (contacts delivered, forms completed, “engagement rates”) presented as impact.
  • Unclear cohort definitions (who is included, when they enter/exit, and whether high-risk cases are mixed with routine support cases).
  • Inconsistent recording (different workers record outcomes differently, undermining comparability).
  • No reconciliation route (a commissioner cannot sample a case file and see where the dashboard figure came from).

A dashboard is only as credible as its audit trail. If the service cannot show how a metric is produced from routine records, the metric becomes a presentation tool rather than evidence.

What a commissioner-auditable dashboard looks like

1) Cohort rules that match pathway stages

Start with cohorts that align with the pathway and contract logic, for example:

  • Post-discharge cohort (first 4 weeks after discharge from acute care).
  • High-risk escalation cohort (active early warning plan and step-up thresholds in place).
  • Routine support cohort (stable support with scheduled reviews).

Each cohort should have entry criteria, exit criteria, and a clear reporting period. This prevents misleading averages and helps commissioners interpret “what good looks like” at each stage.

2) Measures that reflect outcomes, intensity and system impact

A defensible set of metrics usually includes three layers:

  • Outcome change (stability indicators, independence progression, participation ladder progress).
  • Support intensity (minutes/visits per week, prompting levels, step-up episodes).
  • System outcomes (crisis contacts, escalation frequency, step-down success, avoidable re-escalation).

This combination supports value for money arguments and reduces the temptation to treat “more activity” as “more impact”.

3) Definitions that can be applied consistently

Every metric needs a one-line operational definition that front-line staff can follow. For example, “step-down success” might be defined as: “support intensity reduced in line with plan and maintained for four weeks without crisis escalation or safeguarding concern requiring step-up.” When definitions are vague, recording becomes inconsistent and auditability collapses.

4) Reconciliation routes built into governance

Commissioners trust dashboards when providers can demonstrate routine controls:

  • Monthly reconciliation sampling (select cases and trace dashboard entries back to notes, risk reviews and logs).
  • Data completeness checks (missing baselines, missing review dates, inconsistent coding).
  • Exception review (high contact with low change; repeated step-ups; stalled step-down).

These are not “nice-to-haves”. They are the mechanism that turns a dashboard into defensible evidence.

Operational examples (building auditability into real practice)

Example 1: Crisis reduction metric that can be tested in case files

Context: The contract expects reduced crisis use. The service previously reported “fewer incidents” without definition, which commissioners challenged.

Support approach: The service defines a crisis contact metric (crisis line calls, urgent responses, unplanned emergency contacts) and links it to early warning plans for a defined high-risk cohort.

Day-to-day delivery detail: Staff record early warning indicators at each contact and code whether escalation thresholds were reached and what action was taken (step-up visit, partner liaison, crisis plan activation). Managers review crisis logs weekly and sample-check whether coded events are supported by dated notes and risk review updates.

How effectiveness/change is evidenced: Dashboard shows crisis contacts per person per month trending down within the high-risk cohort, alongside improved time-to-intervention. File sampling demonstrates that each recorded crisis contact and each step-up action is visible in notes and review records, making the metric auditable.

Example 2: Independence progression shown through prompting levels

Context: Commissioners ask whether support is building independence or maintaining reliance. Contact counts alone cannot answer this.

Support approach: The service uses a graded prompting scale (independent / prompted / supported) across selected routine tasks and appointments, with planned step-down milestones.

Day-to-day delivery detail: Each visit records the prompting level for agreed tasks and any adaptations made (visual cues, rehearsal, coping tools). Supervision includes checks for “prompt creep” and ensures temporary step-ups are documented with review dates and safeguarding rationale where relevant.

How effectiveness/change is evidenced: Dashboard shows the proportion of tasks completed independently increasing over time and average support minutes reducing without a rise in incidents. Case file sampling confirms the prompting levels are evidenced in routine notes and reflected in review decisions.

Example 3: Participation metrics that capture meaningful engagement

Context: A service reports “group attendance”, but commissioners want to know whether participation is meaningful and sustained, not just “present”.

Support approach: A participation ladder is defined (attends / stays full duration / participates / attends independently / sustains for four weeks) and linked to coping plans and positive risk-taking.

Day-to-day delivery detail: Staff document preparation (travel plan, coping strategies), support level, debrief learning, and risk assessment updates based on observed competence. Managers check monthly that ladder coding matches narrative records and that risk decisions are documented when independence increases.

How effectiveness/change is evidenced: Dashboard shows ladder progression and sustained engagement rates, and audit checks confirm that each ladder stage is supported by dated notes and updated risk assessments rather than optimistic interpretation.

Explicit expectations that must be met

Commissioner expectation

Commissioners expect dashboards to be auditable, comparable and meaningful. They will look for clear cohort definitions, stable metric definitions, and evidence that reported figures can be reconciled to routine records. They also expect providers to use dashboards for improvement: identifying drift, investigating variation across localities or cohorts, and documenting corrective actions through governance.

Regulator / Inspector expectation (e.g. CQC)

Inspectors expect data use to improve care without undermining rights or safety. They will test whether staff understand plans, whether risk is managed proportionately, and whether outcome measurement drives inappropriate pressure or restrictive practice. Inspectors also look for governance integrity: that incidents, safeguarding concerns and learning are reflected in plans and that dashboards do not mask unmet need or delayed escalation.

Practical governance template for dashboard assurance

A simple, repeatable assurance cycle makes dashboards defensible:

  • Weekly operational review of escalation and high-risk exceptions.
  • Monthly governance meeting reviewing cohort trends, data completeness, and variance.
  • Monthly reconciliation sampling tracing selected dashboard entries back to case files.
  • Quarterly metric review to remove vanity measures and refine definitions with commissioner feedback.

When this cycle is embedded, dashboards become not a reporting add-on but a credible window into day-to-day delivery and pathway impact.