DiagnosticMind · ← All editions · PT

The Rehearsal Is the Risk

What a recovery drill demonstrates, and what it only performs

A continuity exercise is run. It completes inside its window. It is signed off, filed, and counted toward the obligation that required it. The organisation records itself as recovery-ready. This shape — the drill that passes, and the readiness inferred from its passing — is treated as the system working exactly as intended. It rewards a closer look, because the moment a manual drill is signed off is also the moment it is most likely to have demonstrated nothing it is credited with.


What the drill is graded on

An exercise is graded, almost everywhere, on whether it was run. Was the scenario executed? Were the steps followed? Did the window close with the boxes ticked? These are questions about the activity. None of them is the question the activity exists to answer: does the system recover, repeatably, without depending on the particular people who happened to be in the room?

The grade and the claim then drift apart, quietly. The grade records that the exercise was completed. The claim — made later, in a board paper or a regulator's file — records that the organisation is recovery-ready. Those are different statements. Only the first was tested.


The successful drill that proves nothing

Consider a recovery rehearsal performed by hand and judged a success. What, precisely, has it established? That on the day, with those people, under a scenario known in advance, a recovery-shaped sequence could be carried through to completion. That is a real fact. It is a fact about the availability and skill of the people present. It is not, on its own, a fact about the system.

The experienced operator who senses the wrong step before it is taken, the colleague who improvises around the gap the runbook never anticipated — their presence is what the drill actually measured. Change the people, move the hour, alter the scenario, and the same drill can return a different outcome. A rehearsal that succeeds proves the rehearsers were available. The recoverability of the system is inferred from that, not shown by it.


The rehearsal is the hazard

There is a second cost, and it is the one almost never named. A manual recovery exercise is not a neutral observation of the system. It touches production. It fails things over, reroutes, restores from backups, exercises the paths that are dangerous precisely because they are used so rarely. The more manual and the more ambitious the test, the more the test is itself an operation capable of causing the outage it was meant to rehearse.

This produces a quiet inversion. The exercise commissioned to reduce operational risk becomes, for the hours it runs by hand, a source of operational risk. Organisations sense this, even when they do not name it. It is why the most consequential drills tend to be run least often, scoped most narrowly, and booked into the quietest window that can be found. The rarity is read as prudence. It is at least as often an admission that the test is hazardous — and a test too dangerous to repeat is too infrequent to be evidence of anything.


The question with no box to tick

The diagnostic move is not to grade the drill more strictly. It is to ask what the drill is for. Two questions separate a rehearsal that produces evidence from one that produces a record:

Neither question has a box on the exercise report, because the report was built to capture completion, not establishment. The questions sit outside the form. That is precisely why they are rarely asked, and why the gap they would expose stays comfortable.


A drill graded on completion measures attendance. Whether the system recovers — repeatably, safely, without the particular people who were in the room — is a separate question. The exercise was never scored on it. The readiness was assumed anyway.