Find out whether your recovery plan has ever carried real work, because testing the parts isn’t the same evidence.
How do companies pay for work nobody approved?
By not asking for the whole thing at once.
Building a disaster recovery site was one of the main reasons I was hired, and it became my first major responsibility. There was no budget for it.
I’d money for machines. I’d nothing for people, and I couldn’t hire anyone to do the work.
So I sent the people I already had out to the site periodically to build a piece of it, and I managed it that way. A trip here, a trip there, across years.
There was never a project. Never a headcount request, never an approval for the thing as a whole, never a line item anyone could point at and cut.
It existed because somebody kept sending people out there.
Is that how IT resilience usually gets built?
In companies under a couple hundred people, frequently yes, and it’s the opposite of how it gets sold.
The vendor version is a program with a start date, a budget, and a completion milestone. Somebody signs, work happens, and at the end you have a capability.
The version that happens is a person who decides it matters and keeps chipping at it. No single trip needs approval, so no single trip gets refused.
That approach has an obvious weakness, which is that it depends entirely on one person continuing to care. It also has one advantage nothing else offers. It survives budget cycles, because there’s nothing on the budget to cut.
Why was the site only forty miles away?
Because that was the limit of the technology.
The site sat about forty miles from the main computer room. That distance wasn’t a preference or a compromise between risk and cost. It was as far as we could go and still keep the systems copying properly.
Connectivity started as a few circuits, became a faster one, and eventually became something much faster than that. Every improvement changed what was possible, and bandwidth was the binding constraint on how good the copy could be.
What’s worth noticing is what a hard constraint does to a decision. Forty miles was argued about, measured, and understood by everyone involved, because we’d no choice about it.
Today distance is nearly free. A company can put a copy of its data anywhere, and because the tradeoff disappeared, almost nobody reasons about it. When a constraint is forced, people think about it. When it vanishes, they stop.
What does a disaster recovery site need besides computers?
Somewhere for people to sit and work.
Ours had a duplicate machine running, receiving data, building logs. That part gets all the attention and it’s the part everyone plans.
It also had laptops, computers, and a full twenty-desk setup, ready to go.
That was deliberate. A disaster that takes out a computer room usually takes out the building the computer room is in, and the people who run the business need somewhere to be. A recovery plan that restores the systems and leaves forty people with nowhere to sit has solved half the problem.
What is the difference between a tested system and a proven one?
We tested that site. We never failed over to it live.
The testing was real. Systems were exercised, the copy was verified, the equipment was checked. That work wasn’t theater and it caught things.
What never happened was the business running on it. No day where the main site was gone and the company operated from forty miles away.
My word for that is fortunately, and I mean it.
Testing establishes that the parts work under known conditions. A live cutover establishes that the whole thing carries real load, real users, and real decisions made under pressure by people who have never done it before. Only the second one is proof, and almost nobody has it.
Should you run a full failover test?
Most companies should and most won’t, and the reasons are legitimate.
A genuine failover test means taking the business down deliberately. It means accepting that something might not come back, and it means somebody has to authorize a self-inflicted outage on a normal Tuesday.
The version that works for a smaller company is narrower. Fail over one system and not everything. Do it during a quiet period. Have the people who would run the recovery do it, instead of the person who designed it.
That last detail matters more than the scale. A recovery executed by its designer proves the design. A recovery executed by whoever is on duty proves the documentation, and the documentation is what you’ll have.
What should you tell your board about recovery readiness?
The truth about which kind of evidence you hold.
We have a recovery site and we test it regularly is accurate and honest. We’re covered implies something stronger, and the gap between those two statements is where companies discover their plan was a belief.
Saying it plainly costs nothing and it changes the conversation. A board told that the parts are tested and the whole has never been exercised will usually ask what a full test would take. That’s exactly the question you want asked.
Every company with a backup strategy sits somewhere on that line. Knowing where’s most of the value.
