The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

We Built a Disaster Recovery Site Without a Budget

This entry is part 18 of 21 in the series The Operations Room
TL;DR: A tested system and a proven system are different things, and most companies say covered when they mean tested. I built a disaster recovery site over several years with no budget and no headcount, tested it regularly, and never once ran the business on it. Knowing which kind of confidence you have changes how you talk about your own readiness.

Find out whether your recovery plan has ever carried real work, because testing the parts is not the same evidence.

How do companies pay for work nobody approved?

By not asking for the whole thing at once.

Building a disaster recovery site was one of the main reasons I was hired, and it became my first major responsibility. There was no budget for it.

I had money for machines. I had nothing for people, and I could not hire anyone to do the work.

So I sent the people I already had out to the site periodically to build a piece of it, and I managed it that way. A trip here, a trip there, across years.

There was never a project. Never a headcount request, never an approval for the thing as a whole, never a line item anyone could point at and cut.

It existed because somebody kept sending people out there.

Is that how resilience usually gets built?

In companies under a couple hundred people, frequently yes, and it is the opposite of how it gets sold.

The vendor version is a program with a start date, a budget, and a completion milestone. Somebody signs, work happens, and at the end you have a capability.

The version that happens is a person who decides it matters and keeps chipping at it. No single trip needs approval, so no single trip gets refused.

That approach has an obvious weakness, which is that it depends entirely on one person continuing to care. It also has one advantage nothing else offers. It survives budget cycles, because there is nothing on the budget to cut.

Why was the site only forty miles away?

Because that was the limit of the technology.

The site sat about forty miles from the main computer room. That distance was not a preference or a compromise between risk and cost. It was as far as we could go and still keep the systems copying properly.

Connectivity started as a few circuits, became a faster one, and eventually became something much faster than that. Every improvement changed what was possible, and bandwidth was the binding constraint on how good the copy could be.

What is worth noticing is what a hard constraint does to a decision. Forty miles was argued about, measured, and understood by everyone involved, because we had no choice about it.

Today distance is nearly free. A company can put a copy of its data anywhere, and because the tradeoff disappeared, almost nobody reasons about it. When a constraint is forced, people think about it. When it vanishes, they stop.

What does a disaster recovery site need besides computers?

Somewhere for people to sit and work.

Ours had a duplicate machine running, receiving data, building logs. That part gets all the attention and it is the part everyone plans.

It also had laptops, computers, and a full twenty-desk setup, ready to go.

That was deliberate. A disaster that takes out a computer room usually takes out the building the computer room is in, and the people who run the business need somewhere to be. A recovery plan that restores the systems and leaves forty people with nowhere to sit has solved half the problem.

What is the difference between a tested system and a proven one?

We tested that site. We never failed over to it live.

The testing was real. Systems were exercised, the copy was verified, the equipment was checked. That work was not theater and it caught things.

What never happened was the business running on it. No day where the main site was gone and the company operated from forty miles away.

My word for that is fortunately, and I mean it.

Testing establishes that the parts work under known conditions. A live cutover establishes that the whole thing carries real load, real users, and real decisions made under pressure by people who have never done it before. Only the second one is proof, and almost nobody has it.

Should you run a full failover test?

Most companies should and most will not, and the reasons are legitimate.

A genuine failover test means taking the business down deliberately. It means accepting that something might not come back, and it means somebody has to authorize a self-inflicted outage on a normal Tuesday.

The version that works for a smaller company is narrower. Fail over one system rather than everything. Do it during a quiet period. Have the people who would run the recovery do it, rather than the person who designed it.

That last detail matters more than the scale. A recovery executed by its designer proves the design. A recovery executed by whoever is on duty proves the documentation, and the documentation is what you will have.

What should you tell your board about recovery readiness?

The truth about which kind of evidence you hold.

We have a recovery site and we test it regularly is accurate and honest. We are covered implies something stronger, and the gap between those two statements is where companies discover their plan was a belief.

Saying it plainly costs nothing and it changes the conversation. A board told that the parts are tested and the whole has never been exercised will usually ask what a full test would take, which is exactly the question you want asked.

Every company with a backup strategy sits somewhere on that line. Knowing where is most of the value.

The Guides That Get Your Book Written, Published, and Sold

Four short, practical guides on writing, publishing, and selling your book, plus the occasional note when there's something worth your time. No fluff, no daily inbox clutter. Drop your email and they're yours.

We use MailerLite to manage our list and send these emails. Your address is used only to send you what you signed up for. We will not sell it, share it, or use it for anything else, and you can unsubscribe anytime.

Frequently Asked Questions

How far away should a recovery site be?
Far enough that one event cannot affect both locations, considering power grids, flood zones, and weather systems rather than distance alone. Network performance used to cap this and rarely does now, so the limit is your recovery objectives instead of the technology.
How often should disaster recovery be tested?
At least annually for the full plan and more often for individual components. Test after any significant change, since a plan validated against last year’s systems describes a company that no longer exists.
Can a small company afford a recovery site?
Most cannot afford a second building and most do not need one. Hosted infrastructure in a different region covers the systems, and the remaining question is where people would work, which can be answered with an agreement rather than a lease.
Who should run a recovery test?
Whoever would run the real one, which is rarely the person who designed the plan. A test executed by the designer validates the design. A test executed by the on-duty person validates the documentation, and the documentation is what exists at three in the morning.
What gets forgotten in recovery planning?
Where people will physically work, how they will reach each other when the directory is on a failed system, who has authority to declare a disaster, and what happens if the event runs past a single day.
How do you fund resilience work with no budget?
In pieces small enough that no single one requires approval, using capacity you already have. The approach is slow and it survives budget cuts, because there is no line item to remove.

📁︎ Technology

🏷︎ Disaster Recovery🏷︎ IT Operations🏷︎ Risk Management

📝 Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

One comment

Was this useful?

Leave a comment