Latest
When a Client Thinks the Ghostwriter Used AIThe Clients Who Pay and VanishWhat an AI Detector Score on Your Manuscript Is WorthWhen Your Memoir Should Be a NovelWhat Belongs on a Copyright PageThe One-Hour Call Before I Quote Your BookThe Work You Would Never Have StartedWhen Your Own Memoir Sounds Like BraggingMonthly or Milestone: How Ghostwriting Gets BilledWhat It Costs to Fix an AI-Written ManuscriptThe Quotation Marks That Get Authors SuedThe Hugging Face AI Agent Attack: An Operations ReadingBehind the Book: The Mysterious Island, Neb’s SideHow to Organize Decades of Memories Into a MemoirWhy Rotten Tomatoes Sucks: The Score Does Not Mean What You ThinkWhy Amazon KDP Sucks: They Terminated My Account OvernightIngramSpark: How I Publish Now and WhyWhy Fiverr Sucks for Ghostwriting: The Buyer’s SideWhy eBay Sucks Now: A Seller’s Numbers and a Buyer’s WarningThe Ghost Story TraditionThe Gothic TraditionThe Christmas Ghost Story TraditionBooks to Give a WriterResurrection as a Narrative StructureThe Beach Read ArgumentWhy It’s a Wonderful Life Failed on ReleaseWhat to Read in SpringWhat to Read in SummerWhat to Read in OctoberHow Warner Bros. Dismantled a $17 Billion Cartoon EmpireThe Imaginary Scarcity TrapThe Graph That Goes Vertical Is Usually Somebody Else’sSubstack Is Not Collapsing. The Promise Was.The Disasters That Happen to Ordinary PeopleToba: The Winter That Almost Ended UsJay Stifflemire: Nothing Ever Gets Written DownGeorgie-Ann Getton: I Forgot I Had Free WillAI Detection Cannot Be Evidence, and Publishing Is Using It That WayAI Consciousness Left Philosophy and Entered the LaboratoryThe Office Block Where the Bedrooms AreThe Web Got Fenced: What AI Search Costs Small SitesBlack Tuesday: The Web Ring War Nobody Outside It NoticedWhat the AI Visibility Industry Sells, and What the Evidence SaysBlack Tuesday: The Original ring-master.net Page, 2000Behind the Book: Peacekeeper, The Dissolution WarsBehind the Book: Real World SurvivalBehind the Book: Publish Your BookBehind the Book: ReincarnationBehind the Book: Sell Your BooksBehind the Book: Selling on eBay
The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

We Built a Disaster Recovery Site Without a Budget

This entry is part 18 of 21 in the series The Operations Room
TL;DR: A tested system and a proven system are different things, and most companies say covered when they mean tested. I built a disaster recovery site over several years with no budget and no headcount, tested it regularly, and never once ran the business on it. Knowing which kind of confidence you have changes how you talk about your own readiness.

Find out whether your recovery plan has ever carried real work, because testing the parts isn’t the same evidence.

How do companies pay for work nobody approved?

By not asking for the whole thing at once.

Building a disaster recovery site was one of the main reasons I was hired, and it became my first major responsibility. There was no budget for it.

I’d money for machines. I’d nothing for people, and I couldn’t hire anyone to do the work.

So I sent the people I already had out to the site periodically to build a piece of it, and I managed it that way. A trip here, a trip there, across years.

There was never a project. Never a headcount request, never an approval for the thing as a whole, never a line item anyone could point at and cut.

It existed because somebody kept sending people out there.

Is that how IT resilience usually gets built?

In companies under a couple hundred people, frequently yes, and it’s the opposite of how it gets sold.

The vendor version is a program with a start date, a budget, and a completion milestone. Somebody signs, work happens, and at the end you have a capability.

The version that happens is a person who decides it matters and keeps chipping at it. No single trip needs approval, so no single trip gets refused.

That approach has an obvious weakness, which is that it depends entirely on one person continuing to care. It also has one advantage nothing else offers. It survives budget cycles, because there’s nothing on the budget to cut.

Why was the site only forty miles away?

Because that was the limit of the technology.

The site sat about forty miles from the main computer room. That distance wasn’t a preference or a compromise between risk and cost. It was as far as we could go and still keep the systems copying properly.

Connectivity started as a few circuits, became a faster one, and eventually became something much faster than that. Every improvement changed what was possible, and bandwidth was the binding constraint on how good the copy could be.

What’s worth noticing is what a hard constraint does to a decision. Forty miles was argued about, measured, and understood by everyone involved, because we’d no choice about it.

Today distance is nearly free. A company can put a copy of its data anywhere, and because the tradeoff disappeared, almost nobody reasons about it. When a constraint is forced, people think about it. When it vanishes, they stop.

What does a disaster recovery site need besides computers?

Somewhere for people to sit and work.

Ours had a duplicate machine running, receiving data, building logs. That part gets all the attention and it’s the part everyone plans.

It also had laptops, computers, and a full twenty-desk setup, ready to go.

That was deliberate. A disaster that takes out a computer room usually takes out the building the computer room is in, and the people who run the business need somewhere to be. A recovery plan that restores the systems and leaves forty people with nowhere to sit has solved half the problem.

What is the difference between a tested system and a proven one?

We tested that site. We never failed over to it live.

The testing was real. Systems were exercised, the copy was verified, the equipment was checked. That work wasn’t theater and it caught things.

What never happened was the business running on it. No day where the main site was gone and the company operated from forty miles away.

My word for that is fortunately, and I mean it.

Testing establishes that the parts work under known conditions. A live cutover establishes that the whole thing carries real load, real users, and real decisions made under pressure by people who have never done it before. Only the second one is proof, and almost nobody has it.

Should you run a full failover test?

Most companies should and most won’t, and the reasons are legitimate.

A genuine failover test means taking the business down deliberately. It means accepting that something might not come back, and it means somebody has to authorize a self-inflicted outage on a normal Tuesday.

The version that works for a smaller company is narrower. Fail over one system and not everything. Do it during a quiet period. Have the people who would run the recovery do it, instead of the person who designed it.

That last detail matters more than the scale. A recovery executed by its designer proves the design. A recovery executed by whoever is on duty proves the documentation, and the documentation is what you’ll have.

What should you tell your board about recovery readiness?

The truth about which kind of evidence you hold.

We have a recovery site and we test it regularly is accurate and honest. We’re covered implies something stronger, and the gap between those two statements is where companies discover their plan was a belief.

Saying it plainly costs nothing and it changes the conversation. A board told that the parts are tested and the whole has never been exercised will usually ask what a full test would take. That’s exactly the question you want asked.

Every company with a backup strategy sits somewhere on that line. Knowing where’s most of the value.

Frequently Asked Questions

How far away should a recovery site be?
Far enough that one event cannot affect both locations, considering power grids, flood zones, and weather systems and not distance alone. Network performance used to cap this and rarely does now, so the limit is your recovery objectives instead of the technology.
How often should disaster recovery be tested?
At least annually for the full plan and more often for individual components. Test after any significant change, since a plan validated against last year’s systems describes a company that no longer exists.
Can a small company afford a recovery site?
Most cannot afford a second building and most don’t need one. Hosted infrastructure in a different region covers the systems, and the remaining question is where people would work, which can be answered with an agreement instead of a lease.
Who should run a recovery test?
Whoever would run the real one, which is rarely the person who designed the plan. A test executed by the designer validates the design. A test executed by the on-duty person validates the documentation, and the documentation is what exists at three in the morning.
What gets forgotten in recovery planning?
Where people will physically work, how they’ll reach each other when the directory is on a failed system, who has authority to declare a disaster, and what happens if the event runs past a single day.
How do you fund resilience work with no budget?
In pieces small enough that no single one requires approval, using capacity you already have. The approach is slow and it survives budget cuts, because there’s no line item to remove.

📝 Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

One comment

Was this useful?

Leave a comment