We Lifted a Running Computer Off the Floor

This entry is part 16 of 21 in the series The Operations Room
TL;DR: A hard constraint will hand you a solution you would never have picked from a list of options, and the solution is often better than what you would have chosen freely. We had a floor tile buckling under a live cabinet, no spare capacity to move the workloads, and a rule that nothing could go down. So we lifted the cabinet while it was running.

Write down which constraints are fixed before anyone starts proposing solutions.

Why do constraints produce better solutions?

Because a team with room to maneuver picks the obvious answer, and the obvious answer is obvious to everybody.

Give a group unlimited time and budget and they will schedule the outage, move the systems, do the work, and move them back. That is a fine plan and it requires no thought.

Remove the outage from the table and something different happens. The obvious answer is gone, so the team has to look at the actual problem rather than at the standard procedure for problems of that shape.

Most good engineering I have watched came out of somebody being told they could not do the normal thing.

What happens when a data center floor starts to fail?

The consultant running our computer room came to me one day with a problem.

A raised floor tile under the main cabinet was buckling. Not cracked, not worn. Buckling, which meant it could give way at any moment.

Above that tile sat the cabinet holding eight machines running the business.

This had to be fixed immediately, so we took the next weekend. A week to plan, a weekend to execute.

How do you replace a floor tile under live equipment?

The obvious approach failed before we started.

Move the workloads elsewhere, take the cabinet down, replace the tile, bring it back. That is what anyone would propose and it is what I would have proposed.

We did not have the spare capacity to relocate eight machines. We moved a few and could not move the rest.

Then my boss added the constraint that settled everything. Zero downtime. The business would not stop for this.

Now the standard answer was unavailable, and so was the fallback. What remained was one option nobody would pick from a menu.

Lift the cabinet while it was live and running. Replace the tile underneath it. Set it back down.

What does an operation like that look like?

I walked into the computer room and there was an enormous inflatable bladder under the cabinet, with bracing everywhere.

It looked like a nightmare. It was also clearly not going to fall over, which is the impression good rigging gives you.

They lifted it. They replaced the tile. They set it back down. It went very well, and it was expensive.

Nothing about the operation was improvised. The firm was US Technical Services, and the people doing the work were ex-Marines and ex-special forces. They had clipboards. They had a written plan and they followed it exactly.

A dangerous operation was boring to watch, which is what competence looks like from the outside.

Why does written procedure matter more than skill?

Skill fails quietly under pressure and procedure does not.

The people lifting that cabinet were highly capable, and that is not why it worked. It worked because somebody had written down the sequence, everyone knew their part, and nobody had to make a judgment call while several tons of equipment sat on an air bladder.

The failure mode in this kind of work is never the risky operation. It is the risky operation performed by people improvising, where two competent people each make a reasonable decision that contradicts the other.

Most companies have never written down what happens during their equivalent moment. They find out how well they improvise while the thing is already buckling.

What should a small company write down before it needs to?

The three or four moments where being wrong is expensive.

Not a binder. A page each for the things that would hurt. Restoring from backup. Failing over to a second site. Bringing the business up after an extended outage. Handling the discovery that something physical is about to fail.

Each page says who does what, in what order, and who decides. Nothing else goes in, and writing one takes an afternoon.

The value is not the paper. The value is that writing it forces the decisions in advance, when there is time to think, rather than at the moment when everyone is standing around looking at a problem.

How do you decide what your boss should see?

My boss wanted to watch the lift. I told him no.

He asked why, and I told him the truth, which was that watching it would probably upset him. He accepted that and stayed out of the computer room.

I did not hide the operation from him. He knew what we were doing, he had set the constraint that shaped it, and he had approved the cost. What I declined was the viewing.

There is a version of managing upward that means controlling what the person above you knows, and that version ends badly. This was different. He had every fact and I made a judgment about what would help him and what would only make him anxious about a decision already made.

He trusted the judgment. That is worth more than any report, and it only exists because of everything that came before it.

The Guides That Get Your Book Written, Published, and Sold

Four short, practical guides on writing, publishing, and selling your book, plus the occasional note when there's something worth your time. No fluff, no daily inbox clutter. Drop your email and they're yours.

We use MailerLite to manage our list and send these emails. Your address is used only to send you what you signed up for. We will not sell it, share it, or use it for anything else, and you can unsubscribe anytime.

Frequently Asked Questions

Should zero downtime ever be a hard requirement?
Only when the cost of an outage exceeds the cost of working around it, because the workaround is usually expensive and sometimes risky. Stated casually, zero downtime removes safe options and forces exotic ones.
How do you plan work that cannot be rehearsed?
Write the sequence, assign every step to a named person, define who calls a stop, and agree the abort conditions before starting. Rehearsal is unavailable for one-time physical work, and a written sequence covers most of what rehearsal would have caught.
What belongs in an operations runbook?
Who does what, in what order, who decides, and what stops the work. Skip the explanation of why the system exists. The document is for somebody executing under pressure, not for somebody learning.
Is it worth paying more for a vendor with military discipline?
For high-consequence physical work, frequently yes. What you are buying is procedure followed exactly rather than technical knowledge, and the price difference is small against the cost of an improvised mistake.
How much notice does failing infrastructure give you?
Sometimes weeks and sometimes none. Physical warnings like a buckling floor or an unusual noise deserve immediate attention, because the visible symptom usually appears late in the process rather than early.
When should you keep an executive away from an operation?
When their presence adds anxiety and no information, and only when they already have every relevant fact. Withholding facts is a different thing and it damages the relationship that makes the judgment possible.

📁︎ Technology

🏷︎ Infrastructure🏷︎ IT Operations🏷︎ Risk Management

📝 Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

0 comments

No comments yet. Yours can be the first.

Was this useful?

Leave a comment