When a digital transformation fails, the post-mortem usually blames strategy or vision or leadership. In my experience, the real cause is almost always something far more boring. The testing wasn’t good enough.
I led transformations at a national retailer for two decades, and now I ghostwrite books for the leaders who run them. The failure point was rarely the grand plan. It was the unsexy work of testing that nobody wanted to fund.
Here’s the pattern I saw over and over. We’d build a new system. The programmers tested their pieces and did a pretty good job. We tested our pieces and did a pretty good job. Everything worked in testing. Then we went live, thousands of real people logged on, and the system fell over.
The culprit was usually load. We’d Oracle queries that weren’t tuned, and they were fine with a handful of test users. Under real load, with the full weight of everyone using the system at once, those same queries crawled.
Here’s why it happens and why it’s so hard to catch. A database query that scans a little data runs instantly when ten people use it. The same query against the full production dataset, hit by thousands of users at once, can grind to a halt. The behavior isn’t linear. A system that’s perfectly snappy with test traffic can collapse completely under real traffic, and there’s often no warning until you cross the threshold.
We tuned those queries after go-live instead of before. There was no good way to replicate the real load of a big system with a lot of people on it.
Your new system works perfectly in testing. Then a thousand real users log on at once and it falls over. That’s load testing, and almost nobody does enough of it.Share on X
That’s the lesson I’d press on anyone running a transformation. Do far more load testing than you think you need. Simulate the real volume, the real number of concurrent users, the real size of the production data, before you go live. A system that passes every functional test can still die the moment real volume hits it. Finding that out in production, with the entire company watching the new system fail on day one, is the worst possible time and the worst possible way.
This connects directly to the component-by-component approach I describe in digital transformation is plumbing: when you move one piece at a time, each piece meets real load on its own. A load problem shows up small and contained instead of taking down everything at once.
The other weak spot was failure-mode testing. We tested how the system worked when everything went right. We rarely tested what happened when something went wrong. What breaks when a server drops? When a connection times out? When a piece of data is malformed, or a disk fills up, or a dependency the system relies on is suddenly unavailable? Those paths went largely untested, because testing the normal path felt like enough, and the normal path is what everyone naturally thinks to check.
It’s not enough. Systems spend plenty of time not in the normal state. Things fail constantly in real operation. Hardware dies. Networks hiccup. Data arrives malformed. Those are exactly the moments you need the system to fail gracefully instead of catastrophically. A system that handles the happy path beautifully, then corrupts data or crashes hard the first time it meets a malformed record, was never finished.
Testing only the happy path is testing for a world that doesn’t exist. The same lesson shows up in security, where, as I describe in the boring security work that keeps you safe, the backup you never tested is the one that fails when you finally need it. Untested failure handling is the same trap in a different costume.
Testing is the technical failure point. People are the human one, and their resistance was more predictable than you might think. It tracked almost perfectly with one thing: whether the person believed the change threatened their job.
Someone who thought the new system put their position at risk resisted hard. They dragged their feet, found problems, declined to learn the new way, because every step toward the new system felt like a step toward their own replacement. Someone who saw that we were training them, investing in them, keeping them employed, welcomed the change and often became its strongest advocate. Same transformation, opposite reactions, and the difference was entirely about what the person thought it meant for them, not about the technology at all.
People don’t resist change. They resist losing their job. Show them you’re training them, not replacing them, and the resistance mostly disappears.Share on X
That taught me that managing the human side of transformation is mostly about addressing the fear honestly. People are protecting themselves. That’s rational. Show them they’re safe, that the change comes with training and a place for them in the new world, and most of the resistance evaporates. This is exactly why the order of operations matters so much, why you start with people before you touch the technology, which I lay out fully in people, process, technology.
Which failures kill digital transformations?
Weak load testing, no failure-mode testing, and unaddressed fear among the people whose jobs are changing. All three are preventable, and all three get skipped because they’re tedious.
From Conversations With Influencers
Kader Sakkaria and Imran come at transformation cloud-first and API-first, defining customer experience by the channel the customer arrives through. They laid it out on Conversations With Influencers.
None of these failure modes are mysterious. Weak load testing. No failure-mode testing. Unaddressed job fear. All three are preventable with effort nobody wants to spend, because it’s boring and it delays the exciting launch. That’s the pattern across every failed transformation I’ve seen. The work that would have prevented the failure was known and available. It got skipped because it was tedious and slowed things down.
When I ghostwrite a transformation book, the failures are often the most valuable material, because the executive learned more from what broke than from what worked. The accurate account of why a system fell over, and what they’d test differently next time, is worth more to a reader than any success story. You can see how I work with technology leaders on the technology ghostwriting page.
Frequently Asked Questions
Usually poor testing, not bad strategy. Teams test whether a system works normally and skip two things. Load testing. That shows how it behaves under real volume. Failure-mode testing. That shows what happens when things break. Systems that pass every functional test still die when thousands of real users hit them at once.
It’s testing how a system behaves under realistic volume, not just a handful of test users. Database performance isn’t linear: a query that’s instant for ten users can collapse under thousands hitting full production data. If you don’t load test thoroughly with real volume, you discover the problem in production on day one, the worst possible time.
Testing what happens when something goes wrong, not just when everything works. What breaks when a server drops, a connection times out, a disk fills, or data arrives malformed? These paths often go untested because the normal path feels like enough. It’s not, because systems fail constantly in real operation and need to fail gracefully.
Almost always because they think it threatens their job. Resistance tracks that fear precisely. Someone who believes the change puts their position at risk fights it; someone who sees they’re being trained and kept on welcomes it and often champions it. Address the fear honestly, with training and a place in the new world, and most resistance disappears.
I’ve watched digital transformations fail for the same three preventable reasons over and over. The first is weak load testing, where a system passes every test with a handful of users but falls over the moment thousands of real people log on, because untuned database queries only reveal their problems under real volume. The second is skipping failure-mode testing, so nobody checks what happens when a server drops, a connection times out, or data arrives malformed, and the system that looked finished actually never was.
The third is unaddressed fear, since resistance to change tracks almost perfectly with whether someone believes the new system threatens their job, and that fear mostly disappears once you show people they’re being trained and kept, not replaced. Moving one component at a time also helps a great deal, because each piece meets real load and real failure conditions on its own instead of everything collapsing together on launch day.
Yes. I led transformations for two decades and ghostwrote three on the subject. The failures are often the most valuable material, because leaders learn more from what broke than from what worked. You can see how I work on the technology ghostwriting page.
