This entry is part 46 of 53 in the series Technology
TL;DR:Cleaning infected machines one at a time inside a connected network is a plan to clean the same machines forever. We lived that loop: disinfect machine A, and while cleaning machine B, machine C reinfects A. The cycle only broke when we stopped treating machines and started treating the network, isolating the workstation population and separating clean systems from dirty ones.
One infection at the retailer taught me more about incident response than any course ever did, because it refused to die. It was a network-propagating piece of malware, and our first response was the obvious one: find an infected machine, clean it, move to the next.
The malware had a different plan. Clean machine A. While you’re cleaning machine B, machine C reinfects machine A. Clean C, and by then the infection has looped back through D and E and is sitting on A and B again. We were playing whack-a-mole against an opponent that never got tired, never went home, and worked every machine at the same time while we worked one at a time.
The arithmetic of losing
The failure was mathematical. Our cleaning rate was serial: one technician, one machine, some number of minutes. The malware’s infection rate was parallel: every infected machine was a transmitter, all the time. When the propagation rate exceeds the cleaning rate, the infected population grows while you work. You can staff up, and we did, but you’re racing an exponential with a linear tool.
Every IT team that’s fought a network worm knows this feeling, and most learn the answer the same way we did: the hard way.
We tried the obvious countermove first: more hands. Pull technicians from other work, clean faster. It helped the way bailing helps a holed boat, meaning it changed the numbers without changing the outcome. Parallel infection beats serial cleaning at any staffing level you can afford, because the malware recruits every machine it captures while each technician remains exactly one technician. The staffing response fails mathematically before it fails operationally, and recognizing that early saves you the week we spent proving it.
How do you stop a malware reinfection loop?
The cycle broke when we changed the unit of response. Instead of cleaning machines inside a hostile network, we isolated the workstation network itself, then split the population: a clean segment for disinfected machines, a dirty segment for everything else. Machines moved one direction only, dirty to clean, and only after verification. The transmitters could no longer reach the cured.
Once separation was in place, the math flipped. Every cleaned machine was permanently subtracted from the infected population instead of temporarily. The dirty segment shrank monotonically to zero, and the incident, which had run for a long, demoralizing stretch, ended in an orderly march.
Never fight an infection on ground the infection controls. Take the ground first.
Containment precedes eradication. That ordering is now standard incident-response doctrine, but doctrine reads like a slide until you’ve watched a cleaned machine light up again ninety seconds after you left the desk. The lived version of the rule is simpler: never fight an infection on ground the infection controls. Take the ground first.
The same ordering governs modern incidents. Ransomware response starts with isolating segments, not restoring files. Compromise response starts with cutting attacker access, not resetting passwords one account at a time while the attacker holds the mailbox that receives the resets. Every generation of defenders rediscovers the ordering, usually mid-incident.
When executives ask me what makes a security war story worth publishing, this is my example. Not because the malware was exotic. Because the failure mode, serial defense against a parallel attacker, is universal, and a reader who absorbs it’ll recognize it in incidents that haven’t happened yet.
For more from this series, see The Cybersecurity Hub: breaches, audits, and hard-won security lessons from four decades in the trenches.
Frequently Asked Questions
Why do cleaned machines get reinfected?
I learned this firsthand when we cleaned an infected machine only to watch it get reinfected minutes later. The problem is that a cleaned machine gets returned to a network that still has other infected machines actively spreading the malware. Every infected system is transmitting all the time, so while you clean one machine, another can reinfect the ones you already fixed. Our cleaning rate was serial, one technician working one machine at a time, but the infection rate was parallel, with every infected machine spreading it at the same time. That mismatch meant the infected population kept growing even as we worked.
What is the right order for malware incident response?
Containment first, eradication second. Isolate the affected network segment, separate verified-clean systems from infected ones, then clean the remaining population. Cleaning before containing means cleaning repeatedly.
How do you contain a network worm outbreak?
In our case, the outbreak only stopped once we quit cleaning individual machines and isolated the entire workstation network instead. We split the population into two segments, a clean one for disinfected machines and a dirty one for everything still infected, and only let machines move from dirty to clean after verification. Because the segments were separated, the infected machines could no longer reach the ones we’d already cured. That changed the math completely, since every cleaned machine was now permanently removed from the infected population instead of being at risk of reinfection. The dirty segment shrank steadily until it reached zero, and an incident that had dragged on for a long time finally ended in an orderly way.
Richard Lowe is a professional ghostwriter and author with 113+ books authored and 54+ ghostwritten. Before writing full time he spent 33 years in enterprise technology, including 20 years as Director of Computer Operations and Technical Services at Trader Joe's. He writes nonfiction, fiction and memoir, and works with executives and experts on books that build authority.
The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.