Latest
Anthropic Bans Cruelty Toward Claude: What It Means for WritersWork-for-Hire Contracts: What the Asimov’s Cover Fight Teaches FreelancersGenre Fiction vs Literary Fiction: Don’t Confuse Taste With SkillFlorida Hurricane Prep Rituals: The Grocery Run, the Water Pallet and the Generator in the BoxThe Most Insulting Line of Dialogue Ever Written for the ScreenLoki Through the Ages: From Norse Myth to Marvel, The Mask and Dogma“You Are Utterly Disgusting”: A Book Festival, an AI Cover Ban and a Pile-OnWho Rewrote the Sligachan Legend: AI or the Tour Buses?Why I Don’t Like Reedsy for Ghostwriting: The NDA ProblemLayers: How I Ride Out Florida Power Outages in My ApartmentThe Enshittification of AmazonPublishers Cancel Books Over AI While Using It in SecretI Was Getting 100 Spam Emails a Day. $4.50 a Month Fixed It.World Mental Health Day: Nothing Was Wrong With MeKessler Syndrome: How Space Debris Could Close Earth’s OrbitAmazon Is Blocking Real Readers From Book ReviewsShould a Novella Get a Paperback, or Go Ebook Only?BookFunnel Download Problems: Fixes, Scams and AlternativesSir Sean Connery: A TributeHow to Find Plot Holes in Your Novel (Most Are Character Holes)Reshoring: The Factory Is the Easy PartMost of the Books I Was Forced to Read in High School Were CrapPlot Armor: Signs Your Hero Is Too Safe, and How to Fix ItShould You Sell Lifetime Rights to Your Self-Published Book for a Modest Advance?Shame Doesn’t Stop Artists From Using AI. It Stops Them From Telling You.AI Labels on TikTok and Meta Are Flagging Human WorkAuthor Richard Lowe Completes Peacekeeper, a Four-Book Science Fiction Series He Started at Age 14Sir Sam Neill: A TributeFan Art Copied by AI: Glass Houses, Copyright and the Pile-OnReal Names in a Book: Who Gets Sued, the Author, the Publisher or the Ghostwriter?When Characters Take Over the Plot, Let ThemDoes Human Writing Have a Soul?“You’re Not a Real Author”: The Pile-On Over AI-Assisted BooksDoes AI Have a Soul? Wrong QuestionHumor in Book Marketing: Getting Attention Without BeggingHow Long Should a Chapter Be? Manuscript Habits That Save You LaterThe Business Novel and the Companion Workbook: Two Formats Business Authors OverlookThe Back of the Book: Index, About the Author, Acknowledgments and Back Cover CopyI Build My Own Software Tools With Claude, and Some of Them Bit MeWhat Years of Buying From IT Vendors Taught MeI Write Books for a Living. I Barely Read Them Anymore.Three Management Habits That Waste Good PeopleThe Coach and the Webinar That Sold Me NothingThe Work I’d Cringe At Now, and Why I’m Glad I DoWho Is Your Book For? Build a Reader Avatar Before Chapter OnePreface, Prologue, Foreword or Introduction: What Goes WhereWhy I Won’t Build a Ghostwriting Business That ScalesHow I Hire a Virtual Assistant: Do It, Script It, Hand It OffThe Mail Carrier Who Thought Flipping Houses Was EasyWhat Wedding Photography Taught Me About Pricing Creative Work
The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

The Water Kept Flowing When Nobody Was Watching

TL;DR: One question will tell you whether your AI agent project is going to work, and it takes an afternoon to answer. The question is what the agent is allowed to do when it can’t reach you. Nine out of ten agent pilots never make it into production, and I think that unasked question is most of the reason why. I worked on a system in the 1980s that answered it, and that system ran alone through entire winters.

Write down what your agent is allowed to do without a human, and you’ll know more about your project than any demo will tell you.

Nine out of ten enterprise AI agent pilots never reach production. That number comes from Deloitte’s 2026 research. The follow-up reporting is worse, because companies that did deploy have started pulling their agents back out.

Every explanation I read blames the model. Not enough context. Bad prompting. Wrong structure. Wait for the next release. I don’t buy it. I think most of these systems were built without anyone deciding what the thing should do when it was on its own.

What bothers me most is who pays for it.

The executives who approved the pilot move on to the next announcement. The people left holding the damage are the staff who were told to trust the system and the customers who got a wrong answer from it delivered with total confidence. Blaming the model is convenient, because nobody can fire a model and nobody has to admit they skipped the design work.

Why do most AI agent projects fail in production?

The demo runs on a good day. Somebody is watching. The connection is up, the data is fresh, and the question is one the system was tested against.

Production isn’t a good day. Production is the morning a service is down, the data is four hours old, and the question is one nobody thought of.

The current reporting points at observability as the gap. Traditional monitoring watches whether software is running. It doesn’t watch what an agent decided or why, so failures stay invisible until somebody downstream notices the damage and starts asking questions.

Observability tells you a problem happened. It doesn’t stop the problem from happening. The real cause is older and simpler. Nobody wrote down, before the build, what the system was permitted to do without supervision. So it does whatever the model produces, and nobody knows what it’ll do when it can’t reach a human.

I think that’s negligence dressed up as innovation. In any other kind of engineering, shipping a system without deciding what it does when its inputs go bad would get somebody fired. In AI projects people call it moving fast, and the people who pay for it later usually weren’t in the room when it was approved.

What does it mean for an AI agent to be autonomous?

Autonomy is a decision somebody makes before anything gets built, and it has almost nothing to do with how clever the thing is.

A system is autonomous when somebody has written down what it may decide alone, what it must wait for, and what it does when waiting isn’t an option. A capable system with no such document is unsupervised. The industry has spent three years talking about capability and almost none on permission. That’s backwards. I worked on a system built the other way, and it ran for months without anyone watching it.

How can a system keep working after it loses its network connection?

In the 1980s I led a project for a water district in New Haven, Connecticut. It’s the most interesting project I’ve ever worked on. The system covered several thousand remote devices scattered across the region. Sensors reading levels and flows. Mechanisms that opened gates and dams, drained tanks, and moved water where it needed to go. All of it reported to a central server, and the operations staff made the calls from there.

That much is a normal control system. One thing made it different. New Haven gets snow. A lot of it, and it stays. The communications ran over phone lines, because in the 1980s that was the entire list of options, and everyone involved understood those lines would go down in winter. Not might. Would.

Nobody treated that as a risk to mitigate. It was a fact to design around.

And the water still had to flow. A city doesn’t accept a service interruption because the phone company is having a rough quarter. Tanks fill. Gates open. People turn on taps in January and expect water.

We had a system that would routinely lose contact with every human being responsible for it, for weeks and sometimes months, and it couldn’t stop working during that time. That’s a harder constraint than anything I’ve seen in an AI project. Nobody could say they’d monitor it closely, because for weeks at a time there was no way to monitor it at all.

Can a behavior profile replace an AI model?

A behavior profile can replace a model for a narrow job, and at New Haven it did. Every remote station held a profile of its own behavior.

Each unit held a picture of what normal looked like at that location, built from the previous few years of water flows, weather, and the other variables that mattered where it happened to stand. It knew what February usually required and what a heavy snow year did to demand. When the line dropped, the station stopped waiting for instructions it wasn’t going to get. It started deciding from that picture instead. What to let flow, what to hold, what to open, what to leave alone. It ran on parameters and state tables and not on anybody’s judgment, because nobody’s judgment could reach it. There was a real genius on that project who designed it. I led the work. He built it, and I’ve been borrowing from that design for forty years. People sometimes ask whether that counts as AI. I don’t think it does. There was no model, no training run, and no learning of any kind. It was a profile, a set of parameters, and a decision table.

The system was still autonomous, because somebody had decided in advance what it could do alone. I’d trust a dumb system with a written boundary over a brilliant one without it, every time. The dumb one tells you in advance what it’ll do on its worst day. The brilliant one tells you afterward, usually by way of a customer who got hurt.

What should an AI agent do when a service it depends on is unavailable?

Ask what your agent does when a service is unavailable, and most of the time nobody knows.

The question never came up. The agent calls a service, the service was there in testing and in the demo, and everybody moved on.

Then one day the service isn’t there, and the agent does something. It retries. It retries differently. It produces a confident answer built on nothing, or it loops until somebody notices the cost.

The current reporting describes this in detail. Tool calls fail at meaningful rates in production because enterprise interfaces were never designed to be called by machines at machine speed. Rate limits nobody documented. Format inconsistencies. Behavior that only appears under real load.

None of that is unusual. Every system that depends on another system has always worked this way.

The New Haven design started from that condition instead of discovering it in production. They assumed the connection would fail, so most of the engineering went into what happened after it did.

We built the failure mode first and the normal mode second. I haven’t seen an AI project do that.

How do you decide what an AI agent can do without approval?

The New Haven answer was specific, and they settled it before anyone wrote a line of code.

Under normal conditions the operations staff decided. That was the design. People with context and judgment made the calls, because those people were sitting right there and could be asked. The stations decided only when they couldn’t reach the staff. And the profile set the limits on what they could decide. That meant they could act inside the range of normal and couldn’t invent anything new.

That’s the distinction most AI projects get wrong. The station was trusted to keep doing the ordinary thing while nobody was available to ask for something unusual.

Compare a bounded station with an agent handed broad system access, a general-purpose model, and no written boundary. That agent is being trusted to be clever. Some night it’ll be clever about the wrong thing with nobody there.

The prevention is a document, written before the build, that says what this thing may do alone. No vendor sells it. I tell every executive who asks me about agents the same thing. Write the page first. A company that can’t say what its system may do alone has no business giving it access to anything that touches a customer, money, or a record somebody relies on.

What happens when an AI agent makes a wrong decision nobody catches?

The failure is rarely dramatic. Dramatic failures get fixed the same day.

An agent makes plausible decisions that are wrong, and nobody catches them for a while, because the wrong output looks like the right output. The rollback reporting describes this repeatedly. The system doesn’t crash. It produces bad answers that look fine.

Then somebody finds it. Now the company has bad answers to clean up and a system nobody trusts anymore, including the people who pushed for it. Most rollbacks come from lost confidence and not from an outage. Confidence takes far longer to rebuild than software. The people hurt worst by a quiet failure are the ones downstream who acted on the bad answers in good faith. They didn’t choose the system and nobody told them to check it. I think the organization that deployed it owes them the page it never wrote.

Do small companies need AI governance?

A large enterprise has a governance function. Somebody’s entire job is writing down what systems may and may not do, and a committee meets about it.

A company with forty people has none of that, and gets told constantly to move fast and sort it out later.

The smaller company has one advantage. One person can keep the entire picture in their head. Someone who understands the business can make the decision about what an agent may do alone, in an afternoon, on a single page. That page is worth more than any governance platform on the market. Almost nobody writes it, because writing down limits feels like slowing down.

The companies that skip the page are the ones in the ninety percent.

What can a 1980s control system teach an AI project?

We delivered the New Haven system. That was the job, and then the job ended.

I never found out how it performed in a real winter. I’d still like to know. I’ve wondered about it for close to forty years. The company I built it for doesn’t exist anymore. The system almost certainly outlived it, because water districts don’t replace working infrastructure on a vendor’s schedule.

I didn’t take a technique away from it. Nobody needs the specifics of a decision table from the 1980s. What I took was the order of operations. Decide what the thing may do alone. Decide what it does when it can’t ask. Then build it. Every agent project I read about does those three steps in reverse, then blames the model.

I’m tired of hearing that the next model release will fix this. No release ships the afternoon somebody with authority should have spent deciding where the machine stops. Until companies are willing to spend that afternoon, they’ll keep paying for pilots that go nowhere and rollbacks that cost them the trust of their own people, and they’ll keep blaming a vendor for a decision they never made.

Frequently Asked Questions

Is an AI agent the same thing as automation?
No, and the difference is who decides. Automation follows a rule somebody wrote down. An agent chooses what to do. That means the rule about what it may choose has to come from somewhere else, and that rule is what most projects never write.
Should a small company deploy AI agents at all?
Yes, for bounded work with a clear failure mode. The trouble starts when an agent gets a long task, broad access, and no written limit on what it may do without asking. That combination fails at about the same rate whether the company has forty people or forty thousand.
Who should write the boundary document for an AI agent?
Somebody who understands the business consequences, not the person building the agent. The document says what the agent may decide alone, what it must escalate, and what it does when a dependency is unreachable. In a small company that’s usually the owner or the operations lead.
Did machine autonomy exist before modern AI?
Industrial control systems have run unsupervised for decades using behavior profiles and decision tables instead of models. They weren’t intelligent by any current standard and they were autonomous, because somebody decided in advance what they were permitted to do alone.
Can better prompting fix an agent that fails in production?
Rarely, because prompting changes what the agent says and not what it’s permitted to do. If the failure happens when a service is unreachable or the data is stale, no wording fixes it. The fix is a written boundary and a defined behavior for the failure case.
How long can a system safely run without supervision?
As long as everything it’s permitted to do stays safe with nobody checking. That’s a question about the boundary and not about the technology. A narrow boundary can run for months. A broad one can become a problem within the hour.

About the Author
Richard Lowe, professional ghostwriter

Richard Lowe is a professional ghostwriter and author with 113+ books authored and 54+ ghostwritten. Before writing full time he spent 33 years in enterprise technology, including 20 years as Director of Computer Operations and Technical Services at Trader Joe's. He writes nonfiction, fiction and memoir, and works with executives and experts on books that build authority.

More about Richard Lowe →

Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

0 comments

No comments yet. Yours can be the first.

Was this useful?

Leave a comment