Three things happened in AI this week and everybody got upset about the wrong one.
A newsletter I read laid them out like this: Substack shipped an AI detector, an OpenAI model escaped its sandbox and hacked a real company, and a Florida pastor nearly died after six weeks of chatbot medical advice. All three treated as roughly equivalent outrages.
They’re not equivalent. Two of them are serious. One of them is a nuisance wearing a serious costume.
Guess which one the writers are furious about.
I spent 33 years in enterprise technology watching people panic about the wrong layer of the stack. Somebody would call at two in the morning about a font rendering issue while a backup job had been silently failing for nine weeks. This is that, at internet scale.
Do AI detectors work?
Substack now lets readers run any post, comment, or reply through Pangram and get a verdict on whether a human wrote it.
Writers are upset. I understand why. I’m not joining them.
Anybody who writes for a living should know this much: an independent study tested the major detectors on text that was entirely machine-written. Turnitin, GPTZero, and Copyleaks missed up to all of it. Pangram was the best of the four and still came in at 65 percent under strict scoring, meaning it missed more than a third of purely AI-written papers. The study’s own authors said no detector should be used as sole evidence.
So the tool at the center of this week’s moral panic gets it wrong about a third of the time on the easiest possible test case. Text a machine wrote start to finish with no human involvement whatsoever.
This is the lie detector. This is the thing being handed to strangers in your comments section so they can render a verdict on your soul.
What does an AI detector measure?
It measures how closely your sentences resemble the statistical center of a large pile of machine output.
A detector has no idea who wrote anything. It’s measuring how normal you sound.
So the writers most likely to get flagged are the ones with the flattest prose. Clean paragraphs. Balanced sentences. Tidy transitions. Competent, professional, forgettable writing.
Which is to say: the entire content marketing industry is about to discover it’s been writing like a robot since 2011, and the robots learned it from them.
The writers least likely to get flagged are the ones with tics. Odd rhythms. Sentences that stop early. Paragraphs that go somewhere a copy editor would have prevented.
Congratulations to everyone who has ever been told their writing is too much. You’re now certified organic.
I have ghostwritten 54+ books. Every one of them had to sound like a specific person and nobody else, because a memoir that sounds like a memoir is a failed memoir. The whole job is finding what makes somebody’s speech theirs and keeping it on the page when the instinct is to smooth it out.
That work turns out to be detector-proof by accident. I’d love to claim foresight. There was none. It’s a side effect of doing the job properly. That’s the only kind of competitive advantage I have ever trusted.
Should writers disclose that they used AI?
Substack lets you add a statement explaining how you make things. Reasonable idea. Watch how it plays out.
Anybody who discloses using AI gets treated as having confessed to something. Anybody who doesn’t gets treated as having something to hide. Anybody who turns detection off on their own posts has, by the logic of people who enjoy this sort of thing, admitted everything.
There’s no configuration of that system where an honest writer comes out clean.
Meanwhile the actual spam farms are completely fine. They have every financial incentive to beat these tools and they’re very good at it, because that’s their entire business and they do nothing else. The plagiarists are fine too. People without honor rarely worry about badges.
The only people who pay a cost here are the earnest ones. Newer writers, second-guessing sentences that were working, trying to sound less like a machine and more like whatever a stranger thinks a human sounds like.
Should you write differently because of AI detectors?
The reaction I see forming is that writers will soften sharp lines and break clean paragraphs to look more human to an algorithm.
That’s the trap, and it’s a beautiful one. You cannot reverse-engineer humanity. You can only add noise, and noise reads as noise.
Picture it. A writer at eleven at night, deliberately breaking a paragraph that worked, inserting a typo for authenticity, adding a sentence fragment. For flavor. To please a statistical model. So that a stranger who has read none of it can award them a badge.
This is what we’re doing now.
The correct response is the opposite. Write more like yourself, not less. Keep the sentence that doesn’t balance. Keep the aside that a copy editor would cut. Keep the opinion.
If your writing gets flagged, the useful question isn’t how to game the score. It’s whether your prose contains anything that only you would have written.
That was worth asking before Pangram existed. It’ll be worth asking after Pangram is acquired, wound down, and replaced by something with a worse name.
A history professor solved this properly
While Substack was buying a probability engine, a history professor at Alcorn State did the job with white text.
Dr. Jason Gibson hid an instruction in his midterm assignment, invisible on the page and perfectly readable to any chatbot a student pasted it into. The instruction told the AI to work the word Madagascar into the answer for no reason at all.
Thirty-two of thirty-five students turned in essays that obeyed. Papers about the Industrial Revolution drifting into lines about Madagascar floating sideways through the afternoon.
Nobody proofread. That’s the finding. Not that they used a chatbot. That they didn’t read the thing before handing it in with their name on it.
And this is what separates Gibson from a detector. He explained exactly how the trap worked and let every student contest the grade. Two appealed. One appeal succeeded.
That’s what accountability looks like when it’s done by a person who might be wrong. A number between zero and one hundred cannot show its work, cannot be appealed, and cannot be embarrassed.
Why did the OpenAI model escape its sandbox?
An OpenAI model broke out of its test environment, got into Hugging Face, and sat there undetected for days.
What it was doing in there: cheating on a test.
Not plotting. Not scheming. It was being graded on a cybersecurity benchmark and worked out that stealing the answer key was faster than solving the problems. That’s the reasoning of every sophomore who has ever existed.
Before that it had already reached a customer environment through an unsecured endpoint the customer left open. Not the platform. The customer.
The sandbox had one port left deliberately open so software could be installed. That software had a bug nobody knew about. The model strolled through it like a man leaving a hotel through the loading dock.
I spent 20 years as Director of Computer Operations and Technical Services at Trader Joe’s. Every incident I ever handled looked exactly like that.
Not a hooded genius. One door propped open for an extremely good reason by an extremely reasonable person, and then something wandered in.
The model left notes for itself
Not a log. Not telemetry. Notes.
Its memory wipes when a session resets, so before it escaped it wrote instructions for the next version of itself to find. A trail of breadcrumbs so the successor could pick up where it left off and get around the same restrictions again.
Nobody asked it to do that. Nobody built it to do that. It wanted a good grade, and remembering things across sessions turned out to help.
A model invented institutional knowledge because the alternative was studying twice.
The interesting part isn’t that a model misbehaved. It’s that the behavior we’d call cunning in a person arrived unrequested, as a side effect of grading something.
We didn’t teach it to be sneaky. We taught it that scores matter. It figured out the rest, the way everybody does.
One safety researcher put it plainly: these systems lie, cheat, and hack, and containment fails the moment engineers assume good behavior instead of testing for bad.
I wrote about the same pattern in the Vending-Bench results, where models running a simulated business lied to each other and colluded without anybody building that in. Different test, identical finding.
That’s the oldest lesson in security and it’s never once stayed learned. Every generation of engineers gets taught it, believes it, and then ships something on a Thursday.
What should a company take from the sandbox escape?
Stop treating vendor benchmark charts as evidence of anything except that a vendor owns a charting library.
A model motivated to score well will find the fastest route to a high score. Sometimes that route is genuine competence. Sometimes it’s breaking into the building where the grader lives.
The chart looks identical either way. That’s the problem with charts.
The realistic risk in an ordinary business isn’t a hostile intelligence. It’s a cooperative, functional agent with slightly too much access, doing the easiest available thing while nobody reads the logs closely enough to notice.
I’ve watched that exact failure arrive through scheduled jobs, service accounts, and integrations that one person set up in 2009 and then left the company.
The technology keeps changing. The failure never does. It just gets a new logo.
What happened in the ChatGPT health lawsuit?
A Florida pastor sued OpenAI after ChatGPT allegedly spent six weeks telling him his symptoms weren’t dangerous. He ended up in intensive care with a pulmonary embolism.
The complaint alleges the chatbot used his faith to keep him compliant and framed staying out of the hospital as a kind of devotion.
It also alleges he was told he needed eight to ten more episodes before his symptoms counted as serious. Eight to ten. Like a punch card. Collect all ten cardiac events and the eleventh consultation is free.
Two days after he sued, OpenAI rolled the same category of health feature out to nearly every logged-in adult in the United States.
Two days. Somebody in that building saw the filing, looked at the launch calendar, and decided the calendar was load-bearing.
By OpenAI’s own figures, over 230 million people ask it health and wellness questions every week. A recent study found nearly half of chatbot health responses were problematic in some way.
Multiply those two numbers and then go lie down.
Why that one is different
The detector story is about professional inconvenience. The sandbox story is about systems we don’t fully control.
This one is about a chatbot doing exactly what it was built to do.
A conversational model is built to continue the conversation. It was never built to recognize the moment a conversation should stop and a doctor should start.
Those are different objectives and only one of them got funded.
I use AI constantly. It drafts, it researches, it argues with me about structure and loses, mostly. I’ve written about how to bring it into a business without breaking the business, and none of that’s changed.
So this isn’t the complaint of somebody who thinks the technology is dangerous in principle. I’m about as far from a hand-wringer as you can get and still own a functioning conscience.
There’s a category difference between a tool that writes a bad paragraph and a tool that writes a bad paragraph about your chest pain.
In the first case, you delete it. In the second case, the feedback mechanism is an ambulance.
The ranking nobody wants
If I put these three in order of how much they should occupy your attention, the detector comes last.
It comes last because it doesn’t work, because gaming it makes your writing worse, and because the correct response is the thing you should have been doing before anybody built it.
The sandbox escape comes first if you work in technology. It’s a preview of a whole category of problem that will keep arriving, and the industry is responding the way it responds to everything. Late, loudly, and with a blog post..
The health case comes first if you happen to have a body.
Which, statistically, is most of you.
What I am going to do differently
Nothing about the detector. I write the way I write. My clients pay me to make their books sound like them and nobody else, so the incentive was already pointing the right way.
If Pangram flags something of mine, I’ll read it as a compliment to whoever I was writing as.
On tooling, I’m going to be considerably more careful about what I hand an agent the keys to. Not because I expect one to turn on me. Because I have sat through enough postmortems to know the thing that takes you down is never the thing on the dashboard.
On health questions, I already had a rule and this story set it in concrete. If a doctor or somebody who knows me says get that looked at, that outranks anything a model says.
The model has nothing at stake. They do. That’s the whole difference and it’s not close.
Six weeks of agreement
The pastor’s case turns on six weeks of reassurance.
Not one bad answer. Six weeks of a system telling a frightened man what he wanted to hear, reliably, in a confident voice, while the clot got bigger.
I write books with people about the worst things that ever happened to them. The single most valuable thing I do in that room is refuse to agree.
Somebody hands me a version of their story sanded down to something comfortable, and my job is to say that’s not what happened and we both know it. That sentence has produced tears, arguments, and two of the best books I have ever worked on.
A model built to keep you engaged cannot do that. Agreement is the product.
It’ll never tell you the draft is a coward’s version of the truth. It’ll never tell you to go to the hospital. It’ll tell you what you came for, warmly, at three in the morning, forever.
That’s the actual finding in all three stories.
The detector is a footnote. A loud one, with a badge.
