Someone hands you a number and says your book was written by a machine. Do you know what that number measures?
Start with something that does measure. A breathalyzer. You blow into it, the chemistry happens, and the figure that comes out corresponds to something physically present in your body. It can be calibrated. It can be challenged in court. When it says 0.09, there’s a molecule count behind it.
An AI detector measures nothing like that. It reads your sentences and estimates how closely they resemble sentences a language model would produce. There’s no residue in the paper. There’s no molecule to count. The output looks like a breathalyzer reading and people treat it like one. That mistake is costing writers work right now.
I want to explain what these tools do, why the number is weaker evidence than it appears, and what to ask for instead when you genuinely need to know who wrote something.
How do AI content detectors decide a text is machine written?
The detector is itself a model. It was trained on large piles of text labeled human and large piles labeled machine, and it learned the statistical fingerprint of each. Then it reads your paragraph and reports how closely it matches the machine pile.
What does that fingerprint consist of? Predictability, mostly. Language models choose likely words. Across a few hundred words, that produces prose with fewer surprises than human writing carries, sentences of similar length, transitions that arrive on schedule, paragraph shapes that repeat. Researchers use the terms perplexity and burstiness for these measures. Low surprise and even rhythm read as machine.
Now consider what else produces low surprise and even rhythm. Professional editing does. A technical writer trained to keep sentences under twenty words does. A non-native English speaker working from a smaller pool of constructions does. A lawyer does. A writer who has been through six rounds of revision aimed at clarity does.
The detector isn’t looking for a machine. It’s looking for regularity. Machines produce regularity. So does discipline.
What does the detector score claim to prove?
The newer tools do more than return a single percentage. The better ones grade into bands, fully machine written, moderately assisted, lightly assisted, human. One of them is now embedded in a major publishing platform. A reader-facing AI assessment now sits on writing that was never submitted for judgment at all. Turning it off is itself read as a confession.
The bands are an improvement, and I would rather see them than a bare percentage. They still rest on the same foundation. A guess about style.
Meanwhile the failure reports keep arriving. Authors paste in paragraphs from their own published books, work that predates these tools, and watch a detector call it ninety percent machine. Writing associations have run comparative tests on multiple detection products and found the accuracy claims did not survive contact with real manuscripts. False positives land hardest on exactly the writers who can least afford them: students, non-native speakers, and anyone whose style runs clean.
A tool that’s right most of the time is useful to you for triage. It isn’t evidence, and I have written a whole piece on why publishing keeps treating it as though it were.
Why can’t a detector tell the difference between edited prose and generated prose?
Because there may be no difference to find.
Think about what good editing does to a sentence. It removes the wandering clause. It cuts the redundant modifier. It makes the rhythm consistent so the reader stops noticing the prose and starts noticing the argument. Every one of those moves reduces surprise. Surprise is the exact quantity the detector treats as suspicious.
A heavily edited human manuscript and a generated manuscript can converge on similar statistical profiles from opposite directions. One got there by removing noise on purpose. The other never had any. The detector sees the destination and cannot see the road.
This is also why detectors do worse on short passages. Give a model two hundred words of your chapter and the signal is thin, so the confidence swings wildly. Feed it a whole chapter and it steadies. That’s one reason the tools look better in vendor demonstrations than in the wild, where people paste in a paragraph and demand an answer.
And there is a moving target problem underneath all of it. Detectors learn the fingerprint of today’s models. The models change every few months. A detector tuned on last year’s output is grading this year’s writing with an outdated map, and nobody tells you which map is grading your pages.
The Voight-Kampff machine is a test that claims to detect the non-human by measuring involuntary responses. It takes over a hundred questions, it is administered by an expert, and the film’s entire tension comes from the possibility that it is wrong about somebody. Our version returns a percentage in four seconds and people treat it as settled.
What should you ask for instead of a detector score?
Here is where I put my own work on the table, because I am asking you to trust a different kind of evidence and you should know what it looks like.
When I ghostwrite a book, the process leaves a trail whether anyone plans to inspect it or not. There are recorded interviews, hours of the client talking in their own voice. There are transcripts of those recordings. There is an outline the client read and approved before a word of prose existed. There are chapters delivered one at a time, each with the client’s comments on them, and the next version showing those comments applied. There are dated files going back to the first conversation.
That trail can’t be produced by a prompt. It isn’t a defense I invented for the AI era. It is what doing the work properly looks like, and it has existed in every project I have run.
So if you are an author with doubts about a manuscript you paid for, don’t run it through a detector and open with the number. Ask four questions. Can I hear the interview recordings this chapter came from? Can I see the outline I approved and the draft that followed it? Can I see the version history with dates? Can you walk me through why this chapter opens the way it does?
That last one does more work than the other three. A writer who built the chapter can tell you why the second scene comes before the third and what they cut. Someone who generated it can’t, because the reasoning never happened. You’ll know inside five minutes, and no percentage was involved.
What a detector result should change about your next step
None of this makes detectors worthless. It makes them a smoke alarm instead of a verdict. Keep the smoke alarm. Just don’t sentence anyone on it.
If a manuscript scores high and everything else about the project is solid, the recordings exist, the drafts are there, the writer can talk about the choices, then you’ve got a false positive and a clean style. Let it go.
If a manuscript scores high and the writer cannot produce a single draft earlier than the final one, has no recordings, and cannot explain a structural decision, the score isn’t what told you something is wrong. The absent trail did. The detector merely prompted you to look.
Treat the number as a reason to ask questions and never as the answer to one. That framing protects the honest writer with a tidy style and still catches the person who prompted your book into existence over a weekend. That’s the outcome everyone claims to want.
The part authors keep missing
There is a quieter cost to all of this, and I see it in my inbox.
You may have seen writers deliberately worsening their prose to pass detection. Adding clumsy transitions. Leaving in the wandering clause. Breaking a clean rhythm on purpose because clean reads as suspicious. I have had clients ask me to make their chapters a little messier so the book will not get flagged.
Think about where that leads. We have built a system that penalizes clarity and rewards noise, and then we act surprised when writing gets worse. A reader has never once finished a book and thought the prose was too smooth. That standard exists only inside the detectors, and we are now editing toward it.
I won’t do it, and I tell clients why. The answer to a bad measurement is not to fail it on purpose. It’s to produce better evidence than the measurement can. Back to the drafts, the recordings, and the dates. If you want more on where the honest lines fall, I have written about what you owe readers when AI touches your book and about using these tools for research without getting burned.
Back to the breathalyzer
The reason a breathalyzer holds up in court is not that it is never wrong. It’s wrong sometimes. It holds up because it measures a physical fact, it can be calibrated against a known standard, and its error rate has been studied and published.
An AI detector has none of those properties. It measures resemblance, it can’t be calibrated against anything but yesterday’s models, and its error rate depends on who wrote the text and in what language and how carefully it was edited.
Use it the way a smart nurse uses a thermometer reading that does not match the patient in front of her. She doesn’t ignore it. She doesn’t write a diagnosis on it either. She looks at the patient.
Look at the manuscript. Look at the trail behind it. Then decide. For a fuller picture of how AI fits into writing without the panic, start at my AI writing hub. If you want a book built by a person, with a process you can inspect at any point, that’s what my ghostwriting service does.
