Latest
What an AI Detector Score on Your Manuscript Is WorthThe One-Hour Call Before I Quote Your BookWhen a Client Thinks the Ghostwriter Used AIMonthly or Milestone: How Ghostwriting Gets BilledWhen Your Memoir Should Be a NovelWhat It Costs to Fix an AI-Written ManuscriptThe Clients Who Pay and VanishThe Quotation Marks That Get Authors SuedThe Work You Would Never Have StartedWhen Your Own Memoir Sounds Like BraggingWhat Belongs on a Copyright PageThe Hugging Face AI Agent Attack: An Operations ReadingBehind the Book: The Mysterious Island, Neb’s SideHow to Organize Decades of Memories Into a MemoirWhy Rotten Tomatoes Sucks: The Score Does Not Mean What You ThinkWhy Amazon KDP Sucks: They Terminated My Account OvernightIngramSpark: How I Publish Now and WhyWhy Fiverr Sucks for Ghostwriting: The Buyer’s SideWhy eBay Sucks Now: A Seller’s Numbers and a Buyer’s WarningThe Ghost Story TraditionThe Gothic TraditionThe Christmas Ghost Story TraditionResurrection as a Narrative StructureBooks to Give a WriterThe Beach Read ArgumentWhy It’s a Wonderful Life Failed on ReleaseWhat to Read in SpringWhat to Read in SummerWhat to Read in OctoberHow Warner Bros. Dismantled a $17 Billion Cartoon EmpireThe Imaginary Scarcity TrapThe Graph That Goes Vertical Is Usually Somebody Else’sSubstack Is Not Collapsing. The Promise Was.The Disasters That Happen to Ordinary PeopleToba: The Winter That Almost Ended UsJay Stifflemire: Nothing Ever Gets Written DownGeorgie-Ann Getton: I Forgot I Had Free WillAI Detection Cannot Be Evidence, and Publishing Is Using It That WayAI Consciousness Left Philosophy and Entered the LaboratoryThe Office Block Where the Bedrooms AreThe Web Got Fenced: What AI Search Costs Small SitesBlack Tuesday: The Web Ring War Nobody Outside It NoticedWhat the AI Visibility Industry Sells, and What the Evidence SaysBlack Tuesday: The Original ring-master.net Page, 2000Behind the Book: Peacekeeper, The Dissolution WarsBehind the Book: Real World SurvivalBehind the Book: Publish Your BookBehind the Book: ReincarnationBehind the Book: Sell Your BooksBehind the Book: Show Don’t Tell
The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

AI Detection Cannot Be Evidence, and Publishing Is Using It That Way

This entry is part 8 of 11 in the series AI for the Worried
TL;DR: AI detection tools are being used as evidence in publishing, and they cannot bear that weight. The accuracy figures are measured on balanced test sets and then applied to text somebody already found suspicious. That’s a different problem entirely. The same manuscript can score differently on a second run. The vendors often cannot say what their tools key on. And published research shows detectors flag non-native English speakers at enormous rates. Keep your drafts.

An author can now lose a publishing contract because a stranger uploaded their manuscript to a website and got a number back.

Not because anyone proved the book was written by a machine. Because a percentage appeared on a screen, and the percentage was high enough that the people holding the contract decided the risk was theirs to avoid.

Whatever you think about AI in publishing, that’s not how evidence works. This piece is about why these tools cannot do what they’re being asked to do, argued from how they function, not from anyone’s opinion about machines and art.

The number on the box is not the number that matters

Detection companies advertise accuracy in the high nineties. Some claim better than 99 percent. Those figures are usually honest as far as they go, and they’re close to meaningless in practice.

Here’s why. An accuracy figure comes from running the tool against a test set built for the purpose, half human writing and half machine writing, balanced on purpose so the score means something.

Nobody uses it that way. A reader doesn’t feed a detector a random sample of ten thousand books. A reader feeds it the one book that felt wrong. Every text that reaches these tools has already passed through a human filter that selected it for being suspicious.

That changes the question completely. The advertised figure answers “how often is this tool right on a balanced sample.” The question that matters is “given that a suspicious-seeming manuscript was flagged, what are the odds it was machine written.” Those are different numbers, and the second one depends on how common machine-written manuscripts are among the ones people find suspicious. Nobody knows that figure, and the detection companies don’t publish it, because it’s not a property of their tool.

Nobody knows that figure, and the detection companies do not publish it, because it is not a property of their tool.
Share on X

Anyone who has worked with medical screening knows this problem. A test with excellent accuracy, applied to a population where the condition is rare, produces mostly false positives. The test isn’t broken. It’s being asked a question it wasn’t built to answer.

Why do AI detectors give different results for the same text?

Because most of them aren’t measuring a fixed property. They’re producing an estimate, and estimates move.

This is the part that should end the conversation on its own. Multiple reported cases involve the same document, unchanged, run twice, returning materially different scores. In one documented instance a set of published opinion pieces was flagged, then re-run, and produced results that didn’t match the first pass.

A measurement that doesn’t repeat isn’t a measurement. It’s a reading. If a scale gave you a different weight each time you stood on it, you wouldn’t use it to make decisions, and you’d certainly not use it to make decisions about somebody else’s career.

Formatting alone can move a score. Strip the front matter, change the line breaks, remove the chapter headings, and the number shifts. Which means the result depends partly on what was done to the file before it was uploaded, and by whom, and with what intent.

Nobody can tell you what the tool is looking at

Ask a detection company what specifically triggered a flag and you won’t get a straight answer. In at least one public exchange, the head of a detection company was asked directly and could only speculate, offering correct grammar and coherent argument as possibilities.

Sit with what that means for an accused writer. There’s no rebuttal available. You cannot address the evidence because nobody will say what the evidence is. You’re left proving a negative to people who have already made a decision, using a number nobody can explain, produced by a process nobody will describe.

In any other context we’d call this an accusation without particulars. Courts throw those out. Publishing appears to accept them.

The finding that should have stopped all of this

In 2023, researchers at Stanford led by Weixin Liang published a study in the journal Patterns with a title that leaves little room for interpretation: GPT detectors are biased against non-native English writers.

They ran essays written by non-native English speakers through seven widely used detectors. More than 60 percent were misclassified as AI generated. One detector flagged 98 percent of them. Essays by native speakers, run through the same tools, were classified correctly at high rates.

The mechanism isn’t mysterious once you see it. Detectors lean heavily on a property called perplexity, which is roughly a measure of how surprising the next word is. Machine writing tends toward the predictable, because predicting the likely next word is what the machine does. Second-language writing also tends toward the predictable, because a writer working outside their first language reaches for the safe construction, the common word, the sentence pattern they’re confident in.

The detector cannot tell those apart. It was never able to. It sees low lexical variety and flags it.

The same logic extends past second-language writers. Anyone writing in plain, unornamented prose is at elevated risk. So is anyone writing deep inside a genre with strong conventions, which brings us to the next problem.

What these tools may be detecting is convention

When flagged passages get published, they’re worth reading carefully, because they’re rarely what you’d expect.

Phrases cited as evidence of machine authorship have included lines about a pulse thundering in someone’s ears and drowning out the crowd. That’s not a machine fingerprint. That’s a romance cliché, and it’s been one for forty years.

The reasoning behind flagging such a phrase is that it appears often in suspected machine writing and rarely in confirmed human writing. But consider what that comparison captures. Machine models were trained on enormous quantities of genre fiction, so they reproduce genre convention faithfully. A human writer working in that genre reproduces the same conventions, for the same reason: it’s what the form sounds like.

If a detector keys on conventional phrasing, then the more thoroughly a writer has absorbed the conventions of their genre, the more suspicious they appear. That’s an inversion of what anyone wants. It penalizes fluency in a form.

There’s a version of this critique that cuts the other way and deserves stating. Some of the books at the center of these controversies were, by wide agreement, not very good. Flat prose, tired images, dialogue that doesn’t land. That’s a real editorial judgment and it’s worth making. It’s also a completely different claim from “a machine wrote this,” and collapsing the two is how careers get destroyed over a taste dispute.

The chain of custody nobody is asking about

Set the technology aside and look at what has to happen for one of these accusations to reach a publisher.

Someone obtains a copy of the manuscript. Sometimes that copy is pirated. They may modify it, stripping the parts that would skew a result. They upload another person’s copyrighted work to a third-party service, without the author’s knowledge or consent, and that service processes it on servers the author has never heard of.

Then a screenshot of the resulting number travels further and faster than any correction ever will.

Every step of that would be inadmissible anywhere that takes evidence seriously. Unknown provenance, unverified handling, an instrument that doesn’t reproduce, and an interested party operating it. In publishing it’s a Tuesday.

And there’s a commercial dimension worth naming plainly. Every accusation is free advertising for the detection company whose logo appears on the screenshot. Search interest in these tools spikes in step with each controversy. That doesn’t mean anyone is acting in bad faith. It does mean the incentives all point one direction, and nobody involved is neutral.

What should an author do about AI detection accusations?

The uncomfortable answer is that you cannot win the argument after it starts, so the work is in what you keep beforehand.

Keep your drafts, dated. Not the final manuscript. The bad version from eighteen months ago, the one with the abandoned subplot. Cloud documents with version history are ideal because the timestamps aren’t yours to fake. A file that shows a chapter growing over six weeks is worth more than any denial.

Keep the debris. Notes, outlines, the research you did, the emails where you argued with your editor about chapter nine. The texture of a real process is difficult to manufacture after the fact and instantly recognizable to anyone who has made something.

Write down what you use and when. If you use a machine for spelling, grammar, research or brainstorming, put it in writing before anyone asks. A disclosed process is a position. An undisclosed one becomes an admission the moment somebody goes looking, whatever it was.

Get it into the contract. If you work with a publisher, a collaborator or an editor, agree in advance what tools are acceptable and who’s responsible for what. Two of the recent cases involved manuscripts that passed through another person’s hands, and the resulting blame moved in circles.

Don’t run your own work through a detector to reassure yourself. You’ll get a number, the number will mean nothing, and if it’s high you’ll have talked yourself into a panic over a reading that doesn’t reproduce.

The part that should worry publishers more than authors

Look at the sequence in these cases and something odd emerges.

Agents read the manuscript and championed it. Editors read it and bid on it. In one case fourteen publishers competed and the price climbed past two million dollars. Every professional in the chain read the words and wanted them.

Then a number appeared, and the same professionals reversed. The manuscript hadn’t changed. Nothing new was learned about how it was written. A screenshot arrived and the judgment of every experienced reader who had touched the book was set aside in favor of it.

That should trouble anyone in publishing far more than the possibility that a machine wrote a novel. If professional editorial judgment collapses the moment a percentage appears, the judgment wasn’t doing much work in the first place. And the question worth asking is what editors were doing with these manuscripts, because flat prose and tired metaphors are ordinary editorial problems with ordinary editorial fixes. Marking them as suspicious instead of fixing them is a strange way to run a publishing house.

Where this leaves us

None of this is an argument that machine-written books are fine, or that nobody is passing off generated text as their own. People are, the volume is rising, and it’s a real problem for readers and writers both.

It’s an argument that the tool being used to police it doesn’t work well enough to police anything. A detector that cannot reproduce its own results, cannot explain what it measures, and misclassifies the majority of second-language writers isn’t a smoke detector. It’s a coin with a logo on it.

The cost of using it anyway lands on specific people. Careers that took a decade end in an afternoon, and the correction, if it ever comes, reaches a fraction of the audience that saw the accusation.

If you’re writing a book and this frightens you, the honest reassurance is small but real: keep your drafts, write down your process, and make something only you could have made. I’ve written elsewhere about whether AI can write your book and about the phrases that make writing sound machine-made, both of which are more useful than any detector. And if you want somebody in your corner who documents the whole process as a matter of course, that is what I do.

Source

The cases behind this piece were reported in a video essay that traces the sequence in each one and links its own sources in the description. I’ve left the authors unnamed here on purpose, since repeating an accusation with a name attached is the exact harm this article is about, but the reporting is worth watching:

Frequently Asked Questions

Are AI detectors accurate?
Not in the way the advertised figures suggest. Accuracy claims in the high nineties come from balanced test sets built for measurement. In real use, detectors are only run on text somebody already found suspicious, which is a filtered sample and produces a completely different error profile. Several tools also return different scores for the same unchanged document on a second run, and a measurement that doesn’t reproduce isn’t a measurement.
Why do AI detectors flag human writing as AI?
Most rely on perplexity, roughly how predictable each next word is. Machine text is predictable because predicting likely words is what the machine does. But plain prose is predictable too, and so is genre writing that follows convention, and so is writing by someone working in a second language. The detector sees low variety and flags it. It’s no way to tell the causes apart.
Do AI detectors discriminate against non-native English speakers?
Published research says yes. A 2023 Stanford study led by Weixin Liang, in the journal Patterns, ran essays by non-native English speakers through seven common detectors and found more than 60 percent misclassified as AI generated, with one tool flagging 98 percent. Essays by native speakers were classified correctly at much higher rates. The likely cause is that second-language writing reaches for safe, common constructions. That’s also what detectors read as machine-generated.
What does a percentage score from an AI detector mean?
Less than it appears to. It’s not a proportion of the text written by a machine, and it’s not a probability in any calibrated sense. It’s an internal confidence figure whose derivation the companies frequently decline or are unable to explain. There’s no agreed threshold at which a score becomes evidence. That’s why published accusations have ranged from 60 percent to 97 percent with entirely different outcomes.
How can a writer prove they wrote their own book?
Before an accusation, not after. Keep dated drafts including the bad early ones, ideally in a cloud document whose version history you cannot edit. Keep notes, outlines, research and correspondence with editors. Write down which tools you use and for what, in advance. A manuscript that visibly grew over months, with abandoned material still traceable, is far more persuasive than any denial issued afterward.
Should I run my own manuscript through an AI detector?
No. You’ll get a number that doesn’t reproduce and cannot be interpreted, and a high score will frighten you over nothing. There’s also a practical risk: uploading an unpublished manuscript to a third-party service means handing your unreleased work to a company whose data handling you haven’t reviewed. Spend the time on your drafts instead.

📝 Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

0 comments

No comments yet. Yours can be the first.

Was this useful?

Leave a comment