An author can now lose a publishing contract because a stranger uploaded their manuscript to a website and got a number back.
Not because anyone proved the book was written by a machine. Because a percentage appeared on a screen, and the percentage was high enough that the people holding the contract decided the risk was theirs to avoid.
Whatever you think about AI in publishing, that is not how evidence works. This piece is about why these tools cannot do what they are being asked to do, argued from how they function, not from anyone’s opinion about machines and art.
The number on the box is not the number that matters
Detection companies advertise accuracy in the high nineties. Some claim better than 99 percent. Those figures are usually honest as far as they go, and they are close to meaningless in practice.
Here is why. An accuracy figure comes from running the tool against a test set built for the purpose, half human writing and half machine writing, balanced on purpose so the score means something.
Nobody uses it that way. A reader does not feed a detector a random sample of ten thousand books. A reader feeds it the one book that felt wrong. Every text that reaches these tools has already passed through a human filter that selected it for being suspicious.
That changes the question completely. The advertised figure answers “how often is this tool right on a balanced sample.” The question that matters is “given that a suspicious-seeming manuscript was flagged, what are the odds it was machine written.” Those are different numbers, and the second one depends on how common machine-written manuscripts are among the ones people find suspicious. Nobody knows that figure, and the detection companies do not publish it, because it is not a property of their tool.
Anyone who has worked with medical screening knows this problem. A test with excellent accuracy, applied to a population where the condition is rare, produces mostly false positives. The test is not broken. It is being asked a question it was not built to answer.
Why do AI detectors give different results for the same text?
Because most of them are not measuring a fixed property. They are producing an estimate, and estimates move.
This is the part that should end the conversation on its own. Multiple reported cases involve the same document, unchanged, run twice, returning materially different scores. In one documented instance a set of published opinion pieces was flagged, then re-run, and produced results that did not match the first pass.
A measurement that does not repeat is not a measurement. It is a reading. If a scale gave you a different weight each time you stood on it, you would not use it to make decisions, and you would certainly not use it to make decisions about somebody else’s career.
Formatting alone can move a score. Strip the front matter, change the line breaks, remove the chapter headings, and the number shifts. Which means the result depends partly on what was done to the file before it was uploaded, and by whom, and with what intent.
Nobody can tell you what the tool is looking at
Ask a detection company what specifically triggered a flag and you will not get a straight answer. In at least one public exchange, the head of a detection company was asked directly and could only speculate, offering correct grammar and coherent argument as possibilities.
Sit with what that means for an accused writer. There is no rebuttal available. You cannot address the evidence because nobody will say what the evidence is. You are left proving a negative to people who have already made a decision, using a number nobody can explain, produced by a process nobody will describe.
In any other context we would call this an accusation without particulars. Courts throw those out. Publishing appears to accept them.
The finding that should have stopped all of this
In 2023, researchers at Stanford led by Weixin Liang published a study in the journal Patterns with a title that leaves little room for interpretation: GPT detectors are biased against non-native English writers.
They ran essays written by non-native English speakers through seven widely used detectors. More than 60 percent were misclassified as AI generated. One detector flagged 98 percent of them. Essays by native speakers, run through the same tools, were classified correctly at high rates.
The mechanism is not mysterious once you see it. Detectors lean heavily on a property called perplexity, which is roughly a measure of how surprising the next word is. Machine writing tends toward the predictable, because predicting the likely next word is what the machine does. Second-language writing also tends toward the predictable, because a writer working outside their first language reaches for the safe construction, the common word, the sentence pattern they are confident in.
The detector cannot tell those apart. It was never able to. It sees low lexical variety and flags it.
The same logic extends past second-language writers. Anyone writing in plain, unornamented prose is at elevated risk. So is anyone writing deep inside a genre with strong conventions, which brings us to the next problem.
What these tools may be detecting is convention
When flagged passages get published, they are worth reading carefully, because they are rarely what you would expect.
Phrases cited as evidence of machine authorship have included lines about a pulse thundering in someone’s ears and drowning out the crowd. That is not a machine fingerprint. That is a romance cliché, and it has been one for forty years.
The reasoning behind flagging such a phrase is that it appears often in suspected machine writing and rarely in confirmed human writing. But consider what that comparison captures. Machine models were trained on enormous quantities of genre fiction, so they reproduce genre convention faithfully. A human writer working in that genre reproduces the same conventions, for the same reason: it is what the form sounds like.
If a detector keys on conventional phrasing, then the more thoroughly a writer has absorbed the conventions of their genre, the more suspicious they appear. That is an inversion of what anyone wants. It penalizes fluency in a form.
There is a version of this critique that cuts the other way and deserves stating. Some of the books at the center of these controversies were, by wide agreement, not very good. Flat prose, tired images, dialogue that does not land. That is a real editorial judgment and it is worth making. It is also a completely different claim from “a machine wrote this,” and collapsing the two is how careers get destroyed over a taste dispute.
The chain of custody nobody is asking about
Set the technology aside and look at what has to happen for one of these accusations to reach a publisher.
Someone obtains a copy of the manuscript. Sometimes that copy is pirated. They may modify it, stripping the parts that would skew a result. They upload another person’s copyrighted work to a third-party service, without the author’s knowledge or consent, and that service processes it on servers the author has never heard of.
Then a screenshot of the resulting number travels further and faster than any correction ever will.
Every step of that would be inadmissible anywhere that takes evidence seriously. Unknown provenance, unverified handling, an instrument that does not reproduce, and an interested party operating it. In publishing it is a Tuesday.
And there is a commercial dimension worth naming plainly. Every accusation is free advertising for the detection company whose logo appears on the screenshot. Search interest in these tools spikes in step with each controversy. That does not mean anyone is acting in bad faith. It does mean the incentives all point one direction, and nobody involved is neutral.
What should an author do about AI detection accusations?
The uncomfortable answer is that you cannot win the argument after it starts, so the work is in what you keep beforehand.
Keep your drafts, dated. Not the final manuscript. The bad version from eighteen months ago, the one with the abandoned subplot. Cloud documents with version history are ideal because the timestamps are not yours to fake. A file that shows a chapter growing over six weeks is worth more than any denial.
Keep the debris. Notes, outlines, the research you did, the emails where you argued with your editor about chapter nine. The texture of a real process is difficult to manufacture after the fact and instantly recognizable to anyone who has made something.
Write down what you use and when. If you use a machine for spelling, grammar, research or brainstorming, put it in writing before anyone asks. A disclosed process is a position. An undisclosed one becomes an admission the moment somebody goes looking, whatever it was.
Get it into the contract. If you work with a publisher, a collaborator or an editor, agree in advance what tools are acceptable and who is responsible for what. Two of the recent cases involved manuscripts that passed through another person’s hands, and the resulting blame moved in circles.
Do not run your own work through a detector to reassure yourself. You will get a number, the number will mean nothing, and if it is high you will have talked yourself into a panic over a reading that does not reproduce.
The part that should worry publishers more than authors
Look at the sequence in these cases and something odd emerges.
Agents read the manuscript and championed it. Editors read it and bid on it. In one case fourteen publishers competed and the price climbed past two million dollars. Every professional in the chain read the words and wanted them.
Then a number appeared, and the same professionals reversed. The manuscript had not changed. Nothing new was learned about how it was written. A screenshot arrived and the judgment of every experienced reader who had touched the book was set aside in favor of it.
That should trouble anyone in publishing far more than the possibility that a machine wrote a novel. If professional editorial judgment collapses the moment a percentage appears, the judgment was not doing much work in the first place. And the question worth asking is what editors were doing with these manuscripts, because flat prose and tired metaphors are ordinary editorial problems with ordinary editorial fixes. Marking them as suspicious instead of fixing them is a strange way to run a publishing house.
Where this leaves us
None of this is an argument that machine-written books are fine, or that nobody is passing off generated text as their own. People are, the volume is rising, and it is a real problem for readers and writers both.
It is an argument that the tool being used to police it does not work well enough to police anything. A detector that cannot reproduce its own results, cannot explain what it measures, and misclassifies the majority of second-language writers is not a smoke detector. It is a coin with a logo on it.
The cost of using it anyway lands on specific people. Careers that took a decade end in an afternoon, and the correction, if it ever comes, reaches a fraction of the audience that saw the accusation.
If you are writing a book and this frightens you, the honest reassurance is small but real: keep your drafts, write down your process, and make something only you could have made. I have written elsewhere about whether AI can write your book and about the phrases that make writing sound machine-made, both of which are more useful than any detector. And if you want somebody in your corner who documents the whole process as a matter of course, that is what I do.
Source
The cases behind this piece were reported in a video essay that traces the sequence in each one and links its own sources in the description. I have left the authors unnamed here on purpose, since repeating an accusation with a name attached is the exact harm this article is about, but the reporting is worth watching:
The Guides That Get Your Book Written, Published, and Sold
Four short, practical guides on writing, publishing, and selling your book, plus the occasional note when there's something worth your time. No fluff, no daily inbox clutter. Drop your email and they're yours.
We use MailerLite to manage our list and send these emails. Your address is used only to send you what you signed up for. We will not sell it, share it, or use it for anything else, and you can unsubscribe anytime.
