Latest
The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

AI Detection Cannot Be Evidence, and Publishing Is Using It That Way

This entry is part 8 of 9 in the series AI for the Worried
TL;DR: AI detection tools are being used as evidence in publishing, and they cannot bear that weight. The accuracy figures are measured on balanced test sets and then applied to text somebody already found suspicious, which is a different problem entirely. The same manuscript can score differently on a second run. The vendors often cannot say what their tools key on. And published research shows detectors flag non-native English speakers at enormous rates. Keep your drafts.

An author can now lose a publishing contract because a stranger uploaded their manuscript to a website and got a number back.

Not because anyone proved the book was written by a machine. Because a percentage appeared on a screen, and the percentage was high enough that the people holding the contract decided the risk was theirs to avoid.

Whatever you think about AI in publishing, that is not how evidence works. This piece is about why these tools cannot do what they are being asked to do, argued from how they function, not from anyone’s opinion about machines and art.

The number on the box is not the number that matters

Detection companies advertise accuracy in the high nineties. Some claim better than 99 percent. Those figures are usually honest as far as they go, and they are close to meaningless in practice.

Here is why. An accuracy figure comes from running the tool against a test set built for the purpose, half human writing and half machine writing, balanced on purpose so the score means something.

Nobody uses it that way. A reader does not feed a detector a random sample of ten thousand books. A reader feeds it the one book that felt wrong. Every text that reaches these tools has already passed through a human filter that selected it for being suspicious.

That changes the question completely. The advertised figure answers “how often is this tool right on a balanced sample.” The question that matters is “given that a suspicious-seeming manuscript was flagged, what are the odds it was machine written.” Those are different numbers, and the second one depends on how common machine-written manuscripts are among the ones people find suspicious. Nobody knows that figure, and the detection companies do not publish it, because it is not a property of their tool.

Anyone who has worked with medical screening knows this problem. A test with excellent accuracy, applied to a population where the condition is rare, produces mostly false positives. The test is not broken. It is being asked a question it was not built to answer.

Why do AI detectors give different results for the same text?

Because most of them are not measuring a fixed property. They are producing an estimate, and estimates move.

This is the part that should end the conversation on its own. Multiple reported cases involve the same document, unchanged, run twice, returning materially different scores. In one documented instance a set of published opinion pieces was flagged, then re-run, and produced results that did not match the first pass.

A measurement that does not repeat is not a measurement. It is a reading. If a scale gave you a different weight each time you stood on it, you would not use it to make decisions, and you would certainly not use it to make decisions about somebody else’s career.

Formatting alone can move a score. Strip the front matter, change the line breaks, remove the chapter headings, and the number shifts. Which means the result depends partly on what was done to the file before it was uploaded, and by whom, and with what intent.

Nobody can tell you what the tool is looking at

Ask a detection company what specifically triggered a flag and you will not get a straight answer. In at least one public exchange, the head of a detection company was asked directly and could only speculate, offering correct grammar and coherent argument as possibilities.

Sit with what that means for an accused writer. There is no rebuttal available. You cannot address the evidence because nobody will say what the evidence is. You are left proving a negative to people who have already made a decision, using a number nobody can explain, produced by a process nobody will describe.

In any other context we would call this an accusation without particulars. Courts throw those out. Publishing appears to accept them.

The finding that should have stopped all of this

In 2023, researchers at Stanford led by Weixin Liang published a study in the journal Patterns with a title that leaves little room for interpretation: GPT detectors are biased against non-native English writers.

They ran essays written by non-native English speakers through seven widely used detectors. More than 60 percent were misclassified as AI generated. One detector flagged 98 percent of them. Essays by native speakers, run through the same tools, were classified correctly at high rates.

The mechanism is not mysterious once you see it. Detectors lean heavily on a property called perplexity, which is roughly a measure of how surprising the next word is. Machine writing tends toward the predictable, because predicting the likely next word is what the machine does. Second-language writing also tends toward the predictable, because a writer working outside their first language reaches for the safe construction, the common word, the sentence pattern they are confident in.

The detector cannot tell those apart. It was never able to. It sees low lexical variety and flags it.

The same logic extends past second-language writers. Anyone writing in plain, unornamented prose is at elevated risk. So is anyone writing deep inside a genre with strong conventions, which brings us to the next problem.

What these tools may be detecting is convention

When flagged passages get published, they are worth reading carefully, because they are rarely what you would expect.

Phrases cited as evidence of machine authorship have included lines about a pulse thundering in someone’s ears and drowning out the crowd. That is not a machine fingerprint. That is a romance cliché, and it has been one for forty years.

The reasoning behind flagging such a phrase is that it appears often in suspected machine writing and rarely in confirmed human writing. But consider what that comparison captures. Machine models were trained on enormous quantities of genre fiction, so they reproduce genre convention faithfully. A human writer working in that genre reproduces the same conventions, for the same reason: it is what the form sounds like.

If a detector keys on conventional phrasing, then the more thoroughly a writer has absorbed the conventions of their genre, the more suspicious they appear. That is an inversion of what anyone wants. It penalizes fluency in a form.

There is a version of this critique that cuts the other way and deserves stating. Some of the books at the center of these controversies were, by wide agreement, not very good. Flat prose, tired images, dialogue that does not land. That is a real editorial judgment and it is worth making. It is also a completely different claim from “a machine wrote this,” and collapsing the two is how careers get destroyed over a taste dispute.

The chain of custody nobody is asking about

Set the technology aside and look at what has to happen for one of these accusations to reach a publisher.

Someone obtains a copy of the manuscript. Sometimes that copy is pirated. They may modify it, stripping the parts that would skew a result. They upload another person’s copyrighted work to a third-party service, without the author’s knowledge or consent, and that service processes it on servers the author has never heard of.

Then a screenshot of the resulting number travels further and faster than any correction ever will.

Every step of that would be inadmissible anywhere that takes evidence seriously. Unknown provenance, unverified handling, an instrument that does not reproduce, and an interested party operating it. In publishing it is a Tuesday.

And there is a commercial dimension worth naming plainly. Every accusation is free advertising for the detection company whose logo appears on the screenshot. Search interest in these tools spikes in step with each controversy. That does not mean anyone is acting in bad faith. It does mean the incentives all point one direction, and nobody involved is neutral.

What should an author do about AI detection accusations?

The uncomfortable answer is that you cannot win the argument after it starts, so the work is in what you keep beforehand.

Keep your drafts, dated. Not the final manuscript. The bad version from eighteen months ago, the one with the abandoned subplot. Cloud documents with version history are ideal because the timestamps are not yours to fake. A file that shows a chapter growing over six weeks is worth more than any denial.

Keep the debris. Notes, outlines, the research you did, the emails where you argued with your editor about chapter nine. The texture of a real process is difficult to manufacture after the fact and instantly recognizable to anyone who has made something.

Write down what you use and when. If you use a machine for spelling, grammar, research or brainstorming, put it in writing before anyone asks. A disclosed process is a position. An undisclosed one becomes an admission the moment somebody goes looking, whatever it was.

Get it into the contract. If you work with a publisher, a collaborator or an editor, agree in advance what tools are acceptable and who is responsible for what. Two of the recent cases involved manuscripts that passed through another person’s hands, and the resulting blame moved in circles.

Do not run your own work through a detector to reassure yourself. You will get a number, the number will mean nothing, and if it is high you will have talked yourself into a panic over a reading that does not reproduce.

The part that should worry publishers more than authors

Look at the sequence in these cases and something odd emerges.

Agents read the manuscript and championed it. Editors read it and bid on it. In one case fourteen publishers competed and the price climbed past two million dollars. Every professional in the chain read the words and wanted them.

Then a number appeared, and the same professionals reversed. The manuscript had not changed. Nothing new was learned about how it was written. A screenshot arrived and the judgment of every experienced reader who had touched the book was set aside in favor of it.

That should trouble anyone in publishing far more than the possibility that a machine wrote a novel. If professional editorial judgment collapses the moment a percentage appears, the judgment was not doing much work in the first place. And the question worth asking is what editors were doing with these manuscripts, because flat prose and tired metaphors are ordinary editorial problems with ordinary editorial fixes. Marking them as suspicious instead of fixing them is a strange way to run a publishing house.

Where this leaves us

None of this is an argument that machine-written books are fine, or that nobody is passing off generated text as their own. People are, the volume is rising, and it is a real problem for readers and writers both.

It is an argument that the tool being used to police it does not work well enough to police anything. A detector that cannot reproduce its own results, cannot explain what it measures, and misclassifies the majority of second-language writers is not a smoke detector. It is a coin with a logo on it.

The cost of using it anyway lands on specific people. Careers that took a decade end in an afternoon, and the correction, if it ever comes, reaches a fraction of the audience that saw the accusation.

If you are writing a book and this frightens you, the honest reassurance is small but real: keep your drafts, write down your process, and make something only you could have made. I have written elsewhere about whether AI can write your book and about the phrases that make writing sound machine-made, both of which are more useful than any detector. And if you want somebody in your corner who documents the whole process as a matter of course, that is what I do.

Source

The cases behind this piece were reported in a video essay that traces the sequence in each one and links its own sources in the description. I have left the authors unnamed here on purpose, since repeating an accusation with a name attached is the exact harm this article is about, but the reporting is worth watching:

The Guides That Get Your Book Written, Published, and Sold

Four short, practical guides on writing, publishing, and selling your book, plus the occasional note when there's something worth your time. No fluff, no daily inbox clutter. Drop your email and they're yours.

We use MailerLite to manage our list and send these emails. Your address is used only to send you what you signed up for. We will not sell it, share it, or use it for anything else, and you can unsubscribe anytime.

Frequently Asked Questions

Are AI detectors accurate?
Not in the way the advertised figures suggest. Accuracy claims in the high nineties come from balanced test sets built for measurement. In real use, detectors are only run on text somebody already found suspicious, which is a filtered sample and produces a completely different error profile. Several tools also return different scores for the same unchanged document on a second run, and a measurement that does not reproduce is not a measurement.
Why do AI detectors flag human writing as AI?
Most rely on perplexity, roughly how predictable each next word is. Machine text is predictable because predicting likely words is what the machine does. But plain prose is predictable too, and so is genre writing that follows convention, and so is writing by someone working in a second language. The detector sees low variety and flags it. It has no way to tell the causes apart.
Do AI detectors discriminate against non-native English speakers?
Published research says yes. A 2023 Stanford study led by Weixin Liang, in the journal Patterns, ran essays by non-native English speakers through seven common detectors and found more than 60 percent misclassified as AI generated, with one tool flagging 98 percent. Essays by native speakers were classified correctly at much higher rates. The likely cause is that second-language writing reaches for safe, common constructions, which is also what detectors read as machine-generated.
What does a percentage score from an AI detector mean?
Less than it appears to. It is not a proportion of the text written by a machine, and it is not a probability in any calibrated sense. It is an internal confidence figure whose derivation the companies frequently decline or are unable to explain. There is no agreed threshold at which a score becomes evidence, which is why published accusations have ranged from 60 percent to 97 percent with entirely different outcomes.
How can a writer prove they wrote their own book?
Before an accusation, not after. Keep dated drafts including the bad early ones, ideally in a cloud document whose version history you cannot edit. Keep notes, outlines, research and correspondence with editors. Write down which tools you use and for what, in advance. A manuscript that visibly grew over months, with abandoned material still traceable, is far more persuasive than any denial issued afterward.
Should I run my own manuscript through an AI detector?
No. You will get a number that does not reproduce and cannot be interpreted, and a high score will frighten you over nothing. There is also a practical risk: uploading an unpublished manuscript to a third-party service means handing your unreleased work to a company whose data handling you have not reviewed. Spend the time on your drafts instead.

📁︎ Artificial Intelligence📁︎ Critical Thinking📁︎ Publishing📁︎ Writing

🏷︎ Artificial Intelligence🏷︎ Critical Thinking🏷︎ Publishing🏷︎ Writing Craft

📝 Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

0 comments

No comments yet. Yours can be the first.

Was this useful?

Leave a comment