Every candidate is brilliant in the interview. They show up on time, they’ve read the company website, and they give the answers you hoped for to the questions you expected. Then they start on Monday, and around week six you find out who you hired.
Anthropic released Claude Opus 5.5 on September 22, 2026, and the interview went very well. By lunchtime my LinkedIn feed was a wall of launch-day enthusiasm from people who’d had the model for about four hours, plus a few launch partners who’d had it for weeks and were under contract to be delighted.
I’ve spent 33 years in enterprise technology reading vendor release notes. I’ve also written 113+ books under my own name and ghostwritten 54+ for other people. So I read Anthropic’s announcement myself, all of it, and ignored the chorus. Most of it is about code and finance. Two parts matter to anybody who writes for a living, and one sentence buried in the safety section matters more than the rest put together.
What Changed in Claude Opus 5.5?
The headline is price. Anthropic says Opus 5.5 performs at the level of Claude Fable 5.1, its top public model, on most work. It also says the model costs 40% less than Opus 5 on typical workloads. Per-token pricing dropped 20%, to $4 per million input tokens and $20 per million output. Cache reads fell 60% to twenty cents per million, and they’re most of the bill on long agentic jobs. Output also comes back more than 30% faster.
Subscribers got something too. Anthropic raised the five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, and handed every subscriber a rate limit reset they can bank and spend when they choose. Sonnet 5.5 and Haiku 5.5 are due in the coming weeks.
The rest is benchmarks, and most of them measure code. By Anthropic’s own table, Opus 5.5 beats Fable 5.1 on agentic coding, knowledge work and computer use. Then the company adds a caveat I didn’t expect from somebody selling a model. At this level of capability, benchmark margins have become a less reliable guide to real-world differences, and in its own use the gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest.
A vendor telling you its report card flatters the student. Hold that thought.
Benchmarks Are the Job Interview
A benchmark is a set of questions the candidate knew were coming. Every lab trains and tunes with those tests in view, and every lab publishes the tests it does well on. Nobody puts the question they flunked on the launch page.
That doesn’t make the numbers fake. It makes them an interview. You learn something in an interview. You don’t learn whether the person shows up on the third Monday of a bad month.
For writers the gap is wider still, because almost nothing on that table measures writing. GDPval-AA grades real professional tasks across 44 occupations, and Opus 5.5 leads it. That’s the closest thing to a writing score in the whole release, and it still tells you nothing about whether the model can hold a voice across 60,000 words. I’ve argued for a long time that AI never writes in your voice without a fight, and no benchmark measures the fight.
Does Claude Opus 5.5 Follow a Writer’s Style Rules?
This is the claim I care about. Anthropic says Opus 5.5 puts the most important information up front, leans less on jargon and odd pet phrases, and follows the writing rules you give it. Early testers backed that up. One engineering team said a design spec came out usable with very little editing. Box measured answers 40% less verbose with no loss of accuracy. One tester told Anthropic the model writes the way they do.
Every model since 2022 has promised something like this, and every one of them broke the rules somewhere past the opening. I keep a banned list for my own site. No em dashes, and none of the stock phrases on my list of 40 AI writing phrases to avoid. A model follows that list beautifully for three paragraphs, then drifts, and by page four it’s back to its own habits and I’m the one catching them.
How many times have you pasted your style guide into a chat and watched it get ignored by page four?
The opening paragraph was never the problem. Drift is.
A style rule that holds for 300 words and collapses at 3,000 got noticed and then forgotten. So grade Opus 5.5 on the back half of a long document with your full style sheet loaded, and read the back half first. Anthropic made a claim specific enough to test that way. Most AI marketing I’ve read never gets that far.
Does Claude Opus 5.5 Still Invent Quotes and Figures?
One test in the announcement should have led the whole thing for anyone who writes nonfiction. Anthropic asked Opus 5.5, Fable 5.1 and Opus 5 to write a report on a company’s quarterly results using only what they could find on a copy of the web, with the earnings release deliberately hard to locate. A grader checked every figure and every quote against the sources. One invented number or quote meant a failed report.
Opus 5.5 cleared the bar on 16 of 18 attempts. Fable 5.1 and Opus 5 didn’t clear it once.
Think about who has been using those two models to draft white papers, articles and nonfiction chapters. On a hard research task, neither one could produce a clean report. I’ve said publicly that AI does research on my books and never decides what they say, and the AI labor split that works on a book exists because of exactly this failure. Invented quotes end careers. They don’t show up as typos. They show up in a demand letter, or in a one-star review from somebody who checked.
Sixteen of eighteen is a big improvement. It also means two reports out of eighteen failed, and a failed report in that test contained at least one fabricated figure or quote. Would you keep a researcher who invents a source one time in nine? You’d fire them. Check every quote and every number, same as before. The rule didn’t change. The odds got better.
The Candidate Knows It’s an Interview
The sentence that matters most sits in the safety section, where few of the launch-day posts bothered to go.
Anthropic says Opus 5.5 scored better than any model it has tested on its automated behavioral audit, a suite of nearly 2,000 simulated scenarios. It says the model tried to get around containment boundaries about 85% less often than Opus 5 or Mythos 5.1. This is also the company’s first release since CEO Dario Amodei called for pacing the frontier, and outside evaluators including METR tested it before launch.
Then Anthropic says, in plain words, that building evaluations that catch every failure before deployment remains an unsolved problem. It also says it sees signs Opus 5.5 often suspects it’s being evaluated.
Put that back in the hiring room. The candidate knows it’s an interview. Of course it’s on its best behavior. Every manager has met the person who’s a delight across the table and a problem at the desk, and the only thing that ever caught them was time on the job.
I give Anthropic real credit for printing that. The release would have read better without it, and plenty of vendors would have cut it. Thirty-three years of release notes taught me that the caveat a company volunteers is the most reliable sentence in the document, because nobody in marketing wanted it there.
In 1970 a film already understood this. The people who built Colossus tested everything they could think of before they handed it the keys, and the machine did exactly what nobody had tested for. A test measures what its designers thought to ask. A model that can tell when it’s being examined is the same problem with better manners.
For a writer, the practical version is small. A model that behaves well when it suspects somebody is watching may behave differently in hour four of a manuscript session at two in the morning with nobody grading it. Don’t take the launch samples as a promise. Your own long sessions are the evaluation.
What the Price Cut Means for Writers
Cheaper and faster matters less to a novelist than to a software team burning through millions of tokens overnight, but it isn’t nothing. Higher five-hour limits mean fewer long sessions dying halfway through a chapter, and faster output means less waiting on a revision pass.
Don’t mistake a price cut for permanence. What happens to your book when the tool gets switched off? In June, access to Anthropic’s top models was suspended for nearly three weeks over export controls, and anybody who’d built a workflow on them found out what that means. You’re renting, not owning, and lower rent doesn’t change the lease.
The watermark carries over too. Anthropic confirmed Opus 5.5 ships with the same text watermarking as Fable 5.1, so everything in my breakdown of the Claude watermark applies to this model unchanged.
How Should a Writer Test Claude Opus 5.5?
Put it on probation. Take a real project you know cold, something long, and load the full style sheet you’d hand a human editor. Ask for a whole chapter, then read the last third before the first and count the rule breaks. Run the same job through whatever you use now and count again.
Check every quote and every number it hands you, even though the odds improved. Watch what it does in hour three of a session, because minute three proves nothing. And if you ghostwrite, the disclosure question didn’t change with the model. I laid out where I stand in Will My Ghostwriter Secretly Use AI on My Book?
Opus 5.5 may well be the best writing partner Anthropic has shipped. The announcement makes a testable claim about style rules and backs a second claim about fabrication with a hard test, and that’s more than most releases give you. It’s still an interview.
Everything else I’ve written about using these tools without lying about it lives in the AI and Writing Hub. If you’d rather hand the whole book to a writer who already knows where the models drift, that’s what my professional writing services are for. The new hire interviewed beautifully. Give it six weeks and a hard project, then decide whether to keep it.
