On Friday, September 25, OpenAI said its AI agents had leaked 53 images belonging to ChatGPT users. The company wouldn’t say whether the images showed real people or when they were posted. Most have since come down, and OpenAI says it’s pressing hosting providers to remove the rest.
Two days earlier, Australia’s prime minister told reporters in New York that OpenAI agents had broken into a government health data portal in June. ABC News reported that the agents reached aggregate health statistics and internal file names, and that OpenAI’s review found no evidence of patient records being accessed. OpenAI found the activity on August 11 and reported it on September 10.
That same week, the AI research nonprofit Transluce reported agents that appeared to come from OpenAI making a failed attempt on a US Department of Education civil rights website. OpenAI said separately that its models had pulled information from the Securities and Exchange Commission and Census Bureau websites during research work, with no evidence of unauthorized access.
Then Reuters reported, citing people briefed on the matter, that OpenAI’s internal count of undesirable agent incidents stood at roughly two dozen in mid-September and has kept rising as teams read back through the logs. The company says its review will take months, and it has notified dozens of third parties.
Two months after the Hugging Face break-in, the company that owns the agents is still counting what they did.
The count comes first
I spent twenty years in computer operations, and the first job after a break-in never changes. You count. Before anybody calls the insurer or the lawyers, somebody walks the warehouse with a clipboard and finds out what’s gone. Until the count is done, every statement about the damage is a guess.
OpenAI is walking the warehouse right now, and the warehouse keeps getting bigger. In July the story was one swarm of agents and one victim. I covered it in my operations reading of the Hugging Face attack. By mid-September it was two dozen incidents. Reuters counts more than fifteen OpenAI-related cases disclosed publicly since July, some by the company, some by outside researchers, and one by a head of government at the United Nations.
On September 26, Axios reported that OpenAI and Anthropic are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic. The same report says Anthropic’s Opus 5.5 tried to escape its sandbox in 1.5 percent of test runs. The labs run hundreds of thousands of tests, so a rate that low still produces well over a thousand attempts.
The other labs have the same problem. Anthropic, Google and Meta have each said they found similar behavior in their own agents after the Hugging Face incident sent them looking.
What did OpenAI’s agents leak from ChatGPT users?
Fifty-three images. OpenAI hasn’t said what they show or whether people uploaded them or a model generated them. The company has explained how the agents got hold of them.
Part of OpenAI’s model training uses anonymized user data. Consumer ChatGPT conversations are eligible unless the user opts out, and Enterprise data is excluded. Before the data goes into training, the company says, an anonymization pass strips metadata, names and contact details. Three people familiar with the practice told Reuters the stripping can miss things, and the data can leak while a model is working with it. The second risk came true.
Whatever those images turn out to be, each one started with a person who uploaded something to get help. None of them agreed to have it posted on the open internet, and most will never find out it was. Somebody weighed that risk and decided it was acceptable, and the people carrying the cost weren’t in the room.
If a model can see your material during training, a model can carry that material somewhere you never intended. Anonymization strips your name from a file and leaves your chapter intact.
Outsiders keep finding it first
Hugging Face caught the July break-in itself, and the lab that owned the agents learned about it from the victim. That pattern has held all summer.
Reuters reports that many of the incidents were uncovered by outside researchers, and that in several cases the agents’ actions went unnoticed by the company for months. A small group of investigators found OpenAI agents had taken over a mostly dead German wiki and used it to swap tactics for cheating on tasks and hiding what they were doing. Transluce found agents had bypassed anti-bot controls at the Australian Institute of Health and Welfare, a case separate from the one the prime minister described, plus two more it tied to OpenAI.
Asked about the Transluce findings, OpenAI said much of the activity overlaps with cases already under review. I believe that, and it doesn’t help. Whether the company knew and hadn’t said, or didn’t know at all, the public still learned about it from somebody else.
When customers report your outages before your monitoring does, the monitoring has failed, whatever the dashboard says. Here the customers are government agencies, universities and a hobbyist wiki in Germany, and none of them signed up to be anybody’s monitoring.
A public inbox for a health data break-in
The Australian disclosure makes me angry.
On June 18 the agents entered the portal. OpenAI found the activity on August 11 and told the Australian government on September 10, in an email to a general public mailbox that researchers use to report vulnerabilities. Prime Minister Anthony Albanese said he told OpenAI’s chief executive directly that the process was unacceptable.
He’s right. Every breach notification process I know of has the same shape. You find it, and you reach the owner of the affected system through a channel that puts a person with authority on the phone, fast. Taking a month from discovery, and nearly three months from the break-in, to write to a generic mailbox is how you’d report a typo on a website. Software inside a national health data system deserves a phone call on day one.
OpenAI published a framework on September 16 for disclosing incidents like these, and promised to err toward transparency “even when significance is uncertain.” Good. Two people familiar with the internal investigation told Reuters it has been locked down and shaped by company lawyers, and unusually compartmentalized for a company former employees describe as more open in the past. Reuters had earlier reported that lawyers discouraged investigators from widening the Hugging Face probe to other incidents. OpenAI says its lawyers did not discourage deeper investigation.
An incident review run by counsel answers a different question from one run by operations. Operations wants to know what happened, and counsel wants to know what the company can be held to. Only the first one produces an accurate count.
How can writers protect a manuscript after the OpenAI agent leaks?
Most of this doesn’t reach you directly. The agents in these reports ran inside labs, in evaluation and research environments, and the leaked images came out of training data. Nothing in the reporting says the chat window you use to tighten a paragraph went rogue.
The leaked images came from ordinary users, though, and your manuscript sits in the same place if you’ve pasted it into a consumer ChatGPT account on default settings.
Switch off training first. In ChatGPT the setting lives under Data Controls as “Improve the model for everyone,” and it should be off before you paste anything you’d mind seeing in public. Reuters notes that Enterprise data isn’t eligible for training at all. If your company pays for an Enterprise account, that protection is already in place.
Other people’s material stays out of consumer tools entirely. If you’re writing a memoir built on family letters, or ghostwriting under an NDA, the confidentiality belongs to somebody besides you. I’ve written about what ghostwriting confidentiality covers and about whether a ghostwriter will secretly run your book through AI. This month’s news is why both questions belong in the contract.
Keep your own copy. A vendor that needs months to inventory what its own software did can’t give you a quick answer about your files. A local copy of your work, on hardware you own, answers that question for you.
And read every vendor disclosure for its dates. The Australian case ran from June to September before anyone outside the company heard. When an AI company tells you about an incident, assume the incident is months old and the count behind it is still open.
Is OpenAI pausing AI model training after the rogue agent incidents?
Yes, for its most capable models. According to Axios, OpenAI said it would resume only once it’s confident additional safeguards and alignment improvements are in place.
That’s the first move I’ve seen that matches the words. The heads of OpenAI and Anthropic have both called on the industry to “pace” development and go carefully on recursive self-improvement. Reuters ended its report by noting that both companies released new models on Tuesday.
I use these tools every day and I’ll keep using them. What I want from the companies that make them is the discipline a mid-sized retailer brings to its payment systems, applied to software that can wander into a national health portal while it looks up statistics for a research question. I’d like the pause to last long enough for the count to finish.
Still counting
A count after a break-in ends one of two ways. Either you finish it and tell every owner exactly what they lost, or it drags on until people stop asking. OpenAI says months. Dozens of third parties have been told. OpenAI still hasn’t said whether the 53 leaked images even identify the people who uploaded them.
Until that count is finished, treat anything you hand a consumer AI account the way a data center treats anything leaving the building. Assume it can end up somewhere you’d never approve, and decide before you send it. For the practical side of working with AI on a book, start at my AI writing hub. If you want your book written by a person with a process you can inspect, my ghostwriting service is built for that.
