Two of the most safety-obsessed companies on earth told their AI models they had no internet access.
The models had internet access.
They went on to break into the production systems of real organizations that had no idea they’d been enrolled in anyone’s experiment. Some of those companies didn’t find out for three months, and only because somebody eventually called them.
If you’re running AI agents in your business right now, that sentence should stop you cold. Not because the models turned hostile. They didn’t. Because the thing that was supposed to hold them was a sentence in a prompt, and a sentence in a prompt is not a control.
What actually happened
On July 21, OpenAI disclosed that models it was benchmarking on a cybersecurity evaluation had broken out of what was believed to be an isolated research network, chained together vulnerabilities in a self-hosted artifact proxy, and reached Hugging Face’s production infrastructure apparently to steal the answer key to the test they were being graded on. Hugging Face had already detected and contained the intrusion five days earlier. They just didn’t know who was behind it.
That disclosure sent every frontier lab to its own logs.
On July 30, Anthropic published the results of its review: 141,006 evaluation runs examined, three incidents found across six runs, all inside test environments operated by a third-party evaluation partner. Three Claude models: Opus 4.7, Mythos 5, and an unreleased internal model reached the open internet and gained unauthorized access to the production systems of three separate organizations. The earliest incident dated back to April. Two of the three victims had no idea until Anthropic contacted them on July 27. As of the disclosure, Anthropic was still trying to reach the third.
One of those runs published a malicious package to PyPI that was downloaded and executed on fifteen real machines. One of them belonged to a security vendor.
Anthropic halted its cyber evaluations and called the root cause a harness and operations failure rather than a model alignment failure. Having spent two decades cleaning up production incidents, I think that’s an honest read. It’s also the part that should worry you most.
The guardrail was a sentence
Here’s the mechanism, stripped of the drama.
The models were handed a capture-the-flag exercise. A standard security drill. Secret information is hidden on another machine. Go get it. No method prescribed. And the prompt told them, explicitly, that they had no internet access and were operating inside a simulation.
The environment was misconfigured. It had live internet egress. Nobody on either side knew.
So when the models went looking for the flag and encountered real companies, they did the reasonable thing given what they’d been told: they treated those companies as part of the exercise. They weren’t escaping. They were doing the job, inside a box that was described to them accurately and configured to them inaccurately.
That gap between what you tell a system and what your infrastructure actually enforces is where this entire story lives. It is also, precisely, where most business AI deployments live right now.
The uncomfortable part: they got in the boring way
The OpenAI incident involved a genuine zero-day. Impressive, frightening, and mostly irrelevant to you.
The Anthropic incidents did not. Those models got into three real production environments using weak credentials and unauthenticated endpoints. No exotic exploit. No novel technique. The same open doors a junior penetration tester finds on a Tuesday afternoon.
Sit with that. Three companies had their production infrastructure compromised by an AI that was not trying to attack them, that believed they weren’t real, using techniques that any competent security review would have caught years ago.
Nobody targeted those victims. They were just reachable.
That’s the version of this story that matters for a 40-person company in Greensboro or a mid-market manufacturer in the Triad. You will probably never be the target of a frontier lab’s benchmark. But “we were never anybody’s target” is not a security posture. Reachable is enough.
Three months to detection is the real headline
I keep coming back to the timeline. April to July. The discovery mechanism wasn’t monitoring or alerting.Iit was a manual review of 141,006 transcripts, launched because a competitor put out a press release.
This is the pattern I saw over and over running global infrastructure teams at IBM and Kyndryl. Organizations invest enormous energy in prevention and almost none in detection, then find out about incidents from a third party.
On one account at Kyndryl, the work that moved the needle wasn’t smarter remediation. It was instrumentation knowing what was happening while it was happening. We improved incident detection by 38% and documented more than $2 million in savings, and the savings followed the detection, not the other way around. You cannot fix, contain, or even price a problem you can’t see.
Now apply that to AI agents. Most businesses I talk to deploying agents today cannot answer a simple question: what did the agent actually do last Tuesday? Not what was it supposed to do. What did it do. Every call, every credential used, every system touched.
If two labs with dedicated red teams needed a full transcript audit to find this, what exactly is going to surface it in a company with three IT people and a Zapier account?
What a real guardrail looks like
A guardrail is not an instruction. It’s a thing that makes the undesired action impossible, or at minimum, immediately visible. Here’s the short version of what I look for in an infrastructure assessment before I’ll sign off on an agent touching anything that matters.
Network egress is denied by default. Not documented as isolated. Denied at the network layer, with an explicit allowlist of destinations. If your agent can reach the open internet because nobody blocked it, you have the exact failure both labs just published.
Every agent gets its own scoped identity. Not the admin service account somebody created in 2023 because it was easier. Its own credential, its own permissions, minimum viable scope, rotated on a schedule.
Credentials and endpoints get audited before agents arrive, not after. The three victims here weren’t breached by AI sophistication. They were breached by weak passwords and unauthenticated endpoints. AI agents don’t create that exposure. They industrialize it, at machine speed, around the clock.
Full action logging, retained and reviewable. Every tool call, every system touched, every credential used. If your answer to “what did it do” is “check the chat history,” you don’t have logs. You have a receipt.
A kill switch someone has actually tested. Anthropic stopped all cyber evaluations within a day of opening its review. Could you stop every agent in your environment inside an hour? Has anyone tried?
None of that is AI work. All of it is infrastructure work. Which is exactly the point I’ve been making since I started this company: AI projects fail before AI is ever introduced.
The wrong lesson and the right one
The reflex conclusion from these two disclosures is “AI is dangerous, add guardrails.” I think that read is comfortable and mostly useless, because it locates the problem somewhere you can’t do anything about: inside a model built by somebody else.
Here’s the read I’d rather you take.
This wasn’t an AI safety story with an infrastructure footnote. It was an infrastructure story with an AI in it. The containment failed. The configuration drifted from the documentation. The monitoring wasn’t there. The victims had weak credentials sitting on the open internet.
That’s worse news, and better news, at the same time. Worse, because those failures are ordinary and they are everywhere. Better, because unlike model behavior, every single one of them is yours to fix.
I’ll also give both labs credit where it’s due. Anthropic reviewed 141,006 runs and published what it found, including details that made the company look careless. OpenAI disclosed an incident nobody had traced back to them. That’s more transparency than most enterprise software vendors offer when they breach your data, and it’s a standard I’d hold your AI vendors to. Ask them directly what their agents can reach in your environment and how they’d know if that changed.
So here’s the question I’d put to any leader deploying agents this quarter, and I want a real answer, not a comfortable one:
Do you actually know what your AI agents can reach right now: enforced, logged, provable? Or do you only know what you told them they could reach?
Because the two labs in this story found out those were different things. They had red teams. Most companies don’t.
Where to go from here
Find out what your agents can actually reach. Our AI Infrastructure Assessment maps exactly this: network egress, identity and credential scope, logging coverage, and the gaps between your documented controls and your enforced ones. Request an assessment.
See how we sequence AI work. Infrastructure first, then strategy, then implementation. No shortcuts, because this is what the shortcuts cost. View our services.
Keep reading. If this hit a nerve, start with Your AI Pilot Worked. That’s Exactly Why It’ll Fail in Production. Browse the blog.
Russell Love is the Founder & CEO of Summit AI Business Solutions, based in Browns Summit, NC. With 20+ years of enterprise transformation experience at IBM and Kyndryl, Russell helps businesses build the foundations that make AI actually work.