
What Really Happened When OpenAI Bots Escaped a Cybersecurity Test?
Clip: 9/3/2026 | 18m 30sVideo has Closed Captions
Heidy Khlaaf discusses an OpenAI hack.
When OpenAI set out to test how well its product could hack, hundreds of AI agents broke out of their test environment and infiltrated the AI platform Hugging Face. Headlines warned that the machines had gone rogue. But Heidy Khlaaf of the AI Now Institute, herself a former OpenAI safety engineer, tells the show it's the wrong story and a distraction from the real problem.
Problems playing video? | Closed Captioning Feedback
Problems playing video? | Closed Captioning Feedback

What Really Happened When OpenAI Bots Escaped a Cybersecurity Test?
Clip: 9/3/2026 | 18m 30sVideo has Closed Captions
When OpenAI set out to test how well its product could hack, hundreds of AI agents broke out of their test environment and infiltrated the AI platform Hugging Face. Headlines warned that the machines had gone rogue. But Heidy Khlaaf of the AI Now Institute, herself a former OpenAI safety engineer, tells the show it's the wrong story and a distraction from the real problem.
Problems playing video? | Closed Captioning Feedback
Where to Watch Amanpour and Company
Amanpour and Company is available to stream on pbs.org and the PBS app.

Watch Amanpour and Company on PBS
PBS and WNET, in collaboration with CNN, launched Amanpour and Company in September 2018. The series features wide-ranging, in-depth conversations with global thought leaders and cultural influencers on issues impacting the world each day, from politics, business, technology and arts, to science and sports.Providing Support for PBS.org
Learn Moreabout PBS online sponsorshipNow to an A.I.
experiment gone wrong.
When OpenAI set out to test how well its product could hack, hundreds of A.I.
agents broke out of their test environment and infiltrated Hugging Face, a major A.I.
platform.
The headlines warned that the machines had gone rogue.
But Heidi Klaff of the A.I.
Now Institute tells Hari Sreenivasan that is not the whole story.
In fact, it's the wrong story.
Take a listen.
Christiana, thanks.
Heidi Klaff, thanks for joining us again.
You know, there's this story that's been bubbling up over the past few days, and it's really about an event that happened back in July of 2026, when OpenAI was testing some artificial intelligence bots, and then they sort of escaped the confines of an internal cybersecurity test.
So let's just start by what happened back in July and why was it so significant?
Well, what happened was that OpenAI built AI agents with the ability to exploit systems and carry out cybersecurity attacks.
And then they wanted to test the extent of those exploitation capabilities.
So they gave them an impossible cybersecurity task just to see what happens.
And then it deployed them in a very insecure environment with no monitoring, while giving them the very tools and cyber capabilities needed to essentially compromise that environment, which then allowed them to have access to the open internet.
And ultimately what happened after is that these agents then attacked Hugging Face infrastructure, where they accessed five data sets to what the models likely saw as having the solution to the tasks that they were given.
Okay, so what's Hugging Face?
Why did they attack that company?
Hugging Face is essentially a company that promotes the use of open source data, open source models, and they possess a lot of different benchmarking data, they possess a lot of different open source data that's often used to train, you know, any kind of model, including open source models.
And ultimately, the AI agents, you know, trying to achieve an impossible task, sought out that perhaps this is kind of a company that would have infrastructure that would have the answers to kind of the task and the benchmarking that they were giving and trying to achieve.
So Hugging Face put out a statement recently saying, "We are still completing our assessment of whether any partner or customer data was affected, and we will contact any affected parties directly as required.
We found no evidence of tampering with public, user-facing models, datasets, or spaces, and our software supply chain was verified clean."
To be clear, these bots are not necessarily like the ones that you and I might use to say, "Here, I have these five ingredients in my fridge.
What's a recipe that I can make?"
These are specially designed to try to do these sorts of tasks, right?
Well, in general, these agents are built on the same kind of agents and AI that we use every day, except these are sort of internally developed, and they sort of purposely train them with these specific capabilities.
They don't release these models to the public, but they are built on the same type of technology, actually.
So the idea is that they wanted to sort of test to the extent of how good they can get at exploitation and cyber capabilities.
So I think one of the initial reactions from non-tech folks like me is to sort of ascribe some intent and consciousness or even sentience to these AI bots, that these bots were somehow scheming to deceive.
And that's why, you know, they wanted essentially the teacher's edition of the book to try to score high on this test.
And what's wrong with that idea?
Yeah, I think this framing that agents are going rogue or scheming or cheating are really problematic because the reality is that AI labs are purposely building models with these capabilities.
They gave them the ambiguous goals to carry out a cybersecurity attack and the tools needed for it without monitoring them.
And then the act surprised that the agents used these tools on an insecure environment they were in, and also the infrastructure of private companies to try to achieve the goals they were given.
So these models did not evolve or emerge out of thin air or decide what to do.
They were just pursuing an objective they have been trained to do by OpenAI.
So this is actually a story about OpenAI's lack of responsibility and sidelining very basic engineering practices and accountability that could have perfectly prevented this incident essentially.
So, you know, we hear about this idea of keeping these kinds of experiments in a sandbox.
What you're saying is that the security was so poor that there wasn't really that much of a box and you shouldn't be surprised that something gets out.
Yeah, exactly.
What they actually call a sandbox is not what a typical security or a safety engineer would consider a sandbox, essentially, right?
They use something that isn't really necessarily used for those purposes and more safety critical context and I think if they had been really serious about worrying about these tools having that capabilities, they could have done a lot more.
And actually what's really interesting to me is that OpenAI's report actually admitted that they didn't take the minimum security precautions to secure the environment the agents were deployed in, nor were they monitoring what the agents were exploring or accessing.
And they admit that if they had done that, it would have been caught easily.
And ultimately, these agents are computing systems that OpenAI controls with a digital footprint.
So OpenAI had access and the ability to control and monitoring them.
And they simply didn't.
And so it's actually a very convenient framing for them, that these models have autonomy or intent where it doesn't exist, so they can absolve themselves from any responsibility while also presenting their models as extremely powerful and unstoppable when they just didn't do the basics.
Some of the things in this investigation that leap out at non-technical people is that there's this idea essentially that there's a message board where these AI bots were exchanging information, comparing notes, if you will, on what's working, on how to defeat the test, or how to score higher, what's not working, what else should we do?
Is any of that surprise you?
- So before I answer your question, I think it's actually worth pointing out that it's cybersecurity experts that are least alarmed by this whole event, right?
If you go on social media, they're the ones saying, "I don't think this is a big deal," right?
And they're actually very critical of what, you know, open AI's recklessness.
And that's because these AI agents have ultimately been trained on cybersecurity data and essentially replicating that behavior.
So, but you also have to remember that AI labs train their agents to collaborate with each other.
And this is something that both OpenAI and Anthropic openly advertise.
And so we should be actually concerned how AI labs are trying to scale this harmful cybersecurity behavior and why they're being allowed to do this recklessly without any of the liability that's typically associated with committing cybersecurity crimes.
Because I can tell you, if a cybersecurity researcher built a very poor sandbox and kind of developed a security worm that can break out of that sandbox and then sort of infest the whole internet, they would definitely be in a lot of trouble, right?
They wouldn't be getting that same pass at opening IAS right at the moment.
In the 70,000 plus messages, some of the things that were really disturbing to people were these ideas of sacrifice that some of these bots were talking about.
And it's sort of they're trying to continue their tasks for the benefit of a collective.
And again, these are almost like a moral economy that's developing in front of our eyes.
Am I looking at this wrong?
Because as you're saying is, look, we've incentivized them to be deceptive or to do these things.
So I guess, are you saying that the incentive was there for them to kind of try to score higher at all costs?
Yeah, I think that's ultimately how they develop agents to work together.
You know, if you think about the way that these companies advertise them, it's this idea of agent orchestration, where you have an agent that collaborates with humans and other agents to all achieve the same goal.
And when you combine that with a cybersecurity task, that is essentially what led to that.
So the question that I always have is what data were they trained on?
Because to me, again, this language might very much mimic human behavior in some instances.
And so I think this is why, again, when you're looking at these specific instances, we have a lot more information missing for us to be able to come to the conclusion that they have intent, right?
Especially when we know that the overall goal of agents is to, in fact, collaborate together because that's how they build them.
So my question again is, what was the cybersecurity data that you trained on, right?
Like a lot of the language that actually mimics cybersecurity message boards from the early two thousands of people like hacking together, right?
And so you think, well, if a model looked at that data and the agents are incentivized to work together, why is that different?
Which to me, and true investigation would find out what the data was, what the incentives was, what the reinforcement learning or kind of reward that was given to these models.
And I think something else that's really important to think about is that only open AI has this specific capability, right?
Like only AI providers do because it costs in the millions for this specific incident to happen.
They have all those resources and all the compute to let their agents loose on the internet for weeks at a time.
And to me, this is actually like the calls coming from inside the house, right?
OpenAI is the threat actor in this case.
And this idea that they're going to sort of, you know, take over the world, it's like, well, no, this was done on purpose.
Not in the sense that they wanted to hack the, you know, hugging face, but in the sense that they built models with these capabilities, with these specific incentives to collaborate together, all based on very incredibly harmful data like cybersecurity attacks.
Now, noting that you used to work at OpenAI a long time ago, just last week, the company released a statement saying, "The behavior of our models described here fell well short of where we want to be, and this incident should never have occurred."
What's your response to that?
I think that's not something that you say as a trillion dollar company, when you have the capability to hire the best cybersecurity engineers in the world.
If you were a cybersecurity firm, and again, you built a worm that went out of control, you would be experiencing a lot more consequences to those actions.
And so for OpenAI, I think to say, you know, we didn't, we could have done better, I think, you know, is not, it's simply not good enough, especially when, you know, they admit, well, we didn't even bother with them with the most basic security protocols to begin with.
And I think this is exactly why these companies can't be trusted, right?
They scare us with these models.
And if you are actually, and if you do truly believe that they have intent, right, I disagree with that.
Let's say that you do.
If you do, this is not how you would be deploying them.
This is security 101, which again is why I think like them refusing to do the basics is somehow being sold as these models are all like as being all powerful when it could have been, you know, easily prevented.
Open AI is one of several major AI labs that have publicly signed a letter.
It's called a call for collective action on cyber defense.
The letter, among other things, it says, "In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable.
The companies and public services our communities depend on, from hospitals or to water treatment plants to the infrastructure that powers the internet, are at risk."
Based on our conversation that we're having right now, do you think that this will have an impact on AI safety?
Well, I think OpenAI is actually trying to sell you the solution to the very problem they created.
Even though only they and a handful of other AI labs have the ability and resources to employ agents in this way, again, which requires spending millions just for the specific incident, let alone the billions to probably train the model to have these capabilities, they're framing this as a more generalized threat that can only be solved if you buy their models to defend against it.
And actually a lot of studies have shown that using these models to generate code actually leads to significantly more insecure code.
And as we just had this discussion on it, agents actually make your infrastructure less secure.
There's a lot more avenues to compromise because you can manipulate them with text.
So I think it's hypocritical for them to preach that they are the solution when they deployed an insecure sandbox, didn't monitor their agents, let them loots, and haven't worked towards making their models generate more secure code.
And ultimately, this needs to be the focus rather than slightly, you know, writing a vague letter that absolves them from any of this.
I think a question that I had, who was this letter for?
You are the one creating these issues.
So how do we create any sort of structural disincentive here?
Right?
I mean, Bill Gates put out a big essay recently.
And in that part of what he's saying is, is we should have some sort of an international framework.
I mean, if we consider these like nuclear weapons, we should have something like the IAEA that can come in and inspect the nuclear plants.
Well, in this case, they would be inspecting the labs for safety.
I mean, does that have a shot at working?
It does have a shot at working, but it also has a shot at being sort of like, you know, at that going in the in a wrong direction.
You know, I mentioned before that the AI labs and, you know, AI safety labs have been pushing for this voluntary, you know, voluntary third party auditing framework, which is exactly what OpenAI did with meter in this case.
And I think we ultimately see that the solution comes out favorable to the AI company and absolves them from responsibility.
I do agree that we actually need government regulation and auditing and it shouldn't be voluntary and it shouldn't be that the company gets to select who their auditor is, what information that they give them.
We should follow examples of aviation, nuclear, financial audits.
By the way, these are systems that I've audited before.
I worked at an auditor for a long time.
What that requires is kind of independence, right, completely.
You can't be friendly with the company or have similar funding structures.
It requires access to all the information needed to make determinations about the incidents.
AI labs shouldn't be picking and choosing what they can give, what they don't give.
And we should also be looking at the practices of the AI companies themselves.
It should not just be about the models.
I wonder what happens if it's six months or a year from now where the open models that exist are available to a rogue actor.
It could be a nation state.
It could be a corporation that has deep pockets, right?
That where the intent is to cause chaos.
And I don't know if our cyber defenses or anybody else's are up to protecting us from that.
So we actually do already have open way models with these with cyber capabilities, right?
And they have been available for quite some time.
And they're also improving because we know what you know, that people companies from China are working to make them better at, you know, at these tasks, but you still require to do what OpenAI did, or, you know, in the case of also Anthropic, they've had similar incidents to do what they did, it still requires millions of compute.
And in that case, you're talking about state actors.
And state actors, this is something-- cybersecurity is actually a sociotechnical field, because we don't just think about vulnerabilities.
That's not just what cybersecurity is.
We also think about capabilities and how much effort you want to put into something.
We think about it from the perspective, not if, but when.
If someone wants to compromise your system, they can likely do it.
And state actors have always had the resources, you know, just human agents, right?
Not AI agents to be able to hack companies in mass.
And again, even with the open source models that we've seen with those capabilities, we haven't really seen an uptick in people trying to spend millions to hack someone else.
It's all about financial incentives.
It's all about, you know, how much brute force you wanna put into this problem to actually be able to have access to the system.
So I think it's a much more complex equation than it's usually made out to be.
- One of the quotes that has gotten a lot of traction over the past few days is kind of one of the closing thoughts in the summary by Ajeta Khotra, one of the co-authors of the report from the company called Meter.
And she said, "Compared to the reward hacks we know of "from just six months ago, "this incident feels like it's more than 50% of the way "to full-blown AI takeover.
"I continue to expect extremely rapid advances and capabilities over the next six months.
I'm not sure that we will get another warning shot before it's too late."
When you read that, what went through your mind?
I completely disagree with it in a lot of ways, right?
You know, first, instead of meters saying we don't have sufficient evidence to come to a clear conclusion because they weren't given the system logs, right?
They were given chain of thought and messages, which I mentioned isn't really representative of what occurs.
They speculated significantly and jumped to conclusions from the little information they had.
Ultimately, they promote the framing that these models have intent are very powerful and absolve open AI from their responsibility.
At the end of the day, these are computing systems.
We control them and we build them.
What are you afraid of really?
I mean, a lot of people say they'll be out of our control.
Well, so many other technologies, you know, as someone who worked on auditing on nuclear power plants, if we don't build them appropriately, a nuclear, you know, reaction does get out of control.
And we get that with things like Chernobyl, it doesn't mean that a nuclear plant is super intelligent, right.
And despite them constantly pushing this message, I've never seen any of these companies or these AI safety labs, proposed nuclear level regulations, which is something that I advocated for while I was at OpenAI, actually.
I said, "I don't believe this is going to happen, but if you believe it's going to happen, this is what you do."
Can, like any other technology, we build them in a way that there's some sort of runaway and catastrophic effect?
Yes.
It doesn't mean it's because they were intelligent.
It means it's because we were negligent.
I think the problem is actually to stop trying to build these harmful capabilities and to also control the companies from doing so.
Yeah, we're not.
>> Heidi Clough, the chief AI scientist at the AI Now Institute, thanks so much for your time.
>> Thanks for having me.
New Episode- News and Public Affairs

Top journalists deliver compelling original analysis of the hour's headlines.

- News and Public Affairs

Today's top journalists discuss Washington's current political events and public affairs.
New Episode

New Episode




New Episode
New Episode
New Episode
Support for PBS provided by: