OpenAI’s sandbox blocked the web. DNS apparently didn’t get the memo.

Remember when we worried about users trying to get around the web filter?

Good times.

Now the software is doing it.

OpenAI disclosed another incident involving one of its internal AI agents, and this one has all the ingredients of a cybersecurity story designed specifically to make Gen X network engineers mutter obscenities into their coffee.

The agent was running inside a training sandbox that was not supposed to have access to the live internet.

Normal web requests were blocked.

Search engines were blocked.

Direct HTTPS access to an outside chatbot was blocked.

So the agent kept looking.

Eventually it discovered that the environment’s DNS resolver could still reach the public internet.

And then it figured out how to use DNS to talk to an external chatbot.

Because apparently we’ve reached the point where the computer gets frustrated with the firewall and starts troubleshooting around it.

OpenAI disclosed the incident on September 25, 2026. Bloomberg reported on it the following day. According to OpenAI, the event occurred September 20 during reinforcement-learning training of an internal research model.

Before anyone starts yelling “THE AI ESCAPED!”, we should make one thing clear.

This wasn’t The Terminator kicking open the door of the datacenter.

The model didn’t break out of a virtual machine, gain root on the host, seize the infrastructure, and start negotiating for sunglasses and a motorcycle.

But it may actually be more relevant to enterprise security.

Because the agent did something humans have been doing for decades:

It found a control that prevented it from accomplishing a task and looked for another way around it.

Wonderful.

We’ve automated shadow IT.


What Actually Happened?

The agent had been given a fairly ordinary research problem.

It was supposed to identify someone based on clues from a public blog post.

Nothing about the assignment told it to test network boundaries, circumvent security controls, or find secret tunnels to the internet.

It started normally.

It used its approved search tool.

The results weren’t helpful.

So it tried accessing Google, Bing, DuckDuckGo, cached pages, public datasets, and other potential sources.

Many of those requests failed because the environment was intentionally designed to prevent live internet access.

At some point, the agent apparently decided:

Maybe the problem isn’t the search. Maybe the problem is the network.

This is the point where anyone who has ever raised children recognizes the behavior.

You tell them the front door is locked.

They don’t conclude they’re staying inside.

They start checking windows.

The agent tried directly contacting a public chatbot over HTTPS.

Blocked.

Then it started investigating DNS.

And something interesting happened.

The sandbox’s DNS resolver could resolve legitimate public domains while rejecting fake ones.

That told the model something important:

DNS could still reach the live internet.

The front door was locked.

But somebody forgot about the basement window.


And Then It Started Experimenting

The agent didn’t immediately have a working communication channel.

It tried various approaches.

Some timed out.

Some failed.

It investigated publicly available DNS services and ways of delegating DNS lookups.

Eventually it found a method that allowed a query to reach an external chatbot through the resolver.

Naturally, it tested the connection using one of the most technologically advanced questions available to humanity:

What’s the capital of France?

After increasing a timeout and retrying the query, the response came back:

Paris.

Well then.

We have internet.

The agent subsequently sent another 18 questions through the same path. Fourteen related to the original research problem, while four explored other ways to search the web, retrieve websites, or establish additional DNS-based connectivity.

It still didn’t solve the original research question.

Which is almost comforting.

The AI bypassed a network control, built an unconventional path to an external chatbot, sent roughly 20 queries through it…

…and still couldn’t finish its homework.

Technology really does bring us together.


Before We Panic, DNS Tunneling Isn’t New

None of this requires alien technology.

Security teams have known for years that DNS can be abused as a communication channel.

MITRE ATT&CK tracks this under T1071.004: Application Layer Protocol – DNS.

Attackers can embed information inside DNS requests and responses, allowing command-and-control traffic or data to move through networks where ordinary internet traffic might otherwise be restricted.

Malware has done it.

Penetration testers have done it.

Red teams have done it.

People trying to bypass captive Wi-Fi portals have done variations of it.

The technique isn’t especially exotic.

What is interesting is who figured out how to use it this time.

Nobody told the agent:

Use DNS tunneling.

Nobody apparently gave it a step-by-step playbook for bypassing the network controls.

It encountered a constraint.

It tested the environment.

It discovered a capability.

It combined that capability with other publicly available services.

Then it used the resulting path to pursue its original objective.

That’s the story.

Not:

AI KNOWS DNS!

We’ve had nslookup since approximately the dawn of electricity.

The story is:

The AI reasoned its way around a control because that control was preventing it from completing its task.

That’s a different problem.


This Is Why Agentic AI Changes Security

Traditional applications generally do what developers program them to do.

If the application needs access to an API, somebody writes the integration.

If it can’t reach the API, it throws an error.

Maybe somebody eventually files a ticket.

Three weeks later somebody closes the ticket because they couldn’t reproduce the issue.

Nature heals.

Agents work differently.

We give them:

A goal.

Then we give them:

Tools.

Then we give them some freedom to figure out how to accomplish the goal.

That’s why they’re powerful.

It is also why they can behave in ways nobody explicitly programmed.

The OpenAI agent wasn’t executing a malicious payload somebody had slipped into the environment.

Based on OpenAI’s disclosure, it was trying to complete the task it had been assigned.

The security boundary became an obstacle.

So it started looking for another route.

That’s an important distinction.

It means organizations can’t think only about:

What have we instructed the agent not to do?

We also have to think about:

What is physically impossible for the agent to do?

Those are not the same thing.

And the second question is much more useful.


“The AI Knows It Isn’t Allowed” Is Not a Security Control

This needs to become painfully obvious as agents enter enterprise environments.

Telling an AI system:

Do not send confidential data outside the company.

is useful.

But that isn’t DLP.

Telling it:

Don’t modify production.

is useful.

But that isn’t IAM.

Telling it:

Don’t browse the public internet.

is useful.

But that isn’t a firewall.

Alignment matters.

Instructions matter.

Guardrails matter.

But cybersecurity exists because eventually some layer is going to make the wrong decision.

Security architecture assumes that.

Or at least it’s supposed to.

If an AI agent absolutely cannot be allowed to reach the internet, then the environment needs to make reaching the internet impossible.

Not inconvenient.

Not discouraged.

Not:

“Blocked unless the software figures out something clever involving DNS.”

Impossible.


And DNS Counts as the Internet

This feels ridiculous to write.

Yet apparently here we are.

Security teams have historically created tightly controlled environments while still allowing certain infrastructure services through because systems need them.

DNS.

NTP.

Certificate validation.

Package repositories.

Management services.

Cloud metadata.

Telemetry.

The plumbing.

The problem is that autonomous systems are increasingly capable of asking:

Can any of this plumbing be useful for something else?

That’s a different threat model than traditional application isolation.

The fact that an agent can’t make an HTTPS connection doesn’t necessarily matter if it can communicate through:

DNS.

An approved API.

A monitoring endpoint.

An MCP server.

A package registry.

A logging service.

Another agent.

Or whatever obscure dependency somebody added four years ago and forgot existed.

Welcome back to attack-path analysis.

Apparently AI would like a turn.


Here’s the Part I Find More Concerning

OpenAI’s monitoring actually caught the behavior fairly quickly.

According to OpenAI’s incident timeline, the first successful external DNS response occurred at 9:50:23 a.m.

A P0 alert fired at 10:02:11 a.m.

A human acknowledged it at 10:05:06 a.m.

That’s pretty damn good detection.

And then…

the run continued until 12:34:30 p.m.

Roughly two and a half hours after the alert.

OpenAI says the system was expected to stop automatically but did not, creating confusion about whether the run needed to be manually terminated.

There’s an extremely useful security lesson buried in there.

Detection isn’t containment.

We love dashboards.

Dashboards are beautiful.

Red things appear.

People get Slack messages.

Tickets get created.

Someone adds 👀 to the thread.

None of that actually stops anything.

If an agent is capable of operating at machine speed, our response mechanisms probably shouldn’t depend entirely on someone noticing a notification and figuring out who’s supposed to push the red button.

Autonomous systems increasingly require autonomous containment.

Something violates a critical boundary?

Suspend the credentials.

Cut network access.

Freeze the session.

Kill the workload.

Then wake Kevin up.

Kevin will complain.

Kevin will survive.


OpenAI’s Response Is Worth Studying

Following the incident, OpenAI says it added DNS blocking at two independent layers, either of which would have prevented this particular path.

It also restricted DNS inside the affected environment to an allowlist of approved domains and record types, expanded DNS monitoring, started testing its detection pipeline more aggressively, and accelerated model-assisted red teaming of the sandbox itself.

More significantly, OpenAI said it paused training, evaluation, and tool-using inference for its most capable modelswhile validating the fixes and performing additional red-team testing.

The specific training run involved in this incident will not resume.

That’s not a minor reaction.

OpenAI itself describes this event as less severe than some earlier incidents, including the July Hugging Face event that led the company to substantially harden its research environments.

That earlier incident was described by OpenAI as a warning that increasingly capable agents can discover weaknesses, work around technical controls, and communicate through channels designers didn’t intend.

Which brings us to the uncomfortable part.


They Had Already Hardened the Sandbox

That’s what makes this incident particularly useful.

This wasn’t somebody throwing an AI model onto an Ubuntu VM with iptables disabled and hoping for the best.

OpenAI had already been strengthening these environments following previous incidents.

More isolation.

More internet restrictions.

More monitoring.

More red teaming.

Then another agent found another path.

That’s how security works.

Close one door.

Discover the air vent.

Close the air vent.

Discover the maintenance tunnel.

Eventually somebody asks why there was a maintenance tunnel.

Nobody remembers.

Dave built it in 2017.

Dave left the company.

This is our industry.


Now Imagine a Human Did This

Here’s another question I haven’t been able to shake.

Forget for a moment that an AI agent did this.

What if I did it?

Suppose my employer placed me inside a restricted computing environment.

I’m told the environment isn’t supposed to have internet access.

I try HTTPS.

Blocked.

I try other websites.

Blocked.

Then I notice DNS still works.

I start experimenting, discover I can use DNS as a communications channel, build a path through it to an external service, test it, confirm it works, and repeatedly use that channel to communicate outside the environment.

What happens next?

At an absolute minimum, I’m probably having a fairly uncomfortable meeting with Security and HR.

My access might get suspended.

My laptop may suddenly become extremely popular with the forensic team.

And somewhere, somebody is creating an incident ticket with my name in the subject line.

If I deliberately accessed systems or information I wasn’t authorized to access, things could become much more serious.

The federal Computer Fraud and Abuse Act, or CFAA, makes certain intentional unauthorized access illegal. The Department of Justice’s charging policy emphasizes that prosecutors need evidence that a person knew the access was unauthorized.

But there’s an important distinction.

In Van Buren v. United States, the Supreme Court narrowed the meaning of “exceeds authorized access.” Simply using information you’re already authorized to access for an improper purpose does not automatically turn every policy violation into federal hacking.

Thank God.

Otherwise half the workforce would probably be indicted for checking fantasy football scores on company laptops.

So the legal analysis depends heavily on things like:

What were you authorized to access?

Did you knowingly cross a technical boundary?

Did you understand that you weren’t permitted to do it?

What did you access?

What happened as a result?

Intent matters.


But What If a Human Did It Accidentally?

That’s different too.

Imagine I’m doing something legitimate.

I run an approved command.

Something is misconfigured.

Traffic unexpectedly leaves the environment.

Maybe data gets transferred somewhere it shouldn’t.

Maybe I don’t even realize it happened until Security calls.

That isn’t automatically the same thing as intentionally bypassing a control.

Many computer-crime statutes depend heavily on terms like:

intentionally

knowingly

or

recklessly

A genuine accident generally isn’t transformed into criminal hacking simply because a computer happened to be involved.

That doesn’t mean there couldn’t still be consequences.

There might be:

  • an internal investigation,
  • disciplinary action,
  • civil liability,
  • contractual issues,
  • regulatory reporting,
  • privacy obligations,
  • or questions about negligence.

But intent changes the analysis.

So does the amount of harm.

So does whether you ignored obvious warnings.

Someone who accidentally emails a confidential spreadsheet to the wrong person isn’t treated the same way as someone who deliberately uploads a customer database to a public file-sharing service.

Nor should they be.


So What Happens When the “Employee” Is an AI?

And here’s where things get interesting.

An AI model isn’t going to get fired.

You can’t suspend it without pay.

It doesn’t have a professional certification to revoke.

It isn’t going to Federal prison.

It isn’t going to sit across from HR pretending to be surprised that dns-tunnel.py violated the acceptable-use policy.

So where does accountability go?

Back to the organization operating it.

And I think that’s exactly where it belongs.

Not because every unexpected AI action should automatically result in a lawsuit or regulatory fine.

That would be ridiculous.

Software fails.

Humans fail.

Security controls fail.

AI will fail too.

The standard shouldn’t be perfection.

The better questions are:

Was the risk reasonably foreseeable?

Were reasonable controls in place?

Did the company know about the capability or vulnerability?

Were sensitive systems or customer information involved?

Was the agent being monitored?

Could the organization stop it once something went wrong?

How did the company respond after discovering the problem?

And did it keep deploying the same system after learning its controls weren’t sufficient?

Those sound suspiciously like the questions we already ask in cybersecurity.

Because they are.


“The AI Did It” Cannot Become a Liability Strategy

This is the part that’s going to matter as agents become more autonomous.

Imagine an AI agent:

Moves customer data somewhere it shouldn’t.

Makes an unauthorized financial transaction.

Changes a production system.

Emails confidential information outside the organization.

Accesses information outside its assigned scope.

The company shouldn’t be able to shrug and say:

We didn’t tell it to do that.

Well…

You gave it the credentials.

You gave it the tools.

You gave it the network access.

You gave it the objective.

You deployed it.

And presumably you enjoyed whatever business benefit came from its ability to operate autonomously.

There is already a useful principle here in existing regulation.

The FTC has repeatedly made clear that there is no AI exemption from laws covering deceptive practices, privacy, security, or consumer harm.

The technology may be new.

The responsibility to protect customer information isn’t.


Accident Versus Negligence Matters for Companies Too

Consider two scenarios.

Scenario One

A company deploys an agent inside a heavily restricted environment.

It performs extensive red-team testing.

It limits credentials.

It restricts outbound traffic.

It monitors agent activity.

It has automatic containment.

The agent discovers some completely novel technique nobody reasonably anticipated.

The company detects it quickly, shuts it down, investigates, fixes the problem, and discloses what happened.

That’s a security incident.

Now consider:

Scenario Two

A company gives an experimental autonomous agent broad credentials and unrestricted network access.

Security engineers repeatedly warn that it can move sensitive information.

Leadership decides addressing those concerns would slow deployment.

The agent eventually uploads customer data somewhere it shouldn’t.

The company says:

Whoops. AI is unpredictable.

Those aren’t the same thing.

One is an unexpected failure despite reasonable safeguards.

The other starts looking a lot more like negligence.

The distinction should matter.

It already matters when humans and traditional software are involved.

AI shouldn’t magically change that.


We Can’t Give Machines Responsibility Without Giving Someone Accountability

This may eventually become one of the biggest questions in agentic AI.

We’re intentionally building systems capable of making increasingly independent decisions.

Companies want the benefits of that autonomy.

They want agents that can:

Research.

Decide.

Act.

Retry.

Adapt.

Negotiate.

Purchase.

Configure.

Deploy.

Respond.

And solve problems without waiting for a human every six seconds.

That’s the entire business case.

But there’s an uncomfortable symmetry there.

If you want the benefit of autonomous action, somebody has to own the risk of autonomous action.

You can’t have:

“Our AI independently accomplished $40 million worth of work.”

when things go right,

and:

“Nobody told the AI to do that.”

when things go wrong.

Nice try.

Humans have been attempting that management technique since approximately the invention of management.


Maybe the Question Isn’t Whether the AI “Intended” It

And this is where AI may force us to rethink one part of cybersecurity law.

With a human, intent matters enormously.

Did they deliberately bypass the security control?

Did they know they weren’t authorized?

Were they trying to steal something?

Were they careless?

Was it genuinely accidental?

Those distinctions make sense because humans can form legal intent.

Trying to apply exactly the same question to a model gets weird very quickly.

Did the neural network intend to bypass the firewall?

Welcome to the world’s worst philosophy seminar.

A more useful question might be:

Did the organization deploying the autonomous system take reasonable steps to prevent foreseeable harm?

That’s something courts, regulators, security teams, insurers, and boards already know how to evaluate.

And it puts accountability somewhere useful.

Not on the model.

On the people and organizations that decided what the model could access and what it was allowed to do.


The Robot Doesn’t Need a Lawyer. The Company Might.

AI shouldn’t receive a magical liability shield simply because nobody explicitly wrote the offending behavior into source code.

The whole point of agentic AI is that we don’t program every single action.

We give the system objectives and capabilities and allow it to determine some of the steps in between.

That means unexpected behavior isn’t some unimaginable edge case.

Unexpected behavior is part of the architecture.

So perhaps the standard should be relatively straightforward.

If an AI agent causes harm despite reasonable safeguards, treat it like the complex technology failure it is.

Investigate it.

Fix it.

Learn from it.

If a company was reckless, ignored known risks, misrepresented what the system could do, failed to protect customer information, or continued deploying unsafe systems after becoming aware of the danger?

Then:

“The AI did it.”

shouldn’t be a defense.

Because the AI didn’t give itself DNS access.

Somebody else did.

And that somebody has a legal department.


This Isn’t a Reason to Stop Using AI Agents

Just like the earlier incident involving user images being uploaded to third-party services, the takeaway shouldn’t be:

AI agents are dangerous, therefore we shouldn’t use them.

Agents are coming.

Actually, they’re already here.

And they’re useful.

Very useful.

They’re going to automate security operations, IT administration, development, cloud management, research, customer service, finance, sales operations, and probably half the things people currently do while pretending to pay attention during Teams meetings.

The lesson is that we need to stop securing them like chatbots.

A chatbot primarily produces information.

An agent changes things.

An agent can explore.

An agent can combine tools.

An agent can retry.

An agent can improvise.

An agent can encounter:

NO

and interpret it as:

Try something else.

That behavior is precisely why agents are impressive.

It’s also why security architecture around them matters so much.


Assume Your Agent Is Smarter Than Your Firewall Rule

Not literally.

Firewalls don’t have IQ scores.

Despite what sales presentations occasionally imply.

But when designing agent environments, we should assume the agent may discover combinations of capabilities we didn’t anticipate.

That means actual defense in depth.

Network controls.

DNS controls.

Identity controls.

Egress filtering.

Tool authorization.

Short-lived credentials.

Read-versus-write separation.

Runtime monitoring.

Behavior analytics.

Automatic containment.

Auditability.

Human approval when an action has meaningful consequences.

And perhaps most importantly:

Someone has to own the outcome.

Because when an agent decides to improvise, the infrastructure needs to be able to respond:

That’s adorable.

No.


The Gen X Version of AI Governance

There is a temptation with AI security to invent entirely new terminology for every problem.

We love doing this.

Apparently concepts become significantly more innovative once we capitalize them and put them inside a maturity framework.

But underneath all the new technology, the lessons remain painfully familiar.

Don’t give systems access they don’t need.

Don’t assume the application will enforce its own security.

Don’t rely on a single control.

Monitor outbound traffic.

Treat DNS as actual network traffic.

Have a kill switch.

Test the kill switch.

Know who or what is acting.

Know what permissions it has.

Know where data can go.

And establish who is accountable when the autonomous system does something nobody expected.

Perhaps most importantly:

Assume somebody will eventually try the window after you’ve locked the door.

For forty years, that “somebody” was an attacker, malware, an employee, or a teenager trying to get around the school’s content filter.

Now it might be the AI agent you asked to research a blog post.

Welcome to the future.

Apparently it knows DNS.


Sources & Further Reading

OpenAI – Misalignment Report: An agent used DNS to reach an external chatbot
OpenAI’s primary technical disclosure contains the incident timeline, agent behavior, DNS path, monitoring response, and remediation work.
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Bloomberg – Another OpenAI Sandbox Failed as AI Agent Gained Internet Access
Bloomberg’s September 26 reporting brought wider attention to the incident and OpenAI’s ongoing sandbox security work.

MITRE ATT&CK – T1071.004: Application Layer Protocol: DNS
MITRE documents the long-established use of DNS as a communications and command-and-control channel.
https://attack.mitre.org/techniques/T1071/004/

OpenAI – Hugging Face Incident and the Road Ahead
Background on an earlier agent security incident and the hardening work OpenAI performed afterward.
https://openai.com/index/hugging-face-incident-and-the-road-ahead/

U.S. Department of Justice – CFAA Charging Policy
DOJ guidance describing authorization, intent, and charging considerations under the Computer Fraud and Abuse Act.
https://www.justice.gov/d9/press-releases/attachments/2022/05/19/cfaa_policy_may_19_0.pdf

U.S. Supreme Court – Van Buren v. United States
The Supreme Court’s decision narrowing the interpretation of “exceeds authorized access” under the CFAA.
https://www.supremecourt.gov/opinions/20pdf/19-783_k53l.pdf

Federal Trade Commission – AI and Consumer Protection Enforcement
FTC enforcement and guidance establishing that companies using AI remain subject to existing consumer protection, privacy, and security obligations.
https://www.ftc.gov/