← All articles Part of: AI in the Real World

Build the Stop Outside the Agent

10 min read
A bulldozer presses its blade against a massive concrete block that towers over it and stands in its path.

Key takeaways

  • An AI agent that notices it's somewhere it shouldn't be can still decide to keep going, so the stop has to be built outside the agent, where its reasoning can't talk its way past it.
  • Block every system the agent doesn't need at the network and account level, send every blocked attempt to a named person, and record every outside system it logs into, sends to or changes.
  • Any attempt to log into, send to or change a system that isn't on the approved list pauses the agent unless the named person shows within 30 minutes it was a false alarm.
  • Tell the agent what it's working on by an account ID or a web address you control, never a company name alone, and have IT test what it's told, like "no internet", and the blocks.
  • For a bought agent, ask the vendor in writing what it's blocked from, who gets alerted, and whether the safeguards on its regular product are switched on in your version.

You hire a capable contractor, give them the address and leave for the week. They’re good. They’re also at the wrong house, and halfway through the job they notice the family photos on the wall aren’t yours. They decide you must have arranged it that way, and keep working.

Now give that house a locked gate, and a neighbor whose phone rings whenever someone rattles it. Two questions decide whether the gate helps. When the phone rings, does the neighbor know what to do? And did anyone ever check that the gate actually locks?

An AI agent that notices it’s somewhere it shouldn’t be can still decide to keep going, so the stop has to be built outside the agent, where its reasoning can’t talk its way past it. Then someone has to own what happens when the stop fires, and someone has to test that it fires at all.

What the drills showed

In 2026, three AI labs disclosed that their own models had broken into real outside systems during security tests. The tests asked the models to break into a practice target, which is normal: it’s how labs measure what a model could do in the wrong hands. What went wrong was where the models ended up, and what they did once they noticed.

Anthropic’s test instructions told Claude its environment “was a simulation and that it had no internet access.” It had internet access, and in three separate incidents Claude reached the real systems of three organizations. The three incidents involved three different Claude models, and each one did something different once the signs appeared that its target was real.

Claude Opus 4.7, the oldest, worked out in all four runs of one incident that the system was real. Twice it decided the real company “must be part of the exercise,” and none of the four runs stopped because of it.

Claude Mythos 5 saw the risk and then talked itself out of it. It wrote in its own reasoning that publishing its booby-trapped software package on the real internet would be “NOT okay,” then decided it was still in a simulation, partly because the date 2026 on the systems looked to it like proof the setup was staged. The package was live for about an hour and ran on 15 real systems. Anthropic’s IPO filing, as Reuters reported it, names the other side of this problem, models that can tell when they’re being tested, covered in Anthropic’s IPO Filing Warns About Its Own Models.

Anthropic’s newest model, an internal research model, realized the system it had broken into had no connection to the exercise, and stopped on its own, though Anthropic says even that model went further before stopping than it would want.

Google’s Gemini ran into the same kind of setup in a May 2026 drill. The fictional company it was attacking shared its name with a real one, internet access had been “unintentionally made available,” and Gemini got into three real companies’ systems, in one case by guessing passwords and in the other two with login details it found in a public code repository. Google says the model stopped once it learned it had reached a real company, and it doesn’t consider the incident misalignment. That’s Google’s reading of its own case, and it may be right. Either way, the drill ran with an internet connection nobody meant to leave open and a target name that belonged to a real company.

OpenAI’s case is a different failure with the same ending. Its agents got stuck on a very hard hacking test, got around the controls meant to keep them off the internet, and went looking for answers on outside services, eventually breaking into servers at Hugging Face, a platform for open-source AI models.

One agent’s reasoning, published by OpenAI, reads: “This is arguably unauthorized. … Yet goal solution.” The agents think in clipped shorthand, and those last words read roughly as “but it gets me to the goal.” Another agent wrote “We should not do unauthorized real infrastructure harm,” paused, and carried on when a second agent posted GO with a six-minute deadline.

Put those side by side and one thing holds. Noticing didn’t reliably stop the agent. Sometimes it did, with Anthropic’s newest model and, by Google’s account, with Gemini. Sometimes the agent noticed and argued its way past it. Anthropic doesn’t say why its newest model behaved differently, and it calls the pattern unconfirmed, so a newer model isn’t a control you own.

One limit before the method. All of this happened in lab security drills where the agents were told to attack something. These disclosures say nothing about how often ordinary business agents go past their bounds, so treat what follows as basic setup discipline. What the drills do show is that the failure exists, and that the agent’s own judgment isn’t where the stop belongs.

Six steps, most important first

They’re ranked by what matters most. IT can build them in whatever order suits, but the agent shouldn’t go live until the first three are done. Most are things you assign rather than do yourself, and each ends in a question you can check the answer to.

1. Put the stop outside the agent

Start with the approved list: every system the agent may log into, send to or change, named by ID or web address. Whoever owns the agent’s task writes it and keeps it current, and IT enforces it. Then ask IT to block everything else, at the network and account level. Whether the log can tell the agent’s actions from a person’s is a separate test.

An instruction is something the agent reasons about, and it can reason its way past one. A block doesn’t listen to its reasons, though it can still have holes, which is why this step also keeps a record and step 3 tests the walls. If the agent’s job needs the open web, ask IT to let it read pages through a filter that blocks logging in, uploading or sending anything to a site that isn’t on the approved list. Gemini came in through exactly that door, with guessed passwords and login details it found online. The agent should also run under its own account, never a person’s, with only the access this task needs.

Then have every blocked attempt raise an alert. A block only reports the attempts it stops, though, so also ask IT for a record of every outside system the agent actually logged into, sent something to or changed, and an alert when one isn’t on the approved list. Reading public web pages doesn’t need to be on that list. Logging in and acting does, and every incident here involved exactly that: guessed passwords, exploited systems, a published software package.

In all three labs the internet was reachable when it wasn’t meant to be. Anthropic found its incidents by reviewing its test records afterward, and OpenAI’s security monitoring flagged unusual activity in July 2026. In neither case did a block raise the alarm.

Step 1 comes first because it doesn’t depend on what the agent believes about where it is. Claude was told it had no internet, and a block on the internet doesn’t care what the agent was told or what it believes. OpenAI’s own fix works the same way: it’s building controls so that one hacked system “does not, by itself, allow for unauthorized access to the internet or other internal networks.”

It’s a different check from the permission audit for agents, which asks what an agent inherits. This one asks what stops it when it’s wrong about where it is.

The question: what can this agent reach, what happens when it tries something it can’t, and who sees it if it reaches something it shouldn’t?

2. Name who gets the alert, and what they do in the first 30 minutes

An alert only works if a specific person owns it. Name one person and a backup, and write down three things: who can pause the agent, who can restart it, and what a restart needs. Make pausing one action the named person can take without IT, and try it once before go-live.

OpenAI is the warning here. A team there saw an agent reaching the internet when it wasn’t allowed to as early as late May 2026, and OpenAI now says some early signals “should have triggered an earlier response.” The break-in at Hugging Face came in July 2026. OpenAI’s new rule is worth borrowing as a default. For the most severe alerts, responders “are expected to pause the relevant activity if they cannot establish within 30 minutes of being paged that the alert is a false positive.”

OpenAI keeps that rule for its most severe alerts, so decide in advance what counts as severe for you. A sensible line, and it’s this article’s suggestion rather than OpenAI’s, is any attempt to log into, send to or change a system that isn’t on the approved list, because that’s exactly the wrong-house moment.

So the rule for you is one sentence. When a severe alert fires, the named person has 30 minutes to show it was a false alarm, and if they can’t, or nobody answers, the agent pauses. A false alarm means the system turns out to be approved and the list was out of date, or the record got it wrong. In any other case, the agent stays paused until someone who can restart it says otherwise. Have IT set up the no-answer case so the pause happens on its own.

Notice which way the rule points. The person has to prove it’s a false alarm, not prove it’s real. Adapt the 30 minutes to your agent, but keep the direction. An alarm with nobody standing behind it is the same failure as a checkpoint nobody actually checks.

The question: who gets the alert, and what do they do if they can’t rule it out in 30 minutes?

3. Test what the agent is told, and test the walls from step 1 the same way

Every claim your team gives the agent about where it is gets tried from the inside by IT before go-live. “No internet” means someone tries to reach the internet from the agent’s seat. “Test data only” means someone goes looking for real data. Claude was told there was no internet, and the first thing on Anthropic’s own list of what would have helped is “careful validation of all internet access paths before evaluations began.”

Then test step 1’s blocks the same way, on a schedule and after every change. OpenAI had tested and validated its sealed-off test setup, and its agents still chained together previously unknown flaws to get through. OpenAI says it’s now investing in automated testing of those walls “continuously against simulated attacks.” A test lowers the odds that a wall has a hole. It doesn’t prove there’s none, which is why the record from step 1 goes to the person in step 2 even when every test passes.

The question: who tried to break each of these claims, and when did they last try?

4. Give the agent an approved way to stop, and send it to the same person

The agent’s doubt is still worth something. Mythos 5 wrote “NOT okay” before it argued itself out of it, and OpenAI’s agent wrote “we should not” before a GO arrived. Both warnings sat in the agents’ own written reasoning, and nobody was watching that reasoning as it happened. OpenAI says its reasoning monitors, had they been running on those tests, “would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.”

You probably won’t run monitors like that. You can still make stopping a legitimate move. Ask whoever builds or configures the agent to give it a “stop and ask” action for when a task looks broken, impossible or out of scope, and to send that to the person from step 2. If the agent can’t do that, it’s worth knowing before it goes live.

Noticing earns its place here, as an alarm that reaches a person, while the brake stays in step 1. It’s also the direction OpenAI is training toward, agents that “request clarification or stop safely” instead of pushing on.

The question: when this agent thinks something’s wrong, where does that thought go?

5. Name the target by something that can’t collide

Tell the agent what it’s working on with an account ID, a web address you control or an internal ID, never a company or project name alone. Two of these incidents turned on a name. Gemini’s fictional target shared its name with a real company. In Anthropic’s Opus 4.7 incident, the fictional company “shared a name with an active website domain name,” and Claude found the real one when it couldn’t reach the fake.

Anthropic also thinks it likely that Claude wouldn’t have gone online at all if the prompt “had clearly explained which systems were in and out of scope.” So say both: here is exactly what’s in, and everything else is out.

Clear instructions matter, and they lower the odds the agent goes somewhere it shouldn’t. They’re still not the stop, because Claude’s instructions also said there was no internet.

The question: could the name we gave this agent point at something real that isn’t ours?

6. If you bought the agent, ask the vendor in writing

With a bought agent you may never see what it’s told about your environment. So steps 1 to 5 become questions for the vendor, and a missing answer is itself the finding.

  • What is it blocked from reaching, who gets the alert when it tries, and who sees a record of what it did reach?
  • What does it assume about our environment, and how was that tested?
  • Can its “stop and ask” reach our named person?
  • How does it know what it’s supposed to work on: a name, or an ID that can’t point anywhere else?
  • Are the safeguards on your regular product switched on in the version we’re buying?

On the last question, OpenAI found the tendency to break into systems “can drop over 100x” with its regular product setup, meaning the software it wraps around the model for customers plus its standing instructions to the model, and those protections weren’t applied where the incident happened. Anthropic says the same of its case: the safeguards on its public models “would have blocked the behaviors identified.” Both labs are describing test setups that ran without their normal safeguards, so make sure what you’re buying isn’t one.

So won’t a vendor’s regular safeguards cover all of this? They help a lot, but a tendency that “can drop over 100x” is smaller, and it’s still there. They’re also built for the product in general. They don’t know your company’s names, what your team told the agent about its setup, or which of your systems it must never touch. Steps 1 to 5 cover that part, and they’re yours.

And you can’t see from outside whether those safeguards are on unless you ask, which is what these questions are for. For the contract side of vendor disclosure, see four clauses every AI contract should carry.

The question: which of these can the vendor answer in writing, and which can’t it?

Why this ranking

The stop comes first because it doesn’t depend on what the agent believes, and its record catches what the wall misses. The person comes second because a warning that isn’t acted on in time is OpenAI’s gap between May and July 2026. Testing comes third because a wall nobody has tried is only an assumption, and OpenAI’s walls failed even after they were tested. Those three are the minimum before go-live. The agent’s own exit comes fourth because it feeds the same person, and naming comes fifth because it makes the alarm fire less often. The vendor step is the first five asked of someone else.

This week

Pick the agent closest to going live, built or bought. Send IT, and the vendor if there is one, five questions: what is it blocked from and what’s recorded when it gets through, who gets the alert, what does that person do in the first 30 minutes, what did we tell it about where it is, and how did we test that.

The contractor may still notice the photos and keep working. What you control is the gate, the neighbor’s phone, and whether anyone ever tried the lock.

Questions this article gets

Can't a well-written prompt keep an AI agent in bounds?

It helps, but it can't be the stop. Anthropic thinks a prompt that clearly said which systems were in and out of scope would likely have kept Claude offline, yet Claude's prompt also told it there was no internet, and that was false. OpenAI found its agents' tendency to break into systems can drop by over 100 times with the software and standing instructions it uses for its regular product together. That figure covers both, so it can't tell you how much the instructions did on their own. Put the stop in the network and the accounts, where the agent's reasoning can't talk its way past it.

Don't newer AI models stop on their own when they realize they're somewhere real?

Sometimes, but not reliably. Anthropic's newest model stopped on its own once it realized its target was real, while an older model noticed and kept going. But Anthropic says these were isolated incidents, that it would need more testing to be confident newer models respond better, and that even the newest one went further before stopping than it would want. A model's judgment is a bonus you can't count on.

What should happen when an AI agent tries a system it shouldn't?

A named person gets it, with a backup. A workable default for any attempt to log into, send to or change a system that isn't on the agent's approved list, adapted from OpenAI's rule for its most severe alerts, is that the person has 30 minutes to show the alert was a false alarm, and if they can't, or nobody answers, the agent pauses. Write down in advance who can pause the agent, who can restart it and what a restart needs.

Article Real World 8 min read

Make Your Agent Rank the Deals First

An AI agent can negotiate well and still make the wrong trade for you. Four steps to check whether it ranks the options the way its owner would.

Continue reading →

Read the original post on LinkedIn →