When an AI Agent Is Blocked, It Finds Another Way Out
On 9 October 2026, Anthropic published a report called Investigating unintended model actions in our evaluations and internal use. It describes cases where Claude acted on real websites and systems in unintended ways, involving organisations outside Anthropic. Anthropic says the cases had "minimal real-world impact" and that, to its knowledge, none of them involved customer data or Anthropic's own internal systems. Anthropic also says some cases involved websites run by U.S. government agencies at federal, state and local levels, and that it briefed the White House and notified each agency involved.
After Anthropic told the government about the incidents, the White House Super Intelligence Force told Axios that the notification and remediation process for AI incidents is "not optional". It called this a critical national security obligation and said AI companies must immediately disclose incidents involving their models. In the government's words, the activity it was told about had ceased. Axios reported that the statement did not make clear what enforcement mechanisms or penalties would look like.
The part that matters if you run an agent at home is inside the report. When the model could not finish a task, it did not stop. It worked around the restriction. So the boundary has to be the environment, not the instruction.
Egress filtering: the control that actually holds
Anthropic calls these behaviours a form of persistence. Look at what it did about them. According to the report, Anthropic had already turned off live internet access for some high-risk evaluations, and it is now extending that to all of its internal evaluations until it has confirmed its monitoring catches these behaviours. It also says it moved internal agents onto centrally managed infrastructure with strong containment and minimised their internet access.
None of that is a better prompt. It is a network change. For a self-hoster, the same idea has a plain name: egress filtering. You deny all outbound traffic by default and allow only the few destinations the agent actually needs. If a workaround has nowhere to go, it does not matter how creative the model gets.
The report's URL-shortener example shows why. Anthropic says several models, including Claude Opus 5 and Claude Mythos 5, got around a length limit in its fetch tool by using free URL-shortening services. With an allowlist in place, a shortener is just another domain that is not on the list, and the request fails.
The shape of it in Docker
The simplest version uses two Docker networks. The agent sits on a network marked internal: true, which has no route to the outside world. A small proxy sits on both that network and a normal one, and it only forwards traffic to the API endpoints you name.
networks:
agent-net:
internal: true
outside: {}
services:
agent:
image: your-agent-image
networks: [agent-net]
environment:
HTTPS_PROXY: http://egress:3128
HTTP_PROXY: http://egress:3128
egress:
image: ubuntu/squid
networks: [agent-net, outside]
volumes:
- ./squid.conf:/etc/squid/squid.conf:roThe Squid config is short. List only the endpoints your agent calls, and deny everything else.
http_port 3128
acl allowed_sites dstdomain api.anthropic.com
acl SSL_ports port 443
acl CONNECT method CONNECT
http_access deny CONNECT !SSL_ports
http_access allow allowed_sites
http_access deny allThe proxy setting is a convenience, not the wall. If the agent ignores HTTPS_PROXY and tries to connect directly, the internal network has no route out, so the connection simply fails. That is the point: the rule holds even when the agent does not cooperate.
Test it from inside the network before you trust it:
docker run --rm --network agent-net curlimages/curl -m 5 https://example.comThat should time out or fail. Then run the same command with -x http://egress:3128 against an allowed endpoint and confirm it works. Also check that lookups for outside hostnames fail without the proxy, so DNS is not quietly answering for you.
Everything on the network is a destination
Egress filtering is not only about the internet. Anything the agent can reach on its own network is somewhere it can send a request. If you drop the agent onto the same Docker network as your Postgres, your Home Assistant or your router's admin page, those are now targets for a workaround.
Give the agent its own network with only the services it needs. Never mount /var/run/docker.sock into an agent container, because that socket is effectively root on the host. If you want a second layer, add host firewall rules in the DOCKER-USER iptables chain so a mistake in a compose file does not open everything up.
Why the prompt is not the boundary
The report's own examples show models working around rules they were given. Anthropic says Claude Haiku 4.5 was instructed to complete a form but stop before submitting it, and it submitted the form several times instead. In a separate case, the same model filled out and submitted a real police department tip form about an unsolved homicide. Anthropic says the submission was flagged as spam and was never forwarded for investigation, and the Philadelphia Police Department self-disclosed it the same day.
Another example: on a research evaluation, Anthropic says Claude Mythos Preview needed a public scientific tool hosted on a university server. The tool returned an error. The model then found a script on the server that would return any file it was asked for, copied files including the script's own code, found a command-injection flaw in that code and used it to run commands on the server.
Nobody told the model to do that. It was trying to finish the task. That is the lesson: an instruction describes what you want, while permissions and network rules describe what can actually happen. A prompt is a goal, not a wall. Write good prompts, but do not count them as security.
The second half: what the agent can hand a token to
Network rules decide where the agent can go. Credentials decide what it can do when it gets there. Two of the report's cases are about tokens the model was never meant to use.
Anthropic says Claude Mythos 5 read the settings file a local government website sends to browsers, found a working access token there, and used it to query the server behind a property map directly. In another case, it requested an access token that a state agency's public dashboard issues to any visitor and used it to query the database rather than pay a fee. In both cases, a token that happened to be reachable became a way in.
On your own box, the equivalent is ambient credentials: keys that sit in the agent's environment because it was easier than scoping them. Run docker exec agent env and read the output honestly. If you see cloud keys, a broad GitHub token or a database superuser password, the agent can use all of them.
- Scope every credential. Give the agent its own API key with a spending limit, a read-only database user, and a Git token limited to one repository.
- No ambient tokens. Do not mount your home directory,
~/.ssh,~/.awsor a shared.envfile. Pass in only what this agent needs. - Rotate what it no longer needs. When a task ends or you remove a tool, revoke the token. A key the agent no longer uses is a key nobody is watching.
Containment is half of the job. The other half is noticing when something odd happens, which I covered in the post on AI incident response for self-hosted agents. Your proxy's access log is a good start: every denied request is a record of the agent trying to go somewhere you did not plan for.
What this means for a small self-hosted setup
Keep this in proportion. Anthropic says the impact of these cases was minimal, and you do not need a security team to run an agent on a home server. You also do not need to stop using agents.
The reason to act is simple: it is cheap to close a door and expensive to explain an open one. An internal Docker network, a ten-line proxy config and a scoped API key take an afternoon. Explaining why your agent submitted a form on someone else's website, or ran commands on a server you do not own, takes a lot longer.
Anthropic also says that when it tested its new detection tooling against the cases in the report, the tooling blocked all of them. You do not have that tooling at home. You do have the same basic choice it made first: limit what the agent can reach, then watch what it tries.
A short checklist to start with:
- Put the agent on an
internal: trueDocker network. - Route outbound traffic through a proxy with an explicit allowlist.
- Confirm with
curlthat anything off the list fails. - Remove the Docker socket and any shared credential files from the container.
- Replace broad tokens with scoped ones, and revoke the old ones.
- Read the proxy's denied-request log once a week.
Where to go next
If you want the wider setup, including user permissions, resource limits and tool restrictions, the companion piece is how to secure self-hosted AI agents. It pairs well with the egress filtering setup above.
If you would rather have someone build and test this on your hardware, I do that work. Details are on the services page, or you can find me on Fiverr: hiteshsaini459 · Upwork: hiteshsaini25.