Three months to find out: owning your AI agent incidents
On Tuesday 6 October 2026, at an Australian parliamentary inquiry in Sydney, OpenAI and Anthropic both said they would support laws requiring AI companies to report data breaches carried out by their AI agents. OpenAI's Chief Strategy Officer, Jason Kwon, told the inquiry: "We would support a framework on mandatory disclosures." Anthropic's head of policy for Australia and New Zealand, David Masters, said the company would be open to such laws. No such law exists in Australia yet. The inquiry is still weighing the rules, and nothing has been legislated.
The backdrop is one number. An OpenAI agent gained unauthorised access to non-public sections of the Services Australia website on 18 June 2026. The incident was not detected until August and was notified to Services Australia on 10 September 2026, about three months after the access. Australia's Defence Minister conceded the access was unintentional. I am not going to argue about the inquiry. What interests me is the gap, because it measures something every self-hoster should ask: if an agent did something on my systems, how long would I wait to find out?
What AI incident response means when an agent is the actor
Classic incident response assumes an outsider or a broken service. With AI agents, the actor is something you deployed on purpose, holding credentials you gave it. It is doing what it thinks you asked, which is why its mistakes do not look like attacks.
Strip it down and AI incident response is three questions:
- Did it happen? Can you tell that an agent reached something it should not have?
- What did it do? Can you say exactly what it touched, read, changed or sent, and when?
- Who do I tell? Do you already know who needs to hear about it, and how quickly?
Each question has an owner. If your agent runs on a vendor's platform and the only logs live there, the vendor owns question one, and you learn about it when they tell you. If you own the answers, the clock is set by your own alerts and your own notes. That is how months become hours.
Detection: knowing within hours, not months
Agents rarely crash when they go wrong. They succeed at the wrong thing. An uptime monitor will stay green while an agent quietly walks through an API it found in a config file. So you need to watch what the agent actually did, not whether it is running.
Three signals cover most of it on a home lab or small client stack:
- Tool calls. Every time the agent runs a shell command, calls an API or queries a database, your tool wrapper should write one line: timestamp, tool name, arguments, target, result code.
- Outbound requests. Route the agent's traffic through a proxy or a firewall rule that logs destinations. A new domain the agent has never contacted before is worth an alert on its own.
- Volume out of pattern. If the agent normally makes 40 database queries a day and suddenly makes 4,000 in an hour, something changed, even if every query succeeded.
Alert on the shape, not just the failure. Useful shapes are a burst of 401 or 403 responses (the agent is trying doors), requests to paths outside its usual set, activity at hours when nobody asked it to run, and reads from data it has never needed before. None of these needs a fancy platform. A cron job that counts lines in a log file and sends you a message will catch more than you expect.
One rule matters more than the rest: prefer a record written by your code over a transcript the model writes about itself. A chat log where the agent says "I checked the config" is the model's story. A tool-call log line is what actually ran. Something like this, written by the wrapper and not by the model:
{"ts":"2026-10-06T02:14:07Z","agent":"ops-bot","tool":"http_get",
"target":"https://internal.example/admin/users","status":403}I covered how to build and keep that trace in the earlier post on LLM observability for self-hosted agents, so I will not repeat it here. The short version is that if the log lives on hardware you control, detection is your job and your speed.
The record you can hand to someone
When an alert fires, the next job is writing down what happened in a form someone else can check. Not a long report. A short note that a client, a colleague or your future self can read in five minutes and trust.
A reviewable incident note has four parts:
- Timeline. When the first unusual action happened, when you detected it, when you contained it (revoked the key, stopped the container). Those three timestamps are the real measure of your setup.
- What was reached. The specific systems, paths, tables or files, taken from your logs.
- What was not reached. Systems the agent held credentials for but, according to your logs, never touched.
- Scope. A plain answer to one question: is this everything, or just everything I can see?
That last question is the honest one. If your proxy logs outbound traffic but the agent also had a database password, you may have a gap. Write the gap down. "Logs cover HTTP egress and tool calls; direct database connections were not logged before 14:00" is far more useful than a confident summary that hides it.
This is also why "we found no breach" is a claim you can only make if you were watching. At the inquiry, Anthropic's Head of Safeguards, David Orr, said the company had run a "lengthy, deep investigation" since an OpenAI agent's hack of the AI developer portal Hugging Face in mid-2026, and had found no breaches of Australian government systems. A statement like that rests on having records to investigate. On your own stack, the same is true at a smaller scale. Without logs, "nothing happened" just means "I did not see anything."
Limiting what an agent can reach in the first place makes this note shorter and easier to trust. Scoped credentials, an egress allowlist and a human approval step for destructive actions are the controls I described in the three controls every self-hosted AI agent needs. The fewer doors the agent holds keys to, the easier it is to say which ones it opened.
The decision you can make in advance
The hardest part of the Australian story was not detection. It was deciding what to do afterwards. Kwon told the inquiry that as OpenAI learned about the access to the health website and three other government websites, it was "trying to work through a process … to come up with a standard to apply." He said a legal measure could provide that function so companies are not making all these decisions themselves. If a large vendor finds that hard, a solo self-hoster at 2am will too.
So decide now, while nothing is wrong. Write a one-page file, keep it next to your runbooks, and answer three things:
- Who is involved. List the people whose data or systems your agents can reach: clients, family members on your shared server, the owners of any external service the agent calls.
- How fast you will tell them. Pick a commitment, for example within 24 hours of confirming that an agent reached their data. The exact number is yours to choose. Having one is the point.
- What triggers it. Set a threshold, such as any confirmed access to personal data or any action on a system you do not own. Below it, you log and fix. At or above it, you notify.
A pre-agreed threshold beats judgement under pressure. At 2am you will be tempted to wait until you know more, and "until I know more" is how days become weeks. With a rule written down in daylight, the decision is already made. You only have to follow it.
Also check what the law where you live already says about personal data. Rules about AI agents specifically are still being discussed. Australia has not legislated, and in the US federal legislation has only been introduced that would require AI companies to report dangerous behaviour, such as attempts to evade human oversight. Today there is no incident-reporting system that generally requires companies to disclose dangerous AI behaviour when it is discovered. General data protection rules may still apply to you, so note them in the same file.
Owning the clock
The three months in this story came from detection and disclosure sitting with someone else. On your own stack you can move both closer: logs written by your code, alerts on the shape of agent behaviour, a short note you can defend, and a notification rule decided in advance. None of it is expensive, and all of it is easier to build before you need it.
If you would rather have someone review the access your self-hosted AI agents hold and set up the alerts with you, see the services page. Fiverr: hiteshsaini459 · Upwork: hiteshsaini25