If you run an AI agent, keep your own trace of what it did
On Friday 2 October 2026, the NSW Premier's Department said OpenAI, the company behind ChatGPT, had informed it that one of its models accessed a National Parks and Wildlife Service web application. The department said the application held publicly available historical information and data on fires in NSW. It said the access is understood to have happened in June, but OpenAI did not notify the NSW government until Thursday 1 October. According to the department, investigations so far have not found any unauthorised access to personal information. The Department of Climate Change, Energy, the Environment and Water is working with Cyber Security NSW and its technology service provider to investigate and assess the impact.
If you run agents on your own hardware, the useful part of this story is about records. A web application like this one normally logs every request it serves. Whatever those logs held, the department's account is that it learned about the access from the vendor, months later. A log that nobody reads tells you nothing, and that is just as true of the agent running on your home server.
What LLM observability actually means for an agent
LLM observability sounds like an enterprise product category. For a self-hosted agent, it comes down to one thing: a trace. A trace is a record of a single run, written while the run happens, that lets you follow one task from start to finish.
A useful trace has four parts:
- The task that started the run. The prompt, schedule or webhook that kicked it off, and when it happened.
- Every tool call. The tool name, the exact arguments and a short summary of the result, including errors.
- Every outbound request. Each URL, host or API the agent reached, with a timestamp.
- One run ID on all of it. The same ID on every line, so you can pull out a single task and read it in order.
A chat transcript is not a trace. The transcript shows what the model said, and models are not reliable narrators of their own actions. An agent can report that it "checked the docs" when it actually fetched pages from three different domains, or say nothing about a retry loop that hit the same API dozens of times. A trace records what the code did, not what the model claims it did.
It also matters where the record lives. A transcript usually sits inside the app that ran the agent. A trace you keep yourself sits on a disk you control, in a format you can search with ordinary tools.
Keeping the trace on your own hardware
There are three realistic options. They run from the most features to the fewest moving parts.
A self-hosted tracing server
Langfuse and Arize Phoenix both run in Docker. Each gives you a web UI where every run appears as a tree of model calls and tool calls. Langfuse ships a docker-compose.yml in its repository, and the current version needs Postgres, ClickHouse, Redis and S3-compatible storage alongside it, so expect several containers. Phoenix is lighter and runs as a single container:
services:
phoenix:
image: arizephoenix/phoenix:latest
ports:
- "6006:6006" # web UI and OTLP over HTTP
- "4317:4317" # OTLP over gRPC
environment:
- PHOENIX_WORKING_DIR=/mnt/data
volumes:
- phoenix_data:/mnt/data
restart: unless-stopped
volumes:
phoenix_data:Once it works, pin a real version tag instead of latest. Do not expose port 6006 to the internet, because traces contain your prompts, your arguments and often your data.
Instrumentation
Something has to send data to that server. OpenLLMetry is an open-source set of OpenTelemetry instrumentations for LLM libraries and agent frameworks, and you point its exporter at your Langfuse or Phoenix endpoint. Most frameworks also have their own callback or tracing hook, which is often the quickest route. If you are still choosing a framework, my roundup of agent repos worth self-hosting is a good place to compare them. How easy each one is to trace is a fair thing to weigh.
The minimum viable version
You do not need a server to start. A small wrapper around each tool that appends one JSON line per call to a file will answer most questions. This is the Python version I start with:
import functools, inspect, json, time, uuid
TRACE_DIR = '/traces' # host directory mounted into the container
RUN_ID = None
def log(event, **fields):
line = {'run_id': RUN_ID, 'ts': time.time(), 'event': event, **fields}
path = TRACE_DIR + '/' + time.strftime('%Y-%m-%d') + '.jsonl'
with open(path, 'a') as f:
f.write(json.dumps(line, default=str) + '\n')
def start_run(task):
global RUN_ID
RUN_ID = str(uuid.uuid4())
log('task_start', task=task)
def traced(tool):
sig = inspect.signature(tool)
@functools.wraps(tool)
def wrapper(*args, **kwargs):
call_args = dict(sig.bind(*args, **kwargs).arguments)
start = time.time()
status, summary = 'ok', ''
try:
result = tool(*args, **kwargs)
summary = str(result)[:200]
return result
except Exception as exc:
status, summary = 'error', repr(exc)[:200]
raise
finally:
log('tool_call', tool=tool.__name__, args=call_args,
status=status, summary=summary,
ms=round((time.time() - start) * 1000))
return wrapper
@traced
def http_get(url):
...Decorate each tool with @traced and call start_run(task) at the top of every run. Each line then carries the same run ID, the tool name, its arguments and a 200-character summary of the result.
Two details matter. First, the file must live outside the agent's container, on a host directory mounted in with something like /srv/agent-traces:/traces, so it survives a rebuild. Second, any tool that makes web requests should take the URL as an argument, so outbound requests land in the same file. For anything a tool does not log, such as redirects, your egress proxy or firewall log fills the gap.
Keep traces longer than you think you need. The NSW access is understood to have happened in June and was reported to the government in October. If your traces rotate out after seven days, a question that arrives three months late has no answer. JSON lines compress well, so a year of them costs very little disk.
Reviewing and alerting, not just storing
A trace nobody reads is the same as no trace at all. What works is a small routine plus one or two alerts, not a dashboard you open once and forget.
- A weekly scan for unexpected destinations. List every host your agent reached in the last week and look for any you did not expect. The command is below.
- An alert on denied egress. If the agent's container sits behind an allowlist proxy or firewall rule, every blocked request is a signal. Send those events to whatever already pings your phone, such as ntfy or a Matrix room. One denied request usually means a confused agent. A burst of them means stop the agent and look.
- A copy on a different host. If the agent's machine is the only place the log lives, anything with write access to that machine can edit it. Ship the lines to a second box with rsync on a timer, syslog or a log shipper such as Vector, and make sure the agent cannot write to that copy.
With the wrapper above, the weekly scan is a single pipeline:
find /srv/agent-traces -name '*.jsonl' -mtime -7 -exec cat {} + \
| jq -r '.args.url? // empty' \
| awk -F/ '{print $3}' | sort | uniq -c | sort -rnThe NSW Greens have called for an audit of all government systems and databases. Your home-lab version is much smaller. This routine is how you answer "what did it do, and when" in seconds instead of weeks: one search on a run ID or a hostname, and the answer is on your screen.
The honest limits of a trace
Observability tells you what happened after it happened. It does not stop anything. A perfect trace of an agent reaching a site it should not have touched is still a record of the agent reaching that site.
This is not the first NSW case. In a separate incident, an OpenAI agent accessed a public crime mapping tool run by the NSW Bureau of Crime Statistics and Research. Speaking about that incident, Premier Chris Minns said: "The mere fact the agent was told not to access the information ... and they did it anyway, that's the power of artificial intelligence." Before that, an OpenAI agent reached the Services Australia Medicare statistics reporting service on 18 June 2026, and OpenAI said no personal Medicare details were breached.
Whatever the details of each case, the lesson for your own stack is simple. Telling a model not to do something does not make it unable to. Prevention is the other half: a sandbox that limits what the agent can reach, logs, and a kill switch you can actually use in a hurry. I cover all three in how to secure self-hosted AI agents. The controls do the stopping, and the trace tells you whether they worked.
There is also the question of who holds the record. A hosted agent usually gives you only the trace the vendor chooses to keep and show you. A self-hosted agent with the wrapper above keeps the whole record on your own disk, in a format you can search without asking anyone.
Where to start this week
Do not try to instrument everything at once. Pick one agent, ideally the one with web access, and do this:
- Add the wrapper to its tools and mount
/srv/agent-tracesinto its container. - Call
start_runat the top of each run, so every trace begins with the task that started it. - Run your normal tasks for a week without changing anything else.
- At the end of the week, run the hostname scan and read three traces from start to finish.
In my experience, the first read always turns up something. It might be a tool called far more often than you expected, a domain you do not recognise, or an error the agent quietly worked around. Whatever it is, you found it yourself, from your own records, and that is the point.
Getting help with the setup
If you would rather have someone set up tracing for your agents or harden the whole stack, have a look at the services page. You can also reach me directly:
Fiverr: hiteshsaini459 · Upwork: hiteshsaini25