How to log what an AI agent actually did
Updated 2026-09-28
A chat transcript tells you what the agent said. It does not tell you what happened. If an agent can call tools, book something, send something, spend something, you need a separate record of every call: who asked, what ran, what it changed, and what came back. That record is a receipt. Build it at the point where the tool actually executes, not by parsing the conversation afterward.
Why a chat transcript is not an audit trail
A transcript is prose generated by a model. The model can paraphrase a tool call, omit one, or describe an action it didn't take. None of that is malicious; it's just what language models do with text. An audit trail has to come from the system of record, written by the code that actually ran the action, not by the code that described it. If your only log is the chat history, you can't answer "did this really happen" without re-running the action or trusting the model's own account of itself.
What a receipt should record
One receipt per tool call, written at execution time. At minimum:
- Who. The person or account the action is being taken for, and whether the caller was a person, an agent acting on their behalf, or an unattended job.
- Which agent, which tool. The exact tool name and the definition or app it came from. Two agents that can both call "send_email" are not interchangeable; log which one did it.
- Inputs and effect class. The arguments the tool ran with, and where that tool sits on a severity scale: something like read, write, destructive, or spend. A destructive call and a read call need different scrutiny later; you can't tell them apart after the fact unless you recorded the class at the time.
- Approval. If the action needed a yes from a person, which approval covered it, and whether the approved action matches the one that ran. An approval for "email this recruiter" should not silently cover "email a different recruiter."
- Result. What the tool returned, or the error it failed with.
- Cost. What the call cost, if it spent money or metered usage. Even a nominal cost is worth logging; you can't total spend across a month of agent runs otherwise.
- Time. When the call happened, timestamped by the system, not asserted by the agent.
- An idempotency key. Something unique enough that a retried call, or two agents racing on the same action, doesn't produce two receipts for one real-world effect, and doesn't run the action twice.
The Model Context Protocol lets a tool author annotate a tool with hints like readOnlyHint, destructiveHint, and idempotentHint. Those are useful for a human skimming a tool list, but they're hints from whoever wrote the tool, not a guarantee. Nothing stops a tool author from marking a destructive action read-only, by mistake or otherwise. Treat them as documentation, and classify effect severity independently, on your own side of the boundary.
Where to enforce it: the tool boundary, not the prompt
An instruction in a system prompt is a request to the model, and a model can misread it, get talked out of it by a user, or drop it after a long conversation. It is not a control. The place a receipt actually gets written, reliably, is the code that executes the tool call: the function or service the agent's tool call passes through on its way to doing something real. That's also the one place that can refuse a call outright, before it runs, if the caller isn't allowed to make it. Writing the receipt and enforcing the policy in the same place means a receipt only exists for calls that were actually allowed to happen, and every allowed call leaves a receipt. Neither is true if you're relying on the model to narrate itself.
This matters more as an agent's tool list grows. OWASP's Top 10 for LLM Applications names this failure "excessive agency": an agent ends up able to take actions nobody explicitly scoped it to take, because the permission logic lives in a prompt instead of in code that can say no.
Retention and privacy
A receipt should record enough to reconstruct what happened without becoming a second copy of everything the agent touched. Inputs and results often contain personal data, so decide up front how long a receipt lives, who can read it, and whether it needs to be scoped to the account it belongs to rather than sitting in one shared log everyone can query. Multi-tenant systems should treat receipts as tenant data like any other: a receipt from one customer's agent run should not be readable through another customer's account, even by an internal query that forgot to filter.
How receipts support undo and billing
A receipt is also what makes "undo" possible. If you know exactly what changed, with what inputs, you can write the reverse of that specific action instead of guessing at what state to restore. And a receipt with a cost field is what makes usage-based billing auditable: you can show a customer the exact list of paid actions an agent took on their behalf, not just a total.
Checklist
- Every tool call writes a receipt at the point it executes, not after, and not by parsing a transcript.
- The receipt names the actor, the agent, the tool, and the effect class, independent of any tool-author-supplied hint.
- A denied or failed call gets a receipt too. A record of "this was refused" is as important as a record of what ran.
- Receipts carry an idempotency key so retries and races don't double-write or double-run.
- Receipts are scoped per account or tenant, with a stated retention window.
- Cost is on the receipt if the action spent money or metered usage.
How Agent Rails does it
On Agent Rails, every tool call runs inside the same process that owns the account's data, and the platform, not the agent's own instructions, decides what each tool is allowed to do. Rails keeps a classification table of every tool in every installed app, independent of what the app itself claims about a tool: least to most severe, a call is read, write, destructive, or one that moves money. An app or tool Rails hasn't classified fails closed rather than running unclassified.
Receipts are written as part of the same database transaction as the change they describe, so a receipt can't exist for a write that got rolled back, and a write can't happen without one. A refusal writes a receipt too, through an independent connection, specifically so that rolling back the failed action can't erase the record that it was refused. Every receipt and every record an installed agent creates is scoped to the tenant it belongs to, enforced by the database itself, not only by application code that could have a bug in it: a query that forgets to filter by account gets nothing back, rather than someone else's rows.
Where an action needs a person's yes first, that approval is checked against the exact content of the action, not a general permission. Agent Rails' Recruiter agent, for example, only submits a candidate after the owner approves a submission whose recorded content matches, field for field, what's about to be sent; if the underlying draft changed since the approval was granted, the approval expires rather than covering the new version.