Learn

How to log what an AI agent actually did

Updated 2026-09-28

A chat transcript tells you what the agent said. It does not tell you what happened. If an agent can call tools, book something, send something, spend something, you need a separate record of every call: who asked, what ran, what it changed, and what came back. That record is a receipt. Build it at the point where the tool actually executes, not by parsing the conversation afterward.

Why a chat transcript is not an audit trail

A transcript is prose generated by a model. The model can paraphrase a tool call, omit one, or describe an action it didn't take. None of that is malicious; it's just what language models do with text. An audit trail has to come from the system of record, written by the code that actually ran the action, not by the code that described it. If your only log is the chat history, you can't answer "did this really happen" without re-running the action or trusting the model's own account of itself.

What a receipt should record

One receipt per tool call, written at execution time. At minimum:

The Model Context Protocol lets a tool author annotate a tool with hints like readOnlyHint, destructiveHint, and idempotentHint. Those are useful for a human skimming a tool list, but they're hints from whoever wrote the tool, not a guarantee. Nothing stops a tool author from marking a destructive action read-only, by mistake or otherwise. Treat them as documentation, and classify effect severity independently, on your own side of the boundary.

Where to enforce it: the tool boundary, not the prompt

An instruction in a system prompt is a request to the model, and a model can misread it, get talked out of it by a user, or drop it after a long conversation. It is not a control. The place a receipt actually gets written, reliably, is the code that executes the tool call: the function or service the agent's tool call passes through on its way to doing something real. That's also the one place that can refuse a call outright, before it runs, if the caller isn't allowed to make it. Writing the receipt and enforcing the policy in the same place means a receipt only exists for calls that were actually allowed to happen, and every allowed call leaves a receipt. Neither is true if you're relying on the model to narrate itself.

This matters more as an agent's tool list grows. OWASP's Top 10 for LLM Applications names this failure "excessive agency": an agent ends up able to take actions nobody explicitly scoped it to take, because the permission logic lives in a prompt instead of in code that can say no.

Retention and privacy

A receipt should record enough to reconstruct what happened without becoming a second copy of everything the agent touched. Inputs and results often contain personal data, so decide up front how long a receipt lives, who can read it, and whether it needs to be scoped to the account it belongs to rather than sitting in one shared log everyone can query. Multi-tenant systems should treat receipts as tenant data like any other: a receipt from one customer's agent run should not be readable through another customer's account, even by an internal query that forgot to filter.

How receipts support undo and billing

A receipt is also what makes "undo" possible. If you know exactly what changed, with what inputs, you can write the reverse of that specific action instead of guessing at what state to restore. And a receipt with a cost field is what makes usage-based billing auditable: you can show a customer the exact list of paid actions an agent took on their behalf, not just a total.

Checklist

How Agent Rails does it

On Agent Rails, every tool call runs inside the same process that owns the account's data, and the platform, not the agent's own instructions, decides what each tool is allowed to do. Rails keeps a classification table of every tool in every installed app, independent of what the app itself claims about a tool: least to most severe, a call is read, write, destructive, or one that moves money. An app or tool Rails hasn't classified fails closed rather than running unclassified.

Receipts are written as part of the same database transaction as the change they describe, so a receipt can't exist for a write that got rolled back, and a write can't happen without one. A refusal writes a receipt too, through an independent connection, specifically so that rolling back the failed action can't erase the record that it was refused. Every receipt and every record an installed agent creates is scoped to the tenant it belongs to, enforced by the database itself, not only by application code that could have a bug in it: a query that forgets to filter by account gets nothing back, rather than someone else's rows.

Where an action needs a person's yes first, that approval is checked against the exact content of the action, not a general permission. Agent Rails' Recruiter agent, for example, only submits a candidate after the owner approves a submission whose recorded content matches, field for field, what's about to be sent; if the underlying draft changed since the approval was granted, the approval expires rather than covering the new version.