Learn

When should an AI agent ask before acting?

Updated 2026-09-28

An agent should ask before it does anything hard to undo, and act on its own for everything else. The hard part is making that consistent: sorting your agent's tools by how bad it is if one runs by mistake, then deciding per tool whether it runs automatically, asks first, or never runs at all. Do that sorting once, in code the agent can't talk its way around, not in the system prompt.

Sort tools by effect, not by how they look

Two tools that look similar in a tool list can carry very different risk. "Save a draft" and "send an email" might be one line apart in your tool schema and be nothing alike in consequence. Group tools by what happens if they run: a read tool looks something up and changes nothing; a draft or write tool changes state the person can see and undo; a send tool reaches someone outside the account; a spend tool moves money; a delete tool removes something for good. The exact set of buckets matters less than having one at all, and applying it to every tool, not just the obviously scary ones.

A policy per tool: auto, ask, never

Once tools are sorted, each one gets a policy. A read tool usually runs automatically. A write tool might too, if it's easy to see and undo. A send or spend tool asks first, every time, for that specific action. And some actions should never run through the agent at all, no matter who asks: not because the model can't be trusted with the request, but because the action needs a channel a model shouldn't be able to trigger on its own, like moving a large sum or changing someone's login credentials.

The useful part of this policy is that it's legible to the person who installed the agent. "This agent can look things up and draft messages on its own. It will never send anything, submit anything, or spend money without asking you first, for that exact thing" is a sentence a non-technical buyer can read and trust, because it names what never happens rather than promising good behavior in general.

Approve the exact action, not a standing yes

"The agent may email candidates" is a blanket consent: it's given once and covers every email forever, including ones the person never saw. "Send this email, to this person, with this subject and this body" is consent to one action. The second kind is what actually protects the person granting it, because it can be checked: before the action runs, compare what's about to happen to what was approved. If the two don't match exactly, that isn't a stale approval you can proceed on anyway; it's a new action that needs its own yes.

This also means an approval should expire, and should be tied to a specific version of the thing being approved. If a draft changes after someone approves it, the approval for the old draft should not cover the new one.

Why the check belongs at the tool boundary

Telling a model "always ask before sending" in its instructions is advice, not a gate. A long conversation, a clever prompt, or a genuine model mistake can all make an instruction stop holding, and there's no way to prove from outside that it held. The check that actually works runs in the code between the model's decision to call a tool and the tool executing: it looks up whether this exact call has a matching approval, and it refuses the call if not, before anything happens. The model never gets the chance to talk its way past a check it can't see or influence. This is the same idea OWASP's Top 10 for LLM Applications describes under "excessive agency": permission has to be enforced outside the model, not requested of it.

Show the person what they're approving

An approval prompt that just says "Recruiter wants to submit a candidate, approve?" asks for trust it hasn't earned. Show the actual content: which candidate, which role, which fields are going where, what the person is attesting to. If the action reaches a form or an account outside your system, show what that form will actually contain, not a summary of it. The point of an approval is that a person looked at the real thing before it happened.

Expiry and revocation

An approval that sits open forever is a standing grant with extra steps. Give it a lifetime, and let the person revoke it before it's acted on. If the agent hasn't executed the approved action by the time it expires, treat it as if it were never approved, not as a pending item still worth running. And if the underlying draft or state moved on since approval, the approval should void itself rather than quietly cover the new state.

Checklist

How Agent Rails does it

Every tool an installed agent can call is classified by Rails itself, not by the agent's own manifest: least to most severe, a call is read, write, destructive, or one that moves money, and an app or tool Rails hasn't classified yet fails closed instead of running unclassified. Consent for ordinary use is given once, at install, from a plain-language disclosure the person reads and accepts; Rails treats a small set of effects as irreversible enough that they need a fresh yes every time, specifically actions that move money, change a third-party account, or message someone who isn't the installation's own owner.

Recruiter, one of Rails' installed agents, states its "never" plainly in its own install disclosure: it will not send a message to someone, or submit a candidate to an employer, without the owner's yes for that exact recipient or that exact submission. In code, that approval is checked as a content match: preparing a submission produces a row containing a hash of the exact packet, role, and form fields about to be sent, and approving it only succeeds if that hash still matches what's live. If the underlying draft changed, or the approval sat too long, the digest no longer matches and the approval is void; Recruiter has to prepare a fresh one. The approval also records who prepared it, who approved it, and through what channel, and the owner can revoke it any time before it runs.