An agent that can issue refunds, merge code or delete records will eventually do one of those things for the wrong reason: a misread instruction, a poisoned document, a bug in a tool. Putting a person in the loop is the obvious safeguard. It is also easy to do badly. Ask too often and people approve without reading. Ask too vaguely and they cannot tell what they are agreeing to.

This post covers which actions deserve a person's decision, and how to build an approval step that still means something on its hundredth use.

When an agent should stop and ask

Approval is a scarce resource. Each request spends a little of someone's attention, and attention that is overdrawn turns into reflex. So the first design task is to ask rarely.

Good candidates for approval share some of these traits.

  • Irreversible. Deleting data, sending money, sending an email to a customer, publishing.
  • High in value or wide in reach. A large transfer, a change to production, an action affecting many records.
  • A change to access. Granting a permission, creating a credential, adding a user.
  • Outside the usual pattern. A destination, amount or resource the agent has not handled before.
  • Beyond a limit. The agent may refund up to a set amount; anything above waits.

Poor candidates are reads, reversible low-value writes, and anything a hard rule could decide. If the answer should always be no, refuse it in policy. If it should always be yes, grant it. Approval is for the cases in between, where context decides.

Limits before approvals

The most effective way to cut approval volume is to put limits on permissions, so that routine calls pass on their own and only the exceptions reach a person. A permission to refund up to 50 with approval above that produces far fewer requests than a permission that asks every time. Limits are covered in the agent boundary.

What a rubber stamp looks like

It helps to recognise the failure before designing against it.

  • The request says "Agent wants to continue. Approve?" with no details.
  • Requests arrive dozens of times a day for trivial actions.
  • Approval is a message in a chat channel, where anyone can react and nothing is recorded.
  • The agent can retry until someone says yes.
  • A request nobody answers goes through after a timeout.
  • The approver is whoever happens to be online, not someone accountable for the agent.

Each of these teaches people that approving is the path of least resistance.

Designing a request that gets read

Say exactly what will happen

The request must carry the real parameters of the action, not a summary the agent wrote about itself. An approver needs to see which agent, acting for whom, wants to do what, to which resource, with which values.

{
  "agent": "support-assistant",
  "on_behalf_of": "user_4821",
  "action": "billing.create_refund",
  "details": { "order": "A-20931", "amount": 420, "currency": "EUR" },
  "reason_for_approval": "amount above the limit of 50",
  "expires_in_seconds": 300
}

The details should come from the system that will enforce the decision, taken from the actual call. A model's description of its own intent can be wrong or manipulated.

Bind the approval to that exact action

An approval for "a refund" that the agent can then spend on a different, larger refund is not an approval. The decision must be tied to the action and its parameters. If anything changes, the agent has to ask again. A short-lived, single-purpose token is a clean way to carry this: it names the approved action and expires within about a minute.

Send it to the right person, out of band

The approver should be someone accountable for the agent: its owner, or a group that answers for it. The request should reach them on a channel the agent does not control, and they should answer as an authenticated person. OpenID Connect Client-Initiated Backchannel Authentication (CIBA) is the standard for this shape of flow: the client starts a request, the person is asked on a separate device or channel, and the client waits for the outcome.

Two rules follow. The agent's own credential must never be able to approve. And the agent should not be the thing that renders the approval prompt, because then the agent controls what the person sees.

Make silence a refusal

Every request needs an expiry, and an expired request must be refused. Fail closed. If an unanswered request is treated as consent, an attacker only has to ask at three in the morning.

Make refusing as easy as approving

Give approve and refuse equal weight. Let the approver add a reason. A refusal should end that attempt, so the agent cannot simply loop.

Record the decision

Store who decided, what they saw, when, and what the agent did next. This is what lets you review approvals later and spot the ones that should have been limits.

Operating approvals over time

Approval design does not end at launch. Review it like any other control.

  1. Measure the volume per approver. If someone gets more requests than they can read, tighten limits or grant the routine cases outright.
  2. Look at approval rates. A request type that is approved every time is a candidate for a policy rule. One that is often refused points to an agent doing something it should not attempt.
  3. Check time to answer. Requests answered in a second or two are probably not being read.
  4. Reassign when owners change. An approval route that ends at someone who has left is a route to nobody.

For actions taken by people, the equivalent control is asking for fresh authentication before something sensitive. See step-up authentication.

How AuthFI does it

In AuthFI's AI Auth, every agent is registered with a named owner, or a group that answers for it if it works on its own. That is where approval requests go.

You mark actions as sensitive. When an agent attempts one, the action waits. The request says which agent, acting for whom, wants to do what, and with which details, and the owner approves or refuses with those details in view. Permissions carry limits, such as a highest amount or an allowed list, which AuthFI checks on every call; past the limit the answer is no or a person is asked.

Three properties come from the product's own description:

  • only a signed-in person can decide, never the agent's own credential
  • a request nobody answers lapses and is refused
  • the approved action runs on a token bound to that one action for 60 seconds

An agent acting for a person asks through OpenID CIBA in poll mode: the request goes to the person on a separate channel and the agent polls until they answer. A request is pending, approved, refused or expired. People answer requests from their workspace, and the request, who decided and when are written to the audit record. More detail is on the agent approvals page.

Key takeaways

  • Ask rarely. Use limits and policy for the routine cases so that approval is reserved for irreversible, high-value or unusual actions.
  • Show the real parameters of the action, taken from the call itself, and bind the approval to exactly those parameters.
  • Route requests to an accountable person on a channel the agent does not control. The agent must never be able to approve for itself.
  • No answer is a refusal, and every decision goes on the record.