An AI agent takes instructions from text, and some of that text is written by people you do not trust: a web page it reads, a ticket it summarises, a file in a repository. If the agent can reach everything its host can reach, then anyone who can influence its input can reach it too. Asking the model to behave is not a control.

A boundary is the answer to a narrower question than "is this agent safe?". It asks: if this agent goes wrong, what is the most it can do? This post covers how to draw that line and where to enforce it.

What a boundary is

A boundary is the complete set of things an agent can reach, defined by you and enforced by something the agent does not control. It has three properties.

  • It is explicit. Everything the agent may do was granted by someone. The default is nothing.
  • It is enforced outside the agent. Not in the prompt, not in the agent's own code, and not in a library the agent could skip.
  • It is observable. Every refusal is recorded, so you learn what the agent tried.

A system prompt that says "never call the payments API" is an instruction, not a boundary.

Scope along four dimensions

Least privilege for agents is the same principle as for people, stated more precisely because agents act faster and do not hesitate.

Which resources

Start with the API or tool, then narrow. An access token should name the one API it is for (the aud claim), and a resource server should refuse a token issued for a different audience. A token that works everywhere is a skeleton key.

Which actions

Read and write are different grants. So are "create a refund" and "list refunds". Grant tools one at a time. Connecting an agent to a tool server should not hand it every tool that server offers.

Which arguments

This is the dimension most often missed. An agent permitted to issue refunds can issue a refund of any size unless the permission itself carries a limit. Useful limits include a highest amount, an allowed list of values, or a fixed value the agent cannot change. The limit must be checked by the system granting access on every call, not by the agent.

A permission written with its limits might look like this. The format is illustrative; the idea is that the ceiling lives with the grant.

agent: support-assistant
grants:
  - tool: billing.create_refund
    limits:
      amount: { max: 50 }
      currency: { allowed: [EUR, GBP] }
    approval: required_above_limit
  - tool: tickets.read

For how long

A credential should live about as long as the task. Access tokens that expire in minutes mean a leaked token has a short useful life, and revocation does not depend on reaching every place the token was copied to. For a single sensitive action, go shorter still: a token bound to that one action and its exact details, valid for about a minute.

Why short-lived credentials matter more for agents

Agents keep credentials in places people do not: environment variables, tool configuration, conversation context, logs of their own reasoning.

Short lifetimes change the question from "has this ever leaked?" to "has this leaked in the last fifteen minutes?". Combined with an agent that authenticates using a private key it never sends anywhere, there is no long-lived secret to exfiltrate at all. For the mechanics, see what is AI agent identity.

Where the boundary is enforced

A boundary is only as strong as the weakest place it is checked. There are three useful layers, and they catch different failures.

  1. At the token issuer. The agent only receives a token for what it was granted. This stops it from asking for more.
  2. At the gateway or resource. Each call is checked against the grant and its limits, with the arguments in view. This stops a valid token being used for the wrong thing.
  3. On the host. The machine the agent runs on refuses connections to destinations nobody granted. This stops everything that never goes near a token at all.

The third layer is the one most deployments lack. A compromised agent does not have to use the SDK, the proxy or the gateway you gave it. If its container can open a connection to any address, it can post data to an attacker's server directly, or call the cloud metadata address (169.254.169.254) and pick up the node's credentials. Controls that the workload must cooperate with can be bypassed by not cooperating. Controls in the operating system kernel cannot be skipped from inside the process.

How to roll a boundary out

Turning on default-deny for a running workload is a reliable way to cause an outage. A staged approach works better.

  1. Inventory. List the agents you run and what each one actually calls and connects to.
  2. Classify. Decide which workloads are agents. They get tighter treatment than ordinary services.
  3. Draft rules from observation. Use what the agent really did as the starting allow-list, then remove what it should not need.
  4. Watch before enforcing. Run every rule in a mode that records what it would refuse. Read the results.
  5. Enforce one workload at a time. Keep a way to switch enforcement off quickly.
  6. Add approval for what is left. Some actions are allowed but still should not be an agent's decision alone. See human approval for agent actions.

How AuthFI does it

AuthFI describes the boundary in two parts: what an agent is granted, and what the machine it runs on will let through.

On the granting side, AI Auth issues each agent a 15-minute token that carries only what it was granted. A permission carries its limit, such as a highest amount, an allowed list or a fixed value, and AuthFI checks that limit on every call. Tool calls go through the MCP gateway, where a tool the agent was not granted is refused and the tool's own credential never reaches the agent. A sensitive action uses a 60-second token tied to that action's details.

On the machine side, Zero-Code enforces the boundary with eBPF programs in the Linux kernel of each node. For a workload classified as an AI agent:

  • outbound connections are limited to the destinations that were granted
  • the cloud credentials address is always refused, and no grant can open it
  • the agent cannot open a raw network socket, read another process's memory or load a kernel program

The agent's code is not changed, and there is no SDK or prompt instruction for it to respect. Every rule starts in watch-only and shows what it would refuse; you switch enforcement on one workload at a time, and the cluster owner can turn it off on a node.

The stated limits: the host-level boundary needs Kubernetes on Linux (x86) and a kernel with the BPF security module enabled. Each node reports which protections actually attached, so a node that cannot enforce a rule says so.

Each refusal is written to the audit record with the layer that refused it.

Key takeaways

  • A boundary is what an agent can reach when it stops following instructions. Prompts are not part of it.
  • Scope by resource, action, argument and time. Limits on arguments are the dimension most often left out.
  • Enforce in layers: at the issuer, at the gateway and on the host. Only the host layer covers traffic that bypasses your tooling.
  • Roll out in watch-only first, then enforce one workload at a time.