What Does "Least Privilege" Mean for an Autonomous Agent?

The meeting had been going for forty minutes and the spreadsheet had one unfilled column.

The meeting had been going for forty minutes and the spreadsheet had one unfilled column.

It was an ordinary access-design session — the kind where a platform team maps a new caller onto the existing role model before it gets a credential. Every other row was filled in. Owner: the developer-experience team. Environment: production. Data classification: internal, some customer PII. Then the column headed job function, which is where the role assignment comes from, because that's how the model works: you name what the thing does, and the roles follow.

Somebody had typed engineering-assistant and somebody else had deleted it.

The agent's job was "whatever an engineer asks it to do in a Slack thread." Sometimes that meant reading a runbook, sometimes querying the warehouse. Twice in the pilot it had opened a pull request, and once it had, entirely reasonably, decided that the way to answer a question about a flaky test was to inspect the CI credentials configuration. All of that was the same job function. None of it was the same permission set.

The security engineer in the room said the thing that ended the meeting: "I can't write a role for this. The role is a function of the prompt."

She was right, and the implication is bigger than one spreadsheet. Least privilege as we practice it is not a principle — it's a derivation, and one of its inputs has gone missing.

The derivation that no longer has an input

Every enterprise access model runs the same pipeline, whether or not anyone wrote it down. Someone has a job; the job implies a set of actions; the actions are grouped into a role; the role is granted, and it is correct to the degree the grouping matches the work. Reviews check for drift between the two.

The load-bearing assumption is in the second step: a principal's action set is predictable in advance from what the principal is for. True of a person — a payroll analyst will run payroll, read compensation data, and file exceptions, and will do that for years. True more strongly of a conventional service: a batch job's action set isn't merely predictable, it's enumerable from source code.

A model-driven agent breaks this exactly. Its action set is chosen at runtime, from natural language, by a component whose output is not enumerable in advance and not stable across invocations. There is no job function to derive from, because the job is supplied per-request by whoever is talking to it. You can enumerate the tools it holds — a real and useful bound — but the tool list is a superset far larger than what any single task needs, and it grows every time someone adds an integration.

What's interesting is that this is a return, not a novelty. The phrase comes from Saltzer and Schroeder in 1975, and their formulation was not about people at all: every program and every privileged user of the system should operate using the least set of privileges necessary to complete the job. Program first. They were explicit that the hard part was granularity — that a protection domain should ideally change as a program moves between phases of its work, and that systems of the day couldn't express that, so privilege was granted for the union of everything a program might do.

RBAC was the detour. It made least privilege tractable by binding privilege to a durable human role, which worked because human roles are durable. We spent forty years optimizing a special case, and agents have put us back in front of the general problem the original paper described: privilege that changes during execution, at a granularity finer than the identity.

The three ways teams paper over it

Before the better frame, the three answers I see in the wild, in rough order of popularity.

Give the agent the operator's permissions. The intuitive one: the agent works for Alice, so it should do what Alice can do. Nothing more, which sounds like restraint.

It isn't, and the reason is temporal rather than about scope. Alice's permission set is safe because Alice is attached to it — she exercises maybe two percent of it in a given week, she notices when something looks wrong, and she is asleep for a third of the day. Handing that set to a process that runs unattended at machine speed is a privilege escalation even though the set is unchanged: a permission previously gated by human attention and working hours is now available continuously to something that can be talked into using it. It compounds when the operator is an administrator, which is exactly who gets pilot access first.

Give it a static service-account role. The disciplined-looking answer, failing on a longer timescale. The agent gets agent-assistant with three scopes. Then someone wants it to file Jira tickets, so a scope is added. Then read the wiki. Then check deploy status. Each addition is small, justified, and approved by someone who can see only that request. Nothing is ever removed, because removal means proving a negative about a caller that cannot be interviewed.

The difference with agents is the rate. A conventional service acquires a permission when someone ships code; an agent acquires one when someone has an idea. Eighteen months of that and you have a functional administrator held by a process taking instructions from text.

Ask the model to limit itself. System prompts urging conservative tool use, a planning step that "checks whether this action is in scope," a supervisor model reviewing the plan. These improve the median outcome and are worth having for that. They are not a boundary, for a structural reason rather than any comment on model quality: the model is the untrusted input path — the component processing attacker-influenceable text. Making it the enforcement point puts the decision inside the blast radius of the thing you're defending against. Same shape as a client-side permission check.

So: not the human's set, not a static set, not a self-imposed set. What's left is to stop attaching privilege to the identity at all.

Privilege that belongs to a unit of work

The reframe is small to state and awkward to build: authority attaches to a task, not to a principal.

The agent's own identity carries essentially nothing — enough to authenticate and to be told no. When a task begins, something deterministic decides what that task requires and issues a credential bounded to it. When the task ends, the credential is meaningless. Standing privilege between tasks is approximately zero, so "what is this agent allowed to do?" has no answer, which is the point. The answerable question is "what is this task allowed to do?"

Three things have to be true for that to be more than a slogan.

The authorization decision happens at delegation time, made by something that isn't the model. When a task is created, a deterministic component — policy engine, authorization server, a plain function with a table in it — maps the task to a bounded permission set. It may take the model's proposed plan as input; it must not take it as the decision. That's the same distinction that separates a request from an authorization.

The token is bounded to the task, not the calendar. Below, because it's where the practical pain is.

The task must be nameable inside the token. The one people skip, and without it the rest is theatre. If the credential says payments:write, you have narrowed nothing meaningful — the agent can move any amount to any account, which is the whole problem restated in a smaller font.

Naming the task is what RFC 9396 (Rich Authorization Requests) exists for. Instead of a scope string, the request carries an authorization_details array of structured objects: a type naming a schema, plus fields like actions, locations, identifier, datatypes. The concrete version is "initiate one transfer of £2,400 from account X to sort code Y," expressed as data the authorization server can evaluate and the resource server can check against the actual request body.

Two things about RAR bite. The type value is a contract between your authorization server and your resource server, not a standard vocabulary — RFC 9396 defines the envelope, and the semantics of payment_initiation are yours to define and keep consistent across every service that reads it. That's a schema-governance problem dressed as an authorization mechanism, and it's the real cost of adoption.

And an authorization server that doesn't understand a type must reject the request, not pass it through. A permissive implementation that forwards unrecognized authorization details turns a fine-grained mechanism into an attacker-controlled claim in a signed token. I have seen exactly that shipped, on the reasoning that unknown types would be "handled downstream."

RFC 8707 resource indicators handle the coarser half: the resource parameter binds the issued token to a specific API's identifier, so a token minted for one downstream isn't spendable at another. Mundane, and it's what makes the narrowing survive a leak. Both mechanisms are inert unless the resource server validates the constraint — a token carrying beautiful authorization_details that no service reads is a comment.

Time-boxing, and what it actually breaks

If authority belongs to the task, its lifetime should be the task's lifetime. Not eight hours because that's the SSO session, not one hour because that's the default in the library.

The argument is about what the credential does when nobody is using it. A token held by a process actively making a call has a small exposure surface. The same token sitting in a queue message, a retry buffer, or a log line for the remaining fifty-eight minutes of its life is a standing grant that happens to expire eventually. Agent credentials spend the overwhelming majority of their lifetime idle, and every second of that is risk buying nothing.

Then reality intervenes, in four places.

Long-running tasks. A migration agent runs for six hours. You cannot issue a six-hour task token and claim to have time-boxed anything. The honest answer is decomposition — the task is not one authorization, it's a few hundred, re-requested as the work proceeds — which means the agent framework has to treat credential acquisition as a normal step in the loop rather than something that happens at startup. Most frameworks don't.

Retries. The token expired during a backoff. Now the retry path must distinguish "failed, retry with the same credential" from "failed, re-authorize, and the answer may now be no." Teams reliably discover this in production and reliably fix it by lengthening the lifetime.

Human-in-the-loop pauses. An agent asks for confirmation. The human is at lunch. The token that was going to perform the action dies during the wait. The tempting fix — issue the token after approval — is correct, and it means the approval has to carry enough structure to re-derive the request, which is more engineering than a Slack button.

Refresh. The reframing worth taking away: for an agent, a refresh is not a rubber stamp, it's a re-authorization opportunity — the one moment where a deterministic component gets to look at a still-running task and ask whether it should continue. Has the budget been consumed, has the initiating user's own access changed, has the agent been suspended, has the plan drifted from what was approved. Treating refresh as a policy checkpoint rather than a lifetime extension is nearly free, because you're already making the call.

Note the cost: your authorization server now holds task state, which is exactly the statefulness stateless tokens were meant to avoid. That's not a rounding error, and it's the honest reason short task tokens are rarer than they should be.

Chains compose by intersection, and almost nothing enforces it

Here is the part I think is most under-appreciated, and it matters more than any individual scope grant.

An agent calls an MCP server, which calls an internal service, which reaches a database. Ask what that chain is permitted to do. The intuitive answer, and the one most implementations encode, is that each hop authorizes independently against whatever credential it holds — so effective authority is the union of the participants, and the last hop's credential decides.

The correct answer is the intersection: what the originating user may do, and what the agent may do on that user's behalf, and what this task authorizes, and what each intermediary may do. A chain cannot produce more authority than its most restricted link. Every hop is an opportunity to reduce and never to add.

flowchart TB
    subgraph U["Union — evaluated at the last hop only"]
        UA[User: read + write orders] --> UB[Agent: broad tool access]
        UB --> UC[MCP server: service credential<br/>full DB access]
        UC --> UD[(Database)]
        UD --> UR["Effective authority =<br/>the service credential"]
    end
    subgraph I["Intersection — evaluated at each delegation"]
        IA[User: read + write orders] --> IB[∩ Agent policy: read only]
        IB --> IC["∩ Task: order #4471 only"]
        IC --> ID["∩ MCP audience + scope"]
        ID --> IR["Effective authority =<br/>read order #4471"]
    end

Say plainly what enforcing the right-hand side requires, because it's more than a diagram.

Every hop must present the inbound authority as the basis for the outbound one, rather than authenticating as itself — delegation rather than a service credential, which is the mechanism the token-exchange conversation is about and which I won't relitigate here.

The narrowing has to be enforced by the issuer, not requested politely by the caller. If a hop can ask for a scope its inbound token didn't carry and get it, you have union semantics with delegation-shaped decoration.

And the hard one: intersection is only computable over comparable permission vocabularies. Intersecting orders:read with orders:read is trivial. Intersecting a RAR object describing one payment against a downstream service's RBAC roles against a database's row-level policy is not a lattice operation — there's no general algorithm, because the three systems share no semantics. In practice you get intersection only inside the domain where one authorization server owns the vocabulary, and at every boundary out of it you fall back to the coarsest common denominator, usually a scope string. I don't have a good answer for this, and I don't think anyone does. It's the main reason chain-wide least privilege stays aspirational even in teams doing everything else right.

There's a temporal hole too. Intersection is evaluated when the token is minted; the participants' authority can change afterwards. Revoke a user mid-chain and tokens already issued for her still carry her authority until they expire — another argument for lifetimes measured in the task, since short tokens are how you bound a problem you can't otherwise solve.

When scoping is imperfect, blast radius is the control that still works

All of the above will be imperfectly implemented. Design for that rather than around it.

Limits on rate and value, enforced below the agent. Not "the agent shouldn't send more than fifty emails," but a counter in the resource server or gateway that refuses the fifty-first. The number is a policy decision made by a human at design time; the enforcement is deterministic. Unglamorous, and the control that most reliably converts a catastrophe into an incident.

Irreversibility is the axis that matters, not the verb. The useful tiering isn't read/write:

  • Reversible and invisible outside the system — reads, drafts, anything undoable before anyone notices.
  • Reversible with effort — writes to systems with soft deletes, version history, restorable state.
  • Irreversible or externally visible — sending mail to a customer, moving money, publishing, deleting from a store with no history, changing permissions.

That cuts across HTTP verbs in ways that surprise people. DELETE on a soft-delete store is tier two. POST /messages is tier three the instant the message leaves your network, because no authorization system can un-send it. Reversibility is a property of the resource, not the method, so you have to have the conversation resource by resource. Nobody wants to. It takes an afternoon per service and produces the only classification that actually predicts how bad a mistake will be.

The irreversible tier requires human confirmation, and that's an authorization decision. This is where I'd push hardest against how it usually gets built. Confirmation for irreversible agent actions tends to be treated as a UX affordance — a Slack card with an Approve button, owned by the product team, easily disabled when it gets annoying. It should be a property of the authorization model: no token an agent can obtain is sufficient for a tier-three action, and completion requires a separate authorization bound to that specific action and that specific human.

The difference shows up when someone asks to turn it off. A UX affordance gets removed in a sprint planning meeting because it's slowing people down. An authorization rule requires a policy change, which is reviewable and hard to do accidentally. Same user experience, entirely different failure mode under pressure.

The counterarguments I find hardest

It's a round trip and complexity most teams will not pay for. Task-scoped authorization means an authorization call per unit of work, an authorization server that understands your task vocabulary, a RAR schema per resource type, and an agent framework that re-acquires credentials mid-loop. Against that, a static service account is one config change and works this afternoon. The trade is worth it above a certain blast radius and clearly not below one, and I don't have a crisp threshold to offer. If your agent only reads public documentation, build none of this.

Over-narrow scoping fails loudly and gets widened badly. This is the failure I'd bet on. An agent scoped to exactly the permissions of its first three tasks will fail on the fourth, at 2am, in front of a customer. Somebody with production access will widen the grant to stop the failure, and under time pressure they will widen it generously and permanently. You end up worse off than with a moderate static role, because the widened grant carries the credibility of having been added by an engineer solving a real problem.

The only mitigations I trust are making narrowing reversible and observable — denied-but-logged shadow mode before enforcement, per-permission usage telemetry so widening can be walked back on evidence rather than nerve — and accepting a slightly looser initial scope in exchange for it holding under stress.

And there is no good answer for a genuinely open-ended task. "Investigate why revenue dropped last Tuesday" has no derivable permission set, because determining what data is relevant is the task. You can't scope it without answering the question first. The options are all unsatisfying: constrain it to a curated read-only surface and accept it will miss things; make it request access per step and accept the latency and approval fatigue; or grant broad reads with hard tier-three blocks and heavy limits, while being clear with yourself that you chose containment over least privilege.

I don't think better engineering closes that gap. Least privilege requires knowing what the work needs, and there is a class of work whose definition is figuring out what it needs. That was always true — it's why incident responders have broad access — and for humans we handled it with accountability rather than restriction. There is no agent equivalent of a career's worth of consequences.

Where this leaves the spreadsheet

The column that couldn't be filled in was the right thing to be stuck on. It was reporting that the model underneath had run out of inputs, not that someone lacked imagination.

What replaces it is less tidy: no durable role, a near-empty standing identity, authority minted per task and expiring with it, chains that narrow at every hop, and a structural stop in front of anything you can't undo. Some of that is buildable today from parts that already exist and are mostly unused — RFC 8693 for the delegation, 9396 for saying what the task actually is, 8707 for keeping the result spendable in one place. The rest — chain-wide intersection across systems that share no permission vocabulary, and open-ended tasks generally — has no good answer yet, and I'd rather say so than sell a diagram.

For disclosure: the model I work on (ClavionX, design-phase, not battle-tested) treats an agent as a registered first-class object with an owner, a stated purpose, and an explicit REGISTERED → ACTIVE → SUSPENDED → RETIRED lifecycle, realized as an ordinary OAuth2 client and governed by a reusable agent policy rather than per-object grant types — machine-to-machine only, never authorization code — with delegation as plain RFC 8693, subject staying the human and actors accumulating. That solves the governance half: which agent, whose, under what policy, revocable how. It does not solve the scoping half. I mention it mainly to be clear about which half anybody's product is solving, mine included.

The scoping half is the interesting one, and it is open.