Token Exchange (RFC 8693): The Grant Type Nobody Teaches

Somewhere in your architecture there is a service that holds a token it should not be using.

Somewhere in your architecture there is a service that holds a token it should not be using.

It got that token honestly. A user logged in, the frontend called the orders service, and the orders service needed something from the pricing service, so it forwarded the token it already had. Pricing needed inventory, so it forwarded the same token again. Four hops later a token minted for a single browser session is being replayed across half the estate, and every service that touched it could, if it wanted, use it anywhere the original audience is accepted.

Nobody designed this. It's what happens when the only two options a team knows are "forward the user's token" and "call downstream as ourselves," and both are wrong in ways that only show up during an incident.

The third option has been standardized since 2020, is implemented by most serious authorization servers, and almost nobody reaches for it. RFC 8693 defines a grant type whose entire job is this: turn one token into a different token, more narrowly scoped, addressed to a specific downstream audience, that still records whose request this originally was.

It reads as an obscure corner of OAuth until you start building anything with agents in it. Then it stops being optional.

The two bad options, stated precisely

Forwarding the user's token is convenient because it requires no work. The user's identity is preserved perfectly — sub is still Alice at every hop, so authorization downstream is easy to reason about.

The costs are all in blast radius. That token's aud names the first service, and everything downstream is now accepting tokens not addressed to it, which means every service in the chain has to accept tokens addressed to every other service. You've built one trust domain wearing the costume of a microservice architecture. Any service that logs a request header, any sidecar that dumps traffic during debugging, any dependency that gets compromised — all of them now hold a credential valid everywhere. And nothing in the token says which services touched it, so an audit log downstream shows Alice's request arriving with no record of the four processes that shaped it.

Calling downstream with your own client credentials fixes the blast radius and destroys the identity. Inventory sees a request from orders-service, which has broad permissions because it serves every customer. The user's identity is gone by the second hop. Authorization downstream degrades to "is this a service I trust," which is not authorization, and any user who can influence what orders-service asks for has just borrowed all of its permissions.

Teams usually pick the second option after a security review flags the first, then rebuild the lost identity by hand — a X-On-Behalf-Of: [email protected] header, or a user ID in the request body. That header is unsigned. It is an assertion by a service that a user exists somewhere upstream, and the downstream service either trusts it blindly or does nothing with it. I have seen this pattern in production at three different companies, and in all three, nobody could say what would happen if that header were wrong.

Token exchange is what you build when you want the identity and the containment, and you want the answer to be cryptographic rather than conventional.

What the grant actually does

A service presents a token it holds, plus its own credentials, and asks for a new token. In return it gets one that names a specific audience, carries whatever subset of scopes was requested, and — this is the part that matters — retains the original subject while recording the caller as an actor.

The request is a POST to the token endpoint, form-encoded like every other grant:

POST /oauth2/token HTTP/1.1
Content-Type: application/x-www-form-urlencoded

grant_type=urn:ietf:params:oauth:grant-type:token-exchange
&subject_token=eyJhbGciOi...            # the token representing WHO this is for
&subject_token_type=urn:ietf:params:oauth:token-type:access_token
&audience=https://inventory.internal    # who the new token is FOR
&scope=inventory.read                   # narrowed, not inherited
&requested_token_type=urn:ietf:params:oauth:token-type:access_token

The client authenticates normally — private_key_jwt, mTLS, whatever your platform mandates — so the exchange is bound to a known caller rather than being a thing any holder of the subject token can do.

Two parameters carry most of the design weight.

subject_token is the answer to on whose behalf. In the common case it's the user's access token, but the type is negotiable: an ID token, a SAML assertion, or a platform-issued OIDC token from GitHub Actions or your Kubernetes API server. That last case is worth noticing, because it means the same grant that handles user delegation also handles workload identity federation. Both are "I hold a credential from one trust domain and need one from yours."

actor_token is the answer to who is doing it, and it's optional in a way that trips people up. If you omit it, the authenticated client is implicitly the actor. You supply it explicitly when the caller is relaying for someone else — when service B, called by service A, needs the resulting token to record both.

The response looks like an ordinary token response, with one oddity worth flagging because it breaks naive clients:

{
  "access_token": "eyJhbGciOi...",
  "issued_token_type": "urn:ietf:params:oauth:token-type:access_token",
  "token_type": "Bearer",
  "expires_in": 300
}

issued_token_type is required, and the server is allowed to give you something other than what you asked for. Ask for a JWT, get an opaque token — that's a legal response, not a bug, and if your client assumes it can parse the result you'll find out in production.

Delegation and impersonation are different operations

The spec supports two semantics through the same endpoint, and the distinction is the single most important thing to get right.

Delegation preserves both parties. The new token says Alice, via orders-service. Downstream sees the user and the intermediary, and can make decisions about either.

Impersonation erases the caller. The new token says Alice, full stop, indistinguishable from one issued to Alice's own browser session. Anything downstream sees a request that appears to come directly from the user.

Impersonation is occasionally the right answer — a legacy resource server that can only reason about a sub claim leaves you no choice. But it destroys attribution at exactly the moment attribution matters. When a support engineer's tooling impersonates a customer to reproduce a bug, and something gets deleted, the audit log says the customer deleted it. That's not a logging gap you can fix later; the token itself contained no other information.

The mechanical difference is the act claim. Impersonation omits it. Delegation adds it, and the structure is nested rather than flat:

{
  "iss": "https://auth.example.com/tenant-42",
  "sub": "[email protected]",
  "aud": "https://inventory.internal",
  "scope": "inventory.read",
  "exp": 1754400000,
  "act": {
    "sub": "orders-service",
    "act": {
      "sub": "agent-7f3a"
    }
  }
}

Read that from the outside in. sub is Alice and stays Alice at every hop, forever. The outermost act is the current actor, the one that made this call. Each nested act is a prior actor, further back in the chain. So the token above says: this is Alice's request, currently being carried by orders-service, which received it from agent-7f3a.

The nesting order is the detail people get backwards, and getting it backwards is worse than not having the claim, because now your audit log is confidently wrong about who did what. The most recent actor is at the top. Each exchange wraps the previous chain rather than appending to a list.

What a chain looks like end to end

sequenceDiagram
    participant U as Alice (browser)
    participant A as Agent
    participant O as Orders service
    participant AS as Authorization server
    participant I as Inventory

    U->>A: authenticated request
    A->>AS: exchange(subject=Alice's token,<br/>audience=orders)
    AS-->>A: sub=alice, act={agent}
    A->>O: call with narrowed token
    O->>AS: exchange(subject=that token,<br/>audience=inventory, scope=inventory.read)
    AS-->>O: sub=alice, act={orders, act:{agent}}
    O->>I: call with narrowed token
    Note over I: sees Alice, and the full path<br/>that carried her request

Two properties are worth staring at.

The token Inventory receives is only valid at Inventory. If it leaks, it buys an attacker one audience, a handful of scopes, and five minutes. Compare that to the forwarded-token case, where the same leak yields a credential accepted everywhere.

And Inventory can authorize on Alice while still knowing an agent is involved. Those are different facts and you frequently need both: Alice may read inventory, but perhaps an agent acting for Alice may not write it. Without the actor chain, that policy is inexpressible — you either allow Alice's authority to be exercised by anything holding her token, or you don't allow the delegation at all.

Narrowing is the whole point, and it's the part people skip

Every exchange is an opportunity to reduce authority, and an implementation that returns the same scopes it received is doing bookkeeping rather than security.

Three things should shrink monotonically as you move down a chain:

Audience. Each token names exactly one downstream. Multiple audiences in an exchanged token means you've decided to skip the exchange for the next hop, which is the forwarding problem with extra steps.

Scope. The requested scope must be a subset of what the subject token carried. This is where you catch design mistakes: if orders-service needs inventory.write to serve a request that Alice initiated with orders.read, something upstream is doing more than the user asked for. The exchange makes that visible as a rejected request rather than an invisible privilege gain.

Lifetime. Exchanged tokens should be short — single-digit minutes, because they're consumed immediately by the next hop. A five-minute token whose only holder is a service making a call right now has a very different risk profile from a one-hour token that sits in a session store.

The spec is deliberately quiet about all of this. It defines the mechanism and leaves the policy to you, which means a compliant implementation can happily return an unrestricted, long-lived, multi-audience token and still pass every conformance test. The grant type gives you the ability to narrow. Whether you actually narrow is a decision your platform has to make and enforce.

Depth limits, and the question of who may act for whom

Nothing in RFC 8693 stops a chain from growing indefinitely. Each hop wraps the last, the token gets larger, and the semantics get harder to reason about. Real deployments cap it — three or four actors is a reasonable ceiling — and the cap needs to be enforced by the authorization server, because no individual service knows how deep it sits in the chain.

The harder question is authorization of the exchange itself: which clients are allowed to act for whom.

The spec offers may_act, a claim placed in a token stating who is permitted to act as its subject:

{
  "sub": "[email protected]",
  "may_act": { "sub": "orders-service" }
}

It's elegant and it's rarely what you want operationally, because it requires knowing the delegation graph at issuance time — at login, before anyone knows which services this session will touch. Most platforms instead resolve delegation permission at exchange time from configuration: this client, holding a token for this kind of subject, may obtain a token for that audience.

Which is a policy decision, and policy decisions have a way of becoming per-client special cases. This is the part of token exchange that gets ugly at scale, and it's worth designing before you have thirty clients rather than after.

ClavionX's approach here is worth mentioning only because it illustrates the shape of the answer rather than the answer itself: delegation permissions come from a named policy assigned to a class of callers, not from fields set on individual clients. Agents get an agent policy that permits TOKEN_EXCHANGE (and deliberately never AUTHORIZATION_CODE, since there is no human to redirect); MCP servers get a policy that can require a delegated user, refusing tokens where no real person sits behind the chain. Change the policy once and every governed object recompiles. There's no agent-specific code path in the runtime — an agent is an OAuth client and delegation is plain RFC 8693, which is precisely why the chain composes when agents start calling agents.

The general principle survives without the product: make delegation permission a property of a class of callers, expressed once, rather than a field somebody remembers to set on each new client. The per-client version works fine until the twentieth client, and then it is load-bearing configuration that nobody can review.

What to check in your implementation

Four things, all of which I have seen wrong in systems whose owners believed they had implemented token exchange:

Does the exchange validate the subject token fully? Signature, expiry, issuer, and audience. An authorization server that accepts an expired subject token because "we're issuing a fresh one anyway" has turned an expired credential into a valid one.

Do downstream services validate aud? If they don't, none of the narrowing matters — you've paid for containment and left the door open. This is the single most common way a correct token-exchange deployment provides no security benefit at all.

Does anything log the act chain? The chain exists so that an incident investigation can answer "who was involved." If your audit records store sub and drop act, you've built the mechanism and thrown away the evidence.

Can a service exchange a token it received into one addressed back to itself, or to a peer at the same level? That's lateral movement inside your own delegation system, and it's usually possible by default because nobody wrote the rule forbidding it.

Why this is suddenly urgent

Token exchange has existed for years as a solution to a problem most teams tolerated. Service-to-service calls forwarded tokens, security reviews grumbled, and the world continued, because the number of hops was small and the services were all written by people down the hall.

Agents change the arithmetic in three ways. Chains get longer and more dynamic — an agent decides at runtime which tools to call, so the delegation graph is not knowable at design time. The intermediaries are no longer all yours: an MCP server may be operated by someone else entirely. And the thing deciding what to do next is influenced by text it read from a document, an email, or an API response, which is a property no service-to-service call ever had.

In that world, "who is this request ultimately for, and what path did it take" stops being a nice property of your audit log and becomes the input to an authorization decision. A resource server needs to distinguish Alice asked for this directly from Alice's agent asked for this after reading a web page, and treat them differently.

You cannot make that distinction with a forwarded token, and you cannot make it with a service credential. You can make it with a token whose subject is Alice and whose actor chain is intact — which is exactly the artifact this grant type exists to produce.

It's the least glamorous entry in the OAuth family, and it's the one the next few years of architecture will be built on.