Identity Isn't Your Authorization Engine (Beyond a Point)

The request that started it was two sentences long, and both were reasonable.

"Can the token include which projects the user can see? We're making a call back to the API on every page load just to filter the project switcher, and it's slow."

Reasonable. The identity platform already knows the user, already knows the tenant, already mints a token on every login. Adding a projects claim is a mapping change. Ship it.

Then: "Archived projects shouldn't appear." Fine — filter on a flag.

Then: "Users on the Starter plan only see the three most recently active." Now the claim depends on billing state and activity timestamps.

Then: "Guests see projects they've been explicitly invited to, but not projects their team owns." Now it depends on a membership graph with two kinds of edge.

Then: "When a project moves to a different team, the old team keeps read access for thirty days." Now it depends on an event and a clock.

Then a support ticket: a user was removed from a project at 09:02 and still had it in their switcher at 09:11.

Six requests in, the identity platform needs to know what a project is, what archived means, how a plan tier maps to a quota, the difference between a member and a guest, and the retention semantics of a team transfer. It needs a synchronized copy of the project table. And it has acquired a nine-minute window in which revoked access still works.

Nobody made a bad decision. Each step cost less than the one before and obligated more. That's the shape of this boundary: it doesn't get crossed, it gets eroded.

This isn't the definitional argument

Authentication Is Not Authorization covers the conceptual distinction, and it's not the interesting part. Everyone can recite it. The interesting question is the one that comes after you already agree: some authorization genuinely belongs in the identity platform. Tenancy does. Coarse roles do. Whether the user cleared MFA in the last five minutes definitely does.

So the boundary isn't a wall. It's a line through a continuum, and it sits in a different place for a five-person startup than for a company with 900 tenants — which is why "put authorization in the app" is useless advice on its own, and why teams that hear it still end up with a projects claim. It's the same erosion that turns an identity store into the company's user database, one sympathetic field at a time. What follows is where the line goes, why it moves, and what breaks when you put it in the wrong place.

A test with three questions

Take a concrete authorization decision — not a category, an actual yes/no question — and ask:

1. Does answering it require domain state the identity platform would have to replicate?

The load-bearing one. "Is Alice an admin of tenant 44?" needs only identity data. "May Alice approve this expense?" needs the amount, the submitter, the state, and the reporting line — four facts the identity platform can only get by holding a shadow copy. A shadow copy is a synchronization problem, and synchronization problems fail stale-and-permissive far more often than stale-and-restrictive.

2. Does it change at a different rate than identity data?

Identity facts change at HR speed — a hire, a transfer, a termination, a group added during a quarterly review. Days to months. Domain authorization changes at application speed — a document shared, a project archived, a ticket reassigned. Seconds. Co-locate two datasets whose rates of change differ by three orders of magnitude and the propagation machinery built for the slow one becomes the correctness bound on the fast one. Your projection pipeline was designed for group membership and is now the delivery mechanism for unshare.

3. Is its blast radius the authentication path?

Anything evaluated at issuance runs inside the login. A defect in project-visibility logic should degrade a project switcher; living in claim mapping, it can fail token issuance instead — a bug in a domain feature taking down authentication for everyone, including the people trying to log in and fix it. A project list has no business sharing a failure domain with the front door.

Applied:

Decision Domain state? Rate Blast radius Where
Which tenant is this user in? No Months Auth path anyway Identity
Is this user active or terminated? No Months Auth path anyway Identity
Did they authenticate with a phishing-resistant factor? No Per session Auth path only Identity
Does this tenant's plan include the reporting module? Billing tier Months Feature gate Border — see below
May Alice read document 8812? Yes Seconds One request Application
May Alice approve this expense? Yes, plus a graph Seconds One request Application

Three yeses is unambiguous. The interesting rows have a single yes — and there the test doesn't resolve the question, it prices it. A plan-tier entitlement is domain state, but low-cardinality, slow-changing domain state, and putting it in a claim buys a feature gate with no network call. Often the right trade. Make it knowingly.

Why the line moves

Two variables move it, and neither is architectural taste.

Cardinality. A claim expresses a decision whose answer space is small and enumerable: five roles, twenty groups, a plan tier. A relationship expresses one whose answer space is subjects × objects — user:alice is editor of doc:8812 is one fact out of hundreds of millions. Claims scale with the number of kinds of access; relationships scale with the number of instances. Once the answer depends on which specific object is named, claims have stopped being a sane encoding — not philosophically, but because the encoding is now O(objects) inside a bearer token.

Coupling across applications. A fact three applications need is worth centralizing; a fact one application needs is taxed by centralization without benefit. This is why the same decision sits on different sides in different companies: "which regions may this user operate in" is a domain rule in a single product and a genuine identity attribute where twelve products enforce it. It's the composability test applied to a decision instead of a feature.

Failure mode one: the token stops fitting

Everyone has heard that tokens get big. Fewer have seen how it actually lands, which is nonlinear and on your most important user.

The arithmetic: a group name is 30–60 bytes of JSON, JWTs are base64url-encoded so multiply by 1.33, and the practical ceiling isn't a spec value but the default header buffer of whatever proxy sits in front of you — commonly 4KB or 8KB, for all headers combined. In a cookie the ceiling is lower still, 4096 bytes.

Then the distribution. Group membership isn't uniform: most enterprise users are in a handful of groups, a long tail sit in hundreds, and whoever holds the most memberships in a large directory is almost always someone senior with a decade of accumulated access. So it works in staging, works for 99% of production, and breaks for one customer's global administrator — the person most likely to be on a call with your CEO — with a 400 reading Request Header Or Cookie Too Large that mentions nothing about groups. Diagnosing that from the symptom is hard, because the request never reached your application.

The deeper point isn't the byte budget. It's that "just put it in the JWT" is not a design, it's a deferral, and what it defers is a cardinality decision. Groups-per-user grows monotonically in every organization, because adding a group is a two-minute helpdesk action and removing one requires knowing what breaks. You are betting that a number nobody governs stays under a limit nobody monitors. If you take one operational action from this article: emit token size as a metric at issuance, with percentiles. Two lines of code, and it turns a mystery outage into a chart with a slope.

Failure mode two: the decision is as stale as the token

This one gets waved away, and it deserves stating plainly:

An authorization decision baked into a 15-minute token is a 15-minute-stale authorization decision. Not "usually fresh." Not "eventually consistent." A hard ceiling on how fast a revocation can take effect, chosen — usually accidentally — by whoever picked the token TTL for unrelated reasons.

Fifteen minutes is a fine window for "is Alice an admin." It's a terrible one for "may Alice see this project," because the expectation attached to unshare is that it's immediate, and the person clicking it files a security incident when it isn't. Staleness tolerance is a property of the decision, not of your token infrastructure. Identity facts tolerate minutes; sharing decisions tolerate roughly zero.

The escape hatches all cost something. Short-lived tokens hand your availability budget to the token endpoint. Introspection makes every decision current at the price of a network call on the hot path — less expensive than most teams assume, but no longer offline. A revocation list buys fast negative decisions and nothing else. It's the JWT-versus-opaque trade-off in different clothes. None is wrong; what's wrong is a window you didn't choose. Write the number down per decision class, and the decisions demanding a small number turn out to be exactly the ones the test already excluded from the token.

Failure mode three: the identity platform becomes a small, bad programming language

This one kills maintainability rather than availability.

Once domain-shaped decisions live in identity, the configuration surface must grow to express them. Mapping a field isn't enough; now you need a conditional. Then a conditional over a list. Then a lookup. Each step is a small extension to a config format, and collectively they reinvent an expression language — one with no debugger, no test framework, no type system, and evaluation semantics defined by whatever the implementation happens to do on a null. The Extensibility Trap traces that progression generally; authorization is its most common entry point, because authorization is where customers have the most specific requirements.

The counter-move is to keep enrichment deliberately, almost frustratingly inexpressive. Disclosure: I work on ClavionX, whose constraints make this concrete. Claim enrichment there (ADR-0009) is a field-reference model — a claim name bound to a single placeholder like {{user.department}} or {{authEvent.acr}} — with no conditionals, no loops, no computed expressions, no network calls, and no database lookups at issuance. Values must resolve to a string, number, boolean, or array of strings; nested objects are rejected at configuration-load time, explicitly so uncontrolled growth can't produce tokens that blow past header limits. Protocol claims can't be overridden at all.

The interesting part isn't the restriction, it's the stated consequence: if conditional logic is required, it must be encoded in the projected user data upstream, not in the claim expression. That's a cost transfer, and an honest one — some requirements get harder. What it buys is issuance that stays deterministic, does no I/O, and can't be broken by a customer's rule, which is what makes it safe inside a login path that must keep working when the control plane is unreachable (ADR-0002; see Identity Is Mostly Read Traffic). The principle holds regardless of platform: the identity runtime's enrichment should be a projection, not an evaluation.

What genuinely belongs on the identity side

The purist reading of all this is wrong. These are identity's job:

  • Tenancy. Not negotiable. It's the isolation boundary, needed before any domain lookup can be scoped; sourcing it elsewhere is a circular dependency.
  • Lifecycle state. Active, suspended, terminated — the only authorization decision with a legal deadline attached, and identity-sourced by definition.
  • Coarse roles and group membership. Low cardinality, slow, cross-application, and — critically — inputs to policy rather than policy itself. If the only expression of an access rule is a string comparison against a group name, you can't enumerate who has access without reading source code.
  • Coarse entitlements. Plan tier, licensed modules, seat class. Domain-ish, but low-cardinality and slow, and the alternative is a billing lookup per request.
  • Authentication assurance: acr, amr, auth_time. The strongest case here, because identity is the only system that can produce them. An application can require step-up for a dangerous action; it cannot itself know whether a phishing-resistant factor was used ten minutes ago. Identity asserts the evidence, the application sets the bar — the model behind step-up authentication done properly.

One more, often missed: administrative authorization over the identity platform itself is identity's job. Who may create users, rotate a client secret, or read the audit log concerns identity's own resources, and needs the rigor of any delegated administration model. ClavionX draws that as a hard separation: administrative authorization (a group-to-permission model over the control plane) and application authorization (the scopes and claims your application consumes) are distinct concepts sharing no data and answering different questions. Conflating them is how an application role quietly becomes a platform privilege.

Where policy engines actually sit

The systems that handle what identity shouldn't fall into two families, and they solve different halves.

Relationship engines (Zanzibar-style ReBAC — SpiceDB, OpenFGA, the Google original) answer "is there a path from this subject to this object through these relations?" They solve the cardinality problem: hundreds of millions of tuples, sub-10ms checks, and a consistency token so you can demand read-your-writes when the answer must reflect a share from 200ms ago. Use them when access is derived from structure — folders, teams, ownership, inheritance.

Policy engines (OPA/Rego, Cedar) answer "given these attributes and this context, does policy permit it?" They solve the expressiveness problem: business rules, conditions, time windows, amounts. Use them when access is derived from rules over data you already have.

Plenty of systems need both. The mistake is thinking either replaces identity — they consume it:

flowchart LR
    IdP["Identity platform<br/>sub, tenant, role, group<br/>acr / amr / auth_time<br/>lifecycle state"]
    PDP["Policy / relationship engine<br/>rules + relationship graph<br/>per-object, per-action"]
    App["Application<br/>enforces, owns domain state"]

    IdP -->|"claims as inputs"| PDP
    App -->|"subject + object + action"| PDP
    PDP -->|"permit / deny"| App
    App -->|"writes relations"| PDP

Note that the arrows into the decision point come from both sides: identity supplies who and how strongly authenticated, the application supplies which object and what action. Neither alone suffices, which is the structural reason this can't collapse into one box. And the policy engine is a peer of the identity platform, not a component of it — bundling it inside the IdP puts a service with domain-rate write traffic into the failure domain of the login path, and recreates the shadow-copy problem in a new location.

The case for centralizing, taken seriously

The counter-argument is not naive, and I've watched teams regret ignoring it.

One audit trail. When authorization lives in eleven services, "who could access this record last March" is eleven queries against different schemas with different retention policies, and the reconstruction is approximate. Centralized, it's one query — a real compliance advantage, and where most centralization efforts start.

One revocation point. Split authorization means split revocation. Disabling an account in identity does not delete the relationship tuples granting Alice edit access to 4,000 documents; if she's reinstated in a different role, they're still there. Someone must own that reconciliation, and if nobody is named, nobody does it.

One place to answer "who can do what." Access reviews are hard with one system. With two, the reviewer needs a joined view somebody has to build and keep correct.

The split costs more besides. Debugging becomes distributed: "why was this denied" spans a token, a policy decision, and an application check, and unless all three log a shared correlation ID with their inputs, the answer is unrecoverable. Two change processes, two rotations, and a seam between them where the interesting bugs live — usually a claim the policy engine expected and identity stopped sending.

Three mitigations do most of the work. Emit decisions from every enforcement point into one stream, with inputs and policy version, not just outcomes — recording decision inputs is the one part you cannot add retroactively. Have the policy engine subscribe to identity's lifecycle events, so a termination propagates instead of needing a cleanup job. And keep one document stating, per decision class, which side owns it — the erosion in the opening story happens fastest where nothing is written down.

The honest exemption: with one application, one tenant, and no second identity source in prospect, this separation is overhead bought against flexibility you may never need. Put it in the token. Just know which of the three questions you answered "yes" to, and what your window is.

Back to the project switcher

The right answer was never a projects claim. It was: the token carries sub, tenant_id, the user's role, and the assurance level; the application asks its own store — or a relationship engine — which projects that subject can see right now; and the page-load cost gets fixed with a cache keyed on a relationship version, invalidated when membership changes.

That answer is more work in week one and less in year two, which is precisely why it loses arguments. The way to win it isn't to appeal to purity. It's to ask the three questions out loud, in the ticket, before the mapping change merges: does this need domain state we'd have to copy, does it change faster than identity data, and does a bug in it break login? Three noes and it's a claim. One yes and you're making a trade. Three yeses and you're teaching your identity platform what a project is — and once it knows, it can never be allowed to forget.


Disclosure: I work on ClavionX, and used two of its constraints above to make general points concrete — a deliberately inexpressive claim-enrichment model, and a hard separation between administrative and application authorization. Both carry real costs; there are requests they can't satisfy as directly as a scriptable hook would. The argument stands independently of the product, and applies just as much if you run Keycloak, Entra, Auth0, or something you built.