Threat Modeling an Identity Platform

The whiteboard had eleven boxes on it and we had been arguing for two hours.

It was a threat modeling session for an identity platform, run properly: a facilitator, a printed STRIDE table, sticky notes in four colors. We went component by component. The token endpoint: could it be spoofed, tampered with, repudiated, disclosed, denied, elevated? Six answers. The session store: six more answers. The SAML handler, the JWKS endpoint, the admin API. Sixty-six cells, most of them filled with some variant of "TLS" or "input validation" or "we log it."

The document that came out of that session was thirty-one pages and, as far as I know, nobody ever opened it again. Nine months later we shipped a bug where an assertion signed by one customer's identity provider was accepted for a user in a different customer's tenant. It was not a component bug. Every box on that whiteboard behaved exactly as designed. The bug lived in the arrow — in the tacit agreement between two boxes about what the data crossing between them meant.

That is the thing about identity platforms that generic threat modeling handles badly. In an identity system, almost nothing interesting happens inside a component. It happens at the moment one side stops verifying and starts believing. The vulnerabilities that have actually defined this field — signature wrapping, IdP mix-up, code interception, tenant confusion — are all failures of an assumption held across a boundary, not failures of a component's internal logic. If your threat model's unit of analysis is the component, you will enumerate sixty-six cells and miss all of them.

So this article proposes a different unit of analysis, and a structure you can reuse. Not "what can go wrong with this box," but: for each boundary, what crosses it, what does each side assume about the other, and where has that assumption historically failed?

A trust boundary is an asymmetric belief

The standard definition — a boundary is where the privilege level changes — is true and not very useful. Here is a sharper one that I have found actually generates findings:

A trust boundary is a place where one party accepts a statement it cannot independently verify, on the strength of something it believes about the other party.

Two properties fall out of that, and both matter.

First, the belief is usually asymmetric. The relying party believes the identity provider is honest. The identity provider mostly does not care whether the relying party is honest, because it isn't the one being harmed. When you write the boundary down, you must write down two assumptions, one per side, and they will not be mirror images. The exploitable one is almost always the side that is trusting more and checking less.

Second, the belief is inherited by everything downstream. If a resource server trusts the token, and the token trusts the assertion, and the assertion trusts the redirect, then a flaw at the redirect is a flaw at the resource server. Identity is the layer everything else derives its authority from, which is why identity boundary failures are rarely "medium severity."

Here is the map I would draw first for any identity platform. It is deliberately small — five boundaries, not eleven boxes.

flowchart LR
    U["User agent<br/>(browser, mobile app)"]
    IDP["Identity provider<br/>(authorization server)"]
    RP["Relying party<br/>(app / resource server)"]
    T1["Tenant A state"]
    T2["Tenant B state"]
    CP["Control plane<br/>(configuration authority)"]
    RT["Runtime<br/>(protocol execution)"]
    AG["Agent"]
    RES["Downstream resource"]

    U ---|"B1: credentials, redirects,<br/>codes, cookies"| IDP
    IDP ---|"B2: assertions, tokens,<br/>metadata, keys"| RP
    T1 ---|"B3: issuer, keys,<br/>subject identifiers"| T2
    CP ---|"B4: projected config,<br/>policy, keys"| RT
    AG ---|"B5: delegated authority,<br/>tool calls"| RES

    classDef b fill:#f6f6f6,stroke:#666,stroke-width:1px
    class U,IDP,RP,T1,T2,CP,RT,AG,RES b

Five boundaries. Now do each one the same way, every time: what crosses, what each side assumes, and the named historical failure that proves the assumption is not free.

B1 — Browser ↔ identity provider

What crosses: credentials, MFA challenges, authorization codes, state and nonce values, session cookies, and — critically — redirects.

What the IdP assumes: that the user agent faithfully relays messages between the parties it was told to relay them between, and that a request arriving with a session cookie was intended by the human who owns that session.

What the browser assumes: that the page asking for the password is the page it appears to be.

Both are weak, and the weakness has a name: the user agent is the transport, and the transport is attacker-influenced. In every other protocol boundary you control the channel. Here, the channel is a program running on a machine you don't own, executing markup from origins you didn't choose, following redirects on your behalf.

The historical failure: redirect handling. The redirect_uri is the one field where the authorization server hands attacker-influenceable data the power to relocate a credential. The class covers open redirects chained through a legitimately registered URI, prefix-matching registration (https://app.example.com/cb matching https://app.example.com/cb.evil.com), and the covert-redirect family from 2014. The mobile variant — a malicious app registering the same custom URL scheme and intercepting the authorization code — is the reason PKCE (RFC 7636) exists, and it is worth remembering that PKCE was originally a mobile patch that we later concluded everyone needed, including confidential web clients. That reversal is itself a lesson about boundary assumptions: "our channel is trustworthy because it's a server-side web app" was never the property that made the code safe.

The reusable question: which values crossing this boundary can an attacker influence, and which of those are used to make a routing or trust decision rather than a display decision?

B2 — Identity provider ↔ relying party

What crosses: assertions and tokens, plus the metadata that says how to validate them — issuer URLs, signing certificates, JWKS documents, discovery documents.

What the RP assumes: that a document bearing a valid signature from a configured key was issued by that IdP, for this RP, about this subject, recently.

What the IdP assumes: usually nothing. This is the most asymmetric boundary in the whole system, which is precisely why it is the most attacked.

Notice how many separate claims are packed into the RP's one-sentence assumption. Valid signature — over what, exactly? For this RP — checked how? Recently — against whose clock? Each conjunct has its own failure history.

The historical failures, and there are several worth naming separately because they break different conjuncts:

  • XML Signature Wrapping. The signature is valid; it just covers a different element than the one your code reads. The 2018 round (Duo's research, and CVE-2017-11427 and neighbors) added the variant where XML comment handling caused text-node truncation, so [email protected]<!---->.evil.com verified as one string and was read as another. Breaks valid signature over the thing I read. There is a detailed walkthrough in xml-signature-wrapping-explained.md.
  • Accepting unsigned or partially signed messages. Breaks the assumption before it starts; see why-unsigned-saml-is-dangerous.md.
  • IdP mix-up. When an RP federates to more than one IdP, an attacker starts a flow at an honest IdP and delivers the response to the RP as if it came from a different, attacker-controlled one — or the reverse. The RP validates the code or token against the wrong IdP's endpoints and leaks it. This was serious enough that the fix became a spec: RFC 9207 requires the authorization response to carry an explicit iss. Breaks issued by the IdP I think I'm talking to.
  • Audience and recipient not checked. A perfectly genuine assertion, minted for a different relying party, replayed at yours. Breaks for this RP.

The reusable question: for each conjunct in "signed, by the right issuer, for me, about a subject I can identify, and still fresh" — which line of code enforces it, and what happens if that line is deleted? If deleting it breaks no test, the conjunct isn't enforced, it's assumed.

B3 — Tenant ↔ tenant

This boundary does not exist in single-tenant deployments, which is why it is so consistently underestimated: it appears only when a platform becomes multi-tenant, and it appears inside code that was written when there was nothing on the other side.

What crosses: ideally nothing. In practice: shared issuers, shared signing keys, shared caches, shared subject identifiers, and shared code paths.

What tenant A assumes: that a token good for tenant A is not good for tenant B, and that its users, its policy, and its keys are not reachable from another customer's configuration.

What the platform assumes: that every query carries a tenant predicate.

The historical failures cluster into three shapes:

  • Issuer confusion. A platform issues tokens with one issuer for all tenants and puts the tenant in a claim. Now the RP's validation of iss and signature passes for a token from any tenant, and the only thing preventing cross-tenant access is whether the RP remembered to also check the tenant claim. Many did not. The Azure AD multi-tenant class of bugs — validating the signature and the issuer but not the tid — is the canonical public example, along with the "nOAuth"-style variants where an unverified email claim was used as the cross-tenant join key.
  • The dropped predicate. A repository method loses its tenant filter in a refactor, and the bug is invisible until two tenants happen to have overlapping identifiers. This is the mundane version and it is far more common than the exotic one.
  • Shared cache keys. The lookup is correct; the memoization in front of it is keyed on user ID without the tenant segment.

The structural defense is to make the boundary impossible to cross by construction rather than by predicate. For disclosure, the platform I work on — ClavionX, design-phase rather than battle-tested — takes the strict version: every tenant gets its own OIDC issuer URL and its own signing keys, so a token minted for tenant A fails validation against tenant B's issuer before any claim inspection happens, and there is no "shared trust by default" to forget to check. Its runtime state keys carry a tenant segment as a construction rule, not a convention, so a cross-tenant read requires building a key that doesn't exist rather than omitting a WHERE clause. Neither of those makes tenant confusion impossible. What they do is move the failure from silent acceptance to loud rejection, which is the only reliable difference between a boundary and a hope.

The reusable question: if a tenant predicate were silently removed from any single query in this system, would anything fail loudly? If the answer is "only if two tenants collide," you have a filter, not a boundary.

B4 — Admin plane ↔ runtime

This is the boundary that gets left off the diagram, because both sides are us. It is also the boundary with the highest blast radius on the entire map.

What crosses: configuration. Federation connections, signing keys, client registrations, claim mappings, policy, user records, group memberships.

What the runtime assumes: that the configuration it has been given is authentic, current, and complete.

What the admin plane assumes: that whoever is making this change is authorized to make it — and, more subtly, that a configuration change is a lower-severity operation than an authentication.

That second assumption is the interesting one and it is wrong. An administrator who can add a federation connection can mint any user. Not "can escalate to admin" — can become any user in that tenant, silently, through an entirely legitimate protocol flow, leaving audit records that look like a normal login. Compare the ceremony around a password reset with the ceremony around adding an IdP, in almost any platform, and you will find the more dangerous operation is the less guarded one.

The historical failure class is admin-plane privilege escalation that never touches the authentication path: a support tool that can impersonate users; a tenant admin able to attach a globally-scoped resource; an admin API whose authorization check verifies the token's signature and scope but not that the target tenant is one the caller administers. The SolarWinds-era Golden SAML technique belongs here too — steal the IdP's signing key from the management side and every downstream assertion is genuine forever, because the boundary you compromised is the one that defines validity rather than the one that checks it.

Two structural properties help, and both cost something. The first is direction: make configuration flow one way. In ClavionX's ADR-0002 the runtime never synchronously calls the control plane during a request; state arrives only as projections, and a missing projection fails secure rather than reaching back for a fresh read. That removes a whole category of "runtime request reaches into admin state" attacks, and it buys availability during a control plane outage. It costs you eventual consistency you now have to reason about explicitly — a revoked entitlement has a propagation window, and that window is a number you must be able to defend.

The second is audience separation on a shared runtime. Its ADR-0010 states the invariant plainly: a shared issuer is not a shared trust domain. Admin tokens carry an admin audience, live in a separate session namespace, and enter through a separate ingress; a tenant-audience token must never satisfy an admin API and vice versa. That is a policy boundary rather than an infrastructural one, and it is honest to say policy boundaries are easier to erode — one convenience endpoint that accepts either audience and the invariant is gone.

The reusable question: which configuration changes are equivalent in power to authenticating as an arbitrary user, and do they require more ceremony than a password reset does?

B5 — Agent ↔ resource

What crosses: delegated authority — a token obtained on a human's behalf, carried by software that decides for itself what to do next.

What the resource assumes: that the caller's instructions originate from the principal named in the token.

That assumption was safe for thirty years, because software did what it was compiled to do. It is not safe now. An agent's next action is a function of text it read, and some of that text came from the data it was asked to process. Prompt injection is not a new vulnerability class; it is an old one — confused deputy — arriving through a new channel. The deputy holds authority delegated by the user and is induced by a third party to exercise it. Norm Hardy's 1988 compiler bug, with an LLM in the middle. the-confused-deputy-comes-back-with-agents.md works through the mechanics; least-privilege-for-autonomous-agents.md covers why the usual scoping derivation loses its input.

What matters for the threat model is narrower: this is the first boundary on the map where authority and intent decouple. Everywhere else, if the token is valid, the request reflects the principal's intent. Here it may not, and no amount of token validation detects the difference — the token is genuine, the signature verifies, the audience is right, and the action was chosen by a web page the agent read.

So the controls that work are not authentication controls. They are the ones that assume the deputy will be confused and bound what a confused deputy can do: authority that shrinks at each delegation hop rather than being passed along whole (RFC 8693 exchanges narrowing audience and scope, with the actor chain recorded); the downstream service being able to see that it is talking to an agent rather than a human, so it can apply different policy; and structural human checkpoints on the small set of actions where being wrong is unrecoverable.

The reusable question: for every action this agent can take, what is the worst outcome if the instruction to take it came from the data rather than from the user?

The whole thing on one page

Boundary What crosses The trusting side's assumption Named failure when it breaks
Browser ↔ IdP Credentials, codes, cookies, redirects The user agent relays faithfully; this request was intended Open redirect / prefix-matched redirect_uri; mobile code interception (→ PKCE)
IdP ↔ RP Assertions, tokens, keys, metadata Signed, by the right issuer, for me, about this subject, fresh XML signature wrapping; comment truncation (CVE-2017-11427 class); IdP mix-up (→ RFC 9207); audience not checked
Tenant ↔ tenant Issuer, keys, identifiers, code paths A token for A is not good for B Multi-tenant issuer confusion (tid unchecked); dropped tenant predicate; unsegmented cache key
Admin ↔ runtime Configuration, keys, policy The config I hold is authentic and current Golden SAML; admin-plane escalation; audience-confused admin API
Agent ↔ resource Delegated authority, tool calls The caller's instructions come from the principal Prompt-injected confused deputy

Five rows. That table is the artifact I would actually keep — not the thirty-one pages. It fits on a screen, a new engineer can read it in four minutes, and every row generates a specific question about specific code.

What this doesn't catch, honestly

I would rather you use this than STRIDE-by-component, but I am not going to pretend it is complete.

It only covers boundaries you drew. The tenant boundary didn't exist until the platform had two customers; the agent boundary didn't exist three years ago. The failure mode of any boundary map is a boundary that came into existence after the map was drawn, and no methodology fixes that — only re-drawing it when the architecture changes does. A map with a date on it and a rule about when to redraw beats a better map with no date.

Time is a boundary this framing draws badly. Most identity incidents involve staleness: a token still valid after a termination, a cached key after a rotation, a projection lagging behind a revocation. That is a boundary between now and five minutes ago, and it doesn't appear as an arrow between two boxes. The compensating practice is to attach a number to each crossing — maximum propagation lag, maximum token lifetime — and treat those numbers as part of the model.

People are a boundary too. Help desk resets, offboarding processes, the vendor with support access. These are real crossings with real assumptions, and they fail more often than cryptography does. the-hardest-problem-in-identity-is-recovery.md is about the largest of them.

And a threat model is not a control. This is the part worth saying plainly, because it is where the ceremony goes wrong. A document listing assumptions has changed nothing. The assumption becomes real only when it is attached to something that fails when it's violated: a test that sends a wrapped assertion, a test that replays tenant A's token at tenant B, an architecture rule that fails the build when a repository method lacks a tenant parameter. If the output of a threat modeling session is prose, its half-life is one sprint. If the output is tests and build rules that carry the boundary's name, it survives the people who wrote it.

That is the real difference between the thirty-one pages and the five rows. The pages described a system. The rows describe a set of beliefs — and beliefs, unlike components, can be written down, argued about, and checked.