Identity Is Not Your Integration Layer

The incident started at 09:12 on a Tuesday, and for the first forty minutes nobody on the call believed it was an identity problem.

Logins were timing out. Not all of them — maybe one in six, and only for users in two of the company's business units. The identity team's dashboards were clean: CPU flat, database healthy, Redis fine, token endpoint p50 unchanged. The runtime was, by every metric it collected about itself, working perfectly.

Somebody eventually thought to look at outbound calls rather than inbound ones. The login path was making a synchronous HTTPS request to a licence-management service that nobody on the call had heard of. That service was making its own call to a vendor API. The vendor was having a partial regional degradation and returning after 9.8 seconds instead of 40 milliseconds. The identity runtime's client had a 10-second timeout, which meant that for the affected users, login was waiting almost ten seconds and then — because the licence check was configured to fail closed — refusing.

The call for that vendor's status page had 300 people on it, in a different company. Nobody there knew that a change in their p99 was, at that moment, deciding whether several thousand people in an unrelated organisation could open their email.

The post-incident question that mattered wasn't "why did the vendor degrade." It was: when did the identity service acquire an outbound dependency list, and who reviewed it? The answer was that nobody had, because nobody had ever seen the whole list at once. It had been built one reasonable ticket at a time over four years, and the total was seven synchronous outbound calls on the authentication path — of which two were on every login, three were conditional, and two nobody could explain at all.

This article is about that shape. Not about whether you should run customer code inside the login path — that's The Extensibility Trap, and it's a different argument with a different answer. Not about which attributes belong in the identity store — that's Identity Isn't Your User Database. This is about position in the dependency graph, and the structural rule that follows from it: the component with the most dependents can afford the fewest dependencies of its own.

That's not an aesthetic preference. It's arithmetic, and it has teeth.

Availability composes multiplicatively, and identity multiplies it across the estate

Everyone knows the first half of this. If your service calls three others synchronously and needs all three to succeed, your availability is the product of theirs. It's the first thing anyone learns about distributed systems, and it's usually where the analysis stops.

Do it concretely. Suppose your identity runtime is genuinely good: 99.99% on its own merits, which is about 52 minutes of downtime a year. Now add the three integrations that showed up in tickets last year, each of which was individually sensible:

Component Availability Annual downtime
Identity runtime itself 99.99% 52 min
CRM (token enrichment) 99.9% 8.8 h
HR system (manager lookup) 99.5% 43.8 h
Licence service (seat check) 99.9% 8.8 h
Composed login path 99.29% ~62 h

Sixty-two hours. You built a four-nines service and you are operating a two-nines service, and the difference is entirely made of systems your team doesn't own, doesn't monitor, and can't page.

That table is familiar enough. Here's the half that isn't.

Availability numbers are not costs. Availability multiplied by blast radius is a cost. The same 99.9% dependency has wildly different consequences depending on where in the graph you attach it. Attach the CRM call to the sales application: 8.8 hours a year during which the sales application is degraded, affecting the people who use the sales application. Attach the identical call to the login path: 8.8 hours a year during which every application in the estate is degraded, because every application authenticates through the same runtime.

If you have 200 applications, adding a dependency to identity is — in expected user-impact terms — roughly 200 times more expensive than adding the same dependency to any one of them. That multiplier is exactly your fan-in.

This is the reframe I'd want an architecture review to internalise, because it converts a vague instinct ("identity should be simple") into a number you can put on a slide. Nobody in the room will argue that a CRM callout is fine if you tell them it converts a 52-minute-a-year system into a 62-hour-a-year system for every product the company ships.

Criticality tiers propagate upstream, and nobody gets told

There's a second-order effect that is, in my experience, the single most under-appreciated fact about integration in identity.

Most organisations tier their services. Tier 0 is "the company stops if this stops": pager coverage 24/7, change freezes, formal capacity planning, no unreviewed deploys. Tier 3 is an internal reporting tool with a business-hours SLA and a maintenance window on Saturdays.

Now: a service's real criticality tier is the maximum of the tiers of everything that synchronously depends on it. Not the tier written on its wiki page. The tier it inherits from its callers.

The moment identity makes a synchronous call to that Tier 3 licence service, the licence service is Tier 0. Nothing about it changed. Its team wasn't consulted. Its runbook still says a Saturday maintenance window is fine. Its on-call rotation still ends at 6pm. It has been silently promoted to a component that can stop the entire company, and the promotion happened in a pull request against someone else's repository.

That is the mechanism behind the classic story in The Extensibility Trap — the HR team taking their perfectly-sanctioned Saturday window and taking down SSO. It isn't that the HR team did anything wrong. It's that criticality flowed upstream along a dependency edge and nobody notified the destination.

So here's a concrete governance rule worth adopting: you may not add a synchronous dependency to the identity path without the owning team formally accepting the tier they are being promoted to. Not a Slack message: an accepted change of SLO, pager rotation and change-management policy. Most of these integration requests die right there, which is the point — the request was only cheap because the cost was being externalised onto a team that hadn't been asked.

Why the pull is real, and why every individual request is reasonable

None of this happens because people are careless. It happens because identity is genuinely, structurally the most convenient place in the architecture to put things, and the convenience is not imaginary.

The identity runtime is the one system that, at the moment of login, simultaneously has:

  • an authenticated user — the hard part is already done, no other system has this yet;
  • a connection to the directory, so it already knows about groups and org structure;
  • an execution point that fires on every login, which is exactly the trigger a lot of business processes want;
  • credentials for half the estate, because federation and provisioning already required them;
  • a team that is good at integration, because that's most of what identity work is.

Given that, look at how reasonable these sound:

"Enrich the token from the CRM." The application needs the account tier to render the right navigation. It's one field. The identity platform is already minting the token. Adding a claim is a mapping change.

"Provision a mailbox on first login." We don't want to pay for mailboxes for people who never sign in, and the only system that knows "this is the first sign-in" is the one doing the signing in.

"Notify the ticketing system when a contractor authenticates." Compliance wants a record correlated with their access request. Identity is where the event happens.

"Check and decrement a licence seat at issuance." We're over-provisioned and finance wants the count to be accurate at the moment of use, not eventually.

Every one of those has a real business need behind it, a named stakeholder, and no obvious alternative home. And every one of them would be rejected instantly if it were proposed in its honest form:

"I'd like to add a synchronous third-party network call to every single request made by every employee in the company, and make the company's ability to work contingent on that third party's afternoon."

That's the same sentence. The reason it doesn't feel like the same sentence is that the request arrives scoped to one attribute, while the cost arrives scoped to the whole graph. The asymmetry between how the request is framed and how the cost lands is the entire phenomenon.

flowchart LR
    subgraph In["Inbound: ~200 dependents"]
        A1["Sales app"]
        A2["Support portal"]
        A3["Data platform"]
        A4["...196 more"]
    end
    A1 --> IDP
    A2 --> IDP
    A3 --> IDP
    A4 --> IDP
    IDP["Identity runtime<br/>Tier 0 · 99.99%"]
    IDP -->|"sync"| CRM["CRM · Tier 2"]
    IDP -->|"sync"| HR["HR system · Tier 3"]
    IDP -->|"sync"| LIC["Licence svc · Tier 3"]
    IDP -->|"sync"| MBX["Mail provisioning · Tier 2"]

The diagram is the argument. Every arrow leaving the identity node is inherited by every arrow entering it. Fan-in on the left multiplies the cost of fan-out on the right, and the two halves are usually reviewed by different people in different meetings.

Latency composes too, and it composes worse

Availability at least has the courtesy to fail visibly. Latency doesn't.

The point I want to make here is narrow, because I've made the general case elsewhere: The Identity Latency Budget walks through where 200 milliseconds actually goes, and Understanding Tail Latency in Authentication covers why serial chains and fan-in amplify rare slow events into common experiences. I don't want to re-derive either.

What's specific to the integration question is this: an outbound call's median cost is paid on every login, and its tail cost is paid by whichever unlucky users hit it — but both are multiplied by the same estate-wide fan-in as the availability cost. A 40ms CRM call that nobody would notice inside the sales app becomes 40ms added to the first interaction every employee has with every product, every morning.

And the tails don't average out. In a serial chain, the journey's tail is driven by the tails of its slowest hops, not by anyone's median. Adding a fourth dependency to a login path doesn't add its p50 to your p99; it adds another independent opportunity for a p99-class stall, and those compose toward the worst hop's p99.9.

There's a nastier version. A slow dependency hurts more than a dead one. A dead dependency fails fast — connection refused, and you're into your fallback in a millisecond. A dependency that has degraded to 9.8 seconds holds a request thread, a connection, and a slot in whatever pool you're using, for nearly the full timeout, on every affected login, while the retry storm you configured makes it worse. That's how the incident at the top of this article worked, and it's the standard shape: the identity service was healthy and simultaneously unable to serve, because all of its capacity was parked in await.

How to tell it's already happened

This is the part I'd actually run as an exercise. The transition from "identity platform with a couple of integrations" to "integration layer that also does authentication" has no announcement, but it has symptoms, and they're observable today without instrumentation:

  • Your identity service's outbound dependency list is longer than its inbound API surface. More systems it calls than endpoints it exposes. This one is nearly diagnostic on its own.
  • You cannot answer "what does a login touch?" in one breath. If it takes a whiteboard, the answer is already too long — and if two senior engineers give different answers, it's worse than too long, it's unknown.
  • Incidents in unrelated systems page the identity team. The CRM degrades and identity gets the first page, because identity is where it manifests. You are now the detection layer for other people's outages.
  • Onboarding a new customer requires the identity team to write code. Not configure — write. Integration work has become bespoke per-tenant work in the most shared component you own. (The Hidden Cost of Customer-Specific Features covers where that road goes.)
  • Your integration tests need six services running. The test suite is a proxy for the dependency graph. If you can't test login without a CRM stub, login depends on a CRM.
  • A deploy freeze on someone else's system blocks your release. Their change-management calendar is now your change-management calendar.
  • You have a retry queue. Hold that thought, because it's the important one.

The failure mode that makes it irreversible: orchestration acquires state

Most integrations start stateless. Call a system, read a value, put it in a token, done. If it fails, you've lost nothing but the value.

Then someone asks for provisioning, and everything changes — because provisioning writes.

Consider "create a mailbox on first login." The login path calls the mail provider. The call times out after 5 seconds. Did the mailbox get created? You don't know. So now you need to know, which means you need to record that you attempted it, which is a row somewhere. Then you need to retry, which means a queue. Then you need to not create two mailboxes when the user refreshes and logs in again, which means an idempotency key. Then the mail provider succeeds but the licence seat allocation fails, and you have a user with a mailbox and no licence, which is a state no one designed, so you need a compensating action to release the mailbox — and now you're writing a saga.

Read that paragraph again and notice what happened. The identity platform now owns a distributed transaction across systems it does not control. It has cross-system state: a half-provisioned user, a pending retry, an in-flight compensation. It has to answer questions like "what is the correct behaviour when the compensating action also fails," which is the hardest question in distributed systems, and it has to answer it in the login path.

This is the point of no return, and it's worth being precise about why. Stateless integrations are removable: delete the call, accept that a claim is missing, ship it Tuesday. Stateful orchestration is not removable, because other systems now depend on identity having done the writing. The mailbox provisioning pipeline has no other trigger. The licence count has no other reconciliation. Removing identity from the middle means building the thing that should have existed in the first place, while the current thing is load-bearing in production. That project needs a quarter and an owner, and it never gets one, because it delivers no visible feature.

"We'll refactor it later" stops being true at exactly the moment orchestration acquires state. Everything before that is a call you can delete.

And it drags the worst question along with it. If provisioning fails, does login succeed? Fail closed and a mail-provider incident locks out every new starter — and eventually every user, once someone extends the check to run on more than first login. Fail open and you've issued a token asserting entitlements that don't exist, which downstream systems will happily use for authorization. That dilemma is unpleasant enough when it's about a missing attribute; when it's about a partially-completed write across three systems, there is no correct answer available, which is the tell that the architecture put you somewhere you shouldn't be.

The resolution: emit, don't call

The rule is short. Identity should emit events outward and consume projections inward. It should not make synchronous outbound calls in the request path.

Concretely, the four requests from earlier resolve like this:

Request Wrong home Right home
Enrich token from CRM Sync call at issuance Attribute synced ahead of time (SCIM, scheduled push); token reads local state
Provision mailbox on first login Sync call in login path Provisioning service consuming a user.first_authentication event
Notify ticketing on contractor auth Sync call in login path Ticketing consumes the authentication event stream
Licence seat check Sync call at issuance Depends — see the next section; this one may be genuinely synchronous

The general shape: events out, projections in. Provisioning belongs to a provisioning system that subscribes to identity events. Enrichment belongs to attributes pre-computed and synced ahead of login. Orchestration belongs to a workflow engine — something explicitly designed to be slow, to retry for hours, to hold saga state, and to page a human when a compensation fails. Those are all legitimate, well-understood systems. None of them should be your authentication runtime, and all of them are allowed to have properties (slow, occasionally failing, stateful) that an authentication runtime is not.

If you're wondering how the events get out, that's the path in Why Every Identity Team Eventually Builds an Event Bus — the bus is the mechanism, the boundary is the reason.

The honest part: this moves complexity, it doesn't delete it

I want to be careful here, because the emit-don't-call answer is often presented as if it were free, and it isn't.

You have traded a synchronous failure for eventual consistency, and you now own:

A lag window. Between the HR system recording a promotion and the projection landing in identity, there is a period — seconds, usually; minutes when a consumer is behind; hours when nobody noticed a consumer was dead — during which the user's attributes are stale. For authorization-relevant attributes, that window is a security parameter, and you should be able to state it out loud.

The "logged in before the sync landed" case. A new starter authenticates at 08:58; their entitlement projection arrives at 09:03. What happens in between? This case is guaranteed to occur, it will occur on someone's first day, and it needs a designed answer rather than a discovered one.

Two disciplines make this manageable:

Fail secure on missing projected state. If the projection isn't there, the answer is "no," not "let me go and fetch it." The temptation to add a synchronous fallback read is enormous and it is exactly the thing you just removed — worse, it's a fallback that will be exercised for the first time during the incident where hammering the authoritative store is the worst available move. A missing projection means denied, logged, and visible. (The general form of that question is worked through in What Should an Identity Platform Do When Its Database Is Down?.)

Make the lag visible and bounded. Every projection carries the timestamp of the source event it was built from. Consumer lag is a first-class metric with an alert on it, expressed in seconds of staleness rather than messages queued — because "40,000 messages behind" means nothing and "eleven minutes stale" means something to everyone including your auditor. If you can't show a per-tenant staleness number on a dashboard, you don't have eventual consistency, you have hope.

The complexity is real. It's also located correctly: in a pipeline that can be monitored, backfilled, replayed and reasoned about offline, rather than in a request path where the only available responses are "wait" and "fail."

The cases that genuinely are synchronous

The argument would be dishonest if it claimed everything can be pre-computed. Some things can't, and pretending otherwise is how architects lose credibility with the people who have to ship.

Two categories are real:

A risk or fraud signal evaluated at authentication time. The whole value of "is this login anomalous" is that it reflects this login — the device, the IP, the velocity, the impossible-travel calculation. A pre-computed risk score from four hours ago answers a different question. If you're doing adaptive authentication, some part of it is synchronous by definition.

An authoritative check at the moment of issuance. Some licence and entitlement models genuinely require that the count is correct at the instant a token is minted, usually because a contract says so. "Eventually consistent seat counts" is not always an acceptable answer to a vendor audit.

For these, the containment rules matter more than the principle:

  1. A timeout tight enough to be a real budget, not a safety net. 200–300ms, not 10 seconds. The timeout is a latency budget line item, and it should be set from your login budget backwards, never from the dependency's p99 forwards. The incident at the top of this article had a 10-second timeout, which is another way of saying it had no budget at all.
  2. A circuit breaker that opens fast and stays open. Once the dependency is unhealthy, stop calling it. Every call you make into a degraded system after you have enough evidence to know it's degraded is pure cost, borne by users.
  3. The fallback must be a decision, not an error. This is the rule I'd tattoo on the design doc. When the risk engine times out, the answer is not a 500. The answer is a pre-agreed decision: "treat as elevated risk and require a second factor," or "treat as normal and log for later review." Someone with authority to accept the risk chooses it in advance, in writing, and it is tested. A fallback that hasn't been decided is a fallback that will be improvised by whoever is on call, at 3am, under pressure, and it will be wrong.
  4. Scope it to the smallest possible population. If the licence check must be authoritative, it should run at issuance for licensed applications, not on every authentication in the estate. Most "must be synchronous" requirements shrink dramatically when you ask which specific requests they actually apply to.

Two or three genuinely synchronous dependencies, each with a tight timeout, a breaker and a decided fallback, is a system you can operate. Seven, discovered during an incident, with a 10-second timeout and a fail-closed default, is not.

Where the line actually falls

Disclosure: I work on ClavionX, and the next paragraph is the same rule applied to our own internals rather than a product claim.

ClavionX splits into a control plane and a runtime, and the constraint is absolute: the runtime never synchronously calls the control plane during request execution, under any operational condition including failure. Configuration reaches the runtime as projections built from the control plane's domain events, and a missing projection fails secure rather than reaching back for a fresh read. That's the whole of the argument above, turned inward — the runtime is the most-depended-upon component, so it is given the fewest dependencies, including on us. The sanctioned answer for "I need an attribute from an external system in the token" is the same one I'd give anyone: pre-compute it via SCIM or a scheduled sync so that at login time it's a local read. It's a real constraint with real costs.

The general rule is the one to take away, and it's independent of anyone's product:

The component with the most dependents must have the fewest dependencies.

It sounds like a slogan. It's arithmetic. Availability composes multiplicatively down the chain, criticality propagates upward along dependency edges without asking permission, and identity sits at the top of the graph where both effects are multiplied by the size of your estate. Every integration you accept into the login path is not one integration — it's one integration times the number of applications in the company, and it silently promotes someone else's Tier 3 service to Tier 0 without their knowledge.

So the useful question when the next "while you're there, could you also..." arrives isn't can we do this? It's almost always yes. The question is: what tier does this system have to be, if we say yes — and has anyone told them?

Emit the event. Let something that's allowed to be slow do the work.