Every Identity Bug Is a Trust Bug

The deploy went out at 09:12 on a Tuesday. It changed one line in the assertion-consumer path: a normalization step so that a customer whose IdP emitted NameID with a trailing space would stop failing. Small, obviously correct, reviewed in four minutes.

At 09:41 the on-call engineer opened the admin console to look at the error rates and was redirected to a login page. She typed her password. The login page redirected her to the admin console. The admin console redirected her to the login page. She watched this happen eleven times before she understood what she was looking at.

The status page was behind the same SSO. So was the incident management tool, the runbook wiki, and the dashboard that would have told her which tenants were affected. The company had done the good, modern thing and put every internal system behind one identity provider, and that identity provider was now, for a subset of tenants, issuing sessions that its own relying parties would not accept.

They fixed it at 10:06. The redeploy took six minutes. That is where most retrospectives of this incident would end, and it would be recorded as a fifty-four minute partial outage.

It was not. Between 09:12 and 10:06, the service issued roughly 41,000 sessions with a malformed subject identifier. Twelve of them, in a tenant that used the subject as a lookup key against a locally-cached profile table, matched the wrong user. Those sessions were valid for eight hours; the last one expired at 18:04, a full working day after the "outage" ended. Nobody knew they existed until a customer's own security team asked, four days later, why a support agent's audit trail showed her opening tickets she had no memory of opening.

The fifty-four minutes cost them a bad morning. The twelve sessions cost them a disclosure, an eleven-week security review from that customer's team, and a renewal that came with a new set of contractual terms.

The asymmetry nobody puts on the architecture diagram

Almost every failure in a software system degrades a capability. Search gets slower. The recommendations widget doesn't render. Reporting is stale for an hour. The system is diminished but standing, and the tools you would use to understand and repair it are unaffected, because they sit somewhere else.

Identity does not fail that way, for a structural reason: it is the only component that every other component consults before doing anything at all. That is not an accident of your architecture, it is the definition of the thing. Authentication is infrastructure, and infrastructure failures are multiplicative rather than additive. An identity outage is not one product being down. It is every product being down at the same instant, with the same symptom, from the same cause, and there is a particular trap inside that:

The systems you need in order to fix an identity outage are, in a well-run organization, behind the identity system. Single sign-on for internal tools is correct security practice. It is also a circular dependency that nobody draws, because the arrow from "admin console" to "IdP" looks like every other arrow on the diagram until the day it is the only arrow that matters. The better your access hygiene, the tighter the loop. Teams with a sloppy legacy VPN and a shared password vault sometimes recover faster than teams with immaculate zero-trust posture, which is an uncomfortable thing to say out loud and worth saying anyway.

There is a second-order version that catches teams who have thought about the first. The break-glass credentials live in a vault whose unseal keys are held by people who need Slack — to be told which key to use — and Slack is behind SSO. Dependency chains for emergency access are almost never traced to depth two.

Blast radius is a calculation, not a feeling

"Blast radius" gets used as a mood word. For identity it's arithmetic, and doing the arithmetic changes decisions.

The exposure of an identity defect is approximately:

radius = (tenants affected)
       × (relying parties federated to those tenants)
       × (artifacts still outstanding when you fixed it)

The first two terms are the ones everyone counts, and they make identity bugs large. A defect in a shared code path is not scoped to a feature; it's scoped to your customer list. One tenant with forty federated applications means a single bad assertion path is forty application-level incidents in someone else's estate, and you will hear about it from all forty support queues at once.

The third term is the one that gets forgotten, and it is the one that makes identity bugs long.

A bug that stopped happening twenty minutes ago is still live inside every artifact it issued. Fixing the code stops production of bad tokens. It does nothing to the ones already in browsers, in mobile keychains, in server-side session stores, and in the refresh-token families that will keep minting fresh access tokens from a bad ancestor for as long as the family is allowed to live. If your access tokens are one hour and your refresh tokens are thirty days and rotation extends the family, the honest end of the incident window is not the redeploy — it is the redeploy plus the longest chain the artifact can produce. This is the same property that makes blue-green deploys strange for identity platforms: you are never draining requests, you are draining artifacts whose lifetimes are measured in hours to months.

Which yields a useful exercise. For each credential your system issues, write down: how long is it valid, can it be enumerated, and can it be revoked in bulk by the property that identifies the bad ones? That last clause is where most designs fail. Teams have revocation — per user, or global. Very few can express "revoke every session minted between 09:12 and 10:06 by the SAML path for tenants using connection type X," which is exactly the query an incident needs. Being able to answer it converts a nine-hour tail into a four-minute one, and it is a data-model decision (stamp issuance metadata onto the session record) made long before the incident, not a thing you can improvise during one.

The two failure directions are not symmetric

This is the central point, and the reason "identity bug" and "ordinary bug" should be different words in your head.

Identity gets exactly two things wrong. It can refuse someone it should have admitted — call it a false negative — or it can admit someone it should have refused. The instinct is to treat these as two sides of one dial. They are not remotely comparable events, and the differences are mechanical.

Locked out legitimate users Admitted the wrong principal
How you learn about it Immediately, from everyone You don't. Someone else tells you, later
Self-reporting Yes — users cannot use the product and say so No — the request succeeds and looks correct in every log
Visible in error rates Yes, as failures No. It is a success, by every metric you have
Effect of deploying the fix Incident over Incident begins
Time direction of the work Forward: restore service Backward: bound the exposure, indefinitely
Cost A bad day, some SLA credits Forensics, notification, disclosure, audit, renewal risk

The row that matters is effect of deploying the fix. For an availability failure, the fix is the resolution: users log in, the queue drains, you write the retro. For a correctness failure, the fix only stops the bleeding forward in time. The actual work — determining who got in, to what, for how long, and whether you can prove it — runs backwards from the moment of discovery, and it has no natural boundary. You do not get to declare it over. Someone else does.

And there is a specific mechanism that makes that backwards work fail, which experienced engineers routinely discover at the worst possible time: your log retention is a hard bound on how tightly you can scope an incident. If a subject-mapping defect has been live for ninety days and your authentication logs roll at thirty, you cannot say "eleven accounts were affected." You must say "we can confirm eleven in the last thirty days and cannot rule out others before that." Those are entirely different sentences to a customer's security team. The first is an incident; the second is an unbounded one, and unbounded incidents are what turn into notifications covering every user because you could not prove a smaller number.

That is the real argument for keeping authentication events longer than your other telemetry, and for keeping them in a form you can actually query by issuance properties rather than only by user. It is not a compliance checkbox. It is the difference between a scoped disclosure and a total one, and it costs storage you are otherwise tempted to trim. (The retention/cost trade is genuinely painful, and worth reading alongside the hidden cost of audit logs.)

What that asymmetry should buy

If one direction costs a bad day and the other costs a disclosure, they should not receive equal engineering attention. Three concrete consequences.

Spend test effort where the failure is silent. Most auth test suites are heavy on the happy path and on obvious rejections, because those are easy to assert. The dangerous cases are the ones where the system returns 200 and is wrong: an assertion from tenant A accepted for a user in tenant B, a token whose audience belongs to another relying party, a subject identifier that collides after normalization, a signature validated against the wrong key because the kid was absent and the library picked the first one it had. These are not exotic — several are on the list in seven identity bugs every team eventually ships. They share a property: no test that only checks "did the request succeed" can find them. They need assertions about who the resulting session belongs to and what it is scoped to, which is a different test shape than most teams write.

The cheapest structural fix in this article is a fixture change: put a second tenant in every test database, holding data no test is permitted to see, and assert its absence. With one tenant in the database, a missing tenant filter and a correct one return identical results — a single-tenant fixture set is incapable of detecting the highest-severity bug class you can ship.

Fail closed where the failure is silent; fail open where it is loud. Not globally — globally is the wrong altitude, and I've argued the per-operation version of this elsewhere in what an identity platform should do when its database is down. But the asymmetry above gives you the tiebreaker for cases where reasonable people disagree. If a path's failure mode is "someone can't work for ten minutes," bias toward availability. If a path's failure mode is "we may have admitted the wrong principal and won't find out," bias toward refusing, even at a real availability cost — because you can recover from the refusal and you cannot recover from the admission. Missing entitlement data, an unresolvable tenant, an assertion you cannot fully verify, a session whose binding you cannot confirm: refuse. The user files a ticket, which is a functioning feedback loop. The alternative has no feedback loop at all.

But notice where that reasoning inverts. Failing closed is not free and is not always safe, because a fail-closed control can itself become the attack. Account lockout is the canonical case: a control designed to prevent unauthorized access, implemented so that any anonymous party can deny service to any account for free. "Fail closed" is a good default for ambiguity about identity; it is a bad default for rate-based suspicion about behaviour, where the attacker controls the trigger. Knowing which of those two you're in is the actual skill.

Identity incidents escalate on a different track

Engineering leaders should understand this part explicitly, because the organizational dynamics are not the ones you've trained on and they arrive faster than you expect.

They arrive from every channel at once. An ordinary incident is reported by the users of one feature. An identity incident is reported by everyone, with mutually contradictory descriptions — "the app is down," "I'm getting a blank page," "it says my password is wrong," "it logged me in as someone else." Support cannot triage it, because the symptom set doesn't resolve to a component. The first fifteen minutes go to establishing that these are all one incident, and they are pure loss.

The customer's security team gets involved, not just their IT team. For a normal outage, your counterpart is an operations person who wants an ETA. For an identity incident — particularly any hint of the false-positive kind — your counterpart becomes someone whose job is to assess whether you are still an acceptable vendor. They ask different questions, they document the answers, and the conversation does not end when service is restored. Expect a written incident report with a defined structure, follow-up questions weeks later, and a re-run of the security review you thought you'd finished. If you've been through an enterprise security questionnaire, imagine one written specifically about the thing that just happened to you.

There is a legal clock running in parallel with the technical one. This is the part that surprises engineering leaders most. Unauthorized-access clauses in enterprise agreements commonly require notification within 24 to 72 hours of becoming aware, and regulators impose their own — GDPR Article 33 requires notification of a personal data breach to the supervisory authority within 72 hours of awareness, with reasons required if you're late. Those clocks start at awareness, not at confirmation. So the moment an engineer says "I think we may have crossed sessions between tenants," a timer has begun, and the incident commander now has a deliverable with a legal deadline that has nothing to do with restoring service.

The operational implication is specific and worth deciding before you need it: your incident process needs a step that notifies legal and account management on suspicion of a correctness failure, not on confirmation. Teams reliably get this backwards, because engineering culture correctly discourages raising alarms before you understand the problem. Here that instinct burns the clock you're being measured against. The right shape is a low-cost early signal ("we are investigating a possible cross-tenant issue, no findings yet") that starts the parallel process, followed by the technical work at its own pace.

What this should change about how you build and ship

This is the payoff, and it should be concrete or it isn't worth writing.

Synthetic login probes, per tenant and per connection. Not a health endpoint. A real end-to-end flow — redirect, credential or assertion, callback, token issuance, one authenticated call — executed continuously against production for every tenant, and for every federated connection within a tenant.

The reason is arithmetic. Tenant-specific breakage is invisible in aggregate metrics. A tenant representing 0.3% of your login volume going to 100% failure moves global login success from 99.95% to 99.65% — inside the normal daily variance of most systems, well below any threshold you could alert on without constant noise. That tenant is completely down and your dashboards are green. This is the standard way identity incidents come to be discovered by the customer rather than by you, and the only detector that works evaluates each tenant independently: the alerting predicate is per-tenant, not global.

The per-connection part matters just as much. Tenants have several IdP connections, and a certificate expiring on one breaks that connection alone; probing "can someone log into tenant X" with a local test account tells you nothing about the SAML connection that is actually failing. Probe the paths customers use, which means test principals on the customer side — an annoying provisioning ask that good customers agree to once you explain why.

Canary by tenant, not by percentage of requests. Percentage-based canaries assume failures are randomly distributed across requests. Identity failures are almost never traffic-shaped; they're configuration-shaped — they hit a connection type, a protocol variant, an attribute mapping, a tenant with an unusual setting. Against that, a 1% request canary is the worst possible instrument: it exposes 100% of tenants to a 1% chance of failure, producing intermittent, unattributable errors everywhere instead of a clear failure somewhere. Users experience it as "it sometimes doesn't work," which is both harder to diagnose and, for authentication, more corrosive to confidence than a clean outage.

Routing entire tenants to the new version — internal tenants, then design partners, then a widening ring — gives you the opposite properties: a contained population, a definitive signal, and an attributable rollback. It costs you a router that can pin a tenant to a version, and version-skew discipline on anything shared between the rings.

A break-glass path that does not depend on what it fixes. A locally-authenticated administrative route, on separate infrastructure, that can disable a connection, roll back a config, or revoke a token family without a working federation path. Three properties make it real rather than theoretical: it is exercised on a schedule (an unexercised break-glass path is a break-glass path that has silently broken); its use generates loud, unsuppressable alerts to people other than the user; and its dependencies are traced to depth two, so the credentials aren't in a vault that needs SSO and the runbook isn't in a wiki that needs SSO.

Rehearse "the IdP is down." Actually run it: block the connection in a staging environment that mirrors production federation, and walk through it with the on-call rotation. The findings are always the same and always surprising to the team having them — the runbook is behind SSO, the escalation contact list is in a tool behind SSO, and the person with the break-glass token left in March. This exercise has the best findings-per-hour of anything in this list.

Ship identity changes on their own. Not because identity code is special in some mystical way, but because rollback needs to be unambiguous. When a release contains one auth change and eleven feature changes and login success drops, the first twenty minutes go to determining what to roll back, and that determination happens while every product is degraded. A release containing one identity change has exactly one hypothesis and exactly one action. You pay a slower cadence for identity work; you buy a rollback decision that requires no analysis. Given that identity outages are total rather than partial, that trade is heavily in your favour.

A related discipline costs nothing: make the deploy of an identity change and the flip of its behaviour two separate events. Ship the code inert behind a flag, verify the fleet is on the new build, then enable per tenant. That separates "the deploy broke it" from "the change broke it" — diagnosed and reversed differently — and makes reversal a config change rather than a redeploy, at the moment when minutes are expensive.

Trust is asymmetric in time, too

The failure modes are asymmetric. So is the recovery of the thing they damage.

A customer's confidence in your identity layer is built by a long unbroken run of nothing happening and destroyed by one event. This isn't sentiment; it has a mechanism. Before the incident, your authentication layer was a settled question and nobody in the customer's organization was assigned to think about it. Afterwards someone is, with a standing agenda item and an obligation to report on it. The cost isn't the apology — it's that you converted an unmonitored dependency into a monitored one, and monitored dependencies get re-evaluated at renewal.

That is the honest economic argument for the slow discipline above. Per-tenant canaries, solo releases, and synthetic probes across every connection are over-engineering for a reporting feature — genuinely, unironically over-engineering, and you should not do them there. They are correct for identity because the expected cost of a defect is not "a bad day" but "a quarter of remediation work with a customer whose security team is now watching," and because you cannot buy that trust back on any timescale that matters to the current fiscal year. Investments that look irrational against incident frequency look obvious against incident cost multiplied by recovery duration.

The counterpoint: paralysis is also an outage

Taken too far, everything above produces a team that ships nothing, and I have watched that happen more often than I have watched a cross-tenant leak.

The failure looks respectable. Every change requires a rehearsal, a sign-off, and a window. Solo releases mean a queue; the queue means batching; batching means larger changes, which are riskier, which justifies more process — a feedback loop that ends in quarterly identity releases of forty commits, precisely the shape of change most likely to cause the outage the process exists to prevent. Meanwhile the routine certificate rotation becomes a project, the library with the known CVE stays two majors behind because upgrading it is "too risky to schedule," and the team's real security posture degrades while its process posture improves.

The discriminator is the blast-radius calculation, applied honestly per change:

Apply the full discipline Don't
Shared token issuance, session, and validation paths Internal tooling with a handful of employee users
Anything on the federation/assertion-consumer path Admin UI presentation over an unchanged API
Tenant resolution, attribute mapping, subject identifiers A new optional connector nobody has enabled yet
Key management, rotation, JWKS publication Docs, error copy, non-security log lines
Anything that can produce a silent wrong answer Anything whose failure is loud, contained, and self-reporting

The right-hand column is not "less important work." It is work whose failures are recoverable in minutes by the person who caused them, and treating it with ceremony borrowed from the left-hand column buys nothing and costs a lot — including, eventually, the goodwill you need for people to take the left-hand column seriously.

The two questions that sort a change reliably: if this is wrong, who finds out — us, or someone else? And: if this is wrong, does deploying the fix end it? Two "us"/"yes" answers and you're in normal engineering. A "someone else" or a "no" and you're in the other regime, where the boring discipline is the cheap option.

Where this lands

Identity is the only subsystem whose failure is simultaneously total, self-obscuring, and legally timed. Total, because everything consults it and the tools you'd use to repair it are usually behind it. Self-obscuring, because its worst failure mode is a successful request that nobody logs as a problem. Legally timed, because the clock on disclosure starts at suspicion and runs independently of your restoration work.

None of that makes identity engineering harder in the sense of requiring more cleverness. The protocols are written down, the crypto is library calls, and most of the bugs are mundane — a missing tenant filter, a cached TTL nobody chose, a normalization step that seemed obviously correct on a Tuesday morning. What differs is the cost function. Elsewhere a defect costs you the time to fix it. Here, one class of defect costs you the fix, plus the time to prove how far it reached, plus the limit of how far back your logs let you prove anything, plus a conversation with someone whose job is deciding whether to keep working with you.

Design for the direction of the failure, not just its likelihood. Count the artifacts still outstanding, not just the minutes of downtime. Probe every tenant and every connection separately, because your aggregate metrics are structurally blind to the failure that will actually happen to you. And keep the discipline scoped to the surfaces that earn it, so that it survives contact with a roadmap.

Users forgive slow. What they don't forgive is being told, weeks later, that you are still working out who was in their account.