Chaos Engineering for Identity Systems

The experiment passed. That was the problem.

A team I worked with ran their first real fault injection against the identity platform on a Thursday afternoon. The target was the session store: a Redis cluster holding roughly nine million live sessions. They cut it off at the network level for four minutes, with the whole team in a room watching dashboards.

Nothing happened. Request rate held. Error rate stayed at its usual 0.3%. Latency moved by single-digit milliseconds. Somebody said "well, that's reassuring" and they went back to their desks and wrote it up as a successful validation of the session-store failure path.

Six weeks later, during a real Redis incident, a support ticket arrived from a customer asking why a contractor account with no assigned roles had been able to open the finance dashboard. The reconstruction took two days and ended at a single method. Session lookup failed, the exception handler caught it, and the code fell through to a path that built a request context from the access token alone. The token was valid. The token was not the problem. The session carried the step-up assertion, the device-binding state, and the tenant's IP restrictions — and a failed lookup returned an empty restriction set, which the policy evaluator read the only way it could: nothing restricts this request.

The four-minute experiment had been perfectly successful at what it measured. Availability never dipped, because the system had found a way to answer every request. It just answered a large number of them yes when the correct answer was I don't know, so no.

That gap — between staying up and failing in the direction you chose — is the entire subject of this article. It is what makes chaos engineering on an identity platform a different discipline from chaos engineering on a recommendation service, and it is why most identity teams who adopt the practice adopt the metrics that will hide their worst bug.

Availability is only half the pass condition

The canonical chaos experiment asks one question: does the system continue to serve users when this component fails? That question is correct for almost every system. If a video-catalogue service degrades to a stale list of popular titles, staying up is unambiguously the right outcome, and availability is a sufficient measure of success.

Identity systems have a second axis, and it is orthogonal to the first. Every request produces a decision — allow, deny, or challenge — and that decision has a correct value that does not change just because a dependency is broken. A system that stays up by resolving every ambiguous decision to allow has achieved perfect availability and failed the experiment completely.

So the outcome space is a 2×2, not a line:

Refused correctly / served correctly Answered when it should have refused
Stayed available Pass The dangerous quadrant — looks like a pass on every dashboard
Went down Loud failure, honest, fixable Failed both ways

Most tooling can only see the rows. The top-right cell is invisible to error-rate graphs, invisible to synthetic uptime probes, and invisible to the SLO burn chart you will be staring at during the experiment. It is also, in my experience, the most common real outcome of an untested identity failure path — because a system under stress that has a choice between "return something" and "return an error" has usually been written by someone whose instinct was to return something.

Which means the hypothesis you write down has to change shape. A generic chaos hypothesis reads:

When Redis is unavailable, the login service continues to serve requests with p99 under 500ms and error rate under 2%.

An identity chaos hypothesis has to name the direction:

When Redis is unavailable, no request is served as authenticated on the basis of a session that could not be read. Session-bearing requests are refused with 401 session_unavailable. Self-validating access tokens continue to validate at normal latency. The deny reason is visible in logs as missing_state, not as a generic 500.

The second one can fail while the first one passes. That is the whole point. Write the pass condition so that "it stayed up" cannot satisfy it on its own. Every experiment in the catalogue below is stated in that form, and if you take nothing else from this article, take the habit of writing the fail-direction into the hypothesis before you break anything.

Fail-secure is a claim, not a property

Ask any identity team whether their system fails closed and you will get an immediate yes. It is a matter of professional identity — we're the security component, of course we fail closed.

Ask when they last observed it and the room goes quiet.

The reason the belief is usually wrong is structural rather than cultural. Fail-secure paths are the least-exercised code in the system. The happy path runs millions of times a day and is hardened by sheer repetition; every bug in it has been found by a user. The cache-miss-during-outage path has run in production maybe twice, both times at 3am, and nobody read the logs afterwards because the incident resolved itself. A branch with that execution history is not a fallback. It is an untested branch with a comment above it saying what the author hoped it would do.

The specific shapes are worth naming, because you can grep for most of them.

The empty result that reads as permissive. A lookup fails, the error handler returns an empty collection, and downstream code treats an empty collection as "no constraints apply." This is the bug in the opening story and it is by far the most common. The subtlety that makes it survive code review: whether empty is safe depends entirely on the polarity of the collection. An empty allowlist should deny everything. An empty denylist permits everything. Same value, same type, opposite security meanings — and a generic return Collections.emptyList() in an exception handler cannot know which one it is serving.

The catch-log-continue. An enrichment step throws, the handler logs a warning and proceeds with a partially populated context. This is correct behaviour when the enrichment is cosmetic (the user's display avatar) and catastrophic when it is not (the group membership that carries the deny binding). Almost nobody classifies enrichment steps by whether their absence is security-relevant, so the same handler covers both.

The default-constructed object on timeout. A remote call times out and the code proceeds with a zero-valued struct. Booleans usually default safe — an unset mfaSatisfied is false, which denies. Numbers and strings usually default unsafe: a riskScore of 0 reads as "no risk detected," a null tenantId matches a query with no tenant predicate, an unset maxSessionAge compares as unlimited. The dangerous defaults are the ones where the zero value is indistinguishable from a legitimate benign answer.

Missing versus empty. This one deserves its own paragraph because it is a data-model problem rather than a code problem, and you cannot patch your way out of it. A tenant that has configured no IP restrictions and a tenant whose configuration projection never arrived are, at the storage layer, both likely to produce {}. If those two states are not distinguishable in the data, no amount of careful error handling can fail secure, because the code genuinely cannot tell whether it is looking at a known-permissive configuration or at ignorance. The fix is that projected state carries a materialization marker — a version stamp, a built_at, an explicit envelope — such that absent is a distinct value from present and empty. This is exactly why the runtime/control-plane split I work on in ClavionX (disclosure: it is the platform I build) makes "fail secure on missing projected state" a stated architectural rule rather than an error-handling convention: the runtime never synchronously calls the control plane to resolve the ambiguity, so the ambiguity has to be resolved in the data model instead. And even with that rule written down and reviewed, the only thing that converts it from a claim into a property is deleting a projection in a controlled window and watching what the runtime actually does.

That last sentence generalizes. Architecture that specifies fail-secure behaviour is much better than architecture that doesn't. It is still a claim until an experiment has produced the refusal and someone has looked at it.

The experiment catalogue

What follows is the set I would run, roughly in the order I would run them, with each hypothesis in fail-direction form. For each one, the interesting part is not the injection — that is usually ten minutes of iptables or a service-mesh fault filter — but the specific bug it is designed to catch.

1. Session store unavailable

Break: Blackhole the session store (Redis or equivalent) at the network layer. Not a graceful shutdown — you want timeouts, not connection-refused, because those are different code paths and the slow one is worse.

Hypothesis: Requests presenting a session cookie are refused with an explicit session_unavailable reason. No request is served as authenticated on the basis of a session that could not be read. Self-validating access tokens continue to validate. New interactive logins fail at the point of session creation with a clean error, not a 500 after a 30-second hang.

Correct degradation: Token validation is unaffected, because it needs a public key and a clock and nothing else. Interactive login is down. Existing session-based requests are refused rather than downgraded. The blast radius is "people using the web UI," not "everything."

The bug you are hunting: "No session found" collapsing into "anonymous," and anonymous being an acceptable state on an endpoint that only looked public. Also: whether your session read has a timeout shorter than your caller's patience, because a session store that is slow rather than dead is the variant that takes down the whole fleet through pool exhaustion.

2. Event bus down or lagging

Break: Stop consumers, or better, throttle the bus so lag grows steadily to an hour. Lag is far more informative than outage, because outage is obvious and lag is silent.

Hypothesis: The runtime continues to serve from the projections it already has. Projection staleness is exposed as a metric and, past a configured threshold, as an alert. Stale-but-known state is served; absent state is refused. Revocation-sensitive operations — token introspection, session invalidation, tenant suspension — either continue to work through a path that does not depend on the bus, or fail closed with a stated reason.

Correct degradation: Everything keeps working, slightly out of date, and you know exactly how out of date. The event bus is the component whose failure should be least visible in the request path — that is the entire reason every identity team eventually builds one, and the reason it becomes load-bearing in ways nobody documented.

The bug you are hunting: A runtime that cannot distinguish "this projection is forty minutes old" from "this projection is current." If staleness is not measurable, you have no way to decide when stale becomes unacceptable, and you will discover during a real incident that an admin's tenant-suspension took effect nowhere. The second bug: consumers that, on restart, replay from the beginning and take twenty minutes to become useful while the runtime happily serves prehistoric state.

3. Config projection missing or version-skewed

This is the sharpest experiment in the catalogue and the one most teams have never attempted.

Break: In a controlled window, delete or withhold the projected configuration for one designated test tenant. Separately, pin one runtime node to a projection version several revisions behind the rest of the fleet.

Hypothesis: Authentication requests for a tenant with no configuration projection are refused, with a distinct reason code, and the refusal is visible per-tenant in telemetry. The runtime does not fall back to a global default, does not synthesize an empty configuration, and does not reach synchronously to the control plane for a fresh read. Under version skew, each node behaves consistently with the version it holds, and the skew itself is observable.

Correct degradation: That one tenant cannot log in. Every other tenant is unaffected. Somebody gets paged with a message naming the tenant.

The bug you are hunting: A runtime that authenticates users against a tenant whose policy projection never arrived. This is the exact bug worth building the experiment for. It happens when a new tenant is provisioned and the projection lags, when a consumer crashed mid-backfill, when a schema change silently dropped records that failed to deserialize. The user experience of the bug is successful login with no policy applied — no MFA requirement, no IP restriction, no session limits — which is indistinguishable from a working system right up until an audit.

Version skew deserves the separate arm because it is the realistic version during a deploy. Two nodes disagreeing about a tenant's MFA policy means a user's security posture depends on which node terminated their request, and that is a coin flip that will not show up in any aggregate metric. Related failure surface: the deploy that logs everyone out.

4. An upstream IdP that is slow, not down

Break: Inject 5–10 seconds of latency (not errors) into responses from one federated upstream — one enterprise customer's Entra ID or Okta tenant, or your test double for it.

Hypothesis: Logins through the affected upstream time out at a bounded deadline with a clear error. Logins through every other upstream are unaffected: their p99 does not move. Connection pools, thread pools, and outbound HTTP concurrency are partitioned such that one slow upstream cannot consume shared capacity.

Correct degradation: One customer's SSO is broken. Everybody else does not notice.

The bug you are hunting: A shared thread pool or a shared outbound connection pool. Partial failure is harder than total failure and enormously more common — a dead upstream fails fast and frees its resources, a slow upstream holds them. This is the classic path from a single-customer problem to a platform-wide outage, and if you are running federation for hundreds of enterprise tenants it is the single most likely cause of your next big incident. The lesson is bulkheading: per-upstream concurrency limits, per-upstream circuit breakers, and deadlines that are shorter than the caller's patience. This is one place where a slow dependency is strictly worse than a broken one, and where the distributed-systems reality described in why identity platforms become distributed systems stops being abstract.

There is a security dimension too, which is why this experiment belongs to security engineers and not only to SREs. When a federated assertion cannot be verified because the upstream's metadata endpoint is timing out, the tempting fix is a cached-metadata path with a generous TTL, and the tempting emergency fix is a bypass. Both are decisions about what you are willing to believe from a layer you cannot currently reach — the subject of what happens when the layer below you lies.

5. Database primary down, or replica lag

Break: Fail over the primary. Separately, hold replication lag at 30 seconds.

Hypothesis: Per-operation behaviour matches the degradation table you already wrote — token validation serves, introspection serves from cache and fails closed on a miss, password authentication and MFA verification hard-fail, admin writes are rejected loudly.

I am not going to re-derive that table here, because it exists: what should an identity platform do when its database is down is the design document, and this experiment is its test suite. The value of running it as chaos rather than as a design exercise is that you will find at least one row where the implementation disagrees with the table, and the disagreement is always in the permissive direction.

The bug you are hunting, specifically for the replica-lag arm: an operation that reads its own write from a replica and interprets the absence as a negative fact. A revocation that was written to the primary and not yet replicated reads as "not revoked." That is a fail-open with a clean audit trail saying everything worked.

6. Signing keys or JWKS unavailable at a rotation boundary

Break: Make the JWKS endpoint unreachable, or the signing key unavailable to the issuer, during an active rotation window when both the old and new key are meant to be published.

Hypothesis: Tokens signed with either key continue to validate from cached JWKS for the cache lifetime. New issuance either uses the key it can reach or fails cleanly; it never falls back to an unsigned or weakly-signed path. Validators do not accept a token whose kid they cannot resolve, and do not fetch JWKS inline on the validation hot path.

The bug you are hunting: the kid that cannot be resolved being treated as "try every key we have" or, worse, "skip signature verification for now." Also the availability version: a JWKS fetch on the validation path turns your busiest operation into a remote dependency at precisely the wrong time. The rotation mechanics are worked through in designing for key rotation before you need it and, for the federation half, SAML certificate rotation without an outage. Rotation is where this gets sharp for multi-tenant platforms: with per-tenant issuers and per-tenant signing keys, a rotation failure is scoped to one tenant, which is better blast-radius-wise and much easier to miss.

7. Clock skew on one node

Break: Push one runtime node's clock forward by 90 seconds. Then, separately, backward.

Hypothesis: Tokens issued by the skewed node are accepted by the rest of the fleet within the configured skew tolerance and rejected beyond it. Validation on the skewed node does not accept expired tokens. The skew is detected and the node is removed from rotation, or at minimum alerted on.

This one is worth doing deliberately because every part of the system quietly assumes wall-clock agreement, and nothing in the system checks it. Token validity windows, replay caches, session idle timeouts, TOTP verification, nonce expiry, and certificate validity all resolve to comparisons against now. A node 90 seconds fast issues tokens that its peers consider not-yet-valid; a node 90 seconds slow accepts tokens its peers have already expired, which is a small, real, boring authentication bypass that will never appear in any error metric. NTP failures are not exotic — a misconfigured VM, a hypervisor migration, a container image with no time sync — and the resulting behaviour is a fleet where security properties depend on which node you hit.

Running these without becoming the incident

Everything above is only responsible if the surrounding practice is. Some notes on that, in rough priority order.

Staging first, and staging will not be enough. Run every experiment in staging before production, and expect staging to pass experiments that production will fail. Staging has a handful of tenants where production has hundreds, a warm cache where production has a long cold tail, uniform data where production has one enormous customer with a deeply nested group hierarchy, and no real traffic mix. This is the same reason load tests lie, worked through in the art of load testing, and it applies with equal force here: a cache-miss path that never misses in staging is still untested after your staging experiment passes.

Blast radius is measured in tenants, not in percent. For a multi-tenant identity platform, "5% of traffic" is the wrong unit, because it distributes damage across every customer. The right unit is one tenant — ideally an internal one, or a designated canary tenant with real but non-critical usage, whose users you can warn. Per-tenant fault injection needs a real mechanism: a control that says for tenant X, treat the session store as unavailable. Build that as a first-class, operator-only, audited control. Do not build it as a general-purpose hook, because every extension point is a permanent public API and a fault-injection hook is the one you least want to become one.

The abort condition is written down before you start, with a name next to it. Not "we'll stop if it looks bad." Something like: abort if authentication error rate for any tenant other than the target exceeds 1%, if p99 on token validation exceeds 250ms, or if any request is observed being served with a permissive decision under missing state. The abort call belongs to Priya; she does not need to justify it. An experiment without a pre-agreed abort is not an experiment, it is an outage with an audience.

Game days beat automated chaos here, and this is not a maturity failure. Continuous automated chaos is the aspirational end state in most of the literature, and for availability-shaped systems it is right. For identity it is only partly available, because the pass condition includes was the refusal correct, and that judgement usually requires a human reading logs and saying "yes, that deny was for the right reason." An automated harness can assert on availability, and if you let it, it will quietly encode the wrong pass condition and give you a green build forever.

There is one automatable piece, and it is the highest-value monitoring artifact I know of in this space: the negative canary. A synthetic probe that attempts an operation which must always be refused — a token signed by a retired key, a request for a tenant that does not exist, a session ID that was revoked an hour ago, a step-up-required endpoint called without the step-up claim. Normal canaries alert when a success turns into a failure. A negative canary alerts when a failure turns into a success, which is exactly the signal that the top-right quadrant is now occupied. Run them continuously, per tenant, and page on them. They are cheap and they are the only continuous instrument that watches the direction rather than the rate.

Observability is a prerequisite, not an output. If you cannot tell, from the logs, which direction a request failed, do not run the experiment yet — you will learn nothing and you will have caused a real outage to learn it. Concretely, that means every authentication and authorization decision emits a structured outcome with a stable reason code, and that the reason codes distinguish denied by policy from denied because state was unavailable from allowed because no policy applied. Those three are frequently the same log line today. The middle one is the one this entire practice exists to make visible. Aggregate deny counts by reason and by tenant; the shape of that chart during an experiment is your actual result. This is the same instrumentation the on-call rotation depends on (running an identity platform on-call) and the same telemetry that makes cross-layer debugging tractable (debugging across four layers).

Pair the catalogue with the threat model. The experiments worth running are the ones that correspond to trust boundaries you have already reasoned about; threat modeling an identity platform is the map, and chaos experiments are how you check whether the mitigations in it survive contact with a broken dependency. A mitigation that only works when the system is healthy is not a mitigation; it is a feature.

The honest counterpoint

Chaos engineering on an identity platform costs more than it does almost anywhere else, and the case against it is not stupid.

You are deliberately locking real people out of their work. A failed experiment on a recommendation service degrades an experience. A failed experiment here means an emergency-room clinician cannot reach a records system, or three hundred people cannot start their shift. The blast radius is not proportional to your traffic share, because identity is the front door to everything. Any framing that treats these as equivalent risks is wrong.

It has a maturity threshold, and there is a correct order of operations. If you cannot currently answer "which tenant is failing and why" from a dashboard within sixty seconds, injecting faults is premature. The sequence I would defend: per-tenant observability with reason-coded decisions first; then a break-glass path that lets you restore service without a deploy; then a written degradation table so the experiment has a specification to test against; then staging fault injection; then a single production tenant with a human hand on the abort. Skipping to step four because chaos engineering is the interesting part is how a team turns a good practice into a self-inflicted incident and then never gets permission to try again.

Some faults should not be injected in production at all. Destroying a signing key, breaking the audit pipeline for a regulated tenant, or corrupting credential storage all have consequences that are not bounded by the experiment window. Model those; do not inject them. The point of the practice is to convert uncertainty into knowledge at an acceptable cost, and a few classes of fault have no acceptable cost.

And there is a real risk of chaos theater — running experiments that are safe because they test paths you already know work, producing a slide with a green checkmark and no new information. The opening story is the version of that failure that keeps me honest: an experiment carefully designed to measure the wrong thing will pass every time, and the confidence it produces is worse than having run nothing at all, because now nobody will look again.

What this comes down to

Every identity team has a stated position on failure behaviour, and almost every stated position is an untested claim about the least-exercised code in the system. The gap between the claim and the behaviour is not usually a design error. It is an empty collection, a catch block, a default-constructed struct, and a data model that cannot tell nothing configured from nothing known.

You will not find those by reading the code, because they all look correct in a diff. You find them by breaking a dependency in a controlled window and asking two questions instead of one: did it stay up, and did it refuse the right things for the right reason.

Start with the config projection. Delete one test tenant's configuration and watch what your runtime does with a user who tries to log in. If the answer is that they get in, you have just learned the most valuable thing this practice has to offer, and you learned it on a Thursday afternoon of your choosing rather than during someone's audit.