What Should an Identity Platform Do When Its Database Is Down?
The first time I watched this happen, the postmortem took four hours and produced one useful sentence.
The first time I watched this happen, the postmortem took four hours and produced one useful sentence.
Primary Postgres failed over. Fifty seconds of unavailability, which is a nothing event — the kind of thing you'd expect to see as a small dip on a dashboard and never think about again. Instead the identity service returned 500s for eleven minutes, and because every product in the suite validated tokens by calling it, all of them went down too. When the database came back the service stayed down for another six minutes, because forty thousand queued retries arrived simultaneously and knocked over the connection pool.
The useful sentence, from someone at the back of the room: "We never decided what this thing should do when it can't reach storage. So it does whatever the framework does."
That's the whole subject. Every identity platform has degradation behaviour. Almost none of it was chosen. It's the emergent sum of exception handlers, default timeouts, and whatever your HTTP client does when a connection pool is exhausted — and the result is uniformly worse than what a mediocre engineer would design in an afternoon, because the failure mode of "unhandled" is always fail loudly and identically for every operation.
Identity is the wrong place for uniform failure. The operations you serve differ enormously in how much they need fresh state.
Fail-open and fail-closed are the wrong frame
The instinct is to hold a philosophical position: we're a security system, so we fail closed.
It sounds rigorous and it doesn't survive contact with the operation list. Fail closed on token validation and you take down every application in the company because a database that has nothing to do with token signatures is unreachable. Fail open on password verification and you've built an authentication bypass with a very unusual trigger condition.
The frame that actually works asks a different question, per operation:
What is the worst thing that happens if I answer from stale data?
For token validation, the answer is that you accept a token revoked in the last few minutes. Bad, bounded, and already true — you accept revoked JWTs for their remaining lifetime under normal conditions too.
For password verification, the answer is that you authenticate someone whose password was changed or whose account was disabled. Unbounded, because you can't know how stale is stale, and the whole point of the operation is to be a security boundary.
Those are different answers, so they get different behaviour. Which means the deliverable isn't a philosophy. It's a table.
The table
Assume the primary datastore is unreachable and you have a cache that may or may not hold what you need. Here's the shape I'd defend in a design review:
| Operation | Behaviour | Why |
|---|---|---|
| Validate a signed token (JWT) | Serve from cached keys | Needs the JWKS, not the database. Should not touch storage at all. |
| Token introspection (opaque) | Serve from cache; fail closed on miss | A cached "active" answer is stale by seconds. An uncached one is unknown, and unknown is not active. |
| Refresh token exchange | Fail closed | Rotation requires a durable write. Serving without it either breaks detection of reuse or issues tokens you can't revoke. |
| Password authentication | Fail closed | Requires current credential state and current account state. No safe stale answer exists. |
| MFA verification | Fail closed | Same, plus replay counters need writes. |
| SSO to an existing session | Serve if the session store is healthy | Often a different datastore. If it is, this survives. |
| New user registration | Fail closed | Pure write. Nothing to degrade to. |
| Admin/config writes | Fail closed, loudly | The most dangerous thing you can do here is accept the write and lose it. |
| Consent lookups | Serve from cache; deny on miss | An absent consent record is a "no." |
| Audit event emission | Buffer, bounded; then shed the request | See below — this is the one people get wrong. |
Three rows in that table are worth arguing about properly.
Token validation should never have been coupled to the database
Under load, token validation is 90–95% of your request volume. It is also the operation with the weakest dependency on fresh state: verifying a JWT signature requires a public key and a clock.
If your validation path touches the database, that's a design accident — usually one of two. Either you're loading tenant configuration on each validation (fixable with a projection and a cache), or you're checking a revocation list synchronously (a real decision, with real costs, that you should make deliberately rather than inherit).
Getting this right is the single highest-leverage change available, because it converts your worst incident class from "everything is down" into "logins are down." Those are very different incidents. In the first, every product in the suite is dark and every executive is in the channel. In the second, users who are already working stay working, and the blast radius is people arriving at the front door during the outage window.
The corollary is that your signing keys must be cached with a long TTL and refreshed in the background, never fetched inline on a cache miss during validation. A JWKS fetch on the hot path is the same bug in a different coat: you've made a high-volume, low-dependency operation depend on a remote system being reachable at exactly the wrong moment.
Introspection is the interesting case
Opaque tokens are usually dismissed here — "they need a database call per request, so they're dead in an outage" — and that's only true of a naive implementation.
Introspection results are cacheable. A 30-second cache on introspection responses removes almost all database traffic under normal operation, and during an outage it means you can serve every token you've seen recently from local state. Your revocation latency becomes cache TTL plus propagation, which is a number you chose, rather than "the remaining life of the JWT," which is a number the token chose for you.
The rule that matters: a cache hit serves, a cache miss fails closed. Serving unknown tokens as active during an outage is how you turn a database incident into a security incident, and it will be defended in the moment with "it's only for a few minutes." A cached affirmative is a fact that was true recently. A cache miss is not a fact at all.
Audit logging is where fail-closed goes to die
Here's the question that splits a room: your audit pipeline is down and someone tries to log in. Do you let them?
Compliance says no — an unlogged authentication is an unauditable one, and for privileged operations that's genuinely unacceptable. Availability says yes, obviously, because you're about to take down authentication for the entire company over a logging problem.
Both are right, which means the answer is per-operation again, and it's the one row of the table that must be decided by someone who can sign off on the compliance consequences.
What works in practice is a bounded buffer with an explicit overflow policy, and different policies by event class. Ordinary authentication events buffer, and if the buffer fills, they're dropped with a counter incremented and an alert fired. Privileged operations — permission grants, admin changes, key rotations, anything a regulator will ask about — block on a durable write, because the entire value of those records is that they exist.
The failure I've seen most often is the unbounded buffer, which converts a logging outage into an out-of-memory crash roughly fifteen minutes later. That's strictly worse than either honest answer, and it's the default in every implementation where nobody thought about it.
Slow is worse than down, and this is the part people miss
Everything above assumes a clean failure: the connection is refused, you know immediately, you take the degraded path. That's the easy case, and it's not the case you'll get.
The one you'll get is a database that responds in eight seconds instead of four milliseconds. Nothing is down. Health checks pass. Every request is technically succeeding.
And your service is dead, because connections are held for two thousand times longer than budgeted, the pool exhausts, requests queue, callers time out at thirty seconds and retry, and the retries add load to a system that is already the bottleneck. The classic metastable failure: retries become the load, and removing the original trigger doesn't fix it, because the retry storm now sustains itself.
Three things prevent this, and all three have to be in place beforehand:
Timeouts shorter than your caller's patience. If your database timeout is 30s and your caller times out at 10s, you're doing work nobody will read while holding resources nobody can use. Database timeouts belong in the hundreds of milliseconds for identity operations. Anything slower has already failed from the user's perspective.
Circuit breakers on every remote dependency. After N consecutive failures, stop trying and take the degraded path immediately. This is what makes the table above fast instead of merely correct — degradation that takes eight seconds per request to decide is not degradation, it's a slower outage.
Load shedding you can turn on. When you're saturated, rejecting 30% of requests immediately keeps 70% working. Serving all of them slowly serves none of them. This is deeply counterintuitive to everyone in the incident channel and needs to be a switch that exists before you need it, because nobody builds it at 3am.
Recovery is a separate design problem
The eleven-minute outage in my opening story became seventeen because nobody had thought about the moment the database came back.
Everything that was retrying hit at once — synchronized, because they'd all failed at the same instant and all backed off on the same schedule. Every cache had expired during the outage, so the first wave was 100% misses. Connection pools filled instantly. The database, freshly recovered and with cold buffers, fell over again.
The mitigations are unglamorous and well known, and the reason they're absent is that recovery gets no attention in design reviews: jitter on every backoff so clients don't synchronize; request coalescing so a thousand simultaneous misses for the same key become one query; staggered cache warming for the handful of keys you know will be needed — signing keys, tenant configuration, the ten largest tenants' settings; and gradual pool ramp rather than opening every connection the instant the socket accepts.
Write "how do we come back" into the same design document as "how do we degrade." They're the same feature, and only one of them ever gets built.
Making degradation a mode rather than an accident
The version of this that works long-term isn't a scatter of try/catch behaviours. It's an explicit read-only mode: a state the platform can be put into, deliberately, by an operator or by a health check, in which write operations are rejected with a clear error and read operations are served from cache.
That gives you three things you don't otherwise have. Consistent behaviour, because there's one code path rather than fifty exception handlers. A testable artifact — you can enter read-only mode in staging on a Tuesday afternoon and watch what happens, which is the only way anyone ever finds out that the admin console becomes unusable rather than degraded. And a usable operator tool, because sometimes you want to shed writes deliberately, during a migration or a suspected data-integrity problem.
Architecture helps here more than any amount of error handling. Strict separation between the runtime path and the control plane is worth building for exactly this reason: ClavionX's runtime never calls the control plane synchronously — configuration arrives as projected state through events, and the runtime fails secure when the projection is missing rather than reaching back for a fresh read. The property that buys you is the one that matters during an outage: the control plane can be entirely down, and token validation and session lookups keep working, because they were never on the same fate line to begin with. That's less an argument for a particular product than for the general shape — anything on the hot path that synchronously depends on your management plane will eventually take an outage it had no business taking.
What to actually do
Write the table. Yours, for your operations, with your team's names on the decisions. It's an hour of work and it will be the most-referenced page in your runbook.
Then test it, because a degradation path that has never executed is a hypothesis. Kill the database in staging during business hours, with the team watching, and check the behaviour against the table row by row. You will find at least one operation that fails in a way nobody predicted — the usual culprits are a health check that itself queries the database and takes the whole instance out of rotation, or an admin endpoint that returns 200 with empty results instead of an error.
The goal isn't surviving a database outage unscathed. That's not available. The goal is that when it happens, the outage is the shape you chose: logins fail, existing sessions continue, tokens keep validating, admin writes are cleanly rejected with an honest error, and the whole thing recovers when storage does — instead of a uniform wall of 500s and a four-hour postmortem to discover that nobody ever decided anything.