Identity Platforms Are Mostly Caching Problems

Here's a claim that sounds wrong until you've operated one: an identity platform is not primarily a cryptography system, a protocol implementation, or a database application. It's a cache coherence problem wearing a security costume.

Here's a claim that sounds wrong until you've operated one: an identity platform is not primarily a cryptography system, a protocol implementation, or a database application. It's a cache coherence problem wearing a security costume.

The cryptography is well-understood and mostly library calls. The protocols are specified in RFCs you can read. The database schema is not especially interesting. But the question of what is cached where, for how long, and what happens when it's wrong — that's where the real engineering lives, and it's where the genuinely dangerous bugs come from.

The reason is a ratio.

The ratio that drives everything

Count the writes an identity platform receives in a day. A few thousand password changes. Some hundreds of admin configuration updates. A trickle of user provisioning.

Now count the reads. Every API call in every application that trusts you triggers a token validation. Every page load checks a session. Every token issuance reads tenant config, client config, signing keys, and policy. Every JWKS fetch from every resource server, repeatedly, forever.

The ratio in a mature deployment is somewhere north of 10,000:1.

No amount of database tuning fixes a 10,000:1 read amplification. You cache, or you buy a much larger database and cache anyway.

But identity data has a property that makes caching genuinely dangerous rather than routine: stale identity data is a security failure, not a stale UI. A cached product listing that's five minutes out of date is a minor annoyance. A cached authorization decision that's five minutes out of date means a user you fired at 2pm still has access at 2:04.

That tension — cache aggressively or die under load, cache carelessly and create security holes — is the actual design problem.

What's actually cached

Six categories, and they have wildly different risk profiles.

Signing keys. Read on every token operation, changed only during rotation. The most cacheable thing in the system and the most catastrophic to get wrong: a validator caching a key set that no longer contains the key a token was signed with rejects valid tokens, which looks like a total outage.

Tenant configuration. Password policy, MFA requirements, token lifetimes, branding, federation settings. Read on every single request in a multi-tenant system. Changed by admins occasionally. High cache value, moderate risk — a stale MFA policy means someone skipped a factor they should have been challenged for.

Client metadata. Redirect URIs, allowed grants, scopes. Read on every authorization and token request. This one carries real security weight: a stale allowlist of redirect URIs means a URI you removed because it was compromised might still be accepted.

Sessions. Read on every authenticated request. This is usually the whole point of a session store rather than a cache layered over one, but the distinction blurs when you add local caching on top — and that's where "logged out but still authenticated" bugs are born.

Authorization data. Roles, groups, permissions. Read constantly, and the single most dangerous thing to serve stale, because it's the data that directly answers "is this person allowed to do this."

JWKS, from the outside. Every resource server that validates your tokens caches your public keys, on infrastructure you don't control, with TTLs you didn't choose. You have no way to force an invalidation. This is worth pausing on, because most teams don't think of it as their cache — but it's the one that breaks rotation.

The layering that everyone converges on

Almost every mature platform ends up with the same two-tier structure, usually after trying one tier and finding it insufficient.

flowchart TD
    Req["Request"] --> L1["Local in-process cache\n(nanoseconds, per-instance)"]
    L1 -->|miss| L2["Distributed cache\n(~1ms, shared)"]
    L2 -->|miss| DB["Database\n(~10ms, authoritative)"]

Local caches are absurdly fast and impossible to invalidate precisely — every instance has its own copy, and telling all of them to forget something requires a broadcast mechanism you probably don't have. Distributed caches are invalidated in one place and cost a network round trip.

The pattern that works: short-TTL local cache for the hottest, most stable data (keys, tenant config), backed by a distributed cache with longer TTLs and explicit invalidation.

The local tier's TTL is not a performance parameter. It's a security parameter — it's the maximum time your system can serve a decision based on data that has already changed. Choose it that way. Sixty seconds might be fine for branding config and completely unacceptable for a revocation list.

The four problems that will actually bite you

Thundering herd

Your signing key cache expires. Two hundred runtime instances discover the miss within the same millisecond and all hit the database simultaneously.

This is the classic failure and it has a nasty property: it happens at expiry, which means it's synchronized across instances that all started at roughly the same time and therefore all cache-filled at roughly the same time. Your traffic pattern isn't smooth — it's a spike every TTL interval.

Fixes are well-known: jitter the TTLs so expiry spreads out, use a lock so only one instance refills, or refresh proactively in the background before expiry so a miss never happens on the request path. For signing keys — which change on a schedule you control — proactive refresh is clearly correct. There's no reason to ever take a cache miss on a key you knew was expiring.

Negative caching

A token arrives referencing a key ID you don't have. You look it up, find nothing, return an error.

If you don't cache that negative result, an attacker (or a misconfigured client in a retry loop) can send a thousand requests with bogus key IDs and generate a thousand database lookups. If you do cache it, a key that legitimately appears a moment later — say, mid-rotation — stays invisible for the duration of the negative TTL.

The resolution is asymmetric TTLs: cache negatives briefly (seconds), positives longer. And on a miss for something that should exist, refresh from source once before concluding it doesn't — that single retry is what makes key rotation survivable.

Invalidation across a boundary you don't own

An admin disables a user. How long until every runtime instance in every region stops honoring that user's sessions?

If the answer is "up to the local cache TTL," then your security posture includes a window during which a disabled user still works. That window might be entirely acceptable — sixty seconds is not a scandal — but it needs to be a number you chose, documented, and can state to a customer's security team without checking the code first.

The version of this that hurts is external JWKS caching. When you rotate a signing key, every resource server that validates your tokens is caching your old key set on infrastructure you don't operate. You cannot invalidate it. All you can do is publish the new key well before you start using it, so that by the time a token signed with it appears, every consumer's natural refresh cycle has already picked it up.

This is why key rotation is a scheduled, phased operation rather than an atomic one — publish, wait longer than the longest plausible consumer TTL, then start signing.

sequenceDiagram
    participant AS as Authorization Server
    participant RS as Resource Servers
    AS->>AS: Generate new key, add to JWKS
    Note over AS,RS: Wait — longer than any consumer's cache TTL
    RS->>AS: Routine JWKS refresh, picks up new key
    AS->>AS: Begin signing with new key
    Note over AS,RS: Old key stays published until old tokens expire
    AS->>AS: Retire old key

The cache that becomes the source of truth

This one is subtle and it's the one that causes outages.

You cache sessions in Redis. Redis is fast and everything works. Then Redis restarts, and every user on the platform is logged out at once — because Redis wasn't a cache, it was the only copy of that state, and nobody wrote that down.

The test is simple and worth running deliberately: what happens if you flush the entire cache right now? If the answer is "brief latency spike as it refills," it's a cache. If the answer is "everyone gets logged out" or "authorization fails until we restore," it's a database with an eviction policy, and it should be designed, backed up, and operated as one.

Where the reasoning leads

The thing that makes identity caching different from ordinary application caching is that the failure modes point in opposite directions, and both are unacceptable.

Cache too little and the system falls over — 10,000:1 read amplification against a database is not survivable at scale, and the latency makes every login feel slow.

Cache too much and you get security holes with a time delay: revoked tokens that still validate, disabled users who still authenticate, removed redirect URIs that still redirect. These bugs are especially unpleasant because they're intermittent by nature and they resolve themselves before you can reproduce them.

So the discipline is per-datum rather than global. For each thing you cache, three questions, answered explicitly and written down:

  1. What's the maximum time a stale value is tolerable? That's your TTL, and it's a security decision.
  2. What invalidates it, and how does that invalidation reach every tier? If there's no mechanism, TTL expiry is your invalidation, and see question 1.
  3. What happens if it's gone? Refill, or fail? Fail-open and fail-closed are both defensible and the choice is different per datum — but it must be a choice.

Teams that answer those three questions for signing keys, tenant config, client metadata, sessions, and authorization data end up with a system that's fast and predictable. Teams that add caching reactively, wherever the profiler pointed, end up with a system that's fast and occasionally, inexplicably, wrong.

The second failure mode is much harder to debug, because by the time you're investigating, the cache has expired and the evidence is gone.