Identity Is Mostly Read Traffic — Design Accordingly
A team I spoke to had a login endpoint with a p99 of 420ms and no idea why.
They had done the obvious work first. Tenant config was cached. The password hash was Argon2id tuned to about 90ms, which they knew was deliberate and correct. They had a read replica. They had traced the request and confirmed there was no accidental N+1 lurking in the credential check. And still: 420ms, reliably, every weekday between 08:45 and 09:30, and a support queue full of people saying the app "feels slow when I sign in."
The profile, when they finally got a good one, was embarrassing in the way that good profiles usually are. Roughly a third of the tail was three UPDATE statements.
One set last_login_at on the user row. One reset the failed-attempt counter to zero. One touched the session's sliding expiry. Each of them costs almost nothing in isolation. All three ran inside the login transaction, against the primary, and at 09:00 several thousand people were updating rows in the same few pages of the same two tables at the same time. The replica they had added — the one that was supposed to absorb the read load — was absorbing nothing, because the login path was not a read.
That's the shape of the problem this article is about. Identity systems are overwhelmingly read-dominated in traffic and frequently write-dominated in cost, because of two habits: designs that turn reads into writes, and designs that compute at read time what could have been decided once, at write time. Both are invisible in code review. Both show up as a latency graph with no obviously guilty query on it.
The ratio is a budget, not a caching argument
The read/write asymmetry in a mature identity deployment lands somewhere north of 10,000:1 — every API call validating a token, every page load checking a session, every issuance reading tenant config, client config, keys and policy, against a trickle of password changes and admin edits. I've argued elsewhere that this makes identity platforms mostly caching problems, and I don't want to re-run that argument.
I want to point at a different consequence of the same number, one that almost nobody spends.
An asymmetry that large is a budget. If reads outnumber writes 10,000:1, you can make a single write dramatically more expensive and the aggregate barely moves. Call one read one unit of work:
| Cost of a write, relative to one read | Increase in total system work |
|---|---|
| 1× | +0.01% |
| 10× | +0.09% |
| 100× | +1.0% |
| 1,000× | +10% |
| 10,000× | +100% |
You can do a hundred times more work on every write than you do on a read and pay one percent for it. You can do a thousand times more and pay ten percent. In exchange, you get to delete work from a path that runs ten thousand times more often.
Almost no identity system spends that budget. Most spend it backwards.
The reason is that the default data model everyone starts from is write-optimized, and nobody thinks of it that way. Third normal form is a write-time optimization: store each fact exactly once so that changing it is a single-row update with no risk of divergence. That is an excellent property for the 0.01% of operations that change things. It is paid for, every single time, by the 99.99% that only want to know the answer.
What "write-optimized" costs the read path
Here is the canonical example, and it is in more systems than you'd think.
A user belongs to 14 groups. Groups nest four levels deep. Each group carries role bindings, each role carries scopes, and the tenant has an entitlement policy that grants or denies on top. "What is this person allowed to do" is therefore a recursive traversal joined against three or four other tables, producing a couple of hundred rows that get collapsed into a claim set.
Warm, with the right plan, that query takes 12ms. That looks fine in a dashboard. But it is the query whose cost is most sensitive to things you don't control: buffer cache state, planner choice as row counts drift, one tenant with an unusually deep hierarchy, replication lag causing a fallback to the primary. When it degrades it doesn't degrade to 20ms; it degrades to 300ms, for one tenant, at their busiest hour. And it runs on every login and every token refresh, which is to say it runs on the read path forever.
The alternative is to stop storing the inputs to the decision on the read path and start storing the decision.
When a group membership changes, recompute the affected users' effective entitlement documents and write them down. The login path then does one key lookup and gets a finished answer: no joins, no recursion, no policy evaluation, no planner. You have moved work from a path that runs constantly to a path that runs rarely, which is exactly the trade the ratio was offering you.
This is not free, and the bill is worth naming honestly:
- You now own a projection, with all that implies: rebuild procedures, backfill after a bad deploy, a version stamp so consumers can discard stale copies, and reconciliation so a dropped event becomes a delay rather than permanent divergence. That's the argument in Building Identity Like Kubernetes and Why Most Identity Systems Need (At Least) Four Data Stores, and it's real work.
- You have a staleness window, and for authorization data that window is a security parameter you have to be able to state out loud.
- You've inverted the fan-out. One write can now touch a very large number of derived records. More on that below, because it's the sharpest edge on this whole design.
What you get back is a read path with almost no variance in it, which is the property users actually perceive. Nobody notices a fast median. Everybody notices the login that took two seconds.
The reads that are secretly writes
This is the part I find most teams have never audited, and it's why the 10,000:1 ratio is often a fiction in their own system.
The ratio describes traffic shape. It says nothing about what your implementation does with that traffic. A surprising amount of identity code performs a write on operations that are, from the outside, plainly reads:
| Operation the caller thinks is a read | The write it actually performs | Why it hurts |
|---|---|---|
| Session validation | Sliding-expiry touch | A write on every authenticated request — the single hottest path you have |
| Successful login | last_login_at update |
Row and page contention concentrated in the morning surge |
| Failed login | Attempt-counter increment | Writes on the exact path an attacker is trying to flood |
| Token refresh with rotation | Insert new token, revoke old, transactionally | Legitimately a write, and unavoidable — but often on the primary, synchronously |
| Introspection with usage tracking | last_used_at on the token record |
Turns a cacheable read into an uncacheable write |
| Risk-based auth | "Device seen before" record | A write per sign-in, usually to the same store as everything else |
| Login audit record | Append to the audit trail | The one write here that is genuinely non-negotiable |
None of these were bad decisions. Each was a small, reasonable feature: show the user their last login, expire idle sessions, lock accounts after five failures. Collectively they mean the busiest path in the system takes a write lock, and your read replica is decorative.
The useful question to ask of each one is narrow and answerable: does the correctness of this response depend on that write being durable before you reply?
For last_login_at, no. It can be asynchronous, batched, or coalesced. For device-seen records, no — a lost one costs you one extra MFA prompt. For refresh-token rotation, absolutely yes; reuse detection is the whole point and it must be atomic. For the audit record, yes, if your compliance posture says audit is fail-closed. Sorting the list this way usually collapses seven synchronous writes into two.
The sliding session touch deserves its own note, because it's the one with the highest volume and the easiest fix. You do not need to persist a touch on every request; you need the idle window to remain meaningfully tracked. The platform I work on, ClavionX, records this as an explicit rule rather than an optimization someone remembered to apply: session touch is throttled, and a touch is skipped entirely if less than a threshold has elapsed since the last one, with that threshold set at roughly 10–30% of the idle timeout (ADR-0007). The idle window stays accurate to within that threshold, and write amplification on the hottest path in the system drops by an order of magnitude or more. I mention it as a worked example of the rule, not a product claim — the arithmetic is the same wherever you implement it.
The same reasoning shows up in a less obvious place. Validating a client secret with Argon2id is CPU work, not a database write, but it has the identical shape: an expensive operation on a path that repeats constantly with the same input. ClavionX's answer (ADR-0008) is a derived-validation cache — a short-lived entry keyed by an HMAC of the presented secret, recording that this secret validated for this client at this security_version, invalidated immediately on secret rotation, client disable, tenant suspension or a security_version advance. The secret isn't stored and the hash isn't weakened. The full derivation just stops happening thousands of times per minute for the same unchanged credential. Pay once, at the boundary where the thing actually changed.
Where the read-heavy assumption breaks
If the argument stopped here it would be too clean. There are four situations where "identity is mostly reads" is straightforwardly false, and each one has bitten a real system.
Password spray and credential stuffing invert the ratio. These attacks look like reads at the API boundary and behave like writes underneath: every failure increments a counter, and lockout counters are hot single rows keyed per user or per IP — precisely the contention pattern a distributed attacker generates for free. They are simultaneously a CPU exhaustion attack, because failed passwords are hashed too. Two defences follow: keep those counters in a store built for high-contention increments with TTLs rather than in your user table, and reject what you can before the hash. (The lockout counter itself is a weapon; see Account Lockout Is a DoS Vector.)
Token issuance is a write, and it's on the hottest path. Authorization codes, PKCE verifiers, nonce and replay records, refresh-token families — all created and consumed at protocol speed. This is the distinction that matters most and that gets blurred constantly: the read-heavy claim is about configuration and identity data, not protocol state. They are two separate taxonomies with opposite profiles, and conflating them is how session state ends up in the configuration store. Worth designing the split explicitly: in ClavionX, projected configuration is written only by the event processor from the control plane's domain events, while operational protocol state is written directly by the runtime and never travels through the event pipeline at all (ADR-0003). Same Redis, deliberately different keyspaces, different owners, different failure semantics.
Fan-out inverts the arithmetic. One admin action — disabling a group, changing a role's scopes, suspending a tenant — can invalidate or rewrite tens of thousands of derived records. Your write path is not sized by write rate; it's sized by the largest single fan-out it must absorb. And revocation is the case where the fan-out must also be fast, because the whole point is that access stops. A design that materializes aggressively must have an answer for "an admin just changed something that affects 40,000 users" that is better than "it'll catch up eventually." Usually that answer is a version stamp on the aggregate — invalidate one tenant-level or group-level version rather than 40,000 individual documents — which trades a wider blast radius on refill for a bounded revocation time.
Writes cluster; reads don't. The 10,000:1 ratio is an average over a day. During the 02:00 SCIM sync it might be 3:1. During a quarterly access review, a compliance team modifies more entitlements in four hours than the platform sees in the other three months. Sizing the write path for the average is how you get the failure mode where logins queue behind a bulk import, which is the single most common version of this whole problem. And there's a cold-start variant: a newly provisioned tenant has no projection yet, so its first requests take the slow path. That's acceptable if you decided it — ClavionX loads projections lazily and accepts elevated first-request latency for cold tenants rather than pre-loading thousands of tenant configs at startup (ADR-0008) — and it's an outage if you didn't.
Separate the paths as failure domains, not just as modules
Everything above pushes toward one structural conclusion: the read path and the write path in an identity platform have different availability requirements, different scaling curves, different consistency needs and different acceptable deploy cadences. Systems that acknowledge that in the deployment topology behave very differently from systems that acknowledge it only in package names.
flowchart LR
Admin["Admin / SCIM / HR sync"] --> CP["Write path<br/>normalized, transactional<br/>consistency over throughput"]
CP -->|"domain events"| Proj["Projection builder<br/>pays the cost once"]
Proj --> RM["Read models<br/>denormalized, per-tenant<br/>the answer, not the inputs"]
RM --> RT["Read path<br/>logins, tokens, sessions<br/>availability over freshness"]
RT -.->|"never synchronously"| CP
The dotted line is the load-bearing part. If the read path can call the write path synchronously, it has inherited the write path's availability, and every argument above collapses — because the fallback will be exercised for the first time during exactly the incident where hammering your administrative database is the worst available move. ClavionX makes this an absolute rule rather than a guideline: the runtime never synchronously calls the control plane during request execution, under any operational condition including failure, and a missing projection fails secure rather than reaching back for a fresh read (ADR-0002). I'm disclosing that as the system I work on; the reasoning is general, and other platforms arrive at the same rule independently. The related question of what "fail secure" should mean per code path is worked through in What Should an Identity Platform Do When Its Database Is Down?.
There is a cost, and you should expect the ticket. An admin changes a setting, refreshes, and sees the old value, because the console they're looking at reads a projection that hasn't caught up. Every projection-based identity system gets this bug report. The correct fix is not to make the runtime read authoritative state — it's to have the admin UI read from the write side, which it can, because it's already talking to the write side to make the change. Read-your-writes for administrators; bounded staleness for the runtime. Two different consistency contracts for two different audiences, chosen deliberately.
The last consequence is one that catches people out. A 10,000:1 ratio means your load tests exercise the read path and essentially nothing else. The write path — the projection builder, the fan-out, the bulk import, the reconciliation — is the least-tested code in the system and the code most likely to be running during your worst hour. If you take one operational habit from this: run your next load test with a full SCIM sync executing concurrently, and watch the login p99. That number, not the clean one, is your actual latency.
So: count your writes on the read path and justify each one. Move the cost of a decision to the moment the decision changes, not the moment someone asks about it. Size the write path for its largest fan-out and its worst burst, not its average rate. Keep the two paths in separate failure domains, with no synchronous edge from the fast one to the slow one. And treat the ratio as what it actually is — a standing offer to trade work you do ten thousand times for work you do once.
Most identity systems decline that offer without ever noticing it was made.