Designing for the Next Million Logins: What Breaks First, in Order
The outage started nine minutes after the capacity increase.
The team had done everything right, by the usual standards. A large customer was going live on Monday and their own estimate said traffic would roughly triple. So on Friday afternoon the platform team took the runtime from six instances to twenty — a comfortable margin, applied early, with a whole weekend to watch it. The deploy was clean. Health checks green. Login p50 unchanged at 96ms.
At 09:14 on Monday, login p99 went to eleven seconds. Not errors. Latency. The database was at 22% CPU and serving every query it received in under 4ms. The application dashboards showed no slow queries, no GC storms, no CPU pressure anywhere. The only anomaly anyone could find was a metric nobody had a panel for: pg_stat_activity sitting flat at 497 out of a max_connections of 500, for hours, like a painted line.
Twenty instances. A connection pool with a maximum of 50 per instance. A database that accepts 500 connections and reserves three for superusers.
They had caused the outage by adding capacity. The scale-up was the incident, and it was waiting patiently in the configuration for three days before Monday morning traffic arrived to trigger it.
I have watched some version of this happen enough times to believe the ordering is not random. When an identity platform goes from where it is now to ten times that, the things that break tend to break in a particular sequence, for reasons that are mechanical rather than mysterious. And the thing at the top of most teams' preparation list — raw CPU for the password hash — is almost never the thing that breaks first.
This article is about that order, and about the property that produces it.
Linear costs are budgetable; non-linear costs are incidents
Here is the distinction I want to organize everything else around.
Password hashing at 100ms of Argon2id per verification is expensive, and it scales linearly and visibly. At 50 logins per second you need roughly 5 cores of pure hashing capacity; at 500 you need 50. You can put that in a spreadsheet, and the spreadsheet will be right. Better: the cost announces itself gradually as you approach it, because queueing latency rises smoothly with utilization long before anything falls over. (The queueing arithmetic — expected wait scaling as ρ/(1−ρ) — is worked through in Understanding Tail Latency in Authentication, and I won't re-derive it here.)
A cost that grows linearly and telegraphs its approach is a budget item. You forecast it, you buy it, and if you get it slightly wrong you get a warning first.
Every failure in the list below is different in kind. Each one is flat, flat, flat — and then a step function. There is no gentle region where the graph bends and someone notices. The system is fine at 6× and unrecoverable at 7×, and nothing in the intervening dashboards distinguishes those two states. That is what makes them incidents rather than line items: not that they are expensive, but that they arrive without a ramp.
So the useful question to ask before a 10× is not "what is expensive?" It's "what is currently flat that will stop being flat, and at what multiple?"
| Password hashing CPU | The four below | |
|---|---|---|
| Shape | Linear in login rate | Step function |
| Warning | Latency rises smoothly with utilization | None until the threshold |
| Forecastable | Yes, on a napkin | Only if you measure the ratio |
| Failure mode | Slow, then slower | Fine, then outage |
| Fix under pressure | Add instances | Usually a config or architecture change |
Teams prepare for the first column because it's the one they can compute. That's exactly why it isn't what breaks.
1. Connection pools
This is first, and the reason it's first is an inversion that stays counterintuitive even after you've been bitten by it: the connection pool is sized per instance, so horizontal scaling multiplies your connection count against a database whose limit is fixed. Adding capacity to the stateless tier consumes a scarce resource in the stateful tier. Every other scaling lever you have makes this one worse.
The arithmetic. Your ceiling is instances × pool_max versus max_connections − reserved. Six instances at 50 is 300 against 497: fine, 1.7× headroom. Twenty instances at 50 is 1,000 against 497: an outage. Note that the traffic never entered the calculation. The threshold is crossed by the deploy, not by the load — the load merely reveals it by causing enough concurrent checkouts to actually claim the connections the config permits.
Worse, most pools hold a minimum idle set. Twenty instances with minimumIdle: 10 occupy 200 connections at three o'clock in the morning with zero traffic. You can exhaust a database from a completely idle fleet.
The mechanism, and why it presents as latency. A saturated pool does not return an error. It queues. The caller waits for a connection to be returned by whoever has it, and acquisition wait is unbounded while query time stays exactly where it was. This is why the database looks healthy during the incident: it is healthy. It is serving 4ms queries all day. The three seconds your users are experiencing were spent in a queue inside your own process, on a metric that most default dashboards do not chart.
What it looks like in production. Active connections pinned at a flat ceiling — a horizontal line, not a noisy one, which is the tell. Application p99 climbing while database p99 does not. Then, at the very end, errors that finally name the problem: HikariPool-1 - Connection is not available, request timed out after 30000ms, or from the other side, FATAL: sorry, too many clients already / remaining connection slots are reserved. By the time you see the second message you have been failing for a while, because the first thing that happened was that everything got slow.
The fix is smaller pools, not bigger limits. Little's Law gives the pool size you actually need: concurrency = throughput × service time. At 4,000 queries per second with a 2ms mean, that's 8 concurrent connections. Most pools are sized 5–10× larger than the workload can use, because 50 felt like a safe number and nobody did the multiplication. Shrinking the pool improves throughput under load — a smaller pool queues in your application, where waits are visible and boundable, instead of thrashing the database's own scheduler.
Beyond that, put a transaction-mode pooler (PgBouncer, RDS Proxy, ProxySQL) between the fleet and the database so that connection count decouples from instance count entirely. This is the one item in this article that I think is worth building before you need it, and I'll come back to why in the counterpoint section.
Two related multipliers to check while you're in there: read replicas do not help if the login path performs writes, and it usually does — see Identity Is Mostly Read Traffic for the audit of reads that are secretly writes. And every additional data store you talk to has its own pool with its own per-instance multiplier (Why Most Identity Systems Need (At Least) Four Data Stores is the map of how many of those you're likely to have).
2. Cache hit ratio
The second failure is the one whose arithmetic people consistently get wrong, and the error is always in the same direction: treating a hit ratio as a percentage that degrades smoothly.
It does not degrade smoothly, because the underlying variable isn't a percentage — it's whether your working set fits in memory. While the set of keys actually being requested in a given window fits, nearly every request after the first is a hit and the ratio sits near its ceiling, barely moving. The moment the working set exceeds available memory, eviction starts removing entries that are about to be requested again, and the ratio falls off a cliff. There is no meaningful middle.
The arithmetic that matters. Backend load is driven by the miss rate, and the miss rate is the small number, so small absolute changes in hit ratio are enormous multiples of load. 99% to 95% sounds like a 4% degradation. It is a 5× increase in requests reaching your database — 1 in 100 became 5 in 100.
Now compose that with the growth that caused it:
| Today | At 10× traffic, same cache memory | |
|---|---|---|
| Requests | 1,000/s | 10,000/s |
| Hit ratio | 99% | 95% |
| Requests reaching the database | 10/s | 500/s |
| Multiple of today's backend load | 1× | 50× |
Ten times the traffic produced fifty times the database load. That is the whole failure in one row, and it is why teams who sized their database for "10× headroom" find it saturated at 10× traffic. Their headroom calculation assumed the cache would keep working exactly as well as it does now.
There's a latency multiplier stacked on top. In identity, a miss is rarely one row read. It's a tenant config load, a client metadata lookup, a signing-key fetch, and possibly a policy compile — the difference between roughly 1ms and several tens of milliseconds, as itemized in The Identity Latency Budget. So a 5× increase in misses is also a large increase in time spent per miss, which occupies connections, which is how failure 2 detonates failure 1.
The leading indicator is not hit ratio; it's eviction rate. Hit ratio is a lagging, averaged number that stays reassuring while the cliff approaches. Evictions per second sits at exactly zero for as long as the working set fits, and becomes non-zero the instant it doesn't. That transition — evicted_keys moving off the floor, or your local cache's eviction counter doing the same — is the earliest signal available, and it typically precedes user-visible symptoms by days or weeks. Alert on "evictions > 0 for 15 minutes", not on a hit-ratio threshold.
Why growth outruns memory faster than you expect. The working set in a multi-tenant identity platform scales with active tenants × their active users × their distinct clients and devices, not with request volume. Signing up a hundred small tenants adds very little traffic and a great deal of working set. This is also why load tests reassure you falsely: a fixture with five tenants has a working set of approximately nothing and a 100% hit rate by the second minute, which The Art of Load Testing treats as one of the two standard ways a load test lies to you.
The deeper design question — what is cached, for how long, and what a stale entry means when the datum is an authorization decision — is Identity Platforms Are Mostly Caching Problems. Here I only want the capacity consequence: the cache is a load-bearing structural element, and its failure mode is a cliff, not a slope.
3. A single hot tenant
The third failure is the one your dashboards are structurally incapable of showing you, because aggregation is the thing that hides it.
Growth is not uniform. "10× traffic" is a statement about a sum, and sums in multi-tenant platforms are dominated by their largest terms. In practice a 10× year is often one enormous customer growing 100× while the long tail grows 2–3×. The tenant that will hurt you is frequently one that was 3% of traffic when you did the capacity plan.
Why aggregate metrics can't see it. A tenant that is 2% of your traffic and experiencing 100% failure sits comfortably inside your p99 — global p99 is p98-and-better for everyone else. Their users are having a total outage; your SLO dashboard is green; your on-call has nothing to look at. Per-tenant p99 is a fundamentally different graph from global p99, not a drill-down of it.
The mechanism that makes this a hard cliff rather than a load problem. Three of them, actually, and they compound:
- The shared cache becomes a noisy neighbor. One tenant's working set can evict everyone else's. Their growth degrades other tenants' hit ratios, which is failure 2 arriving for customers who did nothing.
- Per-tenant work is not uniformly sized. A tenant with 300,000 group memberships and a four-level nested hierarchy runs a different query plan from one with forty groups. Their cost per login can be 50× the median while their request count looks ordinary.
- The sharding key is the real ceiling. If you partition by
tenant_id— and most identity platforms eventually do — then a tenant is indivisible. When one tenant's load exceeds what one shard can serve, horizontal scaling stops working for that tenant specifically, and no amount of added capacity helps. This is the point where the fix stops being operational and becomes a migration.
What it looks like. One shard or one node at 85% CPU while its siblings sit at 20%. A support escalation from one named customer that your metrics contradict. Rate limits sized "per tenant" that a single tenant now legitimately exceeds all day, which turns your capacity control into a customer-facing outage (Rate Limiting Identity Is a Capacity Problem is the fuller argument).
The number to track is your largest tenant's share of total traffic, and its growth rate relative to the platform's. If the biggest tenant is 8% today and growing 3× annually while the platform grows 1.6× annually, then in two years they are 8% × 9 / 2.56 ≈ 28% of everything, on a shard they can't be split off. That's a two-year warning available today from data you already have — the sort of thing Why Enterprise Customers Cost More Than Consumer Users treats from the commercial side.
4. The audit and event pipeline
The fourth is the one that surprises people most, because it's the part of the system nobody considers performance-critical.
Audit writes scale with everything. Not with logins — with logins × events-per-login, and both factors grow. Ten times the logins with events-per-login drifting from 14 to 20 (because you added MFA, risk scoring, and consent evaluation in the same period that you grew) is a 14× increase on the pipeline. The economics of that volume are The Hidden Cost of Audit Logs; I'm not re-deriving them. What matters here is the capacity mechanism.
It is synchronous somewhere nobody remembers. Everyone knows not to INSERT into the audit table inside the login transaction, and most teams removed that years ago. What replaced it is "publish to the event bus and return", which is asynchronous only while the producer's buffer has room. When the downstream consumer falls behind and the buffer fills, the producer's send() blocks — and that block is executing on the authentication request thread. The abstraction that was supposed to decouple the two systems turns out to have a rope in it, and you discover the rope at the worst moment.
The same shape hides in a transactional outbox whose relay is behind (the outbox table grows, its index degrades, and the login transaction that writes to it slows down), and in any "fire and forget" call implemented with a bounded queue whose full-behavior is block rather than drop.
What it looks like. Login p99 spikes that correlate perfectly with consumer lag and not at all with login rate — the tell that the causality runs backwards from what you assume. Producer-side errors like Expiring N record(s) for auth-events-0: 30000 ms has passed since batch creation. A backlog that recovers linearly at a rate slower than it accumulated, so the ten-minute stall takes forty minutes to drain, during which every additional minute of peak traffic extends it.
The fixes, in order of leverage: make the producer's full-buffer policy an explicit decision — for high-volume, low-forensic-value events, dropping is correct and should be counted, not blocked on. Classify events at emission so the low-value 80% can be sampled or aggregated rather than carried at full fidelity. Give the pipeline enough sustained headroom that it can drain a burst faster than it accumulates, because a pipeline sized for the mean can never catch up. And decide out loud whether audit is fail-open or fail-closed for the auth path, since the default in most codebases is "fail-closed by accident, discovered during an incident" (What Should an Identity Platform Do When Its Database Is Down? covers how to make that choice per path).
The second tier
These arrive later, or in a different order depending on your architecture, but each has the same flat-then-step shape.
Token and session store memory grows faster than your user count. Session storage scales with users × devices × concurrent sessions × lifetime, and the multipliers move independently of user growth. The day a mobile app ships, per-user device count goes from 1.2 to 3.5 and refresh-token lifetime goes from an 8-hour browser session to 30 days, so storage per user increases roughly 10× with zero change in user count. Add refresh-token rotation, where old token records are retained for reuse detection, and a portion of your storage is now proportional to issuance rate × retention window — decoupled from user count entirely. See The Cost of Session Storage. The failure is an eviction policy quietly logging people out, or an OOM in a store that everyone assumed was a cache and is actually the only copy of that state.
JWKS fetch amplification is normally trivial and occasionally unbounded. Steady-state fetches scale with resource servers × their instances / TTL, which is a small number that has nothing to do with your user traffic: 200 services × 20 instances every 5 minutes is 13 requests per second. The step function is key rotation. An unknown kid triggers a synchronous fetch, and if every validator's cache invalidates in the same window, fetch rate briefly equals request rate across your entire ecosystem, against an endpoint sized for 13 per second. The Cost of Key Rotation and Designing for Key Rotation Before You Need It cover the phased publication that avoids this; the capacity point is that this endpoint's load is uncorrelated with everything else you're monitoring.
Lock contention on last_login-style updates. Per-user row locks are not the issue — each user owns their row. The issue is physical: at 09:00 several thousand updates land in the same handful of index pages of the same two tables, generating page-level contention, index churn, and write-ahead-log volume that scales with login rate rather than with anything anyone considers a write workload. Shared counters (per-tenant, per-IP) are strictly worse because they are single hot rows. This is the mechanism Identity Is Mostly Read Traffic opens with; at 10× it stops being a tail-latency curiosity and becomes the reason your primary can't keep up.
Synchronized token expiry. Deploy at 09:00, every client gets a one-hour token, and at 10:00 they all refresh together. The herd's amplitude scales with fleet size while your capacity is sized on the mean, so the peak-to-mean ratio gets worse precisely as you grow, and every mass restart re-synchronizes it. Jitter the lifetimes. It's covered in The Art of Load Testing; I mention it because it's the cheapest item on this entire list to fix and one of the most common to still be present at 10×.
The scaling checkpoint
None of the above requires a load test to predict. Every one of them is a ratio you can compute this afternoon from data you already have, and the ratio tells you which failure you'll meet first.
Build one table. Current value, value at 10×, and the hard limit.
| Ratio | How to compute it now | You are in trouble when |
|---|---|---|
| Connection headroom | (instances × pool_max) ÷ (max_connections − reserved) |
> 0.6 today, or > 1.0 at your planned instance count |
| Pool right-sizing | p99 in-use connections ÷ pool_max |
< 0.2 — your pool is oversized and you're hoarding the limit |
| Working set vs memory | distinct cache keys touched in 24h × mean value size ÷ cache maxmemory | > 0.5, and immediately if evictions/sec > 0 |
| Miss amplification | (1 − hit ratio) × request rate vs database capacity |
project it at your target rate with hit ratio 4 points lower |
| Largest-tenant share | biggest tenant's requests ÷ total, plus its growth rate vs platform growth | > 15%, or growing faster than the platform |
| Per-tenant tail | p99 computed per tenant, worst-tenant vs global | worst-tenant p99 > 3× global p99 |
| Audit fan-out | events emitted ÷ logins, times login rate, vs pipeline sustained throughput | > 0.5 of sustained throughput at peak |
| Backlog drain rate | stop a consumer for 10 minutes; time the recovery | drain takes longer than the burst that caused it |
| Session bytes per active user | store size ÷ 30-day active users, tracked as a trend line | trending up while user count is flat |
Two of those are experiments rather than queries, and they're the two worth actually doing. Stopping a consumer for ten minutes and timing the drain tells you whether your pipeline can ever catch up — a property no steady-state metric reveals. And computing per-tenant p99 for your top twenty tenants, once, will usually surprise you enough to change what you monitor. Both fit inside an afternoon; deliberately breaking things to learn this is the subject of Chaos Engineering for Identity Systems.
The honest counterpoint: measuring is cheap, pre-solving is not
Everything above could be read as an argument for building for 10× today. It isn't, and I want to be precise about the difference, because premature scaling work has a real and recurring cost: complexity you operate every day, in exchange for a failure you may never meet.
A tenant-sharded, multi-region, fully decoupled architecture at 5,000 logins per day is not a prudent investment. It's a permanent operational tax paid on a forecast, and — this is the part that gets missed — it's a tax paid on your guess about which bottleneck matters, which is the exact judgment this article argues teams routinely get wrong. Pre-solving the wrong thing costs you twice: once in complexity, and once in the false confidence that you've done your scaling work.
So I'd split the list explicitly.
Worth doing in advance, because retrofitting them under load is genuinely hard:
- Connection topology. Introducing a transaction-mode pooler is a day of work when nothing is on fire and a delicate surgery when everything is. It also removes the multiplication that makes every future scaling action dangerous, which is unusually high leverage for the effort.
- Getting audit off the synchronous path, with an explicit full-buffer policy. This is architectural. Changing it during an incident means changing durability semantics under pressure, which is how audit records get lost.
- Choosing a partition key you could actually shard on later. You don't need to shard. You need to not have a schema that forbids it, and to know whether your largest tenant fits in one node.
- Jitter on token lifetimes and cache TTLs. Trivially cheap, and it eliminates a category of self-inflicted spike outright.
A dashboard is sufficient, because the fix is cheap once the number moves:
- Cache sizing. More memory is a config change and a restart. What you need in advance is the eviction alert, not the capacity.
- JWKS and key-rotation load. Prefetch and phased publication are changes you make when rotation is scheduled, and rotation is always scheduled.
- Session store growth. Watch bytes-per-active-user as a trend. Act when the trend bends, which it will do months before the store fills.
- Hot-tenant detection. Per-tenant percentiles cost you cardinality in your metrics system and nothing else. You need the visibility early; the isolation work only when a specific tenant justifies it.
The general rule: spend engineering effort in advance on things whose fix requires changing a contract — a topology, a durability guarantee, a partition key. Spend a dashboard on things whose fix is a number in a config file.
The thing that actually happened
Go back to the opening. The team was ready for Monday. They had capacity-planned for a triple, correctly forecast the CPU they'd need for the hashing load, and provisioned generously. Every number in their plan was right.
They were destroyed by a ratio that appeared in none of it: 20 × 50 > 497. It was not a load problem, it was not visible in any test, and it was fully determined three days before it fired.
That's the pattern worth taking away. The costs teams prepare for are the ones that grow smoothly, because smooth growth is what a capacity plan knows how to represent. The costs that take you down are the ones that are flat until they aren't — a pool that's fine at 497 connections and fatal at 500, a hit ratio that holds at 99% until the working set exceeds memory by one byte, a tenant that fits on a shard until they don't, a buffer with room until it hasn't.
You cannot budget for a step function. You can only find out where the step is, which is nearly always a division you can do today with numbers you already collect.
Do the division before you add the instances.