Rate Limiting an Identity Platform Is a Capacity Problem, Not a Security One

At 09:02 on a Monday, an identity platform I know of started returning timeouts on /token for every tenant it served. Not one tenant. All of them.

The incident channel did what incident channels do and went looking for the attack. The traffic graph had a clean 6× step at 08:50 and stayed there. Somebody pulled the source IPs, expecting a botnet, and got corporate egress ranges in Ohio. Somebody else pulled the usernames, expecting a spray, and got a normal-looking distribution of real employees at one customer. There was no attack. There was a customer who had run a company-wide forced password reset on Friday afternoon, which invalidated every session in their estate, which meant that on Monday morning eleven thousand people all typed a new password into a login form inside the same twenty minutes.

Every one of those was a legitimate, correct, authorized request. Each one cost about 90 milliseconds of Argon2id. Multiplied out, one customer's routine administrative action asked for roughly four times the CPU the entire platform had provisioned for password verification, and the platform did what an unbounded queue in front of a saturated CPU pool always does: it accepted all of it and served none of it. Queue wait crossed four seconds. The mobile SDK timed out at three. Every timed-out request had already been hashed — the work was complete, the answer was correct, and nobody was still listening when it arrived. Then the SDK retried, and the retries queued behind the retries.

The global rate limiter, set at 8,000 requests per second, never fired. Peak was 3,100.

That's the whole argument of this article in one scene. The limit that would have saved them is not a security control and cannot be denominated in requests. It is a capacity control, it has to be denominated in the cost of the work, and it has to be enforced per tenant — because the thing that took the platform down was one customer consuming a shared, finite, deliberately expensive resource, entirely in good faith.

A quick scope note so we're not litigating the wrong thing: this article is not about detecting credential stuffing, locking accounts, or deciding whether a login attempt is malicious. Those are real problems with their own literature, and Account Lockout Is a DoS Vector covers why the obvious answer there is usually the wrong one. Here, intent is irrelevant. The capacity layer cannot tell an attacker from a Monday, and — this is the useful part — it doesn't need to.

The arithmetic nobody writes down

Start from the parameter you already chose deliberately. Argon2id, tuned so a single verification takes about 90ms of CPU on your production instance type. That number is not an accident or an inefficiency; it is the security mechanism, and the identity latency budget spends a large chunk of itself on it for good reason.

Now turn it into capacity.

90ms per verification, single-threaded
  → 1 core sustains 11.1 verifications/second
  → 500 logins/second peak needs 45 cores of pure hashing

Forty-five cores, running flat out, producing nothing but "yes" and "no." And you cannot provision 45 cores for a 500/s peak, because at 100% utilization queue wait goes to infinity. The queueing curve is unforgiving here and I won't re-derive it — Understanding Tail Latency in Authentication walks through why ρ/(1−ρ) means a CPU-bound auth service wants to live at 30–50% utilization. Take the friendlier end: you need roughly 100–110 cores to serve a 500/s login peak with a tail you'd be willing to publish.

Two things follow that most capacity plans miss.

The memory parameter bounds your concurrency independently of your cores. Argon2id at m=64MiB means every verification in flight holds 64 mebibytes of scratch. Forty-five concurrent hashes is 2.9 GiB of resident memory doing nothing but resisting GPUs. If your auth pods are provisioned at 4 GiB, your real concurrency ceiling is about 60 hashes regardless of how many cores the node advertises — and you will hit it as an allocation failure or an OOM kill, not as a graceful slowdown. Worse, Argon2's memory-hardness means it saturates memory bandwidth, so per-core throughput degrades as you add concurrent hashers on the same socket. The 11.1/second figure is a single-hasher measurement; measure it again at your actual target concurrency before you multiply.

The asymmetry is about three orders of magnitude. An attacker — or a badly-behaved SDK, or eleven thousand people in Ohio — sends a ~400 byte POST. Generating and sending it costs the sender maybe 20 microseconds of CPU. Servicing it costs you 90 milliseconds. That's a leverage ratio around 4,500:1. A single laptop on a home connection can sustain 2,000 of those requests per second on about 6 Mbit/s of upstream, which asks for 180 cores of hashing against your provisioned 45. You do not need a botnet. You need one enthusiastic person, or one customer with a maintenance window.

And here is the part that makes this a capacity problem rather than a security one: those two scenarios are byte-for-byte identical at the layer that has to decide. Same endpoint, same payload shape, same plausible credentials, same corporate IP ranges. Any defence that requires first establishing malice has already lost, because by the time you have established it you have done the work.

Requests are the wrong unit

Here's the practical core of the article, and it's a single sentence: rate limits on an identity platform should be denominated in cost, not in requests.

The reason is that the spread between the cheapest and most expensive operation on an identity API is enormous — larger, I think, than on almost any other kind of service. Pick a base unit of 100 microseconds of service CPU and lay the operations out against it:

Operation Typical service cost Cost units Notes
JWT validation, cached JWKS ~60 µs 1 ES256/RS256 verify plus claim checks. Pure CPU, no I/O.
Introspection, cache hit ~200 µs 2 See introspection isn't as slow as you think.
Introspection, cache miss → store read ~1.5 ms 15 Cost is I/O wait, not CPU — it consumes a different resource.
WebAuthn assertion verification ~2 ms 20 Signature verify is microseconds; the lookups dominate.
Authorization code → token exchange ~3 ms 30 Signing dominates.
RFC 8693 token exchange ~5 ms 50 Validate, apply policy, re-sign.
SAML response processing ~8 ms 80 XML canonicalization and signature verification.
Password login (Argon2id) ~95 ms 950 The outlier.
SCIM bulk import, per record ~15 ms + writes 150+ Plus projection fan-out downstream.
Admin write with fan-out 50 ms + N 500–50,000 Sized by blast radius, not by request count.

Read the two bolded rows together. A password login costs roughly a thousand times what a token validation costs. So a limit that says "5,000 requests per second per tenant" is not one limit. It is a permission slip for either 5,000 token validations — about half a core-second — or 5,000 password logins, which is 450 core-seconds of demand arriving every second. The same number authorizes two workloads that differ by a factor of a thousand in what they ask of you.

Every request-denominated limiter I have seen in production has this bug. It is usually invisible for years, because the traffic mix is stable: 99% cheap validations, 1% expensive logins, and the request limit is implicitly a cost limit at that mix. Then the mix shifts — a forced password reset, a session-store flush, an SSO outage at an upstream IdP pushing everyone to local passwords, a mobile release that stopped caching tokens — and the limiter goes on faithfully permitting the same request rate while the actual demand goes up 50×.

Denominating in cost fixes this structurally. You assign each route a cost class, you charge the tenant's budget the cost at admission time, and the limit becomes a statement about the resource you actually own: this tenant may consume 20 core-seconds per second. Under that limit, 200,000 token validations and 210 password logins are the same amount of permission, which is correct, because they are the same amount of work.

Three implementation notes that matter more than they look:

Charge the estimate at admission, reconcile the actual after. You know the route and the tenant before you do the work, so you know the cost class. Charge it up front — that's the entire point, since a budget you debit after the work is finished has already let the work happen. Then, on completion, apply the delta between the estimate and the measured cost. Over a few seconds the bucket converges on truth, and a class whose estimate is systematically wrong shows up as a persistent drift you can alert on.

Cost classes are not a static table. Reclassify from production telemetry on a schedule. Argon2 parameters change, hardware changes, and a tenant with a 400-group hierarchy makes token issuance cost four times what the table says.

Different resources need different budgets. The introspection row above is 15 units of I/O wait, not CPU. If you collapse everything into one currency you'll eventually let a tenant exhaust the connection pool while their CPU budget sits half-used. Two or three budgets — CPU, database concurrency, and outbound calls to third parties like SMS or an upstream IdP — is usually enough, and it's a lot fewer than the number of dashboards you already have.

Admit or reject before you schedule, not after

The Monday incident had a second failure stacked on the first, and it's the one that turned a throughput shortfall into a total outage.

A throughput shortfall is survivable and bounded: you can serve 500 logins per second, 2,000 arrive, and 1,500 people should get a fast, honest "try again in a moment." What actually happened is that all 2,000 were accepted into an unbounded queue. Arrival exceeded service, so the queue grew without bound, so wait time grew without bound, so every request — including the ones inside the 500 you could have served perfectly — eventually exceeded the client's deadline. An unbounded queue in front of a saturated pool doesn't degrade the excess; it converts a partial outage into a complete one. Nobody gets served, and you burn 100% of your CPU proving it.

So the queue must be bounded, and the bound is computable rather than a matter of taste. Little's Law gives it to you directly: if your hash pool services 500/s and you are willing to tolerate 200ms of queue wait, the queue may hold 500 × 0.2 = 100 items. Item 101 is rejected immediately, at a cost of microseconds, before it consumes anything expensive.

Two refinements are worth knowing, because they're the difference between a bounded queue that works and one that merely fails politely.

Timeouts alone do not save you. This is the trap the Monday incident fell into. A 3-second client timeout with a 4-second queue wait means every request is dequeued after its client has left, and then hashed anyway — 90ms of CPU spent on an answer with no recipient. Your goodput is zero while your utilization is 100%, which is the most demoralizing shape a dashboard can take. The fix is cheap and startlingly effective: stamp each request with its deadline at admission and re-check it at dequeue, before starting the expensive work. Dropping a stale request costs nothing and returns a core to someone who's still waiting.

Under overload, serve the queue LIFO. This one is counterintuitive enough that it's worth stating plainly. In a FIFO queue during sustained overload, the item you dequeue is the oldest — which is to say, the one most likely to have already been abandoned. Every request gets served late and nobody is happy. Under LIFO, the item you dequeue is the freshest, and it is the one most likely to still have a client attached. The old items at the bottom age out and get dropped for free at the deadline check. Some requests are treated very unfairly, but more requests get a useful answer, which is the metric that matters. Switch to LIFO only when queue wait crosses a threshold; FIFO is correct when you're healthy.

An adaptive variant is better still: rather than a fixed depth, shed on measured queue wait — the CoDel idea applied to admission. Sustained wait above target means shed regardless of queue depth, which adapts on its own when the work per item changes.

The global limit is a tax on your largest customer

Now the multi-tenant half, which is where most of the design mistakes live.

Suppose you have one global limit: 5,000 cost units per second across 400 tenants. It has two properties, both bad.

It punishes success first. Your biggest customer is 40% of your traffic. When the limit binds, shedding is proportional to arrival rate, so they absorb 40% of the rejections — the largest absolute number of failed logins, at the account with the largest contract, at their peak hour, which correlates with the peak hour of everyone in their timezone. And they are the tenant closest to any per-tenant ceiling you later add, so every subsequent tightening lands on them again. You have built a mechanism whose reliability cost scales with revenue.

It has no answer for the noisy neighbour. A global limit says nothing about who consumed the budget. One tenant's SCIM sync, one tenant's forced password reset, one tenant's SDK bug — the limit trips and the failures land on everyone, including the 399 tenants whose traffic did not change at all. That is the specific failure the Monday scene describes: the tenants who went down were mostly not the tenant who caused it. In a multi-tenant identity platform this is the worst failure mode you have, because identity is a hard dependency for everything the customer runs. You didn't degrade a feature. You logged their entire company out.

The reflex fix is hard partitioning: give each of the 400 tenants 12.5 units per second and be done. Don't. Identity traffic is spiky and heavily correlated with local business hours, so per-tenant demand is near zero for most of the day and 20× its mean for twenty minutes. Hard partitions size everyone for their peak, which means the system as a whole runs at maybe 15–20% utilization while individual tenants get throttled during the only period they care about. You've bought isolation with most of your capacity, and you still throttle people.

The shape that actually works is reservations plus a shared burst pool:

  • Each tenant gets a reservation — a floor they can always get, sized from their observed baseline (their p95 over a trailing window, say, with some headroom). This is the number you can honestly put in a contract.
  • Reservations are deliberately oversubscribed against total capacity — sum them to roughly 50–60% of what you can serve. That's safe precisely because peaks don't coincide across timezones, and it's the assumption to monitor rather than assume.
  • The remaining capacity is a shared burst pool, allocated by weighted fair queueing across whoever is currently over their reservation.
  • Under contention, tenants are served strictly in order of how far they are over their reservation. A tenant inside its floor is never shed. A tenant at 8× its baseline is shed first, no matter who they are.

That last rule is the one that makes the noisy-neighbour problem go away, and it's worth noticing what it does not require: no judgement about whether the burst is legitimate, no attack attribution, no rules engine. "You are furthest above your own normal" is a complete decision criterion, computable in a few microseconds, and it produces the right answer whether the cause is malicious, accidental, or a genuinely great sales quarter.

The scheduling algorithm here is not novel; it's just borrowed from an adjacent field. Deficit Round Robin was invented to fairly schedule variable-size network packets, and a cost-denominated per-tenant limiter is exactly DRR where the packet size is CPU cost. Each tenant gets a quantum of cost units per round, unused quantum carries as a deficit, and a tenant whose next request costs more than its remaining deficit waits a round. Thirty-year-old algorithm, small implementation, and it solves both halves of this article at once — cost denomination and fairness — because it was designed for exactly the case where the unit of work has wildly variable size.

flowchart LR
    R["Request"] --> C["Cost classifier<br/>route + tenant → units"]
    C --> B{"Tenant<br/>within<br/>reservation?"}
    B -->|"yes"| Q["Bounded queue<br/>depth = rate x target wait"]
    B -->|"no"| P{"Burst pool<br/>available?<br/>(DRR by overage)"}
    P -->|"yes"| Q
    P -->|"no"| S["Shed: 429 + Retry-After<br/>cost: microseconds"]
    Q -->|"full, or deadline expired"| S
    Q --> W["Worker pool<br/>Argon2id, isolated"]
    W -->|"measured cost"| C

The feedback edge from the worker pool back to the classifier is the reconciliation loop — the estimate gets corrected by what the work actually cost.

Where the limiter lives, and why that's the same question

There's a rule that decides limiter placement, and it falls out of the cost table above: the limiter must be cheaper than the thing it protects.

A central Redis round trip to check and decrement a bucket costs perhaps 0.5–1ms including its own tail. Against a password login at 950 units, that's an 1% overhead for exact, globally-consistent accounting — an excellent trade, and you should take it. Against a token validation at 1 unit, the limiter costs sixteen times as much as the operation it is limiting, and you've built a system whose dominant cost is asking permission. Worse, you've put a network dependency on the hottest, most availability-critical path you own, which is the mistake what should identity do when the database is down is largely about.

So the answer is split by cost class, and the split is not a compromise — each side is right for its own traffic:

Cheap, high-volume (validation, introspection) Expensive, low-volume (login, SCIM, admin)
Where Local, in-process token bucket per node Central, or a dedicated admission service
Accuracy Approximate; bounded overshoot Exact enough to hold a contract
Sync Async reconciliation, 100–250ms leases Synchronous check
Overhead ~1 µs ~1 ms, i.e. ~1%
Failure mode Fail open, keep serving Fail closed to a conservative local budget

The local-bucket approach needs one piece of arithmetic stated honestly. Worst-case overshoot during a lease interval is bounded by the number of nodes times the per-node burst, so you tune the lease interval against how much overshoot you can absorb — and you size the real capacity for the overshoot, not for the nominal limit.

The subtler problem is skew. Load balancers do not distribute one tenant's traffic evenly, especially with connection reuse and sticky sessions, so a tenant whose traffic lands on three of twenty nodes gets throttled at 150 units/second while 850 units of their own budget sit unused elsewhere. This is the most common way a per-tenant limiter produces support tickets. The fix is to allocate leases proportional to recently observed demand per node rather than evenly: a node serving 40% of a tenant's requests leases 40% of the budget, refreshed every couple of hundred milliseconds from a central allocator that is off the request path. It's a small control loop, and it turns a limiter that fires wrongly into one that mostly doesn't.

Sometimes the honest answer is "buy more cores"

Rate limiting is rationing, and rationing is what you do with a resource you have decided not to expand. Deciding that is a legitimate engineering choice, but it should be a choice, and there's a real failure mode where a limiter becomes the permanent substitute for a capacity plan — every quarter the limits get tighter, every quarter more legitimate traffic is shed, and the graph everyone looks at (error rate) stays flat because sheds are working exactly as designed.

Some tests for which situation you're in:

  • Who is being shed? If sheds concentrate on tenants far above their own baseline, the limiter is doing its job. If the median tenant is being shed during normal business hours, you are under-provisioned and the limiter is hiding it.
  • Is the shed rate correlated with growth or with events? Event-correlated sheds are the mechanism working. Growth-correlated sheds are a capacity plan expiring in slow motion.
  • What does headroom cost versus what does the shed cost? For a 500/s login peak you need ~100 cores. Doubling that is a rounding error next to a single enterprise renewal, and the art of load testing makes the case that this conversation belongs with whoever owns the forecast, six months early.

And there's a third option that is neither limiting nor provisioning, which is making the work cheaper. The 950-unit row in that table is a design decision, not a law of nature. A WebAuthn assertion costs 20 units — roughly a fiftieth. Moving a tenant to passkeys doesn't shave the login cost, it deletes it, and it deletes the entire class of incident this article opens with, because there is no expensive-CPU asymmetry left to exploit. That is a capacity argument for passwordless, and it's one I rarely see made in the business case, which is usually written entirely in the language of phishing resistance. Both are true. Only one of them shows up on the cloud bill.

What to put on the dashboard

If you take one operational change from this, make it this: your existing alerts will not catch the failure described here, because the system is behaving correctly right up until it isn't.

Error rate is a lagging indicator by design — a well-tuned limiter converts overload into deliberate 429s, which either look like errors (and cry wolf constantly) or get excluded (and go silent during the exact event you care about). The signals that actually lead:

  • Shed rate, per tenant, per cost class. The single most informative number in the system. A tenant whose shed rate crosses zero for the first time is a story, not a metric.
  • Queue wait, p99, separated from service time. Wait and work have opposite fixes and every dashboard showing only total duration makes them indistinguishable.
  • Cost units consumed versus reservation, per tenant. Your leading indicator for a capacity conversation, and a genuinely useful input to a renewal.
  • Burst pool utilization. If it's routinely empty you've oversubscribed reservations too aggressively. If it's routinely full you're leaving capacity on the floor.
  • Estimate drift per cost class. Persistent divergence between estimated and measured cost means the table lies, and the limiter is enforcing a number that stopped being true.
  • Goodput, not throughput. Responses that arrived before their client gave up. This is the only metric that would have shown the Monday incident for what it was: 100% CPU utilization, and zero useful work.

None of this requires knowing whether anyone meant you harm. That's the point, and it's why this belongs to the SRE half of the org rather than the security half. Your password hash is expensive on purpose. Expensive-on-purpose is a finite resource. Finite resources need admission control, denominated in the resource, allocated fairly among the people who paid for it — and enforced before the work is scheduled, because after is just a slower way of saying yes.