Rate Limiting the Login Endpoint Without Punishing Real Users

The ticket said "login broken for entire Manila office." It arrived at 01:14 UTC, which is 09:14 in Manila, and by the time anyone read it there were nine more like it.

Nothing was broken. Three weeks earlier someone had shipped a per-IP limit on POST /login — sixty attempts per minute, a number chosen because it was ten times what a human could type and nobody could think of an objection. It had been quiet since, because the customers logging in during European mornings egressed through a large SASE vendor with hundreds of points of presence and their traffic spread across enough addresses that no bucket ever filled. Manila was one customer, twenty-two thousand employees, four egress IPs, and a security team that had spent two years consolidating internet breakout onto exactly those four addresses so it could be logged.

The arithmetic is not subtle once you write it down. Roughly 60% of that workforce authenticates inside the first twenty minutes of the working day: about 13,200 logins over 1,200 seconds, eleven per second, spread across four addresses. That's 165 attempts per minute per IP against a limit of sixty. From 09:00 the buckets were full and stayed full, and the system rejected roughly two-thirds of every login attempt from that company for the better part of an hour. Retries made it worse, because the mobile client retried immediately on any non-2xx and a 429 came back in fifteen milliseconds instead of the nine hundred a real login took.

The limiter was working exactly as specified. Every single request it rejected was a real employee typing a real password correctly.

What made it a bad limiter was not the threshold. It was the key. Sixty per minute is a reasonable number for a person, and the limiter had no idea what a person was — it had an IP address, and it had quietly assumed an IP address was a person. That's true for some of your users, false for a great many, and false in a way that correlates precisely with your largest enterprise customers.

A rate limiter is an identity claim made before authentication. The key you choose is that claim, and every key available to you is either forgeable, shared, or both — so there is no single correct key, only a set of wrong ones whose errors you can arrange not to overlap. The design work is not picking a number. It is choosing what you are willing to assume about an anonymous request, deciding what it costs when that assumption is wrong, and making sure the cost lands on the attacker rather than on the customer who consolidated their egress because you told them to.

This is the practitioner's half of that problem: what to key on, which algorithm, where the counter lives, what happens when the counter store is down, what you do when a limit trips, and how to tune it against real traffic without generating a false-positive queue nobody works. Two arguments I'm deliberately not making, because they're made properly elsewhere: Rate Limiting an Identity Platform Is a Capacity Problem covers cost-denominated limits, per-tenant reservations and fair queueing — the admission-control layer everything below sits underneath — and Account Lockout Is a DoS Vector covers why counting failures on the account and stopping at N hands anyone who knows a username a free denial-of-service primitive. Both treated as read.


The key is the whole design

Start by being honest about what each candidate key actually identifies, who shares it, and what it costs an attacker to get a fresh one. That last column is the one that decides everything, and it's the one nobody puts in the design doc.

Key What it identifies Legitimately shared by Cost to the attacker of a new one
IPv4 /32 A network egress point 1 to ~50,000 people Cents. Residential proxy pools are sold by the gigabyte.
IPv4 /24 An allocation, loosely Everything above, ×256 Same, plus a little targeting
IPv6 /128 Nothing useful Usually one device Free. See below.
IPv6 /64 One customer's subnet One household or one office VLAN Free-ish; a residential /56 contains 256 of them
ASN An operator Millions Meaningful. There are only ~75,000 routed ASNs.
Account A user (allegedly) Service accounts, shared mailboxes Zero — it's the victim's, not theirs
Tenant A customer Everyone at that customer Zero, if the attacker is targeting that tenant
client_id An application Every user of that app Zero — it's public by construction
Device / session cookie A browser profile One browser Free. Discard and re-request.
TLS fingerprint (JA3/JA4) A client stack Everyone with that browser build Low, and falling — mimicry libraries are commodity
Account × IP A user at a place Rarely more than a few Zero for the account half, cents for the other

Read the right-hand column top to bottom and the shape of the problem appears. Every key an anonymous requester supplies is cheap for them to replace; every key that isn't cheap to replace is shared by people you care about. No row is both attacker-expensive and user-exclusive, and that's not an accident of this table — it's the structure of the problem. Pre-authentication you have attributes of the connection and attributes the requester asserted, and the requester chooses both.

The IPv6 row deserves its own paragraph

Per-address rate limiting on IPv6 is not a weak control. It is not a control at all, and a surprising number of limiters ported straight from IPv4 have this bug sitting in production, invisible because IPv6 traffic is still a minority of login volume in many estates.

A /64 contains 2⁶⁴ ≈ 1.8 × 10¹⁹ addresses, and a /64 is the smallest thing anybody gets. A residential customer typically receives a /56 (256 subnets) and a business a /48 (65,536). So an attacker on an ordinary consumer connection controls something on the order of 10²¹ source addresses that all route to them. A limit of five attempts per address per hour authorises about 10²² attempts per hour from one broadband line. They will not use them all; they need perhaps ten thousand, and each one arrives from an address your counter store has never seen.

That last clause is the second failure and it's worse than the first. Every novel address is a new key. A limiter keyed on /128 under IPv6 attack does not throttle anything — it faithfully allocates one counter per attempt and hands you an unbounded-cardinality write workload against a store you sized for a few million keys. The limiter becomes the amplifier.

Key on the allocation, not the address: /64 at minimum, /56 or /48 for anything you'd treat as one actor. That occasionally aggregates two households behind one operator's unusual addressing plan, which is a much better trade than the alternative. If you take one mechanical change from this article, make it this one — and go and read the code that derives the key, not the config file.

Corporate NAT and CGNAT: the shared-address problem doesn't have a clean answer

The Manila incident is one instance of a general condition. A single IPv4 address legitimately carrying thousands of distinct humans is the normal state of large-enterprise egress, of carrier-grade NAT in networks that ran out of IPv4 space, of university and hospital networks, and of every mobile operator that anchors a region's traffic at a handful of gateways. The same infrastructure that makes impossible travel full of false positives makes per-IP rate limiting full of them, for the same reason: the address describes the network, not the person.

The obvious remedy — allowlist your customers' known egress ranges — fails operationally rather than technically. It works for a quarter. Then egress ranges change when a customer migrates to a SASE vendor and nobody tells you; the list is maintained by whoever handled the last escalation; "known egress" for a global SASE deployment is a few hundred prefixes the vendor reserves the right to change; and an allowlist is a permanent exception with no expiry, so after two years a meaningful fraction of your login traffic is exempt from the control and nobody can say which fraction or why. I have never seen this scale past about forty customers, and I've seen it fail at twelve.

The practical middle is to type the address rather than exempt it. IP intelligence feeds classify addresses by connection type — residential, mobile, business, hosting, VPN/anonymiser — and that classification is a far better predictor of expected concurrency than the address is of location. Set the limit per type:

Hosting, cloud and anonymiser gets a tight limit — almost no legitimate interactive login originates from a datacentre, so this is where a low per-IP threshold is both safe and useful. Mobile gets a loose per-address limit, because a carrier gateway is thousands of handsets, with the compensation on the account axis. Business likewise: loose per-address, tight per-account, and lean on tenant-scoped budgets from the capacity piece rather than on the IP. Residential sits in the middle — the only class where "one address ≈ a few people" holds, and also where residential-proxy traffic hides, which is rather the point of residential proxies.

This is less a solution than a redistribution of the error into a place where it costs less. Be clear-eyed: the classification is a purchased dataset with the same staleness and coverage problems as geolocation, and an attacker who buys residential egress lands in your most permissive class on purpose. It buys you a lot on the hosting row and very little on the residential one.


The composite key is the workhorse

Key the primary failure counter on the pair (account, source) — account identifier and IP prefix together — rather than on either alone. It's worth understanding why rather than just adopting it.

Against the Manila problem: twenty-two thousand employees behind four addresses now occupy twenty-two thousand buckets rather than four, so the per-IP aggregate is irrelevant and each user gets their own budget. Against the lockout-DoS problem: an attacker hammering [email protected] from their own address fills (alice, attacker-prefix) and cannot touch (alice, alice's-prefix), so Alice logs in normally while the attacker exhausts a counter that constrains only the attacker. That second property is the cleanest answer I know to the objection that any account-scoped counter is a weapon: a counter scoped to (victim, attacker) can only be exhausted by the attacker, and its exhaustion only harms the attacker.

It generalises, too — everywhere you're tempted to key on an unauthenticated assertion, the composite rescues you. A real example from a security review of the platform I work on: the token endpoint's throttle bucket was keyed on the client_id in the request body, evaluated before the client authenticated. client_id is public by construction, so anyone could send garbage token requests for a client they didn't own and blockade that client's real traffic. The rule falls straight out:

Never key a limiter solely on an unauthenticated assertion of the identity you are protecting. You may key on unverified input only when exhausting that bucket harms nobody but the presenter. Otherwise, composite it with something the presenter has to actually hold.

(client_id, source prefix) restores that property: to blockade a client you must now also occupy the address, and occupying an address is the one thing on the table that costs the attacker anything at all.

What the composite still misses

It misses the attack that matters most, and pretending otherwise is how teams end up with an expensive limiter and a flat takeover rate.

A distributed spray sends one attempt per (account, source) pair: ten thousand accounts, one attempt each, from a pool rotating every request. Every composite counter reads 1. Every per-account counter reads 1. Every per-IP counter reads a handful. The limiter is functioning perfectly and observing nothing, because the attack's whole design is to stay under a per-key threshold — and it can, because the key space is cheap and the number of keys is the attacker's free parameter. Same structural blindness that makes lockout thresholds useless against spraying, one layer down.

So the composite key needs a second axis that counts across keys rather than within them. The questions that detect a spray are cardinality questions:

  • How many distinct accounts has this prefix, or this ASN, failed against in the last hour?
  • How many distinct sources have failed against this account in the last hour? (The lockout-DoS detector, and a good one.)
  • What fraction of this tenant's login attempts in this window are failures, against its own trailing baseline?
  • Is the same password appearing across many accounts? Countable over a keyed hash without retaining candidates.

Those are aggregations, not counters, and they don't belong in the synchronous request path — they belong in a streaming job that publishes verdicts the request path reads. Which gives the layered shape:

flowchart TB
    R["Login attempt"] --> T{"Source type<br/>residential / mobile<br/>hosting / corporate"}
    T --> C["Composite counter<br/>(account, prefix)<br/>synchronous, exact"]
    C --> A["Account counter<br/>generous ceiling<br/>failures only"]
    A --> V{"Verdict cache<br/>from stream job"}
    V -->|"clean"| P["Verify credential"]
    V -->|"flagged source<br/>or flagged account"| E["Escalate:<br/>delay, challenge,<br/>step-up"]
    C -->|"exceeded"| E
    S["Stream job<br/>distinct-account counts,<br/>tenant failure ratio"] -.->|"publishes, async"| V
    P -.->|"outcome events"| S

The dotted edges are the point: the expensive, cross-key reasoning runs off the request path and its conclusions arrive as a cheap lookup. Anything that needs to consult more than two counters synchronously is going to fail the latency argument in the identity latency budget before it fails the security one.


Algorithms, judged on how they behave under attack

The textbook comparison of rate-limiting algorithms is written for API quotas, where the failure mode is a billing dispute. For a login endpoint the properties that matter are different: what happens at a window boundary when someone is deliberately aiming at it, how much memory the structure costs multiplied by your key cardinality, and whether it can be evaluated atomically in one round trip.

Fixed window. A counter per key per calendar minute. Trivially cheap — one INCR with a TTL — and it has a boundary an attacker finds in a few probes. With a limit of 10 per minute, ten requests at 11:59:59.900 and ten more at 12:00:00.100 is twenty requests in 200 milliseconds, which is not "2× the limit"; it is an instantaneous rate of 6,000 per minute. If the thing behind the counter is a password hash at 90ms of CPU, that's 1.8 core-seconds arriving in a fifth of a second, per key, from every key the attacker holds. Fine for coarse, generous ceilings. Not fine as your primary control.

Sliding window log. A timestamp per attempt in a sorted set; drop entries older than the window, count what's left. Exact, no boundary artefact, and the memory is the problem: in Redis a sorted-set entry is realistically 60–80 bytes once you count skiplist and dict overhead. At a limit of 10 and five million active keys that's up to 50 million entries, call it 3.5 GB, against roughly 600 MB for plain counters over the same keys. Six times the memory for exactness on a threshold you picked by intuition.

Sliding window counter. Keep the current and previous fixed-window counts and interpolate: prev × (1 − elapsed/window) + curr. Two integers per key, no boundary cliff, and the error is bounded — it over-counts when the previous window's traffic was back-loaded, by at most that window's count. For a control whose threshold is a judgement call, that's noise. The right default for most login counters.

Token bucket. Capacity B, refill r per second. The property to internalise, because it reframes the whole tuning conversation: the burst capacity is for your users and the refill rate is for your attacker. A real user who typos twice, checks their password manager and succeeds on the fourth try consumes burst and never notices. An attacker doesn't care about your burst — they care about the sustained rate, which is exactly r.

So tune the two separately and do the arithmetic on r. Capacity 5, refill 1 per 30 seconds feels tight and grants 2,880 guesses per key per day. Refill 1 per 5 minutes grants 288. Against a 30-bit password that's irrelevant either way; against the top 1,000 passwords, 288/day exhausts the list against one account in three and a half days — well within patience for a targeted attack on one executive. Against a spray the bucket never binds at all, because the attacker spends one token per key. Write those three numbers down before you argue about the threshold: they tell you which attacks the control can and cannot touch, which is the question worth asking of any control.

GCRA — the generic cell rate algorithm, from ATM traffic shaping — is token bucket rearranged so the state is a single timestamp (the theoretical arrival time) rather than a level plus a last-updated stamp. Eight bytes per key, no background refill task, equivalent semantics, one compare-and-set to evaluate. If you're writing the limiter rather than configuring one, write this; it is strictly better than the naive token bucket and almost nobody reaches for it.

Leaky bucket as a queue smooths output by making excess wait rather than fail. Mostly the wrong shape for an interactive request with a client timeout attached — but see the tarpitting discussion below, where making the attacker wait is exactly the point.

Probabilistic structures, with the honest verdict. A count-min sketch at width 2,048 and depth 4 costs 32 KB and overshoots any key's count by at most ε = e/w ≈ 0.13% of the total stream volume at 98% confidence — over a million events, ±1,300. So a sketch is a fine instrument for finding heavy hitters ("which prefixes are in the top 0.1% of failure volume") and a useless one for enforcing a threshold of five. Sketches for the long tail of keys you can't afford to materialise, feeding the detection axis; exact counters wherever the threshold is small. Getting this backwards — a sketch enforcing per-account limits — produces a control that occasionally blocks the wrong person for reasons that are unreproducible by construction.

For the cardinality questions above, HyperLogLog is the right tool: Redis's costs 12 KB dense for a 0.81% standard error, and stays in a sparse encoding of a few hundred bytes until a key holds thousands of elements. That asymmetry is what makes it affordable — a million source prefixes cost the sparse size, and only the handful that genuinely touched thousands of distinct accounts get promoted. The memory concentrates on exactly the sources you want to be watching.


Cardinality is the cost, not throughput

This is the part that surprises people who have sized a rate limiter before, and it's the reason limiters fall over in ways that look nothing like "too many requests."

Redis will do a couple of hundred thousand INCRs per second per core without complaint. Request rate is not your constraint and never will be. Your constraint is the number of distinct keys alive at once, because each one costs memory whether or not it's being touched:

key string   "rl:pw:t_4f2a:8e11c0d3a97b:2001:db8:1a2b::/64"   ~48 bytes
robj + dictEntry + expires entry                              ~80 bytes
value (two ints, embedded)                                    ~0
                                                        -------------
                                                       ~128 bytes/key

Five million live composite keys is roughly 640 MB. Fifty million is 6.4 GB and you are now sizing a cluster around your rate limiter, which is a strange place to have arrived.

Now the failure mode. If any component of your key is attacker-controlled and unbounded, your limiter is a memory-exhaustion primitive. Key on the identifier as typed — not resolved, not hashed to fixed width — and an attacker sending 100,000 requests per second with a random username in each mints 100,000 new keys per second at ~128 bytes: 12.8 MB/s, 46 GB an hour. Long before that fills anything it evicts every legitimate key under allkeys-lru (so the real limits stop working) or refuses writes under noeviction (so, depending on your fail policy, either every login is throttled or none are). Either way the attacker disabled the control by feeding it, and no alert you have is named "rate limiter key cardinality."

Three cheap mitigations: hash every attacker-influenced component to a fixed width (8–16 bytes of a keyed hash, which also keeps raw identifiers out of a store that isn't your user store); only materialise a per-account key for identifiers that resolve, charging a shared per-source bucket otherwise — carefully, so it doesn't become an enumeration oracle, per the next section; and cap live keys per tenant and per prefix, alerting on the derivative, because a tenant whose live-key count goes up 40× in five minutes is a story you'd rather not learn about from a memory alert at 3am.

And a fourth, really a correctness bug that surfaces under exactly this load: the increment must be atomic. A read-modify-write across two round trips admits far more than the limit under concurrency — twenty parallel requests all read 4, all decide they're under the limit of 5, all proceed. It is one of the most common defects in hand-rolled limiters, invisible at low traffic, and it appears precisely when an attacker parallelises. One INCR, one Lua script, or one GCRA compare-and-set. Never two round trips.


Where the counter lives

Three plausible homes, and the choice is not "which is best" but "which key belongs where."

Edge / CDN API gateway Identity service
Sees Raw connection, TLS fingerprint Route, tenant, client_id Account, credential type, tenant policy
Latency added ~0 (already in path) ~0.2 ms 0.5–1.5 ms if central store
Accuracy Per-POP, so ×POPs Per-node, so ×N Exact, if shared store
Right for Volumetric floods, hosting-ASN blocks, /64 caps Per-client_id, per-tenant admission (account, source), failure counters, policy
Can't do Know who the account is Know if the credential was right Absorb a flood — it's already in your capacity

The distributed-counter trade is the one people get wrong. A local per-node counter costs a microsecond and is wrong by a factor of N: a limit of 10 per minute across 24 uncoordinated nodes is really 240 per minute, drifting with load-balancer behaviour. A shared store is exact and costs a round trip per login — for a password login at ~90ms of hashing that's ~1% overhead and obviously worth it, and for a token validation at 60 microseconds it's a 16× overhead and obviously not. The capacity piece works through the lease-based middle ground and the load-balancer skew that makes naive even-splitting produce support tickets.

What it doesn't cover, and what matters more here, is the tail. A synchronous limiter check is a serial hop on the login path, so its p99 lands inside your login p99, and percentiles across serial hops don't compose the way intuition says. A limiter store with a p99 of 8ms and a p99.9 of 400ms — entirely normal for a Redis instance sharing a node with something noisy — puts a 400ms stall into one login in a thousand. So: give the check a timeout in the low single-digit milliseconds, and decide now what happens when it fires, because it will fire during exactly the incident where you care.

Fail open or fail closed, per key type

The instinct is to answer this once for the whole limiter. Wrong granularity — it depends on what the counter protects, and the rule is short enough to memorise:

Fail open where the counter protects the user. Fail closed where the counter protects the system.

Failing open on a security counter costs a bounded window of extra guesses against a password the attacker probably wasn't going to find anyway. Failing closed on that same counter rejects every login in the estate because a cache node restarted — the outage you built the limiter to avoid, self-inflicted. Failing open on a capacity counter, meanwhile, lets the hash pool take unbounded arrivals, which is the one failure with no partial version.

Counter Protects Store unavailable →
(account, source) failure counter The user's account Open, degraded to a local per-node approximation with a conservative threshold
Per-account ceiling The user's account Open, and raise the assurance requirement instead — step-up rather than block
Per-tenant cost budget The platform Closed to a static local budget of total / N
Global hash-pool admission The platform Closed. This one is never allowed to fail open.
Hosting-ASN / anonymiser caps The platform, mostly Closed. Almost no legitimate interactive login is here.

"Degraded to a local approximation" is doing real work in that table and is worth building deliberately: an in-process per-node counter with a threshold of limit / N rounded up, used only when the shared store is unreachable. You lose cross-node precision and keep the control — a much better failure than either extreme, and the same graceful-degradation reasoning as what an identity platform should do when its database is down, where the answer is a decision table rather than a global switch.

One more placement trap, because it's the most common way a limiter is bypassed in practice. If you derive the client IP from X-Forwarded-For without pinning how many proxy hops you trust, the attacker sets the header and picks their own bucket — and, worse, picks someone else's, turning your rate limiter into a targeted denial-of-service tool aimed at whichever legitimate address they name. Count hops from the right-hand end of the chain, with the trusted-hop count configured explicitly per deployment. This is a real finding from a real platform's security review, and it belongs to the class covered in what happens when the layer below you lies.


Successes and failures are not the same event

A login endpoint sees two outcomes with wildly different meanings, and a limiter that counts them the same is throwing away most of the signal it has.

Count failures aggressively; count successes generously. A hundred successful logins in an hour from one account is a service account, a shared mailbox, or a bad integration — a capacity question, not a security one. A hundred failures is one of exactly two things, both worth acting on. Run two counters with thresholds an order of magnitude apart, and put the tight one on failures.

Never let a successful attempt reset a counter the attacker can influence. If a success clears the source-scoped failure counter, an attacker with one valid account of their own — a free trial, say — interleaves a success every N failures and runs indefinitely at whatever rate they like. A success may clear the counter for its own (account, source) pair and nothing wider. Source- and tenant-scoped counters decay by time only.

The counter's existence must not be observable. Everything in the credential-stuffing piece about classification oracles applies here, and a rate limiter is unusually good at leaking. Three specific leaks:

  • Existence via cost. If you only materialise a per-account counter for accounts that resolve — the memory mitigation above — then throttling behaviour after N attempts tells the attacker the account is real. Charge a shared per-source bucket for unresolved identifiers at the same observable threshold with the same response, so only the storage differs.
  • Existence via timing. A short-circuited rejection returns in 3ms; a real failed password returns in 95ms. That gap is trivially measurable and reveals which limit fired and, transitively, whether the account exists. Rate-limited responses need the same constant-time floor as unknown-user responses — the familiar dummy-hash mitigation, extended to the limiter, where it's frequently missing because the limiter was added later by someone else.
  • Scope via message. "Too many attempts for this account" versus "too many requests from your network" tells an attacker which axis they saturated and therefore which one to rotate. One message, one status, for every throttled outcome.

The platform I work on makes this a spec-level requirement rather than an implementation detail: at the identifier stage, "no such user" and "rate limited" render the same string — "We couldn't sign you in. Please try again." — behind a constant-time floor; and the password-attempt limit is keyed per (tenant, user) and shared across all concurrent login sessions for that user, so opening five browser tabs doesn't multiply the budget by five. That second detail is the kind of bug you only find by asking "what is the actual key, and what can the attacker do to make more of them?"


What to do when it trips

The 429 is the least interesting response available and, for a signal this noisy, usually the wrong one. The useful framing is a ladder ordered by what it costs a real user who tripped it by accident against what it costs an operator running ten million attempts.

Response Cost to a real user Cost to an operator at 10M attempts Use when
Monitor only (count, don't act) Zero Zero Always, first, for weeks
Progressive delay (0 → 1s → 3s → 10s) Barely noticed at the low end Caps sustained rate per key; trivially parallelised around Default for the composite counter
Silent tarpit (hold connection 2–5s, then respond) Slow login, no error Caps concurrency, not just rate Detected automation, retry storms
Proof of work (~200ms client-side) 200ms, invisible ~556 CPU-hours ≈ $15–30 spot Capacity protection, not credential protection
CAPTCHA 10–30s, ~5–15% abandonment ~$10–30k of solver spend Rarely. See below.
Forced step-up / OOB verification A real interruption Blocks session, not validation High-risk signal, high-value account
Hard 429 block Login is impossible Rotate the key, continue Last resort, narrow keys only

Four things worth drawing out.

Degrading beats denying, and the reason is arithmetic rather than kindness. Whatever threshold you pick, some legitimate users cross it — that's the false-positive budget below. If the response is a hard block, each of them is a support contact and a push toward the recovery path, the weakest and most attacked surface you own. If the response is three seconds of delay, each of them is a person who noticed nothing. Against an attacker the two differ very little; against your users they differ enormously. That's why the threshold and the response are one decision, not two: tight-and-soft is coherent, loose-and-hard is coherent, and tight-and-hard is how you get the Manila ticket.

Tarpitting is more powerful than it looks, and it's the counterintuitive answer to the retry storm. A fast rejection is a gift to a badly-behaved client: it frees the socket so the client can immediately try again. Holding the connection for two seconds before rejecting caps that client at 0.5 requests per second regardless of its retry logic, because a client blocked on a socket is a client not sending. The caveat: hold it cheaply — an async handle, not a thread, with a hard cap on concurrent held connections, or you've implemented a self-inflicted Slowloris.

Proof of work is a capacity control mis-sold as a security control. 200 milliseconds of memory-hard work is invisible to a human and costs a funded operator running ten million attempts about twenty dollars of spot compute; it will not deter anyone. What it does do is convert your Argon2 cost into their CPU cost before you spend yours — a real property, and an availability one, belonging to the conversation the capacity piece is having rather than this one. CAPTCHA is the same argument with worse economics: solver services publish price lists around a couple of dollars per thousand solves, so at a 0.5% stuffing hit rate that's roughly $0.40 of solver cost per validated credential pair — comfortably inside the margin on a corporate SSO account, paid without a human involved, while the cost lands on your users as abandonment and lands hardest on those with assistive technology, poor connections and old devices. Challenges earn their place only when they force per-target engineering instead of a call to an existing service, and that property decays on a schedule set by your vendor's market share.

The best response is often on a different axis. A throttled attempt that would have succeeded is the most interesting event your login endpoint produces, and the right answer is usually not a block but a step-up: authenticate the person, then require an additional factor bound to the action. Step-up done right and adaptive MFA are the mechanisms; the point here is that "throttle" and "challenge" are different verbs and the limiter should be able to emit either.


429, Retry-After, and the client that turns a limit into an outage

Return 429 Too Many Requests, always with Retry-After. Not 403, which clients cache and treat as terminal; not 503, which tells load balancers and service meshes to eject your node; never 200 with an error in the body, which defeats every retry policy in every client library. The RateLimit header family is worth emitting on machine-facing endpoints like /token and worth not emitting on interactive login, where the only party reading it carefully is the one probing your thresholds.

Now the dynamic that turned Manila from a bad morning into a bad day, and the most under-anticipated thing in this whole area.

A mobile client with no backoff retries immediately on any non-2xx. A normal login round trip is about 900ms, so a client in a retry loop against a healthy server generates about 1.1 requests per second. A 429 from the edge comes back in 15 milliseconds. The same client now generates 66 requests per second — a 60× amplification of exactly the traffic you were shedding. Fifty thousand affected handsets is 3.3 million requests per second at your edge, and the graph doesn't look like a rate limit working; it looks like a DDoS from your own customers.

Mitigations, in order of how much you control them. Serve 429s from the cheapest layer you have, and check that layer can absorb the amplified rate — if the 429 is generated by the identity service, the amplification lands on the identity service. Escalate Retry-After per key (1s, then 5, 30, 120): a client that honours it backs off, and a client that ignores it becomes distinguishable, which is a useful classifier you got for free. Add jitter, because a uniform Retry-After: 30 synchronises every throttled client onto the same second and gives you a thundering herd on a schedule you published. Tarpit the ones that ignore it — the only lever that works on a client you can't change, and you will have several, because some fraction of your install base is a version you shipped two years ago. And fix your own SDK (backoff with full jitter, a retry budget rather than a retry count, a circuit breaker), then accept that the old versions are the ones that will be in the incident.


Tuning against real traffic

Everything above is structure. This is the part that decides whether the thing you built is a control or a support-ticket generator, and it is almost always skipped, because the number gets chosen in a design review by whoever speaks first.

Measure the distribution before you pick a threshold

Instrument attempts per account per day for a fortnight and plot the shape. Every workforce IdP I've seen numbers looks roughly like this, and the shape matters more than the exact values:

Percentile Login attempts per account per day
p50 1
p75 2
p90 3
p99 8
p99.9 23
p99.99 60+
max thousands

Two things to notice. First, the distribution is extraordinarily skewed — the median user makes one attempt and the tail runs four orders of magnitude past it — so any threshold you pick is either far above the median (useless against a targeted attack) or somewhere in a tail populated entirely by legitimate weirdness. Second, and more useful: the extreme tail is not human. The accounts making thousands of attempts a day are a stale credential in a mail client, a scheduled task with last quarter's password, a mapped drive in a laptop in a drawer — mostly non-human identity using human credentials. Identify and separate that population before you tune anything, or you'll set the threshold to accommodate machines and lose the ability to constrain people.

Budget the false positives explicitly, in humans per day

Here is the calculation nobody does, and it's the one that makes the decision.

Take 500,000 daily active accounts. A threshold set at the observed p99.9 of attempts-per-account-day means, by construction, about 0.1% of account-days cross it: 500 real users per day. Now price that against the response ladder:

  • Hard 30-minute block → 500 blocked users/day, of whom maybe 15–25% contact support → 75–125 tickets a day, forever, plus the ones who just abandon.
  • Progressive delay topping out at 10 seconds → 500 users/day who experience a slow login and never mention it. Zero tickets.
  • Step-up challenge → 500 challenges/day, with the completion-rate and support-contact profile of whatever your second factor is, concentrated on your least typical users.

Same threshold, same detection, three completely different operational realities. And notice what the arithmetic proves: at any threshold tight enough to be useful, a hard-blocking response cannot be affordable, because the false-positive population is structurally larger than the true-positive one. That's the base-rate argument from the impossible-travel piece applied to a limiter — a low-precision signal is fine as long as the action it triggers is proportionate to its precision.

Run in monitor mode, and mean it

Ship the counters, evaluate the thresholds, emit the verdict as a log field and a metric, and enforce nothing. Then look at what would have been blocked, for at least two full weekly cycles plus a month boundary — month-end and quarter-end produce login patterns a fortnight won't show you, and your customers' onboarding days, all-hands meetings and forced password resets are the events that produce the Manila shape.

The metric that matters is not "how many would we have blocked." It is "who would we have blocked, and were they real?" — sampled, and driven to root cause the way you'd drive an alert queue. A hundred sampled would-be-blocks, each resolved to an actual explanation: this tenant's SASE egress, this customer's shared kiosk account, this integration with an expired credential, this genuine spray. Three or four causes will dominate, at least one will be your own infrastructure, and one you can eliminate outright. That exercise is worth more than any amount of threshold arithmetic — and then make it permanent, with an owner, rather than a launch activity. A limiter is a detector, and a detector with no ongoing precision measurement decays silently as traffic changes underneath it.

The endpoints everybody forgets

Almost every team that tunes /login carefully leaves the neighbours open, and the neighbours are frequently the cheaper attack.

  • MFA verification, the worst one, and worth the arithmetic. A six-digit OTP has a million values; five guesses per code is 5-in-10⁶ and feels safe. But if resends are unlimited, each new code buys five fresh guesses. A hundred resends against one account is 500 guesses — 0.05%. Now spray it: 10,000 accounts × 500 guesses is 5,000,000 attempts against a 10⁶ space, an expected five successful OTP guesses, on a system whose per-code limit works exactly as documented. Cap attempts and resends per account per rolling window, not per code, and invalidate outstanding codes when a new one is issued. Same reasoning for push approvals, where the resource exhausted is the user's patience — MFA fatigue is that attack.
  • Password reset. Unlimited requests is a free email-bombing service aimed at your users, a paid-SMS drain, and a reset-token guessing surface. Limit per account, per source, per tenant.
  • Registration. Spam-account creation, a mail-reputation problem, and — if signup says whether an address is taken — the cleanest enumeration oracle you own.
  • /token. Most likely to be hit by machines, most likely to be keyed on an unauthenticated client_id, most likely to be excluded because "it's server-to-server." Refresh grants especially: a client library bug becomes a sustained flood with valid credentials.
  • Magic links and device-approval URLs. Reachable by GET, therefore fetched by mail-security scanners, therefore spending budget on behalf of a user who hasn't clicked yet.

The tuning procedure

If you're starting from nothing, or auditing something you inherited, this is the order.

1. Audit the keys you actually have — the code that derives them, not the config file. Is the IPv6 key a /128? Is the client IP derived from a header with an explicit trusted-hop count? Is any component attacker-controlled and unbounded in length? Is any key an unauthenticated assertion of the identity being protected? Each is a defect with a known exploitation path and a small fix.

2. Make (account, source-prefix) the primary failure counter. IPv4 /24 or /32 by connection type, IPv6 /64 or shorter. Sliding-window counter or GCRA, one atomic operation, TTL-bounded. This single key defeats both the NAT problem and the lockout-DoS problem, and it's the cheapest structural improvement available.

3. Add the cross-key axis asynchronously. Distinct accounts per source, distinct sources per account, tenant failure ratio against its own baseline. Streaming job, HyperLogLog, verdicts published to a cache the request path reads in microseconds. Without it you are blind to spraying, which is the attack you actually have.

4. Measure your traffic and publish the distribution — attempts per account per day, per tenant, at p50/p90/p99/p99.9 — classifying the non-human tail separately before you pick anything.

5. Pick the threshold and the response together, and price the false positives in humans per day. If the arithmetic says more than a handful of support contacts a day, the response is too hard for the precision you have. Move down the ladder, not up the threshold.

6. Run in monitor mode for two weeks minimum, spanning a month boundary, and drive a hundred sampled would-be-blocks to real root causes.

7. Decide fail-open/fail-closed per counter, and test it by killing the counter store in a game day and watching login success rate. User-protecting counters fail open to a degraded local approximation; platform-protecting counters fail closed to a conservative static budget.

8. Fix the client before you tighten the server — backoff with full jitter, honour Retry-After, retry budgets not counts — then assume for eighteen months that most of your install base ignores all of it, and make sure your cheapest layer absorbs a 60× amplification.

9. Repeat all of the above for MFA verify, password reset, registration and /token, in that order. The OTP resend arithmetic is the one most likely to be a live vulnerability in your system right now.

10. Keep "who did we block and were they real" on the dashboard permanently, with a name next to it, alongside per-tenant throttle rate — because the tenant that consolidated its egress on your advice will hit the limit first, at 09:00 their time, on a graph that is aggregated across regions and looks completely fine.


None of this makes the underlying problem go away, and it's worth being straight about that. A rate limiter on a login endpoint is a guess about identity, made before you know anything, using attributes the requester chose. It buys you time and it protects your capacity. It does not stop credential stuffing, because stuffing is not rate-limited by attempts; it does not stop spraying, because spraying is designed to stay under every threshold you can afford to set; and it will never be better than the quality of the keys available to you before authentication.

What it can be is fair — in the specific sense that the cost of being wrong lands on the party that chose to be anonymous rather than on the party who was where they always are, on the device they always use, typing the password they always type, on the Monday morning when their whole company came online at once.

That's the standard to hold the design to. Not "did it block the attack," which you mostly can't know, but: when this fires wrongly — and it will, hundreds of times a day — who pays, and how much? Answer that first, and the threshold picks itself.