The Economics of Credential Stuffing

The takeover was six weeks old by the time anybody looked at it properly. A customer's account had been drained through a payout change, the fraud team had reversed what they could, and the postmortem was the usual archaeology: pull the sessions, pull the device fingerprints, work backwards until the story starts.

The story started earlier than anyone expected. Five weeks before the takeover, a single login attempt had come in against that account from an ASN nobody recognised. The password was correct. The push went out, the user ignored it, the attempt timed out, no session was established. As far as the system was concerned, nothing had happened, and no alert fired. The event didn't appear on the authentication dashboard, which counted successful logins, and it didn't appear on the abuse dashboard, which counted blocked attempts. It fell between the two because it was neither: a login refused by a control working exactly as designed, after the password had already been confirmed valid.

Somebody in the room said "well, MFA did its job," and moved on to the payout controls. That sentence has stayed with me, because it is half true in a way that costs money. MFA did stop that login. It also, in stopping it, published a fact: this pair works here. Five weeks later somebody used that fact — a phishing page prepared for one person, a help-desk call, a SIM port, the route hardly matters — and the second attempt worked.

The thing we had built and instrumented as a wall was, from the other side, a measuring instrument, and we had left it switched on for anyone who wanted to use it.

Credential stuffing is not an attack on your authentication; it is a manufacturing process whose product is validated credential pairs, and your login endpoint is the machine that manufactures them. Once you model it that way — as a business with a cost line, a depreciating raw material, and a conversion step — you can see that the term almost every defence targets, the monetary cost of a single attempt, is already so close to zero that moving it accomplishes nothing. The terms that actually bind are the freshness of the list, the cost of classifying your responses, and the difficulty of converting a validated credential into money. Nearly every login control that teams argue about in review meetings moves the smallest term in the equation.

This is a defender's article, deliberately about arithmetic rather than tooling; there's nothing here you couldn't reconstruct from a quiet afternoon with your own logs.

A scope note, because this blog has circled the subject from several directions. The supply-side argument — that credential independence, not credential strength, is what breaks stuffing, and why password managers were the decade's least-credited security win — is made in full in Why Password Managers Were the Decade's Biggest Security Win, and I'm going to assume it rather than repeat it. Lockout as a control belongs to Account Lockout Is a Denial-of-Service Vector; rate limiting as a capacity discipline belongs to Rate Limiting an Identity Platform Is a Capacity Problem; the mechanics of checking a corpus without transmitting secrets belong to How to Check Passwords Against Breach Databases Without Sending Passwords. This piece is the attacker's profit and loss statement, and what a defender can do with it.


The equation, written out

Start with the unit of work. It is not "an attack." It is one credential pair, tried once, against one target. Everything an operator does is a decision about how to spend attempts.

The expected value of one attempt looks roughly like this:

EV = ( V × pvalid × psurvive × pconvert ) − cattempt − cburn

Five terms, and it is worth being pedantic about what each one is, because the pedantry is where the defensive insight lives.

V — the realisable value of a taken-over account here. Not the balance; what an attacker can actually extract, net of your controls and their labour. A loyalty account with transferable points, a retail account with a saved card, an email account (a skeleton key to everything else), a corporate SSO account with a VPN behind it. This spans orders of magnitude across targets, which is why target selection is an operator's first decision.

pvalid — the probability that this pair is a live credential here. The hit rate, and it is small. Publicly circulated figures for stuffing runs land in the low fractions of a percent — commonly characterised somewhere between a tenth of a percent and a couple of percent depending on freshness and targeting, which you should treat as an order of magnitude rather than a measurement. Google's Password Checkup telemetry around 2019 reported roughly 1.5% of observed logins using a credential present in a breach corpus: a different quantity, same direction. The precise number doesn't matter. What matters is that a fraction-of-a-percent hit rate is extremely profitable when the denominator costs nearly nothing.

psurvive — given a valid pair, the probability the attempt yields a usable session. MFA, device reputation, step-up, bot mitigation, session restriction. Note that this term is downstream of validation: a defence sitting here reduces sessions, not validated pairs. Hold that thought; it's the crux of the article.

pconvert — given a session, the probability of turning it into money. Payout holds, new-payee cooling-off, transaction risk, shipping-address friction, KYC at cash-out. Usually owned by a different team, and — I'll argue below — often the most tractable term here.

cattempt — the marginal cost of one attempt. An IP that isn't already blocked, sometimes a solver call, a share of the engineering that keeps the client plausible. The term the industry spends most of its defensive energy on, and essentially zero.

cburn — the amortised cost of degrading the asset by using it. Nobody writes this one down. Every attempt against a well-instrumented target risks triggering a response — forced resets, corpus-wide invalidation, indicator sharing — that reduces the value of the entire list, everywhere. Attackers manage it the way a fund manages market impact.

Two structural facts fall out immediately.

The revenue side is a product of four terms, three of them small. Multiplicative structures are brutal: to halve the attacker's revenue you can halve any term, but to end the attack you need one at zero, and only one of them can reach zero.

The cost side is dominated by fixed costs, not marginal ones. Acquiring the list, understanding a specific target's login flow, building a plausible client, standing up the classification pipeline — engineering projects with one-time costs amortised over millions of attempts, against which the marginal cost per attempt is a rounding error. That is the cost structure of a commodity manufacturer, and it produces the behaviour you'd expect: enormous volume, thin per-unit margins, extreme sensitivity to anything that raises fixed cost per target, near-total indifference to anything that raises marginal cost per attempt.

Once you see that, most of the login-security debate reorganises itself.


Why per-attempt friction is a dead end

Here is the uncomfortable part for anyone who has spent a quarter tuning a rate limiter.

The monetary cost of one attempt has been effectively zero for years, and it is not the binding constraint on anything. Residential proxy capacity is a commodity market sold by the gigabyte. Solver services for interactive challenges are a mature industry with published price lists and API clients, cheap enough to be a line item rather than a decision. Compute is free at this scale — a stuffing client is an HTTP request, not a hash computation. I'm deliberately not quoting figures, because they drift and the point doesn't depend on them: the correct characterisation is not "cheap" but "not the binding constraint," and those are different claims. A cost that isn't binding can be doubled without changing behaviour.

Work through what marginal friction actually does. Suppose you raise the cost of an attempt by an order of magnitude. The operator's response is not to stop; it is to spend less on attempts that were going to fail anyway. They tighten targeting — pairs whose email domains suggest a demographic that plausibly holds an account here, or the list intersected against a customer roster leaked from your own marketing vendor. That raises pvalid by a factor that comfortably absorbs the cost increase, and now you face fewer attempts at a higher hit rate. Your attempt-volume graph goes down. Your takeover count doesn't.

That's a general property of controls that tax a cheap input: when a resource is abundant, taxing it changes allocation, not outcome. The attacker was wasteful because waste was free. Making waste expensive makes them efficient, and an efficient adversary is not obviously an improvement.

Two honourable exceptions. The first is capacity: stuffing traffic will happily consume your password-hashing budget and take the site down as a side effect. Defending that is necessary work — it just isn't an anti-stuffing control, it's an availability control, and conflating them produces the classic outcome where you're proud of a system that stays up while takeovers continue at the same rate. That argument is made properly in the capacity piece.

The second is that friction raises fixed costs when it forces per-target engineering. A challenge an operator must build something specific to defeat — rather than call an existing service for — is a genuine increase, because it hits the amortised term. That's why bot-mitigation vendors have real value, and also why that value decays: once a bypass exists for your vendor it exists for every customer of your vendor, and your per-target fixed cost collapses back into the industry's shared amortisation pool. You are buying time, on a lease, at a renewal price set by your vendor's market share.


The list is a depreciating asset, and that sets the clock

Now the part that I think is genuinely under-modelled by defenders.

A combolist is inventory, and inventory depreciates. A freshly assembled corpus — credentials from a breach not yet widely traded, or a fresh harvest from infostealer logs — is worth dramatically more per pair than the same corpus after every operator in the market has run it against the top few hundred targets. The reason is mechanical: the first operator to run a list against your site finds every valid pair in it. The second finds the ones the first didn't monetise, minus the ones you've since forced to reset. By the tenth pass the list is a public good with the value of one.

For the attacker, that means the scarce resource is not attempts, or proxies, or engineering. It is time and exclusivity. Hence tiered credential markets — early, narrow distribution at high prices, broad distribution later at commodity prices — and hence the operators with the best economics being the ones closest to the source of fresh material. It's also why infostealer logs have become such a significant input: unlike a historical dump they arrive continuously, they carry session cookies and exact URLs alongside the credentials, and freshness is intrinsic rather than a decaying property.

For the defender, this converts an ambient problem into a scheduled one, which is a genuinely different way to allocate effort. Your exposure to stuffing is not a smooth function of your posture. It is a series of spikes, each keyed to the appearance of new corpus material containing your users' credentials — and the window in which that material is most dangerous to you is the window in which it is most valuable to the person holding it, measured in weeks rather than years. Between spikes your exposure decays on its own, because the pairs that work have been found and the rest discarded.

The consequence rarely gets acted on: the refresh rate of your breach-corpus checking matters more than its existence. A team checking against a corpus updated quarterly has a control whose effectiveness is decided by where in the quarter the breach landed. A team checking against a continuously refreshed corpus is intersecting the attacker's inventory while it still has value. Same control, same code path, wildly different position in the attacker's calendar.

And the corollary that catches people out: corpus checking is a subscription, not a migration. The common failure is a one-time sweep at rollout, resets forced on the hits, project closed. Every account that passed that day is unverified against every corpus published since. Since you cannot bulk-check your own stored hashes — salted and slow by design, which is the whole point — the only moment you'll hold the plaintext again is the next successful login. That makes the login-time check not a nice-to-have but the only mechanism that keeps the population current. The k-anonymity range query exists so you can do it without transmitting anything; this depreciation clock is why you run it on every authentication and not only at set-time.

There's a second-order idea in cburn too, worth stating even though few organisations are positioned to act on it. If a campaign against you reliably produces a forced-reset sweep across every account whose credential appears in the corpus that fed it, then attacking you destroys inventory the attacker intended to use elsewhere. You become an expensive target not by being hard but by being a place where lists go to die — one of the few defences that gets stronger as more organisations adopt it, because the depreciation compounds across targets.


Classification is the product

Now the reframing that I think earns this article its place.

An attacker holding ten million pairs does not want access to your site. Most of those pairs will never touch your site again. What they want is a label on each one: valid here, invalid here. A labelled pair is a different grade of inventory, worth substantially more than an unlabelled one, because the buyer's uncertainty is gone.

So run the process backwards. What is the manufacturing step? Not the login. It is the classification of your response. The attacker sends a pair, observes what comes back, and decides which bucket it goes in. Everything else — proxies, browser automation, solver calls — is logistics in service of that one measurement.

Which means your login endpoint has a second job most designs never articulate. Its first job is to admit legitimate users and refuse everyone else. Its second is to deny the attacker a clean signal about which case just occurred. Those jobs are in tension, and almost every login flow I've reviewed optimises the first while giving away the second for free.

The channels are more numerous than the error message everyone thinks about:

  • The error taxonomy. "No such user" versus "wrong password" is the textbook case and most teams have fixed it. Fewer have checked that the account exists question isn't answered elsewhere — registration ("this email is taken"), password reset ("we've sent a link" versus "no account found"), a public profile route, an invitation flow, an API returning 404 versus 403.
  • Timing. A correct password gets hashed and verified; a nonexistent user often short-circuits before the KDF runs. With Argon2 or bcrypt at sensible parameters that gap is tens of milliseconds — enormous, trivially measurable over a few samples, invisible to anyone reading the code. The mitigation is verifying against a dummy hash for unknown users, which everyone has read about and a surprising number of systems don't do, because the fast path got optimised later by someone who didn't know why the slow path existed.
  • The shape of what comes back. Response size, header set, redirect target, whether a CSRF token or a partial-authentication cookie is issued. A partial-auth cookie set before the second factor is a perfect oracle: its presence is the answer.
  • Downstream behaviour. Whether this attempt triggers throttling, whether the second try differs, whether a challenge appears at all. Anything that only happens to real accounts labels real accounts.
  • The second factor itself. The big one, and the hardest to fix, because it isn't a bug. If a valid pair causes a push or an SMS to be dispatched and an invalid one doesn't, the presence of the challenge is the classification. The oracle's output is delivered to the user's phone, and the attacker doesn't even need to see it — only that your server behaved as though it had somewhere to send it.

Uniformity isn't free, and I'd rather say so than pretend. Perfect indistinguishability is unachievable and past a point undesirable: you must eventually tell a real user something actionable, or support cost and abandonment will pay for the attack on the attacker's behalf. Generic errors make legitimate failures harder to self-diagnose. Constant-time responses mean paying full KDF cost for garbage traffic, which is the capacity problem again. Suppressing "this email is already registered" makes signup worse for everyone.

So the useful framing isn't "eliminate the oracle." It's: know what your oracle is, decide deliberately what it costs to consult, and make sure the cheapest signal isn't the most informative one. A system where an attacker must complete a full challenge flow to learn a pair's validity is meaningfully more expensive to classify against than one where a 200-byte difference in the response says it. That is a fixed-cost increase per target — the kind that actually bites.

Here's the pipeline, with the defences placed where they actually act rather than where we usually talk about them:

flowchart LR
    L["combolist<br/><i>depreciating</i>"] --> A["attempt<br/>reaches auth"]
    A --> C{"classification:<br/>is this pair valid?"}
    C -->|labelled valid| S{"survives<br/>controls?"}
    C -->|labelled invalid| X["discarded"]
    S -->|yes| M["session"]
    S -->|no| G["<b>validated pair,<br/>no session</b><br/><i>inventory, upgraded</i>"]
    M --> $["conversion<br/>to money"]
    G -.->|resold, targeted later| $
    style G fill:#fdd,stroke:#c66
    style C fill:#ffe9c7,stroke:#c99

The red box is the whole argument. Every control teams describe as "stopping credential stuffing" sits on the edge between classification and session — MFA, device reputation, step-up, bot scores. None of them touch the classification step, and the output of that step is the product. An attempt that gets blocked at the second factor has still successfully manufactured a labelled pair, and that pair has appreciated, because the buyer now knows it works.


"MFA stopped it" is an incomplete sentence

Which brings us back to the opening scene.

A validated pair that hits MFA and fails is not a defeated attack. It is a successful validation with a deferred payoff, and the deferral is short. The routes from "the password works but there's a second factor" to "we're in" are well-worn: a phishing page built for one named person now known to be a real customer; a real-time proxy relaying the challenge; an MFA-fatigue campaign, which is a numbers game that only makes sense once you know the password is right (the arithmetic is here); a SIM port; or a phone call to a help desk armed with enough correct information to sound like an owner. Each is expensive per target, and each becomes economically rational precisely because validation has already de-risked it. Targeted attacks are expensive; targeted attacks against a pre-verified list are a different business.

So the operational conclusion, small to implement and rare to find implemented:

Your breached-credential response must trigger on the "password was correct" event, not on the "login succeeded" event.

Most systems emit success after all factors pass and failure when any factor fails. The interesting bucket — first factor correct, subsequent factor unsatisfied, session never established — collapses into "failure" with a reason code nobody alerts on. Neither a success nor a block, so it appears on neither dashboard, which is exactly what happened in the incident I opened with.

Treat it as what it is: a confirmed disclosure of a live credential to an unknown party. The credential is compromised whether or not a session was created, so the response should look like any other credential compromise — invalidate the password, require a change, notify the user with a message that says what actually happened, raise the assurance requirement on that account for a while. Weight it by context: a correct password from a device and network you've never seen, on an account whose credential also appears in a corpus, during elevated attempt volume, is about as clear a signal as authentication ever produces.

The counterargument is real. Users mistype codes, abandon pushes, get new phones, travel. Fire a reset on every abandoned challenge and you generate a wall of support contacts and train users to expect reset prompts, which is its own security problem. So the rule needs a discriminator — new device and unrecognised network and no prior success from this context — rather than the raw event. That's a tuning problem, and a much better one to have than not knowing the event happened.

The same logic runs one step further out. If your bot-mitigation layer blocks an attempt before the password is checked, you have prevented a validation — arguably the only place in the stack where blocking prevents manufacturing rather than merely preventing access. Worth knowing when you decide where in the request path your defences sit. A control that runs after credential verification cannot protect the credential. It can only protect the session.


The conversion step, which is somebody else's budget

Now the term with the best returns and the worst organisational odds.

A validated credential is worth nothing until it converts. At the end of the chain is a payout, a shipment, a points transfer, a resale, a gift-card purchase — a step that touches money rails, involves human labour, and is far less anonymous than an HTTP request from a residential proxy. Compared to everything upstream, conversion is slow, expensive, attributable, and often reversible.

Which makes V × pconvert — the term most authentication teams never think about — frequently the cheapest to move. A new-payee cooling-off period, a hold on the first payout to a fresh destination, a cap on points transfers within a window of a credential change, a delay on shipping-address changes, a re-authentication requirement bound to the specific action rather than the session: each directly reduces the realisable value of a taken-over account, and each is invisible to the overwhelming majority of users who don't do those things in that pattern.

None of which is an argument for building fraud logic into your identity platform — that boundary is real and its reasons are load-bearing, and I've argued the separation at length in Identity Isn't Your Fraud Engine. The point is about where in the estate the highest-return control sits, not which service hosts it.

And the honest organisational observation: conversion-side controls are under-deployed against takeover not because anyone analysed the economics and declined, but because the identity team owns the login endpoint, the payments or risk team owns the payout hold, and takeover lands in the identity team's incident reviews. The team holding the incident does not hold the highest-leverage lever, so it optimises the levers it does hold — the ones on the small terms.

If you take one organisational action from this article: put the login-side and conversion-side owners in the same review, with the same takeover number. That hour is usually the most productive anyone spends on this problem, and it tends to end with a two-week change on the payments side that outperforms a quarter of login work.


The ledger

Sorting every common control by which term it moves clarifies a lot of arguments.

Control Term Honest assessment
Breach-corpus check at set-password time pvalid Real, cheap, necessary. Governs only new credentials — the existing population is unaffected.
Breach-corpus check at login, continuously refreshed pvalid The strongest routinely-available control. The only mechanism that reaches credentials set years ago, and its value is set by corpus freshness.
Password manager adoption / credential independence pvalid → 0 per account Structurally the largest effect and not yours to deploy. Your job is not to obstruct it.
Phishing-resistant credentials (passkeys, WebAuthn) removes the term The only defence that eliminates rather than compresses. See the closing section.
Response uniformity, timing equalisation, error taxonomy classification cost Underrated, cheap, mostly a design discipline rather than a project. Raises fixed cost per target.
Bot mitigation before credential verification pvalid observed, and capacity Genuinely prevents manufacturing. Effectiveness decays; renew the assumption, not just the contract.
Device / session reputation psurvive Good, and it composes well with everything else. Weakest exactly when it matters most: a genuinely new device.
MFA / step-up psurvive Large effect on sessions, zero effect on validation. Necessary, insufficient, and routinely mis-scored as a complete defence.
Session scope restriction (authenticated but read-only pending step-up) pconvert Underused. Decouples "we let you in" from "we let you move money."
Payout holds, cooling-off, transaction risk V × pconvert Frequently the highest-return term and almost always in another team's backlog.
Forced reset on corpus hits after a campaign cburn Destroys the attacker's inventory beyond your own perimeter. Rare, and structurally deterrent.
Password composition rules nothing Governs a variable that stuffing does not read. Forty years of evidence.
Forced periodic rotation nothing, or worse Moves no term; degrades credential quality and pushes users toward predictable mutation.
Marginal per-attempt friction cattempt Moves the smallest term. Changes attacker allocation, not attacker outcome.
Account lockout on failure count negative Never fires against an attack producing one failure per account, and hands anyone who knows a username a free denial-of-service primitive.

Two rows cut against instinct. MFA sits in the middle of the table, not the top — which is not a claim that MFA is unimportant (deploy it everywhere) but a claim about which term it moves, and the observation that a defence acting only on psurvive leaves the manufacturing line running at full speed. If your entire anti-stuffing strategy is "we have MFA," your position is that you're content to be the industry's credential-validation service as long as the pairs get cashed somewhere else.

And lockout is the only row with a negative sign, for a reason this lens makes plain: it converts an attack whose cost the attacker pays into an attack whose cost you pay — the exact opposite of the direction every other row moves.


Measuring it without fooling yourself

Architects have to justify this work, so it matters what goes on the dashboard. Most of what gets reported here is noise dressed as a metric.

Attempt volume is not a security metric. It is set by the attacker's list size, their target selection, and how many operators happen to be running this month. It can rise by a factor of forty while your exposure falls — that's the opening chart of the password managers piece. "We blocked 40 million malicious login attempts this quarter" tells your board about the attacker's inventory, not your posture. Block counts are the same number wearing a badge, since blocks scale with attempts: a quarter where blocks doubled is indistinguishable, from outside, between "we got better" and "someone bought a bigger list."

The metric that means something is the validity rate of attempted pairs. Of the pairs presented in traffic you believe to be automated replay, what fraction verified against a stored credential? That is pvalid — the attacker's own hit rate, computed by you — and it measures your population's exposure directly. It moves when your users' credentials stop appearing in corpora and doesn't move when the attacker buys more proxies. Track its trend rather than its level, because the level depends on classification choices you'll keep changing.

Then the nasty confound. The better your pre-verification bot mitigation, the blinder you are to your own validity rate, because the traffic that would have told you never reaches credential verification. You destroyed your own measurement in the course of defending yourself. That's a fine trade, but know you've made it and recover the measurement elsewhere: sample the blocked population and evaluate out-of-band, instrument the layer below the blocker, or lean on the corpus-membership rate of your active population — the cleanest available proxy, measurable at every successful login with no attacker traffic at all.

Three other signals, each with a caveat:

  • Distinct usernames per source, per window. A stuffing run produces one failure per account and thousands per source; per-account counters are structurally blind to that, which is the spray argument from the lockout piece. Source attribution across residential pools is weak, so treat the ratio as a campaign detector, not an identifier.
  • The correct-password-then-blocked bucket, as a first-class series. Its rate of change is your best early indicator of a targeted follow-up wave.
  • User-reported takeovers, understood as badly lagged and badly biased. Reports arrive weeks late, only from users who noticed, disproportionately where the loss was visible. Useful for postmortems, useless quarter-over-quarter — and expect the profile of a successful takeover to worsen over time even as total harm falls, because the remaining victims are adversely selected.

The general discipline: a metric the attacker controls is not a KPI. Attempt volume, block counts and solve rates all have the attacker's hand on the dial. Validity rate and corpus-membership rate have your users' hand on it, and your users are the population you're trying to change.


What to change

Not a summary. Seven things, in the order I'd do them, roughly by return on effort.

1. Instrument the correct-password-but-no-session event, then decide what it triggers. Emit it as its own event type with its own reason code, graph it, and write the rule that responds — password invalidated and change required when the context is unrecognised, with a discriminator strict enough not to punish someone whose phone battery died. A day of work, and it closes the gap that produced the incident at the top of this article. Nearly every system I've looked at has this event; almost none alert on it.

2. Move breach-corpus checking from set-time to set-time plus every login, against a continuously refreshed corpus. Login is the only moment you hold the plaintext of a credential set before you had the check, so it's the only path that reaches your legacy population. Use the range-query construction or an offline corpus. Act on hits proportionately — authenticate and interrupt, don't block — and make refresh cadence an operational commitment with an owner, because freshness is the control's actual strength.

3. Audit your validation oracles deliberately. Take three attempts — nonexistent user, valid user with wrong password, valid pair — and diff everything: status, body length, headers, cookies, redirect target, and above all wall-clock timing over enough samples to see the distribution. Repeat for registration, password reset, and any API touching account existence. Fix what's free, write down what isn't, and make sure the cheapest-to-observe signal isn't the most informative. Confirm specifically that unknown users still incur a KDF-equivalent delay; that mitigation is famous and frequently removed by a well-meaning optimisation.

4. Put a validity rate on the dashboard and retire the block count. Of pairs presented in traffic you consider automated, what fraction verified — alongside the corpus-membership rate of your active population. Then note explicitly where bot mitigation has blinded the first measurement and how you're recovering it. Both numbers are uncomfortable at first. That's what makes them useful.

5. Get conversion-side controls into the same review as login-side ones. One meeting, one takeover number, both owners present. What can a taken-over account extract in the first hour, what holds exist, and what would a payout delay or new-payee cooling-off cost in legitimate friction? Session scope restriction is the identity-side half of the same idea: authenticated is not the same permission as authorised to move money, and most systems merged them by default rather than on purpose.

6. Delete the controls that move nothing. Composition rules, forced rotation, and failure-count lockout on ordinary accounts. Each carries an argument that predates the attack we actually face, each costs real support volume or credential quality, and none appears on the profitable side of the ledger. The rotation conversation with your auditor is a separate battle with its own article; fight it, but don't wait on it to fix what you already control.

7. Stop obstructing the defence you didn't deploy. Allow paste, allow long credentials, use the standard autocomplete tokens, keep field names stable, publish /.well-known/change-password. Every control that fights a password manager pushes a user back into the reused-credential population — the only population stuffing works against. One-line changes that move the term at the top of the ledger, still shipping broken on a remarkable number of login pages.


The only term that reaches zero

Everything above is margin compression. It is worth doing — margin compression against a thin-margin business is a legitimate and effective strategy, and an attack that stops clearing its cost stops being run. But it is honest to name what it is.

Each of those controls multiplies a term by something less than one. Corpus checking shrinks pvalid for the accounts it reaches, and there will always be credentials in no corpus yet. Uniformity raises classification cost, and classification never becomes impossible, only slower. MFA cuts psurvive hard, and leaves the validation line running. Payout holds cut pconvert, and a patient attacker waits out a hold. None of them reach zero, because all of them are defences applied around a shared secret that the user can transmit to anyone, that you store a verifier for, and that is replayable by construction.

There is exactly one move that takes a term to zero, and it is removing the replayable secret. A public-key credential cannot appear in someone else's breach corpus, because the relying party never held anything worth stealing. It cannot be replayed, because each assertion signs a fresh challenge. It cannot be offered to the wrong site, because the origin is inside the signed data rather than inside the user's judgement. pvalid for an account with no password isn't small; it is undefined, in the way that a lookup against a table with no rows is undefined. The manufacturing line has nothing to manufacture. Password Breaches Only Matter If You Have Passwords makes that case in full, and the historical arc explains why it took six decades to get there.

Which sets up the honest framing for anyone building a roadmap. The stuffing defences in this article are what you run while you shrink the population of accounts that still have a replayable secret — and you will be running them for years, because that population shrinks slowly, because recovery paths keep a password alive behind the passkey, and because enrolment coverage is not the same as password removal. Do the work. Sequence it by which term it moves. But hold the two things in your head at once: you are compressing a margin, and someone else's fresh breach can restore it overnight.

The number that ends this, eventually, is not your block count. It is the fraction of your accounts for which there is no longer anything to replay.