The Cost of Password Hashing
The finding in the pentest report was one line: bcrypt work factor of 10 is below current guidance; recommend 12 or higher.
It took a sprint. The change itself was a constant in a config file, plus the rehash-on-login plumbing to convert the existing population. It shipped, the finding closed, and the report went into the compliance folder.
Fourteen months later a FinOps review flagged an autoscaling group called auth-worker at $1,240 a month. Nobody on the call could say what it did. It turned out to be the hashing pool, which had quadrupled in size over that sprint, because bcrypt cost 12 is four times the CPU of cost 10 and the group was sized for the login peak.
Two other prices had been paid at the same time and never attributed. Login p50 went from 340 ms to 590 ms, which the frontend team logged as "backend regression, no owner." And the p99 got worse by more than the p50, because a 320 ms service time means any queueing at all is visible — the quantum of delay is now the work factor itself.
So the purchase cost about $15,000 a year, a quarter of a second on every login, and a measurable dent in the signup funnel. Fine — security costs money. The question nobody asked in either direction is the one this article is about: what did it buy?
The answer, worked below, is roughly this. Against an attacker willing to spend $10,000 on GPU rental, the increment from cost 10 to cost 12 moved the attacker's reachable guess depth from about 290,000 guesses per account to about 72,000. That difference covers roughly 2.6% of the password population — around 36,000 accounts. Meanwhile 4.5% of the same population, around 63,000 accounts, had a password inside the first ten thousand guesses and was crackable for about $1,000 regardless of the work factor, and a breach-corpus check that would have removed most of them had been in the backlog for two years at zero recurring cost.
Password hashing is the one line item in your stack where you deliberately buy CPU-seconds as a security control. It's the only place where spending more compute is the correct answer rather than a bug. And almost nobody prices the purchase — not the dollars, not the latency, not the exchange rate against the attacker, and certainly not the marginal value of the last increment.
This is the compute chapter of a series doing arithmetic on identity infrastructure. The Hidden Cost of Audit Logs set the method; The Cost of Session Storage found that a storage bill was really a write bill; The Cost of Token Introspection vs Self-Validating Tokens found the infrastructure lines landing within a few hundred dollars of each other and payroll dominating both. This one has a different shape from all three, because here the cost is the product. You are not trying to minimize it. You are trying to buy the right amount.
The assumptions, published so you can re-run it
| Parameter | Value | Note |
|---|---|---|
| Registered users | 2,000,000 | Same scenario as the rest of the series |
| Accounts holding a password credential | 1,400,000 | Rest are federated or passkey-only |
| Password verifications | 60/s peak, 18/s average | ~46.7M/month |
| Peak-to-mean ratio | 3.3× | Sets the idle-capacity tax |
| vCPU | $25/vCPU-month | Blended on-demand/committed |
| Instance RAM | $3.50/GB-month | |
| Target CPU utilization | 40% | Because queueing |
| Marginal CPU price | $0.0000241/CPU-second | $25 ÷ 2.592M s ÷ 0.40 |
| Rented GPU, 24 GB consumer class | $0.50/hour | Marketplace mid-range |
| Attacker budget modelled | $10,000 | One offline campaign against a stolen table |
| Attacker perf-per-dollar improvement | 2× per ~2.5 years | Slower than it was; still real |
| Fully-loaded engineer | $16,700/month |
Every hashing figure below is a single-core measurement on modern x86, and they vary by a factor of three across CPU generations and library builds. Re-measure on your own instance type. The structure survives; the constants don't. The GPU figures are worse than that — they're a moving target with a public benchmark culture, and I'll flag where the uncertainty actually matters.
The defender's price list
Start with what a parameter set costs you. CPU time per verification, memory held per verification in flight, and the pool you must provision for a 60/s peak at 40% utilization.
| Parameters | CPU ms | MiB in flight | Provisioned vCPU @ 60/s | $/month | Marginal $/M verifications |
|---|---|---|---|---|---|
| PBKDF2-HMAC-SHA256, 600k iters | ~250 | ~0 | 37.5 | $938 | $6.03 |
| bcrypt cost 8 | 20 | 0.004 | 3.0 | $75 | $0.48 |
| bcrypt cost 10 | 80 | 0.004 | 12.0 | $300 | $1.93 |
| bcrypt cost 12 | 320 | 0.004 | 48.0 | $1,200 | $7.71 |
| bcrypt cost 13 | 640 | 0.004 | 96.0 | $2,400 | $15.42 |
| scrypt N=16384, r=8, p=1 | ~60 | 16 | 9.0 | $225 | $1.45 |
| Argon2id m=12288, t=3, p=1 | 42 | 12 | 6.3 | $158 | $1.01 |
| Argon2id m=19456, t=2, p=1 | 45 | 19 | 6.8 | $169 | $1.08 |
| Argon2id m=47104, t=1, p=1 | 54 | 46 | 8.1 | $203 | $1.30 |
| Argon2id m=65536, t=3, p=4 | 227 | 64 | 34.1 | $850 | $5.47 |
| Argon2id m=2097152, t=1, p=4 | ~2,420 | 2,048 | 363 | $9,075 | $58.32 |
Three things to notice before we go anywhere near the attacker.
Argon2id's CPU cost is very nearly linear in m × t. Calibrating from the middle of the table, about 1.18 ms per MiB-pass. That's the whole model, and it means the five parameter sets OWASP publishes as alternatives — m=47104,t=1 through m=7168,t=5 — are iso-cost for you. They were constructed that way. Hold that thought; it becomes the most useful decision in this article.
The last row is why nobody follows RFC 9106's headline recommendation. The RFC's first option, 2 GiB with t=1, is not expensive because of CPU — although $9,000 a month is real. It's unaffordable because 60 verifications in flight at 2 GiB each is 120 GiB of resident memory, and there is no instance shape where that is a sane thing to buy for an interactive login path. That recommendation is written for key derivation at rest, not for a login form, and the RFC's second option (64 MiB, t=3, p=4) is the one that belongs in an authentication service.
The p parameter is a latency knob, not a security knob — and it's worth exactly nothing at peak. p=4 splits the same total CPU work across four cores, so wall-clock time falls ~4× while your bill doesn't move. That's a genuinely good trade when cores are idle. At your peak, when all cores are busy, p>1 buys you nothing at all: there is no spare core to parallelize onto, and you've just made the scheduler's job harder. So p improves your p50 and does approximately nothing for your p99, which is the metric that actually governs your provisioning and your client timeouts. Set it, but don't count it.
The attacker's price list
Now the other side of the trade. One rented 24 GB consumer-class GPU at $0.50/hour, published-benchmark ballparks, 2026:
| Scheme | Hashes/sec/card | $ per million guesses |
|---|---|---|
| MD5, unsalted | ~1.6 × 10¹¹ | $0.0000009 |
| SHA-256, unsalted | ~2.2 × 10¹⁰ | $0.0000063 |
| PBKDF2-HMAC-SHA256, 600k | ~18,300 | $0.0076 |
| bcrypt cost 8 | ~22,400 | $0.0062 |
| bcrypt cost 10 | ~5,600 | $0.0248 |
| bcrypt cost 12 | ~1,400 | $0.0992 |
| scrypt 16 MiB | ~3,000 | $0.046 |
| Argon2id m=19456, t=2 | ~2,500 | $0.056 |
| Argon2id m=65536, t=3 | ~500 | $0.281 |
The MD5 row is there to set the scale. A billion guesses against an unsalted MD5 table costs about a tenth of a cent. Against Argon2id at 64 MiB it costs $281. That's the eleven orders of magnitude the last twenty years of KDF design bought, and it is the reason this whole discipline exists.
But look at what happens when you put the two price lists side by side, per hash:
bcrypt cost 12 costs you $7.71 per million verifications and costs the attacker $0.099 per million guesses. Their hash is 78 times cheaper than yours.
That number surprises people, and it should be stated plainly because the folklore has it backwards. The asymmetry in password hashing does not come from your hardware being better than the attacker's. It comes from arithmetic on both sides being run over very different quantities. A user logs in perhaps 30 times a month; you buy 30 hashes. An attacker needs 10⁵ hashes to reach that same user. Their per-hash cost advantage of 78× is swamped by a guess-count disadvantage of 3,300×, and the net leverage in your favour is about 43:1.
That is the entire defense, and it is much narrower than most people assume.
It's also worth naming where a chunk of the attacker's 78× comes from: they don't have a p99. They run batched, at 100% utilization, with no latency SLO, no headroom, and no idle capacity held against a peak. Your 40% utilization target and your 3.3× peak-to-mean ratio are, between them, an 8× tax the attacker simply doesn't pay. Roughly a tenth of their advantage is silicon; the rest is that you have users and they don't.
The exchange rate, and why bcrypt's never improves
Divide one price list by the other. Attacker dollars imposed per defender dollar spent:
| Scheme | Defender $/M | Attacker $/M | Attacker $ per defender $ |
|---|---|---|---|
| PBKDF2-HMAC-SHA256, 600k | $6.03 | $0.0076 | 0.0013 |
| bcrypt cost 8 | $0.48 | $0.0062 | 0.0129 |
| bcrypt cost 10 | $1.93 | $0.0248 | 0.0129 |
| bcrypt cost 12 | $7.71 | $0.0992 | 0.0129 |
| bcrypt cost 13 | $15.42 | $0.198 | 0.0128 |
| scrypt 16 MiB | $1.45 | $0.046 | 0.032 |
| Argon2id m=12288, t=3 | $1.01 | $0.053 | 0.052 |
| Argon2id m=19456, t=2 | $1.08 | $0.056 | 0.052 |
| Argon2id m=47104, t=1 | $1.30 | $0.067 | 0.052 |
| Argon2id m=65536, t=3 | $5.47 | $0.281 | 0.051 |
Two structural facts fall out, and they're the most useful things in this article.
Within an algorithm, the exchange rate is constant. Every bcrypt row is 0.0129. Every Argon2id row is 0.052. Cranking a work factor buys you more security; it never buys you a better deal. You are moving along a line, not onto a better one. This is why "just raise the cost factor" is a weak answer to a pentest finding — it's the only lever with a guaranteed-flat return.
Between algorithms, the exchange rate varies by 40×. Argon2id is a 4× better buy than bcrypt and a 40× better buy than PBKDF2, at every point on both curves. PBKDF2 is the worst purchase on the page by a wide margin, and it's still the default in a distressing number of frameworks and the only option in several compliance-driven configurations.
Which produces an uncomfortable, checkable result for anyone currently sitting on bcrypt 12 and planning to "upgrade" to OWASP's Argon2id defaults.
| Configuration | Your monthly bill | Attacker guess depth at $10,000, per account |
|---|---|---|
| bcrypt cost 12 | $1,200 | 72,000 |
| Argon2id m=19456, t=2, p=1 | $169 | 128,500 |
| Argon2id m=65536, t=3, p=4 | $850 | 25,400 |
Migrating from bcrypt 12 to OWASP's mid-range Argon2id parameters cuts your bill by 86% and makes the attacker's job 1.8× easier. It is a cost optimization wearing a security upgrade's clothing, and I have watched two teams ship it believing the opposite. The correct destination is the third row: 71% of the bcrypt-12 bill for 2.8× the attacker cost. Same project, same migration mechanics, one different constant — and the difference between the two is entirely invisible unless you have done this arithmetic.
Memory versus time: a bet on which attacker you face
So within Argon2id, how should you split a fixed budget between m and t?
The folklore says memory hardness is strictly better, because memory is what defeats GPUs. The GPU column above says otherwise: all four Argon2id rows have the same exchange rate. That isn't a rounding artifact. A GPU running Argon2id is memory-bandwidth-bound, so its throughput falls as 1/(m × t) — the same denominator as your CPU cost. Trading t=3, m=12288 for t=1, m=47104 moves nothing against a rented GPU.
The m² advantage that memory-hardness papers describe is real, but it appears in a different attacker's budget. An attacker building or renting purpose-designed hardware pays for memory as silicon area, and must hold m bytes for a duration proportional to m × t. Their area-time product scales as m² × t while your cost scales as m × t. Every doubling of m, with t halved to keep your bill flat, halves a hardware attacker's efficiency advantage and does nothing to a GPU attacker's.
The same logic explains bcrypt's flat exchange rate, and it's the real indictment. bcrypt's working set is 4 KiB. Four kibibytes fits in the block RAM of an FPGA thousands of times over, so a purpose-built bcrypt cracker replicates cores until it runs out of die, and its efficiency advantage is enormous and independent of the cost factor. Raising bcrypt's cost slows every attacker by the same factor, including you. Raising Argon2id's m slows the well-capitalized attacker faster than it slows you. That is the entire difference between the two algorithm families, expressed as economics rather than as cryptography.
There's a second GPU-side effect at large m that the bandwidth model misses: VRAM capacity. At 64 MiB a 24 GB card holds 384 concurrent lanes and stays comfortably bandwidth-bound; at 2 GiB it holds twelve and falls off the bandwidth curve into capacity starvation. The quadratic starts to bite on commodity hardware somewhere between those two points, and exactly where depends on a card generation you can't predict.
So the practical rule, stated honestly:
Prefer memory to passes, not because it's a better deal against the attacker you can see, but because it's free to choose and it's the only dimension that degrades the attacker you can't. Then stop when either your RAM budget or your tradeoff-resistance margin binds.
The margin caveat matters. Argon2id at t=1 has less headroom against time-memory tradeoff attacks than at t≥2, and the guidance to keep t at 2 or 3 exists for that reason rather than for cost. m=65536, t=3 is a defensible place to land: it takes the memory dimension seriously, keeps the tradeoff margin, and costs less than bcrypt 12. m=1048576, t=1 would be a better exchange rate against an ASIC and a worse bet against a cryptanalytic result, and you are not in a position to price that risk. Neither am I.
What one increment actually buys, in accounts
Exchange rates are abstract. Convert to the only unit that matters: how many of your users stop being crackable.
An offline attacker with budget B works down an optimized guess ordering. Salting makes their cost per (guess, account) pair, so reachable depth per account is B ÷ (price per guess × account count). At $10,000, bcrypt 12 and 1.4M password rows, that's 72,000 guesses per account.
Then you need a crack curve — the fraction of a population whose password lies within the first N guesses. Yours is the only one that counts, but published studies of consumer and workforce corpora cluster around this shape:
| Guess depth | Cumulative % of accounts |
|---|---|
| 10² | ~1.0% |
| 10³ | ~2.5% |
| 10⁴ | ~4.5% |
| 10⁵ | ~9% |
| 10⁶ | ~15% |
| 10⁷ | ~22% |
| 10⁸ | ~30% |
Roughly 4.5 percentage points per decade in the region that matters. A work-factor doubling halves the reachable depth, which is 0.30 decades, which is about 1.35 percentage points.
At 1.4M password accounts:
| Change | Depth: from → to | Accounts rescued | Monthly cost | $ per account-year protected |
|---|---|---|---|---|
| bcrypt 11 → 12 | 145k → 72k | ~18,900 | +$600 | $0.38 |
| bcrypt 12 → 13 | 72k → 36k | ~18,900 | +$1,200 | $0.76 |
| bcrypt 13 → 14 | 36k → 18k | ~18,900 | +$2,400 | $1.52 |
| bcrypt 12 → Argon2id 64 MiB/t=3 | 72k → 25k | ~28,900 | −$350 | — |
There is the whole shape of the problem in one table. The benefit per increment is roughly constant in accounts. The cost per increment doubles. A concave benefit curve against an exponential cost curve always has an interior optimum, and it is always at a finite, unremarkable work factor. This is the arithmetic reason why "crank it as high as your latency budget allows" is bad advice and why the published guidance settles where it does.
It also gives you a genuinely usable decision rule. $0.76 per account-year is a good purchase if a compromised account carries $50 of expected loss — chargebacks, support handling, breach notification, churn. It's a bad purchase at $2 an account, which is what a dormant free-tier signup is worth. You cannot answer the parameter question without a number for what an account is worth, and most parameter decisions are made by people who have never been given one.
Now the comparison that should reorder your backlog. That table's rescued band sits between 18,000 and 145,000 guesses. Below it, at depth 10⁴, sit 63,000 accounts that no work factor will ever save — reachable for about $1,000 at bcrypt 12 and about $200 at Argon2id's OWASP defaults.
Those accounts are removable, and not with CPU. A k-anonymity breach-corpus check at registration and password change removes most exact matches to leaked credentials, which is the densest part of that region — call it 5 percentage points, or 70,000 accounts, for a bounded engineering cost and essentially zero recurring spend.
| Control | Accounts moved out of reach | Recurring cost |
|---|---|---|
| One work-factor increment | ~18,900 | $7,200–$28,800/year |
| Breach-corpus rejection | ~70,000 | ~$0/year |
| Passkey adoption at 40% | 560,000 rows, if you delete the password | Negative (see below) |
The work factor and the corpus check are not substitutes — they act on different parts of the distribution, and you want both. But if you are choosing what to do next, the corpus check is somewhere around two orders of magnitude the better buy, and it is nearly always further down the backlog than the work-factor ticket, because the work-factor ticket is the one that appears in pentest reports.
The parameters you chose are a liability with a maturity date
Everything above prices a decision made today. The awkward part is that a hash's cost parameters are frozen at write time and verified for as long as the record lives.
Two consequences that people routinely conflate into one.
Your bill is set by the parameter distribution of your active population. It converges to current policy quickly, because rehash-on-login converts 70–80% of monthly-actives within a month.
Your exposure is set by the parameter distribution of every stored row. It converges to current policy never. Six months into a conversion, the active population is done and 30% of all rows are still on the old parameters — including every account that hasn't logged in since 2019, which is disproportionately the population with the weakest passwords and no MFA.
I'm not going to re-derive the migration mechanics; Migrating Password Hashes Without Resetting Every User covers rehash-on-login, wrapped hashing, peppering, and the tail decision in detail. What that article treats as a project, this one treats as a recurring liability, and the missing piece is depreciation.
Attacker perf-per-dollar has roughly doubled every two and a half years — slower than in the GPU boom, but persistent. A doubling of attacker throughput is exactly one work-factor unit. So:
Your stored hashes depreciate at about 0.4 work-factor units per year, and only the ones that log in get topped up.
| Written | At | Effective work factor in 2026 |
|---|---|---|
| 2026 | bcrypt 12 | 12.0 |
| 2022 | bcrypt 12 | 10.4 |
| 2018 | bcrypt 12 | 8.8 |
| 2014 | bcrypt 10 | 5.2 |
That last row is what the histogram at the top of the migration article really means. Four million rows at cost 8 written in 2014 are, in 2026 attacker dollars, worth roughly cost 3 — which is to say, worth nothing. Depreciation is why a dormant tail is not a deferred decision but an accruing one, and it's the honest argument for bulk-wrapping or peppering the tail rather than waiting for logins that will never arrive.
It's also the argument for building needs_rehash against policy rather than algorithm identity from day one. If a parameter bump is a config change you can make one every couple of years and stay level with depreciation. If it's a project, you'll make one every eight years and lose three units in between.
Where the work runs, and the 65% of your bill that is idle
Back to capacity, because the parameter is only half the invoice.
Provisioning is set by peak. Busy CPU at the 18/s average is 4.1 cores for Argon2id at 64 MiB; at the 60/s peak it's 13.6. Provision the peak at 40% utilization and you buy 34 vCPU — $850 a month to serve work whose average demand is 4.1 cores. About 65% of your hashing bill is idle capacity held against a peak, and unlike the session store in the storage piece, the hashing pool genuinely does go quiet overnight. Sessions can't be scaled down because mobile records never drain. Hashing can.
Except that it can't, quite, and the reason is a timing mismatch worth stating precisely:
Password hashing is perfectly scalable and practically unscalable. It's pure stateless CPU, so it parallelizes without limit. But scale-up is measured in minutes — instance boot, image pull, runtime warmup, health check, load-balancer registration, three to six minutes end to end — and the arrivals that hurt you are measured in seconds.
A mass-invalidation herd (the deploy that logs everyone out walks through one, and I won't repeat the derivation) puts hundreds of thousands of logins into a queue inside a minute. A customer running a forced password reset across eleven thousand employees does the same thing on purpose. Autoscaling arrives after the incident is over. Which means the reserve you must hold is set by the largest step change that can occur inside your scale-up latency — and for identity, that step change is very large. So you hold peak capacity, permanently, and admission control (denominated in CPU cost, not requests) is what actually protects you. Capacity is how you avoid the incident; it is not how you exit one.
There is one genuinely interesting escape, and it's the only identity workload where the serverless economics come out ahead.
Password verification is stateless, CPU-bound, short, retry-safe, needs no local data, and has a 3.3× peak-to-mean ratio. Per-invocation pricing at roughly $0.0000166667 per GB-second, at 1.8 GB (about one vCPU) for a 227 ms Argon2id verification:
0.227 s × 1.8 GB × $0.0000166667 = $0.0000068 per verification
+ request charge = $0.0000070
× 46.7M verifications/month = $327/month
Against $850 provisioned. The crossover is computable and general: per-invocation compute is about 3.1× the price of provisioned compute per CPU-second, so it wins whenever your effective utilization — target utilization divided by peak-to-mean — falls below about 32%. Here it's 0.40 ÷ 3.3 = 12%. Serverless wins by 2.6×.
The honest counterargument is the one this article keeps returning to: cold starts land on p99, which is the metric that matters. Buying provisioned concurrency to fix that gives the saving straight back, and you've moved credential verification into a second execution environment with its own supply chain and secrets handling. Most teams should read this as evidence that their provisioned pool is 65% idle, not as a migration plan.
What is not optional is getting the hashing out of your general request-handling loop, and the economic argument is stronger than the operational one. You stop sizing your entire API fleet for the login peak; you get to pick an instance family for memory bandwidth rather than for whatever your web tier needed; and you stop paying an externality that never appears on your own budget line. Argon2id at 64 MiB, touched pseudo-randomly, flushes a shared 32–96 MiB L3 continuously — every co-resident workload's cache miss rate rises, and their p99 rises with it. That is the exact counting error the introspection piece names: a diffuse cost spread across other people's budgets is invisible, and invisible costs grow. Here it's a cache line rather than a CPU percentage, but it's the same failure.
Why the p99 is the number, not the mean
Three reasons, stacked, and they're the reason capacity models built from a benchmark run are optimistic by more than people expect.
Service time sets the quantum of delay. For a pool of c cores at utilization ρ with service time S, mean queue wait is roughly (ρ/(1−ρ)) × S/(2c). At ρ=0.4 and S=227 ms, a 34-core pool waits 2.2 ms on average — fine. But the p99 of that wait isn't 2.2 ms; the moment you queue behind even one in-flight verification, you've waited a full service time. The work factor doesn't just add to your latency, it sets the granularity of your latency. A pool with four cores and a 320 ms service time has a p99 you will be asked to explain.
Small pools pay more per login for the same security. Queue wait scales as S/c. A deployment doing 5 logins/second at bcrypt 12 needs 4 cores and gets a visibly worse tail than one doing 500 logins/second at the same parameters with 400 cores. Same U-curve that showed up in session storage: below a certain scale the fixed structure of the component dominates, and small deployments are systematically worse off at identical configuration.
Your verification cost is multimodal, and the p99 is set by the worst mode. A population mid-migration has native Argon2id records, legacy bcrypt records at two or three cost factors, and — if you took the wrapped-hashing route — records that pay bcrypt plus Argon2id. Your mean verification cost is a weighted average that describes no actual request. So admission control must be denominated in the worst-case cost class or it will admit work it cannot serve, and the per-hash number you feed your capacity model should be the p95 of your parameter distribution.
And one measurement trap: per-hash throughput degrades under concurrency and single-threaded benchmarks won't show it. Argon2id at 64 MiB, t=3 moves roughly 384 MiB of memory traffic per verification. Thirty-two concurrent hashers on one socket is tens of GB/s of sustained random-ish traffic, which is a large fraction of what the socket can deliver. Measured degradation of 1.3–1.6× at realistic concurrency is normal. That factor multiplies with the 2.5× provisioning factor, so a capacity model built from an isolated benchmark can be optimistic by 4× overall. Benchmark at your target concurrency or don't benchmark.
What passwordless does to the optimum — and what it doesn't
The obvious expectation is that as passkey adoption rises, hashing matters less and you should spend less on it. The arithmetic says something more interesting.
Your hashing cost scales with the number of password logins: N_password_users × login_rate × cost_per_verification. The security value of a work-factor increment scales with the number of password accounts rescued: N_password_accounts × 1.35pp × value_per_account. Both are linear in the size of the password population. Divide one by the other and it cancels.
The optimal work factor is invariant to your passkey adoption rate. Adoption changes the size of the bill, not the shape of the decision.
That's a cleaner result than the one I expected to find, and it means the common instinct — "we're going passwordless, so let's not over-invest in hashing" — is reasoning about the wrong quantity. What falls is your total spend, automatically, with no parameter change at all. What doesn't fall is the case for spending the right amount per verification.
Two corrections make it real, though, and the second is the one worth acting on.
The residual password population is adversely selected. The users who don't adopt passkeys skew toward old devices, shared workstations, low engagement, and no MFA — the same population whose passwords sit in the fat part of the crack curve. The residual crack curve is worse than the average curve you started with, which pushes the optimum slightly up, not down.
Adoption cuts your bill but not your exposure, unless you delete the credential. This is the part almost everyone misses. When a user enrolls a passkey and stops typing their password, they stop consuming your CPU immediately — the invoice responds the same month. But the bcrypt row is still in the table, still at whatever parameters it was written with in 2019, still depreciating, still crackable, and now never rehashed, because rehash-on-login requires a login that will never come. Passkey adoption silently converts your active, self-maintaining hash population into a dormant, decaying one. Fifty percent adoption doesn't halve your breach exposure; it halves the fraction of your hashes that are being kept current.
So the move that actually changes your position is deleting the password credential once a user has a working passkey and a recovery path that doesn't depend on it. That's a product decision with real edges — recovery, shared workstations, enterprise fallback, the support call from someone whose phone is in a lake — and Why Passwordless Isn't About Convenience makes the case that removing the asset class, rather than defending it, is the point. Economically, it's the only control on this page that reduces your bill and your exposure at the same time. Everything else trades one for the other.
The framework
Stop treating parameters as a constant to be copied and start treating them as a purchase with a stated threat model. Six steps, an afternoon:
- Name the attacker's budget. Not "an attacker" — a dollar figure. $1,000 is a hobbyist with a spare card. $10,000 is a serious credential-stuffing operation monetizing a dump. $1,000,000 is a state actor for whom none of this arithmetic applies and against whom the correct control is not having the password. Pick one and write it down.
- Convert budget to guess depth.
depth = budget ÷ (attacker $/M × your password row count ÷ 10⁶). Do it for your candidate parameter sets. - Get a crack curve for your own corpus. You can run this offline against a copy of your own hashes with a standard wordlist and rule set, and you should, because the table above is somebody else's population. This is the single highest-value measurement in this article and almost nobody does it.
- Price the band, not the total. An increment moves the depth by 2×; read the percentage-point difference off your curve; multiply by row count. That's what you're buying.
- Multiply by what an account is worth. If nobody in the organization will give you a number for expected loss per compromised account, that is itself the finding, and the parameter conversation cannot be finished until it's resolved.
- Compare against the alternatives at the same price. Corpus rejection, MFA on the accounts in the band, credential deletion for passkey adopters, and — always — the algorithm change, which is the only move on the list that improves the exchange rate rather than the quantity.
flowchart LR
T["Attacker budget<br/>(named, in dollars)"] --> D["Reachable guess<br/>depth per account"]
D --> B["Band moved by<br/>one increment"]
B --> A["Accounts rescued<br/>× value per account"]
A --> C{"Beats the same<br/>spend elsewhere?"}
C -->|no| X["Corpus check · MFA ·<br/>delete the credential"]
C -->|yes| P["Buy the increment"]
P --> R["Re-price annually:<br/>0.4 units/year decay"]
R --> T
The loop closing back on itself is the part that's usually missing. A parameter chosen once and never revisited is not a decision, it's a sediment layer — and it depreciates whether or not anyone is watching.
What to measure
Five numbers. Most teams running password authentication know none of them.
- Your parameter histogram, weighted two ways. By stored row (your exposure) and by daily successful login (your bill). They are different distributions and the gap between them is your migration debt.
- Your p95 verification CPU cost, at production concurrency. Not the mean, not a single-threaded benchmark. This is the input to your capacity model and it is usually wrong by 2–4×.
- Your peak-to-mean password-login ratio. It sets your idle-capacity tax, which is typically most of your hashing bill, and it decides whether the per-invocation option is worth modelling.
- Your own crack curve at four depths — 10³, 10⁴, 10⁵, 10⁶ — measured against your own hashes offline. Without this, every parameter argument in your organization is an aesthetic one.
- Expected loss per compromised account. Owned by someone outside engineering. Without it, step 5 of the framework has no denominator and you will default to whatever number a pentest report contained.
The closing thought
The team in the opening story eventually did the arithmetic. They moved off bcrypt 12 to Argon2id at m=65536, t=3, p=4, which cut the pool from 48 vCPU to 34, cut p50 login latency by 90 ms, and made the attacker's job 2.8× harder — all three at once, which is not a trade you get offered often and which was available the entire time. Then they shipped the corpus check that had been sitting in the backlog, which removed four times as many accounts from reach as the original work-factor increment had, for no recurring cost at all.
The pentest finding had been correct. Cost 10 was below guidance. What the finding couldn't tell them — what no finding of that shape can tell anyone — is that they were being asked to buy more of the worst-value item on the menu, at a price that doubles per unit, to protect a band of their population they had never measured, against an attacker whose budget nobody had named.
Every other line item in this series is a cost you're trying to minimize. This one is a cost you're trying to get right, which is harder, because there's no dashboard that goes green. The nearest thing to a rule I can offer is this: the first dollar goes to the algorithm, the second to the accounts no work factor can save, and only the third to the work factor itself. Most organizations spend all three on the third, because it's the only one that has a number in a report.