Longer Beats Complex: The Entropy Argument, With Actual Math
The policy change was approved unanimously, which should have been the first warning.
Minimum length went from eight characters to twelve. Four character classes became mandatory instead of three. The slide deck had a keyspace calculation on it: the old policy admitted about 2⁵² passwords, the new one about 2⁷⁸. Twenty-six bits of improvement. Sixty-seven million times harder to crack. Nobody argued.
Six months later, the same team did something almost nobody does: they ran the attack against their own corpus. Ten thousand randomly sampled accounts, a standard rule set, a guess budget of one million per account. Under the old policy, 15.2% of sampled accounts fell inside a million guesses. Under the new one, 11.9%.
Three and a bit percentage points. Converted to the unit the slide used — the horizontal shift of the crack curve — that is about 1.4 bits, not twenty-six. Meanwhile password-reset tickets were up about 40%, and the single most productive attack mask, ?u?l?l?l?l?l?l?l?l?d?d?s, now covered a larger fraction of the corpus than any single mask had before the change. The policy had not expanded the space users drew from. It had told them which corner of it to stand in.
That gap — 26 bits claimed, 1.4 bits delivered — is not a rounding error or a bad implementation. It is what happens when you apply a formula that describes a random generator to a population of human choices, and then measure the wrong statistic on top of it. This article is about doing both correctly, because the corrected math points somewhere quite specific, and it isn't where most password policies point.
Fair warning on scope: everything here concerns resistance to guessing. That's one threat among several, and I'll argue at the end it isn't the most important one — but it's the one every password policy claims to address, so it deserves an honest treatment.
Three quantities, all called entropy, and only one of them is about attackers
Take a probability distribution over passwords — the distribution your actual user population draws from, whatever that is. There are at least three numbers you can compute from it, and the folklore uses the word "entropy" for all three interchangeably.
Shannon entropy, H₁ = −Σ pᵢ log₂ pᵢ, is the average number of bits needed to encode a draw. It is the right measurement for a compression problem. It is not the right measurement for a guessing problem, and the reason is that averages are dominated by the bulk of a distribution while guessing is dominated by the head.
Guessing entropy, G = Σ i · pᵢ with the probabilities sorted descending, is the expected number of guesses an optimal attacker needs to find one specific password. Closer to the point, but it's still an expectation, and expectations over heavy-tailed distributions are famously uninformative.
Min-entropy, H∞ = −log₂ p_max, is the negative log of the single most likely password's probability. It is the attacker's success rate on their first guess.
These are not variations on a theme; they can differ by tens of bits on the same distribution. Suppose half your users pick 123456 and the other half generate a genuinely uniform password from a space of 2⁴⁰.
| Quantity | Value | What it says |
|---|---|---|
Shannon entropy H₁ |
21.0 bits | Average encoding cost of a password from this population |
log₂ of guessing entropy G |
≈ 38 bits | Expected guesses to break one specified account |
Min-entropy H∞ |
1.0 bit | One guess breaks half your users |
Same distribution. Three answers: 21, 38, and 1. The 21-bit figure is what a naive per-password entropy estimator would report on average. The 38-bit figure is what you'd get if you asked "how long to crack a random user." The 1-bit figure is the one that describes what actually happens when someone points a script at your login endpoint, because an attacker does not guess randomly and does not guess in an order you chose. They guess in descending order of probability, and they stop when they have enough accounts.
That last clause breaks the average. An attacker compromising a population is not trying to break your account; they are trying to break some accounts, cheaply. Every metric built on an expectation assumes they must finish. They don't — they take the head and leave.
(The word "entropy" also gets used in a thermodynamic, state-space sense — that argument is about counting microstates in a system nobody has enumerated. This one is strictly information-theoretic and strictly about a distribution over guesses. The two are unrelated except by name.)
The rigorous version of "they stop when they have enough" is Bonneau's partial-guessing metrics: λ_β, the fraction of accounts recovered within β guesses each, and G_α, the work needed to reach a fraction α of the population. These are curves, not scalars, and no single number substitutes for them. Studies of large consumer corpora using them land in an uncomfortable place: something like 21 bits of Shannon entropy, but under 7 bits of resistance to an attacker allowed a handful of guesses per account.
The entropy of a generator versus the entropy of a choice
Here is the sleight of hand at the centre of nearly every password strength claim you have ever read.
log₂(95¹²) = 78.8 bits is a true statement about a process: a uniform random draw of twelve characters from the 95 printable ASCII characters. It is not a statement about a password. A password is a string; strings do not have entropy. Only the process that produced them does.
When a password manager generates x7#Kq2mVp!Ld, that string carries 78.8 bits, because the generator is uniform and you know it. When a human types Summer2026!! — also twelve characters, also four character classes, also inside 95¹² — the process that produced it was not a uniform draw. It was a selection from a small mental catalogue, structured by natural language, calendar, keyboard layout, and the shape of the rule they were trying to satisfy. The character-space calculation describes a generator that was never running.
The industry has known this for twenty years. The fiction survives because it's computable: you can put log₂(95¹²) in a policy document and you cannot put "the conditional probability of this string under a PCFG trained on eleven leaked corpora."
The most consequential version of this fiction was the per-character estimator in the original NIST SP 800-63 Appendix A: four bits for the first character, two bits each for characters two through eight, 1.5 bits for nine through twenty, one bit thereafter, plus a six-bit bonus for satisfying composition rules and up to six more for passing a dictionary check. Apply it:
| Password | Appendix-A estimate | Realistic guess number | Real bits |
|---|---|---|---|
P@ssw0rd1! |
4 + 14 + 3 + 6 = 27 bits | inside the first 10⁴ of any rule-based run | < 14 |
Tr0ub4dor&3 |
4 + 14 + 4.5 + 6 = 28.5 bits | leetspeak rules over a common noun | ≈ 18–22 |
correcthorsebatterystaple (generated, 4 words from 7,776) |
4 + 14 + 18 + 5 = 41 bits | uniform over 7776⁴ | 51.7 |
correct horse battery staple (human-chosen four words) |
same ≈ 44 bits | word frequency is Zipfian, syntax constrains order | ≈ 25–30 |
Two things fall out of that table, and the second one is the part most treatments of this topic skip.
The estimator is biased in favour of composition rules by construction. It hands out a six-bit bonus for satisfying exactly the rule that concentrates the distribution. The bonus isn't merely wrong in magnitude; it has the wrong sign.
And the xkcd number does not transfer to a human-chosen passphrase. 4 × log₂(7776) = 51.7 bits is a statement about dice. A person asked to "think of four random words" samples from a few thousand high-frequency words, weighted by natural-language frequency, and tends to produce phrases with grammatical structure — which slashes the effective order space. Work on multi-word passphrases repeatedly finds human-chosen phrases falling tens of bits short of the uniform figure. The comic is right about the mechanism and wrong if read as licence to let users invent their own phrases and claim 44 bits.
The correct reading: passphrases are strong when a generator produces them and merely-better-than-average when a human does. That distinction should appear in your policy, because it changes what you build. A passphrase policy without a generator is a nudge. A passphrase policy with a generator is a control.
Length versus alphabet, with the actual bits
For a genuinely random generator, entropy is length × log₂(alphabet size). Both factors are available to you. They are not equally priced.
| Alphabet | Size | Bits per character |
|---|---|---|
| lowercase | 26 | 4.700 |
| lowercase + digits | 36 | 5.170 |
| mixed case | 52 | 5.700 |
| alphanumeric | 62 | 5.954 |
| printable ASCII | 95 | 6.570 |
| Diceware word (per word) | 7,776 | 12.925 |
The marginal moves, stated precisely:
- Add one character to a lowercase-only password: +4.700 bits. One keystroke, no shift key, no thinking.
- Expand a mixed-case password to full printable ASCII: +0.870 bits per character. So on a twelve-character password, the entire symbol set is worth 10.4 bits — but only if the symbols land in uniformly random positions drawn from the full set, which for a human-chosen password they never do.
- Expand alphanumeric to full ASCII: +0.615 bits per character. If you already allow digits and both cases, symbols are the cheapest bits on the menu, and you're charging the user a lot for them.
Two more lowercase characters roughly match the entire symbol set on a twelve-character password (9.4 bits versus 10.4), cost two keystrokes instead of twelve shift-modified ones, and survive being typed on a phone keyboard, a games console, and a badge-reader terminal in a warehouse.
The same point as strength targets — everything here is a generated password:
| Construction | Bits |
|---|---|
| 8 chars, alphanumeric | 47.6 |
| 8 chars, full ASCII | 52.6 |
| 12 chars, lowercase | 56.4 |
| 12 chars, alphanumeric | 71.5 |
| 12 chars, full ASCII | 78.8 |
| 16 chars, lowercase | 75.2 |
| 4 Diceware words | 51.7 |
| 5 Diceware words | 64.6 |
| 6 Diceware words | 77.5 |
| 7 Diceware words | 90.5 |
Note the pairs. Sixteen lowercase characters (75.2 bits), twelve full-ASCII characters (78.8), and six Diceware words (77.5) are the same security. They are emphatically not the same to a human being, and that difference is the whole reason a policy holds or fails.
What those bits buy in wall-clock time
Assumptions, stated so you can re-run them: an eight-GPU rig of 24 GB consumer-class cards, rented, using the same published-benchmark figures as The Cost of Password Hashing. That's 1.28 × 10¹² MD5/s, 1.12 × 10⁴ bcrypt-cost-12/s, and 4.0 × 10³ Argon2id (m=65536, t=3)/s. Times below are for exhausting the space; halve them for the expected time to a hit.
| Bits | Guesses | MD5, unsalted | bcrypt cost 12 | Argon2id 64 MiB |
|---|---|---|---|---|
| 40 | 1.1 × 10¹² | 0.9 seconds | 3.1 years | 8.7 years |
| 50 | 1.1 × 10¹⁵ | 15 minutes | 3,200 years | 8,900 years |
| 60 | 1.2 × 10¹⁸ | 10.4 days | 3.3 × 10⁶ years | 9.1 × 10⁶ years |
| 70 | 1.2 × 10²¹ | 29 years | 3.3 × 10⁹ years | — |
| 80 | 1.2 × 10²⁴ | 30,000 years | — | — |
Read the columns, not the rows. Against an unsalted fast hash, the threshold where brute force stops being a strategy sits somewhere around 65 bits. Behind a real KDF it sits at about 40. The password does not need to be strong in the abstract. It needs to be strong relative to the cost of a guess — which is why this table is meaningless without the hashing decision beside it, and why I'll come back to how the two combine.
Also note what the table can't show: no realistic human password sits at 40 or 50 bits of the attacker's actual guessing order. Brute force over a keyspace is the attack of last resort. Everything below concerns the attacks that come first.
Composition rules are a constraint, not an expansion
The intuition behind "require four character classes" is that it forces users into a larger space. This is exactly backwards, and it's worth doing the arithmetic because the result is sharper than the intuition.
A composition rule does not expand 52¹² to 95¹². Allowing symbols does that; requiring them removes strings. The compliant set is a strict subset of the space you already had. So the first question is how much the constraint costs in raw keyspace.
Count the twelve-character strings over 95 characters that contain at least one lowercase, one uppercase, one digit, and one symbol. By inclusion–exclusion over the four "missing a class" events:
P(missing lowercase) = (69/95)¹² = 0.0216
P(missing uppercase) = (69/95)¹² = 0.0216
P(missing digit) = (85/95)¹² = 0.2630
P(missing symbol) = (62/95)¹² = 0.0060
pairwise terms sum to 0.0074
union ≈ 0.3048 → compliant fraction ≈ 0.695
So the four-class rule costs 0.52 bits of theoretical keyspace at twelve characters. Run the same calculation at eight characters and the compliant fraction is about 0.456 — a cost of 1.13 bits.
Which is a small number, and that is the point: in pure keyspace terms, composition rules do almost nothing in either direction. They subtract half a bit. Every bit of their claimed benefit must come from moving user behaviour toward the uniform draw — and the record says behaviour moves the other way.
Where the bits actually go: mask attacks
Modern cracking tools don't enumerate a keyspace; they enumerate masks — position-by-position character-class templates — ordered by how often each appears in leaked corpora. The four-class rule is a gift to this technique: it tells the attacker which templates to try.
Take "minimum eight characters, one of each class." Naive keyspace 95⁸ ≈ 2⁵²·⁶. Now count the single most productive mask, ?u?l?l?l?l?l?d?s — capital first, lowercase body, digit, symbol last:
26 × 26⁵ × 10 × 33 = 1.02 × 10¹¹ ≈ 2³⁶·⁶
Sixteen bits below the headline, from one mask. And nobody's lowercase body is random. If the five-letter core comes from a list of the ten thousand commonest words, the digit is one of ten, and the symbol is one of the ten people actually use:
10⁴ × 10 × 10 = 10⁶ ≈ 2²⁰
Twenty bits. A password that satisfies every clause of a strict enterprise policy, sitting at guess number 10⁶ or below.
At the bcrypt-cost-12 rate above, 10⁶ guesses against one salted account takes 89 seconds on one rented rig; against Argon2id at 64 MiB, four minutes. The hashing decision — real, expensive, correctly made — bought about eight bits. The composition rule handed back sixteen.
This is the mechanism behind the opening story. Tightening the rule didn't move users into a bigger space; it moved them into a narrower, more predictable region of the one they already occupied, and it made the top mask more dominant, because a stricter rule admits fewer templates. A rule everyone must satisfy is a rule the attacker can assume you satisfied. Every constraint is one less degree of freedom in the search.
What real distributions look like
The distribution of chosen passwords is not uniform, not normal, and not lognormal. It is close to Zipf: frequency of the r-th most common password falls as p_r ∝ r^(−s), with published fits putting s somewhere around 0.7–0.9 for large consumer corpora.
That functional form has a useful consequence. When s < 1, cumulative coverage of the top N passwords goes as F(N) ∝ N^(1−s) — a power law, which on a log-x plot is a straight line, not a curve that flattens. The crack curve published in The Cost of Password Hashing has exactly this shape: roughly 4.5 percentage points per decade of guesses. Fit a Zipf to it and you get 1 − s = 0.26, so s ≈ 0.74 — comfortably inside the range the literature reports, which is a satisfying piece of internal consistency between an empirical crack curve and a theoretical distribution.
Extrapolating that fit down into the head — the region no crack curve bothers to publish, because it's assumed to be small — gives the numbers that should actually drive policy:
| Guesses per account (β) | Fraction of accounts recovered | At 1,000,000 accounts |
|---|---|---|
| 1 | 0.41% | 4,100 |
| 10 | 0.74% | 7,400 |
| 100 | 1.3% | 13,000 |
| 1,000 | 2.5% | 25,000 |
| 10⁴ | 4.5% | 45,000 |
| 10⁵ | 8.2% | 82,000 |
| 10⁶ | 15.0% | 150,000 |
The first row is the min-entropy argument in operational clothing. One guess per account, against a million accounts, yields about four thousand compromised accounts. That is password spraying. It defeats rate limiting, because one attempt per account per day looks like a typo; it defeats lockout, because nobody locks after one failure; it defeats your work factor, because 10⁶ total hashes is nothing. And it defeats every composition rule you have, because the passwords in that row satisfy them — that's how they became the most common ones under your policy.
H∞ for a corpus like this is about −log₂(0.0041) ≈ 7.9 bits. Put that next to the 78.8-bit figure on the policy slide and you have the honest measure of how far apart the two conversations are.
Two caveats, because the extrapolation is mine and not measured. A pure Zipf overshoots in the deep tail — the fit predicts 27% at 10⁷ guesses where the published curve says about 22% — so treat the head figures as order-of-magnitude, not three significant figures. And the exponent varies by population: a workforce corpus behind an enterprise policy typically has a thinner head than a consumer one, but a more concentrated mask distribution.
Measuring your own population, which is the only measurement that counts
Every number above is somebody else's corpus. The useful version is yours, and it is obtainable.
What you want is a guess-number curve: for each account, the guess index at which a realistic attacker's ordering would have found that password, plotted cumulatively. Not a per-password score. Strength meters collapse a position in an ordering into a scalar, and they usually do it with composition-based rules, which we've just established has the wrong sign. (Any meter rating P@ssw0rd1 above mycatlikestuna is reporting character classes, not guessability — a topic that deserves its own piece.)
You can't run this over your stored hashes at full scale: salting makes cost scale with guesses × accounts, so a 10⁶-deep run against 1.4 million bcrypt-12 rows is 1.4 × 10¹² hashes, four months on the rig above. So sample.
The design. Take a uniform random sample of 10,000 accounts. Run a real guesser against their hashes — hashcat with a mainstream rule set, or a PCFG or neural guesser for an ordering closer to optimal — recording the guess number at each hit, to depth 10⁶.
- Work: 10⁴ accounts × 10⁶ guesses = 10¹⁰ hashes.
- At bcrypt cost 12 on eight rented GPUs: about 10 days, or roughly $1,000 of rented time.
- Statistical power: at an observed rate near 5%, the standard error on 10,000 samples is 0.22 percentage points. That is far more precision than any policy decision needs.
A thousand dollars and a fortnight buys the four numbers — λ at 10, 10³, 10⁴, 10⁶ — that convert every password argument in your organization from aesthetics into arithmetic. Run it annually and after every policy change; the delta is the only evidence a change did anything. The team in the opening story got their 15.2%-to-11.9% figure this way, which is how they knew the change had underperformed its own claim by a factor of nearly twenty.
Two operational notes. Do it in an isolated environment with production-credential handling controls — you are, briefly, manufacturing plaintext. And record the guess number, not just the hit: the curve is the deliverable, a percentage is not.
Blocklists: the control with the best exchange rate, and where to set N
If the distribution is Zipfian with a heavy head, the highest-value intervention is obvious: delete the head.
A blocklist of the top N most common passwords does exactly that, and it has a property no composition rule has — its cost and its benefit are the same number. The fraction of users who will hit a rejection when choosing a password is F(N), because that's the fraction who would have chosen something in the list. And the fraction of accounts you remove from the cheapest part of the attack is also F(N).
| Blocklist size N | Accounts protected | Users who hit a rejection |
|---|---|---|
| 10³ | 2.5% | 2.5% |
| 10⁴ | 4.5% | 4.5% |
| 10⁵ | 8.2% | 8.2% |
| 10⁶ | 15.0% | 15.0% |
That identity makes the trade legible in a way almost nothing else in password policy is: you aren't guessing at friction, you're reading it off the same curve.
Where should N sit? The marginal return is dF/dN ∝ N^(−0.74) — diminishing, with no natural cliff — so the answer comes from friction tolerance rather than from the math. 10⁴ is defensible for a consumer population, 10⁵ for a workforce, where one extra rejection at setup is cheaper than a helpdesk call. Above 10⁶ the rejection rate crosses into "users start gaming the validator," which reintroduces the concentration problem you were trying to remove.
Two details that determine whether it works:
Normalize before comparing. A list that catches password and misses Password1! catches almost nothing, because your composition rule already pushed everyone to the second form. Case-fold, strip leading and trailing digits and symbols, and reverse common leetspeak substitutions before the lookup. This roughly triples the effective coverage of a fixed-size list, and it's the highest-leverage implementation detail here.
Tell the user the truth. "This is one of the ten thousand most common passwords, so it's in the list attackers try first" produces a different next choice than "password does not meet complexity requirements." The second is a lie that teaches nothing, and users answer it with the smallest edit that clears the validator — which is how you get Password1!.
A frequency blocklist is not the same control as a breach-corpus check, and the two are routinely conflated. The k-anonymity breach check queries a billion known-leaked strings, most appearing once or twice; it catches reuse, which is the stuffing threat. The frequency blocklist is ten thousand strings ordered by popularity; it catches the head of the guessing distribution, which is the spraying and cracking threat. Neither contains the other, and a password can pass one and fail the other in both directions. Run both; they cost approximately nothing.
Entropy and work factor multiply — which makes them substitutes at wildly different prices
Now the interaction, which is where the policy argument gets decided.
An offline attacker's total cost to reach an account is (guesses required) × (cost per guess). Take logs and it becomes additive:
log₂(attacker cost) = (effective bits of the password) + (bits of work factor) + log₂(cost of one fast hash)
One bcrypt cost increment halves attacker throughput, which is exactly one bit. One extra character on a generated lowercase password is 4.7 bits. These are interchangeable from the attacker's side and priced nothing alike from yours.
| Move | Bits gained | Your recurring cost |
|---|---|---|
| bcrypt cost 12 → 13 | 1.0 | Doubles the hashing pool: +$1,200/month at 46.7M verifications/month |
| bcrypt cost 12 → 14 | 2.0 | Quadruples it: +$3,600/month |
| +1 character on a generated password | 4.7 | One keystroke, once |
| +1 Diceware word | 12.9 | One word, once |
Buying 4.7 bits through the work factor means 2⁴·⁷ ≈ 26× your hashing bill, forever. Buying it through length costs a keystroke. At the cost figures worked out in the hashing piece, that's roughly $1,200/month becoming $31,000/month — for the same increment a generator gives you free.
Stated as a rule an engineer can carry into a room: the marginal bit is thousands of times cheaper on the credential side than on the hashing side — but only for generated credentials. For human-chosen ones, a length requirement buys perhaps one or two effective bits, not 4.7, and the substitution collapses. Which produces the actual conclusion of this whole analysis, and it is not "make passwords longer":
The high-value move is not lengthening human choices. It's replacing the human generator with a machine one — a password manager, a generated passphrase offered at set-time, a passkey. Everything else on the menu is buying single bits at increasing prices.
There is an important asymmetry in how these two controls act on the distribution, and it's the reason they aren't fully fungible. The work factor is a uniform multiplier: it shifts every account by the same number of bits, including the account whose password is 123456 — which is worth exactly nothing, because that account falls at guess one regardless. Entropy policy reshapes the distribution: blocklists remove the head, generators move mass into the tail. They multiply, but only one of them is uniform, and the non-uniform one is where the head lives. This is the same conclusion the hashing article reached from the cost side — the first dollar goes to the algorithm, the second to the accounts no work factor can save, and only the third to the work factor itself — arrived at independently from the distribution side.
The honest limits, which are large
Everything above is a defense against guessing. Here is what it does not touch.
Phishing. A 90-bit generated passphrase typed into a convincing fake login page is 100% compromised. Entropy is irrelevant; the user handed over the plaintext. There is no length, no alphabet, and no blocklist that changes this outcome by one percentage point.
Credential stuffing. The attacker isn't guessing at all — they're replaying a credential the user already used somewhere that got breached. The password's entropy is irrelevant because it isn't being searched for. Only uniqueness helps, which is a property of the manager, not of the policy.
Keyloggers, infostealer malware, and your own logs. The secret is captured at the keyboard, or in a reverse proxy with body logging enabled, at full entropy. This happens more often than anyone admits.
Which leads somewhere uncomfortable for an article of this length: the entropy debate consumes attention out of all proportion to its share of real compromises. Offline cracking of a stolen hash table is worth defending against properly, but it is not the threat that produces most incidents. Phishing and stuffing are, and both are structurally immune to every calculation above.
The correct posture, I think: get the entropy decisions right because they're cheap and because the wrong ones actively cost you helpdesk volume and user goodwill — then stop optimizing, and spend the remaining effort where the math above has no purchase. The strongest version of that argument is that the asset class itself is the problem, and the endgame is deleting it. The entropy math is how you manage a credential you still have, not a reason to keep having it.
The policy, in a form you can take to a committee
The recommendation the arithmetic supports, with the reason attached to each clause — because a policy without reasons is exactly what ossifies into the policy in the opening story.
- Minimum length 12 for humans, 8 permitted only with a passkey or hardware factor already enrolled. Length is the cheapest real bits available. Twelve is where a generated password clears the offline-cracking threshold behind a modern KDF with margin.
- No composition requirements at all. They cost 0.5–1.1 bits of keyspace, concentrate the realized distribution into a handful of masks worth 16+ bits to an attacker, and increase reset volume. Allow every character, including spaces and Unicode; require none.
- No maximum length below 64, and no character restrictions. A maximum length or a banned-character list tells an attacker how to prune. It also usually indicates the password is being stored somewhere it shouldn't be.
- Frequency blocklist of the top 10⁴ (consumer) or 10⁵ (workforce), applied after normalization — case-folded, edge digits and symbols stripped, leetspeak reversed. Friction equals protection, exactly, and you can read both off the same curve.
- Breach-corpus check at set-time and password change, k-anonymized, with a count threshold rather than a binary. Different control, different threat, near-zero recurring cost.
- Offer a generator inline, defaulted on. This is the single highest-leverage item in the list and the one most policies omit entirely. A five- or six-word generated passphrase, one click to accept, one click to reroll. Every bit of the length argument above is contingent on a machine doing the generating.
- Allow paste, always. Blocking paste breaks password managers, which are the mechanism by which items 1 and 6 actually happen.
- No periodic expiry. Covered at length elsewhere; rotate on evidence. If your framework binds you to it, there's a way to have that conversation.
- Measure λ_β annually on a 10,000-account sample to depth 10⁶, and publish the curve internally. Roughly $1,000 and two weeks. Without it, every subsequent argument about this policy is aesthetics.
- Set the work factor by the cost arithmetic, independently. It's a uniform multiplier on the whole distribution and it is not a substitute for any of the above — the two controls act on different parts of the curve and you want both.
The committee objection you should expect is that dropping composition rules "reduces security," and the answer is the mask calculation: the rule you are defending narrows the search space by sixteen bits and the keyspace argument for it is worth half a bit in the other direction. Bring your own λ_β curve if you have one. Bring the 89-second figure if you don't.
The line worth keeping
Three claims, if you take nothing else.
Entropy is a property of a generator, not of a string — so a policy that constrains what humans type without changing who generates it is measuring something that doesn't exist.
The attacker guesses in descending order of probability and stops when they have enough, so the head of your distribution is the whole game, and every metric based on an average is describing a search that will never be run.
And the cheapest bits are always length from a generator; the most expensive are work factor; and composition rules have a negative price — they cost you keyspace, they cost you helpdesk volume, and they cost you the sixteen bits of mask concentration that made the whole thing tractable in the first place.
The policy in the opening story claimed twenty-six bits and delivered fewer than one and a half. It was not badly implemented and nobody involved was careless. They applied a correct formula to a distribution it didn't describe, and then never measured the result — which is a mistake available to anyone, in any domain, and the only reliable defense against it is the boring one: run the attack against your own data before you believe your own slide.