The Password Policy Arms Race: Did Any of It Work?
In 1979, Robert Morris and Ken Thompson published a short paper in Communications of the ACM describing what they had found when they collected a few thousand passwords from the UNIX systems they had access to and simply tried to guess them.
The paper is worth reading in full, but the number everyone remembers is this: of roughly 3,300 passwords, about 86% fell to a search that a graduate student could have run overnight. Not a clever search. Single characters. Pairs. Short strings of lowercase letters. The system dictionary. The list of names. Words spelled backwards. The whole exercise was, in modern terms, a wordlist attack with no rules engine and no GPU, run on a PDP-11.
That paper also contains the first serious engineering response to the problem, and it is remarkable how much of it we still run. They salted — twelve bits, 4,096 variants, explicitly to defeat precomputation and to stop identical passwords from looking identical across accounts. They deliberately slowed the hash down, iterating a modified DES 25 times and mangling it enough that off-the-shelf DES hardware couldn't be pointed at it. Salting and a tunable work factor, in the same paper, in 1979.
And then they wrote the sentence that has aged the worst, which is that the remaining defence would have to come from making users choose better passwords.
Forty-seven years later, that 86% figure is arguably the most reproducible result in security. Every large corpus published since — RockYou, LinkedIn, Yahoo, the aggregate dumps — has confirmed it in one form or another. The distribution of human-chosen secrets is stubbornly, boringly the same. What changed enormously is the tooling on both sides: forty years of defenders writing rules and attackers writing crackers, each move calling out the next.
So it is fair to ask the question nobody asks cleanly: did any of it work?
The honest answer is that some of it worked decisively, some of it did nothing, some of it was actively counterproductive, and — the interesting part — the field made almost all of those decisions with essentially no evidence, and largely still can't tell you which was which. That last fact is the one with transferable lessons, so it gets the longest section.
I've argued the case against forced rotation at length elsewhere and won't relitigate it here; it appears in the scorecard with a citation and nothing more. Same for breach-corpus checking mechanics, work-factor economics and lockout. This piece is the history and the verdict.
The first two decades, in which length rules did nothing
Here is a fact that most engineers who have written a password policy do not know, and it retroactively invalidates a great deal of 1980s and 1990s security advice.
Traditional UNIX crypt(3) took the first eight characters of your password and discarded the rest. Seven bits from each of eight characters is 56 bits, which is exactly a DES key, because the construction used the password as the key. If your site's policy said "minimum twelve characters," and your users dutifully typed twelve, the ninth through twelfth were thrown away before anything was hashed. The policy was enforced by the password-changing program. The algorithm underneath it had no idea the policy existed.
Windows was worse in a more baroque way. The LAN Manager hash upcased the password, padded or truncated it to fourteen characters, split it into two seven-character halves, and hashed each half independently, without a salt. A fourteen-character password stored as an LM hash is not a fourteen-character password. It's two seven-character passwords with no salt, and — the detail that makes it truly bad — a password of eight characters produces a second half of one character plus padding, which is trivially recognisable, telling the attacker the length before they start.
So for something like two decades, across the two dominant platform families, the single most commonly imposed password rule was measuring a quantity the storage layer did not preserve.
The same class of bug is still live. bcrypt truncates at 72 bytes — generous, but real if you feed it a passphrase or a pre-hashed value. A 2011 flaw in a widely-deployed bcrypt implementation mishandled certain high-bit characters, and the fix shipped a new hash prefix purely so systems could tell which passwords had been hashed by the broken version.
The lesson that generalises: a policy is an assertion about a value; the storage layer is what decides which parts of that value survive. If you have never checked what your hash construction does with the tail of a long input, you do not actually know what your minimum-length rule enforces.
The attack tools arrive
The defence had been published before the attack was automated, which is unusual and worth noticing.
The Morris worm in November 1988 is the moment the theoretical became operational at scale. Among its propagation methods was a password guesser carrying an internal list of a few hundred likely words, plus the system dictionary, plus obvious transformations of the account name — and, crucially, its own re-implementation of crypt, substantially faster than the library one. That is the arms race in a single artifact: the attacker optimised the defender's work factor away.
Then the tools became products. Alec Muffett's Crack circulated from the start of the 1990s and did the thing that mattered: it wasn't a dictionary, it was a rules language, letting you express "take the word, capitalise it, append a digit, substitute 3 for e" as a compact program applied across a whole wordlist. John the Ripper (Solar Designer, 1996) generalised it; L0phtCrack (1997) made it a commercial Windows product aimed squarely at those unsalted LM hashes.
Note what all of these mechanised. Not brute force. Human password-generation habits, expressed as transformations. From 1990 onward, the attacker's model was never "try all strings"; it was "try the strings people actually produce, in descending order of likelihood." Every subsequent defensive rule was proposed against a mental model of brute force that had already been obsolete for a decade.
The defence that won before the attack got famous
Rainbow tables are the best-known password attack in popular culture and they are a strange case, because the defence against them had been shipping for twenty-four years by the time they were named.
The underlying idea is Hellman's 1980 time–memory trade-off. Precompute chains covering the keyspace, store endpoints, and trade storage against search time. Philippe Oechslin's 2003 paper refined the chain structure and gave us the name. For a while, in the mid-2000s, "rainbow table" was security's scare phrase.
Against what, exactly? Against unsalted hashes. A precomputed table maps hash → plaintext, and a salt makes each account's hash a different function, so a single table covers one salt value and you'd need 4,096 of them for 1979-era UNIX and 2^128 for a modern 16-byte salt. Rainbow tables were never a threat to crypt(3). They were devastating against LM and NTLM, against unsalted MD5 and SHA-1 application databases, and against exactly the systems whose designers had skipped the paper from 1979.
Two things fall out of this that are worth carrying.
First, salting is the clearest defensive win in the entire history of the field, and it won prospectively. Morris and Thompson did not add salt because they had observed a precomputation attack. They added it because they reasoned about what an attacker with storage would do. The defence sat in the standard library for a generation before the attack it prevented became a household word.
Second, and less comfortably: rainbow tables have been obsolete for most purposes since roughly 2010, and not because we fixed anything. GPUs got so fast at unsalted hashes that recomputing is cheaper than the storage and lookups a table requires. An attack disappeared because a different attack got better. If your mental model has "rainbow tables" as a current threat, you are two revisions behind in both directions.
The composition-rule era, and the appendix that ate twenty years
The rules everyone hates — one uppercase, one lowercase, one digit, one symbol, no dictionary word, no reuse of the last 24 — came from somewhere specific, and the provenance is genuinely instructive.
NIST published SP 800-63 in 2004. Buried in it was Appendix A, which offered a way to estimate password entropy in bits from a password's length and composition: so many bits for the first character, fewer for the next several, fewer still after that, plus a bonus for using multiple character classes, plus a bonus for surviving a dictionary check. It was a heuristic scoring function, and it was reasonable-looking enough that frameworks, auditors, product managers and directory defaults absorbed it wholesale. If you have ever been told a password needs "at least 30 bits of entropy," you are downstream of that appendix.
Its evidence base was thin, and this is not hindsight bias — it says so. The estimates were adapted in part from Shannon's 1950s work on the entropy of English text, plus stated rules of thumb. There was no corpus of real passwords behind it, because in 2004 no such corpus was lawfully available to anyone. The author, Bill Burr, said publicly years later that he regretted much of it.
The empirical demolition arrived once corpora did exist. Weir and colleagues showed in 2010 that the appendix's entropy score was a poor predictor of how many guesses a real cracker actually needs — you could hold "entropy" constant and vary guess-resistance by orders of magnitude. Carnegie Mellon's large user studies around 2011 found that a policy of sixteen characters with no other requirements produced passwords that resisted guessing better than eight characters with the full four-class comprehensive requirement, while being rated less annoying by the people who had to create them. Better security and lower friction, from deleting rules.
NIST reversed course in SP 800-63B in 2017: no composition rules, no arbitrary periodic expiry, check candidates against lists of compromised values, allow long secrets, allow paste. Revision 4, finalised in 2025, states it in stronger normative language. Thirteen years of enforcement, one appendix, roughly zero empirical support at the point of adoption.
The GPU inflection
Around 2009–2010, general-purpose GPU computing arrived in password cracking and every parameter in every policy became wrong simultaneously.
The exact multipliers depend on hash and hardware and I'd encourage you to benchmark rather than trust anyone's table, including this one. But the shape is not in dispute:
| Era | Hardware | Fast unsalted hash (MD5/NTLM), order of magnitude |
|---|---|---|
| ~1979 | PDP-11 | tens of crypt(3) per second |
| ~2000 | Single desktop CPU | ~10^5–10^6 guesses/s |
| ~2010 | Single high-end GPU | ~10^9 guesses/s |
| ~2025 | Single high-end GPU | ~10^11 guesses/s |
| ~2025 | Rented multi-GPU node, hourly | ~10^12 guesses/s |
That is roughly four to five orders of magnitude in fifteen years, on a curve that no policy document tracked. A password policy calibrated in 2004 to require "enough entropy to resist offline attack" was, by 2013, requiring about one hundred-thousandth of what it thought it was requiring — and nobody re-derived it, because the policy's number was a character-class rule, not a cost model. The rule had no parameter to update.
This is the strongest structural criticism of the whole composition-rule project. It expressed a defence in units that could not be re-tuned when the attacker's economics changed. Work factors have that parameter; a cost of 12 can become a cost of 14 and you can compute what you bought. "One uppercase and one digit" has no dial. When the ground moves, a rule like that doesn't get worse gracefully — it just becomes irrelevant while continuing to be enforced.
Mask attacks: how a defensive rule became an attacker's input
This is the mechanism by which the composition-rule era ended, and it is worth being precise about it, because "complexity rules don't help" is usually asserted rather than explained.
A modern cracker doesn't only take wordlists. It takes masks — position-by-position templates over character classes, ordered by how often each pattern appears in leaked corpora. ?u?l?l?l?l?l?d?d?s means one uppercase, five lowercase, two digits, one symbol, in that order, and its search space is a tiny fraction of the naive keyspace for nine characters.
Where does that ordering come from? From your policy, combined with the entirely predictable way humans satisfy it. Told to include an uppercase letter, people capitalise the first character. Told to include a digit, they append one or two at the end — usually a year or a 1. Told to include a symbol, they append ! after the digits. These aren't stereotypes; they're distributions you can read off any corpus, which is why hashcat ships rule sets that encode them.
So the composition requirement does two things at once. It removes some weak candidates — password no longer validates. And it removes a far larger number of candidates from the attacker's search, because every non-conforming string is now one the attacker can skip. A rule everyone must satisfy is a rule the attacker can assume you satisfied.
I've worked the arithmetic — how few bits a four-class rule actually costs in keyspace, and how many bits the top mask hands back — in the entropy piece; the short version is that the exchange rate is dreadful, and it gets worse as the rule gets stricter, because a stricter rule admits fewer templates. P@ssw0rd1 satisfies almost every enterprise policy ever written and dies in the first second of a rules-based run. mycatlikestuna fails most of them and survives considerably longer.
That's why composition rules score plausibly negative rather than merely useless — a claim I'll hedge properly in the scorecard, because the counterfactual is unmeasurable.
The corpus era: cracking stops being a modelling problem
December 2009. RockYou, a maker of social-network widgets, is breached through SQL injection. Roughly 32 million passwords are exposed — and they were stored in plaintext, so there is nothing to crack. The list is published. Its deduplicated form, rockyou.txt, is still on every penetration tester's laptop today, sixteen years later, and it still works.
I want to be precise about why this was the single most consequential event in the history of password attacks, more so than GPUs.
Before RockYou, building a good cracking attack was a modelling exercise. You reasoned about what people probably do, wrote rules expressing it, and validated against whatever small samples you could get. After RockYou — and then LinkedIn in 2012, Adobe's 153 million records in 2013, Ashley Madison in 2015, and the giant aggregated collections from 2017 onward — it became a lookup exercise. You no longer need a model of human password choice. You have the actual empirical distribution, tens of millions of samples deep, with frequencies attached.
Three consequences follow, and all three are still shaping current guidance.
The top of the distribution is brutally concentrated. A very small number of passwords cover a startling share of any population. That's why a spray of the ten most common passwords across ten thousand accounts is more economical than a deep attack on any one of them, and why lockout thresholds never fire against the attack that actually works.
Blocklists became strictly better than rules. A composition rule is a proxy for "is this password likely to be guessed." A breach-corpus check is a direct measurement of it — this exact string has been seen, this many times, in real dumps. Once you can ask the direct question, keeping the proxy is hard to justify. This is the single clearest example of evidence replacing folklore in the field, and it took roughly seven years from RockYou to NIST codifying it.
And the empirical era of password research began with a crime. This deserves stating plainly, because it explains the previous thirty years. Until 2009, no researcher could lawfully obtain a large corpus of real human-chosen passwords. Studies were run on hundreds of users, or on synthetic sets, or on whatever a sympathetic sysadmin would share. The reason 2004's Appendix A leaned on Shannon's analysis of English prose is that there was nothing better to lean on. The field's evidence base arrived as stolen property, and the guidance improved sharply and immediately once it did.
There is a real methodological problem hiding in that gift, which I'll come back to.
Credential stuffing makes per-site strength almost irrelevant
The corpora didn't only improve cracking. They created an entirely different attack that doesn't involve cracking at all.
If you hold a hundred million email:password pairs from other people's breaches, the highest-return activity is not attacking anyone's hashes. It's replaying those pairs against every other service, because reuse is the norm. Google and UC San Diego's 2017 study of the criminal ecosystem found a meaningful share of leaked credentials matched victims' current Google passwords; Google's Password Checkup telemetry in 2019 found roughly 1.5% of observed logins used a credential known to be breached. Each attempt costs a fraction of a cent and produces exactly one failed login per account — invisible to every per-account threshold ever configured.
Now hold that next to a password policy. Your policy governs the strength of the secret on your site. Credential stuffing does not care about the strength of the secret. It cares whether the user used it somewhere else that got breached. A sixteen-character, four-class, rotated-quarterly password that the user also used on a hobby forum in 2016 is compromised, and no rule in your policy has any bearing on it.
This is the moment the entire premise of per-site password policy quietly collapsed. From roughly 2015 onward, the dominant determinant of whether an account gets taken over is a property of the user's behaviour across other services, not a property of your policy. The controls that matter became reuse detection (blocklists, corpus checks), MFA, and anomaly detection — none of which are password policies.
Forty years, in one picture
flowchart TB
D1["1979 · Salt + iterated hash<br/>Morris & Thompson"] --> A1["1988-1997 · Dictionary + rules<br/>Morris worm, Crack, John, L0phtCrack"]
A1 --> D2["1990s-2000s · Composition rules,<br/>min length, expiry, lockout"]
D2 --> A2["2003 · Rainbow tables<br/>(already defeated by 1979's salt)"]
A2 --> D3["1999-2015 · Real work factors<br/>bcrypt, PBKDF2, scrypt, Argon2"]
D3 --> A3["2009+ · GPU cracking<br/>~10⁴-10⁵× throughput jump"]
A3 --> A4["2009+ · Mask + rule attacks<br/>composition rules shrink the search"]
A4 --> A5["2009+ · Leaked corpora<br/>modelling becomes lookup"]
A5 --> A6["2015+ · Credential stuffing<br/>reuse beats strength"]
A6 --> D4["2017+ · Blocklists, rate limiting,<br/>MFA, passwordless"]
style D1 fill:#e8f4ea,stroke:#4a7
style D2 fill:#e8f4ea,stroke:#4a7
style D3 fill:#e8f4ea,stroke:#4a7
style D4 fill:#e8f4ea,stroke:#4a7
Read the right-hand column of that chain as a single claim: every defence that survived contact with the next forty years was a property of the system, and every defence that failed was a property demanded of the user. Salt, work factor, blocklist, rate limit, second factor, keypair — all system-side. Composition, length-you-must-remember, rotation, hints, security questions — all user-side. That's not a coincidence and it isn't moralising about users; it's that system-side controls have parameters that can be re-tuned when the economics move, and user-side controls have only compliance behaviour, which optimises toward the minimum that satisfies the check.
The scorecard
Before the table, the caveat that makes it honest: nobody ran a controlled trial. There is no world in which half the enterprises adopted composition rules and half didn't, with compromise rates measured over a decade and confounders controlled. What we have instead is a mixture of laboratory studies with real users but artificial stakes, corpus analyses of the populations that happened to leak, attacker-economics arithmetic, and telemetry from a handful of very large providers who don't publish their methodology in full. Every verdict below is a judgement over that evidence, not a measurement. The "evidence" column says how much weight the verdict can bear.
| Intervention | Era | Verdict | Why | Evidence quality |
|---|---|---|---|---|
| Salting | 1979– | Worked, decisively | Killed precomputation outright; the single highest-return line of defensive code ever written in this field. Deployed before the attack it prevents became widespread. | Strong — the mechanism is arithmetic, not behavioural |
| Work factors (iterated/memory-hard KDFs) | 1979, seriously 1999– | Worked, with caveats | Directly multiplies attacker cost. Caveats: it's a parameter with a maturity date, it's paid on every login, and the marginal increment buys less than teams assume. Irrelevant against phishing and stuffing. | Strong for mechanism; weak on how much real-world compromise it prevented |
| Breach-corpus blocklists | 2013– | Worked | Replaces a proxy for guessability with a direct measurement of it. Removes the concentrated head of the distribution, which is where the economical attacks live. Cheap, and it emits telemetry. | Good — measurable rejection rates, direct mechanism against stuffing and spraying |
| MFA | 2000s, mass adoption 2018– | Worked | The largest single reduction in account takeover any provider has reported. Google's 2019 figures put on-device prompts at blocking essentially all automated and bulk-phishing attacks; Microsoft has repeatedly cited >99.9% reduction. Real erosion since, from real-time phishing proxies and push fatigue. | Best available in the field, though vendor-published and not independently replicated |
| Minimum length | 1980s– | Worked modestly | Real once the storage stopped truncating, and the CMU work suggests long-and-unconstrained beats short-and-complex. But the empirical mode of password length in every corpus sits exactly at the enforced minimum, so a minimum sets a floor and simultaneously a ceiling. | Moderate — lab studies plus corpus distributions |
| Rate limiting / throttling | 2000s– | Worked | Converts an unbounded online attack into a bounded one. The one control that touches spraying and stuffing without touching the user. Under-built almost everywhere, because lockout ticked the box. | Moderate; mechanism clear, deployment quality wildly variable |
| Composition rules | ~1990–2017 | Did not work; plausibly negative | Users satisfy them with a small set of predictable patterns, so the policy-conforming set is smaller and better-modelled than the unconstrained one. Feeds mask and rule attacks directly. Costs measurable user friction. | Good against it: NIST reversed on the evidence; the mask mechanism is demonstrable |
| Forced periodic rotation | ~1990–2017 | Did not work | Argued in full here; users transform rather than replace, and the attacker uses the credential in hours, not months. | Good — including a rare longitudinal dataset of per-user password sequences |
| Password hints | 2000s | Actively harmful | Adobe's 2013 breach leaked 153 million records with the hints in plaintext alongside them; because the encryption was deterministic, hints attached to other users sharing the same password could be pooled. A crossword puzzle, assembled by the victim, published by the attacker. | Strong — one enormous natural experiment |
| Security questions | 1990s– | Actively harmful | Google's 2015 analysis found single-guess attack success in the double digits of percent for common questions, while a large fraction of legitimate users couldn't recall their own answers. Harder for the user than for the attacker is the precise inversion of what a credential should be. | Strong — large-scale study on real recovery attempts |
| Account lockout | 1990s– | Did not work; enables a DoS | Blind to spraying and stuffing, and it hands anyone who knows a username a free off-switch. | Good — the mechanism is trivially demonstrable |
| Composition-based strength meters | 2000s– | Mostly theatre | A meter scoring character classes rates P@ssw0rd1 strong and mycatlikestuna weak; both judgements are backwards against real tooling. Guessability-based estimators (zxcvbn and successors) are a genuine improvement, and stringent meters do nudge users longer — so this is theatre with a small real signal inside it. |
Moderate — comparative evaluations plus lab studies on nudging |
The pattern in the verdict column is the article's thesis in compressed form. Everything that worked changed the attacker's cost function or removed the secret from the human. Everything that failed tried to change what the human chose, without ever measuring whether it did.
What "worked" can even mean here
I want to spend a section on why that table is harder to construct honestly than it looks, because the epistemics are the transferable part.
There is no counterfactual. To say composition rules were net negative, you need to know what would have happened without them, on the same population, facing the same attackers. Nobody has that. Corpus comparisons between sites with different policies are confounded by everything: the sites differ in user demographics, in the value of the account, in what else they were doing right.
Survivorship runs in a nasty direction. Everything we know empirically about human password choice comes from corpora that leaked — which over-represent organisations with poor practice. The very best datasets, the plaintext ones like RockYou, come specifically from organisations that weren't hashing at all, which correlates with every other kind of neglect. We calibrated a generation of guidance on a sample selected for institutional carelessness. It's the best sample we have and it is not a random one.
Absence of incident is not evidence of efficacy. The classic form: "we've had a 90-day rotation policy for twelve years and never had a credential breach." That statement is compatible with the policy working, with the policy being irrelevant, and with breaches having occurred and gone undetected — which, given typical dwell times, is not a remote possibility. Controls that produce no telemetry cannot be evaluated by the absence of bad outcomes.
And nobody measured the cost side either. This is the part that should embarrass the field most. Password policy is one of the very few security controls whose price is paid in a currency you can measure trivially — helpdesk tickets, reset volume, minutes of lost work, failed-login rates, abandoned signups. Almost no organisation has a before-and-after on any of it. I've watched teams debate a length requirement for weeks without anyone pulling the reset-ticket time series they already had. If you're going to impose an unmeasurable benefit, the least you can do is measure the measurable cost.
Why a technical field made forty years of decisions on almost no evidence
Five reasons, and every one of them recurs in domains that have nothing to do with passwords.
The data was unobtainable until it was stolen. You could not study password choice at scale without either committing a crime or being one of about five companies. So the field substituted the nearest available proxy — an entropy estimate borrowed from English prose — and then, critically, forgot it was a proxy. Appendix A entered the literature as an estimate and left it as a requirement. Watch for this specific transition: a heuristic that acquires a number, and the number acquires authority the heuristic never had.
Nothing about a password policy emits a signal. A rate limiter tells you how many requests it dropped. A blocklist tells you how many candidate passwords it rejected and which ones. A composition rule tells you nothing at all — it just silently rejects and the user retries. A control that cannot report on its own operation cannot be evaluated, and will therefore survive on narrative alone. If you take one operational principle from this history, take that one: prefer controls that produce evidence of their own effect, and be suspicious of any control that has run for years without generating a single number.
The costs and the authority sat in different budgets. The person who set the policy did not answer for helpdesk volume, and the helpdesk did not have authority over the policy. A control whose benefits accrue to the decision-maker and whose costs land on someone else's cost centre faces no natural pressure to justify itself. This is an organisational design problem wearing a security costume.
Checklists are how the field distributes knowledge, and copying is cheaper than thinking. I've written about how this ossifies in compliance frameworks; briefly, guidance propagates by transcription, each hop strips the reasoning and keeps the number, and the number arrives at your desk as a requirement with no provenance and no attached assumptions.
The incentives are asymmetric and always will be. Nobody has ever been fired for requiring password complexity. The failure mode of over-requiring is diffuse annoyance across thousands of people; the failure mode of under-requiring is being the person who removed a control before a breach. That asymmetry is stable no matter what the research says, which is why "the evidence shows" persuades less than engineers expect in the rooms where these decisions get made.
What an engineer should actually do today
Short, because the argument above narrows it considerably.
- Set a real minimum length — 12 or more for humans — allow at least 64 characters, allow every Unicode character, and delete every composition rule. They cost you friction and hand the attacker a mask.
- Check candidates against a breach corpus at set time, and again at login. Login is the only moment you ever hold the plaintext of a password chosen years ago. Do it without disclosing the password.
- Verify what your hash construction does with long inputs, and don't pre-hash into bcrypt without understanding the 72-byte limit. Know what your length rule actually enforces.
- Use a memory-hard KDF at a work factor you have deliberately priced, and put a calendar reminder on re-deriving it, because the parameter has a maturity date.
- Build layered rate limiting with cross-account correlation, and treat lockout as a narrowly-scoped exception rather than a default.
- Rotate on evidence, never on a calendar — and build the mass-reset capability, which is the half most organisations skipped.
- Delete security questions and password hints. Not soften — delete.
- Push MFA toward phishing-resistant factors, and treat every remaining password as a scheduled deletion rather than a permanent asset to be hardened forever.
- Instrument the cost. Reset tickets, failed-login rate, password-set abandonment, time-to-first-successful-login. If you can't show the cost of a control moving when you change it, you have no basis for claiming the control is cheap.
The line worth keeping
Forty years of password policy produced four genuine wins — salt, work factors, blocklists, and a second factor — and a large body of rules that made users unhappy, made attackers' search spaces smaller, and were never evaluated by anyone.
The uncomfortable observation is that the wins were all available early. Salting and work factors were in the 1979 paper. Blocklists required only a corpus, and became obvious within a few years of one existing. What consumed the intervening decades was an enormous, sincere, well-intentioned effort to solve the problem at the layer where it could not be solved: the human being's choice of a memorable string.
Morris and Thompson can be forgiven for suggesting that layer, since it was the only one they hadn't already fixed. The field's error was spending forty years there without ever checking whether it was working — and the reason it never checked is that it had built controls incapable of telling it.
That is the lesson worth carrying out of the password era and into whatever the next one demands of us. Security folklore survives not because it's persuasive, but because the controls it recommends are silent. Ask of any control you're about to adopt: what number will this produce next quarter that tells me whether it's doing anything? If the answer is none, you are not adopting a control. You are adopting a belief.