Password Strength Meters Mostly Lie

After a credential-stuffing incident that turned into a small forced-reset exercise, a team I was working with did the sensible post-mortem thing: they took the passwords that had actually fallen, and replayed them through their own signup form to see what the strength meter had said about them.

All of them scored strong. Every single one.

There was a moment of stunned silence in the review before someone worked out why, and the reason is so simple it's almost a tautology. The meter was the gate. Nothing got into the database without clearing it. Of course every compromised password had scored strong — a password that scored weak was never allowed to exist. The meter had not failed at ranking these passwords; it had never been asked to. It had been asked "does this string contain an uppercase letter, a digit and a symbol," it had answered correctly, and the answer had no relationship to the thing anyone in that room cared about.

That's the failure worth understanding, and it isn't really about meters. It's about what quantity a meter is supposed to be estimating, and the fact that the widely deployed ones are estimating a different quantity that happens to be easy to compute.

A strength meter is a predictor. There is exactly one thing worth predicting: the number of guesses an adversary must make before this password falls, given the order that adversary actually guesses in. A character-class meter predicts a surface property of the string instead — which classes appear — and that property correlates with guessability only under the assumption that the attacker guesses uniformly at random. No attacker in the history of the field has ever satisfied that assumption.

This piece is about how to estimate the right quantity. Not why composition rules are bad — Longer Beats Complex does the entropy math and the naive-keyspace critique, and The Password Policy Arms Race does the forty-year policy history. Take both as read. This is the narrower engineering problem: you have a text input, a candidate string, a hundred milliseconds, and you want to output a number that means something. How do you compute it, what is it a number about, and where is it systematically wrong?


The quantity: guess number, and why it is adversary-relative

Start with the attacker, because the estimate is defined in terms of them and in no other terms.

An attacker with a hash and a guessing budget produces a sequence of candidate passwords g₁, g₂, g₃, … and tries them in order. That order is not arbitrary and it is not random: a rational attacker enumerates in descending order of estimated probability, because that maximises accounts cracked per unit of compute. Every serious cracking pipeline — a hashcat wordlist plus rule set, a PCFG-generated stream, a Markov chain walked in probability order, a neural model sampled by likelihood — is a machine for producing that ordering.

The quantity you want, for a password p, is its guess number: the index i such that gᵢ = p. How deep into the attacker's list does this password sit? If the answer is 300, the account is gone in the first second of any attack. If the answer is 10¹⁴, it survives an offline attack against a well-parameterised KDF for a long time.

Three consequences follow immediately, and they're where most meter design goes wrong.

Guess number is a property of the pair (password, attacker), not of the password. There is no such thing as the strength of mycatlikestuna in the abstract. There is its position in some enumeration. Change the enumeration — add a wordlist of song lyrics, add the user's pet's name from a social profile — and the number changes by orders of magnitude without the string changing at all. Any meter that reports a single scalar is implicitly asserting an attacker model. It should say which one.

The relevant thresholds differ by four to six orders of magnitude depending on which attacker you're defending against. A password facing only an online, rate-limited login endpoint needs to survive maybe 10³–10⁶ guesses, because that's all the attacker gets before your throttling, lockout policy and detection make the attempt uneconomic. A password whose hash may end up in a dump needs to survive whatever your KDF's cost lets a motivated adversary buy — a number you can actually compute, and which The Cost of Password Hashing walks through in detail. These are not the same instrument. A meter tuned for the online case, deployed as if it covered the offline case, is a thermometer being read as a barometer.

And the estimate is only ever an upper bound on your own model's cost, which makes it a lower bound on nothing. If your estimator says 10⁹ guesses, the honest reading is "at least this many, for an attacker whose model is no better than mine." An attacker with a better model — a bigger corpus, better rules, your specific user population's habits — gets a smaller number. Never the other way round. That asymmetry runs through everything below, and I'll come back to it, because it's the part experienced engineers most often forget when they wire a threshold into a validator.


Two passwords, and what tooling actually does to them

The canonical demonstration. P@ssw0rd1 clears almost every composition rule ever written: upper, lower, digit, symbol, nine characters. mycatlikestuna fails most of them: no uppercase, no digit, no symbol.

Now score them the way an attacker's pipeline does.

P@ssw0rd1 is not a novel string. It is password — rank 1 or 2 in essentially every leaked-credential frequency list ever compiled — with three transformations applied: capitalise the first letter, substitute a→@ and o→0, append a digit. Each of those transformations is a single rule in the standard mangling sets that ship with John the Ripper and hashcat. Capitalisation is c. Leet substitution is sa@ and so0. Appending a digit is $1. Any of the widely used rule files contains all of them and many thousands of combinations besides. So the cost of reaching P@ssw0rd1 is roughly: rank of password in the wordlist, multiplied by the position of that rule combination in the rule file. Both are small. The candidate lands well within the first hundred thousand guesses of a wholly unremarkable attack, and plausibly within the first few thousand.

The composition rules did not make password harder to guess. They specified, quite precisely, which four rules to apply to it — and because every user subject to the same rules performs the same repertoire of minimal edits, they concentrated the population into a region the attacker can enumerate cheaply. That's the mask-attack argument, and Longer Beats Complex has the arithmetic.

mycatlikestuna is a different shape of problem. There's no dictionary that contains it, so a wordlist-plus-rules attack does not reach it at all; rules mangle a word, they don't concatenate four of them. To generate it you need a combinator attack — words drawn from a list, concatenated without separators, in order — and the cost is multiplicative in the number of segments. Four segments from even a modest frequency-weighted list of a few thousand common English words, allowing for the attacker not knowing the segmentation in advance, puts this many orders of magnitude beyond the reach of the attack that kills P@ssw0rd1.

Two honest caveats, because this is where the rant version of this article overclaims.

It is not unbreakable. Frequency-weighted word combinators are a solved technique, my, cat, likes and tuna are all high-frequency, and the phrase is grammatical English, which is itself a massive constraint on the search — a language model prices grammatical four-word sequences far below the product of independent word frequencies. Against an attacker running a decent neural or PCFG guesser with a large offline budget, this passphrase is not a great password. It is merely a far better one than P@ssw0rd1, by a margin large enough that any meter ranking them the other way round is not slightly miscalibrated; it is inverted.

And the second caveat, which is the one that actually determines what you build: the reason a composition meter gets this backwards is not that its formula has a bug. It's that it is computing a function of the character classes present, and guess number is a function of the string's relationship to a corpus and a set of transformations. Those are different mathematical objects. You cannot fix the first by tuning coefficients.


How real estimation works

Here is the mechanism, at the level of detail you'd need to implement or evaluate one. The reference implementation to read is Dan Wheeler's zxcvbn (open-sourced at Dropbox around 2012, written up at USENIX Security in 2016), and the design is worth understanding even if you never ship that particular library, because every credible estimator has the same three phases.

Phase 1: matching

Run a battery of matchers over every substring of the candidate. Each matcher recognises a pattern that an attacker's pipeline can generate cheaply, and emits a match record: (start, end, pattern type, and the metadata needed to price it).

The matcher set that matters in practice:

  • Ranked dictionary. Multiple lists — leaked-password frequency lists, English words by corpus frequency, given names and surnames, television and film vocabulary, place names. The important part is ranked: the match carries the word's rank in its list, because rank is the price. love at rank 200 costs 200 guesses; an obscure word at rank 90,000 costs 90,000.
  • Leet-decoded dictionary. Reverse the substitutions (4→a, @→a, 0→o, 1→l or i, $→s, 3→e) and re-run the dictionary match on the decoded string.
  • Reversed dictionary. drowssap costs the rank of password plus a small constant.
  • Spatial / keyboard walk. qwerty, 1qaz2wsx, zaq!2wsx — adjacency graphs for QWERTY, Dvorak and the numeric keypad, matching runs of physically adjacent keys.
  • Repeats. abcabcabc is one match: base pattern plus a repeat count.
  • Sequences. abcdef, 9876, arithmetic runs in any character range.
  • Dates. All the ways a human writes one: 1/1/2011, 11-11-11, 19871987, bare four-digit years.
  • Regex-ish leftovers. Recent years, all-digit strings of a given length.
  • User-supplied context. More on this below; it's the single biggest omission in deployed meters.

A ten-character password typically produces dozens of overlapping matches. password123 yields dictionary matches on password (positions 0–7), pass (0–3) and word (4–7), a sequence match on 123, a date-ish match on the same span, and more.

Phase 2: pricing each match

Convert each match to an estimated guess count — how many candidates an attacker enumerates before producing this segment in this form.

For a dictionary match, start with the rank, then multiply by the cost of the transformations that were applied, because the attacker has to try those too:

  • Capitalisation. All-lowercase costs nothing — it's the base form. First-letter-capitalised, ALL CAPS, or last-letter-capitalised are the three overwhelmingly common variants, so a small constant factor (the reference implementation uses 2). Anything else — internal capitals in an unusual arrangement — costs the number of ways to choose which letters are upper, i.e. a sum of binomial coefficients, which grows fast and is correctly rewarded.
  • Leet substitution. Count the number of subsets of the applicable substitutions the attacker would have to enumerate. If four characters could plausibly have been substituted, the factor is the number of subset combinations, not 2⁴ blindly — you only count the substitutions that actually appear.
  • Reversal, repetition count, date format ambiguity: each contributes its own small multiplier.

For a keyboard walk, the price is a function of the starting key, the walk length, the number of turns (a straight run down a row is much cheaper to enumerate than a zigzag) and how many keys were shifted. For a sequence, it's the base (digits vs letters vs full ASCII), the length, and whether it ascends or descends. For an unmatched run of characters, you fall back to brute force over the observed character set, which is the only place the naive keyspace formula legitimately appears — and note that it appears as the most expensive option, the price you pay for a segment no pattern explains.

Phase 3: the minimum-cost parse, which is the whole trick

This is the part people skip, and skipping it is what separates an estimator from a rule table.

A password is a concatenation of segments. An attacker generating password123 does not generate it as an atom; they generate it as a structure — word, then digits — where each slot is filled from a source. The total cost is roughly the product of the per-segment costs, times a term for the structure itself (the attacker must also enumerate over possible structures; the reference implementation charges a factorial-ish penalty in the number of segments plus a constant per extra segment, which is a crude but defensible stand-in for the PCFG structure probability).

Crucially, there are many ways to decompose the same string, and they cost wildly different amounts:

  • password + 123 → rank-1 dictionary word, times a cheap sequence, times a two-segment structure penalty. Small.
  • pass + word + 123 → three segments, each cheap, but a heavier structure penalty. Bigger.
  • p + a + s + … → all brute force. Enormous.

The attacker takes the cheapest decomposition available to them, so the estimate must take the cheapest one too. Report anything else and you are pricing an attack no rational adversary would run. That single sentence is the technical heart of guessability estimation, and it is why the computation is a search rather than a lookup.

Concretely: build a graph whose nodes are the positions between characters, 0 through n. Every match from position i to j is an edge from node i to node j+1 with weight log(guesses for that match). Every position also gets a brute-force edge to the next, so the graph is always connected. Then the cheapest parse is the shortest path from 0 to n, and you find it with dynamic programming over prefixes: best[k] = min over all matches ending at k of (best[start] · guesses(match) · structure_penalty(segment count)). Log-space turns products into sums; a single left-to-right pass gets you the optimum.

graph LR
    N0((0)) -->|"dict 'password' ≈ 3"| N8((8))
    N0 -->|"dict 'pass' ≈ 130"| N4((4))
    N4 -->|"dict 'word' ≈ 300"| N8
    N8 -->|"seq '123' ≈ 30"| N11((11))
    N0 -->|"brute force"| N1((1))
    N1 -.->|"…"| N11

Every path from 0 to 11 is a story about how an attacker generated password123. The estimate is the cheapest story, not the most flattering one.

That's it. Matchers, per-match pricing, minimum-cost parse. A few thousand lines, most of which is wordlists, and it produces a number with an actual interpretation: this password is reachable in about N guesses by an attacker whose model resembles mine.

Compare that to if (hasUpper && hasDigit && hasSymbol) score += 1, and it's clear these are not two implementations of the same idea. One estimates a quantity; the other counts features.


Where the numbers come from, and what else is on the shelf

zxcvbn's contribution was that a decent approximation of guess number could be computed in a browser, in a few milliseconds, from a few hundred kilobytes of data. It is not the only or the most accurate way to estimate guessability, and it's worth knowing the landscape because it tells you where your estimate's error comes from.

Probabilistic context-free grammars. Weir and colleagues (IEEE S&P, 2009) showed you can learn password structure from a corpus: derive rules like S → L₈D₃ (eight letters then three digits) with probabilities, and terminal distributions for each slot, then enumerate candidates in descending probability order. This is the formalisation of the "structure plus fillers" intuition the parse step above approximates by hand. If your estimator's structure penalty feels like a fudge factor, that's because it is one — PCFGs replace it with a learned probability.

Markov models. Narayanan and Shmatikov (CCS, 2005) started the line of work modelling passwords as character-level Markov chains. Later refinements — backoff, smoothing, careful normalisation — made these strong guessers, particularly for passwords that aren't clean word-plus-digits structures, exactly the cases PCFGs handle worst.

Neural models. Melicher et al. (USENIX Security, 2016) trained recurrent networks to predict the next character and showed two things that matter here. First, at large guess numbers the neural guesser outperformed the earlier approaches, meaning earlier estimates were optimistic in precisely the region where offline attacks live. Second — and this is the part relevant to meters — the model could be compressed enough to run client-side and produce a monotonic guess-number estimate without shipping wordlists at all.

The practical lesson from the comparison literature. Work by Ur and colleagues comparing guessing approaches against real attack teams found that no single guesser dominates; different methods crack different passwords, and any one of them, used alone, materially overestimates how hard a population is to crack. The recommendation that came out of it — take the minimum guess number across several independent guessers — is the same "cheapest parse" principle one level up. The attacker gets to pick their best tool; your estimate must assume they did.

You will not ship an ensemble of guessers in a signup form. But you should know that when your meter says 10¹⁰, the ensemble number is smaller, and the real attacker's number is smaller still.


The systematic bias, stated plainly

Every guessability estimator is a model of an attacker, and it is therefore wrong in one direction.

It cannot price a pattern it does not know about. If your dictionary set has no Spanish, contraseña derivatives score as unmatched brute force. If it has no anime, no football clubs, no local place names, no lyrics from the song everyone in your market knows, those all score as high-entropy nonsense. The attacker targeting your users has those lists. They are cheap to assemble and widely traded. Your estimate for those passwords is not slightly high; it can be off by six orders of magnitude, and it is high exactly for the population you serve and the attacker does too.

Its rule model is a rough sketch. Real mangling rule sets contain thousands of transformations, ordered by empirical productivity. An estimator charging a factor of 2 for capitalisation and a subset count for leet is approximating that badly in both directions.

And the direction of error is never in your favour. Say it as an operating rule: a guessability estimate is an upper bound on the strength of a password under your model, and no kind of bound at all under a better one. Which means it's usable for rejecting weak passwords — a low estimate is trustworthy, because if your mediocre model finds it in 500 guesses, everyone's does — and much weaker as evidence that a password is strong. Rejection is the sound half of the instrument. Approval is the guess.

That asymmetry should drive your product decisions. Blocking on a low score is well-founded. Declaring a password excellent because it scored 4/4 is a claim your estimator is not entitled to make, and the copy in the UI should reflect that.


The blindness that matters most: context

Here's the omission I find in nearly every deployed meter, including ones built on good estimators.

Real cracking against a specific target does not use a generic wordlist. It uses a targeted one, and the targeting is trivial: the user's first and last name, their email local-part, their username, their birth year, the company name, the product name, the site's domain, the current year, the city the office is in, the name of the team that just won something locally. All of that is either in your signup form or on the front page of your website. Attackers append, prepend, leet and capitalise it with the same rule sets as everything else.

A meter that doesn't know any of this will happily rate Acme2026! — at a company called Acme, in 2026 — as strong. It contains a proper noun not in the standard dictionaries, a four-digit number, a symbol. To a targeted attacker it is roughly a two-guess password.

Good estimators support this. zxcvbn takes a user_inputs array, adds it as a rank-1 dictionary, and prices matches against it accordingly. Almost nobody populates it. The fix is a few lines and it's the highest-value change you can make to an existing meter:

  • Everything on the current form: email, username, display name, first/last name, phone digits, whatever else you're collecting.
  • Everything about you: company name, product names, the domain, brand terms, the mascot, common internal abbreviations.
  • Everything about now: the current year and the previous couple, the current month name.

Then treat a match against that list not as "reduce the score a bit" but, for a match that dominates the parse, as a hard rejection with a specific message: this contains your own email address. It's the one rejection users never argue with, because the reason is self-evident.

One caveat: client-side context matching means the client holds the list. For a signup form that's fine — it's all data the user typed or can read on the page — but an enterprise deployment may be shipping internal terms to unauthenticated visitors. If that bothers you, do the context check server-side only. It belongs there anyway.


The meter is a training signal, and this is the pernicious part

Everything so far treats the meter as an instrument that measures a population. It isn't. It's a control that creates one.

Users do not choose a password and then check it against the meter. They interact with the meter until it turns green, and they do that with the smallest edit that works. This is a rational response to an obstacle with a visible success condition, and it is completely universal. Which means:

The distribution of passwords in your database is a function of your validator. Not influenced by it — determined by it, in its head, which is the part that matters, because the head is where the attacker lives.

Now consider what a composition meter teaches. It rewards adding an uppercase letter, a digit and a symbol. Users comply in the cheapest available way: capitalise the first letter, append a digit, append ! (or substitute a leet character they were already going to use). Every user, independently, performs the same three edits, because the meter told them all to. The result is a population concentrated onto a handful of masks and a handful of rule combinations — the shape a mask attack and a rule-based attack are optimally designed to consume.

That is the sharpest way to state the problem. A composition meter isn't merely a bad measurement; it's a training signal that teaches your entire user base to generate exactly the distribution attackers model best. It doesn't fail to help. It actively transfers structure from your policy into your corpus, where the attacker collects it.

A guessability meter also shapes the distribution — it can't not — but it shapes it towards "add another unrelated word," "don't build it out of your username," "that keyboard pattern is known." Those edits move users away from the modelled region rather than towards it. The direction of the feedback loop is what changed, and that's most of the win.

Which leads to a design rule with more force than it first appears to have: every validator is a generator. Before you ship a rule, ask what a user does to satisfy it in one edit, and assume that's what the head of your distribution will look like in a year.


The meter is UX. The control is somewhere else.

A strength meter runs in a browser, on a string, in JavaScript you shipped. An attacker who wants to register password sends the POST directly. Anyone with devtools can set the field to whatever they like. This is obvious when stated and routinely forgotten in architecture, where the meter ends up in the diagram as if it were a control.

The division of labour that actually holds:

The meter is a communication device. Its job is to change what a cooperative user chooses, in the seconds before they commit. It has no adversarial function whatsoever. Judge it on behaviour change, not on security properties.

Enforcement is server-side, and it is three things. A guessability check recomputed on the server (never trust a score submitted by the client — if your API accepts a strength field, delete it). A breach-corpus check, which is a direct measurement of "has this exact string been seen in real dumps" rather than any kind of model — see Checking Breached Passwords Without Sending Passwords for how to do that without shipping plaintext anywhere. And rate limiting plus anomaly detection on the login path, which is what turns "survives 10⁶ guesses" from an aspiration into a fact; that's a capacity problem more than a security one, as Rate Limiting Identity argues, and it interacts badly with naive lockout, which is its own denial-of-service vector.

If you have those three, the meter's failure modes are annoying rather than dangerous. If you don't, no meter saves you.


Shipping one: the decisions that actually come up

This is the part the audience for this article has to get right, so let's be concrete.

Score and explain, not score alone

A coloured bar communicates a verdict and no information. The user learns that the system is displeased, not why, and their next move is a blind edit.

Naming the matched pattern is what changes behaviour. "This is a keyboard walk followed by a year." "This is a common password with letters swapped for lookalike symbols — attackers try those substitutions automatically." "This contains your company name." Work by Ur and colleagues on data-driven meters (CHI, 2017) found detailed, specific feedback produced meaningfully better passwords than a bar alone, and the mechanism is not mysterious: you can only route around an obstacle you understand.

The nice property of a minimum-cost parse is that the explanation is a free by-product. You already know the winning decomposition and which matcher produced each segment. Rendering "Summer (common word) + 2024 (recent year) + ! (most common symbol)" requires no extra analysis — just don't throw the parse away after taking its cost.

State the bar in guesses, and say which attacker it's for

"Must contain a symbol" is indefensible in a design review because nobody can say what it buys. "Must not fall within the first 10⁸ guesses of a standard rule-based attack" is defensible, testable, and — critically — tunable against a threat model.

Pick the threshold from your actual posture:

  • If the credential only ever faces a throttled online endpoint with decent detection, something around 10⁶ is a coherent bar.
  • If a hash dump is in your threat model — and it is — the bar is set by your KDF cost and an adversary's plausible budget. That's an arithmetic exercise, not a vibe; run it with the numbers in The Cost of Password Hashing, and expect to land somewhere in the 10¹⁰–10¹² region for anything worth protecting.

Then write the chosen number, and the attacker model behind it, in a comment next to the constant. In three years someone will want to change it and the reasoning will be gone otherwise. And note the pleasant property: raising your KDF work factor legitimately lowers the required guess threshold. Strength and work factor are substitutes, priced very differently — the entropy article does that exchange-rate math.

The mapping from guesses to a 0–4 score is a presentation decision, and it should be honest about its scale: zxcvbn's buckets are roughly decades of guesses (under 10³, 10⁶, 10⁸, 10¹⁰, above), which means the bar you enforce and the bar you display should be the same number. A meter that shows "good" for something your server will reject is a bug report generator.

Block versus warn

Get the asymmetry right.

Hard-block on things you measured rather than modelled: a breach-corpus hit, a context match that dominates the parse (their own email, the company name), and an estimate below your online-attack floor. These are cases where your estimate is trustworthy — remember, low scores are the reliable half of the instrument.

Warn but allow in the middle band, where your model says "mediocre" but might be wrong in either direction. A false block here is genuinely costly: the user has now been rejected by an oracle they don't understand and will produce a worse password out of frustration, or leave.

Never block on composition. NIST SP 800-63B removed composition requirements and told verifiers to check candidates against a blocklist of known-compromised and dictionary values instead. That trade — direct measurement in, proxy out — is the right one, and it's also the one your users' password managers are already optimised for.

One more asymmetry: a false positive on a blocklist costs the user one more word. A false negative is silent forever. Bias accordingly.

Where it runs, and what it costs

Bundle size is the real constraint, and it's all wordlists. The estimator code is small; the ranked dictionaries are not. Options, roughly in order of how often they're the right answer: trim each list to its most productive prefix (rank beyond a few tens of thousands contributes little to the parse and a lot to the bytes); store lists as rank-only lookups rather than arrays; compress and decompress in a worker; and lazy-load the whole module on first focus of the password field rather than on page load, since nobody types a password in the first 200ms. The meter must not be on the critical path of first paint for a login page — see The Identity Latency Budget for why login-page weight has consequences well beyond aesthetics.

Don't run it on every keystroke. Debounce — 100–200ms of quiet is plenty — and consider capping the analysed length, since the parse is superlinear in the number of matches and a 60-character pasted passphrase generates a lot of them. If you see input lag, move it to a worker; the API is asynchronous anyway once you're debouncing.

And recompute server-side. Same estimator, same wordlists, same threshold, at set-password time. The client's number is advisory; the server's number is the decision. If those two implementations can disagree, you have a support problem, so pin the same library version and the same dictionaries on both sides, and treat a mismatch as a bug rather than a curiosity.

The trap: a maximum your blocklist would reject

The failure everyone ships at least once. Meter says 4/4. User submits. Server rejects: "this password appears in a breach corpus." The user's reaction is entirely reasonable — you just told me it was excellent — and the trust you spent building the meter is gone.

The cause is ordering. Fix it by running the checks in the same sequence on both sides, with the cheap, certain, measured checks first:

  1. Context match (their email, your company name) — a hard no, with the specific reason.
  2. Breach-corpus check — a hard no, with the specific reason. This is measurement; it outranks any model.
  3. Guessability estimate — the score, the explanation, and the threshold.

And enforce the invariant explicitly: the meter must never display a passing score for a string that any later check will reject. If the breach check is asynchronous and slow, hold the meter at "checking…" rather than optimistically showing green. A momentary spinner is much cheaper than a contradiction.

Don't score what the user didn't generate

If a password manager filled the field with 20 random characters, no matcher fires, the estimator prices the whole string as brute force over the observed character set, and that answer happens to be right. But you're now displaying a meter to somebody who has already solved the problem, and any friction you add here — a length cap, a symbol restriction, blocked paste — is actively harmful. Why Password Managers Were the Decade's Biggest Security Win makes the fuller case; the meter-specific version is: recognise when you're talking to a machine-generated secret and get out of the way.

Better still, offer generation at set time. A "generate a passphrase" button next to the field replaces the human generator entirely, which is worth far more than any amount of scoring the human's output — and it's the move Stop Asking Users to Remember Secrets argues for at the system level.

Don't leak the candidate through your telemetry

A meter is a piece of instrumentation attached to a plaintext password field, so somebody will eventually want analytics on it. Log the score bucket, the guess-number order of magnitude, the types of matched patterns, the number of edits before submission. Never the string, never a prefix, never a hash you could dictionary-attack, and never a "we only log rejected ones" exception — rejected candidates are the user's other passwords, and they're the most attack-relevant strings you could possibly collect. This gets built by a well-meaning frontend engineer roughly once per company. Put it in the review checklist.

The score-bucket telemetry, though, is worth having. A composition rule emits no signal about its own operation, which is why it survived twenty years without evidence; a guessability meter can tell you the distribution of estimates in your population, how many users bounce off a rejection, and whether raising the bar moved the median or just the abandonment rate. Prefer controls that report on themselves.


What to ship

Concretely, in the order I'd do it, and with the honest statement of what each part buys.

1. Put the enforcement layer in first, because the meter is not the control. Server-side check on set-password: breach corpus, context blocklist, guessability threshold. Rate limiting and detection on the login path. If you only ever do this and never touch the meter, you've captured most of the available security.

2. Replace composition scoring with guessability estimation. Use an existing implementation; this is not a build-it-yourself problem, and the value is in the corpora, not the code. Run the same estimator on client and server, same version, same dictionaries.

3. Populate the context list, today. Form fields, company and product names, your domain, the current year. It's a few lines, it closes the largest single blind spot in every deployed meter, and it converts the most embarrassing class of weak password into an immediate, self-evident rejection.

4. Explain the match, don't just score it. The winning parse is already sitting in memory. Render it. "A common word, plus a year, plus the symbol most people pick" changes the next attempt in a way that a red bar never has.

5. State the bar in guesses, and record which attacker it's for. Derive the number from your KDF cost and your threat model rather than inheriting it. Order the checks so a green meter can never be contradicted by the server.

6. Delete every composition rule you have. Not because they're merely useless — because they're a training signal pointing the wrong way, and every one you keep is teaching your users to produce the passwords your attacker's rule file was written for.

And then hold on to the thing the incident review took an hour to work out: a meter that is also the gate guarantees that every password in your database scored well, which makes the score worthless as evidence about anything. The only way to know whether your estimator ranks your population correctly is to run a real guesser against your own hashes and compare the guess numbers to the scores you assigned at set time. That's an afternoon of work with hashcat and a sampled export, it is the only measurement in this entire article that is about your users rather than someone's model of users in general, and I have almost never seen a team do it.

Everything else here is a way of making a better guess. That's the one thing that would tell you whether the guess was any good.