Why Biometrics Aren't Secrets

The question that ended the meeting was administrative, not technical. A security questionnaire from a prospective enterprise customer, item 4.3.1: "For any biometric data held, state the encryption algorithm at rest and the key rotation interval."

The team had an answer ready, because they had built the thing properly by every standard they knew. Face templates encrypted with AES-256-GCM. Keys in a KMS, rotated every ninety days, envelope-encrypted, access logged. Someone typed it into the box and moved on.

An engineer who had joined three weeks earlier asked the question that took the rest of the afternoon: rotated to what effect?

Not a pedantic question. Rotating a KMS key re-wraps the ciphertext. It protects against a stolen data encryption key. It does absolutely nothing about the fact that if those templates ever leave the building, the affected users cannot be issued new faces. The ninety-day interval was answering a question about the storage of the credential while the entire risk lived in the nature of it. The questionnaire had no field for that, because the questionnaire had been written by copying the password section and substituting a noun.

The deeper problem surfaced twenty minutes later, when someone pulled up the matcher configuration and found a similarity threshold: 0.62. Nobody could say where the number came from, what false-accept rate it corresponded to on their actual user population, or how many match attempts a single account permitted per hour. The system had a security parameter that no one owned, in units nobody could convert into a risk.

Both problems come from the same category error, the most persistent one in consumer security: treating a biometric as a secret. It isn't one, in a way that has nothing to do with how carefully you store it. A biometric is a public, unrevocable, fuzzy, low-entropy identifier, and in every system that handles it well it is not an authentication factor at all — it is a local unlock gesture for a secret held somewhere else.

The slogan version, which I'll spend the rest of this article earning: the biometric authorises; the key authenticates.

Most engineers know some of this. The part that's usually missing is what "fuzzy" costs you architecturally, and the arithmetic that shows why the well-designed version is safe for a reason that has almost nothing to do with the sensor.

Four properties, and what each one breaks

Take the definition of a secret that the rest of your platform relies on: a value that is (a) known to a bounded set of parties, (b) replaceable, (c) comparable exactly, and (d) drawn from a space too large to search. A password satisfies all four, badly for (d) and well enough for the others. A private key satisfies all four superbly. A biometric fails all four, and each failure breaks a different piece of your architecture.

Public

Your face is on your passport, your building pass, your LinkedIn profile, and a few thousand frames of other people's photographs. Your fingerprints are on every glass you've put down. Your voice is in every meeting recording. Your iris is resolvable from a sufficiently good photograph at a surprising distance.

This isn't a marginal leakage problem to be managed with better hygiene. The secrecy assumption fails at the definition. There has never been a moment when your fingerprints were confidential, and no policy you write will create one. Any design whose security argument contains the step "and the attacker does not know the user's face" is unsound at the first line.

So a biometric can never serve as the possession proof or the knowledge proof. At most it is evidence that a particular body was in front of a particular sensor at a particular moment — genuinely useful, and not a secret.

Unrevocable

This is the one everybody can recite, and it's usually stated too weakly. The standard framing is "you can't change your fingerprint," delivered as a wry aside. The consequence is structural: a biometric cannot participate in a credential lifecycle, and the credential lifecycle is the load-bearing beam of every identity platform.

Everything else you hold can be rotated, and your operational safety comes from that. Signing keys rotate (designing for key rotation before you need it is an article on its own because the machinery is non-trivial). Password hashes get re-derived under a stronger KDF. API keys get reissued. Sessions expire. Certificates have a validity window precisely so that a compromise has a horizon.

Rotation is how incident response terminates. "We rotated everything in scope" is the sentence that closes a postmortem. For a biometric there is no such sentence. The blast radius of a compromise is not measured in hours to rotate; it is measured in the remaining lifespan of the person. If you hold face templates and you lose them, the remediation plan is: there is no remediation plan. You can invalidate the system's use of them — see cancelable biometrics below — but you cannot invalidate their use by anyone else, forever.

It also breaks retention reasoning. A biometric template's risk does not decay with age the way a password's does — a fingerprint template will still be usable in forty years — so you have created the one asset in your system whose value to an attacker is roughly time-invariant.

Fuzzy

This is the deep one, and it is where the interesting engineering lives.

Password verification is an exact comparison: run the candidate through a KDF with the stored salt and compare byte for byte. Two outcomes, no tunable parameters. The whole reason you can store only a hash is that the comparison is exact — the avalanche property means one input bit flips half the digest, which is precisely what makes the digest safe to store. (The three-way confusion between hashing, encryption and encoding is where this stops being pedantic and starts being architectural.)

Biometric verification cannot work that way, because two readings of the same finger are never identical. Different pressure, different moisture, a cut, a rotated placement, a different sensor. Two photographs of the same face differ in pose, illumination, expression, age, and camera. So the matcher does not compare; it scores. It extracts features, computes a similarity between the probe and the stored template, and compares that similarity to a threshold. That threshold is the security parameter, and it is continuous.

Which gives you the two error rates that every biometric system trades against:

  • FAR (false accept rate, or FMR in the ISO vocabulary): the probability that a different person's presentation scores above threshold.
  • FRR (false reject rate, FNMR): the probability that the legitimate person's presentation scores below it.

They move in opposite directions as you slide the threshold. Plot one against the other and you get the DET curve, which is the honest description of a matcher's quality. The equal error rate (EER) is the single point on that curve where FAR equals FRR — a convenient scalar for comparing algorithms and a terrible operating point for deploying one, because almost no real system wants those errors weighted equally. A phone unlock wants low FRR (a false reject is a user annoyed ten times a day). A border control gate wants low FAR. Reporting a system's EER tells you about the algorithm; the thing you actually deploy is a threshold choice, and that choice is a product decision with a security consequence.

Three things follow that people routinely miss.

Vendor FAR figures are measured under conditions that don't resemble your deployment. A published false-match rate — Apple's often-cited roughly 1-in-50,000 for a single enrolled fingerprint, roughly 1-in-1,000,000 for Face ID, or NIST FRVT results where good 1:1 face algorithms reach around 10⁻⁶ — is measured against zero-effort impostors: a random person presenting their own genuine trait, not trying to look like you. That is not a resistance-to-attack figure. An attacker has your photograph, a latent print, or a resemblance they selected themselves for. And the same algorithms that hit 10⁻⁶ on cooperative, well-lit, frontal images degrade by orders of magnitude on unconstrained captures. Datasheet figure is to deployment figure as lab MPG is to your commute.

Error rates are not uniform across your population. NIST's demographic-effects work found false-match rates varying by up to two orders of magnitude across demographic groups for some face algorithms, and there is a well-known long tail of users whose fingerprints simply don't enrol well — worn ridges, manual labour, age, certain medical conditions. A single global threshold delivers materially different security and different usability to different users, and neither difference will show up in your aggregate dashboards. Some fraction of your users will be locked out of a factor that works flawlessly for everyone in the room where you chose the number.

You cannot hash a biometric template, and the reason is not laziness. People propose it constantly: "store SHA-256(template) and compare digests." It cannot work, because matching must tolerate variation and a cryptographic hash annihilates variation by design. The property that makes a hash safe to store — one input bit flips half the output bits — is exactly the property that makes it useless for a comparison that must succeed on inputs that differ. So a conventional biometric system stores something functionally equivalent to plaintext: a feature vector that the matcher must be able to compute distances against. Encrypt it at rest if you like — it must be decrypted to be compared, so the protection ends at the boundary of the matching process.

Low entropy in practice

The theoretical information content of a fingerprint is often quoted at figures in the dozens of bits; the classic Ratha–Pankanti–style analyses of minutiae correspondence produce numbers around 80 bits for a full match of a well-imaged print with many minutiae. Those numbers describe an idealised comparison of complete templates.

What matters for security is different: the discriminative entropy as used by the deployed matcher at its deployed threshold. Real matchers work on partial overlaps, tolerate rotation and translation, accept a subset of minutiae, and are tuned to keep FRR low. Every one of those accommodations widens the acceptance region and reduces the effective entropy. Published estimates vary widely by methodology, and I'd treat any single figure sceptically, but the direction is unambiguous and large: the effective figure sits well below the theoretical one, commonly cited in the range of tens of bits, and face templates land lower still — the shape of a face has fewer independent degrees of freedom than a ridge pattern, and matchers must tolerate more variation.

Put that in units you use elsewhere. A deployed operating point with a FAR of 10⁻⁵ is, per attempt, equivalent to guessing a secret from a space of about 100,000 — roughly 17 bits. A FAR of 10⁻⁶ is about 20 bits. That is the honest per-attempt strength of the biometric itself. A four-digit PIN is 13 bits. A six-word passphrase is 77. The biometric is closer to the PIN than to anything you would call a credential.

At which point the obvious objection is: so how is any of this safe? Which is the right question, and the answer is the whole architecture.

The arithmetic that actually makes it safe

A biometric's security is not its FAR. It is its FAR multiplied by the attacker's attempt budget.

This is the sentence I would most like people to take away, because it relocates the security of the system to the component that actually provides it.

Do the arithmetic. Take a fingerprint sensor at a FAR of 1 in 50,000. With an attempt budget of five before the device demands a passcode, an attacker with unrelated fingers has roughly a 1-in-10,000 chance — about 13 bits, weak-ish but bounded, and bounded is the operative word. Now remove the budget. At 200 milliseconds per attempt, an unbounded attacker expects success in about 50,000 tries: under three hours. Same sensor, same threshold, same published FAR. The difference between "adequate" and "trivially broken" is entirely a property of who counts the attempts.

Passwords survive being low-entropy for the same reason, which is why the biometric case feels familiar once you see it: an eight-character human-chosen password would be hopeless if it could be guessed at unbounded rate against a live endpoint. It is protected by rate limiting, lockout, and — for the offline case — a deliberately expensive KDF. All three are attempt-budget mechanisms. The biometric has no equivalent of a KDF, because there's no exact comparison to make expensive, so the entire budget must come from the enforcement environment.

That is what the secure enclave is for. Not primarily to keep the template confidential — though it does that — but to be the thing that counts. It holds the template, performs the comparison in a place the host OS cannot observe, enforces a hard attempt limit, falls back to a passcode when the limit is hit, and emits exactly one bit.

Which reframes "is biometric authentication secure?" entirely. You are not evaluating a sensor. You are evaluating whether a tamper-resistant component with a monotonic failure counter stands between the sensor and the secret. If it does, 17 bits is fine. If it doesn't, 20 bits isn't enough for anything.

Template protection: the field that exists, and why you've never deployed it

There is a substantial literature on the problem "store something derived from a biometric that is safe to lose," worth knowing about both because it's clever and because knowing why it isn't in your stack tells you something. ISO/IEC 24745 codifies three properties a protected template should have:

  • Irreversibility: you cannot reconstruct the biometric (or a usable synthetic one) from the stored data.
  • Unlinkability: templates from the same person enrolled in two different systems cannot be correlated.
  • Renewability: you can issue a new protected template from the same trait, so a compromised one can be revoked.

Renewability is the striking one. It is an attempt to give a biometric the property it fundamentally lacks — a lifecycle — by making the stored artifact a function of the trait and a system-specific parameter. Change the parameter, get a new artifact, invalidate the old. Hence "cancelable biometrics."

Two families dominate:

Cancelable biometrics / feature transforms. Apply a repeatable, non-invertible distortion to the feature vector before storage — a random projection keyed per-system (BioHashing), a coordinate warp of the minutiae plane. Matching happens in the transformed domain, so the raw template never exists on the server, and a different key per relying party buys unlinkability and renewability at once.

Biometric cryptosystems / fuzzy extractors. Rather than storing anything comparable, bind a key to the biometric. A fuzzy extractor uses error-correcting codes to derive a stable, uniform key from a noisy input, storing only public "helper data" that corrects the variation without revealing the trait. The fuzzy vault construction hides a secret polynomial among chaff points, recoverable only by presenting enough genuine minutiae. When it works you get close to the property you wanted: a stored value that is useless to a thief and produces a key on a correct presentation.

So why is essentially none of this in production?

Accuracy costs. Every transform or error-correction layer discards information, and the loss shows up directly as a worse DET curve. A scheme that pushes the EER from 1% to 3% is academically modest and commercially fatal — it means triple the support tickets, on a factor whose entire commercial case is that it removes friction.

Interoperability collapses. Only a matcher that knows the transform can match, so you lose the vendor ecosystem, the standard template formats (ISO/IEC 19794), and any hope of an off-the-shelf sensor SDK — and in fingerprint deployments the matcher is usually somebody else's licensed algorithm, which wants a standard template.

The security proofs are conditional in awkward ways. Fuzzy vault schemes are vulnerable to correlation attacks when the same finger is enrolled in two vaults with different chaff, which quietly undermines the unlinkability claim they were built to provide. Several published transforms have been shown to leak more than intended, or to be partially invertible under a known-key model. The field is active and the analyses are moving; that's a poor foundation for a ten-year credential decision.

And most importantly: the problem it solves is one you can avoid entirely. Every one of these schemes is machinery for holding biometric data safely on a server. The dominant architecture of the last decade sidestepped the requirement by not putting it there. Template protection is the sophisticated answer to a question the enclave answered by deletion — which is the same move as removing the password asset rather than hardening it, applied one layer down.

Know these schemes exist. Reach for them only when the biometric genuinely must live server-side — large-scale deduplication or identification (1:N), a different problem from authentication and one most product teams should never be solving.

The architecture that makes biometrics good

Here is the whole design, and it is simpler than the preceding four sections implied.

flowchart LR
    subgraph device["User's device"]
        subgraph enclave["Secure element — tamper-resistant"]
            T["Biometric template<br/>(never leaves)"]
            M["Matcher + attempt counter"]
            K["Private key<br/>(non-exportable)"]
        end
        S["Sensor"] --> M
        T --> M
        M -->|"boolean: verified"| K
        OS["Host OS / browser"]
    end
    K -->|"signature over challenge<br/>+ UV flag"| RP["Relying party"]
    OS -.->|"challenge"| K
    RP -->|"verify signature<br/>against stored public key"| RP2["Session"]

    style enclave fill:#eef,stroke:#333
    style RP fill:#efe,stroke:#333

Trace the trust boundaries. The template crosses none of them. The match happens inside the enclave and its only output is a boolean consumed inside the same enclave. The host OS never sees biometric data — it sees a key that either will or won't sign. The relying party receives a signature and, in the authenticator data, a single bit saying that a local user-verification step succeeded. (The byte-level mechanics of that structure are in Under the Hood of WebAuthn; the point here is what the bit means.)

So the honest security accounting of "biometric login" is:

Component What it contributes
The private key Unforgeability. This is the authentication.
The enclave Non-exportability of the key; the attempt budget on the matcher.
Origin binding in the protocol Phishing resistance.
The biometric A cheap, fast, hard-to-borrow authorisation gesture for using the key once.

Remove the biometric and replace it with a device PIN and the security of the system barely moves. Remove the enclave and the whole thing collapses regardless of how good the sensor is. That asymmetry is the argument in one line: the security rests on the key and the enclave, not on the biometric at all.

Which is why "biometric authentication" is a category error rather than a loose phrase. The biometric authorises a local action. The key authenticates to a remote party. Those are different verbs pointed at different audiences, and conflating them is what produces the questionnaire item at the top of this article.

A corollary most MFA policies get wrong

If UV is a single bit, then the relying party cannot tell what satisfied it. A face match, a fingerprint, a device unlock pattern, or a four-digit PIN all set the same flag. There is no field in the protocol carrying the modality, no field carrying a match score, and none carrying the strength of the fallback.

So an internal policy that says "high-value operations require biometric verification" is, in implementation, a policy that says "require UV=1" — and is satisfied by a user's 4-digit device PIN, or by whatever the platform's fallback happens to be. If your compliance narrative distinguishes biometric from PIN, your protocol does not, and no amount of userVerification: "required" will make it. The honest statement is "verified locally by whatever the device considers sufficient," and the security of that claim rests on the authenticator's integrity — which is the only thing attestation could tell you about, at a cost I've argued elsewhere is usually not worth paying.

I'd rather teams write the honest sentence in the policy than build a control that quietly doesn't exist.

Five ways to get this wrong

1. Server-side matching. You now hold an asset that is unrevocable, time-invariant in value, must be decryptable to be usable, and whose breach you can never remediate. You have also — the part that gets missed — moved the attempt budget from a tamper-resistant counter to your own rate limiter: application code, reachable over the network, with the usual failure modes and the DoS trade-offs of lockout. The FAR arithmetic now runs against your throttle, keyed on an identifier the attacker supplies.

There are legitimate cases — 1:N deduplication in national ID or benefits programmes, access control on a shared door, fraud systems detecting one person enrolling as five. All are identification problems, not authentication problems. If someone proposes server-side matching for logging into an app, the requirement is already solved better by a local gesture over a key.

2. Biometrics as a recovery factor. The worst possible placement, and depressingly common, because it looks like the strongest thing you have. Recovery is the path where all your other assumptions are already suspended: the user has lost the device, so there is no enclave, no attempt budget, and no established authenticator. A "verify with a selfie" recovery flow is server-side matching plus an unauthenticated capture channel plus a total absence of the environment that made the biometric safe. And because the trait is public, the attacker's input is not hard to obtain.

If you must use face verification in recovery, treat it the way the recovery article treats every other signal: as one weighted input into a risk decision, never as a sufficient condition. The passkey recovery piece makes the structural argument — recovery is the account's real security ceiling, and putting an unrevocable public identifier at that ceiling sets it low.

3. Counting it as a second factor when it unlocks the first. A passkey unlocked by Face ID is often described as two factors: something you have (the device) and something you are (the face). That's defensible as a description of the ceremony. It is not defensible as a claim about independent failure modes, because both factors terminate in the same enclave. Compromise the enclave and you have both. There's a stronger version of the mistake: a system that uses the same fingerprint to unlock a password manager and counts the fingerprint as its second factor is running one factor and reporting two.

Factor-counting is the wrong frame anyway. A UV-verified passkey is excellent not because it's "two factors" but because the credential can't be phished and the gesture can't be replayed at rate.

4. Presentation attacks, and the liveness arms race you can't audit. Spoofing is a live field: printed and silicone fingerprints, high-resolution photographs, video replay, 3D-printed masks, and, increasingly, generated media against face and voice systems. Voice is in the worst position by some distance — synthesis quality has moved faster than detection, and a voice sample is trivially obtainable.

Detection is called PAD (presentation attack detection), it's testable under ISO/IEC 30107-3, and it improves. The structural problem is that its current state is unobservable from your side of the API. You call a platform API and get a boolean back. You do not learn which PAD level was applied, whether the sensor has a depth channel, whether the model was updated after last month's published bypass, or whether the user is on hardware whose liveness detection was defeated years ago and never patched. Which is a real argument for treating the UV bit as a good signal rather than a strong one, and keeping other risk signals in the decision (adaptive MFA is the machinery for that).

5. The coercion asymmetry. Jurisdiction-dependent, so I'll stay general and offer no legal advice: in several legal systems, compelling someone to produce a physical characteristic has been treated differently from compelling them to disclose something they know, and the case law is unsettled and varies by country and over time. The design consequence holds regardless of any particular ruling: a biometric can be applied to a device by someone holding both the device and the person, without their cooperation, in a way a passphrase cannot. If your users include journalists, activists, or anyone whose threat model contains device seizure, offering a fast way to disable biometric unlock is a real feature — several platforms ship exactly that, which tells you the threat model is recognised.

Property, consequence, response

Property Consequence Correct design response
Public — the trait is observable and collectable Cannot serve as a secret; "attacker doesn't have it" is never a valid premise Use it only as a local gesture; never as the thing proving identity to a remote party
Unrevocable — no reissue is possible No lifecycle, no rotation, no incident termination; blast radius is a lifetime Never store it where a breach is possible; if you must, use ISO/IEC 24745 template protection with renewability, and accept the accuracy cost
Fuzzy — matching is a threshold over a similarity score Cannot be hashed; stored form is effectively plaintext to the matcher; there is a tunable security parameter someone must own Match inside a secure element; make the threshold an explicit, documented, reviewed decision with a stated FAR — not a default
Low entropy in practice — tens of bits at the operating point Per-attempt strength is comparable to a PIN Enforce a hard attempt budget in tamper-resistant hardware; treat FAR × attempts as the real figure
Fallible per-person — enrolment and error rates vary across users A silent minority is locked out or less protected Always ship a non-biometric path of equal assurance; measure FRR by cohort, not in aggregate
Coercible — can be applied without cooperation Different threat profile from a passphrase under duress Offer fast biometric disable; don't make biometric the sole unlock for high-risk populations

When a vendor proposes server-side biometric matching

You will encounter this — often in fintech onboarding, sometimes in workforce products, occasionally as a "passwordless" pitch that is neither. These are the questions that separate a considered design from a demo, and I'd ask them in this order:

  1. Where does the comparison execute, and where does the template rest? If the answer is "our cloud," everything below matters. If it's "on the device, in the secure element," most of it evaporates.
  2. What is the threshold, what FAR does it correspond to, and on what population was that measured? A vendor who cannot produce a DET curve for a dataset resembling your users is quoting a lab figure. Ask specifically whether the measurement used zero-effort impostors.
  3. What is the attempt budget, who enforces it, and what happens at the limit? If the answer is application-level rate limiting, ask what the limit is per account, per IP, and per template — and what an attacker who controls the account identifier can do with it.
  4. What are the FRR figures broken down by demographic cohort and by sensor? Silence here means the question has never been asked, which means some of your users will fail and you won't find out from the aggregate.
  5. What PAD level is claimed, tested by whom, against what attack instruments, and how are models updated? "We have liveness detection" is not an answer.
  6. What is the remediation plan for a template breach? There isn't one. The value of asking is that it forces the conversation about why the templates exist at all.
  7. What happens to templates on account closure and contract termination? This is also where the regulatory weight lands — biometric data is a special category under GDPR, and several jurisdictions have statutes with private rights of action attached, which changes the economics of a breach considerably.

If questions 1 and 6 don't have satisfying answers, the rest is a discussion about the decor of a building with no foundation.

The honest case for biometrics

Everything above is a set of constraints, not a dismissal. Biometrics earned their place, and the reason is worth stating precisely.

They removed the typing, and the typing was the adoption barrier. The value of passwordless is the removal of an asset class (the argument in full); the reason that removal became achievable at population scale is that the replacement gesture takes 300 milliseconds and requires nothing of the user's memory. Strong asymmetric credentials existed for decades and stayed confined to specialists. What changed wasn't the cryptography — it was that unlocking a key stopped requiring a passphrase. Biometrics are not a security mechanism that happens to be convenient; they are a convenience mechanism that made a security mechanism adoptable, which is a larger contribution than most security mechanisms make.

They are a cheap, high-quality presence signal. A biometric gesture is strong evidence that a specific human being was physically at a specific device at a specific instant. Nothing else in your toolkit produces that as cheaply. It's the correct primitive for step-up on a high-value action, where the question isn't "who are you" — the session already answered that — but "is a person here, now, and did they mean this?"

They are hard to share. Not impossible, but a user cannot casually hand their face to a colleague the way they hand over a password. Credential sharing is one of the most common and least reported control failures in workforce settings, and a gesture with a body attached to it is real friction against it.

The correction, stated precisely

A biometric is not a secret and cannot be made into one. It is public, unrevocable, fuzzily matched, and worth tens of bits per attempt. Every property that makes a secret useful, it lacks.

None of which is a problem, because in a well-built system it isn't doing the job of a secret. It is a local unlock gesture, evaluated inside tamper-resistant hardware that enforces the attempt budget the trait cannot enforce for itself, gating the use of a key that does the actual authentication. The biometric authorises; the key authenticates.

Which gives you the diagnostic to apply to any system you're handed. Don't ask how good the sensor is. Ask two questions: does the template ever leave the secure element, and who counts the failed attempts. If the answers are "no" and "tamper-resistant hardware," a 17-bit factor is doing an excellent job of a well-scoped task. If the answers are "yes, to our servers" and "our rate limiter," then you are storing an unrevocable public identifier as though it were a password — and the ninety-day key rotation interval on the questionnaire is answering a question nobody asked.