Authentication Should Be Fast, Boring, and Predictable

The bake-off had been running for six weeks and was, by any normal measure, already decided. One vendor had a feature matrix with four hundred rows and a demo that showed a login flow being reshaped live on stage. The other had a product you could describe on an index card.

Then a platform engineer from the customer side — someone who had clearly been dragged into the meeting and wanted to leave — asked a question that wasn't on the evaluation sheet.

"What's your p99.9 for token issuance, and does that number change if I turn things on?"

The first vendor's engineer, to his credit, answered honestly. He said it depended on the configuration. He said most customers saw something in the low hundreds of milliseconds. He offered to run a benchmark against a representative tenant.

The room heard "we don't know." He had actually said something more precise and more damning: the product does not have a latency profile. Its customers do.

That distinction is the subject of this article. Boring is not a personality trait that engineers with conservative temperaments impose on software. It is a specification — a set of claims about behavior that can be written down, tested, breached, and put in a contract. Determinism, bounded latency, and no unbounded work in the request path are properties in the same category as "supports SAML" or "SOC 2 Type II." They are things a customer buys. And like any property, you can only sell them if you can state them, which turns out to be the hard part, because every configurable branch you ship makes the statement weaker.

Boring, written as three clauses

"Boring" as a compliment is useless. Everyone agrees they want it and nobody changes a design decision because of it. So let's write it as something a reasonable engineer could argue with.

Clause 1 — Determinism. Given the same credential, the same projected tenant state at version V, and the same wall-clock bucket, the system produces the same decision. Not "usually." Always, and demonstrably.

Clause 2 — Bounded latency. There exists a number D such that the system either completes the operation within D or fails it explicitly. D is enforced by a mechanism, not hoped for by a dashboard.

Clause 3 — No unbounded work. Every loop, traversal, allocation, and outbound call on the authentication path has a bound that is known at build or config time, not at request time.

Each clause is falsifiable. Each one costs something specific. And in every system I've seen where authentication had a bad reputation, at least one of the three was violated in a way nobody had written down.

Clause 1: determinism is a testing strategy, not a virtue

Most teams read "deterministic" as a synonym for "correct" and move on. It's more useful than that, because determinism is the property that makes a class of test possible, and without that class of test you are debugging authentication by reproduction — which is to say, not at all.

The test is decision replay. If the authentication decision is a pure function of its inputs, you can log the inputs alongside the decision, and later re-run the current build against yesterday's inputs and diff the outputs. Every divergence is either a bug or an intended behavior change, and you find out which one before the customer does. This is unremarkable practice in payments and risk engines. It is rare in identity, and the reason is almost always that somebody made the decision impure.

The impurities are worth enumerating, because they're not obvious and they're not all bad:

Impurity on the auth path What breaks Usually fixable by
Outbound call to a system you don't own Replay is impossible; the decision depends on someone's Tuesday Pre-compute the attribute (SCIM, scheduled sync)
Reading "current" state instead of a versioned snapshot Two replays of the same input legitimately differ Stamp the decision with the state version it used
Wall-clock reads scattered through the path Everything is technically nondeterministic Take the clock once at ingress, pass it down
Randomness (jitter, sampling, A/B assignment) Divergence looks like a bug Seed it from the request ID and log the seed
Load-dependent behavior (a fallback that fires only under pressure) The interesting decisions are the unreproducible ones Log which path was taken as part of the decision record
Customer-supplied code All of the above, permanently See The Extensibility Trap

The last row is the one this blog has argued at length, so I won't re-run it. What I want to add is the framing that makes the argument concrete rather than aesthetic: arbitrary code execution in the auth path isn't primarily a security problem or a latency problem. It is the deletion of your regression test strategy. Once a tenant's login depends on a JAR you can't run, you cannot replay their decisions, which means you cannot prove that your next release doesn't change them, which means your release process becomes discovery rather than verification.

The stamped-version row deserves a note too, because it's the cheapest high-value change on the list. If your runtime serves projected configuration — and most mature identity platforms do, for reasons worked through in Why Every Identity Platform Eventually Looks the Same — then a decision made against projection version 41 and a decision made against version 42 are supposed to differ. Without the version in the decision record, that legitimate difference is indistinguishable from a regression, and your replay harness cries wolf until someone turns it off.

There's a second reason determinism is getting more valuable rather than less, and it has to do with who is calling. Human users retry by clicking again, slowly, a couple of times. Machine callers retry immediately, in a loop, with a policy. A nondeterministic decision under human traffic produces a confused user; under machine traffic it produces divergent state across retries of what the caller believes was one operation. As non-human identities become the numerical majority of principals, the cost of an auth decision that isn't a function of its inputs goes up roughly in proportion to retry aggressiveness.

Clause 2: a latency target is not a latency bound

Here is the distinction almost every performance conversation elides.

A target is a statistical description of what the system did last week. A bound is a promise about what it will do next week, enforced by a mechanism that fires when the promise would be broken. The mechanism is almost always the same one: a deadline that converts overrun into an explicit failure instead of an unbounded wait.

The reason this matters is that unbounded waits don't stay local. A login request that hangs holds a connection, a thread, a pool slot, and — critically — a slot in every caller upstream of it. Understanding Tail Latency in Authentication works through why a rare slow event becomes a common user experience through serial and fan-in amplification; the capacity version of the same point is less often stated:

Your variance, not your mean, sets your callers' capacity.

Little's Law gives the arithmetic. Concurrency = throughput × latency. A service doing 2,000 requests per second at a 40ms mean holds 80 requests in flight. Fine. But a caller cannot size its connection pool from your mean — it has to size from your timeout, because a slot is occupied for however long the slowest response takes. If your p99.9 is 40ms, a caller can safely run a 100ms timeout and a small pool. If your p99.9 is 3 seconds, that same caller must either hold pool slots for 3 seconds or cut you off at a timeout that turns your slow responses into their errors. There is no third option.

So a system with a 40ms median and a 3-second tail is not "fast with occasional slowness." It is a 3-second system that is usually early, and everyone who integrates with it must provision for the 3 seconds. This is why a uniformly slower system genuinely feels better and costs less to operate than a fast-but-variable one, and it's why "reduce p99.9" is a different engineering project from "reduce p50" with a different set of moves.

The moves that reduce variance are unglamorous and mostly consist of refusing to do things:

  • Deadline propagation from the top. The journey budget is decomposed downward into per-hop deadlines. No hop may hold a timeout larger than the remaining budget, because a 5-second timeout on hop four is a promise that the whole journey may take five seconds.
  • Bounded queues with fast rejection. An unbounded queue converts overload into unbounded latency for everyone. A bounded one converts it into a fast, honest failure for a few. The second is better for the user and for the system, because a nine-second spinner is what generates the retry that deepens the overload.
  • Admission control before expensive work. Reject what you can reject before the password hash, which is 100ms of CPU you cannot get back. Rate limiting on the credential endpoint is a capacity control before it is a security control; that argument is made properly in Rate Limiting Identity Is a Capacity Problem.
  • Headroom. Expected wait scales as ρ/(1−ρ). At 50% utilization it's about one service time; at 90% it's nine. You cannot run an authentication service hot and have a good p99.9, and this is a budget conversation, not a technical one.

Note what's absent from that list: making anything faster. Variance work is almost entirely about what happens when things go wrong, which is why it never appears on a feature roadmap and why it is the actual product.

Clause 3: find the unbounded loops, because you have some

"No unbounded work" sounds like it needs no elaboration until you go looking, at which point it's an afternoon's audit with a genuinely surprising yield. The auth path collects unbounded work the way a drain collects hair, and almost none of it is visible in a normal code review because each site looks like a small, sensible loop over a small, sensible collection.

A partial inventory, all of which I have seen bite something real:

Unbounded thing Bound it usually needs Failure mode when unbounded
Nested group traversal for effective entitlements Max depth, max total groups One customer's 300k-membership graph is a 4-second login
Claim mapping rules per tenant Max rules, max output size A token that exceeds a header limit at the gateway, not at issuance
Regex in claim mapping or password policy No backtracking (RE2-style) or a step budget Catastrophic backtracking: a 30-character input burning a core for minutes
Audience / scope lists Cardinality cap Quadratic validation against a large authorized set
SAML assertion / XML depth Parser limits, entity expansion off Billion laughs, and the parser is on your auth path
Federated IdP fallback chains One attempt, hard deadline Sequential timeouts stack into a minute-long login
Retries at each layer A retry budget as a fraction of traffic 3 retries × 4 layers = 81 backend requests from one login
JWKS fetch on unknown kid Non-blocking, prefetch on rotation Thundering herd on the path nobody load-tested

The regex row is the one I'd check first if I only had an hour. Claim transformation is exactly the sort of "safe, declarative, no code execution" extensibility that a well-behaved platform offers instead of plugins — and a backtracking regex engine turns that safe declarative surface into a customer-supplied CPU bomb with a config-file delivery mechanism. The fix is boring and total: use a regex engine with linear-time guarantees, or impose a step budget and fail the rule rather than the process. Offering constrained expressiveness is only meaningful if the constraint is actually enforced by the evaluator rather than by the documentation.

Flexibility and predictability are the same currency

Now the part that makes the three clauses hard to keep.

Every optional behavior you ship is a branch. A branch is a state you must test and a latency distribution you must characterize. The combinatorial cost of the first has been argued here already — Configuration Beats Customization makes the entropy case, and I don't want to repeat it. The latency consequence is the one that gets missed, and it's sharper than the testing one because it has an immediate commercial effect.

With N optional behaviors, no two tenants execute the same code path. Your published p99 is therefore not a property of your software. It is a percentile of a mixture distribution, weighted by how your current customers happen to be configured.

The arithmetic is worth doing once, because it's more brutal than intuition suggests. Suppose 95% of your tenants run the plain path at a tidy 60ms p99, and 5% have enabled an option — a directory lookup, a per-request policy compile, a strict federation mode — that puts them at 800ms p99. Your global p99 is not a weighted average of 60 and 800. That 5% cohort occupies the top 5% of the population distribution outright, so your global p99 is entirely inside the slow cohort — it is roughly that cohort's 80th percentile. Your headline number describes only the customers you'd rather not talk about, and it describes them optimistically.

Two consequences follow, and the second is the one that ends arguments.

First, per-tenant percentiles aren't a nice-to-have; they're the only measurement that describes the software rather than the sales pipeline.

Second: a performance number that moves when sales closes a deal is not an engineering metric. Onboard one large tenant on the slow configuration and your global p99 changes without a line of code shipping. If you have published that number, you have published something you don't control. Teams discover this at the worst possible moment — usually when a customer quotes your own datasheet back at you during an incident review.

This is what "flexibility and predictability are the same currency" actually means in practice. It isn't a moral claim about restraint. It's an observation that the set of behaviors you allow determines the set of claims you can make, and those two sets trade off exactly. Each new option buys you a deal and spends a bit of your ability to say anything definite about the system to anyone.

Why "no new features this quarter" is a legitimate roadmap here

Identity is one of the few systems where a roadmap consisting mostly of nothing observable is a defensible plan, and the reason is structural rather than cultural.

Authentication has extreme fan-in. Every service, every page load, every machine caller passes through it. That's the property that makes it infrastructure rather than a feature, argued in Authentication Is Infrastructure. Fan-in has a specific consequence for change: the expected cost of an incident is its probability times its blast radius, and the blast radius of an authentication defect is the entire estate. Since the dominant trigger of incidents in mature systems is change, change cadence is the top-line reliability lever, and reducing it is engineering work with a measurable return, not conservatism.

The trap is thinking "no features" means "no work." The invisible roadmap is full:

  • Variance reduction — the p99.9 work above, which no customer will ever see as a feature and every customer will feel.
  • Deletion — retiring options, collapsing branches, shrinking the occupied configuration set. Each removal restores a claim you can make.
  • Determinism hardening — building the replay harness, stamping versions, removing the last synchronous outbound call.
  • Failure-mode work — the deliberate exercises in Chaos Engineering for Identity Systems, and the on-call ergonomics in Running an Identity Platform On-Call.
  • Capacity ahead of demand, which is what Designing for the Next Million Logins is about.

None of it demos. All of it is what the customer is buying when they buy boring. The organizational problem — how you decline the request that would undo it — is its own discipline, and How to Say No to a Feature Request covers it better than a paragraph here would.

Making predictability visible to someone who is buying

A property nobody can verify is not a property; it's a vibe. If boring is the product, it has to be legible from the outside, and that means publishing things that can embarrass you.

A percentile claim needs four qualifiers or it's unfalsifiable. "Sub-100ms authentication" means nothing. A real clause names the measurement boundary (where the clock starts and stops), the load level it holds at, the exclusions, and the aggregation method:

p99.9 for token issuance ≤ X ms, measured at runtime ingress to response flush, at up to N requests/second/tenant, excluding time in federated upstream IdPs, computed from merged histograms over 5-minute windows.

Every one of those qualifiers is somewhere a vague claim hides. "Measured at ingress" excludes DNS, TLS, and redirects — which is honest only if you say so, because the user's experience includes them, as The Identity Latency Budget walks through line by line. "Merged histograms" matters because averaging per-instance quantiles is the single most common instrumentation bug in this space, and it produces a green dashboard during an outage.

A change cadence is a published artifact. How often does the runtime change? What's the maximum blast radius of a routine deploy? Is there a window where nothing ships? A vendor who can answer this is telling you they've thought about the deploy that logs everyone out; one who can't is telling you their release process is discovery.

A deprecation policy is the honest version of a compatibility promise. Minimum notice period, what counts as a breaking change, whether observable behavior is in the contract or only the documented behavior. Hyrum's Law says customers depend on everything they can observe, so the useful policy states which observable behaviors you are willing to freeze — and, by omission, which ones you aren't.

A stated non-capability is a predictability claim. "We do not execute customer code in the runtime" is a stronger performance statement than any benchmark, because it's the reason the benchmark remains true after deployment. I work on ClavionX, so treat this as disclosure rather than pitch: the design there holds two rules that exist for exactly this reason. The runtime never synchronously calls the control plane — event and projection only, failing secure on missing projected state — which keeps the fast path's availability independent of the slow path's. And there is no customer code execution in the identity runtime at all; the sanctioned alternatives are declarative claim mapping, a policy engine, async webhooks off the critical path, and pre-computing attributes ahead of time via SCIM or scheduled sync. The machine-identity work — agents as registered objects configured through reusable policy objects that compile onto the underlying OAuth2 client, rather than through raw grant types and TTLs — is still in design, and I'd rather say that than imply maturity it doesn't have. The point isn't the specific design; it's that both rules are statements, and statements are what make a percentile claim survive contact with a customer's configuration.

The honest counterpoint: boring loses to useful

Every argument in this article is downstream of one assumption — that the customer's problem is already solved and the remaining question is how reliably. When that assumption is false, boring loses, and it deserves to.

A platform that cannot express a customer's group model, cannot federate with the IdP they actually have, or cannot import their legacy password hashes has a beautiful p99.9 and no deal. Predictability is a multiplier on usefulness, and multiplying zero is still zero. The teams I've watched lose on this weren't disciplined; they were using discipline as a reason not to do hard integration work.

So where does flexibility genuinely earn its variance? The rule that has held up for me is about volume and recoverability, not about principle:

Surface Volume Failure recoverable? Flexibility verdict
Control plane / admin APIs Low Yes — retry, fix, re-run Be as flexible as you like
Provisioning, SCIM, sync jobs Low, batched, off critical path Yes Flexible; this is where enrichment belongs
Enrollment and recovery flows Low per user, high stakes Partially Flexible; genuine product differentiation lives here
Migration and coexistence paths Bounded to a transition Yes Flexible, with a stated end date
Authorization policy evaluation High, but pre-computable Depends Flexible if moved off the request path
Token issuance and validation Extreme fan-in No Rigid. This is where the budget is spent

The interesting consequence is that "no flexibility" was never the right position. The right position is that flexibility is cheap where volume is low and failures are recoverable, and ruinous on the fan-in path — so you locate it deliberately instead of letting it accrete wherever a deal pushes it. A platform with a wildly configurable control plane and an unyielding runtime is not a contradiction. It's the shape the constraint produces, and it's also, not coincidentally, the shape most mature identity systems converge on from very different starting points.

Two more places boring genuinely fails, stated plainly because they're where I'd argue against myself. Security requirements change under you: an authentication path frozen in 2019 is not predictable, it's obsolete, and "we don't ship changes" is a bad answer to a new attack class. And a determinism requirement, taken absolutely, forbids adaptive authentication — risk signals are by construction load- and history-dependent. The reconciliation there is to keep the decision deterministic given the risk score, and treat score computation as a separate, versioned, replayable input. It's more work than either extreme, and it's the honest middle.

What the platform engineer was actually asking

Go back to the bake-off. The question wasn't about milliseconds. He was asking whether the vendor could make a statement about their system that would still be true after he configured it, after they onboarded the tenant next to him, and after their next release.

The first vendor couldn't, and the reason wasn't incompetence. It was four hundred rows of feature matrix, each of which had been someone's reasonable yes, and which collectively meant that no sentence beginning "our system always..." could be completed truthfully.

That's the trade in one line. Every configurable branch buys a deal and spends a claim. You can run out of claims long before you run out of deals, and the day you notice is the day someone asks you a simple question in a meeting and you hear yourself say it depends.