Understanding Tail Latency in Authentication

A team I worked with had a dashboard they were proud of. Login p50: 84ms. Login p99: 210ms. Both well inside budget, both green for months.

A team I worked with had a dashboard they were proud of. Login p50: 84ms. Login p99: 210ms. Both well inside budget, both green for months.

They also had a steady trickle of tickets that said, in various phrasings, "login is slow sometimes." Nobody could reproduce it. The dashboard said it wasn't happening. The tickets kept arriving.

The dashboard was measuring the token endpoint. A login was seven sequential HTTP requests across three services and two redirects, and nobody was measuring the thing the user experienced. When they finally instrumented the journey end-to-end, the p99 was 2.4 seconds and the p99.9 was eleven seconds.

Neither number was a lie. They were measuring different things, and the difference between them is the whole subject of this article. Tail latency in authentication behaves unlike tail latency almost anywhere else in your system, for two structural reasons: authentication is serial — a login is a chain of dependent hops where any one stall delays everything — and it is fan-in — every page load, every API call, every service-to-service request touches it. Those two properties amplify a rare slow event into a common user experience, and the amplification is arithmetic you can do on a napkin.

Most teams have never done it. So let's.

Why "1 in 1000" is not rare

The instinct that p99.9 is an edge case comes from thinking about a single request. Authentication is never a single request.

Serial amplification. Take an OIDC authorization code login: the app redirects to the authorization server, the AS renders a login page, the credential is posted, the AS redirects back with a code, the app exchanges the code for tokens, the app fetches userinfo or validates against JWKS, the app sets its session and redirects to the landing page. That's six or seven sequential round trips, each with its own latency distribution.

If each hop independently has a 1% chance of exceeding its p99, the probability that a login has no p99-class event is 0.99⁷ ≈ 93.2%. So 6.8% of logins contain at least one p99 event. The p99 of the journey is not the p99 of any hop; it's determined by the p93-ish behaviour of the whole chain, and the journey's p99 is closer to the worst hop's p99.9.

Turn that around and it's more useful: to get a 200ms p99 across seven serial hops, each hop needs roughly a p99.85. Per-hop p99s that look excellent produce a journey that doesn't.

Fan-in amplification. Now consider a single-page app that validates or refreshes on API calls. A typical session makes 200 authenticated requests. If token validation has a 1-in-1000 chance of taking 2 seconds, then the chance a session sees zero such stalls is 0.999²⁰⁰ ≈ 81.9%. One in five sessions contains a two-second stall. At 1-in-10,000 it's still 2% of sessions.

Over time. A user logging in twice a day, 250 working days a year, makes 500 logins annually. A 1-in-1000 event hits them, on average, once every two years — but across 10,000 employees that's 5,000 slow logins a year, roughly 14 a day, which is precisely the volume that produces a persistent trickle of unreproducible tickets and a reputation for being slow.

This is the reframe that matters: the tail is not the exception, it is the aggregate experience. Users don't experience your distribution one sample at a time; they experience the maximum over their session, and the maximum of many samples from a fat-tailed distribution is much worse than the median.

Percentiles do not compose, and most dashboards assume they do

Three arithmetic errors are near-universal, and each one produces a dashboard that shows green during an outage.

You cannot average percentiles. If service A has p99 = 100ms in eu-west and p99 = 400ms in us-east, the combined p99 is not 250ms. It depends on the traffic mix and the shapes of both distributions, and it can legally be anything from 100ms to 400ms. Yet every monitoring tool will happily draw you the mean of per-instance p99s, and if one instance out of twenty is sick, its contribution to that mean is 5% — invisible — while its contribution to your users' experience is 5% of all requests being slow.

The fix is to compute percentiles from merged raw data or from mergeable sketches (histograms, t-digest, HDR histograms). Prometheus histograms give you this correctly if you aggregate the buckets and then compute the quantile; avg(rate(...)) over pre-computed quantiles does not. This is the single most common instrumentation bug in this space.

You cannot add percentiles across hops. Summing seven p99s gives you a number that essentially never occurs, because it requires all seven hops to be simultaneously at their worst. Sum the p50s to estimate a typical journey; for the tail, measure the journey directly, end to end, with a correlation ID that spans every hop. If you can only afford to fix one thing after reading this article, make it this: a single trace-level histogram of the complete login journey, from the user's first request to the landing page. Everything else is a proxy.

Per-endpoint percentiles hide per-tenant tails. Your p99 is a population statistic. If one tenant with 300,000 group memberships takes 4 seconds to authenticate, and they are 0.2% of your traffic, they are entirely inside your p99 and 100% of their users are unhappy. Percentile by tenant, or at minimum alert on per-tenant p95 deviation from the global p95. The customer most likely to leave is the one your dashboards cannot see.

Where authentication tails actually come from

The distribution of causes is different from a typical CRUD service, and knowing the list saves you a lot of profiling.

Deliberate CPU burn. You are running Argon2 or bcrypt on purpose, at a cost factor chosen to be expensive. That's 100–250ms of pure CPU per password verification, and it's not a bug. What makes it a tail problem is that it's CPU-bound work on a shared machine: when concurrency exceeds cores, requests queue, and queueing latency grows superlinearly. A box that handles password verification comfortably at 40% CPU falls off a cliff at 75%, because there is no elasticity left — every additional request waits for a core rather than for I/O. This is the most common cause of a login p99 that's fine all week and terrible at 09:00 on Monday.

Two consequences worth internalizing. First, hash verification should be isolated — its own pool, its own instances, its own capacity plan — so it cannot steal cores from cheap operations like token validation. Second, your cost factor is a latency and capacity decision, not just a security one, and it needs to be revisited when hardware changes rather than set once in 2019.

Garbage collection, made worse by the above. Allocating heavily in a JVM or Go service produces pauses. GC pauses correlate with load, and load correlates with everything else being slow, so GC contributes precisely when you can least afford it. In a serial chain, a 300ms pause in any one of seven services is a 300ms pause for the user.

Connection pool exhaustion. This is the classic identity tail and it looks nothing like a database problem. Your pool has 20 connections; queries take 5ms; you can serve 4,000 queries/second. At 3,500/s you're fine. At 4,200/s, requests wait for a connection, and wait time is unbounded while query time stays at 5ms. Your database dashboard shows healthy 5ms queries throughout. Your users are waiting 3 seconds. Always instrument pool acquisition time separately from query time — it is a different failure with a different fix, and conflating them sends teams optimizing queries that were never slow.

Cold cache on a long tail of tenants. Your cache hit ratio is 98%, which sounds excellent, and the 2% is a full config load, a signing key fetch, and a policy compile — maybe 400ms. That 2% is not randomly distributed: it's concentrated in small or infrequently-active tenants, so for those customers the typical experience is the cache miss. Warm caches at deploy time for the tenants you know are active, and measure hit ratio per tenant, not globally.

JWKS and key material fetches. A validator that hits an unknown kid fetches the JWKS synchronously. If that happens on a user-facing request during a key rotation, every service does it at roughly the same moment, and your identity service receives a thundering herd on a path nobody load-tested. Prefetch on rotation, cache with jitter, and never make the fetch blocking without a timeout.

Upstream identity sources you don't control. LDAP and Active Directory bind latency, an on-premises SAML IdP behind a customer's VPN, an SMS or push-notification provider. These have tails measured in seconds and no SLA you can enforce. The important design property is that their tails must not become your tails synchronously: bounded timeouts, circuit breakers, and — where the semantics allow — a cached answer that's a few seconds stale in preference to a fresh answer that takes four seconds.

TLS handshakes and connection setup. A full handshake is 2 round trips; over a trans-Atlantic path that's 200ms+ before any application work. Redirect-based flows are especially exposed because each redirect may target a different host, and a cold connection per host is a fresh handshake. Keep-alive, session resumption, and HTTP/2 connection reuse are worth real attention here — and if your login chain crosses regions more than once, the redirect topology itself is the latency budget.

Retries and timeouts, which usually make it worse. A client that retries after 1 second against a service that's slow because it's overloaded has just increased the load on an overloaded service. Retry amplification through a chain of N services is multiplicative: with 3 retries at each of 4 layers, one user request can become 81 backend requests. Retry budgets (cap retries at a percentage of total traffic), jittered backoff, and no retries on non-idempotent operations. In a serial chain, a retry also adds its full timeout to the user's latency — so a generous timeout with a retry is often worse for the user than a tight timeout with a clean failure.

Utilization is the control knob nobody wants to hear about

The queueing-theory result is unavoidable and worth stating plainly, because it's the thing that turns a capacity conversation into a latency conversation.

For a simple queue, expected wait scales as ρ/(1−ρ) where ρ is utilization. At 50% utilization, wait is about 1× service time. At 80%, 4×. At 90%, 9×. At 95%, 19×. The curve is a hockey stick, and the tail percentiles bend up long before the mean does.

Which means: you cannot run an authentication service at high utilization and have a good p99. If your p99 target is tight and your capacity plan targets 80% CPU, those two documents contradict each other and the capacity plan will win. Identity systems that feel fast are almost always running at 30–50% utilization, and that headroom is not waste — it is the mechanism by which the tail stays short. This is the trade to have explicitly with whoever owns the cloud bill, because "reduce the instance count by 30%" and "keep login under 200ms at p99" are the same decision viewed from two sides.

Two related notes. Variance in service time makes the curve worse — a mixed workload where 5% of requests take 200ms of Argon2 and 95% take 2ms has much worse queueing behaviour than a uniform one, which is the capacity-planning argument for isolating hash verification. And autoscaling does not save you, because scale-up takes 30–120 seconds and your tail event lasts 5 seconds. Autoscaling handles trends; headroom handles tails.

What to actually do

In the order I'd do it, with the highest-leverage first:

Measure the journey, not the endpoint. One end-to-end histogram of the full login, correlation-ID stitched, including client-side time from the browser if you can get it via the Navigation Timing API. Until this exists, every other number you have is about a component rather than a user.

Fix your percentile aggregation. Verify that your p99 is computed from merged histograms rather than averaged from per-instance quantiles. Then look at p99.9 as well — for a serial chain, per-hop p99.9 is what determines journey p99, so p99 alone under-measures the thing you care about.

Separate queue time from work time at every resource: connection pool acquisition, thread pool wait, hash-verification pool. These are the two failure modes and they have opposite fixes, and every dashboard that shows only total duration makes them indistinguishable.

Percentile by tenant. Global percentiles structurally hide your largest and smallest customers, who are exactly the two populations most likely to be having a bad time.

Isolate expensive CPU work into its own capacity so it can't queue behind or ahead of cheap work.

Set explicit timeouts everywhere, and budget them along the chain. If the journey budget is 1 second and there are seven hops, no single hop may have a 5-second timeout — that timeout is a promise that the whole journey can take 5 seconds. Timeouts should be derived from the budget downward, not chosen locally.

Consider hedged requests for idempotent reads. For an operation like a JWKS fetch or a policy read, issuing a second request after the p95 elapses and taking whichever returns first converts a long tail into a modest constant cost — typically 5% extra traffic for a dramatic p99.9 improvement. It works only for idempotent, cheap, read-only operations, so it's not a general tool, but on the specific paths where it applies it's the highest-return trick available.

Shed load rather than queueing it. An unbounded queue converts an overload into unbounded latency for everyone; a bounded queue with fast rejection converts it into a fast failure for a few. For authentication, a fast, honest error with a retry hint is a better user experience than a nine-second spinner, and it's much better for the system, because the nine-second spinner is what generates the retry that deepens the overload.

The point

Average latency describes a request. Tail latency describes a user, and in authentication the mapping between the two is unusually brutal: serial chains multiply the probability of hitting a tail, and fan-in multiplies the number of chances each user gets.

That's why identity teams get "it's slow sometimes" tickets that never reproduce, and why the dashboard is green while the tickets arrive. The dashboard is answering the wrong question. It's telling you about the middle of a distribution when what the user remembers is the worst moment of their week — and with a hundred chances a day, the worst moment is not rare at all.

Measure the maximum your users actually see. Then go look at your utilization target, because that's probably where the fix is.