Why Token Introspection Isn't As Slow As You Think

The argument against opaque tokens takes about four seconds to make and it's the reason most teams never seriously consider them: "A network call on every request? No thanks."

The argument against opaque tokens takes about four seconds to make and it's the reason most teams never seriously consider them: "A network call on every request? No thanks."

It's a good instinct. It's also based on a comparison nobody actually runs. The implied alternative — JWT verification — is treated as free, and introspection is priced at a full uncached round trip on 100% of requests. Neither number is real.

So let's do the arithmetic properly, because when you cost both options honestly, the latency gap is usually somewhere between 0.05ms and 0.5ms, which is smaller than the variance in your database connection pool and far smaller than the engineering cost of the revocation machinery you'll build to compensate for JWTs.

What a naive introspection call actually costs

Start with the worst case, no optimizations. A resource server receives a request, calls the authorization server's /introspect endpoint, and waits.

Within a single datacenter or Kubernetes cluster:

  • TCP + TLS handshake: 1–3ms — but only on a cold connection. With keep-alive and a connection pool, amortized to roughly zero.
  • Network round trip, same AZ: 0.2–0.5ms. Cross-AZ: 0.5–2ms.
  • Authorization server request handling: 0.5–2ms if the token record is in its own cache; 2–10ms if it hits a database.
  • Response parsing: negligible.

So a warm, pooled, same-AZ introspection call against a cache-backed authorization server lands around 1–3ms at p50. Not free. Also not the 50ms that people seem to imagine when they dismiss it — that figure comes from imagining an uncached HTTPS call across the internet, which is what you get if you deploy it carelessly, and which is also what you'd get from a carelessly deployed JWKS fetch.

Now optimize, in the order of payoff.

The cache is the whole argument

Introspection results are cacheable, and the cache key is the token itself. This is the step that changes the economics, and it's remarkable how often the comparison is made without it.

A typical API session makes 50–200 requests with the same access token. Cache the introspection result for even 30 seconds and the hit rate for a normal traffic pattern is 95–99.5%.

Run the numbers on a 99% hit rate:

  • Cache hit (99%): in-process lookup, ~0.01ms
  • Cache miss (1%): pooled call, ~2ms
  • Mean added latency: 0.03ms. p99: ~0.01ms (the p99 request is a cache hit). The misses show up at p99.9.

At 95%:

  • Mean: 0.11ms. Still under a tenth of a millisecond.

For comparison, a single RS256 JWT signature verification costs 30–90µs of CPU (0.03–0.09ms) — every request, no caching possible, because verification is the whole point. Which produces an uncomfortable observation for the JWT camp: a well-cached introspection tier has a lower mean added latency than RS256 verification. Not dramatically lower. But the direction is the opposite of the received wisdom.

EdDSA and ES256 verification are faster than RS256 (roughly 15–40µs), which narrows it, and if you're doing high-volume JWT verification you should be using them for exactly this reason. The point stands: these two options are in the same order of magnitude, and both are noise against a 40ms database query.

Two caching details that matter:

Cache the negative results too, briefly. An invalid or expired token presented repeatedly — which happens constantly, from clients that haven't noticed their token expired — should not generate an introspection call each time. Five seconds of negative caching removes a surprising amount of traffic.

Cache in-process, then in Redis, then call. A two-tier cache means the in-process tier absorbs the repeated-token case with zero network cost, and the shared tier absorbs the cold-start-per-pod case. A pod that restarts doesn't stampede the authorization server; it stampedes Redis, which is a much better outcome. Add jitter to TTLs so that a cohort of tokens cached at the same moment doesn't expire simultaneously.

Cached introspection and short-lived JWTs are the same trade — but one is adjustable

Here's the reframe that makes the choice clearer, and it's the part of the argument that usually flips people.

If you cache introspection for 60 seconds, you have a 60-second revocation window. A revoked token continues to be accepted until the cache entry expires. That is exactly the same property as a 60-second JWT.

The difference is that the cache TTL is a dial you control at runtime, per-service, and the JWT lifetime is a property baked into tokens already issued.

That difference is worth more than it sounds:

  • During an incident, you can set the TTL to zero. Every request becomes a live introspection call — slower, more load on the authorization server, and correct. You cannot do this with JWTs already in the wild; you can only wait.
  • You can set different TTLs per endpoint class. Zero for POST /transfers, 60 seconds for GET /profile. With JWTs, one token lifetime covers every path.
  • Reducing the window doesn't increase token endpoint traffic. Halving a JWT's lifetime doubles refresh traffic against your most stateful component; halving a cache TTL only doubles introspection traffic, which is a cheap, cacheable, horizontally-scalable read.

That last point is the one that gets missed. The standard JWT answer to revocation latency — shorten the token lifetime — pushes load onto the refresh path, which touches the refresh token store and does mutation. The introspection answer pushes load onto a read path with no mutation. Read paths scale better than write paths, so as a mechanism for buying a shorter revocation window, introspection is the cheaper instrument.

Then add push invalidation, and you get something JWTs can't do

If you already emit events for identity changes — and if you have an event bus, you do — then revocation can push rather than wait for a TTL to lapse.

The authorization server publishes a token.revoked or session.terminated event. Every resource server (or gateway) subscribes and evicts the matching cache entries.

flowchart LR
    AS["Authorization Server"] -->|"revocation event"| BUS["Event bus"]
    BUS --> C1["Gateway cache"]
    BUS --> C2["Service A cache"]
    BUS --> C3["Service B cache"]
    C1 -.->|"miss → introspect"| AS

Now your revocation window is event propagation time — typically tens to low hundreds of milliseconds — with a TTL as a safety net for missed events. You can set the TTL generously (5 minutes) because it's no longer the primary revocation mechanism, so your steady-state hit rate goes up and your load goes down while your revocation latency goes down by two orders of magnitude.

This combination — long TTL, push invalidation, TTL as backstop — is strictly better than what a short-lived JWT gives you, on both axes simultaneously. There is no equivalent for JWTs, because there is nothing to evict: the token is valid by virtue of its own signature and no participant in the system holds state you could invalidate.

The honest caveats: you need the event bus to be reasonably reliable (it doesn't need to be perfectly reliable, because the TTL covers gaps); you need per-token or per-session cache keys you can actually target; and a "revoke everything for this user" event needs an efficient eviction strategy, usually a secondary index from subject to cached tokens, or a per-subject epoch counter you bump.

The costs on the other side of the ledger

To make the comparison fair, price the things JWT deployments actually pay for and rarely count.

The JWKS fetch is also a network call, made by every validator, with a cache that has all the same failure modes as an introspection cache — plus one worse one: a stale JWKS after key rotation causes every validation to fail, whereas a stale introspection cache causes a slightly delayed revocation. The blast radius of the two caching bugs is not comparable.

Verification CPU is not free at volume. At 50,000 requests/second with RS256 at 60µs, that's 3 CPU-seconds per second — three cores, continuously, doing nothing but verifying signatures, distributed across your fleet. Not a crisis, but it's more than the cached-introspection path consumes.

Token size costs bandwidth on every hop. A 1.5KB JWT versus a 30-byte opaque token, across twelve internal service calls per user action, is ~18KB of header traffic per action, repeated at every hop. This occasionally crosses into real latency territory when it pushes requests past a congestion window.

Revocation machinery is engineering time. A denylist checked on sensitive paths is a distributed cache with correctness-critical invalidation — which is to say, it's an introspection cache, built by hand, only for the emergency path, usually with less care.

And the operational surface. Every validator needs correct algorithm pinning, kid handling, clock skew tolerance, and JWKS refresh behaviour. Introspection clients need a URL, credentials, a timeout, and a cache. The second list is shorter, and shorter lists have fewer things implemented wrong by teams who aren't thinking about it.

When introspection genuinely loses

This is not a claim that opaque tokens win everywhere. Three cases where they're the wrong answer, and they're decisive:

Validators outside your organization. Requiring a partner to call your introspection endpoint on their hot path couples their availability to yours across an organizational boundary. Signed tokens they can verify offline are the correct design. This is the strongest argument for JWTs and it's sufficient on its own.

Validators in distant regions. A resource server in Sydney introspecting against an authorization server in Frankfurt pays 250ms on every cache miss. You can mitigate with regional introspection replicas, but at that point you're running a distributed authorization data store, and self-validating tokens are simpler.

Extreme fan-out with no gateway. If a single user action touches 40 services and each independently validates, cache miss probability compounds and the authorization server becomes a hub with a very high request rate. Usually the right fix is to validate once at the edge, which is the phantom-token pattern — but if you can't, JWTs handle it better.

What to do

If you're choosing today, and your validators are inside your own infrastructure:

  1. Opaque tokens externally, cached introspection at a gateway. Small tokens, no PII in anything a client holds, instant revocation available at the edge.
  2. Two-tier cache — in-process, then shared — with jittered TTLs in the 30–300 second range, and brief negative caching.
  3. Push invalidation over your event bus, with the TTL as backstop rather than as the revocation mechanism.
  4. A runtime-adjustable TTL, per endpoint class, so an incident can set it to zero.
  5. Measure the hit rate and the p99 of the miss path. If the hit rate is below 90%, your TTL is too short or your traffic pattern isn't what you assumed.

And if you're currently dismissing introspection on latency grounds: measure it. Stand up an endpoint, add the cache, run your actual traffic pattern against it, and look at the number. The most common outcome is a mean added latency around a tenth of a millisecond and a revocation story you no longer have to apologize for in security reviews.

The received wisdom here is a real thing — a network call per request would be bad. It's just that nobody makes a network call per request, and the version everybody actually deploys costs about as much as verifying a signature.