The Cost of Token Introspection vs Self-Validating Tokens

A platform team I know lost an argument to a spreadsheet.

Their FinOps review had flagged a line item: an introspection tier — six instances behind a load balancer, plus a three-node Redis cluster — costing about $900 a month. It appeared under the identity team's cost centre, it had a name, and it was the only thing on the page that could be described in one sentence as "a service we run so that other services can ask a question." The recommendation wrote itself: move to self-validating JWTs, delete the tier, save $10,800 a year.

They did it. It took a quarter, across nine teams and four languages. And the following FinOps review showed the identity cost centre down $900 a month and the platform's total cloud bill up, by an amount nobody could initially attribute, spread across forty services in increments too small for any single team to notice or care about.

The interesting part isn't that they were wrong. It's why the spreadsheet couldn't see it. Introspection's cost is centralized, metered, and named. A JWT's cost is distributed, unbilled, and lands in other people's budgets — three percent more CPU here, a slightly fatter header there, a quarter of an engineer maintaining validation middleware, an audit pipeline that grew for reasons the audit team attributed to growth. Both designs cost money. Only one of them files an expense report.

This is the economics companion to two arguments I've already made elsewhere and don't want to re-run. Why Token Introspection Isn't As Slow As You Think does the latency comparison — the short version is that a well-cached introspection tier adds less mean latency than RS256 verification, which surprises people. JWT vs Opaque Tokens does the operational comparison — token size, debuggability, key rotation, who's on call. Assume both. What's left is the money, and the money produces a conclusion that runs opposite to the folklore.

The folklore says: start with introspection, move to JWTs when you scale. The arithmetic says the reverse. You do not outgrow introspection by getting bigger. You outgrow it by getting further apart.

Let's build the model.

The assumptions, stated up front so you can re-run it

Every number below is derived from these. Substitute yours; the structure is what matters, and the structure is stable even when the prices aren't.

Parameter Value Note
User-facing API requests 1,000/s peak, 400/s average ~1.04B/month
Internal fan-out per request 12 service hops Typical for a decomposed B2B SaaS
Services / pods 40 services, ~200 pods, 3 AZs Polyglot: JVM, Node, Go
Concurrent sessions 500,000 Drives refresh traffic
JWT size ~1.5 KB sub, aud, tenant, roles, exp
Opaque token size 30 bytes
Introspection cache hit rate 99% at 60s TTL See below for where this comes from
vCPU cost $25/vCPU-month Blended on-demand/committed
Cross-AZ transfer $0.01/GB per direction AWS-shaped; GCP and Azure differ
Log ingestion + indexing $0.50/GB Mid-range negotiated rate
Fully-loaded engineer $16,700/month $200k/yr
Target CPU utilization 40% Because queueing

Two of those deserve a defence.

The 99% hit rate isn't optimism. A token is presented repeatedly by the same session — 50 to 200 times over its life is normal for an API-backed UI. Cache the introspection result for 60 seconds and the miss rate is roughly (number of distinct tokens seen per 60s window) ÷ (requests per 60s window). At 400 rps with sessions making a request every few seconds, that lands between 0.5% and 2%. If your traffic is one-request-per-token — a webhook receiver, a batch integration — your hit rate is zero and none of this analysis applies to you. Measure it before you borrow my number.

The 40% utilization target is the one people forget when they price CPU. You cannot buy 2.4 cores of work and provision 2.4 cores; queue wait goes hyperbolic as utilization approaches one. Everything below multiplies busy-CPU by 2.5 to get provisioned CPU, and that factor is why "it's only 60 microseconds" is a misleading way to think about verification cost.

Line item one: the CPU, which is not where you think it is

Start with the number everybody cites. RS256 signature verification: 60µs. ES256: 25µs. EdDSA: faster still.

At 12,000 internal validations per second (1,000 user requests × 12 hops), RS256 verification is 0.72 CPU-seconds per second — under a core busy, about 1.8 cores provisioned, roughly $45/month. That is nothing. If the argument were only about signature math, it would be over, and JWTs would win it walking.

The argument isn't about signature math. Here's the part that doesn't appear in the benchmarks people quote, because the benchmarks measure verify(bytes, key) and production doesn't do that.

A production JWT validation is: split on dots, base64url-decode two segments, allocate and parse two JSON documents, verify the signature, then validate exp, nbf, iss, aud, and alg against expectations, then map claims into whatever object your framework hands the application. In a JVM or Node service, the crypto is frequently the cheapest step. Measured end-to-end with a realistic 1.5 KB token, per-validation cost typically lands between 120µs and 250µs, and the delta over the raw crypto figure is parsing and allocation.

Take 200µs:

12,000 validations/s × 200µs = 2.4 CPU-seconds/s busy
  ÷ 0.40 target utilization    = 6 vCPU provisioned
  × $25/vCPU-month             = $150/month

Still small in absolute terms. But there are two things hiding in it.

First, it's 3.3× the number the benchmark told you, which matters when someone builds a capacity model from a microbenchmark. Second, and worse: that allocation is garbage. Two JSON documents plus a claims object per validation, twelve times per user request, is a steady stream of short-lived objects into the young generation of every service in your fleet. It doesn't show up as CPU cost; it shows up as GC pause frequency, which shows up in p99, which shows up as a latency SLO nobody can explain. I've seen a team chase a 12ms p99 regression across a service mesh for three weeks and find it in a JWT library that parsed the payload twice — once to read kid, once after verification.

The mitigation is real and worth knowing: cache the parsed validation result keyed by a hash of the raw token string, with a TTL bounded by exp. Hit rates are as good as an introspection cache's, for exactly the same reason. Which is a slightly awkward observation, because at that point your self-validating token has a validation cache with a TTL and a revocation window, and the architectural difference between the two designs has narrowed to where the cache miss goes.

Now the other side. Introspection at 1,000 user requests/s, resolved once at the gateway rather than at every hop, with 99% hits: 10 introspection calls/second reaching the authorization server. At ~1.5ms of AS CPU each, that is 0.015 CPU-seconds per second. Not 0.15. Fifteen milliseconds of CPU per wall-clock second.

The compute cost of introspection at this scale is a rounding error. That is not the cost of introspection, and pretending it is has confused this debate for a decade.

Line item two: the fan-out asymmetry nobody prices

Here is the structural difference between the two designs, and it's the one that produces the crossover.

flowchart LR
    C["Client"] -->|"1 token"| G["Gateway"]
    G -->|"1 resolution<br/>(99% from cache)"| AS["Authorization Server"]
    G --> S1["Service A"]
    S1 --> S2["Service B"]
    S2 --> S3["…12 hops…"]
    S3 --> S4["Service L"]
    subgraph JWT["Self-validating: cost ×12"]
        S1
        S2
        S3
        S4
    end

A JWT is validated once per hop. A token is resolved once per request. Introspection's work is proportional to user requests; self-validation's work is proportional to user requests × fan-out. Every service you extract from a monolith adds a hop, and every hop adds a validation, and nobody proposing the extraction is thinking about token verification.

That's the CPU. It's also the bytes, and the bytes turn out to be bigger than the CPU.

The Authorization header travels on every hop. A 1.5 KB JWT versus a 30-byte opaque token is ~1.47 KB of extra header per hop:

4,800 hops/s (average) × 1.47 KB = 7.1 MB/s
                                  = 18.3 TB/month
× ~2/3 crossing an AZ boundary   = 12.2 TB
× $0.02/GB (both directions)     = $244/month

$244 a month to move token bytes that carry no user data. That's larger than the verification CPU line item, it's invisible in every cost dashboard I've seen (it aggregates into "inter-AZ transfer," which nobody attributes), and it scales linearly in both request rate and fan-out. At ten times the scale it's $2,400/month; with a 30-hop fan-out it's more.

Two honest caveats, because this number is the one most likely to be wrong for your deployment.

HTTP/2 HPACK changes it. On a persistent HTTP/2 connection, a repeated header value gets indexed in the dynamic table and subsequent sends cost a byte or two. If a single connection carried one user's token repeatedly, the JWT would be nearly free on the wire. In a service mesh, connections are pooled and shared across all users, so the dynamic table thrashes between thousands of distinct tokens and the indexing mostly doesn't land. The effect is real but partial; if your mesh is HTTP/2 end-to-end with modest per-connection user cardinality, discount this line item by whatever your hpack metrics tell you. If you're on HTTP/1.1 anywhere in the chain — and most people are, somewhere — it's the full number.

Zonal-aware routing changes it too. If your mesh prefers same-zone endpoints, fewer hops cross a billing boundary. That's a real saving and also a real availability trade-off, and it isn't free either.

Introspection's wire cost, for comparison: 10 calls/s × ~2 KB request+response = 20 KB/s ≈ 50 GB/month ≈ $1/month. The opaque token itself, at 30 bytes across 12 hops, is noise.

Line item three: what the HA floor actually costs

So where is introspection's cost? It's the floor.

You cannot run an introspection tier at 0.015 cores. You run it at the minimum footprint that survives an AZ loss and a rolling deploy: call it six instances at 4 vCPU across three AZs, plus a three-node cache tier. That's 24 vCPU ($600/month) plus ~$300/month of Redis. $900/month, and it is the same $900 whether you serve 10 introspections per second or 1,000.

That's the whole shape of the thing. Introspection is a high-fixed-cost, near-zero-marginal-cost design. Self-validation is a near-zero-fixed-cost, positive-marginal-cost design. Any two cost curves shaped like that cross somewhere, and the crossing point is computable.

Per million user-facing requests:

Self-validating JWT Opaque + cached introspection
Validation CPU 12M × 200µs ÷ 0.4 = 1.67 CPU-hr → $0.057 10k misses × 1.5ms = 15 CPU-s → $0.0004
Wire bytes 17.6 GB, ⅔ cross-AZ → $0.235 20 MB → $0.0004
Marginal total ~$0.29 ~$0.001
Fixed infra Denylist cache ~$300/mo Introspection tier ~$900/mo

Set them equal. The introspection design carries $600/month more fixed cost; the JWT design carries $0.29 more per million requests:

$600 ÷ $0.29 per million = ~2,070 million requests/month
                         = ~800 requests/second sustained

Above roughly 800 sustained user-facing requests per second, at a fan-out of 12, the introspection tier is the cheaper design. Below it, JWTs are. Our scenario averages 400 rps, so this particular company should keep its JWTs — on infrastructure cost alone.

Run that back through the folklore and it inverts. "We'll start with introspection and switch to JWTs when we scale" has the curve backwards: introspection's cost is the part that doesn't grow. Scale is the axis on which introspection improves. What actually pushes you toward self-validating tokens is not volume — it's fan-out, geography, and organizational distance, all of which raise the marginal cost of a round trip rather than the marginal cost of a token.

The crossover is also very sensitive to fan-out, in a direction people don't expect. Halve the fan-out to 6 hops and JWT's marginal cost halves, pushing the crossover out to ~1,600 rps. Double it to 24 — a mesh with sidecars that validate independently, which is common — and the crossover drops to ~400 rps, which is below our scenario's average. Microservice decomposition is a tax on self-validating tokens, and nobody puts it in the migration doc.

Line item four: the revocation window has an exchange rate

This is the part I most want to put a number on, because it's the line item that actually decides the architecture and I've never seen anyone price it.

Both designs give you a window during which a revoked token still works. For a JWT it's the remaining token lifetime. For cached introspection it's the remaining cache TTL. Same property; I made that argument in the latency piece and won't repeat it.

The economics question is different: what does it cost to make that window one second shorter?

With JWTs, you shorten the window by shortening token lifetime, and you pay for it on the refresh path.

500,000 concurrent sessions with a 15-minute access token lifetime generate 500,000 ÷ 900s = 555 refreshes/second. Move to a 5-minute lifetime and that becomes 1,667/s — an extra 1,112 refreshes per second, sustained, forever.

Each refresh is not a cheap read. It's a token-endpoint request: authenticate the client, look up the refresh token, rotate it (a write), issue and sign a new access token, and emit audit events. Say 3ms of CPU and, critically, 4 audit events at ~800 bytes.

CPU:    1,112/s × 3ms ÷ 0.4     = 8.3 vCPU     →   $208/month
Writes: 1,112/s sustained against your refresh token store — on the
        one path in your system that is genuinely write-heavy
Audit:  1,112/s × 4 events × 800B = 3.6 MB/s
                                  = 9.3 TB/month
        × $0.50/GB ingested       →  $4,650/month

Shortening a JWT's lifetime bills you through your log vendor. That's the finding. The compute is $208; the audit pipeline is twenty times that, for the same decision, and it lands in a budget owned by a team that will attribute it to growth. This is the exact failure mode from The Hidden Cost of Audit Logs: high-volume, low-forensic-value events getting critical-tier treatment. Token refresh is the single best example of an event class where the aggregate count matters and the individual records almost never do.

The writes deserve their own sentence. Identity is mostly read traffic, and the refresh path is one of the few places it isn't. Tripling the write rate against your most stateful component to buy a shorter revocation window is a strange trade to make silently, and it's what "just use short-lived tokens" means operationally.

With introspection, you shorten the window by shortening the cache TTL, and you pay for it on a read path.

Drop the TTL from 60s to 10s. The hit rate falls — roughly 99% to 94%, since a smaller window catches fewer repeat presentations of the same token — so misses go from 10/s to about 60/s. That's 50 extra introspection calls per second: 0.075 CPU-seconds/s busy, under a fifth of a provisioned core, plus a little cross-AZ traffic.

Extra cost: ~$5/month. No writes. No audit amplification.

Now put them side by side as an exchange rate — dollars per month, per second of revocation window removed:

Design Change Window Monthly cost of the change $/second-of-window removed
JWT 15 min → 5 min 900s → 300s ~$4,860 $8.10
JWT 5 min → 1 min 300s → 60s ~$19,000 (refreshes ×5 again) $79
Introspection 60s → 10s 60s → 10s ~$5 $0.10
Introspection 10s → 0s (live) 10s → ~2ms ~$400 (tier scales to ~1,000 QPS) $40

Two things fall out of that table.

The floor is different, not just the price. JWTs get expensive before they get short. A one-minute JWT lifetime is an aggressive, operationally painful configuration that costs five figures a month in refresh amplification and still leaves a 60-second window. A ten-second introspection TTL is a config change that costs a rounding error. The designs are not competing on the same segment of the curve.

Introspection's expensive move is the one you only make during an incident. Setting the TTL to zero — live introspection on every request — costs real money because the tier has to scale to full request rate. But you make that move for an hour, during a token-leak incident, and then you set it back. It's a spot price, not a run rate. With JWTs there is no equivalent purchase at any price: tokens already issued are valid until they expire, and the only lever is a denylist, which is an introspection cache built by hand for the emergency path.

Which is the honest way to state the JWT revocation cost: you either pay the refresh amplification continuously, or you build denylist infrastructure that costs about what an introspection tier costs and gets exercised twice a year. A cache whose correctness only matters during incidents is a cache that is broken during incidents.

And push invalidation, when you have an event bus, changes the introspection column again — long TTL, event-driven eviction, revocation in tens of milliseconds at lower steady-state cost than a short TTL. That's mechanism, covered in the other article. Economically it just means the introspection column in that table is a ceiling, not a floor.

Line item five: the multi-tenancy multiplier

One cost that only appears in multi-tenant platforms, and it appears on the JWT side.

If you issue per-tenant signing keys under per-tenant issuers — which ClavionX does, and which is the right call for tenant isolation, since one compromised signing key shouldn't be able to mint tokens for every customer on the platform — then every validator now maintains a JWKS cache per tenant, not per platform.

At 300 tenants that's 300 cached key sets in each of 200 pods. The memory is trivial. The behaviour is not:

  • Cold-start cost scales with tenant count. A restarting pod that serves traffic for 300 tenants faces up to 300 cold JWKS fetches, not one.
  • Unknown-kid handling becomes load-bearing. Any tenant's rotation produces unknown-kid misses across the entire fleet. With 300 tenants rotating on independent schedules, "rare event" becomes "continuous background rate," and a validator without negative caching and rate limiting on that path is a DoS amplifier pointed at your own JWKS endpoint — a failure I described in JWT vs Opaque Tokens and which gets strictly worse with tenant count.
  • Rotation cost is superlinear. Rotating keys across 300 tenants × 200 validators is 60,000 cache-coherence events, and the failure mode of getting it wrong is every request for that tenant fails, not a revocation is slightly delayed.

The introspection endpoint, by contrast, is one URL with one credential regardless of tenant count. The tenant multiplier lands entirely on the self-validating side, and it lands on availability risk rather than on the invoice — which is precisely why it doesn't get counted.

Worth noting the architectural reason this is a fair comparison: in a design with strict runtime/control-plane separation, where the runtime never makes synchronous calls to the control plane and works only from projected state, an introspection call is a runtime-local read against projected data. It's not a call into the slow, transactional, config-owning half of the system. If your introspection endpoint reads from the same database your admin console writes to, your numbers will be much worse than mine, and that's a data-flow problem rather than a token-format problem.

Putting the whole bill together

At our scenario's scale — 400 rps average, fan-out 12, 500k sessions, 15-minute tokens:

Line item Self-validating JWT Opaque + cached introspection
Validation CPU $150 ~$1 (AS side)
Wire bytes (cross-AZ) $244 ~$1
Introspection / denylist tier $300 (denylist, mostly idle) $900
Refresh amplification (audit + CPU) baseline at 15 min baseline
JWKS fetch traffic ~$5, spiky at rotation $0
Infrastructure subtotal ~$700/month ~$900/month
Validation middleware, 4 languages ~0.25 FTE = $4,175/month ~0.05 FTE = $835/month
Cache/denylist correctness ownership distributed across 9 teams 1 team, 1 component
Total ~$4,900/month ~$1,750/month

The infrastructure lines are within $200 of each other. The payroll line is 5× apart, and it dominates both.

That's the actual finding, and it echoes the conclusion of What Does One Login Actually Cost You?: the infrastructure cost of an identity decision is usually the part you can measure and the part that matters least. A quarter of an engineer maintaining JWT validation middleware across four languages — bounded cache TTL, unknown-kid handling with rate limiting, hard timeouts, stale-while-revalidate, startup that doesn't hard-fail on an unreachable JWKS, in Java and Node and Go and Python — costs more than every vCPU and every byte in the table above, combined, several times over.

And it's the line item the FinOps spreadsheet structurally cannot contain. Which brings us back to the team at the top.

The counting error, named

Here's the general form, because it isn't really about tokens.

Centralized costs are legible. Distributed costs are not. An introspection tier is a named component in one budget, and a named component in one budget is a target. The same total dollars, spread as a 3% CPU increase across 200 pods owned by nine teams, is literally invisible — below the noise floor of any individual team's capacity planning, attributed to feature growth, never aggregated by anyone because no one owns the aggregate.

This is not a cloud-billing quirk. It's the reason architectural decisions get made badly in large organizations generally: the design whose cost is concentrated looks expensive, and the design whose cost is diffuse looks free, even when the diffuse one costs more. The introspection tier's greatest weakness as a design is that you can see it.

So the arithmetic's real use isn't picking a format. It's arming you against a specific class of bad argument: "we can delete this by going stateless." You can't delete the cost. You can only relocate it somewhere nobody is measuring — and relocating a cost into a place where it's unmeasured is how it grows.

When each one actually wins

Since the infrastructure numbers land within a few hundred dollars of each other at realistic scale, cost should almost never be the deciding input. Here's what should be, with the cost model as a sanity check rather than the argument:

Self-validating tokens win when a round trip is expensive or impossible. Validators outside your organization — requiring a partner to call your introspection endpoint on their hot path couples their availability to yours across a company boundary, and no cache TTL fixes that. Validators in distant regions, where a cache miss costs 250ms instead of 2ms; the marginal cost of a miss is what the whole model turns on, and geography multiplies it by a hundred. Very high fan-out with no gateway. Environments where the identity service's availability genuinely must not be in the request path.

Introspection wins when the round trip is cheap and the revocation window is a product requirement. Same-cluster validators. A gateway you already run. Regulated contexts where "we can revoke in 15 minutes" is not an acceptable answer to an auditor, and where the alternative — one-minute JWTs — costs five figures a month in refresh amplification. Multi-tenant platforms with per-tenant signing keys, where the JWKS coherence problem multiplies by tenant count. Sustained request rates above the crossover with modest fan-out.

And the hybrid wins more often than either, which is why most mature systems run it: opaque at the edge, resolved once at the gateway, JWT internally. The phantom token pattern is not a compromise; on this model it's the cost-optimal shape, because it pays introspection's cost once per request rather than once per hop, and pays JWT's marginal cost only inside a trust boundary where the tokens can be small and short-lived. Its bill is the gateway, which most teams are running anyway.

What to measure

If you want your own version of this model, four numbers and an afternoon:

  1. Your cache hit rate at your intended TTL. Everything in the introspection column is downstream of this. If it's below 90%, your traffic pattern isn't the one I assumed.
  2. Your real per-validation cost, end to end, not the crypto benchmark. Time the full parse-and-validate path in your actual runtime with your actual token size. If it's 3× the number in your library's README, that's normal, and your capacity model is currently wrong by that factor.
  3. Your fan-out. Validations per user-facing request. This is the multiplier on every JWT line item and the one most likely to have doubled since anyone last looked.
  4. Your refresh rate, and the audit events per refresh. Multiply. That product is the price tag on every future conversation about shortening token lifetimes, and it is usually the largest single number in this entire analysis.

Then, before anyone proposes deleting a tier to save money: work out where that cost is going instead, and whether the budget it lands in has anyone watching it.

The tokens were never the expensive part. The revocation guarantee is the expensive part, and you buy it either way — with introspection you buy it in metered infrastructure you can see, and with JWTs you buy it in refresh amplification, audit volume, and payroll you can't.