JWT vs Opaque Tokens: The Operational Trade-off

Almost every article on this subject explains what a JWT is, explains what an opaque token is, notes that JWTs are stateless and opaque tokens are revocable, and stops. That comparison is accurate and nearly useless, because it describes a property you'll read about once and skips the six properties

Almost every article on this subject explains what a JWT is, explains what an opaque token is, notes that JWTs are stateless and opaque tokens are revocable, and stops. That comparison is accurate and nearly useless, because it describes a property you'll read about once and skips the six properties you will actually live with for the next four years.

I want to make a narrower argument, and deliberately not re-litigate revocation — the short version is that JWTs relocate state to the token endpoint rather than eliminating it, and that's a separate discussion. Assume you've had it. Assume you've accepted a bounded revocation window or built a denylist.

What's left is the part nobody writes down: the operational bill. Token size and where it gets truncated. What a 3am debugging session looks like. What key rotation costs. What happens when a partner needs to validate offline. Which team owns the failure when validation breaks. These are the things that determine whether you're happy with the choice in year three, and they don't appear in the comparison table.

The distinction, stated precisely

An opaque token is a reference. 2YotnFZFEjr1zCsicMWpAA means nothing; it's a lookup key into the issuer's storage. To find out what it authorizes, you ask the issuer — RFC 7662 token introspection, or a proprietary equivalent.

A JWT access token is a document. eyJhbGci... decodes to a JSON payload — subject, audience, scopes, expiry — with a signature over it. To find out what it authorizes, you verify the signature and read it. RFC 9068 standardizes what the claims should look like, and if you're issuing JWT access tokens and haven't read it, that's the highest-value thirty minutes available to you right now.

The distinction is about where the authorization decision's data lives at validation time: at the issuer, or in the client's hand. Everything else follows.

Size, and the places it fails silently

An opaque token is 20–40 bytes. A JWT access token starts around 400 bytes for a trivial payload and grows with every claim.

That growth is not linear in usefulness. Here's the progression I've watched happen at three different companies:

Ship with sub, aud, scope, exp — about 450 bytes, fine. Add tenant ID and a role — 500 bytes, fine. A downstream team needs the user's email and display name to avoid a lookup — 600 bytes. Someone puts group memberships in, because authorization needs them and a live lookup per request was too slow — and now token size is a function of how many groups your largest customer's users belong to. At one company that was 340 groups for a particular admin, producing an 8KB token.

Three things break at that size, in this order:

Your reverse proxy. nginx defaults large_client_header_buffers to 4KB per header. Exceed it and the client gets a 400 Bad Request with no useful body. Envoy defaults to a 60KB total header limit, AWS ALB to 16KB per header, and various API gateways sit between 8KB and 16KB. So the failure appears at your infrastructure boundary, not in your application, which means it doesn't show up in your application logs and doesn't reproduce in local development where you have three groups.

The cookie limit, if you're carrying the token in one — 4096 bytes per cookie, and browsers do not report the refusal. The Set-Cookie simply doesn't take effect.

Your bandwidth bill and your tail latency. An 8KB header on every request to every service, in a system where one user action fans out to twelve internal calls, is 96KB of header traffic per action. That's usually not a cost problem; it is occasionally a latency problem, because it can push a request past the initial TCP congestion window and add a round trip.

The pattern that makes this a genuine operational hazard is that it is customer-shaped. It works in dev, works in staging, works for 99.9% of production users, and fails for the largest customer's most senior administrator — the single worst person to be unable to log in. And it fails at an infrastructure layer that reports it as a malformed request.

Opaque tokens have no version of this problem. They are 30 bytes regardless of how many groups anyone belongs to. If you take one thing from this article: do not put unbounded-cardinality data in a token. Groups, permissions lists, and entitlements are unbounded. If authorization needs them, look them up, or put a hash or a versioned reference in the token and cache the expansion server-side.

Debugging: a genuine, underrated JWT win

This is where JWTs earn real affection from the people who operate systems, and it doesn't get mentioned enough because it isn't a security property.

A support ticket says "the API returns 403 for this user." With a JWT, you take the token from the request, paste it into a decoder — or cut -d. -f2 | base64 -d | jq — and read exactly what the authorization server asserted. Wrong tenant. scope missing orders.write. aud names a different service. You have your answer in fifteen seconds without credentials to any system.

With an opaque token you have a meaningless string. You need introspection access, which means client credentials for the introspection endpoint, which means either you have production auth credentials on your laptop (bad) or you're filing a ticket with the identity team (slow). At 3am during an incident, that difference is material. I've watched an outage extend by forty minutes because the only person who could introspect a token was asleep.

There is a real security cost to this transparency, and it's worth naming: a JWT leaks its contents to anyone who obtains it, and to anyone with access to the logs it landed in. If your token contains an email address, a user ID, employment status, or group names, that's PII in every log line, every APM trace, every browser history entry, and every Slack message where someone pasted a curl command. Opaque tokens are inert in a log — bad practice to log, but low-consequence. JWTs are a data-classification question.

The compromise that mostly works: keep identifiers in the token, keep attributes out. sub as an opaque UUID rather than an email address. Tenant as an ID rather than a company name. You keep most of the debugging benefit — you can tell which subject and tenant — and the token stops being a PII carrier.

Key rotation is the JWT operational burden

Opaque tokens have essentially no cryptographic operations to run. JWTs have a signing key, and a signing key means a rotation procedure, and a rotation procedure means an operational commitment that lasts as long as the system does.

The mechanics are well understood — publish a JWKS with a kid, add the new key, wait for validators to pick it up, start signing with the new key, wait for old tokens to expire, retire the old key. What surprises teams is that every validator is now a cache with a correctness requirement, and you have as many of those as you have services.

The failure modes I've actually seen:

A service that fetched the JWKS once at startup and cached it forever. Rotation happened, that service kept validating fine (old tokens), then started rejecting everything the moment new signatures appeared. It had been running for 200 days without a restart, which is normally something to be proud of.

A service that fetched the JWKS on every unknown kid and had no negative caching. An attacker — or, in the real case, a misconfigured client sending garbage tokens — could drive unbounded requests to the JWKS endpoint. That's a DoS amplifier pointed at your identity service, from your own fleet.

A service whose JWKS fetch had no timeout, in a runtime with a bounded thread pool. The identity service got slow, the JWKS fetch hung, the thread pool filled, and a service that was supposed to be independent of the identity service's availability took an outage anyway. The stateless-validation benefit evaporated at exactly the moment it was supposed to pay off.

None of these are hard to prevent. They require that every validator implement the same five behaviours correctly — bounded cache TTL, refresh on unknown kid with rate limiting and negative caching, a hard timeout, a stale-while-revalidate fallback, and startup that doesn't hard-fail on an unreachable JWKS. Doing that once is easy. Guaranteeing it across nineteen services in five languages, some maintained by teams who have never thought about it, is an organizational problem.

Which is the real point: JWT validation looks like a library concern and is actually a platform concern. If every team implements it, some will implement it wrong, and the failure will be attributed to your identity platform. Ship a validated, versioned middleware per language and treat it as a product. If you can't, that's a substantial argument for the gateway pattern below.

Multi-service validation: the case where JWTs are simply correct

Everything above is a cost. Here's the benefit, sized honestly.

An introspection call is a network round trip to the identity service. Inside one cluster, that's maybe 2–5ms at p50, and considerably worse at p99 — and p99 is what matters, because a user action fanning out to twelve services touches the tail repeatedly. It's also a hard dependency: your identity service's availability multiplies into every service's availability, and your identity service's request rate becomes the sum of all API traffic in the company.

JWT verification is an in-process signature check. Around 30µs for RS256, faster for EdDSA. No network, no dependency, no shared bottleneck.

At three services and moderate traffic, this difference is not worth optimizing — introspection with a short cache is simpler and you should probably do that. At forty services across three regions, with partner integrations validating your tokens from their own infrastructure, local verification isn't an optimization, it's the only thing that works. Requiring a partner in another company to call your introspection endpoint on their hot path is a coupling neither side wants.

The threshold, roughly: JWTs win when validators are numerous, organizationally distant, or geographically distant from the issuer. Opaque tokens win when validators are few and close.

Caching introspection results is the obvious middle ground and it's a legitimate technique, but be clear about what it does: it converts your opaque token into a JWT with extra steps and worse debuggability. If you cache introspection for 60 seconds, you have a 60-second revocation window, exactly like a 60-second JWT — except you also still have the round trip on cache misses and the identity service dependency. Sometimes that's the right trade. Often, once you've written it down, you realize you wanted a short-lived JWT.

The pattern most mature systems actually run

The interesting architectures don't choose. They put the boundary at the gateway, and the pattern has a name worth knowing: the phantom token (sometimes "split token") pattern.

flowchart LR
    C["Client"] -->|"opaque token<br/>30 bytes"| G["API Gateway"]
    G -->|"introspect / exchange"| AS["Authorization Server"]
    AS -->|"JWT"| G
    G -->|"JWT access token<br/>internal only"| S1["Service A"]
    G --> S2["Service B"]
    G --> S3["Service C"]

The client — untrusted, external, possibly a browser — holds an opaque reference. It carries no PII, it's tiny, it's revocable instantly because every request passes through a gateway that resolves it. At the gateway, the opaque token is exchanged for a JWT (via introspection, or properly via RFC 8693 token exchange) which is then forwarded internally. Inside the trust boundary, every service does cheap local verification with no dependency on the identity service.

You get instant revocation at the only place it matters — the external edge — plus stateless validation everywhere internally, plus no PII in anything a customer's browser or logs touch, plus JWT debuggability inside your own infrastructure where the transparency is a benefit rather than a leak.

The costs are real. The gateway becomes a mandatory, stateful, latency-critical component — it needs an introspection cache with its own correctness story, and it must be on the path of every request. Internal JWTs need to be short-lived and audience-scoped, or a service that receives one can replay it against a service it shouldn't reach. And you now have two token formats in your system, which is genuinely more to explain to new engineers.

If you already run a gateway that terminates all external traffic, this pattern costs you very little and is close to strictly better. If you don't, adding one to get it is a large project, and a short-lived JWT everywhere is a reasonable answer instead.

The decision, in the order that resolves it

Can you put a gateway on every external request? If yes, phantom token. Stop here; you get both sets of properties.

Do validators outside your organization need to verify tokens? If yes, JWT. Introspection across an organizational boundary is a coupling that fails.

How many services validate, and are they all in one deployment domain? Under five, in one cluster, one language: opaque with a short introspection cache is simpler and you will spend the saved effort better elsewhere. Twenty-plus, or polyglot, or multi-region: JWT, and budget for the validation middleware as a real internal product.

Is your authorization data unbounded in cardinality? Groups, per-resource entitlements, long permission lists — then either token format works, but the data does not go in the token. Decide this before someone else decides it for you in a hotfix.

Who is on call for a validation failure? If the answer is nineteen different teams, JWT validation is distributed and so are its failure modes. That's an argument for centralizing verification at a gateway regardless of format.

The honest summary is that the format is a less important decision than the industry's volume of writing about it suggests. What matters is whether you've decided where validation happens, who owns it, what the token is allowed to carry, and what breaks when the identity service is slow. Teams that answer those questions are fine with either format. Teams that don't, aren't fine with either — and they usually discover it when their largest customer's admin can't log in, and the only clue is a 400 from a proxy nobody remembers configuring.