Identity Isn't Your API Gateway
The outage lasted ninety-four seconds and took out every API in the company.
Not the login page. Every API. The mobile app, the partner integrations, the internal admin tools, the checkout flow, the webhook dispatcher — everything returned 401 for a minute and a half, and then everything came back on its own. The identity platform's own dashboards showed a brief control-plane hiccup during a node replacement; its p99 for token issuance never moved, because almost nobody was logging in at 02:40 on a Thursday. By the time anyone looked, the graphs were flat.
The mechanism took two days to find, and it was a single line in a gateway route template:
introspection:
endpoint: https://idp.internal/oauth2/introspect
cache_ttl: 0
Fourteen months earlier, an auditor had asked how quickly a compromised token could be revoked. The honest answer at the time was "up to sixty seconds, because the gateway caches introspection results." The team that owned the gateway did the obvious thing: they set the TTL to zero, the finding was closed, and the answer became "immediately." Nobody wrote down what it cost.
What it cost was this. Every API request in the estate now made a synchronous call to the identity platform, on the request path, with no fallback. A 99.95% identity platform — perfectly respectable, 263 minutes of unavailability a year — had been promoted into a hard, per-request dependency for a checkout flow that made eleven API calls. 0.9995^11 is 99.45%, so one checkout in 182 was already failing before the incident, quietly, inside a retry that mostly hid it. During the ninety-four seconds, the number was zero.
The interesting part isn't the config line. It's that two teams made a joint decision without either of them knowing it. The security engineer traded availability for revocation latency. The platform engineer who owned API availability was not in the room, did not know the trade existed, and could not have found it in a dashboard afterwards, because the failing component reported itself healthy the entire time.
That is what a merged boundary looks like from the inside. This article is about where the line between an identity platform and an API gateway actually goes, why it keeps getting erased in both directions, and the specific arithmetic that tells you which side of it a given concern belongs on.
Why the pull is real
The gateway is an unusually tempting place to put identity, and the temptation is not stupidity — every reason is a good one.
It already terminates TLS, so it already has the plaintext request. It already sees 100% of traffic, including traffic to services that would otherwise each need their own validation logic. It already has a config language for per-route policy, and a cache. It is already the thing you deploy when you want a rule applied everywhere at once. If you're looking for a place to enforce something universally, the gateway is not a bad guess; it's the first guess, and usually the right one.
The pull runs the other way too, weaker but persistent. The IdP's endpoints are the ones under attack, so someone proposes putting rate limiting inside it. It already needs a public hostname per tenant, so someone proposes it do the routing. Each of these is one component absorbing a neighbour's job because it happens to be standing closest to the problem.
Both pulls produce the same failure, which is that the two components stop having separable lifecycles. You can tell it has happened without reading any architecture diagram, using one question: can you change the gateway's routing without an identity deploy, and rotate an identity signing key without a gateway deploy? If either answer is no, they are one component wearing two names, and the rest of this article is describing your outages.
The distinction that actually resolves it
There is a clean version of this boundary and it is not "auth goes here, routing goes there." It's this:
The gateway is a data-plane component making a per-request decision in microseconds against material it already holds. The identity platform is an issuer and the source of truth for identity lifecycle. Material flows from the issuer to the data plane out of band, on its own schedule. The issuer is never in the synchronous path of an API call.
Validation is a data-plane operation: given this token and this route, admit or reject, using cached keys or a cached introspection result, with no network call in the common case. Issuance is a control operation involving credentials, session state, key custody, audit, and a database — and it happens orders of magnitude less often. Identity is mostly read traffic: a user authenticates once and then makes thousands of API calls under that authentication. Put the expensive once-per-session operation and the cheap thousands-per-session one in the same component and you've coupled the availability of the second to the first, which is where all your traffic is.
flowchart TB
subgraph CP["Control path — asynchronous, cached, allowed to be slow"]
IDP["Identity platform<br/>issuance · key custody · lifecycle · audit"]
IDP -.->|"JWKS, refreshed on a timer"| KC["Key cache"]
IDP -.->|"introspection results, TTL + stale window"| IC["Token cache"]
IDP -.->|"revocation events"| EV["Invalidation stream"]
end
C["Client"] -->|"API request + token"| GW["API gateway"]
KC --> GW
IC --> GW
EV --> GW
GW -->|"admitted"| SVC["Services"]
C2["Browser"] -->|"login, once per session"| IDP
The two arrow styles are the whole design. Solid arrows are the request path and none of them touch the identity platform. Dotted arrows are cache fills, and they are allowed to fail, retry, and lag.
Caching converts an availability dependency into a staleness dependency — but only if you configure the second half
This is the part that the opening incident got wrong, and it is the single most useful idea in this article, because most teams believe they have already done it.
When the gateway calls the IdP on every request, the IdP's availability multiplies into every request. When it calls once per TTL window and serves the rest from cache, something qualitatively different happens: the IdP's availability stops being a term in your availability product and becomes a term in your freshness budget. You are no longer asking "was the IdP up when this request arrived?" but "how old is the answer I'm using?" — a strictly better question, because staleness degrades gracefully and availability doesn't.
Except a TTL alone does not buy you this. A TTL is a freshness control: it says discard this after 60 seconds. It says nothing about what to do when the refill fails, and the default behaviour of almost every gateway is to discard the entry, attempt a refill, fail, and reject the request. So:
| Config | 90-second IdP outage | Revocation window |
|---|---|---|
ttl: 0 |
100% of API requests fail for 90s | 0 |
ttl: 60s, no stale handling |
~0% fail for 60s, then ~100% fail for 30s | ≤ 60s |
ttl: 60s, stale_if_error: 10m |
0% fail | ≤ 60s normally; ≤ 10m 60s during an IdP outage |
ttl: 300s, stale_if_error: 10m |
0% fail | ≤ 5m normally |
The middle row is where most deployments actually sit, and it's the one people believe protects them. It doesn't. It converts a total outage into a delayed total outage, which is worse during an incident, because the correlation with the triggering event is now sixty seconds offset and nobody spots it.
The bottom two rows are the design. The stale window is the availability control; the TTL is the freshness control; they are different numbers with different owners, and a system with only one of them has made half the decision. HTTP has had the distinction since RFC 5861 (stale-while-revalidate, stale-if-error); gateway auth plugins overwhelmingly do not implement it, so you usually have to build it — a cache that keeps the last-known-good entry past expiry, and a circuit breaker that serves it when the upstream error rate crosses a threshold. As a rule: if a cache miss during an upstream outage produces a 401, you have not decoupled from the upstream, you have added a delay to your coupling.
The cost is the right-hand column, and you have to be able to say it out loud: with a ten-minute stale window, a token revoked during an IdP outage keeps working for up to ten minutes past its normal expiry. That's a real security property traded away, and the right response isn't to refuse the trade but to make it at a number someone signed off on, with a lever to collapse the window during an incident. Which is the argument in why token introspection isn't as slow as you think: the TTL's value is that it's a dial, and a dial you've never turned is one you don't know works.
The load arithmetic that surprises people
Here is a fact about gateway-side introspection that changes where you decide to put the cache, and I rarely see it stated:
Introspection load at a gateway is a function of the number of active tokens and the cache TTL. It is not a function of your request rate.
Each active token needs one introspection per TTL window, no matter how many requests it makes inside that window — so the rate is simply active tokens divided by TTL. With 100,000 concurrently-active tokens and a 60-second TTL, that's ~1,667 introspections per second — whether your APIs are doing 20,000 requests per second or 200,000. Doubling your traffic costs the identity platform nothing. Doubling your user base costs it linearly. Halving the TTL to buy a shorter revocation window doubles it, exactly, which means the revocation window has a precisely computable price in queries-per-second and you can put it in a capacity plan.
It's also why the gateway is the right cache location rather than each service: cache locality here is a fan-in property. Twenty services caching independently each see a given token a twentieth as often, so every cache is colder, every one refills separately, and the aggregate introspection rate is up to twenty times higher for the same TTL. Pull the cache to the gateway and the arithmetic above holds with one term. Combined with the phantom-token pattern you get small public tokens, no per-service JWKS handling, and one place to collapse the revocation window during an incident.
One correction to the formula before anyone puts it in a capacity plan: the gateway fleet is not one cache, it's N node-local caches, so the numerator is really active tokens times the number of nodes a given token's requests can land on. With connection reuse and a well-behaved load balancer that's a small multiple; with aggressive scaling and no affinity it's the whole fleet. A shared second tier behind the in-process one collapses it back, at the price of a network hop in the miss path.
Who owns what
| Concern | Gateway | Identity platform | Why |
|---|---|---|---|
| TLS termination for API traffic | ✅ | ❌ | Data plane. The IdP terminates TLS only for its own endpoints. |
| Routing, versioning, request shaping | ✅ | ❌ | Nothing to do with identity; putting it in the IdP makes it a proxy. |
Signature verification, exp, iss |
✅ | ❌ | Per-request, microseconds, against cached keys. |
aud and scope checks |
✅ | ❌ | Enumerable from the route definition at config time. |
| Introspection of opaque tokens | ✅ (cached) | ✅ (serves it) | Gateway caches; IdP is the source of truth. |
| Token issuance and refresh | ❌ | ✅ | Requires credentials, session state, key custody, audit. |
| Signing key custody and rotation | ❌ | ✅ | One issuer, one key lifecycle, published via JWKS. |
| Login, consent, redirect flows | ❌ | ✅ | Stateful multi-request flows; see below. |
| Session lifecycle and logout | ❌ | ✅ | Requires knowing what a session is across applications. |
| Revocation decisions | ❌ | ✅ | Source of truth; gateway consumes the result. |
| IP-keyed rate limiting, geo, bot signals | ✅ | ❌ | Decidable from the request alone — 12.1. |
| Account-keyed rate limiting, lockout | ❌ | ✅ | Requires account knowledge the gateway doesn't have. |
| Per-object entitlement decisions | ❌ | ❌ | Neither — 12.6. |
| User attributes, profile, domain data | ❌ | ❌ | Neither — 12.4. |
The two rows at the bottom are the ones people find surprising, and they are the reason this boundary is genuinely three-sided rather than two-sided. A lot of what teams try to push into the gateway isn't gateway work or identity work; it's application work that got homeless when nobody would claim it.
Validation at the gateway, issuance never
The clean statement is that the gateway may verify a claim of identity and must never originate one. That distinction is sharper than it sounds, and the phantom-token pattern is where it gets tested, because in that pattern the gateway does mint something — an internal assertion for downstream services.
Is that issuance? The discriminator: a component is an issuer if it can produce a credential naming a subject whose authentication it did not itself observe and cannot itself attest to. A gateway translating an opaque token it just introspected into a seconds-long internal assertion is restating a fact the IdP gave it, one hop, with a narrower audience. A gateway that can mint a token for sub=alice because a config rule said so is an issuer, and the moment it becomes one, three things are true that nobody planned for. It holds a signing key, so key custody and rotation now apply to a component that deploys several times a week. There are two things that can create a valid session, so "revoke everything for this user" has two implementations and one will be forgotten. And its compromise is total: a gateway that only validates leaks nothing beyond the traffic passing through it, while a gateway that signs can impersonate anyone until the key is rotated everywhere.
If you build phantom tokens, keep the internal assertion narrow enough that it obviously isn't a general credential: seconds-long lifetime, audience scoped to one internal service, a distinct iss so nothing accepts it as a user token, and a key that is not the IdP's. That's a real cost, and it's why plenty of teams correctly decide the gateway should forward the original token untouched instead.
The JWKS handshake, and the three ways it breaks
The contract between issuer and validator is small enough to write on a napkin, which is why so few teams write it down and so many find out its details during an incident. Rotation itself is a solved problem with a known phased timeline; what follows is specifically the gateway's half of it.
The issuer's obligations: publish every currently-valid key at a stable JWKS URL; emit a kid on every token; add a new key to JWKS before signing with it and keep the old one published until every token signed with it has expired; serve JWKS with cache headers that mean what they say.
The validator's obligations: fetch JWKS over the network rather than holding keys in config; select by kid rather than trying keys in turn; refresh on a timer and opportunistically on an unknown kid; and — the one everyone misses — bound that opportunistic refresh.
Three failure modes, in increasing order of how badly they go.
Keys pinned in gateway config. Someone pasted the PEM into a Helm values file, because it was one key and it worked. Rotation now requires a gateway deploy, so emergency rotation takes as long as your slowest CD pipeline — on a bad day, longer than the incident you're rotating for. This is the purest example of the merge: an identity-lifecycle operation that requires a data-plane deploy, and it silently defeats the phased rotation timeline, because the IdP's careful wait for consumer caches to refresh is a wait for a cache that doesn't exist.
Fetch once at startup, never refresh. The mirror image. Rotation appears to work — the IdP publishes, waits, signs — and then every gateway node that has been up since before phase 1 starts rejecting everything, while nodes that restarted recently work fine. The signature is a partial outage that follows your pod ages, and it self-heals on any deploy, which is why it gets closed as "transient."
Unbounded refresh on unknown kid. The interesting one, because it converts unauthenticated traffic into an attack on your identity platform, originating from inside your own infrastructure.
The logic is sensible in isolation: a token arrives with a kid the gateway hasn't seen, so the gateway refetches JWKS in case a rotation just happened. Now consider 8,000 requests per second bearing tokens with random kid values — trivially generated, no valid signature needed, because the kid lookup happens before verification. Every one is a cache miss. Every one triggers a JWKS fetch. Your gateway fleet issues 8,000 requests per second at the IdP's discovery endpoint, and it is doing so from inside the network, past the edge rate limiter, using well-formed requests that no WAF has any reason to flag. Meanwhile the gateway's own rate limiter doesn't fire, because from its perspective these are cheap requests that are all being rejected correctly.
The fix is three lines of policy and almost nobody has all three:
- A token bucket on JWKS refetches per issuer per node — something like one refetch per five minutes with a small burst. Beyond the bucket, an unknown
kidis simply invalid. - A negative cache on
kid: once you've refetched and thekidis still absent, remember that for 30–60 seconds and stop looking. - Verify cheaply first where you can. Reject structurally invalid tokens, wrong
iss, and expiredexpbefore doing any key lookup. It doesn't eliminate the class — an attacker can send well-formed unexpired garbage — but it removes the trivial version.
One multi-tenant wrinkle, because it's where this contract stops fitting on a napkin. Platforms that give each tenant its own OIDC issuer and its own signing keys — ClavionX does this; I work on it, which is why I've had to think about the gateway side — mean the gateway cannot pin a single JWKS URL. It must resolve issuer → JWKS dynamically, which reopens the refetch question over a much larger key space: an attacker who can vary iss as well as kid tries to make your gateway fetch discovery documents for issuers that don't exist. The answer is an allowlist of issuer prefixes plus the same bounded-refetch discipline, per issuer. Per-tenant keys are a good arrangement — one tenant's key compromise isn't everyone's — but they are not free at the edge, and a gateway configured as though there were one issuer works perfectly until the second tenant onboards.
The revocation window is a sum, and three teams own the addends
Ask a security engineer how long revocation takes and you'll get a number. Ask where the number comes from and you'll usually find it's the token lifetime, because that's the setting with "lifetime" in the name.
The real number is a sum, and its terms live in different config files owned by different people:
| Term | Typical value | Config lives in | Usually owned by |
|---|---|---|---|
| Access token lifetime | 15 min | IdP client config | Identity team |
| Gateway introspection cache TTL | 60 s | Gateway route policy | Platform team |
| Gateway stale-if-error window | 0–10 min | Gateway resilience policy | Whoever wrote it last |
| Downstream service token cache | 30 s | Application config | Each service team |
| Revocation event propagation | 100 ms–∞ | Event bus config | Data/infra team |
For a JWT with no introspection, the window is the token lifetime, full stop — the gateway cannot know about a revocation, because nothing in the request tells it. For introspected tokens, the steady-state window is the introspection TTL, and the degraded-state window is the TTL plus the stale window. If a downstream service caches too, add that.
Nobody owns the sum. That's the actual finding, and it's why the opening incident happened: an auditor asked for one term to go to zero, one team could make one term go to zero, and the sum was never the unit of discussion. The fix isn't architectural — it's that the revocation window is a stated SLO with a named owner, written down per token class, with every contributing config referencing it. Then the conversation about zeroing a TTL happens in the presence of its availability cost, which is the only place it can be evaluated honestly.
Push invalidation reshapes this — an event that evicts gateway cache entries decouples the window from the TTL entirely, letting you run a long TTL for load and a short window for security. That machinery is in token introspection; the gateway-specific note is that the gateway is by far the best subscriber, being one fleet rather than every service, and it can be told to fall back to zero-TTL if the event stream goes quiet for longer than a threshold.
Scope and audience at the gateway; entitlements at neither
The gateway can enforce exactly the checks whose expected value is knowable when the route is configured. That's a precise rule and it draws a clean line.
aud is knowable: this route fronts the payments service, so the token's audience must include the payments service. scope is knowable: this route is POST /transfers, so the token must carry transfers:write. Both are properties of the route, both are static, both are enumerable in a config file, and both can be evaluated in a string comparison with no state.
Audience deserves a paragraph of its own, because it's the check most often skipped and its absence is the most expensive. A gateway that verifies signature, issuer, and expiry but not audience has turned every token it accepts into a universal key. A token minted for the low-value analytics API is now accepted at the payments API, because both are signed by the same issuer and both routes only asked "is this a valid token?" This is the confused deputy at the edge, and the reason it survives review is that the gateway config for each route looks correct in isolation — the mistake is only visible when you notice that no route distinguishes itself from any other. If you audit one thing after reading this, grep your route policies for audience enforcement and count how many have it.
What the gateway cannot enforce is anything whose expected value depends on the request's content: may this user read document 8812, approve this expense, view this tenant's records. Not because the gateway is technically incapable — you can write that code — but because doing so requires the gateway config to encode what a document is and who may see it, and now a product rule change is a gateway deploy. The full argument for why that ends badly is Identity Isn't Your Authorization Engine; the gateway version of it has an extra sting, which is that gateway config is typically owned by a platform team who will be paged for a product team's authorization bug.
The rate-limiting split falls out of the same rule, and the mechanism is worth naming rather than the conclusion. A limiter's placement is decided by what it can key on. The gateway's natural keys are IP, route, and client certificate — present in the request, no state required. Identity's natural keys are account, tenant, and client ID, which the gateway cannot use, because deriving them requires having already validated a token, and the requests you most want to limit — failed logins — don't carry one. That's why an IP-keyed limiter cannot see password spraying: one attempt each against ten thousand accounts from a botnet is unremarkable on every key the gateway has. Same dividing line as 12.1, from the implementation side: the limiter goes where the key lives. Costing the identity half is rate limiting is a capacity problem.
The anti-pattern: login flows at the edge
Everything so far has been about the gateway doing more validation than it should. This section is about the gateway doing authentication, and it's a different category of mistake because it doesn't degrade gradually. It works, demos beautifully, and then hits a wall.
The proposal always sounds like consolidation. The gateway already has an OIDC plugin: turn it on and it handles the redirect, the callback, the code exchange, the session cookie, and the headers injected downstream. No application changes at all. For a while this is genuinely delightful. What breaks, roughly in the order teams hit it:
The gateway becomes stateful. A session cookie at the edge implies a session store at the edge. Either it's in-process — in which case scaling the fleet, deploying, or losing a node logs people out, and the gateway you chose specifically because it was horizontally scalable and disposable is now neither — or it's a shared store, in which case you've added a datastore to your data plane with its own availability, latency, and eviction semantics, sitting in front of every request.
Cookie scope fights routing. Cookies are scoped by domain and path; gateway policy is scoped by route. Different partitions of the same space, and they don't compose. The first time you need two applications on one hostname with different session lifetimes, or one application spanning two hostnames, you're writing rules that are correct for reasons nobody can articulate six months later.
Logout doesn't work. This is the decisive one. The gateway can clear its own cookie. It cannot end the IdP session, so the next request bounces to the IdP, gets an immediate silent re-authentication, and the user is logged back in without a prompt — a bug report that reads "logout button doesn't work" and takes a week to explain. Making it work means implementing OIDC front-channel or back-channel logout in the gateway, which means the gateway now needs to be a registered relying party with a persistent session index. You are building an IdP client inside your data plane. The forgotten side of SSO covers why this is hard even when it's built by people whose job it is.
Downstream services get an identity they can't verify. The gateway injects X-Authenticated-User and services trust it, which is fine right up until any path reaches a service without traversing the gateway. That entire failure class is What Happens When the Layer Below You Lies, and the reason edge-terminated login makes it worse is that there is no token left downstream — the assertion has been reduced to a header with no signature on it, so a service cannot verify it even if it wants to.
Refresh becomes the gateway's problem. Now the gateway holds refresh tokens — the longest-lived, highest-value credentials you have — in a store designed for caching route configs, and must handle rotation, reuse detection, and concurrent-refresh races across nodes that don't coordinate.
None of these are unsolvable. All of them are things an IdP client library already solved, and building them again inside a proxy means building them in a component with the wrong deployment cadence, the wrong state model, and the wrong on-call rotation.
When an edge auth proxy is genuinely right
And yet the pattern persists, because for one class of problem it is unambiguously the correct answer: applications you cannot modify.
A twelve-year-old internal Java application. A vendor appliance with a web UI and no SSO support below the enterprise tier. A Jupyter or Grafana instance that should be behind SSO but whose own auth you don't want to configure per-instance. In every one of these, "add an OIDC library to the application" is not available, and the choice is between an authenticating reverse proxy and no SSO at all. oauth2-proxy, mod_auth_openidc, and the various identity-aware-proxy products exist for exactly this and they are good at it.
The distinction that keeps this from contradicting the previous section is blast radius. The right shape is an auth proxy dedicated to the legacy application — a sidecar or small deployment in front of that one app — rather than the shared gateway growing an auth mode all routes inherit. Concretely:
- One proxy instance per protected application, so its session model, cookie scope, and lifetime are that application's and not everyone's.
- It sits behind the shared gateway, not instead of it. The gateway keeps routing, TLS, and IP-keyed limiting; the proxy does the OIDC dance for its one tenant.
- It strips every inbound copy of the identity headers it sets, and the application is reachable only through it — network policy, not convention.
- It's an adapter with a retirement date, not architecture. When the application is replaced, the proxy goes with it.
The test is whether the pattern is contained. An auth proxy in front of one legacy app is a compatibility shim; the same code path turned on globally at the shared gateway is the anti-pattern above — and the distance between them is one config flag that someone will eventually flip because it worked so well the first time.
The tells, and the failure modes
You have merged the two components if any of these is true:
- Rotating a signing key requires a gateway deploy, or adding a route requires an identity change.
- Onboarding a tenant requires a coordinated release of both.
- The gateway holds a private signing key, or refresh tokens.
- Nobody can state the revocation window without opening two repositories.
- The IdP's config contains route paths, upstream hostnames, or API version prefixes.
- An identity incident and an API incident are always the same incident.
And the failure modes, by misplacement:
| Misplacement | What it looks like when it fails | When you find out |
|---|---|---|
| Zero-TTL introspection at the gateway | Brief IdP degradation becomes total API outage, amplified by fan-out | First IdP maintenance after the config lands |
| TTL with no stale window | Same, delayed by one TTL, so the correlation is missed | During the incident review that reaches no conclusion |
| JWKS pinned in gateway config | Scheduled rotation causes 100% 401s at phase 3 | Your first real rotation, usually the emergency one |
| JWKS fetched once at startup | Partial outage correlated with pod age; self-heals on deploy | Closed as "transient" two or three times first |
Unbounded refetch on unknown kid |
Self-inflicted DoS on the IdP from inside the perimeter | Under attack, or the first time a fuzzer finds you |
No aud enforcement per route |
Any valid token works on any API | Penetration test, or a breach |
| Gateway mints its own tokens | Second signing key with no rotation story; audit can't reconstruct auth | Key compromise, or an auditor asking who signed what |
| Login flows at the shared gateway | Logout doesn't work; scaling logs users out; refresh tokens at the edge | Week two of production |
| Entitlement checks in gateway config | Product rule changes require platform-team deploys | The first time product moves faster than platform |
| Account-keyed limiting attempted at the gateway | Password spraying invisible; per-IP limits shed real users behind NAT | When a customer's whole office gets locked out |
| Routing/TLS/versioning inside the IdP | The IdP is now a proxy with a slow release cadence | When someone needs a header rewritten and it takes a sprint |
What to do on Monday
Four things, in descending order of payoff:
- Give every request-path call to the IdP a stale window. Not a longer TTL — a stale window, a circuit breaker, and a test that proves the gateway serves traffic while the IdP returns 503. If you can't run that test today, you don't know your API's availability.
- Write the revocation window down as a number, per token class, with an owner, and make every config that contributes a term reference it. The sum is the SLO; the terms are implementation.
- Grep your route policies for audience enforcement and count how many have it. Ten minutes, and the highest-value ten minutes here.
- Try rotating a signing key in staging without deploying the gateway. If you can't, you've found the merge, and the rest of this article is describing outages you haven't had yet.
The underlying principle isn't really about gateways. It's the rule that governs every boundary in this series: the component with the most dependents must have the fewest dependencies. The gateway sits in front of every API request in the company, which makes it the component with the most dependents — and that is precisely why it must not synchronously depend on anything, least of all on the identity platform, whose failures look like security events rather than outages and therefore get debugged last.
The identity platform issues. The gateway verifies. Material moves between them on a schedule, cached, with an explicit staleness budget somebody owns. Everything good about the arrangement comes from the fact that neither component has to be up for the other one to work.
Disclosure: I work on ClavionX. It appears above once, for the per-tenant-issuer detail, because that arrangement genuinely changes what a gateway has to do and it's a constraint I've had to design against rather than a feature I'm recommending. The argument holds regardless of what issues your tokens.