Debugging Across Four Layers

The ticket said: "Login works, but only sometimes, and only for some people at Contoso."

Three days of investigation produced the following evidence. The CDN's access log showed a 302 for every request in the flow, no 4xx, no 5xx. The gateway showed the same requests arriving, being forwarded, and returning. The identity platform's audit log showed a clean authentication.succeeded for every affected user, with an assertion issued, signed, and POSTed. The application's log showed the user arriving at the assertion consumer endpoint and being redirected back to the login page with invalid_state.

Four layers. Four logs. Every one of them green, except the last, whose error told you that something was missing without telling you what.

The engineer who eventually found it did not read another log. She got a browser session recording from one affected user, opened the HAR file, and looked at the request that POSTed the assertion. The Cookie header was not there. Not empty — absent. The identity platform had set the flow-state cookie forty milliseconds earlier and the browser had accepted it; it simply declined to send it back on a cross-site POST, because the cookie had been issued with SameSite=Lax and Chrome had shipped the default that made that decision universal. Contoso's users were the ones who had upgraded first. The 3% was a browser-version cohort, not a tenant cohort, and it had looked like a tenant cohort because Contoso's IT department pushed browser updates on a schedule.

The thing worth sitting with is not the SameSite bug, which is well documented. It is this: the evidence that identified the failure did not exist on any server. There was no log line to find, no trace to sample, no metric that could have moved. The failing thing was a header the browser chose not to send, and the only witness to that decision was the browser.

This is the normal condition of cross-layer debugging, not an edge case. What Happens When the Layer Below You Lies argued that layered architecture creates a trust problem at every boundary. Who Owns the Incident When Every Layer Worked Correctly? argued that it creates an ownership problem at every seam. This one is about the craft: given a user who is blocked and four layers that all report success, how do you actually find the failure?

Layers are defined by what they can witness

Start by fixing the decomposition, because the usual one is wrong for this purpose.

The layers in most architecture diagrams are drawn along ownership lines: CDN team, gateway team, identity team, application team. That is the right decomposition for deciding who fixes something, and it is exactly the decomposition 12.11 used. It is the wrong decomposition for finding something, because two components owned by different teams frequently see the identical slice of reality, and one component sometimes sees two.

The decomposition that works for debugging is by unit of observation — the thing each layer's records are keyed by, the primary key of its worldview:

Layer Unit of observation Primary key
1. Browser / client the journey — a sequence of navigations belonging to one user intent tab / document, cookie jar
2. Edge & network the connection and the request source IP, TLS session, connection id
3. Identity platform the authentication transaction account, session, flow state
4. Resource server the credential in use token jti, sub, sid

Four, and not more, because these are the four irreducibly different kinds of record. Your CDN, WAF, load balancer, service mesh and API gateway are five products, five teams and five bills, but for evidence purposes they are one layer: they all observe connections and requests, none of them knows what an account is, and their logs join to each other trivially because they share IPs and request identifiers. Collapsing them is not a simplification, it is an accurate statement about what they can testify to.

Conversely, the browser is a layer even though nobody thinks of it as infrastructure, because it holds evidence — cookie decisions, redirect chains, blocked requests, URL fragments — that no server can reconstruct.

The practical consequence: a bug is findable in a single layer's logs only if its cause and its effect share that layer's primary key. Everything else is a join, and cross-layer debugging is a join problem in a schema where no two tables share a column. That is the actual difficulty. Not log volume, not retention, not tooling.

What each layer emits, and what it hides

1. The browser

Emits: the complete ordered redirect chain with real wall-clock times; which cookies were sent on which request and which were rejected; requests that failed CORS or CSP and never reached a server; the URL fragment; TLS and connection reuse; the exact final URL, including the error parameters your application's error page threw away.

Hides: everything about server-side decisions. It sees a 302 and a Location; it has no idea whether that redirect meant "policy evaluated, MFA not required" or "session cookie unreadable, starting over."

The part people underuse: the browser is the only layer that can testify to a negative. A request that was never made produces no server log anywhere, by construction, and there are a lot of identity failures whose signature is a request that didn't happen. It is also the only layer that sees the URL fragment at all — # is never transmitted — so any flow that returns errors or tokens in the fragment is, server-side, invisible by protocol design.

2. Edge and network

Emits: source IP and ASN, TLS fingerprint and version, exact bytes in and out, request and response headers as they were on the wire, timing to the millisecond, and the status code the client actually received — which is not always the status code your application produced.

Hides: meaning. The edge sees POST /sso/acs → 302. It does not know whether that was a successful login or the fourth failure in a stuffing run. It also hides its own transformations: header rewriting, body buffering, request coalescing, and retries all happen here and are rarely logged as changes.

The part people underuse: the edge is where the browser's own metadata arrives and gets thrown away. Sec-Fetch-Site, Sec-Fetch-Mode and Sec-Fetch-Dest are sent by every current browser on every request and almost nobody logs them. Had the edge in the story above been logging Sec-Fetch-Site: cross-site alongside Sec-Fetch-Mode: navigate on a POST with no Cookie header, the SameSite failure would have been a server-side query rather than a three-day hunt. Two extra fields in an access log format, and an entire class of browser-only failures becomes visible from the server.

3. The identity platform

Emits: the richest semantics in the stack — which policy applied, which factor was satisfied, why a request was refused, which connection the user came through, what was in the token.

Hides: everything before and after itself. It cannot tell you whether the user ever received the redirect it issued, whether the assertion arrived intact, or whether the token it minted was accepted. Its "success" event means I produced a valid artifact, which is a genuinely different claim from the user logged in.

The part people underuse: identity is the only layer that knows the expected shape of the next hop. It issued a state, it set a cookie, it knows the registered redirect_uri. So it is the only layer that can log a specific, actionable absence: not "request rejected" but "expected flow state abc123, received none, cookie header present but did not contain it." Most implementations log the first sentence and discard the material for the second.

4. The resource server

Emits: the token as presented, the validation outcome, the authorization decision, and the business effect.

Hides: how the token was obtained. By the time a bearer token arrives, every fact about the journey that produced it is gone except what someone deliberately put inside the token.

The part people underuse: the resource server is the last place where a signed artifact and a live request coexist, which makes it the only place you can compare the identity platform's clock against your own. Log iat, nbf and exp alongside local time on every validation failure and you get continuous skew measurement for free.

The correlation problem is not what you think

Everyone's first answer is a correlation ID. Mint one at the edge, propagate it in a header, join on it. This works well for a service call chain and it works badly for a login, and the reason is worth stating precisely because most teams discover it only after building the pipeline.

A redirect is not a continuation of a request. It is the termination of one request and the browser's independent decision to start another.

When your identity platform returns 302 Location: https://app.example.com/callback?code=..., the request that carried your traceparent header is over. The browser then constructs a brand-new request to a new host. It attaches: the URL, the cookies scoped to that host, a Referer (which, under the modern default referrer policy, is origin-only — https://idp.example.com/ with no path, enough to know where you came from and useless for knowing which transaction), and a handful of Sec-Fetch-* hints. It does not attach your header. It has never heard of your header.

So in a login flow that crosses four hosts and six redirects, a header-based trace ID does not degrade — it terminates, six times. What you get is not one trace with gaps. It is seven disjoint traces that nothing links.

Only two carriers survive a top-level redirect: the URL and cookies. Cookies are scoped to a host, so they cannot cross the SP↔IdP boundary, which is the boundary you most need to cross. That leaves the URL. Which is why every federation protocol invented its own correlation parameter and put it in the query string:

  • OAuth/OIDC state — minted by the client, echoed verbatim by the authorization server. Opaque to the IdP.
  • OIDC nonce — minted by the client, and this is the interesting one: it is copied into the ID token, so it survives into a signed artifact and therefore reaches layer 4.
  • SAML RelayState — the same idea, and InResponseTo, which joins the assertion back to the request that asked for it.

These are the real join keys, and they change identity at every hop. There is no global trace ID for a login flow. There is a chain of pairwise correlations, and walking it is the skill:

sequenceDiagram
    participant B as Browser
    participant E as Edge
    participant A as App (SP)
    participant I as Identity
    participant R as Resource server

    B->>E: GET /login
    E->>A: GET /login  (trace: t1)
    A-->>B: 302 → /authorize?state=S&nonce=N
    Note over B,A: t1 ends here. Join key is now S.
    B->>E: GET /authorize?state=S
    E->>I: GET /authorize  (trace: t2)
    I-->>B: 302 → /callback?code=C&state=S
    Note over B,I: t2 ends. Join keys: S (to app), C (to token endpoint).
    B->>E: GET /callback?code=C&state=S
    E->>A: GET /callback  (trace: t3)
    A->>I: POST /token (code=C)
    I-->>A: id_token{nonce=N, jti=J, sid=D}
    Note over A,I: N joins the token back to the /authorize request.
    A->>R: GET /api (Bearer, jti=J)
    Note over R: J and D are the only keys that reach here.

Five spans, four different join keys, and the only identifier present in both the first hop and the last is the one you deliberately smuggled through — because nonce and sid are the only fields that make it from the authorization request into the token.

The trick that makes this tractable

state is fully under your control, opaque to the identity platform, and guaranteed to be returned to you byte-for-byte. So stop making it a bare random string.

Make it a signed, compact structure containing the CSRF value and your trace ID:

state = base64url( { "csrf": "9f3a…", "tid": "4bf92f3577b34da6a3ce929d0e0e4736" } )  ‖ signature

Now every layer that sees state — the edge access log, the identity platform's audit record, your own callback handler — can log a field that joins directly back to the trace that started the journey. One field, and the chain of pairwise correlations collapses into something you can query.

Three cautions, because this is a query parameter and query parameters leak. It lands in browser history, in Referer where the policy permits it, and in the access logs of every intermediary including ones you do not operate. Put no PII in it. Keep it small — you are sharing a URL length budget with SAMLRequest and some IdPs truncate aggressively. And sign it, because a state you parse is a state an attacker can craft.

The token-side equivalent — carrying a transaction identifier inside the token so it survives into machine-to-machine hops — is covered in Auditing Agent Actions; the two techniques compose, and together they cover the browser half and the API half of the same journey.

Four clocks, and why timestamp ordering lies

Once you have the records, you will want to sort them by time. Do not do it naively.

The clocks differ in kind, not just in offset. Edge nodes and identity servers are typically NTP-disciplined to within milliseconds because their operators care. The browser's clock is set by a person and can be minutes, or years, wrong — and it is the clock that evaluates cookie Expires, so a user with a badly wrong clock experiences a very specific and very confusing failure where sessions never persist. Your HAR timestamps come from that clock.

Worse, log timestamps mean different things at different layers. Nginx writes its access log line when the response is complete. Envoy records start_time. ALB access logs carry request_creation_time and are emitted at completion. Application frameworks usually log at handler entry. So if a /token call takes four seconds, the layer that logs at completion appears after the layer that logs at start — even though the outer layer strictly contains the inner one. Sort a slow login by raw timestamp and you will produce a sequence in which the response precedes the request, and you will spend an hour theorizing about a race that does not exist.

The rule: normalize everything to start time before ordering. Where a duration field exists, start = logged_time − duration. Where it doesn't, note the layer's convention explicitly in your runbook. And never order across layers by timestamp when a causal key is available — the redirect chain is a causal order, and it is authoritative in a way that clocks are not.

There is a cheap trick for measuring skew retroactively. Every HTTP response carries a Date header from the server. If the browser's HAR records it — it does — then every entry in that HAR contains both the client clock (startedDateTime) and the server clock (Date) for the same instant, to one-second resolution. A HAR from an affected user is therefore a skew measurement instrument, per request, after the fact. If you also echo your request ID in a response header, that same HAR becomes the join between the browser's world and your server logs, which is otherwise the hardest join in the stack.

And sampling makes absence meaningless. Edge logs are frequently sampled; distributed traces are head-sampled at 1% or lower. "There is no record of that request" and "that request did not happen" are different statements, and at the edge you usually cannot distinguish them. The fix is narrow and affordable: never sample the authentication endpoints. /authorize, /token, /callback, the ACS, the introspection endpoint. Identity is a small fraction of total request volume in almost every system, and full-fidelity logs on those specific paths cost little and convert a category of unanswerable questions into queries.

The failures that live at the seams

Some bugs are visible only from a layer that cannot see the cause. These are the expensive ones, and they cluster at four seams.

Browser ↔ edge. Cookies not sent (SameSite, Secure over a mixed-content hop, a Domain that doesn't match, a Path that doesn't cover the callback). CORS preflights that fail, so the real request never leaves the browser. CSP blocking a redirect target. Anything returned in the URL fragment. Every failure in this class has the same signature: the app reports something missing, and no server anywhere logged the request that would have contained it.

Edge ↔ identity. Header rewriting that changes what identity derives from the request — a Host header rewritten by an ingress, so a computed redirect_uri or ACS URL no longer matches what was registered. Size limits: a WAF body limit that rejects a large SAMLResponse, or a proxy's 8 KB header limit that a long Cookie header exceeds and gets a 431 for. This is the seam that produces the single most misdiagnosed identity bug in the enterprise: login fails for exactly the users who are in the most groups. Their assertion is bigger, their session cookie is bigger, their token is bigger, and somewhere a limit sits between the median user and them. It presents as "flaky for 3% of users," it correlates with no browser, no region, and no time of day, and it is completely invisible to the identity platform, which produced a perfectly valid artifact that a proxy then refused to carry.

Identity ↔ resource server. Clock skew on nbf/iat. Audience and issuer mismatch. A JWKS cache that hasn't picked up a rotated key — which fails only on the nodes whose cache is cold, i.e. intermittently, i.e. the worst possible failure shape. Token size against the 4 KB cookie limit, where a session cookie gets chunked or silently truncated. Note that this seam has the same underlying variable as the previous one: per-user data size. Cardinality of group membership is the single best predictor of which of your users will hit seam bugs, and it is worth having that number in your user table.

Resource server ↔ browser. The wrap-around seam, and the most under-instrumented. An SPA receives a 401, silently starts the auth flow again, receives a token that fails the same check, and loops — generating a steady stream of perfectly successful logins in the identity platform's metrics. Login success rate goes up during this incident. The only place it looks wrong is in the browser, where the same user authenticated forty times in two minutes.

A triage procedure

The ordering matters more than any individual step. Each one either localizes the failure to a layer or eliminates half the search space.

0. Refuse to debug an aggregate. "Logins are failing" is not a debuggable statement. You need one user, one tenant, one connection, one approximate timestamp, and one flow. If you cannot get that, getting it is the work — running an identity platform on call covers why the slice, not the platform, is the unit of investigation.

1. Ask whether the flow reached the identity platform at all. One query against the authorize endpoint for that tenant and window. This is the highest-value single question in the procedure because it splits the problem cleanly: if there is no record, the failure is at layers 1–2 and no amount of identity debugging will find it. If there is a record, the failure is at layers 3–4 or at the seam below.

2. Capture browser evidence before it evaporates. A HAR with preserve log enabled, the full final URL including the fragment, and the browser version. This is perishable in a way server logs are not — the user will clear their cookies, or reboot, or "try again in incognito," and the state that produced the bug is gone. Ask first, theorize later.

3. Walk the redirect chain hop by hop and check the join keys. For each hop: was the expected parameter present, and did it match what the previous hop issued? The first hop where the chain breaks is your seam. This is the single most reliable technique in the whole article, and it requires no tooling beyond a text editor. If the broken hop is a SAML one, the per-error decision tree in Debugging SAML: A Field Guide takes over from here.

4. Normalize the clocks before you order anything. Start times, not log times. Then apply causal order, not timestamp order, wherever the protocol gives you one.

5. Name the seam, then ask what each side assumed. Once you have localized to a boundary, the question stops being "what is broken" and becomes "what did the upper layer assume the lower one had done." That is the assumptions register from 12.11, and this is where it earns its keep.

6. Check for the size correlation. If the failure is intermittent and correlates with no obvious dimension, compare the group membership count, token size, or cookie size of an affected user against an unaffected one. It costs two minutes and it resolves a surprising fraction of "flaky" reports.

Symptom → layer → the one discriminating piece of evidence

Symptom Most likely layer The evidence that discriminates
Redirect loop, no server errors 1 ↔ 3 seam HAR: is the session cookie sent on the second /authorize? Absent means cookie policy; present means the identity platform isn't recognizing it
invalid_state / InResponseTo mismatch 1 ↔ 3 seam HAR: Cookie header on the callback request, plus Sec-Fetch-Site
Works in one browser, not another 1 Browser version cohort, not tenant cohort. Diff the Set-Cookie attributes against that version's defaults
Fails only for some users, no pattern 2 ↔ 3 or 3 ↔ 4 seam Group count / assertion size / cookie size of affected vs unaffected user
Intermittent 401 from the API only 3 ↔ 4 seam Which node served it. Cold JWKS cache and skewed clocks are per-node
Destination/redirect_uri mismatch 2 ↔ 3 seam The Host and X-Forwarded-* headers as identity received them, versus what the browser sent
Assertion valid, app never saw it 2 ↔ 3 seam Edge response status and body size. A 413/431/414 with no identity log line is the signature
Login succeeds, user lands unauthenticated 3 ↔ 4 seam Compare sub/sid in the token against the key the app uses to look up the session
Success metrics up, users complaining 4 ↔ 1 seam Authentications per user per minute. A silent retry loop looks like healthy traffic
Nothing anywhere, request never arrived 1 HAR is the only witness. CORS, CSP, mixed content, or a fragment-borne error

The column that does the work is the third one. Most triage runbooks list symptoms and probable causes; the useful artifact lists the one observation that distinguishes two hypotheses, because the failure mode of cross-layer debugging is not having no theory, it is having four plausible theories and no way to eliminate three of them.

What to build before the next incident

Everything above is recoverable after the fact except the instrumentation, so a short list of what actually pays.

Echo the request ID back to the client. One response header. It makes every future HAR joinable to your server logs, which is otherwise the hardest join you have.

Put your trace ID inside state, signed. Described above. It is the only correlation carrier that survives the redirect chain intact.

Log Sec-Fetch-Site, Sec-Fetch-Mode and the presence-and-length of the Cookie header at the edge. Not the cookie value — the length. This makes both the SameSite class and the header-size class visible server-side.

Log absences with their expectation. "Expected flow state abc123, received none" is a debuggable sentence; invalid_state is not. This is the same argument immutable audit logs makes about recording what changed rather than that something changed, applied to the request path.

Turn off sampling on the auth endpoints. Cheap, narrow, and it converts "no record" from ambiguous to meaningful.

Record the layer's own transformations. If the edge rewrote a header, that is an event. Almost nobody logs it, and it is the root cause of an entire seam's worth of bugs.

One architectural note, disclosed as an interested party: at ClavionX the runtime and the control plane are deliberately separated — the runtime never calls the control plane synchronously, consuming projected state via events instead. That is the right call for availability, and it has a debugging consequence worth naming rather than glossing: a configuration change and its effect on a login are two records, in two systems, on two clocks, separated by propagation lag. "The config was correct at the time of the failure" and "the runtime had received the config at the time of the failure" are different statements, and any architecture with an async control plane needs the projection version to appear in runtime logs so you can tell them apart. Good boundaries create this class of question; the answer is to instrument the seam, not to remove it.

Treat a HAR file like a credential

A HAR contains every cookie and every Authorization header from the captured session. Session cookies in it are live until they expire. A bearer token in it is usable by anyone who has the file — and these files routinely get attached to support tickets, pasted into shared drives, and forwarded to vendors.

Chrome now offers a sanitized HAR export that strips cookies and auth headers. Know which one you asked for, because the sanitized version also strips the evidence for the entire first category of failures in the table above. The workable policy: ask for the full export, treat it as a secret in transit and at rest, invalidate the session it captured once you are done, and delete it. If your process cannot do those four things, ask for the sanitized export and accept that some bugs will stay unfound.

The underlying point

The instinct when a user is blocked and every layer is green is to look harder at the layers. It is the wrong instinct. Each layer is telling the truth about the only thing it can see, and the failure is in the space between two of them — which means the answer is never in a log, it is in a comparison between two logs that share no key.

So the work is mostly making that comparison possible in advance: a join key that survives a redirect, clocks you can normalize, absences logged with their expectations, and one piece of evidence per hypothesis that can eliminate it. None of that is sophisticated. All of it has to exist before the incident, because the alternative is what happened at the top of this article — three days, four green dashboards, and a fix that came from a header that wasn't there.