Running an Identity Platform On Call

The page arrived at 03:12. LOGIN_SUCCESS_RATE_LOW — tenant: northwind-fs.

The first thing I did was open the global dashboard, because that is what you do. Aggregate login success rate: 99.94%. Flat. Token issuance rate: normal for the hour. Error rate: 0.06%, which is where it always sits. p99 latency unremarkable. Database healthy, session store healthy, event lag under a second.

Everything was fine, and one customer's entire workforce could not log in.

Northwind was a mid-sized tenant on a federated SAML connection. Their identity provider's signing certificate had expired at midnight UTC. Every assertion since had failed signature validation. For them it was a hundred percent outage, ongoing for three hours, and their European shift had been walking into it since 06:00 local. For the platform it was 0.06% of requests — indistinguishable from the background rate of people typing their passwords wrong.

That gap, between what the aggregate can see and what a customer experiences, is the single most important thing to understand about being on call for an identity platform. Identity does not fail globally. It fails in slices. Almost every alerting design I have inherited was built for the global failure and is structurally blind to the slice, and almost every serious identity incident I have worked lived in a slice for its first several hours.

This article is the runbook I would hand a new on-call engineer: what should page, what should not, what the first five minutes look like, and which dashboards are worth having open at 3am.

The arithmetic that makes global SLOs blind

Take a plausible multi-tenant shape: 400 tenants, a long-tail distribution where the largest tenant is 8% of logins, the tenth-largest is about 1.5%, and the median tenant is around 0.2%.

Now fail one tenant completely and look at what the global number does.

Tenant's share of login volume Global login success rate when they are 100% down Detectable against normal variance?
8% (largest) 92% Yes, screamingly
3% 97% Yes
1.5% 98.5% Probably
0.5% 99.5% No — inside daily noise
0.2% (median) 99.8% No
0.05% 99.95% Not even theoretically

The normal day-to-day variation in aggregate login success rate is on the order of half a percent, driven by things that have nothing to do with your health: a customer's onboarding week, a mobile client release that retries differently, Monday morning versus Friday afternoon password entropy. So a global alert has to sit somewhere around 99% to avoid firing constantly.

At a 99% threshold, you can detect the total failure of your top ten tenants and nothing else. 390 of your 400 customers can each experience a complete, unrecoverable outage without moving a needle you look at. That is not a tuning problem. No threshold on that metric fixes it, because the signal you want is smaller than the noise floor of the metric you are watching. Aggregates average across exactly the dimension where identity failures live.

So the unit of alerting is not the platform. It is the slice — and the slice that matters is login success rate per tenant per connection.

Why "per connection" and not just "per tenant"

Because tenants routinely have several authentication paths and they fail independently. A typical enterprise tenant has a SAML connection to their corporate IdP for employees, an OIDC connection to a partner directory, and local password authentication for contractors and service accounts.

If the SAML connection breaks, that tenant's success rate drops to whatever fraction of their logins came through the other two — say 65% down, 35% still working. A per-tenant alert with a 50% threshold catches it. A per-tenant alert tuned to catch quieter breakages fires on their normal weekday mix shifts. Meanwhile the actual fact — one connection is returning zero successes — is unambiguous, needs no threshold tuning, and points directly at the cause.

Connection-level slicing also gives you the failure mode you cannot otherwise see: the connection that has always been low-volume and has been broken for a week. Nobody reported it because the twelve people who use it assumed it was them.

Small slices need probes, not ratios

Here is where the arithmetic bites a second time. Slicing finely enough to see the failure means slicing finely enough that most slices have almost no traffic.

A median tenant doing 0.2% of a million daily logins is 2,000 logins a day — roughly 80 an hour at peak and, at 03:00 in their timezone, quite possibly zero. You cannot compute a success rate over zero events. A ratio alert on an empty slice either divides by zero, holds its last value, or silently evaluates to "healthy," and every monitoring system I have used defaults to one of the last two.

The consequence: rate-based alerting has a volume floor, and most of your slices are below it. You will find out about their outage when their business day starts and a human opens a ticket.

The fix is to manufacture the traffic. A synthetic probe per connection, running end to end on a fixed interval, converts a statistical problem into a deterministic one. Detection time stops depending on customer volume and starts depending on your probe interval, which you control.

The cost is trivial and worth spelling out because people assume it is not. 400 connections probed once per minute is 6.7 logins per second of synthetic traffic — for most platforms a rounding error against real load, and cheaper than the two-minute-interval variant people propose to save money. Alert on N consecutive failures, not on a rate: three consecutive failures at 60-second intervals gives you three-minute detection on every connection equally, whether that customer does a million logins a day or eleven.

Four things to get right, all of which I have seen get someone paged for the wrong reason:

  • The probe must traverse the real path. A probe that calls an internal health endpoint tests your opinion of yourself. A probe that completes an actual authorization code flow against the real connection, exchanges the code, validates the resulting token against your public JWKS, and then cleans up the session tests the thing customers do. The difference shows up in exactly the incidents that matter — a broken redirect URI, a wrong kid, a clock skew on one pod (one of the seven bugs everyone ships).
  • Probe accounts need explicit, audited exemptions. They will trip lockout thresholds, MFA policies, impossible-travel rules, and password expiry. Exempting them is fine; leaving that exemption undocumented and un-reviewed is how a probe account becomes the most privileged unmonitored credential in the system.
  • Federated probes fail for reasons that are not yours. When the customer's IdP takes a maintenance window, your probe fails all night. If that pages you, you will learn to ignore the probe alerts, which is the worst outcome in this article. Per-connection maintenance windows and a distinct alert class for "remote endpoint unreachable" versus "our validation rejected the assertion" are not optional refinements; they are what keeps the probe credible.
  • Probes must run from where users are. A probe inside your VPC will not see the DNS failure, the CDN misconfiguration, or the regional network partition that is the whole incident.

What should page and what should ticket

The table is the deliverable. The reasoning column is the part that survives contact with a real incident.

Signal Disposition Why
Login success rate for one tenant/connection drops below its own baseline (or N consecutive probe failures) Page Total outage for a real customer. Invisible globally. This is the backbone alert.
JWKS endpoint unavailable or serving a key set missing an active kid Page immediately It breaks every resource server that validates tokens, including for users already logged in. The widest blast radius available in the entire system, and it can be caused by a cache purge or a bad deploy.
Token issuance p99 breaches budget for more than a few minutes Page Issuance is synchronous in every login and every refresh. A p99 breach here is a queue forming; queues in identity go metastable fast (tail latency).
Token issuance rate drops sharply against its weekly baseline Page The best early warning you have. Silence is a symptom — clients failing before they reach you look like nothing at all on error dashboards.
Authentication error rate spike on one dependency (IdP, directory, MFA provider) Page Your degradation posture may be about to be exercised. You want a human watching.
Session store memory above ~85% or eviction rate non-zero Page Evictions are silent logouts. Users experience it as random session loss, and it will be blamed on the last deploy.
Event/projection lag beyond your stated staleness bound Page at the bound, ticket below it Below the bound it is normal operation. Beyond it, config and revocation changes are silently not taking effect — a security-relevant condition, not a performance one.
Audit pipeline lag Ticket, then page at the buffer threshold Lag alone is a backlog. Lag approaching the bounded buffer's capacity means you are about to start dropping records or fail closed on privileged operations — that is a compliance event with a clock on it. Page on time-to-buffer-full, not on lag.
Signing key approaching expiry or rotation overdue Escalating ticket Fully predictable. If it ever pages, your rotation automation is what actually failed.
Federated certificate expiring in 30 days Ticket, on a ladder Someone else's change process is the long pole. The full ladder belongs in SAML certificate rotation.
SCIM sync failure for one tenant Ticket Provisioning is not the login path. A stalled sync means new joiners wait; it does not lock anyone out. Escalate if it is a deprovisioning backlog — that is access outliving employment.
Replication lag Depends — see the rule below The most commonly miswired alert in the set.
Elevated 4xx on the token endpoint Ticket Usually one client with a broken redirect URI or an expired secret. Real, not urgent, and it will page you nightly if you let it.

The replication lag rule. Replication lag matters exactly as much as your read path depends on the replica, and not at all otherwise. Write it as a conditional, not a threshold:

  • If sessions or token state are read from a replica: lag is correctness. A user logs in against the primary and their next request lands on a replica that has not seen the session. Page at single-digit seconds.
  • If only projected configuration is read from replicas: lag is bounded staleness, which you already accept by design. Ticket, until it exceeds the staleness bound you published — then page, because revocations are not landing.
  • If lag is growing monotonically: page regardless of the absolute number. A replica that is 4 seconds behind and holding is healthy. A replica 4 seconds behind and drifting will be 400 seconds behind by breakfast, and the shape of the curve is the signal, not the value.

The first five minutes

The mistake everyone makes, including experienced engineers, is to start diagnosing the cause. It feels like progress. It is the wrong first move here, because identity incidents are usually narrower than they appear — the page says the auth system is unhealthy, the support channel says "everything is down," and the truth is one connection out of four hundred.

Blast radius is not a detail you establish on the way to the cause. It changes what you do next: who you tell, whether you can roll back, whether the fix is a config change on one connection or an all-hands, whether a contractual notification clock has started. Getting the radius wrong in the first five minutes costs an hour later.

So the order is fixed.

flowchart TD
    P["Page fires"] --> R{"Blast radius?"}
    R -->|"One connection"| C1["Federation / cert / metadata<br/>Their side or yours"]
    R -->|"One tenant, all connections"| C2["Tenant config, projection,<br/>rate limit, suspension flag"]
    R -->|"One region"| C3["Regional dependency<br/>Check the other region first"]
    R -->|"Everyone"| C4["Shared path: JWKS, signing keys,<br/>session store, deploy"]
    C1 --> D["Was it our change?<br/>Deploy + config timeline"]
    C2 --> D
    C3 --> D
    C4 --> D
    D -->|"Yes"| RB["Roll back. Diagnose after."]
    D -->|"No"| DEP["Check the layer below<br/>DB, cache, IdP, network"]
    DEP --> POS["Choose degradation posture<br/>deliberately"]
    POS --> COM["Communicate radius,<br/>not cause"]
</antml_diagram_placeholder>

Minute 0–1: Establish blast radius. One connection, one tenant, one region, or everyone? Your login-success-by-tenant panel answers this in one look if it exists, and if it does not, this step takes twenty minutes and you should build the panel tomorrow. Note that "everyone" is the rarest answer and should make you suspect the shared path immediately: JWKS, signing keys, session store, the load balancer, DNS.

Minute 1–2: Was it us? Open the deploy and config-change timeline before anything else. Not because engineers are careless, but because the base rate is overwhelming: the large majority of identity incidents are self-inflicted — a deploy, a config push, a certificate rotation, a policy change, a feature flag, a rate-limit adjustment. If a change landed within the last hour anywhere near the blast radius, treat it as the cause until disproven and roll it back. You are not required to understand the mechanism first. Roll back, restore service, diagnose in daylight. The one exception is a change you cannot roll back — a session format change or a completed database contraction — which is why blue-green deploying an identity platform spends so much of its time on what is reversible.

Minute 2–3: Check the layer below. Database, cache, session store, event pipeline, the customer's IdP, DNS, the network. Two failure shapes to distinguish, because they need opposite responses: down (fast errors, circuit breakers open, degradation path engaged) and slow (everything technically succeeding, connection pools exhausting, latency climbing). Slow is worse and much easier to misread as an application problem. If your dependency dashboard shows healthy-but-slow, you are in a queueing incident, not an error incident, and the fix is shedding load rather than adding capacity.

Minute 3–4: Choose the degradation posture, deliberately. Read-only mode, disable the failing connection and let its users fall back to a secondary method, shed load, extend cache TTLs to ride out a backend outage. Each of these is a decision with a security consequence, and this is not the moment to invent the policy — it is the moment to execute the one you wrote down. The per-operation reasoning belongs to what an identity platform should do when its database is down; on-call's job is to know which row applies and where the switch is.

Minute 4–5: Communicate radius, not cause. "Authentication is failing for tenants on federated connections; existing sessions are unaffected; we are investigating" is a good five-minute update. "We think it's the database" is a bad one, because it is a guess that will be quoted back at you and, if the incident turns out to involve any correctness question, will end up in a customer's written record. State what is broken, for whom, what still works, and when you will update next.

Only now do you go find the cause — and if it spans layers, correlation-ID tracing is the tool, which debugging across four layers covers properly.

The dashboards that matter at 3am

A 3am dashboard is not an analysis tool. It is a triage instrument for someone with degraded judgment who was asleep ninety seconds ago. Six panels. If a panel has never answered a question during an incident, it belongs on a different board.

  1. Login success rate, by tenant and connection, sorted by largest drop. The blast-radius answer. Sorted by delta from that slice's own baseline, not by absolute rate — a connection that normally sits at 96% because of a chronically fat-fingered user population is not the interesting row.
  2. Token issuance rate and latency percentiles. Rate against a weekly-shape baseline catches the silent failures where clients never reach you. Percentiles catch queue formation. Show p50 and p99 together: p50 flat with p99 climbing is contention or a bad node; both climbing is saturation.
  3. Session store health. Memory used against limit, eviction rate, connection count, command latency. Evictions above zero explains a support queue full of "it logged me out" that nothing else on the board accounts for.
  4. Event and projection lag. How far behind is the projection that the runtime reads? This is where config changes, revocations, and tenant suspensions go to be silently ignored — and it is a security signal, not a performance one. (Every identity team eventually builds an event bus covers why this pipeline exists at all.)
  5. Dependency error and latency rates, one row per dependency, error rate and p99 side by side, so that "up but slow" is visible as a distinct state from "down."
  6. The deploy and config-change timeline, overlaid on all of the above.

That last one is the highest-yield panel on the board and it is missing from most of them. The argument is a base-rate argument: if most identity incidents are self-inflicted, then the single most informative fact at minute one is what changed and when. An overlaid change timeline collapses "did something change?" from a cross-team Slack archaeology exercise into a visual coincidence you can see without reading. It has to include everything that can alter behaviour — application deploys, config pushes, feature flags, certificate rotations, policy edits, rate-limit changes, IdP metadata refreshes — because the ones that cause incidents are disproportionately the ones nobody counted as a deploy. Vertical lines on every graph. When the drop starts one minute after a line, you are done in ninety seconds.

The complement is why identity systems fail gradually and then all at once: some incidents have no line, because the change was three weeks ago and a threshold was crossed tonight.

Why identity incidents escalate differently

Four things are true of identity incidents that are not true of a normal service outage, and they change your behaviour in the first ten minutes. Every identity bug is a trust bug develops this argument properly; here is the operational residue.

They arrive from every channel simultaneously, with contradictory descriptions, because everyone is affected and nobody can see the shared cause. Establishing that fourteen reports are one incident is pure overhead you pay before triage begins. This is the practical reason the blast-radius-first ordering wins: it is also the fastest way to reconcile the reports.

Your counterpart becomes the customer's security team, not their IT team. They are not asking for an ETA. They are assessing whether you remain an acceptable vendor, and they are writing the answers down.

A contractual or regulatory clock may already be running. Unauthorized-access notification clauses start at awareness, not confirmation. The moment anyone says the words "cross-tenant," the incident has a legal deliverable attached to it. Your process needs a notify-on-suspicion step, and on-call needs to know it exists.

And the console you would use to fix it may be behind the thing that is broken. This is the one that produces genuine disasters. If the admin UI authenticates through the platform, and the platform is down, you cannot disable the broken connection, flip to read-only mode, or roll back a config. Break-glass access must not depend on the system it is used to fix — a separate authentication path, separate credentials in a separate store, a direct-to-database or direct-to-config-store escape hatch, heavily audited and alerted on use.

I work on ClavionX, where the runtime never synchronously calls the control plane, and I will offer that split as an operational illustration rather than a pitch: because the two planes are separate failure domains, a control-plane outage does not take out token validation, and — more relevant at 3am — a runtime incident does not take out the plane you administer from. Whatever the architecture, the property to verify is the same one, and the way to verify it is to actually try it: log into your break-glass path, quarterly, with the primary path deliberately unavailable. An untested break-glass credential is a rumour.

The counterpoint: over-alerting is the worse failure

Everything above argues for more alerts, more slices, more probes. Taken naively, that argument ends somewhere bad: 400 connections generating slice alerts, a page for every ticket-worthy condition, and an on-call engineer who has learned that pages are usually nothing.

That engineer is a worse outcome than having no alert at all, because a missing alert is a known gap and a learned-ignorable alert is a gap you believe is covered. Alert fatigue is not laziness. It is correct Bayesian updating on a bad prior, and it is fully your fault as the person who configured the page.

The discipline that keeps the list short:

  • Every page must have an action the person woken up can take right now. If the runbook step is "acknowledge and watch," it is not a page. It is a ticket with an anxious tone.
  • Page on symptoms users feel, ticket on causes. Users feel "cannot log in." They do not feel "cache hit rate is 71%." Cause-metrics are for the dashboard you open after the page; as alerts they are a permanent source of false positives.
  • Roll up correlated slices. If forty connections fail at once, that is one page saying "forty connections failing," not forty pages. Design the rollup before you deploy per-slice alerting, or the first shared-dependency outage will demonstrate why.
  • Consecutive-failure and duration thresholds, not instantaneous ones. Almost nothing in identity requires action within thirty seconds; almost everything worth waking someone for persists for three minutes.
  • Review every page monthly and delete on evidence. For each alert that fired: did a human take an action that changed the outcome? If the answer was no twice, delete it or demote it to a ticket. The most valuable thing an on-call rotation produces is not the incidents it resolves — it is the list of alerts it proved were useless.

The candidates for deletion are predictable: CPU and memory alerts on stateless pods (the orchestrator handles it), cache hit rate alerts, individual pod restarts, 4xx rate alerts on the token endpoint, disk space on ephemeral nodes, and any alert whose name contains the word "anomaly" without a defined action. That list is where most of the pages in an untended identity platform come from, and none of them describe something a customer can feel.

The honest tension: per-connection alerting is a lot of alerts, and the discipline above is the only reason it stays survivable. If you cannot commit to the rollup and the monthly deletion pass, deploy per-connection probes feeding a dashboard and a ticket queue first, and promote to paging only the slices where a customer contract makes it worth waking someone. Partial coverage you trust beats full coverage you ignore.

The pre-incident work that actually pays

None of this is buildable at 3am. Four things, in order of return:

Per-tenant, per-connection synthetics. The single highest-leverage investment available. It converts the entire class of silent slice outages from customer-reported to self-reported, and customer-reported identity outages cost you credibility on a different scale than they cost you availability.

A written degradation posture per dependency, agreed with someone who can sign off on the security consequence, and rehearsed. Deliberate fault injection is how you find the rows that are wrong — chaos engineering for identity systems is the discipline for that.

Tested break-glass access. Quarterly, on the calendar, with the primary path deliberately unavailable. Every organization believes it has this. Roughly half do.

Runbooks that name the exact query. Not "check projection lag" — the metric name, the dashboard link, the query, the expected value, and the threshold at which you escalate. A runbook that requires the reader to already know the system is a document for someone who does not need it. The test is whether the second-most-junior person on the rotation can execute it alone at 3am, and the only way to find out is to have them try during business hours.

Where this lands

The through-line is one idea: aggregate health is not health. An identity platform can be at 99.94% and be completely down for the customer whose renewal is next month. The metric that tells you the truth is login success rate per tenant per connection, and for most slices you will have to manufacture the traffic to measure it.

Everything else follows. Alert on slices with a rollup so it stays survivable. Page on what users feel, ticket on causes, and delete the pages that never earned an action. Establish blast radius before cause, because the radius determines the response and identity incidents are usually narrower than the noise around them suggests. Check your own change timeline before you suspect anyone else's system, because it is usually you. Decide the degradation posture in a design review, not in an incident channel. And make sure the door you break the glass on is not the one that is locked.

When something does own the seam between your layers and nobody's dashboard shows it, that is a different problem with a different fix — who owns the incident when every layer worked correctly is the one to read next.