SAML Certificate Rotation Without an Outage
At 08:47 on a Monday, 4,000 people at one of your customers cannot log in. Every SSO attempt fails with a signature validation error.
At 08:47 on a Monday, 4,000 people at one of your customers cannot log in. Every SSO attempt fails with a signature validation error.
Nothing was deployed. Nothing changed on your side over the weekend. What happened is that the customer's IdP signing certificate reached its notAfter date at 00:00 UTC, three years and one day after somebody generated it during an implementation project that has since been reorganized out of existence.
The person who generated it has left. The certificate's expiry was in nobody's calendar. Your side has an alert on your own certificates and none on theirs. The customer's IdP administrator, who can fix this in four minutes, is in a timezone eight hours behind and will not be awake for three hours.
This happens constantly. It is the single most predictable outage in enterprise identity — the failure has a date on it, known years in advance, printed inside the artifact itself — and the industry keeps having it. So the interesting question isn't how to rotate a certificate. It's why an event with a known date keeps arriving as a surprise, and what design removes the surprise.
Why this is harder than rotating your own keys
Rotating your own signing key is a solved problem. You publish a JWKS with both keys, consumers pick up the new one, you switch, you retire the old. You control every step and you can observe the result.
SAML certificate rotation has three properties that make it categorically different:
It requires coordinated action by a party you cannot instruct. The certificate lives on the customer's IdP. Installing a new one is their change, on their change-management schedule, by an administrator whose priorities you don't set and whose contact details may be three job changes out of date.
It's N-to-M. Your SP has one signing certificate, consumed by every customer's IdP. Each customer's IdP has its own, consumed by you. With 300 SAML customers you have 301 certificates to track and 300 relationships to coordinate through, each with a different administrator, a different change process, and a different level of engagement.
The failure is total and instant. There is no degradation. At the moment of expiry, every assertion from that IdP fails validation and 100% of that customer's users are locked out — usually at the start of their business day, since the expiry is at midnight UTC.
The mechanism that makes it safe: two certificates at once
Everything below rests on a single capability, and if your implementation lacks it, nothing else in this article will help.
Your SP must be able to trust two (or more) IdP signing certificates simultaneously, for the same connection.
With overlap, rotation is uncoordinated and safe:
- Customer's IdP admin generates the new certificate. Tells you, or publishes it in metadata.
- You add it to the connection. Both the old and new certificates are now trusted.
- The IdP switches to signing with the new key, whenever it likes — no window, no coordination, no downtime. Assertions signed by either key validate.
- You observe which certificate is actually being used (see below).
- When the old certificate hasn't been seen for a couple of weeks, you remove it.
Without overlap, rotation is a synchronized change: the IdP switches at time T and your SP must switch at exactly the same instant, across a boundary you can't coordinate closely. That is why "rotation requires a maintenance window" — the window exists solely because the implementation can only hold one certificate.
The same is true in the other direction for your own SP signing certificate (used for signed authentication requests and signed logout messages): the customer's IdP must be able to hold two of yours. Many can. Some can't, and for those, your rotation is a coordinated change, which is worth recording per connection so you know in advance which customers are going to be difficult.
The observability step is the one most often missing and it's what converts removal from a guess into a decision: log the certificate thumbprint that validated each assertion. Then "has the old certificate stopped being used?" is a query, not a belief. Without it, step 5 is someone deleting a certificate and hoping, which means in practice nobody deletes it, and you accumulate trusted certificates forever — including ones whose private keys are on decommissioned hardware in a datacenter nobody has visited.
Metadata refresh: the automation that's usually left off
SAML metadata is a document containing entity IDs, endpoints, and certificates, published at a URL. The protocol's intended design is that both sides fetch the other's metadata periodically, so certificate changes propagate automatically.
In practice, most integrations are configured by someone pasting a certificate into a form once, and the metadata URL — if it was ever recorded — is never fetched again. Which converts an automatable event into a manual one, at scale, forever.
If the customer's IdP publishes metadata at a stable URL, use it:
Fetch on a schedule — daily is ample — and diff. A new certificate appearing in metadata should be added to the trusted set automatically. This alone eliminates the majority of expiry incidents, because most IdP administrators do install new certificates ahead of expiry; they just don't tell you, and their metadata reflects the change while your configuration doesn't.
Add, don't replace, on automatic refresh. New certificates from metadata get added to the trusted set. Removals require human review, because a removal is a security-relevant narrowing and metadata fetch failures shouldn't be able to unpick your trust configuration.
Validate the metadata itself. Federation metadata can be signed; verify it if it is. Fetch over TLS with certificate validation, pin the expected entity ID, and treat an unexpected entity ID or endpoint change as an alert requiring review rather than an automatic update — because a metadata document that can silently change your ACS URL or your trusted issuer is a supply-chain path into your authentication.
Handle fetch failure by keeping the last known good. A metadata endpoint that returns a 500 or an HTML error page must never result in an empty certificate list. That failure mode — refresh succeeded, parsed nothing, wrote an empty set — is a way to cause the exact outage you were preventing.
For the customers who don't publish metadata, or publish it at a URL that requires authentication, or whose IdP is behind a corporate firewall your servers can't reach: those connections are manual, and the practical response is to know which ones they are, so your monitoring can treat them differently.
Make expiry a monitored metric, not a surprise
Every trusted certificate, in both directions, has a notAfter field. That's structured data you already possess. It should be a first-class metric, exported like any other.
The alerting ladder that works:
- 90 days: a low-priority ticket, assigned to an owner, per connection. Time for the customer's change process, which for a large enterprise genuinely can take two months.
- 60 days: proactive customer contact, naming the certificate, its expiry date, and what you need from them. This is where most rotations actually get done.
- 30 days: escalate to the account team. Now it's a commercial conversation about a scheduled outage, which is a conversation that produces action.
- 14 days: page someone internally. If it hasn't happened by now, someone senior needs to be personally chasing it.
- Expired: should be unreachable. If you get here, you have a customer-facing outage and a process failure to review afterwards.
Two additions that pay for themselves:
A dashboard sorted by soonest expiry, across all connections, both directions. The single most useful artifact in this whole area. It turns 300 invisible time bombs into a list with dates.
Alert on certificates you no longer see being used. The complement of the above. A trusted certificate that hasn't validated an assertion in 30 days is either a completed rotation whose old certificate should be removed, or a connection that has stopped working without anyone reporting it.
And the detail that catches people: alert on your own certificate's expiry with a longer lead time than theirs, because rotating yours requires action from every customer, and 300 change requests is a project rather than a task. Six months is not excessive for a widely-consumed SP certificate.
When it's already expired: the incident
Since it will happen anyway, have the runbook.
Confirm the diagnosis fast. Signature validation failures across all users of one connection, starting exactly at a UTC midnight, is the signature of this specific failure. Your logs should make this a ten-second determination — which means logging the certificate thumbprint and the specific validation failure reason, not "SAML error."
The fix is on their side. They install a new certificate; you add it to the trusted set. If you have metadata refresh, you may already have it.
Resist the tempting workaround. Somebody will suggest disabling signature validation for this connection "just until they fix it." That converts an outage into a total authentication bypass for that tenant, open to anyone on the internet who can find the ACS URL. The answer is no, and the reason to say it clearly and in advance is that this suggestion arrives at hour three of an outage, from someone senior, under pressure from a customer. Decide the answer now, not then.
There is one legitimate accommodation and it should be a designed feature, not an improvisation: accepting a recently-expired certificate for a strictly bounded grace period, per connection, enabled explicitly, logged loudly, with an automatic expiry on the exception itself. Signature validation still happens — the certificate's key is still proving authenticity — you are only relaxing the expiry check. That's a much smaller relaxation than disabling validation, and having it available as a switch with a 24-hour auto-revert is what prevents someone from reaching for the catastrophic option. Whether to offer it at all is a real judgment call; what's not defensible is having no plan and improvising one at 09:30.
Communicate with real content. "We are seeing SSO failures for your tenant. Your IdP's signing certificate expired at 00:00 UTC today. Here is its thumbprint and the expiry date. Please install a new certificate and send us the new one, or confirm your metadata URL and we'll pick it up automatically." Customers can act on that. "We are investigating an authentication issue" produces another hour of nothing happening.
The design that removes the class of problem
Stepping back — the operational fixes above are all mitigations for a design property. If you're building or choosing an identity platform, these are the capabilities that make certificate expiry a non-event:
- Multiple simultaneously-trusted certificates per connection, in both directions. Non-negotiable; everything else depends on it.
- Automatic metadata refresh with add-only semantics, signed-metadata validation, and last-known-good on failure.
- Per-assertion logging of the validating certificate thumbprint, so rotation completion is observable and stale certificates are removable with evidence.
- Expiry as an exported metric for every certificate in the system, with a laddered alert schedule and a cross-connection dashboard.
- A named owner per connection, internally and at the customer, kept current — because every mitigation above eventually resolves to "email a specific person," and the most common reason rotation doesn't happen is that nobody knows who to email.
- A bounded, logged, auto-expiring grace mechanism so that the option on the table during an incident is a small relaxation rather than a catastrophic one.
None of that is difficult. It's just unglamorous, and it competes for priority against features, right up until 08:47 on a Monday when 4,000 people can't work and the fix is in another timezone.
The certificate told you the date three years ago. The only real question is whether anything in your system was listening.