Debugging SAML: A Field Guide
SAML errors are uninformative on purpose, and knowing why is the first step to debugging them.
SAML errors are uninformative on purpose, and knowing why is the first step to debugging them.
A Service Provider that told you exactly which check failed would be handing an attacker a test harness. "Signature valid, audience wrong" confirms the forgery worked. "Audience correct, expired" confirms the audience guess was right. Every detailed failure message is a bit of an oracle, so implementations collapse fifteen root causes into one string: Invalid assertion. Or, memorably, Error 400.
The detail exists in the SP's server-side logs, which you probably can't see either, because the SP is a SaaS vendor and the assertion is in the customer's browser. So: symptoms, what they usually mean, what they occasionally mean instead, how to tell. Ctrl-F your error string.
First: how to actually look at an assertion
Do this before forming a theory.
The assertion is a base64 form field named SAMLResponse POSTed to your ACS endpoint. In browser devtools, enable Preserve log, run the login, find the POST, read the Payload / Request body. The value is URL-encoded there, so decode that first.
base64 -d resp.b64 | xmllint --format -
(xmllint ships in libxml2-utils on Debian/Ubuntu and libxml2 via Homebrew. If it's missing, python3 -c 'import sys,xml.dom.minidom as m; print(m.parse(sys.stdin).toprettyxml())' will do.)
Do not reformat the XML or paste it through anything that "cleans it up" if you intend to re-verify the signature afterward. Read on for why.
Look at six things, in order:
<saml:Issuer>— is this the IdP you think you're talking to?<saml:Conditions NotBefore= NotOnOrAfter=>— compare against your server's clock, in UTC.<saml:AudienceRestriction><saml:Audience>— byte-for-byte against your configured Entity ID.<saml:NameID Format=>and its value.<ds:KeyInfo><ds:X509Certificate>— which key signed this.<saml:Attribute Name=>— the actual attribute names, not the ones in the docs.
"Invalid signature" / "Signature verification failed"
org.opensaml.xmlsec.signature.support.SignatureException:
Signature cryptographic validation not successful
Signature validation failed. SAML Response rejected.
Usually: certificate mismatch. The IdP is signing with a key you don't have. They rotated and nobody told you — rotation is an identity-team calendar event, not a customer notification — or you were given the encryption certificate, or your metadata was stale on import.
How to confirm: fingerprint the cert from KeyInfo and the one in your config.
# X509Certificate body from the assertion, wrapped in PEM headers
openssl x509 -in from_assertion.pem -noout -fingerprint -sha256 -dates
Mismatched fingerprints and you're done — go get current metadata. -dates catches the variant where the cert is right but expired.
Occasionally: the bytes changed in transit. XML signatures cover a canonicalized byte stream, and canonicalization is exact. A proxy, WAF, or logging middleware that re-serializes the XML — normalizing whitespace, reordering namespace declarations, converting CRLF to LF — invalidates a cryptographically fine signature — as does your own pretty-printer, which is why people lose afternoons to signatures they broke while formatting the evidence. The tell: fingerprints match, verification still fails.
Also occasionally: the signature sits on the Response when your SP requires it on the Assertion, or vice versa. Both are legal. (And a valid signature still doesn't prove you're reading the signed element; that's XML Signature Wrapping.)
"Audience restriction" / "Invalid audience"
The audience 'https://app.example.com' is not valid for this SP
(expected 'https://app.example.com/')
saml2:Conditions AudienceRestriction validation failed
<saml:Audience> must exactly equal your SP's Entity ID. String equality, no normalization.
Usually: a trailing slash, or http vs https. Identical to the eye in a ticket and in an admin UI. Diff them programmatically.
Occasionally, and this is the one worth knowing: the IdP admin put your ACS URL in the Entity ID field, because Entity IDs look like URLs and one that resolves to your login page feels more correct than one that 404s.
It isn't. An Entity ID is an opaque identifier that happens to use URI syntax. It never needs to resolve. urn:example:app:prod is perfectly legal. The HTTPS-URL convention is a namespacing habit borrowed from XML, not a requirement that anything live there. Internalize this and arguments with IdP admins get shorter: you're not asking them to configure a link, you're asking them to copy a string.
How to confirm: diff <(echo -n "$FROM_ASSERTION") <(echo -n "$FROM_CONFIG"), or compare lengths. Trailing whitespace pasted from a PDF is real.
"Assertion not yet valid" / "Assertion has expired"
Assertion is not yet valid: NotBefore=2026-08-03T14:22:10Z now=2026-08-03T14:21:47Z
SAML assertion expired (NotOnOrAfter 2026-08-03T14:27:10Z)
Clock skew between the IdP and your server.
The tell is the failure pattern, not the message. Failing for everyone, always, is probably not skew — that's a stale assertion, a tab left open, a resubmitted cached POST. Failing intermittently, for some users and not others, or only in one region is skew: nodes disagree about the time and which one you land on is luck.
How to confirm: log NotBefore/NotOnOrAfter alongside your server's UTC time at validation, per node, then check drift directly with chronyc tracking or ntpq -p. A node 40 seconds fast rejects assertions every other node accepts.
Skew tolerance should be seconds, not minutes. Sixty is generous. Widening it to five minutes to stop an outage is understandable on a Friday, and it widens the replay window of every assertion you accept — precisely what the validity window exists to bound.
Login "succeeds" but creates a new user every time
No error string. That's what makes this one expensive.
NameID is the user's identifier; its Format says what kind:
<!-- what your SP expects -->
<saml:NameID Format="urn:oasis:names:tc:SAML:1.1:nameid-format:emailAddress">
[email protected]
</saml:NameID>
<!-- what the IdP is sending -->
<saml:NameID Format="urn:oasis:names:tc:SAML:2.0:nameid-format:persistent">
AAdzZWNyZXQxWJb2gAAAA...
</saml:NameID>
If your SP keys accounts on email and the IdP sends a persistent opaque identifier — or worse a transient one, designed to differ on every login — you get a new user row per login. Nobody files a bug, because login works. You find out when someone notices the user count, or a customer asks why their permissions reset. The mirror case: the IdP sends unspecified holding a bare username like alice, your SP matches on email, and every login arrives as a stranger.
How to confirm: log in twice, ten minutes apart, and compare NameID values. Different means transient, and no SP-side code fixes that — the IdP config must change.
Authentication works, authorization is empty
User authenticated but has no groups. Access denied.
KeyError: 'email'
Attribute name drift. Every vendor names the same concept differently, and none of them is wrong:
<!-- ADFS / Entra ID (Azure AD) -->
<saml:Attribute Name="http://schemas.xmlsoap.org/ws/2005/05/identity/claims/emailaddress">
<saml:Attribute Name="http://schemas.microsoft.com/ws/2008/06/identity/claims/groups">
<!-- Okta (default profile) -->
<saml:Attribute Name="email">
<saml:Attribute Name="groups">
<!-- Ping / SAML 2.0 X.500 convention -->
<saml:Attribute Name="urn:oid:0.9.2342.19200300.100.1.3"
FriendlyName="mail">
The third matters: urn:oid: names are standard in Shibboleth and academic federations, and arrive as noise if your mapping layer assumes names are human-readable.
How to confirm: dump every Name from the assertion and compare against your mapping config. Don't trust FriendlyName — it's optional and advisory, and IdPs omit it.
Occasionally: the attribute is correct, but multi-valued groups arrive as repeated <saml:AttributeValue> elements and your parser reads only the first. Everyone gets exactly one group — which reads as a permissions bug, not a parsing bug.
"Destination mismatch" / "Invalid recipient"
SAML Response Destination 'https://app.example.com/sso/acs' does not match
the expected destination 'http://internal-app-7:8080/sso/acs'
Usually: a load balancer or ingress rewriting the Host header. The IdP POSTs to the public URL, TLS terminates at the edge, and your app sees Host: internal-app-7:8080 with scheme http. It builds its expected ACS URL from those and compares against what the IdP correctly sent.
How to confirm: log the computed ACS URL per request. If it isn't the public HTTPS URL, the bug is in your proxy config, not your SAML library. Honor X-Forwarded-Proto and X-Forwarded-Host from trusted proxies only — or better, configure the ACS URL as a static string instead of deriving it from request context.
"It works in staging and fails in production"
Three causes, in descending frequency.
Different Entity IDs. Staging is https://staging.example.com, prod is https://app.example.com, and the IdP has one SP config someone updated in place. Shows up as an audience failure.
Different certificates. Separate IdP tenants, separate signing keys, and you copied working config forward.
Signature validation was disabled in staging. Someone turned it off during integration to unblock a demo, wrote "temporarily" in the commit message, and moved on. Staging validated nothing for eight months, so everything "worked." Production, correctly configured, rejects the same assertions. Check this first, because it inverts your intuition: staging isn't the working case, it's the case with no checks. Grep it for wantAssertionsSigned, validateSignature, wantMessagesSigned, or whatever your library calls it.
Before you file a ticket with the IdP team
Escalating to a customer's identity team is slow and socially expensive. Clear this list first.
- [ ] A decoded assertion from a failing login, not a successful one from last month.
- [ ]
Audiencecompared to your Entity ID as bytes, not by eye. - [ ]
KeyInfocertificate fingerprinted against your configured one. - [ ]
NotBefore/NotOnOrAftercompared to server UTC, plus NTP drift on the node that served the request. - [ ] Computed ACS URL confirmed to match what you published in metadata.
- [ ] Attribute
Namevalues listed, and the ones you need confirmed absent rather than merely renamed. - [ ] Failure reproduced on a second node.
- [ ] One sentence stating what the IdP must change — "the Audience must be
https://app.example.com, no trailing slash" — rather than "SSO is broken."
That last line turns a two-week ticket into a two-hour one. Identity teams run hundreds of integrations; a request naming the exact field and value gets actioned, a symptom report gets queued.
The part nobody likes
A meaningful fraction of these are unfixable from your side. Certificate rotation, NameID format, attribute release, the Entity ID in the IdP's config — all of it lives in a console you don't have, run by a team whose priorities aren't yours. You can diagnose it precisely. You cannot fix it.
So the value of the diagnostic pass isn't the fix. It's sorting the failure into mine (proxy config, attribute mapping, clock drift, ACS URL derivation), theirs (certs, NameID format, attribute release, audience value), or ours together (Entity ID agreement, metadata exchange, signing expectations). Getting that right in the first twenty minutes is most of the job. What follows is either a config change you can ship or an email you can write with confidence — and on a Friday afternoon, knowing which one you're in beats a fix you can't apply.