How to Evaluate an Identity Vendor Without Reading Marketing Pages
Every identity vendor supports SAML, OIDC, SCIM, MFA, and audit logging. Every one of them will tick every box on your RFP. Every one has a page that says "enterprise-grade" and a diagram with a shield on it.
Every identity vendor supports SAML, OIDC, SCIM, MFA, and audit logging. Every one of them will tick every box on your RFP. Every one has a page that says "enterprise-grade" and a diagram with a shield on it.
This is not because vendors are dishonest. It's because feature checklists ask the wrong question. "Do you support SAML?" has one commercially viable answer, and the interesting variation is entirely in the part the question doesn't reach: whether the SAML implementation supports the eleven attribute-mapping quirks your customers' IdPs will present, whether a misconfigured assertion produces an error message a human can act on, and whether adding a connection takes four minutes or four days.
So the goal of an evaluation is not to establish what a platform can do. It's to establish what operating it will feel like in year three, when you have 400 tenants, a compliance auditor, a rotation deadline, and a customer whose IdP does something the spec permits but nobody expected.
Below are the questions that actually differentiate, organized by what they reveal. I work on an identity platform, so read this with appropriate suspicion — and to make that easier, I've marked the questions where my own product's answer is a no, because a list of questions that only flatters the author's product isn't an evaluation framework, it's a brochure with question marks.
Ask operational questions, because features converge and operations don't
"Can you rotate a signing key without a maintenance window? Walk me through it."
The best single question in the set. Every platform will say yes. The follow-up is what matters: how long do both keys coexist, is the JWKS published with both kids throughout, do I initiate it or do you, do I get notified before you rotate yours, and what happens to a relying party that cached the JWKS six months ago?
What you're really testing is whether the platform was designed with the assumption that keys change. Platforms that weren't have a rotation procedure that is a support ticket and a scheduled outage, and you will discover this the week a compliance finding requires rotation.
"When I change a setting, when does it take effect — and how do I know it has?"
Configuration propagation is the most under-asked question in identity procurement. A multi-region platform has a control plane and a runtime, and the delay between them is real: 30 seconds, 5 minutes, or "eventually." That number determines your incident response time. If you disable a compromised client, "eventually" is not an acceptable answer to a security team.
Ask specifically: is there a way to observe that a change has landed everywhere? A platform that gives you a propagation status or a version counter has thought about this. One that says "it's fast" has not.
"What happens when your control plane is down? Be specific about which operations fail."
The right answer is a decomposition, not a reassurance. Authentication continues from cached configuration; new configuration changes cannot be made; new tenant provisioning fails; existing sessions are unaffected. A vendor who can describe this crisply has designed a degradation model. A vendor who says "we have 99.99% uptime" has answered a different question — and note that 99.99% is 52 minutes a year, which for a component every login depends on is a number worth actually thinking about rather than nodding at.
Then ask the same question about your dependencies: what happens when your identity provider's region is unreachable from mine? What's the failover, is it automatic, and does it lose sessions?
"What's in an audit event? Show me the JSON."
Do not accept "comprehensive audit logging." Ask to see a real event for three specific operations: a successful login, a permission change, and a configuration change.
What you're looking for: actor identity (and separately, subject identity — they differ constantly), before-and-after state on changes rather than just an event name, a correlation ID that spans the request, the source IP, and a timestamp with timezone and sub-second precision. An audit trail that says permission_updated without recording what it was before is not usable in an investigation, and the moment you need it is the moment you find out.
Then: how do I get them out, in bulk, continuously, and how long do you keep them? A UI you can search is not an audit capability. You need streaming or scheduled export into your own SIEM, because your retention requirement is probably longer than their default and because correlating identity events with everything else is the entire point.
"What are the limits I'll hit, and what happens when I do?"
Every platform has them and most don't publish them: users per tenant, groups per user, tenants per account, clients per tenant, token size, custom claims, requests per second per tenant, concurrent sessions. Ask for the numbers. Then ask what the failure looks like at the boundary — a clear error, a silent truncation, or a degradation. Silent truncation of group claims is a real behaviour in real platforms and it produces authorization bugs that take weeks to trace.
"Give me p99 latency by region, for the token endpoint and the authorization endpoint, and tell me how it was measured."
Averages are marketing. p99 is engineering. If they can't produce it, they aren't measuring what their users experience. If they can, ask whether it's measured server-side or end-to-end, because a redirect-based login is seven serial hops and the interesting number is the journey.
Ask exit questions, because they reveal how much they're relying on lock-in
This is the set that vendors are least prepared for and that tells you the most about the relationship you're entering.
"How do I migrate off you?"
Ask it directly, early, in the first technical call. Watch the reaction — that's data. Then get specifics:
- Can I export password hashes, in a documented format, with the algorithm and parameters? Some platforms will not export hashes at all, which means leaving means a forced password reset for every user, which means you can't leave. This one question is the single largest determinant of switching cost.
- Can I export users, groups, and their relationships, in bulk, without paginating through an API at 100 records per second?
- Can I export the configuration — clients, connections, policies — as data?
- What can I not take with me? The honest answer always includes something: refresh tokens and active sessions cannot be migrated by anyone, ever, because they're bound to the issuer. Any vendor who doesn't volunteer that limitation either doesn't understand it or isn't telling you.
A vendor who has documented their own exit path is telling you they intend to keep you by being good. A vendor who is vague about it has made a strategic decision, and you're the subject of it.
"What does the price look like at 10× my current volume, and what triggers a tier change?"
The pricing traps in identity are specific and predictable. Charging per monthly active user sounds fair until you have a seasonal business. Charging separately for "enterprise connections" means every SSO customer you win costs you extra, per customer, forever. Charging for machine identities at the same rate as humans is punitive if you have agents or services — and increasingly, you will. Ask which axis grows fastest for your business, and price that axis at 10×.
Also ask what's in the tier above yours, because the honest version of the question is: which of the things I'll eventually need are behind a paywall I can't see yet? MFA, audit export, SAML, and delegated administration are the four most commonly gated behind an enterprise tier.
Ask standards questions, because they predict whether integration will be painful
"Which conformance certifications do you hold, and for which profiles?"
The OpenID Foundation runs a real certification programme with a public register. It's not a guarantee of quality, but it is an objective, verifiable, non-marketing signal — and it's a much better question than "are you standards-compliant," which has no failing answer.
"Where do you deviate from the spec, and why?"
Every platform deviates somewhere. A vendor who names their deviations — "we don't support the implicit flow," "our amr values are non-standard because the spec's registry is inadequate," "we require PKCE even for confidential clients" — is showing you they understand the spec well enough to know where they left it. A vendor who claims no deviations has either not read the spec closely or is not being straight with you.
"How do you version your API, and what's your deprecation policy in writing?"
Identity integrations live for a decade. Ask for the deprecation notice period, ask for an example of a breaking change they've made and how they handled it, and ask whether old API versions are still running. This question predicts how much unplanned work the vendor will generate for your team over five years, which is a real cost that appears in no comparison table.
Ask reality questions
"Show me your last three incidents and the postmortems."
A status page with no history is not evidence of reliability; it's evidence of a status page. Everyone has incidents. What differentiates is whether the postmortem describes a mechanism or offers a reassurance, whether the timeline includes detection latency, and whether customer notification happened during or after.
"Who answers at 3am, and what's the escalation path for an authentication outage specifically?"
Not the SLA document — the actual path. Is there an engineer on call who can act, or a support tier that files a ticket for a team in another timezone? For a component that gates all access to your product, this is a load-bearing detail.
"Is this feature shipped, or on the roadmap?"
Then treat "on the roadmap" as "no." Not out of cynicism — roadmaps genuinely change for good reasons — but because you cannot build a plan on it, and because a vendor's willingness to say "no, we don't do that" is the strongest available signal of general honesty. A vendor who says no to something is a vendor whose yes means something.
The part that actually decides it: run the experiments
Answers are cheap. Trials are not. If you can get a sandbox, these five exercises will tell you more than a month of calls, and each is a few hours of work.
1. Integrate an application from scratch, timed, without contacting support. Note where you got stuck and whether the documentation or the error messages got you unstuck. Error message quality is the most reliable proxy for engineering culture I know of, and you cannot assess it from a demo — demos don't have errors.
2. Deliberately misconfigure something. Break a SAML assertion. Send a wrong redirect_uri. Use an expired token. Read what comes back. If the response is invalid_request with no detail and nothing appears in a log you can see, then every future integration problem — including the ones your customers report — will be debugged blind. This exercise takes twenty minutes and eliminates candidates.
3. Configure something, then measure how long until it takes effect. Actually time it. Compare to the answer you were given.
4. Export everything, then try to reconstruct it. Pull users, groups, config, and audit events through the documented export path. Can you rebuild the tenant elsewhere from what you got? This is your exit rehearsal, and doing it during the trial — when you have no data to lose — is the only time it's free.
5. Run 10× your expected peak against the sandbox if the terms allow, with a realistic mix including cold-cache reads for tenants you've never touched. Watch p99, not throughput.
Where my own product answers badly
To make the framework usable, here are the questions where the platform I work on gives an answer some buyers should reject:
"Can I run custom code in the login flow?" No, deliberately — no uploaded scripts, no plugins, no per-customer hooks in the runtime. The reasoning is that customer code in an authentication path is a shared-fate arrangement: it breaks benchmarks, it becomes a permanent public API, and its failure modes belong to us while its logic belongs to you. The sanctioned alternatives are configuration and standards-based integration at the edges. But if you have a genuine requirement that only executable logic in the middle of authentication can meet, that's a legitimate requirement and this is the wrong platform for it. Some competitors do it well and you should look at them.
"Do you have a twelve-year track record and a 5,000-customer reference list?" No. Maturity is a real and reasonable procurement criterion, and a young platform cannot manufacture it. If your risk posture requires it, that's a rational position, not a failure of imagination.
"Do you support every legacy protocol variant?" No. Deliberate scope limits mean some long-tail integrations aren't supported, and if your estate depends on one of them, the gap is your problem rather than an abstract one.
I include these because an evaluation framework whose questions all favour the author is worthless, and because the questions above are worth asking of every vendor precisely for their ability to produce an honest no. The most useful information in a vendor evaluation is the list of things each candidate cannot do — and the fastest way to get it is to ask questions specific enough that a general reassurance doesn't fit.
The short version
Ignore the feature matrix. It converges. Ask instead:
- Rotate a key with no downtime — how?
- When does a config change take effect, and how do I know?
- What fails when your control plane is down?
- Show me an audit event's JSON, and how I stream them out.
- What limits will I hit, and what happens there?
- p99 latency by region, and how it's measured.
- How do I migrate off you, and what can't I take?
- What's the price at 10×, and what's gated above my tier?
- Which conformance profiles are certified, and where do you deviate?
- Last three incidents, with postmortems.
- Who answers at 3am?
- What can you not do?
Then go break their sandbox and read the error messages. That's the evaluation.