Identity Isn't Your Secrets Manager
The fix took four minutes. Getting permission to apply it took fifty-one.
A payments company, a Saturday, a bad config push to the identity platform's runtime that made token issuance fail for about a third of requests — enough that the API was broken but not enough that anything failed over. The on-call engineer diagnosed it in eleven minutes, which is good, and the runbook was correct, which is rarer: roll back the projection by re-running the control-plane publish, or, if the control plane is unreachable, connect to the runtime database directly with the break-glass account and set the flag by hand. The control plane was unreachable, because the control plane's admin console was behind SSO, and SSO was the thing that was broken.
So: break-glass account. The password lived in Vault, where it should. Vault was configured with the JWT auth method, federated against the company's own OIDC issuer — which is exactly what every zero-trust migration document tells you to do, and which the platform team had been justifiably proud of, because it had eliminated the last set of long-lived Vault tokens in the estate eighteen months earlier.
Vault was up the entire time. Vault was healthy. Vault would not issue a token to a human being, because the human being could not obtain an ID token from the identity platform that was down. The secret needed to fix identity was behind a lock that opened with identity.
They got in eventually — a director had a Vault recovery key shard in a safe at the office, a second shard holder was reachable, and the unseal quorum was a path nobody had walked in two years. Fifty-one minutes. The postmortem's action items were all sensible and all beside the point, because the real finding was one line and nobody wrote it down: the two systems had been merged in the dependency graph even though nobody had merged them in the architecture diagram.
This is the same class of mistake as putting secrets in the identity platform in the first place. Both of them come from the belief that these two systems are the same kind of thing.
The request that keeps arriving
Every identity platform team gets this ticket, usually from a good engineer with a real problem:
Our integration service needs the customer's Stripe restricted key to reconcile their payouts. We already store the customer's identity with you, we already have a per-tenant encryption boundary, and you already hold client secrets. Can we just store the Stripe key as a custom credential on the tenant and read it back at runtime?
Every clause of that is true. The tenant boundary exists. The encryption exists. Client secrets are, in a loose sense, secrets. The vocabulary is shared: both systems "store credentials", both systems "rotate", both talk about "principals" and "least privilege". And workload identity has become the fashionable front door to secrets, which puts an identity platform one config block away from looking like the natural owner.
The answer is still no, and "we're not a secrets manager" is a definition, not a reason. Here is the reason.
The line: one-way custody versus returnable custody
An identity platform issues short-lived, verifiable, self-describing assertions about a principal. A secrets manager stores and returns arbitrary opaque material verbatim — including material it did not create, cannot interpret, and cannot replace.
That difference cascades into almost everything else, and the cascade starts with one property: what the system must be able to do with the bytes it holds.
A well-built identity platform is, to a surprising degree, a system that cannot reveal its own contents. User passwords are Argon2id digests — irreversible by construction (hashing vs encryption vs encoding makes the distinction properly). Client secrets should be stored the same way, because the platform only ever needs to compare a presented value, never reproduce one; the six client authentication methods are ranked partly by how completely they let you avoid holding a reproducible value at all. Refresh tokens are stored as hashes with a family identifier. Passkey registrations hold public keys. TOTP seeds are the awkward exception people forget, and MFA recovery codes should be hashed like passwords.
Count what is left. In a mature identity platform there is essentially one class of material that must survive in usable form: the token signing keys. One class. And the industry's response to that single exception is instructive — we invented an entire tier of hardware (HSM, KMS, PKCS#11) whose defining feature is that the key still never comes back as plaintext; you send it a digest and get a signature. Even our one unavoidable secret is designed so nothing can retrieve it.
A secrets manager's core requirement is the exact inverse. One hundred percent of what it holds must come back as plaintext, to a caller, on demand, forever, because a Postgres password is worthless unless something can present it to Postgres.
| Identity platform | Secrets manager | |
|---|---|---|
| Material it holds | credentials it issued and can re-issue | arbitrary material, mostly issued elsewhere |
| Storage form | hash where possible; keys in HSM/KMS | reversible encryption, always |
| Must return plaintext? | never (signing excepted, and even that is delegated) | always, that is the product |
| Can it replace what it lost? | yes — reset, re-enroll, re-register | no, not without the third party's cooperation |
| Semantic knowledge of contents | full — it defined the format | none — bytes with a label |
| Verification model | verifier checks a signature offline | none; possession is the whole proof |
| Blast radius of full DB theft | offline cracking against KDF cost | immediate, complete, silent |
| Normal read rate | thousands per second | dozens per day |
Row four is the one that decides most arguments. If your identity database is stolen, you can rebuild: rotate signing keys, invalidate token families, force re-enrolment, reset passwords. Painful, bounded, entirely within your authority. If a store of downstream secrets is stolen, you cannot fix anything yourself. Every Stripe key, every partner API key, every customer's SFTP password has to be rotated by a party who does not work for you, on their schedule, through a support process, with an SLA you do not control. The recovery is not an engineering task, it is four thousand emails.
Do the arithmetic once and the conversation usually ends. Four thousand stored third-party credentials, of which perhaps 60% can be rotated through an API and 40% require a human on the other side. Call it 15 minutes each for the automatable ones and 90 minutes of elapsed coordination for the rest: 600 hours plus 2,400 hours, against a breach clock where regulators expect notification in 72 hours and the affected third parties will each independently decide whether to suspend your integration in the meantime. There is no amount of encryption-at-rest engineering that changes that number, because the number is a property of whose credentials they are.
What saying yes actually changes
Teams tend to model "add a secrets table" as a schema change with an encryption function. It is not. Four things change the moment your identity platform can return plaintext material it did not issue.
Your key hierarchy acquires a new root and a new lifetime. Signing keys have a beautiful property: they are rotatable on a schedule with an overlap window, because the thing they protect (a token) lives for minutes. A data-encryption key protecting stored secrets protects material with a lifetime of years, which means the key encryption key above it must be able to re-wrap historical ciphertext, which means you need envelope encryption with a re-wrap job, key versioning in the ciphertext header, and a decision about what happens to a tenant's data when their KEK is destroyed. That is a real subsystem, and it has nothing in common with the JWKS machinery you already built.
Your compliance scope expands to the union of your customers' scopes. This is the part that surprises people. An identity platform's audit scope is bounded by what identity does. The moment you hold a customer's payment-processor key, an assessor can reasonably argue that the systems storing and transmitting it are connected to their cardholder data environment; hold a healthcare integration credential and you have inherited a business-associate conversation. You do not get to declare the blast radius — the contents declare it, and you can't see the contents, because they're opaque bytes with a label. You have accepted an unbounded compliance obligation defined by strings your customers upload.
Breach notification changes category. A stolen table of Argon2id hashes is a serious incident that is nonetheless argued, plausibly, as material that is not immediately usable. A stolen table of returnable plaintext credentials is a disclosure of the third party's secret, and there is no cracking-cost argument to make. It also changes who you must notify: not just your customers but, transitively, every service provider whose credentials you were holding — many of whom have contractual terms about exactly this and none of whom signed an agreement with you.
The insider model changes. An identity operator with database access can impersonate — bad, detectable, and constrained by what the platform's own tokens are accepted for. An operator with access to a returnable secret store can silently be someone else's customer in a third-party system your logs never see. There is no token, no jti, no issuance event. The action happens entirely outside your telemetry.
The retrieval-audit asymmetry
This one is worth its own section because it looks like a retention question and is actually a detection question.
Token issuance is a high-frequency, low-weight event. A mid-sized platform issues tens of millions a month; each individual issuance means almost nothing, and the interesting signal lives in aggregates and anomalies — impossible travel, failure-rate spikes, a client suddenly requesting a scope it never used. Identity Isn't Your Audit System works through what identity should testify to and why it should stream rather than store; I'm not re-deriving that. The point here is different: a secret read is the opposite kind of event, and merging the two streams destroys the property that makes the rare one useful.
| Token issuance | Secret retrieval | |
|---|---|---|
| Volume | 40,000,000 / month | ~900 / month |
| Weight of one event | negligible | potentially an incident |
| Useful alert | on the aggregate | on the individual event |
| Reviewable by a human? | never | yes, individually, and you should |
| Retention that matters | weeks, then stream out | years, and it must be queryable |
| Question it answers | "is the system healthy?" | "who has ever seen this value?" |
The production database password is read maybe four times a year, by an automated job at a predictable hour. That baseline is so tight that a single off-schedule read is an alertable event with a near-zero false-positive rate — the best detection signal in the entire estate, and it exists precisely because the event is rare. Put that read into a pipeline that carries 15,000 events per second and the signal is not degraded, it is gone; nobody builds a per-record human review process on a firehose, and the retention tier tuned for high-volume operational events will not be the one that can answer "who read this value in 2023" four years from now.
There is a second asymmetry underneath. A token issuance is not a disclosure of anything, because the token is short-lived and scoped and observable downstream. A secret retrieval is a permanent, irreversible disclosure: after that call, a copy of a long-lived credential exists somewhere you cannot see, and it exists for as long as the credential does. The correct audit record for it is therefore not "an event happened" but "a copy of this value now exists in the possession of X" — a custody chain, not a log line. Systems built to record custody chains look different from systems built to record traffic.
Rotation is the same word for two different operations
Both systems say "rotate", and the vocabulary collision is why "you already do rotation, just do it for our Stripe keys too" sounds reasonable in a planning meeting.
Rotating a token signing key is a publish problem. Two keys are valid simultaneously; you add the new one to JWKS, wait longer than the longest consumer cache TTL, start signing, wait out the longest token lifetime, retire the old one. Nobody is coordinated, nothing is scheduled with a counterparty, and no user experiences anything at any point — Designing for Key Rotation Before You Need It walks the phases and I won't repeat them. The reason it works is overlap: the protocol was designed so that two answers can be correct at once.
Rotating a database password is a distributed transaction problem. There is exactly one live value. The target system is a participant, not an observer, and it usually cannot hold two passwords at once. Every consumer must cut over inside the window between "the new password is set" and "the old one stops working" — and if that set of consumers is not fully known, and it never is, rotation is a partial outage waiting for the least-recently-deployed cron job to run.
| Signing key rotation | Downstream secret rotation | |
|---|---|---|
| Valid values during rotation | 2 or more, by design | 1 (sometimes 2, if you engineered it) |
| Who must cooperate | nobody | the target system, and every consumer |
| Failure mode | none if the overlap is respected | authentication failures at an unknown time later |
| Reversible? | yes, until you drop the old key | rarely — the old value is gone |
| Bounded by | consumer cache TTL + token lifetime | the slowest consumer's deploy cadence |
| Emergency variant | mass invalidation, ugly but total | same coordination, no time to do it |
The mature workaround in secrets-manager land is to manufacture the overlap that identity gets for free: the alternating-users strategy, where you keep two database roles and rotate them out of phase so there is always a valid credential while the other is being changed. It works, and note what it is — a scheme that doubles the number of principals in the target system in order to simulate a property JWKS has natively. That gap is not a maturity difference between the two products. It is a structural difference between a credential whose verifier you control and a credential whose verifier belongs to somebody else.
The dependency direction, which is the part people get wrong
Now back to the Saturday.
The two systems have a natural relationship, and it is genuinely useful: identity authenticates the human or workload, the secrets manager authorizes the retrieval. That is a good design and I'll defend it in the next section. The failure is not the relationship — it's failing to notice that it creates a cycle at exactly the moment you need both.
flowchart LR
subgraph normal["Normal operation — clean layering"]
W["Workload"] -->|"OIDC token"| V1["Secrets manager"]
V1 -->|"verify via JWKS"| I1["Identity platform"]
V1 -->|"short-lived DB creds"| W
end
subgraph incident["Identity is degraded — the cycle"]
E["On-call engineer"] -->|"needs break-glass password"| V2["Secrets manager"]
V2 -->|"OIDC login required"| I2["Identity platform<br/>(DOWN)"]
I2 -.->|"fix requires the password"| V2
end
The rule that resolves it is a layering rule, and it is stricter than "avoid circular dependencies", because in normal operation there is no cycle — the cycle only exists in the recovery path, which is exactly the path nobody exercises.
A system may depend on the layer above it for its convenience path, but never for its recovery path. If identity mediates access to secrets, then the secrets needed to repair identity must be reachable without identity.
In practice that means every secrets manager needs a second authentication method that shares no dependency with your IdP: a local auth mount with a small number of accounts, or a quorum unseal held on hardware, or a sealed envelope in a physical safe. Cloud-native estates get a variant of the same problem — if the SSO path into the cloud console is down, the root account with hardware MFA is the local auth mount, and the same discipline applies to it.
Three properties make such a path real rather than theoretical, and the third is the one that fails:
- It is exercised on a schedule. An unexercised break-glass path has silently broken. Quarterly, with the actual on-call rotation, not the person who built it.
- Its use is loudly alerted to somebody other than the user. The whole security argument for having it depends on nobody being able to use it quietly.
- Its dependencies are traced to depth two. This is where the Saturday postmortem should have started. The password was in Vault; Vault needed OIDC; that's depth one and they'd thought about it. Depth two: the shard holders had to be told which shard, over Slack, which was behind SSO; the runbook lived in a wiki behind SSO; the list of who held shards was in a spreadsheet in the same wiki. Every Identity Bug Is a Trust Bug makes this point about identity outages generally — the secrets-manager version is worse, because the recovery material is deliberately hard to reach and the difficulty is a feature you cannot dilute.
There is a second, sneakier direction of the same cycle, and it bites at cold start rather than during an incident. If the identity platform stores its own signing keys, database credentials, or per-tenant configuration in the secrets manager — a perfectly sensible thing to want — and the secrets manager authenticates callers via OIDC against that identity platform, then a full-estate cold start has no valid ordering. Identity can't boot without secrets; secrets won't answer without identity. Nobody discovers this during normal operations because something is always already running. You discover it in a region rebuild, or in the DR exercise you were doing to prove you could do a region rebuild.
The fix is boring and must be deliberate: there is a bottom of the stack, and whatever is at the bottom authenticates with material that is not issued by anything higher — an instance identity document, a TPM-attested node identity, a cloud provider's own IAM role, an HSM that requires physical presence. Draw your boot order as a DAG once, with the arrows meaning "cannot start without", and make sure a topological sort exists. It takes an afternoon and it is the highest-value diagram most platform teams don't have.
The composition that actually works: delete the secret
Now the interesting half, because the boundary is not "these systems must not touch". The best available pattern needs both, and it is genuinely better than what most estates run today.
The move is this: stop storing a secret at all, and let the identity assertion be the credential.
The mechanics, concretely. A workload — a Kubernetes pod, a CI job, an agent — is given a short-lived, audience-scoped OIDC token by its own platform. Kubernetes projects a service-account token with a specified audience and a TTL of an hour or less; GitHub Actions mints one per job. The workload presents that token to the target system's trust boundary: AWS STS AssumeRoleWithWebIdentity, or Vault's JWT auth method, or an Azure federated identity credential. That system verifies the signature against the issuer's JWKS, evaluates a trust policy over the claims — issuer, audience, and crucially sub — and returns credentials that live for fifteen minutes.
The long-lived secret ceases to exist. Not "is stored more securely": ceases to exist. There is no value to rotate, no value to leak, no value to appear in a heap dump, no value to be committed to a repository. This is the same structural move as passkeys removing the shared secret rather than hardening it, and it is the single highest-leverage change available to most platform teams — Non-Human Identity Is Now the Majority explains why the population it applies to is bigger than the human one.
It is not free, and the costs are specific:
- The trust policy is now the security boundary, and it is a string match. The overwhelmingly common defect is a condition on
subthat is too loose: matchingrepo:acme/*instead ofrepo:acme/payments:ref:refs/heads/main, or omitting theaudcondition entirely so that a token minted for a different relying party is accepted. This class of misconfiguration has produced real cross-tenant compromises. Worse, the policy lives in the other system, owned by the cloud or platform team, reviewed by nobody who thinks of themselves as working on identity. - JWKS availability moves onto the credential-acquisition path. The verifier fetches and caches your public keys. That is a much softer dependency than a live introspection call — cached keys keep working through an issuer outage — but it makes your key rotation their problem, and a rotation that respects the overlap window is now load-bearing for other people's deploys.
- Token lifetime versus job duration. A projected token with a one-hour TTL and a three-hour batch job produces a failure two hours in, on the retry, in the part of the pipeline nobody watches. Re-fetch the assertion rather than the derived credential, and make sure your SDK actually does.
- Debuggability gets worse before it gets better. "The key is wrong" is a five-minute diagnosis. "The
subclaim doesn't match the trust policy condition because the workload moved namespaces" is an hour with three teams, and it's exactly the multi-layer problem Debugging Across Four Layers is about. - It requires the target to speak OIDC. Which brings us to the limits.
What federation cannot delete, and why the secrets manager stays
Federation works when you control, or can negotiate with, the verifier. Most of the interesting material fails that test.
| Material | Who holds it | Why |
|---|---|---|
| Token signing keys | Identity platform, backed by KMS/HSM | Identity is the issuer; the private key never leaves the boundary, and only signatures come out |
| User passwords, MFA recovery codes | Identity platform, hashed | Only comparison is needed, never retrieval |
| Client secrets you issued | Identity platform, hashed — or eliminated via private_key_jwt |
You are the verifier, so you never need the plaintext back (here's why ops teams end up hating them) |
| Cloud API credentials for your own workloads | Nobody — federate | The verifier supports OIDC; the secret should not exist |
| Database passwords for legacy engines | Secrets manager | The engine accepts only a password; short-lived dynamic credentials help, but a value still exists |
| Third-party SaaS API keys (Stripe, a partner, a 2011 ERP) | Secrets manager, owned by the team that owns the integration | The verifier is someone else and does not accept your tokens; nothing you build changes that |
| TLS private keys for your endpoints | Secrets manager or a certificate manager, ideally short-lived and automated | Must be returned in usable form to a serving process |
| Customer-supplied encryption keys / BYOK | KMS with a customer-controlled root | The point is that you cannot use it unilaterally |
| Break-glass credentials | Secrets manager with a non-federated auth path, or offline | Must be reachable when identity is not |
| SSH host and user keys, code-signing keys | HSM or a dedicated signing service | Same shape as token signing: sign, never retrieve |
Row six is the one from the original ticket, and note that it is unfixable by anyone in the conversation. Stripe will not verify your OIDC token. That is a fact about Stripe, and the honest answer to the requesting team is not "no", it is: "that credential needs a secrets manager owned by your service, and here is the token-exchange path for the cases where the downstream does speak OAuth" (RFC 8693 token exchange handles the delegation half, where the sub stays the original user and actors accumulate in act). How to Say No to a Feature Request has the general shape of that conversation; this is its most common instance.
And note the symmetric obligation. "We're not a secrets manager" is not a licence to be bad at the credentials you did issue. Rotating your own client secrets, supporting overlapping validity during rotation, recording last_used_at so a customer can tell whether a credential is dead, and running the whole lifecycle properly are all squarely identity's job — see The Lifecycle of an OAuth Client. Refusing that work under a boundary banner is using a principle to avoid effort.
The trust root you should decline to hold
There is one more reason to keep these apart, and it is the one I'd put in front of a security committee.
A secrets manager's entire security reduces to a key hierarchy: data encryption keys wrap the secrets, a key encryption key wraps those, and something at the top — a KMS master key, an HSM, a Shamir quorum — protects the KEK. Whoever controls the top of that hierarchy can, by definition, read everything.
An identity platform controls something different but comparable: the ability to become any principal. Signing keys plus a user store means you can mint an assertion for anyone.
Put both in one system and you have built a component where a single compromise yields both the authority to act as any identity and the plaintext of every stored credential in the estate — with no second system's logs to contradict the story, because the system that would have recorded the anomaly is the system that was compromised. That is not a quantitative increase in risk; it collapses two independent failure domains into one. The reason two-person control and separation of duties exist as concepts is precisely to prevent this shape, and no amount of internal partitioning inside one product restores it, because the operators, the deploy pipeline, and the hypervisor are shared.
The auditor's version of the question is short and clarifying: can the vendor read the plaintext? For an identity platform the answer should be a clean and demonstrable no for essentially everything — the hashes prove it, the HSM proves it. Add returnable secret storage and the answer becomes "no, subject to these controls", which is a different sentence with a different amount of trust in it, and it will be the sentence your next enterprise security questionnaire asks about.
Disclosure: I work on ClavionX, and this is a boundary it draws on purpose. Per-tenant issuers with per-tenant signing keys, and a machine-identity model where an agent is a registered first-class object realized as an OAuth client and issued short-lived credentials — not a process that is handed a copy of a stored secret. (That model is in design rather than shipped; treat it as an intended shape.) The relevant part here is the negative space: there is no API that returns customer-supplied plaintext, because adding one would change the answer to the auditor's question for every tenant, permanently, in exchange for solving one team's integration problem.
The tells of a merged design
If you want to know whether the boundary has already eroded, look for these. Each one is a design that has quietly become a secret store.
- A table or field named
custom_credentials,integration_secrets,external_tokens, ormetadatawhere somebody is putting things that need to come back out. - An API endpoint on the identity platform that returns a value that was never returned at creation time — reversible storage always shows up as a
GETthat should not exist. - Encryption-at-rest documentation that has grown a re-wrap procedure. Nothing in a hash-based store ever needs re-wrapping; if you need one, you are custodian of long-lived reversible material.
- The audit pipeline treats a credential read and a token issuance as the same event class, with the same retention and the same alert routing.
- The break-glass runbook's first step requires an interactive login.
- Your DR plan's boot order contains a cycle, and the mitigation is "start Vault first, then… hmm".
- Support has a documented process for helping a customer recover a value they stored with you. That process is proof that you can read it.
- Compliance keeps asking you questions about a regulatory regime that has nothing to do with authentication. That's the contents of the opaque bytes declaring your scope for you.
The rule, in one line
Identity answers who is this, right now, provably — and it should be built so that the answer can be verified without anything being handed back. A secrets manager answers what is the value — and it exists because some systems will only accept a value, and no protocol you adopt can change their minds.
The right integration between them is not storage. It's the one that reduces the second system's inventory: every credential you convert to a federated, short-lived assertion is a secret that no longer needs guarding, rotating, auditing, or notifying anyone about. That is worth doing aggressively, and it will still leave you with a secrets manager — smaller, containing the material you genuinely cannot delete.
Just make sure the door to it doesn't lock behind the system it's meant to rescue.