The Lifecycle of an OAuth Client
Every OAuth deployment I've reviewed has a table of clients, and every one of those tables contains at least one row that nobody can explain.
Every OAuth deployment I've reviewed has a table of clients, and every one of those tables contains at least one row that nobody can explain.
internal-reporting-v2, created four years ago, secret last rotated never, six scopes including one that grants read access to every user record. Nobody currently employed knows what it's for. Its last successful token request was eleven months ago, which is either evidence it's dead or evidence it runs annually, and there's no way to tell which. Deleting it is an unbounded risk borne entirely by whoever presses the button. Keeping it costs nothing that anyone will be blamed for.
So it stays. It will still be there in four more years.
This isn't a hygiene failure. It's what happens when a system treats registration as the whole story — when the only lifecycle operation with a first-class implementation is create, and everything after it is a support ticket. A client is not a config row. It's an object with a lifespan, and the interesting engineering is in the transitions nobody built.
The stages, and where each one actually breaks
stateDiagram-v2
[*] --> Requested
Requested --> Active: approved, credential issued
Active --> Active: rotation · scope change · owner transfer
Active --> Suspended: incident, or unused
Suspended --> Active: reinstated
Suspended --> Retired: nothing broke
Retired --> [*]: purged after retention
Six states, and most platforms implement two: it exists, or it's deleted. The gap between those two is where every operational problem in this article lives.
Registration: the metadata is the whole game
The temptation is to make registration cheap. Self-service, one form, redirect URI and a name, done — it removes you as a bottleneck and developers love it.
Cheap registration is correct. Cheap registration without required metadata is what produces the orphan above, and the marginal cost of demanding the metadata at creation time is roughly zero, because the person creating the client is the one moment in the object's life when someone actually knows the answers.
Four fields, none of them optional:
A named human owner. Not a team alias, not a distribution list. Aliases outlive the people in them and route to nobody; the failure mode is a mail loop rather than an escalation. Store the alias too, but the person is the record. And plan for that person to leave, which is the next section.
A purpose, in a sentence. "What breaks if this client stops working?" Free text, written by someone who knows, read three years later by someone doing an audit. This single field is the difference between "delete with confidence" and "leave it, it might be load-bearing."
An expiry or review date. Covered below — it's the highest-leverage field on the form.
An environment. Which is trivial and constantly omitted, and then a production client is sitting in a table next to forty test clients created during a hackathon, and nobody can filter.
Dynamic Client Registration (RFC 7591) deserves a note here because it's frequently deployed as an open endpoint and that's a mistake in most contexts. Unauthenticated DCR means anyone who can reach your authorization server can mint clients — a registration-spam and reconnaissance surface with no upside outside of specific federated ecosystems. Gate it behind an initial access token. The protocol supports this; the default configurations often don't use it.
Ownership transfer: the transition nobody implements
Here's the sequence that produces most orphans, and it involves no negligence at any step.
An engineer registers a client for a service her team owns. Two years later she moves to a different org. Her account is offboarded correctly — the HR event fires, her SSO access is revoked, the checklist is completed by someone competent. The client she owns is not on the checklist, because nothing joins the two systems.
The client keeps working. Its credential is valid. Its owner field now names a person who no longer works here, and the only signal that anything changed is a field in a table nobody reads.
The fix is a join you almost certainly already have the inputs for. You consume HR or directory events for user offboarding; extend the consumer to query for objects owned by the departing user. Then either force reassignment before the offboarding ticket closes, or flag the client for review with a deadline.
The organizational version is worse and more common: a team is reorganized out of existence and its systems are inherited by people who never asked for them. There's no event for that. The only defence is periodic re-attestation, and the only version of re-attestation that works is one where not responding has a consequence — the client is suspended, not merely marked stale. A review process whose failure mode is a spreadsheet entry is a review process that has already failed.
Credential rotation: the operation that must be overlappable
Most rotation implementations are a replace: generate a new secret, store it, tell the owner. Which means at the instant of rotation, every running instance of that application is holding an invalid credential, and you have an outage until they all pick up the new one.
Predictably, teams learn to avoid rotating. The credential ages. When it's finally forced — usually by an incident or an audit finding — it's a coordinated change across services nobody wants to touch.
Rotation has to support two valid credentials simultaneously, with an overlap window:
- Issue credential B. A remains valid.
- Owner deploys B at their own pace, no coordination required.
- Platform observes which credential is being used — this is the part that's usually missing.
- When A has been unused for the full window, revoke it.
Step 3 is what makes the whole thing safe, and it needs per-credential last-used telemetry, not per-client. Without it, revoking A is a guess, and a guess taken by someone who will be blamed if it's wrong. With it, revocation is a query result: A hasn't been used in 14 days, kill it.
This is also the strongest practical argument for private_key_jwt or mTLS over shared secrets. Asymmetric credentials rotate by publishing a second public key — the platform never holds the private key, the client rotates on its own schedule, and the overlap window is just "two keys in the JWKS." The operational difference is large enough that it usually justifies the migration on its own.
Scope changes: the direction matters enormously
Adding a scope is a privilege escalation for the client. It should require the same approval as the original registration, and it almost never does — in most platforms it's an edit on a form.
The asymmetry worth building in: additions require approval, removals don't. Making removal frictionless is how you get any removal at all. If dropping an unused scope needs the same three-day approval cycle as adding one, nobody will ever do it, and scopes accrete monotonically for the life of the client.
Removal still needs a safety net, because "unused" is a claim about the past. Per-scope usage telemetry answers it properly: if a client hasn't exercised users.write in six months, the removal is well-evidenced rather than hopeful. And a soft-removal mode — where the scope is denied but the denial is logged as a warning for a fortnight before it's enforced — turns a risky change into an observable one.
The general shape here is the same as everywhere else in this article: the operation is easy, the confidence is the hard part, and confidence comes from telemetry you have to collect in advance.
Suspension: the state that makes deletion possible
This is the missing state, and adding it changes the economics of the whole lifecycle.
Deletion is terrifying because it's irreversible and its consequences are unknown. So the rational move for any individual engineer is to leave the client alone, and the orphan table grows forever.
Suspension is reversible. The client's tokens are rejected, its credentials stop working, and everything about it is retained. If something breaks, you reinstate it in thirty seconds and you've had a minor incident. If nothing breaks for ninety days, you have evidence — not a belief — that the client is unused, and retirement becomes an administrative act rather than a leap.
That's the whole trick: suspension converts an unbounded risk into a bounded, recoverable experiment. It's the single highest-value state to add to a platform that doesn't have it, and it's the one that unblocks the cleanup that's been deferred for years.
Two details make it work. Suspension needs to be automatic on an unused-for-N-days rule, because manual suspension is subject to the same nobody-wants-to-press-it problem as deletion. And it needs to notify the owner clearly, with a reinstate link — the goal is a fast, low-drama recovery for the one case in twenty where the client was genuinely dormant rather than dead.
Retirement, and the identifier you must never reuse
When retirement finally happens, one rule matters more than the rest: never reuse a client_id.
It seems harmless — the client is gone, the identifier is free, and some naming scheme somewhere wants it back. But old tokens carry that ID in azp or in an audit trail, downstream systems have cached authorization decisions keyed on it, and a partner's config file still names it. Reissue the identifier to a different client and every one of those becomes a quiet misattribution. The audit record from 2024 now appears to implicate a service that didn't exist then.
Tombstone the ID permanently. Storage is free; forensic ambiguity is not.
The retirement sequence itself has an order that's worth following: revoke outstanding tokens (including refresh tokens, which is the step that gets skipped, and which is how a "retired" client keeps working for another 90 days); disable the credentials; mark retired but retain the record; purge the object only after your audit retention window, keeping the tombstone forever.
The audit trail is the point
Every transition above is a security-relevant event and needs a record: who did it, when, what changed, and — the field people omit — why.
Scope grants and credential issuance are the two that will be asked about. When an auditor or an incident responder asks "who gave this client access to customer data, and when," the answer must be a query, not an archaeology project across Slack history and a ticket system that's been migrated twice.
The failure mode I see most is a platform that logs authentication events beautifully — every token issuance, every failure, nicely structured — and treats configuration events as ordinary application logs, unstructured and retained for thirty days. The config events are the ones that matter for this. Token issuance tells you what happened; the scope grant tells you why it was possible.
ClavionX is one of the platforms that models this as an object lifecycle rather than a row — Agents move through explicit REGISTERED → ACTIVE → SUSPENDED → RETIRED states, and their operational settings come from an assigned policy rather than being edited per-object, so a change applies across a whole class and recompiles. The reason that's worth mentioning isn't the state machine itself, which is unremarkable. It's the second part: per-object configuration is what makes a fleet of clients unreviewable. Thirty clients with individually-set grant types and TTLs is thirty things to check. Thirty clients governed by three policies is three.
What to do on Monday
If you have an existing estate and want the highest return for the least work, in order:
Collect last-used timestamps, per client and per credential. Everything else in this article depends on it. Without it, every decision is a guess and every guess is deferred. This is the one to do first even if you do nothing else.
Add a suspended state. It's what makes cleanup psychologically possible.
Backfill owners on everything, and join the offboarding event stream so you stop generating new orphans while you're clearing the old ones.
Set expiry dates on new clients and let them ride. Renewal inverts the burden: instead of someone proving a client is unneeded, someone must periodically claim it's still needed. Unclaimed clients die on their own, which is the only cleanup mechanism that has ever worked at scale.
Then go look at internal-reporting-v2. Suspend it. Wait ninety days. You already know how this ends.