Why API Keys Refuse to Die
OAuth 2.0 is old enough to drive. The ecosystem is mature, the libraries are good, and every security review in the industry prefers it.
OAuth 2.0 is old enough to drive. The ecosystem is mature, the libraries are good, and every security review in the industry prefers it.
And yet: Stripe, OpenAI, SendGrid, Twilio, Cloudflare, Datadog. Look at what you actually put in your .env file this quarter. A large fraction of the money moving through programmatic APIs is authenticated by a long random string in a header.
The standard explanation is that developers are lazy, vendors are behind, and everyone will get to OAuth eventually. I've believed that. I no longer do. Something that survives fifteen years of unanimous expert disapproval isn't surviving by accident.
The properties nobody talks about
No clock dependency
An OAuth access token carries exp. Validating it requires the resource server and the authorization server to agree, within tolerance, on what time it is.
That sounds trivially satisfied. It isn't. NTP drifts in long-running containers. A VM resumed from a snapshot comes back with a clock from whenever the snapshot was taken. Embedded devices and CI runners are routinely minutes off.
The failure mode is nasty because it's asymmetric and invisible. When a token is rejected as expired, the client cannot tell whether its own clock is fast, the server's is slow, or the token really did expire. The error is identical in all three cases. You end up debugging a distributed systems problem while looking at an authentication error message.
An API key has no exp. There is nothing to compare against a clock. That entire class of failure is structurally absent.
No refresh logic
Tokens expire, so something must renew them. That "something" is where an unreasonable amount of production incident history lives:
- Two threads notice expiry simultaneously and both hit the token endpoint. Now you need a lock, or you're rate-limited by your own IdP.
- A fleet of 400 pods started by the same deploy acquired tokens within the same second — and will all refresh within the same second, an hour later. Synchronized expiry produces a thundering herd on a schedule you didn't choose.
- The token endpoint goes down. Existing tokens work until they don't, and then the fleet fails over a rolling window as each cached token ages out.
None of this is unsolvable. Jitter, singleflight, proactive refresh at 80% of lifetime, stale-if-error caching — the patterns are known. But they are code you have to write and get right, and they fail in ways that only appear at scale or under partial failure. An API key has no lifecycle. There is no refresh path to get wrong.
Trivially debuggable
This one is underrated to the point of being ignored in design discussions, and it may be the strongest argument of the lot.
A failing API-key request reproduces as one line:
curl -H "Authorization: Bearer sk_live_9f3a..." https://api.example.com/v1/orders
Anyone can run it. It is the exact request that failed, byte for byte. Hand it to the partner's engineer, paste it in a ticket, bisect against it.
A failing OAuth request reproduces as: obtain a token with the correct grant, scopes, audience, and client authentication method; confirm the token you got is the same shape as the one that failed; then replay. If the failure was intermittent, or depended on which of two tokens a pod happened to be holding, or on a scope granted at issuance time that you can't inspect, you may not be able to reproduce it at all.
Debuggability is a real engineering property. Credentials that produce reproducible failures cost less to operate than credentials that don't.
Zero dependency at request time
This is the availability argument, and it's the one architects most often get backwards.
flowchart LR
subgraph key["API key"]
C1["Client"] --> R1["Resource Server"]
end
subgraph oauth["OAuth"]
C2["Client"] --> AS["Authorization Server"]
AS --> C2
C2 --> R2["Resource Server"]
end
With client credentials, the authorization server sits in the path to obtaining access. Token caching hides this most of the time — but it converts a hard dependency into a delayed one, which is worse in a specific way: an IdP outage doesn't fail your traffic immediately, it fails it later, gradually, after the incident that caused it has scrolled off the dashboard. And if the resource server validates by introspection rather than local JWT verification, the dependency isn't even delayed. It's per request.
API-key traffic has no such edge. The key is checked against the resource server's own store. Your auth infrastructure can be entirely down and payments still process. That is an availability property, not a security compromise, and it deserves to be evaluated as one.
Comprehensible to the integrator
For a public API, time-to-first-successful-call is a product metric, and it's dominated by how many concepts the integrator must acquire before anything works. An API key requires one: put this string in this header. Client credentials requires grant types, scopes, audiences, token endpoints, client authentication methods, and a token cache — each a place to fail with an error message that assumes you already know OAuth.
Vendors who chose API keys for their public API were frequently not making a security decision. They were making an adoption decision, correctly.
Now the honest bill
None of the above makes API keys safe. The costs are real and they're the ones you already know:
- No expiry. A leaked key is valid forever, until a human notices and revokes it. No automatic ceiling on the damage window.
- Coarse. In most implementations a key means "full account access." No scoping, no least privilege.
- Bearer on every request. The credential crosses the wire constantly, so every proxy log, APM trace, and debug dump is a potential disclosure — the same exposure a client secret has at the token endpoint, just far more often.
- No standard revocation or rotation story. Every vendor invents their own, so every integrator writes bespoke code, so most write none.
- No delegation. A key cannot represent "acting on behalf of user X." There is no clever workaround.
- Possession proves nothing about identity. Correct use and attacker use produce identical requests, so there's no signal to alert on — the same critique that applies to shared secrets generally, and one I've written about elsewhere in the context of rotation.
Since people will use them anyway: build them well
Most of the harm from API keys comes not from the primitive but from implementations that treat it as "a random string in a column." Six things separate a good implementation from a negligent one.
Store them hashed. They're credentials; treat them like passwords — but not like passwords in one specific respect: don't use bcrypt or Argon2. A 256-bit random key has no dictionary to attack, so a slow KDF buys nothing, and it puts a deliberately expensive computation on the hot path of every request, which is a self-inflicted denial-of-service vector. A single SHA-256 is right here. The reasoning that makes bcrypt correct for human-chosen passwords (low entropy) doesn't apply.
Give them a structured, non-secret identifier. Hashing creates an immediate problem: you can't look a key up by its value. The fix is to split it — sk_live_<id>_<secret> — where id is indexed and public and only secret is hashed. That one choice buys four things at once: an O(1) lookup, an identifier you can safely log and put in a support ticket, the ability to revoke a key whose secret half you don't have, and a recognizable prefix that secret-scanning tools can match. Vendors in GitHub's secret-scanning partner program get notified when a key lands in a public repo — and the entry requirement is precisely a distinctive, machine-detectable prefix. A checksum in the suffix cuts the scanner's false-positive rate further.
Support multiple simultaneously-active keys. The single feature that makes rotation possible without downtime, and its absence is why so many keys are never rotated. If a customer must break production to get a new key, they won't.
Scope them. Read-only keys. Per-resource keys. IP-restricted keys. Most integrations need far less than full account access; coarseness is an implementation choice, not a property of the primitive.
Offer optional expiry. The primitive doesn't require it, which is exactly why it should be available: let the security-conscious integrator set 90 days and the embedded device set never.
Record last-used timestamps. Quietly the most valuable of the six. The hardest question in credential lifecycle management isn't "how do I revoke this" — it's "is anything still using this?" A last_used_at column answers it directly, and turns "we're afraid to revoke it" into "nothing has touched this in seven months." Write it coarsely — update only if the stored value is over an hour old — so you're not adding a database write per request.
Where the line actually falls
| Use API keys when | Use OAuth when |
|---|---|
| Server-to-server, no human in the loop | A user is delegating access to their data |
| A single trust boundary — you and the caller | Multi-party: user, client, and resource server are all distinct |
| Public developer API where adoption matters | Enterprise integration where a security review is happening anyway |
| No delegation required | You need "acting on behalf of user X" |
| Availability matters more than fine-grained authz | Short-lived access and revocation windows matter |
| The caller's environment is constrained (embedded, CI, cron) | The caller can hold a key and run a token cache |
The tell for OAuth is delegation. If the answer to "whose data is this?" is anyone other than the party holding the credential, you need a protocol that can express that relationship, and an API key cannot. Everything else is negotiable; that isn't.
The tell for API keys is a single trust boundary plus a hard availability requirement. If your API is called by machines you have a direct contractual relationship with, and it must keep working when your identity infrastructure doesn't, the simpler primitive is not a compromise.
What the persistence is telling us
The interesting thing about API keys isn't that they've survived. It's what they've survived by being: a credential with no protocol, no state machine, no clock, no lifecycle, and no third party. Every one of those absences is a security weakness. Every one is also an operational strength, and the industry has spent fifteen years counting only the first column.
The persistence of API keys is not a failure of OAuth adoption. It's evidence that a simpler primitive is sometimes correctly matched to a simpler problem — and that "more secure" and "right tool for the job" are two different judgments we've gotten into the habit of treating as one.
Make the choice deliberately. Then, whichever you chose, implement it like you meant it.