The Hidden Cost of Audit Logs

Security teams say: log everything.

Security teams say: log everything.

Cloud providers say: thank you for your business.

Nearly every article about identity audit trails makes the same two points — audit logs are important, and you should have them. Both are true and neither is useful, because nobody disagrees. The interesting question isn't whether to keep an audit trail. It's what happens to your infrastructure bill and your authentication latency when you do it naively, which is how most systems do it.

So let's do the arithmetic that usually gets skipped.

The math

Take a mid-sized enterprise identity deployment: 50 logins per second at peak. Not a hyperscaler — a successful B2B SaaS company with a few hundred enterprise customers.

Now count the audit events a single login actually generates in a chatty system. Authentication attempted. Credential verified. MFA challenged. MFA verified. Session created. Token issued. Consent evaluated. Policy applied. Claims mapped. Tenant resolved. Device recognized. Risk scored. Redirect issued. Refresh token issued.

Fifteen or twenty events per login is not unusual. It's what you get when every component logs its own step, which is exactly what happens when audit is added incrementally by different teams.

50 logins/sec 500 logins/sec
Events/sec (@20 per login) 1,000 10,000
Events/day 86 million 864 million
Raw volume/day (@1 KB) ~86 GB ~864 GB
Raw volume/year ~31 TB ~315 TB

Three hundred terabytes a year, from an authentication system. That number surprises people, and it should — but it's not the expensive part.

Storage is cheap. Searchability is not.

Here's the misconception that drives bad architecture: teams reason about audit cost as a storage problem. Storage is genuinely cheap. Object storage runs on the order of a couple of cents per GB per month, and cold tiers are a fraction of that. Even 300 TB in object storage is a manageable line item — real money, but not the one that ruins your quarter.

The cost is in making it searchable.

Log analytics platforms — Splunk, Datadog, Elastic Cloud, most managed logging services — charge primarily on ingestion, typically somewhere between a few cents and a couple of dollars per GB ingested, depending on tier and negotiation. Run 850 GB/day through a platform priced at even $1/GB ingested and you're at roughly $850/day. Call it $300k/year, to index events that in the overwhelming majority of cases nobody will ever query.

That ratio — cheap to store, expensive to index — is the single most important economic fact about audit systems, and it dictates the architecture. You are not paying to keep the data. You are paying for the ability to search it quickly, and you're paying that on all of it, uniformly, whether or not it deserves it.

And it compounds in ways that are easy to miss:

Replication multiplies everything. Multi-AZ for durability. Cross-region for DR. A copy streamed to the customer's SIEM. Another in the data warehouse for analytics. The same event can exist five times, each with its own cost.

Egress is a real line item. Moving 850 GB/day to a customer's SIEM in a different cloud, at typical cross-cloud egress rates around $0.09/GB, is roughly $75/day — about $28k/year purely to move data you've already paid to produce and store.

Index size can exceed data size. If you make user ID, tenant ID, session ID, IP, and device ID all searchable, you're indexing high-cardinality fields. Index overhead frequently runs 1–2× the raw volume. Every field you make searchable is a permanent multiplier.

Compliance retention was written for a different volume. Seven-year retention requirements were largely drafted when systems produced far less data. Seven years of the 500 TPS scenario is over two petabytes, and the framework doesn't say "unless it's expensive."

The mistake almost everyone makes first

Before the cost problem, there's a latency problem, and it's usually in the code from day one:

BEGIN TRANSACTION
  verify_credentials()
  create_session()
  INSERT INTO audit_log (...)
COMMIT

That insert is inside the authentication path. Which means your login latency now includes an audit write, and — worse — your login availability now depends on your audit store being healthy. When the audit table gets large and its indexes get slow, authentication gets slow. When the audit database has a bad afternoon, nobody can log in.

You have coupled the most latency-sensitive, availability-critical path in your infrastructure to a component whose failure should be completely invisible to users.

The fix is architectural, not an optimization:

flowchart LR
    L["Login"] -->|"publish, don't write"| Q["Event stream"]
    Q --> P["Audit pipeline"]
    P --> H["Hot: 30d, indexed"]
    P --> W["Warm: 90d, queryable"]
    P --> C["Cold: archive, retrievable"]
    P --> S["Customer SIEM"]

Authentication publishes an event and returns. Everything downstream happens on its own schedule, with its own failure modes, none of which reach the user. If the audit pipeline is briefly behind, that's a lag metric, not an outage.

The one caveat worth stating: publishing to a queue and writing to a database aren't atomic. If the login succeeds and the publish fails, you've lost an audit record permanently — which for a security-critical trail is a genuine problem. The standard answer is the transactional outbox: write the event to a table in the same transaction as the state change, and let a separate relay publish it. You keep the atomicity without keeping the audit store in the request path.

Not every event deserves the same treatment

This is where most of the savings actually live, and where most systems apply no thought at all.

Audit systems tend to treat all events identically — same storage tier, same retention, same indexing — because that's simplest. But the events are wildly different in value:

Genuinely critical, keep for years, index fully. Permission granted or revoked. Admin role assigned. Password or MFA changed. Account disabled. Recovery completed. Configuration modified. Credential issued. These are what investigations and audits actually need. They're also, notably, rare — often well under 1% of total volume.

Useful, keep for months, index selectively. Successful logins, session creation, federation events. Valuable for pattern analysis and support questions; not usually needed at four-year depth with full-text search.

High volume, low individual value. Token refresh, session validation, claims mapping steps, cache decisions. These dominate volume and are rarely queried individually. They're the events where aggregate counts matter far more than the individual records.

Applying critical-tier treatment to that third category is where audit budgets go to die. If token refreshes are 60% of your event volume and you're indexing them for seven years at the same tier as privilege grants, you're spending most of your audit budget on your least valuable data.

The practical move: classify events by forensic value at emission time, and let that classification drive retention and indexing. It's a small amount of design work that routinely cuts costs by an order of magnitude, because the volume distribution and the value distribution are almost perfectly inverted.

The distinction that fixes the biggest cost driver

Here's the framing I'd most like people to take away, because it resolves the confusion underneath most audit architectures:

Logs are for debugging. Audit is for evidence. They are not the same thing and should not share a system.

Logs Audit
Purpose Diagnose a problem Prove what happened
Audience Engineers Auditors, investigators, courts, support
Sampling Fine — sample aggressively Never
Loss tolerance Acceptable Not acceptable
Mutability Rotate and delete freely Append-only
Retention Days to weeks Months to years
Volume Enormous Should be modest

Most expensive audit systems are expensive because they're actually debug logs wearing an audit costume. Someone applied audit-grade treatment — never sample, never delete, index everything, retain for years — to data that was only ever meant for debugging.

Separate them and both get better. Debug logs can be sampled, aggressively rotated, and kept for two weeks, which is all anyone actually needs them for. Audit events become a much smaller, much more valuable stream that you can afford to treat properly — full retention, immutability, real integrity guarantees.

The test for whether something is an audit event: would this matter in an investigation six months from now, or in front of an auditor? "The cache missed" doesn't pass. "Alice granted Bob admin access" does.

Audit isn't primarily about compliance

Compliance is why audit gets funded. It's not why it's valuable.

Audit is your time machine, and you'll use it far more often for these than for any auditor:

Support: "The customer says they never received the invitation." Audit: sent at 14:32, delivered, opened at 14:47, never actioned.

Security: "How did this account get admin?" Audit: granted by a specific admin, four months ago, from a specific session.

Customer: "Why was our integration suddenly rejected?" Audit: a config change, timestamped, attributed.

Incident: "What was this tenant's MFA policy during the window?" Audit: the before-and-after state of that change.

That last one is the tell for whether your audit system is actually good. Event names are not enough. An entry saying ConfigurationChanged tells you nothing an investigation can use. You need before and after state, the actor, the source, and the timestamp. Teams that build audit for auditors record event names. Teams that have run an incident record state transitions.

Immutability, without over-engineering it

Audit that can be edited isn't evidence, so append-only is non-negotiable — no updates, no deletes, no exceptions.

Beyond that, people reach for hash chaining: each record includes a hash of the previous one, so tampering breaks the chain detectably. It works, and it's genuinely appropriate for high-assurance environments.

For most organizations it's over-engineering. Object storage with a WORM/object-lock policy gives you tamper resistance enforced by the storage layer, at essentially no engineering cost and no ongoing complexity. It's harder to get wrong than a hash chain you maintain yourself, and it satisfies most auditors without a cryptography conversation.

Pick the cheap option unless you have a specific reason not to. A hash chain that's incorrectly implemented — or whose verification tooling nobody has run in two years — provides less assurance than an object-lock policy that simply works.

What to actually do

Measure your events-per-login ratio. It's the most important number in your audit bill and almost nobody knows theirs. Twenty events per login versus five is a 4× difference in every downstream cost, forever. Look at what you emit and ask which of those an investigator would ever want.

Get audit writes out of the authentication transaction. Publish, don't insert. Use an outbox if you need durability guarantees, which for audit you probably do.

Classify events by forensic value, and tier storage and indexing accordingly. This is where the order-of-magnitude savings are.

Separate debug logs from audit events at the source. Different systems, different guarantees, different retention. Mixing them means paying audit prices for debug data.

Index deliberately. Every searchable high-cardinality field is a permanent multiplier. Decide which queries you actually need to answer fast, and accept slower retrieval for everything else.

Stream to the customer's SIEM rather than storing everything twice. Many enterprise customers would rather have events in their own SIEM anyway. That's cheaper for you and better for them — but budget the egress, because it isn't free.

Model the cost before you commit to retention. Seven-year retention at your projected growth is an arithmetic problem you can solve in a spreadsheet today, and a crisis if you solve it in year three.

The closing thought

Good audit systems don't record everything forever. They record the right things, preserve their integrity, and make them discoverable when it matters.

Security isn't measured by how many events you collect. It's measured by how quickly you can answer the question that matters six months later — and a system drowning in token-refresh records is usually worse at that than a smaller one that was designed with the question in mind.

The uncomfortable truth is that "log everything" is not a security posture. It's the absence of one, and it's the most expensive way to arrive at an audit trail nobody can search.