The Cost of Session Storage
At 09:14 on a Tuesday, a cloud provider retired an instance. It was a scheduled maintenance event on a replica node in a three-node session cluster, the kind of thing that happens a few times a year and normally shows up as a five-minute blip on a dashboard nobody is watching.
The replacement replica came up empty and asked the primary for a full resync. The primary forked.
Resident memory on the primary went from 6.1 GB to 11.8 GB in about ninety seconds, against a maxmemory of 12 GB. The eviction policy was volatile-lru, which had been chosen years earlier by someone who reasonably assumed that a store full of keys with TTLs would evict the ones closest to expiring. It does not do that. It evicts the least recently used keys that happen to have a TTL, which in a session store means: the sessions belonging to people who had stepped away from their desks for twenty minutes. Several thousand of them were deleted while perfectly valid.
Those users came back, found themselves logged out, and logged in again. Logging in creates a session, which is a write. Being logged in generates requests, and every request extended the session's expiry, which is also a write. The write rate went up, which meant the forked child's copy-on-write footprint grew faster, which pushed resident memory higher, which evicted more sessions, which produced more logins.
The incident lasted forty minutes and the postmortem said "session store under-provisioned." The remediation was to double maxmemory.
Here is the part that took three more weeks to find. The actual session data — every field of every live session — was 2.4 GB. They were provisioning twelve gigabytes of replicated memory to hold two point four gigabytes of sessions, and it still wasn't enough. The gap wasn't sessions. It was the writes, and the writes came from four lines of middleware that updated last_seen on every single request.
That is the thing I want to convince you of, and it is the whole article in one sentence: session storage is not a storage cost. It is a write cost wearing a storage cost's clothing, and almost every design decision that matters is about how often you write, not how much you keep.
This is part of a series doing arithmetic on identity infrastructure. The Hidden Cost of Audit Logs established the method and the counterintuitive finding for that component — storage is cheap, searchability is expensive. The Cost of Token Introspection vs Self-Validating Tokens priced the stateless-versus-stateful debate and found the infrastructure lines land within a few hundred dollars of each other. This one prices the store itself. I'm going to assume you already accept that JWTs relocate state rather than eliminating it and that a session is not a cache; both arguments are made elsewhere and re-running them here would waste your time.
The assumptions, published so you can re-run it
| Parameter | Value | Note |
|---|---|---|
| Registered users | 2,000,000 | Workforce B2B suite, six applications |
| Daily active users | 900,000 | |
| Peak concurrent session records | ~1,000,000 | Composition below |
| Session-validating requests | 4,500/s peak, 1,800/s average | Across all six apps |
| New sessions created | 90/s peak, 25/s average | |
| Median session record | 1.1 KB serialized | p95 3 KB, p99 8 KB — the tail matters |
| Managed in-memory node | $11.50/GB-month | Roughly cache.r7g on-demand, memory-only view |
| Cross-AZ transfer | $0.01/GB per direction | AWS-shaped |
| Cross-region transfer | $0.02/GB | |
| vCPU | $25/vCPU-month | |
| Log ingestion + indexing | $0.50/GB | Negotiated mid-range |
| Target utilization | 40% CPU, 55% memory | Memory target is lower for reasons we'll derive |
| Fully-loaded engineer | $16,700/month |
Same price anchors as the rest of the series, so the numbers compose.
Session count is a residence-time problem, and mobile breaks the model
The first mistake is reasoning about session storage from user count. Sessions are a population, and populations obey Little's Law: the number of records you hold is the arrival rate multiplied by the mean residence time.
The Deploy That Logs Everyone Out runs that relation in the direction most people need — given a session population, what login rate does it imply, and what happens when the population resets to zero. I'm running it the other way: given your traffic, what population do you actually store, and what does the population cost?
The answer is not one number, because a real deployment has several session species with residence times that differ by three orders of magnitude:
| Class | Records | Mean residence | Behaviour |
|---|---|---|---|
| Web browser sessions | ~500,000 at peak | ~6 h (8 h idle, 24 h absolute) | Created fresh, expire, recreated next day |
| Mobile app sessions | ~480,000 | effectively permanent | 30-day sliding refresh; the same record is renewed |
| Machine / integration | ~20,000 | minutes to hours | High churn, small records |
Two things fall out of that table that people routinely get wrong.
Little's Law only applies to the species that re-creates. Web sessions genuinely follow P = λW: cut residence time and the population falls proportionally. Mobile sessions don't, because the app doesn't create a new record every morning — it refreshes the one it has. For that class the population converges to the number of active devices, and shortening the lifetime changes nothing except how often you rewrite the same row. Teams model their whole store with P = λW and are then confused when halving the timeout barely moves memory. Half their population was never governed by that equation.
Your trough is not near zero. Web sessions drain overnight; mobile records don't. This deployment's peak is about 1.0M records and its 03:00 trough is about 500k, a ratio of 2:1. Compare that to the hashing pool, which genuinely goes to near-idle at night. Session storage is the one identity component you cannot turn off overnight, which means scheduled scale-down — the standard FinOps move for anything cyclical — buys you almost nothing here. You provision for peak and you pay for peak for 24 hours a day.
That's also the honest version of "session count grows faster than user count." It isn't that mobile users hold more sessions, though they do. It's that mobile converts a population governed by residence time into a population governed by device count, which only ever goes up, and which has no timeout policy that meaningfully reduces it.
What a session record actually weighs
Now the bytes, because the byte count sets everything downstream and almost every team quotes a number from memory that hasn't been true for two years.
| Field group | Bytes (binary) |
|---|---|
| Session id, user id, tenant id | 64 |
| Timestamps: created, last_seen, expires, auth_time | 32 |
Auth context: amr, acr, MFA state, IdP, sid |
96 |
| Device: IP, user-agent string, device id, fingerprint hash | 320 |
| CSRF token, anti-fixation nonce | 64 |
| Core subtotal | ~576 |
| JSON serialization overhead (keys, quoting, base64) | ×1.5 → 864 |
| Roles (8, as strings) | +200 |
| Median record | ~1.1 KB |
That's the median, and the median is not what you pay for. Add the extras that real deployments accumulate:
| Addition | Bytes | Who has it |
|---|---|---|
| Group claims from enterprise SSO (120 nested AD groups, as DNs) | +6,000 | Your largest customers |
| Per-application last-active, six apps | +150 | Anyone with a suite |
| Per-tab state (SPA tab ids, draft locks, active workspace) | +900 | Anyone whose users keep tabs open |
| Cached entitlement snapshot | +2,000 | Anyone who put authorization in the session |
The distribution is heavily right-skewed: median 1.1 KB, p99 above 8 KB. And because you provision for the total, your memory bill is set by the tail, not the median. In one deployment I looked at, a single enterprise tenant was 14% of users and 61% of session memory, entirely because of nested group DNs. That's the same structural hazard the four-stores article names for SCIM query cost — a number your largest customer controls, not you — and it shows up here too.
Then there's the overhead the store adds, which is where a genuinely useful piece of trivia lives.
Storing a session as a hash with a dozen small fields is cheap: Redis encodes small hashes as a listpack, a flat contiguous blob, and the memory cost is close to the sum of the bytes. But that encoding has thresholds — by default, more than 128 fields or any single field value longer than 64 bytes converts the entire hash to a real hashtable, with a dict entry and an object header per field. The user-agent string is 200 bytes. The serialized group list is 6 KB.
So the moment you add one long field, the whole record's overhead multiplies — typically 3–5× on the small fields that were previously nearly free. It doesn't degrade gracefully and it doesn't announce itself. It shows up as memory-per-session jumping when you onboard a customer whose group names are long.
Blended across the population, this deployment lands at roughly 1.5 KB resident per session, plus a secondary index — a user_id → {session ids} set, which you need the moment anyone asks for "log out all my devices" or an admin forces a logout — at about 200 bytes amortized per session. Call it 1.7 KB resident per session record, against 1.1 KB of actual data.
One million sessions: 1.7 GB. At $11.50/GB-month that is nineteen dollars and eighty cents.
If the story ended there, this article wouldn't exist. Session data is trivially cheap. Nobody has ever had a budget problem caused by 1.7 GB. The problem is that you cannot buy 1.7 GB.
The write amplification that produces the real bill
Here is the line of middleware that every session library ships with enabled, and that every framework tutorial shows you:
session.last_seen = now()
session.expires_at = now() + idle_timeout
store.save(session) # on every request
That is sliding expiration, and it is the correct behaviour — an idle timeout that doesn't track activity isn't an idle timeout. But look at what it does to the workload:
| Rate | |
|---|---|
| Session creates (the actual state changes) | 25/s average |
| Session reads (the actual work) | 1,800/s average |
| Session writes with naive sliding expiration | 1,800/s average |
| Write amplification | 72× |
Identity is mostly read traffic — except that it isn't, once you enable sliding expiration, because you converted every read into a read-modify-write. A store you sized for reads is now taking 1,800 writes per second, and every one of those writes has to be replicated, persisted, and — this is the part nobody models — paid for in provisioned memory.
Let's price each consequence.
Consequence one: replication bytes
Every write goes to the replica, usually in another AZ. What crosses the wire depends on a detail most people never consider: whether you store the session as one serialized blob or as a hash with individual fields.
| Storage shape | Wire cost per touch | Monthly cross-AZ (1,800/s, ×2 directions at $0.01/GB) |
|---|---|---|
Serialized blob (SET sess:x <1.8 KB>) |
~1.8 KB | 8.4 TB → $168 |
Hash field update (HSET sess:x last_seen …) |
~80 B | 373 GB → $7.50 |
A 22× difference in replication cost between two implementations that look identical in the application code. Most session libraries default to the blob, because it's simpler and because serializing the whole object is how object mappers work. Add a second replica and the blob version is $336/month; add a third AZ and it keeps going.
The same volume lands on your persistence path. AOF with appendfsync everysec writes those 8.4 TB to disk too, and then rewrites the file periodically — which means another fork, which brings us to the expensive consequence.
Consequence two: copy-on-write, which is where the money actually is
When a fork-based store snapshots (RDB save, AOF rewrite, or a replica full resync), it forks. The child shares the parent's pages. Every page the parent writes to during the snapshot gets copied.
The naive model says: if I touch 20% of my keys during the fork, I copy 20% of my memory. That model is wrong, and the direction it's wrong in is the entire reason the incident at the top of this article happened.
Sessions are smaller than pages. At 1.7 KB resident, roughly 2.4 session records share each 4 KB page. Touching one record dirties the page holding its neighbours too. So:
Fork duration during a full resync of a 2.4 GB dataset: ~60 s
Writes during the fork (peak): 4,500/s
Distinct keys touched, N_active = 400k, with repetition:
400,000 × (1 − e^(−270,000/400,000)) ≈ 196,000 keys = 19.6% of the keyspace
Probability a given page contains ≥1 touched key:
1 − (1 − 0.196)^2.4 = 41%
Touching 19.6% of your keys dirties 41% of your memory. That is a 2.1× amplification and it comes from nothing but the ratio of object size to page size.
Now enable transparent huge pages, which many Linux distributions do by default and which Redis logs a startup warning about that nearly everyone ignores. A 2 MB page holds about 1,230 session records:
1 − (1 − 0.196)^1230 ≈ 1.000
With THP enabled, a session store under sliding expiration copies essentially its entire dataset during a fork. Memory doubles. That warning in the log is not stylistic; it is the difference between +41% and +100%, and session workloads — many small objects, high random write rate — are close to the worst case the warning exists for.
So the provisioned-memory chain looks like this:
| Step | Multiplier | GB (per node) |
|---|---|---|
| Session payload, 1M records | — | 1.1 |
| Store overhead + secondary index | ×1.55 | 1.7 |
| Allocator fragmentation under churn (jemalloc ratio 1.2–1.5) | ×1.3 | 2.2 |
| Fork copy-on-write peak, 4 KB pages | ×1.41 | 3.1 |
| Login-surge headroom (+30% records) | ×1.2 | 3.7 |
Provision so steady-state RSS ≈ 55% of maxmemory |
÷0.6 | 6.2 |
| Primary + replica | ×2 | 12.4 GB |
~$143/month per million concurrent sessions, to hold 1.1 GB of data. You are paying an 11× multiple on your payload, and the two largest multipliers in that chain — fragmentation and copy-on-write — are both direct functions of your write rate.
That is the finding I'd most like people to take from this piece:
The touch policy in your session middleware sets the size of your session cluster. Not the number of sessions, and not the size of a session record.
Consequence three: the spiral
Those consequences interact, and the interaction is what turns a maintenance event into an incident.
flowchart LR
T["Sliding expiration:<br/>write on every request"] --> W["High write rate"]
W --> C["Fork copies more pages"]
C --> M["RSS approaches maxmemory"]
M --> E["Eviction of live sessions"]
E --> L["Users bounced to login"]
L --> N["New sessions + more requests"]
N --> W
Every arrow there is cheap and obvious in isolation. The loop is not, and it has no natural damping — the more sessions you evict, the more writes you generate, the longer the fork takes, the more pages it copies.
Two configuration choices decide whether this loop can start at all, and both are usually inherited rather than chosen:
maxmemory-policy. If it's any volatile-* or allkeys-* policy, your session store will silently delete valid sessions under pressure, and it will preferentially delete the ones belonging to users who are briefly idle — which is to say, the ones most likely to come back and immediately log in again. The correct setting for a session store is noeviction. That converts a silent correctness failure into a loud write failure, which is the trade you want: a login that fails with an error is recoverable and alertable; a session that vanishes is neither. This is the operational teeth on the "a session store is not a cache" rule.
Persistence. If you can genuinely tolerate losing every session — and you often can, because that failure is self-healing within one login cycle — then turning off RDB and AOF removes the fork entirely, and with it about a third of your provisioned memory. That's a real saving with a real cost: a primary failover with no persistence and no replica means a full logout event, and you have to have priced the thundering herd before you accept it. My default is a replica for availability and no disk persistence, because a disk snapshot of state that expires in eight hours is protecting you against nothing you couldn't recover from anyway.
Expiry isn't free either
The other half of the write workload is deletion, and TTL-based expiry has mechanics that surprise people who assume "it expires automatically" means "it costs nothing."
Redis expires keys two ways: lazily, when something touches the key, and actively, via a background cycle that samples 20 keys per database per iteration at hz iterations per second, repeating while more than 25% of the sample is found expired. That's a sampling algorithm, and its throughput is bounded — deliberately, so expiry can't monopolize the event loop.
Which produces a failure mode: an expired key still occupies memory until it is collected, and it still counts toward maxmemory.
Now think about arrival patterns. Everyone logs in during the 09:00–09:30 window. With an 8-hour absolute lifetime and no jitter, everyone's session expires during the 17:00–17:30 window. Half a million keys become collectible in thirty minutes, and the active expire cycle drains them at a rate set by sampling, burning CPU — up to 25% of the main thread by design — while memory stays high for the entire drain. If a snapshot happens to overlap that window, you are forking while memory is inflated by keys that are logically already gone.
Worse, the pattern echoes. A synchronized login event produces a synchronized expiry event eight hours later, which produces a synchronized re-login, which produces another synchronized expiry. Mass session invalidation of any kind — a deploy, a key rotation, an incident — installs a resonance in your session store that persists for days.
The fix is one line and I have never seen it in a session library's defaults: jitter your session TTLs. Everyone jitters cache TTLs to avoid stampedes; almost nobody applies the same reasoning to sessions, even though sessions have far more synchronized arrivals than caches do. A ±10% uniform jitter on an 8-hour lifetime spreads a thirty-minute expiry cliff across ninety-six minutes, and no user will ever notice that their session was 7h42m instead of 8h.
The same reasoning applies to explicit cleanup jobs if you're on a store without native TTL. DELETE FROM sessions WHERE expires_at < now() against a relational store is the shape that generates the dead-tuple and vacuum problems described in the four-stores article, which I won't re-derive — but note the economic version of it: a batch delete is a burst of write amplification, and if it runs at 17:15 it runs precisely when your write rate is already elevated.
Pricing the levers
So what actually reduces the bill? Here is every lever I know of, priced against the scenario above.
| Lever | Mechanism | Saving | What it costs you |
|---|---|---|---|
| Lazy touch: write only if >10% of TTL elapsed | Writes drop from 1,800/s to ~140/s | ~$160 replication + ~$45 memory (COW falls from 41% to ~2%) | Effective idle timeout becomes TTL to TTL + 48 min. Must be documented. |
| Field-level update instead of blob rewrite | 1.8 KB → 80 B per touch | ~$160/month | Slightly higher per-record overhead; loses atomic whole-object replace |
| Split hot from cold state | Hot: sid → {user, tenant, exp} at 64 B. Cold: groups, device, claims, in a second key fetched only on step-up or authorization refresh |
Resident memory falls ~60%, ~$85/month | Two lookups on the paths that need cold data; more code |
| Store an index, not the payload | Session store holds identifiers; claims are re-derived from the projected user store on demand | Memory approaches zero; the enterprise-groups tail disappears entirely | A read against another store on every request that needs claims — you moved the cost, and you should measure where it landed |
noeviction + drop disk persistence |
No fork, no COW headroom | ~1/3 of provisioned memory, ~$45/month | Failover becomes a logout event |
| TTL jitter | Spreads expiry load | No direct dollars; removes a class of incident | None. Do it. |
| Shorten session lifetime | Fewer concurrent records | ~$49/month | ~$619/month in login-path cost. See below. |
Two of those deserve to be argued rather than tabulated.
Lazy touch is the highest-leverage change in this entire article, and it is four lines of code. Instead of writing on every request, write only when more than some fraction of the idle timeout has elapsed since the last write:
if now() - session.last_written > 0.10 * idle_timeout:
session.last_seen = now()
store.expire(session, idle_timeout)
With an 8-hour idle timeout the threshold is 48 minutes, so an active session is written roughly once per 48 minutes instead of once per request. Write rate falls by more than an order of magnitude; replication bytes fall with it; and the copy-on-write footprint during a fork falls from 41% of the dataset to about 2%, which comes straight off your provisioned memory.
The honest cost is a semantic one, and you must be able to state it: the effective idle timeout is no longer 8 hours, it is somewhere between 8 hours and 8 hours 48 minutes, because a session might not have been touched since just before the threshold. If your compliance obligation says "15 minutes of inactivity," a 10% threshold means 16.5 minutes, and you will have that conversation with an auditor. Shrink the threshold to 2% and you still get a 30× write reduction with a 30-second imprecision. The knob is continuous; pick a point on it deliberately rather than defaulting to "write always," which is what "we didn't think about it" looks like in production.
"Just use JWTs to avoid the storage" is a relocation, not a saving, and I'm not going to re-argue it because the introspection economics piece already priced the destination: per-hop validation CPU, cross-AZ header bytes at the fan-out multiplier, denylist infrastructure that costs about what a session tier costs, and refresh amplification that bills you through your log vendor. What's worth adding here is the specific comparison. The session store above costs ~$143/month per million sessions. In that article, shortening a JWT's lifetime from 15 to 5 minutes cost ~$4,860/month in refresh amplification for 500k sessions. Session storage is roughly two orders of magnitude cheaper per unit of revocation control than the JWT equivalent, which is a strange thing to discover about the design everyone chose to avoid paying for storage.
The lever that costs more than it saves
Now the inversion, which is the argument I'd most like to leave you with.
The intuition is that shortening session lifetime saves money, because fewer concurrent sessions means less memory. Let's price it. Take the web population from 8h idle / 24h absolute down to 30 min idle / 4h absolute — a change a compliance team might reasonably ask for.
What you save. Mean web residence drops from ~6 h to ~1.5 h. Peak web sessions fall from 500k to ~150k. Mobile is unaffected, because mobile refreshes rather than re-creates.
350,000 fewer records × $143 per million = $50/month
What you spend. At constant user-hours, quartering residence time quadruples session establishment: roughly 3.24M additional authentication events per day, or 37.5/s average.
| Line item | Arithmetic | Monthly |
|---|---|---|
| Full credential verification (25% of events, bcrypt ≈ 250 ms) | 9.4/s × 250 ms ÷ 0.4 = 5.9 vCPU | $147 |
| Silent re-auth at the IdP (75%, ~3 ms) | 28/s × 3 ms ÷ 0.4 = 0.2 vCPU | $5 |
| Audit events | 3.24M/day × 12 events × 800 B = 933 GB/month × $0.50 | $467 |
| Infrastructure subtotal | $619 | |
| Support contacts, at 1 per 50,000 authentications | ~2,000/month × $4 | $8,000 |
You save $50 and spend $619 in infrastructure alone — a 12× loss — before the first support ticket. Include support at even a pessimistically low contact rate and the ratio is over 150×.
As an exchange rate, in the same shape as the revocation-window table in the introspection piece:
| Change | Records saved | Memory saved | Login-path cost | Net |
|---|---|---|---|---|
| 24 h → 8 h absolute | ~170k | $24 | ~$190 | −$166/month |
| 24 h → 4 h absolute | ~350k | $50 | ~$619 | −$569/month |
| 8 h → 30 min idle | (folded into above) | — | dominated by re-auth | negative |
| Lazy touch, 10% threshold | 0 | $205 | $0 | +$205/month |
Every row that shortens a session loses money. The row that changes how you write makes money and changes nothing a user can perceive. That is the whole geometry of session-storage economics: the population is cheap and the writes are expensive, so optimize the writes and leave the population alone.
Which raises the obvious objection: session lifetime isn't a cost decision, it's a security decision. Correct — and Session Timeout Policy Across a Product Suite makes the argument that absolute session lifetime is a compensating control for the revocation you don't have. The economics sharpen that argument into something you can take into the room.
If the reason for a 4-hour absolute lifetime is "a terminated employee must lose access within 4 hours," then you are buying a revocation guarantee, and shortening the session is the most expensive available way to buy it — $619/month, for a guarantee that is still measured in hours. The cheap way is an epoch counter: store a sessions_valid_from timestamp per user, bump it on any credential or entitlement change, and compare it against the session's auth_time on read.
2,000,000 users × 16 bytes = 32 MB, replicated to every validator
Thirty-two megabytes buys you sub-second revocation. Six hundred dollars a month buys you four hours. The epoch check is a local comparison against data small enough to hold in every process, it doesn't change session lifetime, it doesn't generate re-authentications, and it doesn't produce support tickets. If your compliance requirement is genuinely about revocation latency — and it usually is, once you ask — this is the control that answers it, and the timeout conversation becomes a much smaller one about unattended devices.
This is structurally the same finding as the token piece, arrived at from the opposite direction. There, the expensive thing was buying a shorter revocation window through token lifetime. Here, it's buying one through session lifetime. Both are the same error: using an expiry as a revocation mechanism, when expiry is the most expensive revocation mechanism there is.
What it costs at the p99 of a bad day
Steady-state numbers are the ones you can afford to be wrong about. Provisioning is decided by the bad day, so it's worth being explicit about what the bad day looks like, because it isn't simply "more of the good day."
The costly property is that peak session count and peak write rate do not coincide. Peak population is mid-afternoon, when everyone who logged in is still logged in. Peak write rate is 09:15, when the login surge and the full complement of returning sessions overlap. You provision for the union, not for either.
Then the compounding events, each of which multiplies a different term:
| Event | What it multiplies | Effect |
|---|---|---|
| Replica loss → full resync | Fork duration and COW | +41% memory, or +100% with THP |
| Rolling deploy of the app fleet | Local caches lost; every request goes to the store | Read rate ×3, write rate unchanged |
| Mass invalidation (key rotation, cookie-secret change) | Session creates | ×20 creates for 30 minutes, then an echo every lifetime interval |
| Client retry storm on a partial failure | Writes, with no corresponding user activity | Write rate ×2–5, and retries re-touch the same keys, extending the fork |
| AZ failover | Cross-AZ bytes | Traffic that was zonal becomes billable |
| Synchronized expiry cliff | Expire-cycle CPU, uncollected memory | 25% of the main thread, memory held past logical expiry |
Three of those six raise memory during the same window in which fork is copying pages. That correlation is why the 55% memory target in my assumption table is lower than the 40% CPU target most people use elsewhere: with CPU, exceeding your target degrades latency; with memory, exceeding it triggers eviction or an OOM kill, and both of those are step functions rather than curves.
Provision session memory for the fork, not for the sessions. If you take one operational rule from this article, take that one.
The rule of thumb, and why it changes with scale
Pulling it together, per million concurrent sessions per month, at a 1.1 KB median record:
| Configuration | Memory | Replication | Total per million/month |
|---|---|---|---|
Disciplined: lazy touch, field updates, no disk persistence, noeviction |
~$95 | ~$8 | ~$105 |
| Typical: blob rewrite, sliding expiration on every request, RDB snapshots | ~$143 | ~$168 | ~$310 |
| Fat records (enterprise groups in the session, entitlement snapshots) | ~$550 | ~$400 | ~$950 |
| Sessions in a replicated relational store, 3× replication, cross-AZ sync commit | ~$400 + IOPS | ~$300 | $1,500–3,000 |
| Any of the above, replicated cross-region active-active | — | +$1,000 | don't |
Call the defensible headline $100–300 per million concurrent sessions per month for a well-built store, and $1,000+ for a careless one — a 10× spread driven entirely by choices in application middleware, not by infrastructure selection.
But the dominant term changes with scale, and this is where the advice inverts:
Small (under ~100,000 sessions): the HA floor dominates, and nothing else matters. A three-node cluster across three AZs costs $300–900/month whether it holds 10,000 sessions or 200,000. Your cost per session is entirely fixed cost, and every optimization in this article saves you nothing. This is the same shape as the introspection tier's fixed floor from the token piece — small identity components are priced by availability requirements, not by workload. At this scale, put sessions wherever operationally simplest and spend your attention elsewhere.
Medium (100k–5M sessions): write amplification dominates. This is the band where everything above applies, where the difference between a careful and a careless implementation is 10×, and where the single highest-leverage change is the touch policy. It is also the band where nearly every B2B platform lives.
Large (over ~10M sessions): topology dominates and cost per session goes back up. You've sharded, so hot-key and slot-rebalancing costs appear. You're multi-region, so the question of whether sessions replicate across regions becomes the biggest line item on the page — and the answer should be no. Synchronously replicating a session write across regions adds the inter-region round trip to every request, and asynchronously replicating it gives you a session that exists in two places with different expiry states, which is a correctness problem rather than a cost problem. Multi-region identity covers the trade; economically the guidance is simple: keep sessions regional, accept re-authentication on regional failover, and price that as a herd event rather than as a replication bill.
That regional-locality argument generalizes into an architectural one. In a design with strict runtime/control-plane separation — ClavionX's ADR-0002, where the runtime never synchronously calls the control plane and operates only on projected state — the session store is unambiguously runtime-local. It holds ephemeral state derived from an authentication event, in one region, with its own durability class. That framing makes the cost question answerable, because you know exactly what the store is for and exactly what losing it costs. Most expensive session stores are expensive because that classification never happened, and the store gradually accumulated data that belonged in the operational database, the cache, or nowhere.
Cost per session is U-shaped in scale. It starts high because of the fixed HA floor, falls as that amortizes, and rises again as topology and geography reassert themselves. If someone tells you their cost per session and it doesn't include which part of that curve they're on, the number is not transferable.
What to measure
Five numbers, none of which need new tooling, and most teams know none of them:
- Your write-to-create ratio. Session writes per second divided by session creates per second. If it's above 10, sliding expiration is your dominant cost and lazy touch is a four-line fix. If it's above 50, it's also your dominant incident risk.
- Your p99 session size, not your mean. Sort by size and look at the top hundred. If enterprise group claims are in there, you have found both your memory tail and a reason to fetch groups rather than store them.
- Your peak RSS during a fork, versus steady-state RSS. Trigger a manual
BGSAVEat peak in staging under representative write load and watch. If you don't know this number, you don't know your required instance size — you know your dataset size, which is a different and much smaller number. - Whether transparent huge pages are enabled. One command. The difference between +41% and +100% fork memory.
- Your
maxmemory-policy. If it isn'tnoeviction, your session store is authorized to log users out silently under pressure, and it will exercise that authorization at the worst possible moment.
The closing thought
The team in the opening story doubled their memory and the incident didn't recur, which is why nobody looked further for three weeks. What eventually got measured was the write-to-create ratio: 84 to 1. They shipped a lazy-touch threshold at 5% of the idle timeout, switched from blob rewrites to field updates, and moved group claims out of the session and into a lookup against projected state.
Resident memory went from 6.1 GB to 3.4 GB. The fork peak went from +5.7 GB to +0.4 GB. They went back down to the original instance size and had more headroom than before the incident. Nothing a user could perceive changed, except that the effective idle timeout became eight hours and twenty-four minutes instead of eight hours.
The general form is worth stating plainly, because it applies well beyond sessions: when a store's cost doesn't match its data size, stop looking at the data and start looking at the write pattern. Session storage looks like a storage problem because it has "storage" in the name, and the arithmetic says it is nothing of the sort. You are not paying to remember who is logged in. You are paying to write down, over and over, that they still are.