The Art of Load Testing: Why Most Performance Tests Fail Before They Start
Most teams load test a week before launch.
Most teams load test a week before launch.
The best teams start six months before they'll need the capacity.
That gap explains most of what goes wrong in performance engineering, and it isn't a discipline problem — it's a misunderstanding of what the exercise is for. A load test run a week before launch can only answer does this survive today's expectations? By then, the architecture is fixed, the deadline is real, and the only available responses to bad news are "add servers" or "ship anyway."
Capacity planning is a forecasting problem. The number you need isn't your breaking point. It's whether the system will still be standing on the day the business succeeds.
This article is deliberately not about k6, JMeter, or Gatling. The tooling is the easy part and it's well covered. What's poorly covered is how experienced teams think about this.
Your best load test already ran. It was yesterday.
The most common failure is inventing traffic instead of observing it.
Someone writes a script: log in, get a token, call an API, repeat. Ramp to 10,000 TPS. It passes. Everyone's satisfied. Then production melts under a load that's numerically lower than what was tested, and nobody can explain why.
The explanation is almost always that the synthetic test and production traffic have nothing in common except a number.
Your production system is continuously generating the most accurate specification of your workload that will ever exist. Before writing a single test script, extract from it:
- Shape over time. Not average TPS — the actual curve. Where are the peaks, how steep is the ramp, how long do they last?
- Protocol mix. What fraction is password login, SAML, OIDC, token refresh, introspection, SCIM?
- Cardinality. How many distinct tenants, users, clients, and devices appear in an hour?
- Session characteristics. Duration, refresh frequency, concurrent sessions per user.
- The tail. Which tenants are enormous? Which single customer accounts for 30% of your SAML traffic?
Then look for cycles, because identity traffic is unusually rhythmic. Daily peaks at the start of business hours in each timezone you serve. Monday mornings dwarfing Friday afternoons. Month-end and quarter-end spikes in finance-adjacent products. Payroll days. A specific enterprise customer's onboarding wave. Marketing campaigns that your team learns about from the traffic graph.
None of this is guesswork. It's sitting in your metrics right now, and it's a better test plan than anything you'd design from first principles.
Two reasons your load test is lying to you
Even a well-designed test usually reports numbers that are better than reality, for two specific reasons that are worth knowing about.
Coordinated omission
This one is under-appreciated even among experienced engineers, and it systematically hides your worst latencies.
Most load generators work in a loop: send a request, wait for the response, record the time, send the next one. That sounds obviously correct. It isn't.
Suppose your system stalls for two seconds. During that stall, a real user population would have sent hundreds of requests — they don't know your server is struggling, they click when they click. But your load generator is blocked waiting, so it sends nothing. When the response finally arrives, it records one slow request and moves on.
The result: the two-second stall contributes a single bad sample instead of hundreds. Your p99 looks fine. In production, that same stall affects every user who arrived during it, and your p99 is catastrophic.
The fix is to generate load against a schedule rather than a loop — the requests that should have been sent during the stall are counted, with their full waiting time attributed. Most serious load tools support this (constant arrival rate, open workload models); many teams leave it configured the wrong way because the default is the loop.
If your load test has never shown you a p99 that alarmed you, this is the first thing to check.
Your test is running against a warm cache
The second lie is identity-specific, and it's the one I've seen bite hardest.
Load tests typically run with a small, fixed set of test data — 100 users across 5 tenants, because that's what the fixture script created. After thirty seconds, every one of those users, tenants, signing keys, and client configs is in cache. Every subsequent request is a cache hit.
You are measuring a fully cached system. Production has 500 tenants and 200,000 users with a long access tail, which means a meaningful fraction of real requests miss cache and hit the database.
The measured difference is not subtle. As covered in the latency budget breakdown, an uncached tenant config lookup can cost tens of milliseconds against roughly one for a hit. A test with a 100% hit rate and production at 85% are different systems wearing the same version number.
Test data cardinality should approximate production cardinality. If you have 500 tenants, your test needs hundreds of tenants — not five with more traffic each.
Volume isn't what changes. Mix is.
Here's the planning error that catches teams who are doing capacity work: they forecast the number and miss the workload shift.
Identity workloads evolve in a fairly predictable sequence, and each stage stresses something different:
| Era | Dominant traffic | Where it hurts |
|---|---|---|
| Early | Password logins | CPU — deliberately expensive hashing |
| Enterprise arrives | SAML federation | Crypto + XML parsing, per-tenant config lookups |
| Platform maturity | OIDC, token refresh | Signing throughput, key cache, token churn |
| Modern auth | Passkeys | Different crypto profile; hashing pressure drops |
| Agents & automation | Machine identity | Token issuance rate, often far exceeding human volume |
A system sized for password logins is sized for CPU. When that same customer base shifts to passkeys, the hashing load falls away — and the bottleneck moves somewhere you weren't watching. When machine identities arrive, you can see token issuance rates that dwarf your entire human user base, from a handful of clients that never sleep.
So "we'll be at 2× volume next year" is an incomplete forecast. The useful question is 2× of what, and the answer usually comes from sales rather than from a metrics dashboard.
Capacity planning is a business conversation
The naive model is: current TPS → hardware.
The real model runs through the business:
flowchart LR
T["Current traffic"] --> G["Organic growth rate"]
G --> P["Sales pipeline\n(who, how big, when)"]
P --> F["Roadmap features\n(new workload types)"]
F --> M["Safety margin"]
M --> C["Target capacity"]
When sales says "we're closing twenty enterprise customers this year," a capacity planner shouldn't hear more users. They should hear: more SAML connections with per-tenant config, more SCIM sync jobs on someone else's schedule, higher MFA rates because enterprises enforce it, longer audit retention, and more concurrent sessions per user because enterprise users have more devices.
Twenty enterprise customers might represent fewer users than a consumer campaign and considerably more load, distributed completely differently.
This is why capacity planning that lives entirely inside engineering fails. The leading indicators aren't in your metrics — they're in the pipeline, the roadmap, and the onboarding schedule.
Size for six months out, and preferably twelve. Procurement takes time. Architectural changes take longer. If you're sizing for today, then on the day the business succeeds, the platform fails.
Maximum throughput is the least interesting number
Most load tests are designed to find the breaking point. That's the wrong target, because you will never operate there deliberately.
What you actually need to know is how the system behaves in the region around trouble:
How does it degrade? Does latency rise smoothly, giving you time to react — or is it flat until it falls off a cliff?
What happens when a dependency fails? Redis unavailable. Database slow but not down. Kafka backed up. An upstream IdP timing out. These are your real production scenarios, and each has a correct behavior you should have chosen deliberately.
Does it recover on its own? This is the question almost nobody tests, and it's the most important one.
Why identity systems fail gradually, then all at once
Identity has a particularly nasty failure mode worth understanding, because it explains a lot of outages that look inexplicable in hindsight.
Authentication is on everyone's critical path, and clients retry it. When auth gets slow, every client — mobile apps, service-to-service calls, browser tabs — starts retrying. Retries increase load, which increases slowness, which triggers more retries.
The system enters a state where the load is now being generated by the failure itself. And critically: removing the original trigger doesn't fix it. The traffic spike that started it is long over; the retry storm is self-sustaining. Teams sit there watching a system that's failing under load that shouldn't be there, unable to find the cause, because the cause was twenty minutes ago.
Recovery usually requires shedding load — which means you need to have built the ability to shed load before you need it. Retry budgets, exponential backoff with jitter in your client SDKs, circuit breakers, and the operational ability to reject traffic deliberately rather than collapse under it.
Test for this explicitly: drive the system into overload, remove the overload, and see whether it comes back without intervention. If it doesn't, you have a latent outage regardless of what your throughput number says.
The self-inflicted spike nobody sees coming
A specific version of this worth calling out: synchronized token expiry.
Deploy at 09:00. Every client authenticates and receives a token with a one-hour lifetime. At 10:00, all of them refresh simultaneously. You've built a periodic, self-inflicted thundering herd into your own traffic pattern, and it gets sharper every time you do a mass restart or deploy.
The same dynamic applies to cache TTLs expiring in lockstep across instances that started together, and to scheduled SCIM syncs that every tenant runs at midnight because midnight is the default.
The fix is jitter — randomize token lifetimes within a band, stagger scheduled jobs, spread cache TTLs. It's a small change that removes an entire category of periodic spike, and it's invisible in any load test that doesn't model client behavior over hours.
What to instrument
Throughput alone tells you almost nothing. During a test, watch:
Latency distribution — p50, p95, p99, p99.9. The average is actively misleading; it's the tail that generates support tickets.
Saturation, not utilization. Queue depth, thread pool occupancy, connection pool waits. These turn upward before latency does, which makes them your early warning. CPU at 60% with a saturated connection pool is a system in trouble that looks healthy on the dashboard everyone's watching.
Downstream behavior — database connection waits, Redis latency, event bus consumer lag, GC pause frequency and duration.
Error composition. A 2% error rate made of timeouts is a completely different problem from 2% made of rate-limit rejections. The second might be the system working correctly.
What I'd actually do
Start from production data. Extract the real traffic shape, mix, and cardinality before writing any test.
Fix coordinated omission first. If your tool is running a closed loop, your latency numbers are optimistic and you don't know by how much.
Match test cardinality to production cardinality. Hundreds of tenants, not five.
Test the mix you'll have in a year, not the one you have now. Ask sales what's closing.
Test degradation and recovery, not just peak. Kill Redis mid-test. Slow the database. Then verify the system comes back on its own.
Make it continuous. A load test is a snapshot; capacity planning is a practice. Run it against every release and watch the trend — a 15% regression caught in CI is a fix, and the same regression discovered at peak is an incident.
Revisit the forecast quarterly, with someone from the business in the room.
The point
Capacity planning isn't about predicting the future perfectly. It's about making sure success doesn't become your biggest outage.
The best-performing systems aren't the ones that passed yesterday's load test. They're the ones designed around tomorrow's traffic — which is knowable, because it's already visible in today's metrics and this quarter's pipeline, if anyone bothers to look at both at once.