Performance Testing: Loads, Stress, and Soak
Updated: 7 hours ago
Load, stress, and soak get presented as three flavours of the same activity. They answer three different questions, and only one of them is the question most teams think they're asking.
Tolvern Freight load-tested its rates endpoint before peak season. Target was 300 requests per second; they ran it at 600 to be safe. It passed — p99 at 1.1 seconds, no errors, CPU under 40%.
(Composite scenario. Tolvern Freight is an invented company; the failure pattern is a common one in pre-peak load testing.)
Six weeks later the endpoint fell over at 390 requests per second.
The load test had stubbed the carrier API. In production, each rate lookup borrows a connection from a pool of fifty to a real upstream, and at around 380 requests per second those fifty connections saturate. Requests then queue, the queue grows faster than it drains, and p99 goes from 900 milliseconds to thirty seconds in about ninety seconds of wall clock.
The test at 600 was green because it wasn't testing the same system.
The number is only as real as the least realistic thing in the harness
This is the mechanism worth carrying out of that story, and it generalises past stubs.
A performance limit is set by the scarcest contended resource in the path. Not by CPU, usually. By a pool: database connections, upstream sockets, worker threads, file handles, a rate-limited third party, a lock. Everything else has slack.
Which means a performance test measures the real system only if the scarcest resource is present and equally scarce in the harness. Stub the upstream and you have deleted the constraint — the test now measures your own service's ability to serialise JSON, which was never in question.
The trap is that removing the constraint makes the number look better, so the result reads as reassuring rather than suspicious. A load test that comes back faster than production is not good news; it is evidence the harness is wrong.
The test you can run: for your last performance test, list every dependency you stubbed. For each, ask whether it is pooled, rate-limited, or single-threaded in production. Any yes is a constraint your test does not contain, and your headroom number is fiction above it.
Three tests, three questions, and what each one cannot tell you
The question it answers | The axis | What it cannot tell you | |
Load | Does the system meet its target at expected volume? | Volume you predicted | Where it breaks, or how. A pass is a pass at that number only. |
Stress | Where is the cliff, and what does going over it look like? | Volume you didn't predict | Whether normal traffic is comfortable — you're deliberately past that |
Soak | Does it stay healthy over hours or days? | Time, not volume | Anything about peaks. Soak runs at ordinary load on purpose. |
Read the right-hand column, because that's where the reasoning errors live. The most common one at Tolvern is also the most common everywhere: running a load test and drawing a stress conclusion. "It passed at 2× expected, so we have 2× headroom" assumes the response curve is roughly linear up to the number you tested. It isn't. Pooled resources produce a knee — flat, flat, flat, then vertical — and the knee's position is set by pool size, not by traffic.
Load tests interpolate. Only a stress test locates the knee, and the only way to locate it is to keep raising load until something actually breaks.
The test: ask what your system's failure mode is at 150% of capacity. If the answer is "we don't know, we've never gone there", you have never run a stress test, whatever the pipeline stage is called.
Find the knee, then find what set it
A stress test that ends at "it broke at 380 rps" is half-finished. The number moves the moment someone changes a pool size, so the durable output is the constraint, not the figure.
Ramp until failure, and at the failure point capture three things: which resource hit 100% first, what the queue depth did, and whether the system recovered when load dropped. That last one separates two very different systems — one sheds load and returns; the other has a queue it will never drain and needs a restart.
Tolvern's answer was "carrier connection pool, 50, queue unbounded, does not recover". Every part of that is actionable in a way that "380" is not: raise the pool, or bound the queue and fail fast, or cache quotes. They bounded the queue first, which turned a thirty-second timeout for everyone into a fast, explicit rejection for a few — a much better failure.
The test: for the last capacity number you were given, ask what resource produced it. If nobody can name one, the number is an observation, not a finding, and it will silently expire the next time someone edits a config.
The average is the least useful number you will collect
Reporting a mean response time is close to reporting nothing. A system where every request takes 200 ms and a system where 95% take 50 ms and 5% take 3 seconds have similar means and completely different customers.
Useful percentiles, and what each is for:
p50 — what a typical request feels like. Good for spotting broad regressions, useless for finding pain.
p95 / p99 — where your unhappy users live. This is the number to put in a target.
p99.9 — where retries, timeouts, and cascading failures start. Worth watching in stress runs specifically, because it moves first.
The percentile that matters most is the one your callers time out on. If the mobile client gives up at 5 seconds, then the question is what fraction of requests exceed 5 seconds — which is a threshold count, not a percentile at all, and often clearer to report directly.
One caution about aggregation: percentiles from separate load generators cannot be averaged together. A p99 of a p99 is not a p99. Collect raw response times or use a tool that merges histograms properly, or your headline number is arithmetic fiction.
The test: look at your last performance report. If the headline is a mean, recompute it as p99 and see whether the conclusion survives.
Soak is the only one that isn't about load at all
Load and stress both push volume and both finish in minutes. That shared shape hides a category of failure neither can reach: anything that accumulates.
Memory leaks. Connection handles that aren't returned on the error path. Log files filling a disk. A cache with no eviction. Slow, monotonic degradation that looks like nothing at all in a ten-minute run and takes a service down on day four.
A soak test runs at ordinary load for hours or days, and the pass criterion is not response time — it's flatness. Memory flat, handles flat, latency flat. A soak test where p99 climbs 4% an hour has failed even if every request succeeded, because the trend is the result.
This is also the test most often skipped, for an understandable reason: it doesn't fit in a pipeline. The workable compromise is to run it on a schedule rather than per-commit — weekly, or before a release — and to compare the shape of the curve against the last run rather than against a threshold.
The test: plot memory for your longest-running service over its last seven days. If it's a sawtooth that resets on deploy, you have a leak that your deploy cadence is currently hiding.
What to change before your next performance test
One thing, before any tooling decision: name the scarcest pooled resource in the path you're about to test, and confirm it is present and equally scarce in the harness. If it isn't, fix that before you collect a single number — everything downstream inherits the error.
Then decide which of the three questions you're actually asking, and let that pick the test. "Does it hold at peak" is load. "Where does it break" is stress, and it needs a run that ends in failure. "Does it stay up" is soak, and it needs hours.
Tolvern still runs the same load test at 600 requests per second. It now runs against the real carrier sandbox, and it fails at 380 — which is the point.
The intermittent failures that show up when you start running tests under load are usually a different problem wearing performance clothing; start with diagnosing the flake. And if your harness needs realistic volumes of data to be representative at all, that's a test data decision before it's a performance one.
Part of the Software Quality Engineering guide — ShiftQuality's complete map to building quality in.


