Regression Testing Strategies for Fast-Moving Teams
Updated: 6 hours ago
Every fast-moving team eventually splits its suite into "runs at merge" and "runs later". Almost everyone splits it by runtime, because runtime is the number you can see. It's the wrong number.
Tolvern Freight's suite took forty minutes. Merges queued behind it, people stopped waiting, and the team did the obvious thing: took the two hundred fastest tests, called them the merge gate, got it to four minutes, and moved everything else to a nightly run.
Three weeks later a rounding change in the rate calculator shipped. Negotiated-rate customers were quoted to the nearest pound instead of the penny. It was caught the following Tuesday by a customer, five days and about four hundred quotes after the merge.
The test that would have caught it existed. It ran in eleven seconds. It just wasn't in the fastest two hundred.
Sort by what it costs to find out late
The merge gate is not a performance budget. It's a statement about what you are willing to not know at merge time — and that's a risk decision, not a scheduling one.
The useful question for each test is: if this broke and nobody noticed for three days, what would it cost to unwind?
That cost has three multipliers, and none of them is runtime:
Blast radius — one internal screen, or every quote issued?
Silence — does it announce itself with a stack trace, or does it quietly produce wrong-but-plausible output? Silent failures are far more expensive, because time-to-detect is unbounded.
Reversibility — can you redeploy and move on, or is there now bad data to find and correct?
Tolvern's rounding bug scored badly on all three: every negotiated customer, silent, and four hundred already-issued quotes to reissue. That combination belongs in the gate at almost any runtime.
The test you can run: open your merge gate's test list and find the slowest test in it. Ask why it earned its place. If the honest answer is "it happened to be under the cutoff", your gate was selected by a criterion nobody would defend out loud.
What the ranking actually looks like
Test class | Cost if found three days late | In the gate? |
Money and quantities — pricing, tax, rounding, refunds, inventory | Wrong data already sent to customers; manual reissue; possible credit notes | Always, whatever it costs to run |
Permissions and tenancy — who can see what | Potential disclosure; may be reportable | Always |
Data writes and migrations — anything that persists wrongly | Corrupt rows accumulate for three days; a backfill to write | Always |
Core happy paths — can a customer complete the main job | Loud, caught fast, but blocks everyone | Yes |
Integration contracts with other teams | They discover it, not you; cross-team unwind | Yes, if cheap — else run contract tests instead |
Rendering, layout, copy | Visible, embarrassing, trivially reversible | No — post-merge is fine |
Admin and internal tooling | A colleague files a ticket | No |
Slow end-to-end journeys | Usually duplicates cheaper coverage | No — and check they aren't the only coverage |
Two rows deserve a second look. Silent-and-wrong beats loud-and-broken every time — a crash gets a pager, a rounding error gets a customer email a week later. And anything that writes data outranks anything that only displays it, because display bugs end at redeploy while write bugs leave a trail.
The test: for each of the top three rows, confirm you have at least one gate test. Most teams find they cover happy paths thoroughly and money edge cases not at all.
The bug you just fixed is the highest-value test you can write
The cheapest source of good regression tests is your own incident history, and it needs no selection framework at all: a bug that happened once is empirically reachable, which is more than you can say for most tests.
The discipline is narrow and worth being strict about:
Write the test before the fix, and watch it fail. A regression test you never saw red is a guess.
Assert the specific wrong value, not just "no exception". Tolvern's test asserts quote.amount == Decimal("1240.37") — a test asserting amount > 0 would have passed throughout the incident.
Put it in the gate if the bug scored badly on the three multipliers. It already proved it can happen.
The reason this works is selection bias in your favour: production found the gap in your reasoning, so a test written against it covers something your design-time imagination didn't.
The test: take your last five incidents. For each, find the test that now prevents it. Any incident without one will happen again, and the second occurrence is entirely predictable.
Selection is a heuristic, so it needs a safety net
Test impact analysis — running only tests touching changed code — is tempting, and for a large suite it's the difference between four minutes and forty. It is also wrong at exactly the moments that matter most.
Static analysis of what a change "touches" misses reflection, dependency injection, configuration, shared fixtures, feature flags, and anything the database enforces rather than the code. Those are precisely the paths where surprising breakage lives.
So treat impact analysis as an optimisation, never as coverage:
Gate — the risk-selected set above, always, no exceptions and no selection logic.
Impact-selected — additionally run whatever the tool thinks is affected. Free upside.
Post-merge — everything, on main, on every merge if you can afford it. This is your net, and it must run often enough to identify the culprit commit. A nightly full run tells you something broke in the last twenty commits; a per-merge run names it.
The failure mode to avoid is trusting selection so much you stop running the full suite. Then a gap in the tool's model becomes a permanent hole nobody can see.
The test: how long after a merge would you learn that it broke something outside the gate? If the answer is "tomorrow morning", your net has a hole the width of a working day, and every merge after the bad one is now also suspect.
Delete more than feels comfortable
The reason suites reach forty minutes is that nothing is ever removed. Tests accrete; each individually seems worth keeping; nobody is ever thanked for a deletion.
Three categories that can go, and the reasoning for each:
Tests asserting behaviour nobody committed to. If the team would happily change it without telling anyone, it isn't a guarantee — it's a snapshot of a decision, and it will fail on every legitimate change.
Duplicate coverage at a more expensive layer. If an end-to-end test covers only what three API tests already cover, it costs runtime and flake risk for nothing. Delete the expensive one, keep the cheap ones.
Tests nobody trusts. A test that gets rerun on failure has already stopped being a test. Fix it or delete it; leaving it teaches the team to ignore red.
Deleting a test that has never caught anything and fails occasionally is not a loss. It is closing a position that was costing you.
The test: find the three tests that failed most often last quarter. For each, name the real bug it caught. Any with none are costing attention and buying nothing.
What to change this week
Don't rebuild the suite. Re-sort the gate.
Take your current merge gate and, for each test in it, ask the three-day question. Then look at the tests outside it and find the ones that score worst — money, permissions, writes. Move those in, even the slow ones. Move out whatever is only there because it was fast.
Tolvern's gate is now sixty-one tests and runs in five minutes. It contains three tests that take over twenty seconds each, all three about money. It is a minute slower than the version that let the rounding bug through, and that minute is the cheapest insurance the team buys.
Functional regression suites will not catch a slowdown. Performance Testing: Loads, Stress, and Soak covers the tests that do.


