Snapshot Testing: When It Helps and When It Doesn't
Updated: Aug 20
A snapshot test makes one assertion: this output is identical to the output we recorded. Whether that's a test or a rubber stamp depends entirely on something outside the test file.
Tolvern Freight's shipment detail page had a snapshot test. One line to write, 900 lines in the .snap file, and it covered the whole rendered component tree.
A refactor changed the currency helper. Prices went from £1,240.00 to £1240 — thousands separator and decimals both gone. The snapshot test failed, exactly as designed.
The developer ran jest -u, saw 900 lines of diff in the pull request, scrolled past it, and merged. Nobody caught it in review either. It shipped.
The test worked perfectly. It detected the change, failed the build, and demanded a human decision. The human said yes.
The review is the test; the assertion is only the trigger
This is the thing that makes snapshots different in kind from other tests, and it's why they behave so strangely in practice.
An ordinary assertion encodes a judgement made once, carefully, at the time of writing. expect(format(1240)).toBe("£1,240.00") says something specific and permanent about what correct means. It cannot be satisfied by changing the expectation, because changing the expectation is obviously changing the claim.
A snapshot encodes no judgement at all. It says "same as last time" and defers the entire question of correctness to whoever reads the diff when it eventually fails. The correctness check does not live in the test. It lives in a code review, and only happens if someone actually reads it.
Which yields the rule that governs everything else here: a snapshot's value is a function of whether the diff will be read. Not of coverage, not of what it renders — just that.
The test you can run: find the largest snapshot in your repo. Open its last update commit and ask whether anyone could plausibly have verified all of it. If not, that file is recording changes, not testing them.
Size determines whether the diff gets read
Diff length isn't a style preference. It's the mechanism.
Snapshot size | What happens when it fails | Verdict |
Under ~30 lines | Read line by line; wrong values get spotted | Genuine test |
30–100 lines | Skimmed; structural changes noticed, value changes missed | Weak |
Over ~100 lines | -u and move on | Change detector, not a test |
Whole page / full tree | Fails on every unrelated change, trains the team to update reflexively | Negative value |
That last row is worth being blunt about. A snapshot that fails constantly for legitimate reasons doesn't just fail to catch bugs — it actively teaches the team that snapshot failures are noise. It degrades the response to every snapshot in the repo, including the good small ones.
Tolvern's fix wasn't better review discipline. It was deleting the 900-line snapshot and replacing it with four small ones — the price block, the status badge, the address block, the carrier line — each under twenty lines. The price snapshot now fails alone, with a six-line diff, and that diff is unmissable.
The test: for each snapshot over 100 lines, ask what specific thing it protects. If you can't say it in one sentence, split it until you can.
Snapshots are strongest where writing assertions is genuinely tedious
None of this means snapshots are a bad tool. There are jobs where they're the right one, and they share a shape: the output is large, mostly structural, and you care about accidental change rather than specific values.
Good fits:
Serialisation formats. An API response shape, a generated config, a CSV header row. You want to know if it changed; the values are incidental.
Error message catalogues. Long, tedious to assert individually, and unauthorised changes matter.
Compiler or codegen output. Nobody wants to hand-assert an AST.
Small presentational components with no logic — the output is the behaviour.
Poor fits, and the reason each fails:
Anything computed. Prices, totals, dates, tax. Assert the value explicitly; that's a claim you can make correctly once and keep forever.
Anything with a timestamp, ID, or ordering that varies. You'll be fighting nondeterminism, and the fixes make the snapshot flaky or blind.
Whole pages. Too many reasons to change; guaranteed to end up in the bottom row of the table — and usually duplicating cheaper regression coverage.
Anything you'd struggle to review. Which is the same rule again, from the other side.
The distinction underneath: snapshots test that output is stable; assertions test that output is right. Stability is a real property worth protecting — just not the one you want on a number.
The test: for each snapshot, ask whether a wrong value inside it would look obviously wrong in a diff. Where it wouldn't, the value needs its own explicit assertion, snapshot or no snapshot.
The update flag is where the discipline actually lives
Every problem in this article converges on one command. -u is the point where a test becomes a rubber stamp, and it's worth putting friction there rather than relying on intent.
Three practices that hold up:
Never run a bare jest -u. Update by path — jest -u src/pricing. A blanket update rewrites snapshots you weren't even thinking about, including ones failing for real reasons in code you didn't touch.
Fail CI on obsolete snapshots (--ci does this in Jest). An obsolete snapshot means a test was deleted or renamed; the file left behind is dead weight that will confuse the next person.
Make snapshot changes visible in review. If your PR template has a checklist, "snapshot diffs reviewed line by line" belongs in it. If a diff is too long for that to be honest, that's the signal to split it — not to tick the box.
There's a fourth practice worth more than the other three: when a snapshot fails and you're about to update it, say out loud what changed and why. Tolvern's developer could not have completed that sentence — the currency change was invisible in 900 lines. Requiring the sentence is what surfaces the ones you can't answer.
The test: search your history for bare -u runs. Each one updated every snapshot in the repo, and any of those files may now be recording a bug as expected behaviour.
What to change this week
Sort your snapshot files by line count, largest first. Take the top three.
For each, write the one sentence that says what it protects. Where the sentence comes easily, keep it. Where it doesn't — and for the big page-level ones it won't — split it into the two or three things you actually care about, and add explicit assertions for any computed value inside.
Then delete whatever is left over. A snapshot nobody can describe isn't protecting anything; it's just a diff you'll approve without reading, on some Tuesday, in a hurry.
Tolvern's shipment page has four snapshots now, totalling sixty lines. The last time one failed, it was a two-line diff on the status badge — and the reviewer caught it in about four seconds, which is the entire point of the tool when it's pointed at the right thing.


