Add Visual Regression Testing — Test Automation in Practice, Part 8
Updated: Aug 24
Illustrative composite: Tolvern Freight's booking UI. Snapshot Testing argues that the review of the diff is the real test. Visual regression is that idea in pixels, and it fails the same way — just faster.
Before you start
You need:
A component library, Storybook, or stable routes you can render in isolation. Full-page shots of a live app will not survive Step 3.
Somewhere to commit baseline images — and an honest look at your repo's tolerance for binary files.
CI that renders in a container. Baselines taken on a Mac will not match Linux CI: different font rasterisation, different subpixel antialiasing. This is the single most common reason teams abandon visual testing in week two.
Playwright, which has this built in. Percy, Chromatic and Applitools work the same way with hosted review.
About 60 minutes.
What you'll build
Visual coverage of four components with masked volatile regions, a tolerance you chose deliberately, and a review step that a human can complete in seconds.
Step 1: Generate baselines inside the container from the start (10 min)
Do this before writing a single test, because doing it in the wrong order means throwing away every baseline you make.
docker run --rm -v $(pwd):/work -w /work \
mcr.microsoft.com/playwright:v1.48.0-jammy \
npx playwright test --update-snapshots
Every baseline — first and future — comes from the same image CI uses. Never --update-snapshots on your laptop.
Check: generate a baseline in the container, then run the test locally outside it. It should fail on antialiasing differences. Seeing that failure once tells you exactly why the container rule exists.
Step 2: Choose component scope, not page scope (15 min)
This is the decision the whole technique rests on, and it's the same size argument as text snapshots.
A full-page screenshot is roughly 1.4 million pixels. When it fails, the tool shows a human two images and asks "is this correct?" — and after the third unrelated failure, the honest answer becomes "probably, approve it".
// e2e/visual.spec.ts
import { test, expect } from '@playwright/test';
const CASES = ['quote-card', 'status-badge', 'address-block', 'price-summary'];
for (const name of CASES) {
test(`${name} renders as expected`, async ({ page }) => {
await page.goto(`/iframe.html?id=components-${name}--default`);
await expect(page.getByTestId(name)).toHaveScreenshot(`${name}.png`);
});
}
Note getByTestId(name), not page — the shot is of the component, not the viewport. A change to the header now fails the header test alone, with a diff the size of a header.
Check: change one component's padding by two pixels. Exactly one test should fail. If several do, your scope is still too wide.
Step 3: Mask everything that legitimately varies (10 min)
Timestamps, IDs, avatars, animations and anything random will fail every run. Mask them rather than loosening the tolerance for the whole image.
await expect(page.getByTestId('quote-card')).toHaveScreenshot('quote-card.png', {
mask: [page.getByTestId('generated-at'), page.getByTestId('quote-ref')],
animations: 'disabled',
caret: 'hide',
});
The distinction that matters: masking excludes a named region; raising the threshold blinds the entire image. Mask the clock. Don't set a 5% tolerance and lose the ability to detect a broken layout.
Check: run the same test three times with no code changes. It must pass all three. Any failure is an unmasked volatile region — find it before continuing, because it will otherwise be blamed on the tool.
Step 4: The decision point — how much difference is a failure (10 min)
Two plausible settings, and the wrong one in either direction makes the suite worthless.
The signal that decides it: what is the smallest change you would want to be told about?
Too tight (maxDiffPixels: 0) — a font rendering update in a browser patch fails all forty tests. The team approves in bulk. You now have a rubber stamp.
Too loose (threshold: 0.3) — a collapsed layout passes because most pixels are unchanged. You have a suite that only detects total blankness.
Tolvern's setting, and a reasonable default:
// playwright.config.ts
expect: {
toHaveScreenshot: { maxDiffPixels: 100, threshold: 0.2 },
}
threshold is per-pixel colour sensitivity; maxDiffPixels is how many pixels may differ at all. Using both means a small colour shift everywhere passes, while a real displacement of a hundred pixels fails.
Check: shift one element by three pixels and confirm it fails. Then re-run with no change and confirm it passes ten times consecutively. You need both to hold; only one is easy.
Step 5: Make the diff genuinely reviewable (10 min)
- run: npx playwright test visual/
- uses: actions/upload-artifact@v4
if: failure()
with:
name: visual-diffs
path: test-results/**/*-diff.png
Three images per failure — expected, actual, diff — and the diff is the one that matters. If your reviewer has to download a zip and flip between tabs, they will approve without looking. A hosted service earns its cost precisely here.
Check: open a failing pull request as a reviewer would and time yourself. Over thirty seconds to see what changed means the review won't happen properly under deadline.
Step 6: Updating baselines is the discipline (5 min)
Every problem in this tutorial ends at the same command, exactly like jest -u.
Never bulk-update. Update by path: npx playwright test visual/quote-card --update-snapshots.
Updated baselines are a reviewable change. A pull request that alters eleven baseline PNGs needs a sentence saying why all eleven changed.
Say what changed before you approve it. If you can't complete "the padding on the price row went from 8px to 12px", you haven't reviewed it.
Check: search your history for bulk --update-snapshots runs. Each one may have baked a visual bug into the baseline as expected.
The wrong visual suite beside the right one
Full-page, tight threshold | Component-scoped, masked | |
Diff size when it fails | ~1.4M pixels | a few thousand |
Fails on an unrelated change | constantly | rarely |
Volatile content | breaks every run | masked by name |
Baselines generated | on a laptop | in the CI container |
Catches a collapsed layout | yes | yes |
Reviewed properly after a month | no | yes |
You're done when
Every baseline was generated in the same container CI uses.
Three consecutive runs with no code change pass every time.
A three-pixel shift fails, and exactly one test fails.
Every volatile region is masked by name, not absorbed by a raised threshold.
A reviewer can see what changed in under thirty seconds.
Troubleshooting
Everything fails on the first CI run. Baselines came from your machine. Regenerate in the container — Step 1.
One test fails randomly. An unmasked volatile region, or an animation. Add animations: 'disabled' and mask it.
A whole-page test passes despite obvious breakage. Threshold too loose. Step 4.
Forty tests fail after a browser upgrade. Expected — rasterisation changed. Regenerate in the new container image, in one commit, and say so in the message.
Nobody reviews the diffs any more. They're too large or too slow to open. That's Step 2 and Step 5, and it is the failure mode that ends visual testing programmes.
Next
Checking the interface still looks right is a different question from checking a deployment came up healthy — that's Build a Smoke Test Suite. And when you need tests around code you didn't write and can't safely change, start with Add Tests to Legacy Code.


