top of page

Build a Smoke Test Suite — Test Automation in Practice, Part 7

Shawn West
Apr 27
5 min read

Updated: Aug 24

Illustrative composite: Tolvern Freight's shipping platform. A smoke suite is the only suite that runs after deployment against the real environment, and that changes what belongs in it.

Before you start

You need:

  • A deployment pipeline you can add a post-deploy step to.

  • A real environment to point at — staging or production — reachable from CI.

  • A dedicated smoke account whose data you don't mind creating, and a way to distinguish its records from real ones.

  • Agreement on what happens when smoke fails. A suite that goes red and changes nothing is a notification, not a gate. Settle this before writing any code.

About 70 minutes.

What you'll build

Five checks that run after every deploy, finish in under ninety seconds, and answer one question: is this deployment, in this environment, actually working?

Step 1: Ask the deployment question, not the correctness question (10 min)

Tolvern's first smoke suite ran the API test pack against staging. It passed for six weeks, including the afternoon checkout was down.

The tests were correct and useless. They asked "does the code work?" — which the merge pipeline had already answered — instead of "did this deployment come up wired to the right things?"

Those are different failures. Deployment failures look like:

  • A missing or stale environment variable

  • Credentials that expired, or point at the wrong tenant

  • A migration that didn't run, or ran halfway

  • A DNS or ingress change that routed traffic somewhere unexpected

  • A dependency that is healthy in staging and unreachable from production

None of them are reachable in CI, because in CI the config is right by construction. That is the entire reason this suite exists.

Check: for each smoke test you're about to write, ask whether it could have failed in the merge pipeline. If it could, it's a regression test wearing a smoke label — leave it where it was.

Step 2: Choose five flows, and no more (10 min)

Five is not arbitrary. A smoke suite that takes ten minutes stops being run on every deploy, and a suite that isn't run on every deploy isn't a smoke suite.

Tolvern's five:

Check

The deployment failure it catches

Health endpoint reports the build SHA

The deploy silently didn't roll out

Sign in with the smoke account

Auth config, session secret, identity provider reachable

Get a rate quote

The carrier credential is present and valid in this environment

Book, then read the shipment back

Database reachable, migrations applied, writes permitted

An authenticated page renders

The frontend bundle deployed and matches the API

Row one is the one teams forget and the cheapest to add. If /health doesn't return the SHA you just deployed, everything below it is testing the old build and telling you it's fine.

Check: each row names a configuration or environment failure. Any row you'd describe as "checks the logic is right" belongs in CI.

Step 3: Write them thin (15 min)

// smoke/smoke.spec.ts
import { test, expect, request } from '@playwright/test';
const BASE = process.env.SMOKE_BASE_URL!;
const SHA  = process.env.DEPLOYED_SHA!;

test('serves the build we just deployed', async ({ request }) => {
  const r = await request.get(`${BASE}/health`);
  expect(r.status()).toBe(200);
  expect((await r.json()).build).toBe(SHA);
});

test('can quote a shipment', async ({ request }) => {
  const r = await request.post(`${BASE}/rates`, {
    headers: { Authorization: `Bearer ${process.env.SMOKE_TOKEN}` },
    data: { from: 'GB', to: 'FR', weight_kg: 10 },
  });
  expect(r.status()).toBe(200);
  expect((await r.json()).amount).toBeGreaterThan(0);
});

Assert shallowly and on purpose. amount > 0 proves the carrier credential works. Whether the arithmetic is right is a unit test's job and belongs nowhere near a post-deploy check.

Check: point the suite at a URL that doesn't exist. Every test must fail fast with a clear message. If one hangs for thirty seconds, set an explicit timeout now.

Step 4: The decision point — how far to go in production (10 min)

Smoke against staging is easy. Production is where the value is and where the judgement is needed.

The signal that decides it: can this check be undone, and is its footprint identifiable?

  • Read-only checks — health, sign-in, quote. Run in production, always. Nothing to clean up.

  • Writes with an identifiable footprint — book a shipment as the smoke account, then cancel it. Run in production if you can find and exclude those records everywhere they'd otherwise appear.

  • Writes you can't undo — charging a card, sending a customer email, calling a carrier's live booking API. Never in production. Use a sandbox credential or stop at the boundary.

Tolvern books in production against the carrier's sandbox account and tags every smoke record is_smoke = true. That flag has to be respected by analytics, billing, and any alerting on order volume — and getting it into only two of the three is the failure mode. Wire the exclusions before you turn the check on.

Check: run the write check twice, then query production for smoke records. You should find exactly what you created and nothing orphaned. Then confirm one downstream report excludes them.

Step 5: Make failure mean something (10 min)

- name: Deploy
  run: ./deploy.sh
- name: Smoke
  run: DEPLOYED_SHA=${{ github.sha }} npx playwright test smoke/ --workers=5
- name: Roll back
  if: failure()
  run: ./rollback.sh

Two decisions are encoded there. Smoke runs after the deploy, because it is testing the deployment. And failure triggers rollback rather than a notification — which is the only version that changes outcomes at 2am.

If automatic rollback is too aggressive for your context, the alternative is a page, not an email. A smoke failure nobody sees is worse than no smoke suite, because it looks like coverage.

Check: deploy a deliberately broken build to staging. Confirm smoke fails and that the rollback ran. Test the rollback path, not just the detection.

Step 6: Keep it fast enough to always run (10 min)

Ninety seconds is the working ceiling. Past that, someone adds --skip-smoke for a hotfix and it never comes back.

  • Run in parallel — these tests are independent by design.

  • Prefer API checks to UI checks; keep at most one UI check.

  • Set an aggressive per-test timeout. A hanging smoke test is a failure.

  • Resist additions. Every new check must displace one, or justify the seconds.

Check: time the suite. Over ninety seconds, cut something — and cut the check that overlaps most with CI, since that one was never a smoke test.

The wrong smoke suite beside the right one


Regression pack pointed at staging

A real smoke suite

Question asked

is the code correct?

is this deployment healthy?

Could fail in CI

yes — so it's redundant

no

Runs

before deploy

after deploy

Checks the deployed SHA

no

yes

Runtime

8–40 min

under 90 s

On failure

someone looks tomorrow

rollback or page

You're done when

  • /health returns the SHA you just deployed, and the check fails if it doesn't.

  • Every check would have passed in CI — meaning none of them is redundant.

  • A deliberately broken deploy triggers a rollback you have watched happen.

  • Production smoke records are identifiable and excluded from at least one downstream report.

  • The suite finishes in under ninety seconds.

Troubleshooting

Smoke passes but the site is down. You're testing the old build. Add the SHA check — Step 2, row one.

Smoke fails intermittently right after deploy. Instances are still warming. Poll /health until ready before starting, rather than adding a sleep.

Production smoke data appears in reports. The exclusion flag isn't wired everywhere. Find every consumer before running writes again.

The suite keeps growing. Every addition needs a deployment-failure justification. Most don't have one.

It got skipped for a hotfix and stayed skipped. It was too slow. Step 6.

Next

Checking that a deployment is wired correctly is a different question from checking that the interface still looks right — that's Add Visual Regression Testing. And the reasoning behind which tests gate a merge is in Regression Testing Strategies.

bottom of page