Write Your First E2E Test — Hands-On Software Testing, Part 6
Updated: 5 hours ago
Illustrative composite: Tolvern Freight's booking UI. Test Automation ROI argues that UI end-to-end tests are the most expensive coverage you can own and that you should keep very few. This tutorial builds one of the few.
Before you start
You need:
A journey whose silent failure would cost money. Not "the page loads" — a flow where a broken button means no revenue and no error in your logs.
The ability to run the app locally or against a stable environment, with a known test account.
Permission to add attributes to your own markup. Step 3 is not optional and needs a one-line change per element.
Node and Playwright. Cypress works identically; the reasoning is the same.
About 60 minutes.
What you'll build
One test: a customer gets a rate quote and books a shipment. Written so that a redesign changes the page but not the test.
Step 1: Pick the journey by what its silence costs (5 min)
The selection rule is the same one that governs the merge gate: what would it cost to find out three days late?
An end-to-end test earns its maintenance only where the failure is invisible to everything else you own. Tolvern's case: the API is healthy, returns correct quotes, and a JavaScript error makes the Book button inert. Every API test passes. Every monitor is green. Nobody can book anything.
That's the journey. Not the settings page, not the admin table.
Check: name what your chosen journey catches that no API test could. If the honest answer is "nothing", pick a different journey — or write an API test instead and save yourself the maintenance.
Step 2: Get a first pass running (10 min)
npm init playwright@latest
// e2e/book-shipment.spec.ts
import { test, expect } from '@playwright/test';
test('customer can quote and book a shipment', async ({ page }) => {
await page.goto('/quote');
await page.getByLabel('From').fill('GB');
await page.getByLabel('To').fill('FR');
await page.getByRole('button', { name: 'Get quote' }).click();
await expect(page.getByText(/£/)).toBeVisible();
});
Check: it passes. Then break it on purpose — change 'Get quote' to 'Get quotes' — and confirm it fails with a message naming the button. A test whose failure message you haven't read is a test you can't debug at 9am.
Step 3: Replace fragile selectors before you go further (15 min)
This is where most first E2E tests are quietly ruined, and doing it now costs fifteen minutes instead of three rewrites.
CSS selectors like .btn-primary > span:nth-child(2) couple your test to the DOM, which changes on every redesign for reasons that have nothing to do with booking. That coupling — not the browser, not the tooling — is what makes end-to-end suites expensive.
Two selector families survive:
// Semantic — survives restyling, and asserts accessibility as a side effect
page.getByRole('button', { name: 'Book shipment' })
page.getByLabel('Destination country')
// Explicit test IDs — for anything with no accessible name
page.getByTestId('quote-total')
<span data-testid="quote-total">£1,240.00</span>
Prefer the semantic query. Reach for data-testid only when there is genuinely nothing to name — and treat every one you add as a small admission that the markup could be more accessible.
Check: grep your spec for .class, #id, and nth-child. Every hit is a redesign away from failing. Replace them now.
Step 4: The decision point — how the test gets its data (10 min)
Your test needs a customer who can book. Two plausible routes, and the wrong one produces a test that fails for reasons unrelated to booking.
The signal that decides it: is this setup part of the journey you're testing?
Not part of it → create it through the API, fast and reliable.
test.beforeEach(async ({ request }) => {
await request.post('/api/test/customers', {
data: { email: `e2e+${Date.now()}@tolvern.test`, plan: 'standard' },
});
});
Part of it → drive it through the UI, deliberately. If sign-up is the journey, clicking through it is the point.
Getting this backwards means your booking test fails when someone changes the login form. Never drive setup through the UI for a journey that isn't about setup.
Check: count the UI interactions before the first assertion. More than two or three and you're testing setup, not the journey.
Step 5: Wait for conditions, never for durations (10 min)
Every waitForTimeout you write is a flake with a delayed start date.
// Wrong — passes on your laptop, fails on a loaded CI runner
await page.waitForTimeout(2000);
expect(await page.getByTestId('quote-total').textContent()).toContain('£');
// Right — waits for the thing you actually care about, up to a limit
await expect(page.getByTestId('quote-total')).toContainText('£', { timeout: 10_000 });
Playwright's assertions auto-retry, so most explicit waits are unnecessary. The exception is a state change with no visual signal — wait on the network response or a data attribute, not on a guess about how long it takes.
Check: grep for waitForTimeout and sleep. Each one is a future flaky test and its failure will arrive on a busy day.
Step 6: Assert the side effect, not just the screen (5 min)
A confirmation message is the UI's claim that something happened. Take one more step:
await page.getByRole('button', { name: 'Book shipment' }).click();
await expect(page.getByText('Booking confirmed')).toBeVisible();
const res = await request.get(`/api/shipments?idempotency_key=${key}`);
expect((await res.json()).items).toHaveLength(1);
Same reasoning as the four surfaces in API testing: the screen says it worked, the read-back proves it.
Check: stub the booking endpoint to return success without writing. The UI assertion should still pass and the read-back should fail. If both pass, the second assertion isn't reaching the real system.
Step 7: Wire it in without wiring it into everything (5 min)
- run: npx playwright test --reporter=html
- uses: actions/upload-artifact@v4
if: failure()
with: { name: playwright-report, path: playwright-report/ }
Upload the trace on failure. An E2E failure without a trace costs an hour of guessing; with one it costs five minutes.
Check: force a failure and confirm you can open the trace and see the exact step, the DOM at that moment, and the network log.
The wrong first E2E test beside the right one
The usual first attempt | This one | |
Journey | whatever was easiest | the one whose silence costs money |
Selectors | CSS classes, nth-child | roles, labels, data-testid |
Setup | clicks through login | created via API |
Waits | waitForTimeout(2000) | condition with a timeout |
Asserts | a confirmation message | the message and the read-back |
Survives a redesign | no | yes |
You're done when
Renaming a CSS class doesn't break the test.
No waitForTimeout remains in the file.
Stubbing the write makes the read-back assertion fail on its own.
A CI failure gives you a trace you can actually read.
You have one of these, not eight.
Troubleshooting
Passes locally, fails in CI. Almost always timing on a slower runner, and almost always a fixed wait. Step 5.
Fails only when the full suite runs. Tests are sharing a user or a fixture record. Make the test account unique per run.
"Element not found" but it's visibly on screen. Usually an iframe or shadow DOM, or two elements match and Playwright refuses to guess. Narrow the query.
Broke after a redesign. Check what changed: if the journey changed, the test should change. If only the markup changed, your selectors are still coupled — Step 3.
It's flaky and you're tempted to retry. A retry hides the failure without diagnosing it. Run the discriminating experiments first.
Next
If you now want to know whether a deployment is healthy rather than whether the code is correct, that's a different instrument — Build a Smoke Test Suite.
Part of the Software Quality Engineering guide — ShiftQuality's complete map to building quality in.


