Specification by Example: A Hands-On Tutorial
Updated: 2 hours ago
Everyone nods at the rule. The nod is the problem: an abstract sentence lets six people hold six different systems in their heads and never find out, because agreement at that altitude costs nothing. A table costs something — you cannot fill a cell with a shrug. Here's the fork that argument produces, the rows worth writing, and a diagnostic you can run on one rule this week.
The sentence was one line in a backlog refinement, and it took four seconds to get through the room:
Premium customers get free shipping.
Nobody objected. Nobody had a question. It went into the story, the story into the sprint, and the sprint delivered something that did what it said.
Eleven days after release, a customer-service lead forwarded a complaint from a Premium member in Anchorage charged $34 for shipping. The engineer who read it was genuinely puzzled — the code was correct. The rule said Premium customers get free shipping, and the code gave Premium customers free shipping, on the ground service the carrier does not run to Alaska.
The requirement was not badly written. It was written at an altitude where being wrong is impossible, which is a different defect. This is what specification by example is for, and it is neither documentation nor test automation. It is a device for making disagreement visible early enough to be cheap.
The nod is the failure signal
An abstract rule is a compression. "Premium customers get free shipping" compresses a customer status model, a threshold, a carrier zone map, a promotions interaction, and a returns policy into six words. Each reader decompresses them using assumptions from their own corner of the business: ops reads "free shipping" as our standard ground service, the one in the contract; marketing reads "Premium" as anyone who was Premium recently enough to still feel like one; finance reads it as the tier flag on the invoice at time of capture.
All three readings are reasonable. All three produce different software. And none of them will surface in a discussion held at the level of the rule, because at that level everyone is agreeing with their own version.
A concrete example forces the decompression into the open. You cannot write "Premium" in a cell — you must write which Premium, as of when. You cannot write "free" — you must write $0.00 and explain why it isn't $34.00. The row is not the deliverable. The argument you have while filling the row is the deliverable, and the row is what's left over afterwards.
Run this today: take the last requirement your team agreed to without a single question. Write one sentence describing what you pictured, and have someone else from that meeting do the same, separately. Compare. Fast agreement is not evidence of shared understanding — it is the absence of evidence either way.
One fork, walked: five rows and a forty-minute argument
The following is an illustrative composite drawn from patterns common in mid-market retail and fulfilment work. The values and identifiers are constructed to make the mechanism visible; they are not measured from a specific engagement.
Call the retailer Ardwick Supply — 400 SKUs, a Premium membership at $99/year, a warehouse in Ohio, one carrier contract. After the Anchorage incident the team ran a forty-minute session with three people: the product manager who owned the membership programme, the engineer who wrote the shipping calculator, and the QA lead. The instrument was one table, columns agreed before any rows were written — the columns are the model, and getting them wrong quietly is how you end up with a hundred examples encoding the same blind spot.
Five rows came out of that session, with the disagreement recorded rather than smoothed over:
# | Membership status | Subtotal (after promo) | Ship-to | Cart flag | Expected shipping | What the room actually said |
1 | Premium, active | $80.00 | Columbus, OH | — | $0.00 | Instant, unanimous, boring |
2 | Premium, active | $80.00 | Anchorage, AK | — | $18.00 | Ops: contract has no free rate outside the lower 48. PM had never seen the contract |
3 | Premium, expired 3 days ago | $80.00 | Columbus, OH | — | $8.95 | Twenty minutes. Marketing wanted a grace period; billing had none |
4 | Standard | $80.00 list, $62.00 after SPRING20 | Columbus, OH | — | $8.95 | Threshold is $75 — but $75 of what? Nobody knew |
5 | Premium, active | $80.00 | Columbus, OH | oversize (60 lb floor scale) | $0.00 + $45.00 surcharge | Engineer: the surcharge isn't shipping, it's a line item. PM: the customer won't care |
Row 1 is the row everybody expects to write, and it produced no information: it confirmed a shared understanding never in doubt. The main success case is where the least information is — worth writing exactly once, as an anchor, then leaving.
Row 3 is where the session earned its forty minutes, so follow that one to the end.
The disagreement was not about the rule. It was about a fact nobody had checked: what does membership.status mean at the moment an order is placed? The engineer went and looked while the room waited. The status was computed in two places. Billing flipped a member to expired at midnight UTC on the renewal date; the web session cached the tier at login and never refreshed it, so a member who logged in Tuesday and checked out Friday was still Premium in the cart and expired on the invoice. That is not an edge case — that is two systems disagreeing about a fact, found because someone had to put a number in a cell.
The room then made a decision — and this is the part that matters, because a session ending in "interesting, let's look into it" has produced nothing:
Shipping eligibility reads membership status as of the order-placement timestamp, server-side, never from the session cache. The cart shows an estimate and says so.
There is no grace period. Marketing wanted one; the room killed it, because a silent grace period is indistinguishable from a bug to everyone downstream — support cannot explain it, finance cannot reconcile it, QA cannot test a rule nobody wrote down. Win-back gets emails at day 1 and day 14 instead: a marketing mechanism, not a hidden branch in the pricing engine.
The schema gained one column, membership_status_at_order, denormalised onto the order record — because the answer to "was this person Premium?" six months later must be recoverable without replaying a cache.
One row, one afternoon, one schema change, one feature killed on purpose. The alternative history is the grace period getting built quietly by whoever picks up the ticket, and the reconciliation problem surfacing in a finance close eight months later.
Run this today: take your team's most-used business rule and write row 3 — the expired, lapsed, or just-ended version of the actor. If two people give different answers, you have found a real one on your first try.
The discovery move: ask for the row, not the rule
The move is small enough to be forgettable, which is why it needs a name.
When someone states a rule, do not ask "is that right?" — they will say yes, because it is right at the altitude they said it. Ask instead:
"Give me the specific case. What's the actual number, for an actual customer, on an actual day?"
Three parts, all load-bearing. Specific case blocks the retreat into another abstraction. Actual number means someone must know it or go find it — going to find it is where Ardwick discovered the two-place status computation. On an actual day smuggles in time, where a startling share of requirements defects live: statuses expire, prices change, promotions end, and rules stated in the present tense hide all of it.
Then the follow-up that turns one row into two: "Now change exactly one thing so the answer flips." That is how rows 1 and 3 became a pair. A single example tells you what happens; a contrasting pair tells you what the system is keying on. If you cannot construct a near-neighbour with a different outcome, you have not found the rule — you have found a coincidence.
This is discovery work, not testing work, and it belongs upstream of the sprint. If the rules arriving at refinement are consistently too abstract to fill a table, the fix is further back — Writing Good Requirements covers turning a vague ask into something with a failing case.
Run this today: in your next refinement, ask for one specific number instead of one clarifying question. Count how often the answer is "I'd have to check." Each is a fact your team was about to build on without holding.
What makes a row worth writing
Not all examples pay. Three properties separate a row that carries information from one that carries only maintenance cost.
Specific. Real values, not placeholder categories. "A valid amount" is an abstraction wearing an example's clothing — all the ambiguity of the original rule, none of the brevity.
Contrasting. Rows come in pairs differing in one variable. One passes, one fails, and the delta between them names the rule.
Minimal. Every column that does not change an outcome is a column you maintain forever. If the customer's name never affects shipping, the customer has no name in this table.
Applied to the vague phrases that show up in real backlogs:
Vague phrase | The question it hides | A row that settles it |
"Large orders ship free" | Large by what measure, before or after discounts? | Subtotal $74.99 after SPRING20 → $8.95; $75.00 → $0.00 |
"Recent customers get the offer" | Recent as of when, measured from what event? | Last order 2026-05-19 → eligible; 2026-05-18 → not eligible (90-day window, order date not ship date) |
"Refunds are processed promptly" | Prompt is not a value; what is the observable? | Refund requested Fri 16:45 ET → settled by Tue 17:00 ET (2 business days, excluding request day) |
"Users can't order out-of-stock items" | Reserved stock, backorder, or oversell tolerance? | On hand 3, reserved 3, requested 1 → blocked; on hand 3, reserved 2, requested 1 → allowed |
"Premium customers get free shipping" | Which Premium, as of when, to where? | See rows 1–5 above |
Each right-hand cell is a sentence you can hand to an engineer and a sentence you can hand to a tester, and they are the same sentence. That is the property worth having.
Run this today: scan your current sprint for acceptance criteria containing a comparative word — large, recent, fast, prompt, high, sufficient. Each is a missing threshold. Replace one with a contrasting pair and see whether anyone disagrees with your numbers.
Where the rows come from: the boundary catalogue
Teams that struggle to generate examples generate them the same way — imagining a user doing the thing successfully, repeatedly. That produces variations on row 1. The rows that pay come from a small reusable catalogue of positions where systems characteristically disagree with themselves. Work the catalogue rather than your imagination:
Boundary generator | Applied to free shipping | Why it finds things |
Exactly at the threshold | Subtotal $75.00 | Inclusive or exclusive is decided by a > vs >= that nobody discussed |
One unit under | Subtotal $74.99 | Confirms the direction; catches rounding on tax-inclusive pricing |
Expired / lapsed status | Premium ended 3 days ago | Two systems usually compute status in two places |
Zero and empty | Cart of 0 items; $0.00 gift-card order | Division, minimums, and "at least one" assumptions live here |
After a partial reversal | 3 items ordered, 2 returned, subtotal now $41.00 | The rule was written forward-only; refunds run it backwards |
Outside the normal geography | Anchorage, AK; APO address; PO box | Contracts and carriers have exclusions the rule-writer never read |
Two rules interacting | SPRING20 plus Premium plus oversize | Precedence is almost never specified, and is almost always assumed |
The last row is the one most teams never reach and the one that breaks in production most often. Individual rules get reviewed; the interaction of two rules is reviewed by nobody, because it belongs to neither story. Ardwick's row 5 produced a checkout screen showing Shipping: FREE directly above Oversize handling: $45.00 — technically correct, reliably ticket-generating.
Five to ten rows per significant rule is a reasonable range: enough to cover the catalogue, few enough to keep current. A rule needing thirty rows is doing the work of three, and splitting it helps the code as much as the table — the same instinct as the small and testable halves of INVEST.
Run this today: walk one rule through the seven generators, keeping only the rows that produce a real answer. Note which generator produced the first genuine disagreement — in our experience it is almost always "expired / lapsed" or "two rules interacting."
The tool is not the practice
This is where the practice most often goes wrong, and the failure has one shape: a team adopts Cucumber, writes Gherkin, and concludes six months later that specification by example does not work.
The historical record settles it. Ward Cunningham's FIT — the Framework for Integrated Test — executed examples written as plain tables inside ordinary HTML documents, wired to the system through small fixture classes. The primitive was the table, not a syntax. Gherkin came later and wrapped the table in Given/When/Then plus a Scenario Outline with an Examples block (Cucumber Gherkin reference) — convenient, but not load-bearing. Gojko Adzic's framing in Specification by Example puts collaborative specification and refinement at the centre, with automated validation serving a further goal: keeping a living documentation system trustworthy, so the specification and the system cannot silently diverge. Automation stops examples from rotting into a lie; it is not what made them valuable.
Two consequences follow. First, examples written after the code are a different artifact — regression tests, useful but with zero discovery value, because the disagreement they would have exposed was already resolved unilaterally by whoever wrote the implementation. Second, keep implementation detail out of the row. "Given the orders table contains a record with tier_id = 2" is a fixture wearing a business rule's clothing; it breaks when the schema changes even though the rule did not. The columns must be readable by whoever owns the policy — which is also how you place them afterwards: rules varying per feature go in that story's acceptance criteria, universal rules in the baseline, the distinction Definition of Done vs Acceptance Criteria draws.
Run this today: count how many of your automated scenario files were written before the corresponding code merged. Your commit history knows. That percentage is the share of your suite doing discovery; the rest is regression testing — fine, but not specification.
The exception: when examples don't pay
This practice costs real money — session time from three people, plus a table that must stay current or turn actively misleading. Three situations where that cost is not repaid:
The rule is not stable yet. Pre-product-market-fit pricing, a promotion tuned weekly, an early-stage eligibility model. Working out boundary behaviour for a rule that will be deleted in three weeks is ceremony. Wait until something downstream depends on it.
The work is exploratory rather than rule-governed. Design iteration, performance investigation, and most exploratory testing search for the questions rather than answering known ones. Examples close a space down, which is exactly wrong when the job is to open one up.
The behaviour is genuinely non-deterministic or judgement-based. Ranking quality, recommendation relevance, an LLM's summary, a fraud score — there is no single expected value for the cell. The substitute is a property or a bound: no result set may contain an out-of-stock item; the flagged rate stays between 0.8% and 1.4% over a rolling week; no summary introduces a proper noun absent from the source. Those are checkable and they fail loudly. A fabricated "expected output" for a probabilistic system is worse than nothing — it manufactures a false red or green and trains the team to ignore the suite.
One asymmetry survives all three. Even when you skip the table, write down the one number the rule turns on — the threshold, the window, the cutoff — and where it came from. Cheap now, expensive to reconstruct in a year, and exactly what the next person needs when someone in Anchorage gets charged $34.
The five-row diagnostic
You need neither a workshop nor a tool. Take one business rule everyone believes is settled and spend fifteen minutes:
Write the column headers first. What facts does the outcome depend on? If two people list different columns, stop — you have found the disagreement already.
Write row 1, the boring one. Expect unanimity. It is the anchor, and the only row where consensus is meaningful.
Write four boundary rows from the catalogue — at the threshold, one unit under, an expired status, two rules interacting.
Have everyone fill the outcome column independently, on paper, before anyone speaks. Spoken sessions converge on whoever is most senior; simultaneous private answers do not.
Count the rows where the answers differ. That number is the disagreement your team was about to build on.
Zero is plausible for a mature, well-understood rule. One or two is normal, and is the practice paying for itself in fifteen minutes. Four out of five means the rule is not one rule.
Ardwick's engineer had written correct code against a correct requirement. The $34 charge in Anchorage was neither a coding defect nor a testing gap — no test could have caught it, because the expected result was never established. It was a decompression failure: six people reading six words, each silently supplying the missing model from their own department. The table does not prevent that by being a better document. It prevents it by making the room say a number out loud while there is still time to be wrong for free.
Next: if your examples keep colliding, the unit of work is probably too big — Use Cases vs User Stories: When to Pick Which covers choosing the container that fits the rule.
Sources
Gojko Adzic, Specification by Example: How Successful Teams Deliver the Right Software (Manning, 2011) — the key process patterns, collaborative specification, and living documentation.
Gojko Adzic, Bridging the Communication Gap: Specification by Example and Agile Acceptance Testing (Neuri, 2009) — the earlier treatment of examples as a shared-understanding device.
Cucumber, Gherkin Reference — keyword list including Scenario Outline and Examples, and the data-table syntax.
Matt Wynne & Aslak Hellesøy, The Cucumber Book: Behaviour-Driven Development for Testers and Developers (Pragmatic Bookshelf, 2012) — automating examples without letting the tool define the practice.
Ward Cunningham, FIT: Framework for Integrated Test — executable examples expressed as tables in ordinary documents, the precursor to modern example-based specification.
Part of the Software Quality Engineering guide — ShiftQuality's complete map to building quality in.


