top of page

A Practical Testing Workflow for Solo Devs

Shawn West
Jul 8
10 min read

Updated: 2 days ago

Corrected, 4 October 2026. We re-checked this post's figures. One line attributed a coverage target to the test pyramid, which describes a ratio of test types rather than a coverage number; it now describes the target as common advice.

It was a Sunday night. You pushed a small change to the pricing copy, the kind of edit that could not possibly break anything, and went to bed. On Monday a customer emailed to say the checkout button did nothing. You had shipped a broken payment flow to every visitor for fourteen hours, and you found out from a stranger. The fix took four minutes. The refunds, the apology email, and the sick feeling in your stomach took the rest of the day.


Then comes the familiar loop. You resolve to "start testing properly." You open a tutorial. It opens with a test pyramid, coverage gates, a QA hand-off, and a CI matrix across three browsers. You are one person shipping between a day job and dinner. That advice does not tell you where to start — it tells you the bar is so high you will never clear it. So you close the tab, ship untested, get burned again, and feel guilty again.


The problem is not that you are undisciplined. The problem is that almost every word written about testing assumes a team, and you are not a team. This is the workflow for one person: catch the bugs that would actually hurt you, and deliberately ignore the ones that would not. If you have never asked why we bother writing tests at all, start with the fundamentals — this piece assumes you already believe the ledger and just need it to fit one pair of hands.



Why the team advice quietly fails a solo dev


The test pyramid is not wrong. It is a load-balancing model for a group: many cheap unit tests at the bottom, fewer expensive end-to-end tests at the top, so that a team of engineers can change code all day without stepping on each other. Its unit of concern is throughput across many hands.


Your unit of concern is different. You have one pair of hands and a fixed number of evenings. When conventional advice tells you to aim for seventy or eighty percent coverage, it is optimizing a resource you have plenty of (other people's review time) and ignoring the one you are actually short on (your own hours this week). Coverage percentage is the wrong number for you because it counts lines exercised rather than disasters prevented. A suite that covers eighty percent of your settings screen and zero percent of your payment flow scores well and protects nothing.


So drop coverage as a target. The right question is not "how much of my code is tested?" It is "which failures can I not afford?" That reframing is the whole workflow. Everything below is just applying it.


Do this week: open your app and write down, in one line, the single failure that would most embarrass or cost you if it shipped tonight. That sentence is your first test's job description. If you cannot name it, you are not short on tests — you are short on a decision about what matters.


The discovery move: the filter that decides what to test


Before you write a single test, you make one pass over your product and sort it. This is the discovery step, and it is the part solo devs skip — they jump straight to "how do I test this" without first deciding whether this deserves a test at all. Testing fails in discovery, not in code: a suite that tested the wrong things thoroughly is worse than no suite, because it costs maintenance and buys nothing.


The filter is a single question, asked of every feature:


"If this breaks silently for a day, does it cost me money, corrupt data, or destroy the core value of the product?"


If yes, it earns a test. If no, it rides. Note the word silently — the failures that hurt are the ones no one reports until damage is done. A crash on your marketing page is loud and someone will tell you fast; a checkout that charges the wrong amount is quiet and compounds.


Run the filter across your real feature list and it sorts itself:


Area

If it breaks silently for a day

Verdict

What to write

Checkout / payment

Lost sales, wrong charges, refunds, chargebacks

Test it

One end-to-end test through a real purchase

Auth / login / password reset

Users locked out, or worse, logged into the wrong account

Test it

End-to-end happy path + one wrong-password case

Data export / delete

Silent corruption or loss the user cannot undo

Test it

Integration test on the write path, not the button

Account settings (theme, notifications)

Someone's dark mode reverts; a mild annoyance

Let it ride

Nothing until it actually breaks twice

Marketing / about page

Looks off; loud, caught in minutes, reversible

Let it ride

Nothing; eyeball it on deploy


The two "let it ride" rows are the point. Deciding not to test the settings screen is not laziness — it is the same discipline that lets you afford a real test on checkout. Every test you refuse to write is time returned to the two or three that matter. A solo dev who tests everything a little tests nothing enough.


Do this week: make this exact table for your own product. Four to eight rows. If everything lands in "test it," your filter is too loose — re-ask the question with the word silently and be honest about which failures a user would actually report within the hour.


Step 1: Protect the critical path with a few end-to-end tests


The critical path is the one or two flows that make the product worth paying for, plus anything touching money or user data. For a store it is browse → add to cart → pay. For a SaaS tool it is sign up → do the core action → see the result. You almost certainly have one or two of these, not ten.


Write a small number of end-to-end tests that drive these flows the way a user does — a real browser, a real click on the real button, a real assertion that the confirmation appears. These tests are slow, a little fragile, and expensive to maintain, and that is an acceptable price for exactly the flows where a silent break costs you the business. It is the wrong price everywhere else, which is why you are writing two of them and not two hundred. (If you want the discipline for keeping these stable rather than flaky, end-to-end testing without the pain is the companion piece.)


The single rule that makes them worth having: run them automatically before every deploy. A critical-path test you have to remember to run is a critical-path test you will skip on the Sunday night you most need it. Wire it into your deploy command so a red checkout test blocks the ship. If you want an even lighter version of this as a first step, a five-minute smoke test that just confirms the app boots and the core screen loads is a legitimate place to start before you write a full purchase flow.


Do this week: write one end-to-end test for your single most important flow and make your deploy script refuse to ship if it fails. One test that runs automatically beats ten that live in a folder you never open.


Step 2: Make every bug fix permanent


This is the highest-leverage habit in the entire workflow, and it costs almost nothing because the hard work is already done.


Every time you fix a bug, before you fix it, write a test that fails on the broken behavior. Then fix the code until the test passes. You are not inventing a scenario or guessing at edge cases — the bug already happened in the real world, so you know it is real, you know how to reproduce it, and you know exactly what "correct" looks like. The test practically writes itself from the reproduction steps. (This is one more reason to write down a proper repro when a bug lands — a good bug report is a test spec you already paid for.)


Do this for a year and something quietly powerful happens: your test suite becomes a precise record of every way your product has ever actually broken. Not every way it could theoretically break — every way it has. That is a far better suite than any coverage target would have built, because it is shaped by real failures instead of imagined ones. It is the only suite that gets more valuable specifically in the places you have already been hurt.


Do this week: the next bug you fix — even a tiny one — write the failing test first. Just once. You will feel how little extra effort it is when the reproduction is already in front of you, and that is the habit installed.


A week in the life of this workflow


(Developed example — composite scenario.)


Maya runs a small paid newsletter tool by herself. Subscribers pay monthly; the core value is that scheduled issues actually go out on time. She has a day job, so the product gets four or five evenings a week if she is lucky.


She runs the filter across her feature list. Three things touch money or core value: the Stripe checkout for new subscriptions, the "send now / schedule" flow that delivers issues, and the subscriber import that writes email addresses to her database. Everything else — the editor's dark mode, the account-settings avatar upload, the public landing page — she consciously drops in the "let it ride" column. That is the discovery pass, and it takes twenty minutes, not an evening.


For the critical path she writes exactly one end-to-end test, driving a real browser through a real (test-mode) purchase:


def test_new_subscriber_can_pay_and_gets_access(page):
    page.goto("/pricing")
    page.click('[data-testid="subscribe-monthly"]')

    # Stripe test-mode card
    page.fill('[data-testid="card-number"]', "4242 4242 4242 4242")
    page.fill('[data-testid="card-expiry"]', "12/34")
    page.fill('[data-testid="card-cvc"]', "123")
    page.click('[data-testid="confirm-payment"]')

    # The thing that actually matters: they're in.
    page.wait_for_url("/welcome")
    assert page.get_by_text("Your subscription is active").is_visible()

Her deploy script runs this before every push. It is slow — about twenty seconds — and once a month it flakes on a Stripe timeout. She accepts both costs, because this is the one flow where a silent break stops all revenue.


Then, a Tuesday. A subscriber emails: they scheduled an issue for 9:00 a.m. and it never sent. Maya digs in and finds the bug — issues scheduled for a time that included a leading-zero hour ("09:00") were being parsed wrong and silently dropped. Textbook silent failure: no error, no alert, just a customer who did not get what they paid for.


Before touching the parser, she writes the failing test:


def test_issue_scheduled_for_morning_hour_is_queued():
    issue = Issue(scheduled_for="09:00")
    queue = schedule(issue)
    # Before the fix, this issue vanished. It must be queued.
    assert issue.id in queue.pending_ids

It fails, reproducing the bug exactly. She fixes the parser; it passes. That test now lives in the suite forever. The specific way her product hurt a paying customer can never quietly return, and it cost her about ninety extra seconds on top of a fix she was making anyway.


What Maya deliberately does not test, all week: the avatar upload on the settings page (breaks loud, matters little), the landing-page hero animation (cosmetic, reversible in a minute), and the dark-mode toggle (annoying if broken, costs nothing). Three features untested on purpose — so the two that could end her business stay covered. That refusal is the workflow, as much as the two tests she wrote.


Step 3: Let the free tools do the boring half


Some whole categories of bug can be caught without writing tests at all, by tools that run while you type. For a solo dev this is the best return on effort in the entire practice, because the work is configuration once, not maintenance forever.


  • A type checker (TypeScript, or Python type hints with a checker) catches the "passed the wrong shape into this function" class of bug — a large, dull, common category — before you even save.

  • A linter catches the unused variable, the unreachable branch, the await you forgot.

  • Static analysis and a formatter remove a whole genre of "how did that ship" mistakes that would otherwise reach production.


Treat these as tests you get for free. They will never cover your checkout flow — that is still on you — but they will quietly eliminate the low-grade bugs that would otherwise eat the evenings you need for the tests that matter. Where the pyramid-versus-trophy debate lands for a solo dev is exactly here: lean toward the "trophy" shape — a fat middle of type-checked, statically analyzed integration-level confidence — rather than a broad base of hand-written unit tests you do not have time to maintain.


Do this week: turn on a type checker and a linter in your editor and fix what they flag today. That is an afternoon of setup that pays out every single day afterward, with no test to maintain.


When to add more — and the permission to stop


The signals that you have earned more testing are concrete, not moral. Add tests when: the same area breaks twice (the second break is the product telling you it belongs in the "test it" column), when a flow starts touching money or data that did not before, or when you are about to change something risky and want a net under you before you jump.


That last one deserves its own habit. When you are about to build or refactor something genuinely scary — the billing proration logic, the data migration — write the tests as you build, because that is the moment your understanding of "correct" is sharpest and the cost of a mistake is highest.


Everything else can keep riding. And here is the part the tutorials never give you: shipping code you deliberately chose not to test is not negligence, it is triage. A solo dev with three solid tests on checkout, auth, and data — and nothing else — is in far better shape than one who burned out chasing eighty percent coverage and now tests nothing because the suite became a second job. The guilt is optional. Drop it. You are not aiming for a testing culture. You are aiming to never again learn that checkout is broken from a stranger's email.


Sources


  • The end-to-end purchase flow and Maya's newsletter tool are an illustrative composite scenario, not a documented client. They are constructed to demonstrate the filter and the fix-then-test habit; the code is representative, not copied from a real system.

  • The test pyramid originates with Mike Cohn, Succeeding with Agile (Addison-Wesley, 2009). The "testing trophy" framing is Kent C. Dodds' publicly published model. Both are referenced here as named models, not as measured claims — no statistics are asserted about either.


Keep learning. This article is part of the Software Testing Foundations path in the ShiftQuality Learning Center.



Part of the Software Quality Engineering guide — ShiftQuality's complete map to building quality in.

bottom of page