Writing Testable Code: Patterns That Help
The billing service's late-fee function had 94% line coverage, and every test in its file passed. Each test also set up seven mocks: the database, the payments client, the email sender, the clock, the holiday calendar service, the feature-flag client and the audit logger. When finance reported that customers had been charged late fees for payments that fell due on a bank holiday, the team found the bug in ten minutes. The holiday calendar mock in every test returned "not a holiday", because writing a realistic one was too much work. The rule that mattered most had never been exercised.
This example is a composite, built from a common pattern in billing and payments code. The figures are illustrative.
The tests weren't lazy. They were as good as the code allowed. When the business rule (is this payment late, and by how much?) lives in the same function as six pieces of infrastructure, every test of the rule has to fake all six, and the faking is where the rule gets lost.
Why "just write more tests" doesn't fix it
Coverage targets push teams towards testing whatever exists, in whatever shape it's in. Code with tangled side effects can reach high coverage, but only with heavy mocking, and heavily mocked tests have two well-known problems. They check that the code calls its collaborators in a particular order, rather than that it produces the right answer. And they break whenever the implementation changes, even when the behaviour doesn't, which teaches the team to distrust failures.
The fix is in the design, not the test suite. Michael Feathers, in Working Effectively with Legacy Code (2004), defines a seam as a place where you can alter behaviour in your program without editing in that place. Testable code is code with seams in the right places: between the rules and the world.
How tangling happens, and why AI assistants reproduce it
Code gets tangled in the shape of the request. A ticket that says "apply the late fee and email the customer" tends to produce one function that does both, because that's the sentence. AI coding assistants make this pattern faster, not different: ask one for that function and you will usually get exactly that function, with the clock read inline and the email sent from the middle of the calculation. Ask it for tests afterwards and it will typically mock everything the function touches, because that is the only way to test the shape it wrote.
The chain is: request phrased as one action → one function mixing rules and I/O → tests must fake the I/O → fakes default to the easy case → the rule's edge cases are never run.
The discovery move is upstream of the code: when writing the ticket, or the prompt, name the rule separately from the actions. "Rule: a payment is late if received after the due date, where a due date on a non-business day moves to the next business day. Actions: record the fee, notify the customer." That sentence produces different code, from people and from assistants.
Walking the late-fee function apart
Here is the original, simplified.

def apply_late_fee(account_id):
account = db.accounts.get(account_id)
payment = payments_client.latest_payment(account_id)
today = datetime.now().date()
due = account.due_date
if holiday_service.is_holiday(due):
due = due + timedelta(days=1)
if payment is None or payment.received_on > due:
days_late = (today - due).days
fee = min(days_late * account.daily_fee, account.fee_cap)
db.fees.insert(account_id, fee)
email.send(account.owner, "late_fee", fee=fee)
audit.log("late_fee_applied", account_id, fee)
return fee
return 0
Two bugs hide in it. A holiday that falls on a Friday moves the due date to Saturday, not to the next business day. And days_late is counted from today, not from when the payment was received. Neither is visible in tests where the holiday mock always says no and the clock is frozen to one date.
Pattern 1: functional core, imperative shell. Gary Bernhardt's name for it (from his 2012 Destroy All Software screencast and Boundaries talk) is the clearest: put the decisions in pure functions that take values and return values, and keep the I/O in a thin outer layer that gathers inputs and acts on outputs.
# Core: pure. No clock, no network, no database.
def effective_due_date(due, holidays):
while due.weekday() >= 5 or due in holidays:
due += timedelta(days=1)
return due
def late_fee(due, received_on, as_of, holidays, daily_fee, fee_cap):
due = effective_due_date(due, holidays)
settled = received_on or as_of
if settled <= due:
return 0
return min((settled - due).days * daily_fee, fee_cap)
# Shell: gathers inputs, applies the decision. Little logic, few branches.
def apply_late_fee(account_id, db, payments, calendar, clock, notifier):
account = db.accounts.get(account_id)
payment = payments.latest_payment(account_id)
fee = late_fee(
due=account.due_date,
received_on=payment.received_on if payment else None,
as_of=clock.today(),
holidays=calendar.holidays_between(account.due_date, clock.today()),
daily_fee=account.daily_fee,
fee_cap=account.fee_cap,
)
if fee:
db.fees.insert(account_id, fee)
notifier.late_fee(account.owner, fee)
return fee
Pattern 2: time, randomness and environment are inputs. datetime.now() inside a rule is a hidden input that tests can't vary. Passing as_of (and a set of holidays) makes the bank-holiday case a one-line test.
Pattern 3: dependencies arrive explicitly. The shell receives its collaborators as parameters (or through a constructor) instead of reaching for globals. That is the seam Feathers describes: tests, and other environments, can supply a different implementation without editing the shell.
Now the tests that matter need no mocks at all:
def test_friday_holiday_moves_due_date_to_monday():
fri = date(2026, 12, 25) # Christmas Day, a Friday
assert effective_due_date(fri, holidays={fri}) == date(2026, 12, 28)
def test_fee_counts_from_receipt_not_today():
assert late_fee(due=date(2026, 3, 2), received_on=date(2026, 3, 5),
as_of=date(2026, 3, 20), holidays=set(),
daily_fee=5, fee_cap=50) == 15
Both of those tests fail against the original function's logic. The shell gets one or two integration-style tests that check it wires the pieces together, with simple fakes. The mock count per test drops from seven to zero for the rule tests, and the rule's edge cases (weekends, holidays, caps, unpaid accounts) can be listed exhaustively because each test costs one line.
How to find the tangled code in your own codebase
You don't need to read every file. These counts find the knots.
The mock count. For a module's test file, count mock or patch objects per test. A handful of modules will stand out with many more than the rest. As a rule of thumb (not a benchmark), more than three mocks per test usually means a rule is sharing a function with I/O. Start there.
The clock search. Search business-logic directories for direct reads of the current time: datetime.now, Date.now(), new Date(), System.currentTimeMillis, DateTime.Now. Each hit inside a rule is a hidden input and a likely untested edge case.
The "can I run it on a plane?" test. Pick a business rule and try to exercise it in a REPL with no network and no database. If you can't call it without standing up infrastructure, it isn't separated.
The refactor-break test. Rename a private helper or reorder two calls without changing behaviour. Count the tests that fail. Tests that break on a pure refactor are testing implementation, and they'll slow down every future change. (See writing tests that survive refactoring.)
Choosing the right amount of separation
Not everything needs a pure core.
Thin CRUD endpoints with no real rules gain little from extraction. Test them through an integration test against a real (containerised) database.
Orchestration-heavy code, where the logic is the sequence of calls, such as a saga or a retry policy, is legitimately tested with fakes. Prefer simple hand-written fakes over mocking libraries that assert on call order. Mocking, stubbing and faking covers when each fits.
Performance-critical paths sometimes interleave I/O and computation for good reasons. Keep the rule pure where you can, and document where you can't.
Legacy code can't be split in one go. Feathers' approach is to find or create a seam, get a characterisation test around the current behaviour, then extract. The adding tests to legacy code tutorial walks through it.
Making it the default, including for AI-written code
Write rules and actions separately in the ticket or prompt. One sentence for the rule, one list for the side effects.
Ask for the core first. When using an assistant, request the pure function and its tests before the code that calls it. Review the core's tests for edge cases you know about: the assistant can only list the ones the prompt implies.
Add the mock count to code review. A new test with more than three mocks is a question for the author: is there a rule in here that wants extracting?
Ban direct clock reads in domain code with a lint rule, and provide a clock interface.
Track it. The number of rule functions with direct I/O, per module, is a cheap trend line.
Final takeaway
Testability is a design property: whether the rules can be exercised without the world. Code that mixes decisions with side effects can still reach high coverage, but only through fakes that quietly default to the easy case, which is where the real bugs wait. Separate the core from the shell, pass time and dependencies in, and the important tests become short, mock-free and exhaustive.
Your next action: run the mock count across your test suite this week and take the worst module. Extract one rule into a pure function, write the edge-case tests for it, and see how many of them fail against the current behaviour. For the review side of AI-authored code, read What AI-Generated Code Structurally Skips.
Sources
Feathers, M., Working Effectively with Legacy Code, Prentice Hall (2004). Definition of a seam: a place where you can alter behaviour in your program without editing in that place.
Bernhardt, G., "Functional Core, Imperative Shell", Destroy All Software screencast (2012), and "Boundaries" talk (SCNA / RubyConf 2012). https://www.destroyallsoftware.com/screencasts/catalog/functional-core-imperative-shell
Part of the Software Quality Engineering guide — ShiftQuality's complete map to building quality in.


