Write a Useful Integration Test — Hands-On Software Testing, Part 3
Updated: 4 hours ago
Illustrative composite: Tolvern Freight's shipping service, the system used across this path. The word doing the work in the title is useful — most integration tests are slower unit tests, and this tutorial is about the difference.
Before you start
You need:
A function that crosses a seam — one that writes and then reads, spans two tables, or runs inside a transaction. A function that only transforms its arguments is a unit test, and this tutorial won't apply.
A way to run a real instance of your database in tests: Testcontainers, Docker Compose, or a dedicated test database you can reset.
Your migrations runnable against an empty database. If schema lives only in a hand-maintained dev database, fix that first — everything here depends on the test database having the real constraints.
About 55 minutes. Examples are Python with SQLAlchemy and Postgres.
What you'll build
One test for book_shipment() that fails when the database rejects a duplicate booking — a failure the existing mocked unit test cannot produce, because the mock has no constraints.
Step 1: Find the seam, not the function (10 min)
Don't start from "which function has no tests". Start from where your code meets something with rules of its own.
Tolvern's book_shipment() writes two rows in one transaction: a shipments row and a carrier_bookings row. The database enforces a unique index on (customer_id, idempotency_key) and a foreign key that is deferred to commit.
The existing unit test mocked the repository. It passed. It could not have failed for either of those reasons, because a Mock() accepts every write cheerfully.
# The existing unit test — green, and blind
def test_book_shipment_creates_booking():
repo = Mock()
book_shipment(repo, valid_request)
assert repo.save.called # asserts that we asked, not that it worked
Check: for your chosen function, name one failure the database could produce that your current test cannot. A unique index, a NOT NULL, a foreign key, a check constraint, a transaction rollback. If you can't name one, this function doesn't need an integration test — and that's a valid outcome worth five minutes.
Step 2: Get a real database, isolated per test (15 min)
The test must run against the real engine with the real schema. SQLite standing in for Postgres will silently accept things Postgres rejects, which defeats the entire exercise.
# tests/conftest.py
import pytest
from testcontainers.postgres import PostgresContainer
from sqlalchemy import create_engine
from sqlalchemy.orm import Session
@pytest.fixture(scope="session")
def engine():
with PostgresContainer("postgres:16") as pg:
eng = create_engine(pg.get_connection_url())
run_migrations(eng) # the real migrations, not metadata.create_all
yield eng
@pytest.fixture
def db_session(engine):
conn = engine.connect()
tx = conn.begin()
session = Session(bind=conn)
yield session
session.close()
tx.rollback() # every row this test wrote disappears
Two details matter more than they look. Run your real migrations, not create_all — the constraint you're testing may exist only in a migration. And roll back rather than truncating; it's faster, and which isolation strategy you're entitled to is a test data decision worth making deliberately.
Check: write a throwaway test that inserts a row violating a known constraint and assert it raises. If it passes silently, your test database doesn't have the real schema and nothing below will work.
Step 3: Write the test the unit test couldn't (10 min)
def test_duplicate_idempotency_key_is_rejected(db_session):
customer = CustomerFactory()
req = shipment_request(customer, idempotency_key="K-1")
book_shipment(db_session, req)
with pytest.raises(IntegrityError) as exc:
book_shipment(db_session, req)
assert "uq_shipment_idempotency" in str(exc.value)
That last line is the one worth arguing for. Asserting on the exception type alone will pass when any constraint fires — including one you broke by accident in a fixture. Naming the constraint means the test can only pass for the reason you intended.
Check: temporarily drop the unique index in a scratch database and re-run. The test must fail. If it still passes, you're catching an exception thrown somewhere earlier — likely in setup — and the test is green for the wrong reason.
Step 4: The decision point — what stays real (10 min)
book_shipment() also calls the carrier's HTTP API. You now have to decide what to keep real, and both choices are defensible in the wrong direction.
The rule that holds: keep real whatever has rules you are testing; stub whatever merely needs to be present.
The database is the point → real. Its constraints, transactions and types are exactly what a unit test can't reach.
The carrier API is scenery → stub it. It's slow, rate-limited, owned by someone else, and its behaviour is a contract testing problem, not this one.
But stub it at the boundary you own, and assert the arguments:
def test_booking_sends_the_negotiated_rate(db_session, carrier_stub):
customer = CustomerFactory(plan="negotiated", rate_card="TOLV-2026-A")
book_shipment(db_session, shipment_request(customer))
carrier_stub.book.assert_called_once()
assert carrier_stub.book.call_args.kwargs["rate_card"] == "TOLV-2026-A"
Getting this backwards in either direction costs you. Stub the database and the test proves nothing. Keep the carrier real and the test is slow, flaky, and fails when a third party has an outage.
Check: for every dependency in the call path, say which category it's in and why. Any you can't classify is a design question the test has just surfaced for you.
Step 5: Prove the test is isolated (5 min)
Integration tests fail as a group far more often than they fail alone, and the cause is almost always shared state.
pytest tests/integration/test_booking.py --count=3 -q # same test, three times
pytest tests/integration -p no:randomly -q # fixed order
pytest tests/integration --shuffle -q # shuffled
Check: all three modes pass. If shuffled fails and fixed order passes, you have an order dependency — something isn't being rolled back, and the discriminating experiments will name it in about a minute. Fix it now; it will only get harder as the suite grows.
Step 6: Put a budget on it (5 min)
Integration tests are the layer that quietly becomes forty minutes. Decide the ceiling before you have eighty of them.
A workable rule: under 200 ms each, and the container starts once per session, not per test. If a single test takes seconds, it is usually recreating the schema — hoist that to a session fixture.
Check: run with --durations=10. If anything in the top ten is over a second, find out why before adding more.
The wrong test beside the right one
Both are called integration tests. Only one can fail for a database reason.
Mocked repository | Real database | |
Asserts | repo.save.called | the row, and the constraint |
Catches a unique-index violation | no | yes |
Catches a deferred FK firing at commit | no | yes |
Catches a type or precision mismatch | no | yes |
Fails if you delete the schema | no | yes |
Runtime | ~2 ms | ~120 ms |
The 118 milliseconds is the whole price. What it buys is a test that can be wrong.
You're done when
Dropping the constraint makes your test fail, by name.
Every dependency is deliberately real or deliberately stubbed, and you can say which and why.
The file passes in fixed order, shuffled, and run three times consecutively.
Nothing in --durations=10 is over a second.
Troubleshooting
The test sees no data even though the write succeeded. Your test session and the application's session are different connections. Bind both to the same connection, as in Step 2.
Everything passes locally, the suite fails in CI. Usually a container that hasn't finished starting. Wait on readiness rather than sleeping — Testcontainers exposes a wait strategy for exactly this.
IntegrityError fires in setup, not in the assertion. Your factory is reusing a value that should be unique per instance. That's a Sequence problem — see Build a Test Data Factory.
The suite got slow after ten tests. The container or schema is being recreated per test. Move it to scope="session".
Rollback isn't cleaning up. Something in the call path commits, or opens its own connection. That test needs truncation instead — and it's the exception, not the new default.
Next
Proving one service is honest about its own writes is a different job from proving two services agree — that's contract testing. And the reasoning behind which surfaces are worth asserting on at all is in API Testing: A Practical Walkthrough.
The next step up is a full user journey, and Tutorial 6: Write Your First E2E Test walks through one.


