top of page

Microservices Communication Patterns That Scale

Shawn West
May 23, 2025
7 min read

Updated: 2 days ago

Communication patterns are usually taught as a menu — synchronous here, async there, add a circuit breaker. The useful framing is narrower: every pattern on that menu exists to stop a healthy service from making a sick one worse.


A freight logistics company's pricing service went slow one Tuesday — not down, just slow, responding in about four seconds instead of eighty milliseconds. Garbage collection, on one node, for roughly ninety seconds.


(Composite example: a mid-sized freight logistics company, drawn from patterns across several organisations. Not a real company.)


The outage lasted fourteen minutes and took out six services.


Booking called pricing and waited. Its request threads filled up, so it stopped answering the API gateway. The gateway's own calls timed out, so it retried — three times, as configured. Every retry created a new four-second wait against a service that was already struggling, and the retries from six upstream services arrived together.


Pricing recovered in ninety seconds. The system took fourteen minutes, because the traffic that hit pricing on recovery was several times normal — the backlog of everyone's retries, all released at once.


Nothing in this outage was broken


That's the part worth sitting with. Every service behaved as designed. The gateway retried because retrying is correct for transient failures. Booking waited because it needed a price. The only actual defect was ninety seconds of GC.


Distributed communication patterns exist because the aggregate of individually-reasonable reactions is not reasonable. Once you read them that way, the menu stops being a list of nice-to-haves and becomes a set of specific answers to specific amplifications:


Pattern

The amplification it stops

Without it

Timeout

One slow call holding a thread indefinitely

Caller's thread pool fills; caller becomes the outage

Circuit breaker

Continuing to call something that's failing

Every request pays full latency to learn nothing

Backoff + jitter

Retries arriving in synchronised waves

Recovery traffic spikes above normal load

Retry budget

Retries multiplying total load

Illustration: 100 clients × 3 retries = 400 requests at the worst moment

Async messaging

Temporal coupling — A can't proceed until B answers

B's downtime becomes A's downtime


The company had exactly one of these. It had retries.


The test you can run: for your last cascading incident, write the timeline in two columns — what broke, and what the healthy services did about it. If the second column is longer, the patterns above are what you're missing.


The default choice is synchronous, and it's the expensive one


Synchronous request/response feels like a function call, which is precisely the problem — it reads as free and it isn't. It creates temporal coupling: A cannot finish until B answers, so B's availability becomes a ceiling on A's.


Chain three of them and the arithmetic gets unpleasant fast. Three services at 99.9% each, called in sequence, give the caller about 99.7% — before you count the latency, which adds rather than averages.


Async messaging removes the coupling by removing the wait: A publishes, B consumes when it can, and B being down means a queue gets longer rather than a request failing. The cost is real and should be named — you give up the immediate answer, you take on message ordering, duplicate delivery and a queue to operate, and debugging becomes reading a trace instead of a stack.


The line that holds up in practice:


  • Synchronous when the caller genuinely cannot proceed without the answer and a human is waiting. Pricing a quote on screen. Authenticating a request.

  • Asynchronous when the work must happen but not now. Sending the confirmation, updating the ledger, reindexing search, anything a human isn't watching.


Most systems have this backwards by default, because synchronous is what you get if you don't decide.


The test: for each synchronous call in a request path, ask what the user would see if it returned nothing. If the honest answer is "nothing different, later", it should be a message.


Retries are the amplifier, and they need a budget


Retrying is right. Naive retrying is what turned ninety seconds into fourteen minutes, and three properties separate the two.


  • Backoff, with jitter. Fixed-interval retries from many clients re-synchronise into waves — the failing service gets hit by the whole population at once, repeatedly. Exponential backoff spreads them out; random jitter stops them re-converging. Jitter is the part people skip and it's the part that breaks the wave.

  • A budget, not a count. "Three attempts" is per-client, and the service experiences the sum. A retry budget caps retries as a fraction of total traffic (Google's SRE book describes clients that only retry while retries stay under 10% of their requests), so a struggling service never sees more than a modest multiple of normal load no matter how many clients are panicking.

  • Only for retryable things. Retrying a timeout is sensible. Retrying a 400 is just load. And retrying a non-idempotent write can double-charge someone, which is a worse outcome than the failure.


That last point is the one with teeth: retries and idempotency are the same decision. If an operation isn't safe to repeat, the retry policy has to know, or you've traded an availability problem for a correctness one.


The test: find your retry configuration and check for jitter and a budget. Most have neither, and most were written during an incident where a single retry would have helped.


The saga: what to do when there's no transaction


Some operations span services and can't be wrapped in a transaction, because there's no shared database to hold one. Booking a freight order there touches capacity, pricing and invoicing, each with its own store.


Two-phase commit is the textbook answer and is brittle in practice — it needs every participant available simultaneously and holds locks while they coordinate, which is precisely the availability profile you left the monolith to escape.


A saga replaces atomicity with a sequence plus a compensating action for each step. Reserve capacity; if pricing then fails, release the capacity. Charge the account; if invoicing fails, refund. There is no rollback — there is a forward-only path back to a consistent state.


The two things that decide whether a saga works:


  • Every step needs a real compensation, and "delete the row" usually isn't one. Compensating a sent email means sending a correction. Compensating a captured payment means a refund with its own failure modes. If a step has no compensation, it must be sequenced last, after everything that could fail.

  • Intermediate states are visible. For a period, capacity is reserved against an order that may not exist. Someone will see it, so the model has to name that state rather than pretend it isn't there.


Sagas are a coordination protocol; the shape of the data they coordinate is a modelling question of its own.


The test: for your longest cross-service operation, write the compensation for each step. Any step where you can't name one is a step that must move to the end of the sequence.


When to push this into infrastructure — and when not to


A service mesh moves timeouts, retries, circuit breaking and mTLS out of application code into a sidecar. Genuinely useful at a certain size, and genuinely not worth it below it.


It earns its keep when you have enough services that per-service implementations have drifted, when you need uniform mTLS and policy, and when you have someone who will own the mesh as a system. It costs you a new operational surface, an extra network hop, and a debugging story where the thing that dropped your request is a proxy nobody on the team configured.


The honest threshold is organisational rather than technical: a mesh is worth it when maintaining consistent communication behaviour across services has become somebody's actual job. As a rule of thumb, under roughly a dozen services a shared client library does the same work for a fraction of the operational cost — and unlike a mesh, it's a shared library that carries no domain logic, so it doesn't couple anything.


The test: ask whether your services currently agree on timeout and retry behaviour. If they do, you don't need a mesh yet. If they don't, a library is the cheaper first fix.


The quality move: test the reaction, not just the failure


Most failure testing checks that a service copes when a dependency is down. The outage above came from a dependency being slow, and from the healthy services' reactions to it. That is testable, and it is cheap to start:


  • Inject latency, not just errors, into one dependency in a pre-production environment, and watch the callers' thread pools and retry counts.

  • Assert on the retry budget: under injected failure, total requests to the struggling service should stay within a small multiple of normal.

  • Check idempotency for every retried write. The retry that double-charges a customer is a test case, not a review comment.


Timeouts, retry rules and idempotency are also exactly what code generated from a one-line ticket leaves out, because the ticket never said what happens when a call is slow or runs twice (the gaps generated code leaves). Put the timeout and the retry rule in the spec, where the discovery happens.


The test: run one latency-injection experiment against your slowest synchronous dependency and record whether any caller fails before the dependency does.


What to change this week


Don't adopt a pattern. Add a timeout.


Find the synchronous calls in your main request path and check whether each has an explicit timeout shorter than the caller's own. Missing or too-long timeouts are the single most common cause of the cascade above, and they are a configuration change rather than an architecture.


Then add jitter to your retries. It's usually one parameter, it's the specific thing that stops recovery traffic arriving as a wall, and it costs nothing.


The company's pricing service still pauses for GC. It happened twice last quarter. Nothing downstream noticed, because booking now times out at 800ms, falls back to a cached rate band, and the retry budget caps the extra load at a tenth of normal traffic.


Final takeaway


Distributed outages are mostly healthy services amplifying a sick one. Timeouts shorter than the caller's, retries with jitter and a budget, async messaging for work nobody is waiting on, and sagas with real compensations each stop one specific amplification. This week, give every synchronous call in your main request path an explicit timeout, and add jitter to your retries.


Sources


  • Google, Site Reliability Engineering, ch. 21, "Handling Overload" (per-request and per-client retry budgets): https://sre.google/sre-book/handling-overload/

  • Marc Brooker, "Timeouts, retries, and backoff with jitter", Amazon Builders' Library: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/

  • Marc Brooker, "Exponential Backoff And Jitter", AWS Architecture Blog (2015): https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/

  • Michael T. Nygard, Release It!, 2nd ed. (Pragmatic Bookshelf, 2018): https://pragprog.com/titles/mnee2/release-it-second-edition/

  • Martin Fowler, "CircuitBreaker" (bliki, 2014): https://martinfowler.com/bliki/CircuitBreaker.html

  • Hector Garcia-Molina and Kenneth Salem, "Sagas", ACM SIGMOD 1987: https://www.cs.cornell.edu/andru/cs711/2002fa/reading/sagas.pdf

bottom of page