top of page

Data Modeling for Distributed Systems

Shawn West
Nov 25, 2025
6 min read

Updated: 2 days ago

Every rule you learned about modelling data assumes one database. Distribute the data and one of those rules inverts completely — and the teams that get burned are the ones who kept applying it.


A freight logistics company's support team asked a question that took four days to answer: which address are we actually shipping to?


(Composite example: a mid-sized freight logistics company, drawn from patterns across several organisations. Not a real company.)


The customer's delivery address existed in six services. Booking had one, captured at quote time. Operations had one, synced at some point from Booking. Billing had a different one, because billing address and delivery address had been the same field for two years. Support had a fourth, edited directly during a phone call.


None of them were wrong, exactly. Each was a reasonable local copy. What was missing was any statement about which copy other services were allowed to believe — and without that, six copies is not redundancy, it's six opinions.


The rule that inverts


In a single relational database, duplicated data is a modelling error. You normalise it away, because two copies can disagree and the database gives you joins for free, so there's no reason to keep both.


Distribute the data and that reverses. Duplication becomes a design pattern, because the alternative — calling another service at query time — is exactly the temporal coupling that made you split in the first place. Rendering an order list that fetches each customer's name over the network is a distributed join, and it fails in all the ways a local join doesn't.


So you duplicate. The discipline that replaces normalisation is ownership: exactly one service is the system of record for each piece of data, and every other copy is explicitly derived, explicitly stale, and explicitly not authoritative.



One database

Distributed

Duplication

A defect — normalise it away

Expected — copy what you read

Consistency

Guaranteed by the engine

A choice you make per field

The discipline

Normalisation

Ownership

What goes wrong

Update anomalies

Nobody knows which copy is true

The artefact

A schema

A list saying who owns what


The company had done the duplication and skipped the ownership. That's the common failure, and it's invisible until someone asks a question that requires an answer.


The test you can run: pick one field that appears in more than one service and ask who owns it. If two people give different answers, or the answer is "whichever was updated last", that's the four-day question waiting to be asked.


Ownership is a decision, and it's the first one


Before modelling anything, assign the system of record. One service, per piece of data, no exceptions.


What makes it decidable rather than political: the owner is whoever the data changes because of. The delivery address changes because a customer or an agent edits it — that's Support's domain, so Support owns it. The rate card changes because Pricing publishes a new one. Truck position changes because a GPS ping arrives at Operations.


Once assigned, three rules follow and they're worth stating explicitly because teams break them by accident:


  • Only the owner writes. Other services propose changes through the owner's interface. A second writer is a second owner, whatever the diagram says.

  • Copies are read-only and are allowed to be stale. Named staleness is fine — "addresses sync within a minute" is a contract. Unnamed staleness is the problem the company had.

  • The copy carries its provenance. Store when it was synced and from where. It costs two columns and it converts "which is right" from an investigation into a query.


The test: for each duplicated field, name the writer. If more than one service writes it, you don't have a copy — you have a fork.


Consistency is chosen per field, not per system


"Is the system eventually consistent?" is the wrong granularity, and it's why the conversation stalls. Different fields in the same record can and should have different answers, because the cost of being wrong differs enormously.


At the company:


Data

Consistency needed

Why

Mechanism

Capacity on a truck

Strong

Double-booking is a physical problem with a real cost

Owner is authoritative; no cached copies

Delivery address

Strong at dispatch, eventual before

Wrong at dispatch = wrong city; wrong on a screen = mild

Read-through at dispatch, cached elsewhere

Customer name

Eventual, minutes

A stale name on a screen is cosmetic

Event-driven copy

Lifetime freight spend

Eventual, hours

It's a reporting number

Batch


The useful question per field is not "how consistent can we be" but what does being wrong cost, and for how long. That question has an owner, a number, and an answer; "should this be eventually consistent" doesn't.


The test: take three fields you replicate and state the acceptable staleness for each in seconds. Any field where you can't name a number is one where nobody has decided, which means it's whatever the implementation happens to do.


Reads and writes want different shapes


Once data is distributed, read and write patterns often diverge so far that one model serves neither well. Writes need to enforce rules and be consistent; reads need to be fast, wide, and often span services.


Separating them — the write model owning correctness, a read model built for the queries you actually run — is what CQRS names. It's worth doing when the shapes genuinely differ, and it's overhead when they don't.


Where it earns its keep is the cross-service query, which is the hardest thing about distributed data and has no free answer. Three options, all with real costs:


  • Compose at the edge. A gateway calls three services and assembles the result. Simple, no new storage, but latency is the sum and availability is the product — the pattern where one slow dependency becomes everyone's problem.

  • Maintain a read model. Subscribe to events and keep a denormalised view built for the query. Fast, resilient to one service being down, and you now operate a projection that can drift and needs rebuilding.

  • Don't do the query. Frequently the right answer. "Show all orders for customers in this region with unpaid invoices" is three contexts in one sentence, and it's often a reporting question that belongs in a warehouse rather than a live path.


The test: for your worst cross-service query, count the services it touches. Three or more, on a user-facing path, means you're either building a read model or moving it off the path.


When none of this applies


Distributed data modelling is a response to a specific constraint, and applying it without the constraint is pure cost:


  • One service, one database. Normalise. Use transactions. Every rule above is a workaround for something you still have.

  • The data is genuinely local. Plenty of data has exactly one reader and one writer forever. It needs an owner and nothing else.

  • You haven't split yet. Modelling for distribution inside a monolith gives you the constraints without the benefits — no transactions, no joins, and no independent deployment either. Draw the boundaries first.

  • It's a reporting question. Analytics wants a copy of everything joined together, which is the opposite of what services want. That's a warehouse, and forcing operational services to serve it distorts both.


The test: if you're reaching for a saga or a read model inside a single service, stop — you have a transaction available and should use it.


The quality move: make ownership testable


An ownership list is documentation until something enforces it. Two cheap enforcements turn it into a guard: a test, or a permission at the data layer, that fails when a non-owning service writes an owned field; and a staleness check that alerts when a copy is older than its stated contract.


The discovery angle is the one most teams skip. "Who owns this data?" is an intake question, not an architecture-review question. Ask it when the feature is scoped, before anyone, human or assistant, writes the code that copies the field. Generated code is especially exposed here: the ticket said "show the customer's address", and the model will read whichever copy is nearest. What AI-generated code structurally skips is mostly the decisions nobody wrote down, and ownership is one of them.


The test: for one owned field, try writing it from a non-owning service in a test environment. If the write succeeds silently, the ownership rule exists only on paper.


What to change this week


Don't restructure anything. Write the ownership list.


One row per piece of data that exists in more than one service: what it is, which service owns it, who else holds a copy, and how stale that copy is allowed to be. It's a table, it takes an afternoon, and it is the artefact that distributed data modelling actually produces — the schema equivalent for a system that doesn't have one schema.


Expect two findings. Some data will have no owner, which is the company's address problem and the reason those questions take days. Some will have two writers, which is worse and usually older than anyone remembers.


The company's list has forty-one rows. Delivery address is owned by Support, read-through at dispatch and cached everywhere else with a sync timestamp on every copy. The four-day question is now a thirty-second one, and the answer is in the column that says who owns it.


Final takeaway


Distributed data modelling replaces normalisation with ownership: one writer per piece of data, copies that are read-only and openly stale, and consistency chosen field by field according to the cost of being wrong. This week, write the ownership list: one row per duplicated field, with its owner and its allowed staleness.


Sources


  • Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017): https://dataintensive.net/

  • Martin Fowler, "CQRS" (bliki, 2011): https://martinfowler.com/bliki/CQRS.html

  • Eric Evans, Domain-Driven Design: Tackling Complexity in the Heart of Software (Addison-Wesley, 2003). Bounded contexts and model ownership.

  • Chris Richardson, "Pattern: Saga", microservices.io: https://microservices.io/patterns/data/saga.html

bottom of page