Architecture / Microservices / Distributed Systems
Microservices Are an Ownership Decision
Microservices buy independent deployment and ownership at the price of a network, distributed data, and a platform to run it all. How to draw boundaries, keep data consistent without distributed transactions, and recognize a distributed monolith.
On this page
The case for microservices is usually made in technical terms: independent scaling, technology choice, fault isolation. Those benefits are real, but they are not the main reason the pattern exists. The main reason is organizational. Microservices let separate teams change, deploy, and operate separate parts of a system without coordinating every release.
Melvin Conway observed in 1968 that organizations design systems that mirror their communication structure. Microservices take that observation and use it deliberately: draw service boundaries where you want team boundaries, so that most changes stay inside one team's service.
If that is not a problem you have, because there is one team or a few teams that already coordinate easily, most of what microservices cost will not buy you much.
What you buy and what you pay
| You gain | You pay | | --- | --- | | Independent deploys per service | Network calls where there used to be function calls | | Clear ownership and on-call boundaries | No transactions across service boundaries | | Scaling hot paths independently | Versioned contracts between every pair of services | | Failure isolation, if designed for it | Distributed debugging, tracing, and log correlation | | Freedom to choose per-service technology | A platform: CI/CD, discovery, auth, observability per service |
The right column is not optional. A system that adopts the left column without paying for the right one tends to end up as a distributed monolith: all the operational cost of services, none of the independence.
Draw boundaries around data ownership
The most reliable boundary heuristic is ownership of data: each piece of data has exactly one service that writes it. Other services can hold copies, but the owner decides what the data means and when it changes.
That heuristic usually lines up with business capabilities, such as catalog, ordering, billing, and identity, rather than technical layers. A "database service" or a "validation service" is a layer, not a capability, and every feature will need to change it.
Signs that a boundary is in the wrong place:
- Most features require changes in two or more services, released together.
- Services make many fine-grained synchronous calls to each other to serve one request.
- Two services write to the same table, or one service reads another's tables directly.
- A service cannot answer a basic question about its own domain without calling another service.
Do not share the database
A shared database is the fastest way to recouple services that were split for independence. Once service B reads service A's tables, A can no longer change its schema without coordinating with B, and nothing in the code makes that dependency visible.
When a service needs another service's data, it has three honest options:
- Ask for it through the owner's API, at request time. Simple, always current, adds a runtime dependency.
- Keep a local copy built from the owner's events. Fast and available even when the owner is down, but eventually consistent, and the copy needs its own handling for missed or replayed events.
- Move the boundary. If two services cannot work without each other's data, they may be one service.
Consistency without distributed transactions
Inside one database, "create the order and reserve the stock" is one transaction. Across two services with two databases, it is not, and two-phase commit across independently deployed services is rarely practical.
The standard alternative is a saga: a sequence of local transactions, where each step publishes an event or sends a command that triggers the next, and each step that can fail later has a compensating action that semantically undoes it. The idea comes from a 1987 paper by Hector Garcia-Molina and Kenneth Salem on long-lived transactions.
| Step | Service | Action | Compensation if a later step fails |
| --- | --- | --- | --- |
| 1 | Orders | Create order as pending | Mark order cancelled |
| 2 | Inventory | Reserve items | Release reservation |
| 3 | Payments | Authorize card | Void authorization |
| 4 | Orders | Mark order confirmed | None, final step |
Two coordination styles exist:
- Orchestration. A coordinator, often the order service or a workflow engine, tells each participant what to do next and tracks progress. The flow is visible in one place and easy to change. The coordinator becomes an important dependency.
- Choreography. Each service reacts to events and emits new ones, with no central coordinator. Coupling is lower, but the overall flow exists only as the sum of subscriptions, which makes it harder to understand and to change once there are more than a few steps.
Compensations are not rollbacks. Other parts of the system may have observed the intermediate state, so a cancelled order may already have shown up on a dashboard, and an email may have been sent. Design intermediate states that are safe to expose, like pending, and make every step and compensation idempotent, because they will be retried.
Publish events reliably: the outbox
A service that updates its database and then publishes an event has a dual-write problem. If the process crashes between the two, the database says the order exists and no one else ever hears about it. Publishing first has the opposite failure.
The transactional outbox removes the gap. The service writes the event to an outbox table in the same transaction as the business change. A separate relay reads unsent rows and publishes them, either by polling or by tailing the database's change log.
BEGIN;
INSERT INTO orders (id, customer_id, status, total_cents)
VALUES ($1, $2, 'pending', $3);
INSERT INTO outbox (id, aggregate_id, type, payload, created_at)
VALUES ($4, $1, 'OrderPlaced', $5::jsonb, now());
COMMIT;The relay publishes at least once: it may crash after publishing and before marking a row sent. Consumers therefore deduplicate by event ID, which they should be doing anyway. The relay should also preserve order per aggregate, publishing a given order's events in the order they were written, if consumers depend on it.
Contracts between services
Every service API and event schema is a contract with other teams. Treat changes to it the way a library treats public API changes:
- Prefer additive changes, and keep old fields working until every consumer has migrated.
- Use consumer-driven contract tests, where each consumer publishes the parts of the API it relies on and the provider's build verifies them. Pact is a widely used tool for this.
- Deprecate with dates and telemetry. You cannot remove a field safely until you can see that nobody reads it.
- Version events when their meaning changes, rather than changing what an existing field means.
The platform tax
Each service needs a build and deploy pipeline, configuration, secrets, health checks, dashboards, alerts, log shipping, distributed tracing, service-to-service authentication, and a way for developers to run enough of the system locally to work on it. With three services, a team can do this by hand. With thirty, it needs a platform, meaning shared tooling and templates that make the right setup the default.
If that platform does not exist, the first services will each solve these problems differently, and later services will copy whichever one is closest. Budget for the platform as part of the decision to split, not as cleanup afterward.
When splitting is worth it
Splitting a service out tends to pay off when most of these hold:
- A distinct team will own it and be on call for it.
- Its data has a clear single owner and a stable interface to the rest of the system.
- It has a genuinely different scaling, availability, or security profile.
- Its release cadence is held back by, or holds back, the rest of the system.
- The organization already has, or is ready to build, the deploy and observability platform.
When few of these hold, a well-structured single deployable is usually the stronger architecture. That is the subject of The Modular Monolith Is a Real Architecture, the alternative most teams should evaluate first.
Further reading
- Melvin Conway, "How Do Committees Invent?" (1968): the origin of Conway's law.
- Hector Garcia-Molina and Kenneth Salem, "Sagas" (1987).
- Chris Richardson's patterns catalog at microservices.io: sagas, transactional outbox, database per service.
- Sam Newman, Building Microservices.