Distributed Systems / Messaging / API Design / Backend Engineering
How Systems Communicate: Calls, Messages, and Streams
Request/response, queues, logs, webhooks, and streaming connections each trade a different kind of coupling. A practical guide to choosing between them and to the delivery guarantees each one actually gives you.
On this page
- Three kinds of coupling
- Request and response
- HTTP/JSON or gRPC
- Messages: commands and events
- Queues and logs
- Delivery guarantees
- Ordering is per key, not global
- Poison messages and dead letters
- Webhooks and polling across organizations
- Long-lived connections and streaming
- Contracts outlive implementations
- Choosing a pattern
Every integration between two systems answers three questions, whether or not anyone asked them explicitly:
- Who waits? Does the sender block until the receiver has done the work?
- Who knows about whom? Does the sender need the receiver's address, API, and existence?
- What happens when the other side is down? Does the work fail, wait, or get lost?
The communication patterns below are different answers to those questions. None of them is better in general. Each removes one kind of coupling and adds a cost somewhere else.
Three kinds of coupling
It helps to separate coupling into three dimensions, because patterns trade them independently.
| Coupling | Meaning | Reduced by | | --- | --- | --- | | Temporal | Both sides must be running at the same time | Queues, logs, durable buffers | | Knowledge | The sender must know who receives and how to reach them | Publish/subscribe, service discovery | | Format | Both sides must agree on the shape of the data | Schemas, versioning, tolerant readers |
Format coupling never goes away. A message broker removes temporal and most knowledge coupling, but consumers still depend on the event's fields. Most integration breakage in practice is format breakage: a renamed field, a changed enum, a timestamp that switched time zones.
Request and response
A synchronous call is the default for good reasons. The caller gets an answer or an error immediately, errors propagate naturally, and the control flow reads top to bottom. When the caller genuinely needs the result to continue, such as a price before showing a checkout page or an authorization decision before serving a file, request/response is the right tool.
Its costs show up when calls form chains:
- Latency adds. A request that crosses four services pays four network round trips plus four services' queuing and processing time, and the slowest dependency dominates the tail.
- Availability multiplies. If each service in a chain of four is independently available 99.9% of the time, the chain is available about 0.9994 ≈ 99.6% of the time before counting the network. Real failures are often correlated, which can make this better or worse, but the direction is clear: every synchronous dependency is part of your availability budget.
- Load couples. A traffic spike at the edge arrives at every service in the chain at once.
HTTP/JSON or gRPC
Both are request/response. The practical differences:
- HTTP with JSON is universal, easy to inspect with ordinary tools, and native to browsers. Schemas are optional, which is convenient until a field changes shape without anyone noticing.
- gRPC uses HTTP/2 and Protocol Buffers. You get generated clients, a required schema, efficient binary encoding, deadlines as a first-class concept, and bidirectional streaming. Browsers cannot speak it directly without a proxy layer, and payloads are not human-readable on the wire.
A common split is HTTP/JSON at the public edge and gRPC between internal services, but the more important decision is whether the schema is explicit and versioned, whichever transport carries it.
Messages: commands and events
Asynchronous messaging inserts a durable intermediary between sender and receiver. The sender's job ends when the broker has accepted the message.
Two kinds of message are worth distinguishing, because they imply different ownership:
- A command asks a specific receiver to do something:
SendInvoice,ResizeImage. The sender knows who should act and usually cares that it happens. - An event states that something happened:
OrderPlaced,UserDeactivated. The publisher does not know or care who reacts. Adding a new consumer requires no change to the publisher.
Events are how you reduce knowledge coupling. The cost is that the end-to-end behavior of the system is no longer visible in any single codebase. "What happens when an order is placed?" becomes a question you answer by finding every subscriber.
Queues and logs
Brokers come in two broad shapes:
- A queue distributes work. Each message is delivered to one consumer and removed once acknowledged. It suits background jobs and commands.
- A log (or stream) retains messages in order for a configured period. Each consumer group tracks its own position, so several independent consumers can read the same messages, and a consumer can rewind and replay. It suits events with multiple subscribers, and rebuilding derived data.
Delivery guarantees
Every broker offers some version of these guarantees, and the names are less important than what they make your consumer responsible for.
- At most once. A message may be lost but is never delivered twice. Acceptable for metrics samples and other data where gaps are tolerable.
- At least once. A message is never lost once accepted, but may be delivered more than once, for example when a consumer processes a message and crashes before acknowledging it. This is the practical default.
- Exactly once. Usually means effectively once: duplicates may be delivered, but the observable effect happens once, because processing is idempotent or because the broker and the consumer's state are updated in one transaction.
The reliable way to get effectively-once behavior from at-least-once delivery is to make the consumer idempotent, and the simplest idempotent consumer records processed message IDs in the same transaction as its own state change:
async function handleOrderPlaced(message: Message<OrderPlaced>) {
await db.transaction(async tx => {
// Unique constraint on (consumer, message_id).
const inserted = await tx.processedMessages.insertIfAbsent({
consumer: 'inventory',
messageId: message.id,
});
if (!inserted) return; // already handled; acknowledge and move on
await tx.reservations.create({
orderId: message.body.orderId,
items: message.body.items,
});
});
await message.ack();
}If the process crashes after the commit and before ack, the broker redelivers. The second attempt finds the message ID and returns without reserving the items twice.
Ordering is per key, not global
Global ordering across all messages requires every message to pass through a single sequence point, which limits throughput to what that single point can handle. Most systems do not need it. They need ordering per entity: all events for order 812 processed in the order they happened.
Logs provide this by partitioning. The producer chooses a key, the key determines the partition, and each partition is consumed in order by exactly one member of a consumer group.
Two consequences follow:
- Key choice is a design decision. Keying by
order_idgives per-order ordering. Keying bycustomer_idgives per-customer ordering, with coarser partitions. A key with skewed traffic, such as one very large tenant, creates a hot partition that one consumer must handle alone. - Consumers should still tolerate staleness. Retries, replays, and multi-step flows can deliver an older state after a newer one. Including a version or sequence number in each event lets the consumer ignore anything older than what it has already applied.
Poison messages and dead letters
A message that fails every time, because of malformed data or a bug that only this input triggers, will block an ordered partition or loop forever in a queue. Bound the retries. After a fixed number of attempts, move the message to a dead-letter queue together with the error, the attempt count, and the consumer version. A dead-letter queue nobody looks at is just a slower way to lose data, so give it an alert and a tool to redrive messages after a fix.
Webhooks and polling across organizations
When the receiver is another company's system, you lose control of the infrastructure but the same rules apply.
Webhooks are at-least-once HTTP push. Senders should retry with backoff, sign each payload, typically with an HMAC over the body and a timestamp, and include a unique event ID. Receivers should verify the signature, reject stale timestamps to block replays, deduplicate by event ID, and acknowledge quickly: put the work on a local queue and return 2xx, rather than doing slow processing inside the request.
Polling with a cursor (GET /events?after=<cursor>) is the dependable complement. It lets a receiver catch up after downtime, backfill, or reconcile against missed webhooks. Offering both, webhooks for latency and a cursor-based feed for correctness, is a good default for any public event API.
Long-lived connections and streaming
Some interactions are not a single response:
- Server-Sent Events (SSE) stream text events from server to client over ordinary HTTP. They are simple, pass through most proxies, and reconnect automatically with a
Last-Event-IDheader. Token streaming from language-model APIs is typically delivered this way. - WebSockets provide a bidirectional channel. They suit collaborative editing, multiplayer state, and chat where both sides send frequently. You take on connection management, heartbeats, and reconnection with state resynchronization.
- gRPC streams offer the same capabilities between services, with schemas.
For any stream, decide what a client does after a disconnect: resume from a position, or refetch state and start again. "Reconnect and hope" loses data.
Contracts outlive implementations
Whatever the transport, the message format is a contract between teams and between versions of the same service running side by side during a deploy. A few rules keep it from becoming the source of outages:
- Make changes additive. Add optional fields; do not rename or repurpose existing ones.
- Readers ignore what they do not understand. A consumer that rejects unknown fields turns every producer change into a coordinated release.
- Never reuse identifiers. In Protocol Buffers, a removed field's number stays reserved forever.
- Version event types when meaning changes.
OrderPlaced.v2alongsidev1is clearer than av1whose fields quietly mean something new. - Validate at the boundary. Parse incoming messages against a schema and route invalid ones to a dead-letter queue instead of failing deep inside business logic.
Choosing a pattern
| If you need… | Prefer | Watch out for | | --- | --- | --- | | An answer before continuing | Request/response | Long call chains, missing timeouts | | Background work done reliably, once | Queue + idempotent consumer | Unbounded retries, poison messages | | Several teams reacting to the same fact | Events on a log | Invisible end-to-end flows, schema drift | | Per-entity ordering at scale | Partitioned log keyed by entity | Hot keys, stale events on replay | | Integration with external systems | Signed webhooks + cursor feed | Replays, slow receivers, missed deliveries | | Incremental results to a client | SSE or WebSockets | Reconnection and resume semantics |
The useful habit is not picking one pattern for the whole system. It is asking the three questions (who waits, who knows, what happens when it is down) at each integration point, and choosing the coupling you can afford there.
The failure handling that every one of these patterns depends on is covered in Distributed Systems Are About Partial Failure.