Architecture / System Design / Backend Engineering
A Working Method for System Design
A repeatable sequence for designing a system: requirements, rough numbers, interfaces and data, the hard part, failure analysis, and the tradeoffs you are choosing. Worked through on a webhook delivery service.
On this page
System design is often taught as a catalog of components: load balancers, caches, queues, sharded databases. Knowing the components helps, but it is not the skill. The skill is a sequence of questions that turns a vague requirement into a design whose tradeoffs you can name and defend.
This article describes that sequence and applies it to a concrete problem: a service that delivers webhooks, meaning HTTP callbacks to customer-provided URLs whenever events happen in a product. It is small enough to reason about fully and has most of the interesting problems: fan-out, retries, untrusted endpoints, ordering, and security.
1. Pin down what the system must do
Start with behavior, then constraints. Separate the two, because they drive different decisions.
Functional requirements for the webhook service:
- Customers register endpoints (URL plus the event types they want) and can disable them.
- When the product emits an event, every matching endpoint receives an HTTP
POSTwith the event payload. - Failed deliveries are retried for a bounded period.
- Customers can see delivery attempts and manually redeliver an event.
Non-functional requirements, stated as decisions rather than adjectives:
- Durability. An event accepted by the service is never silently dropped.
- Delivery semantics. At least once. Each event carries a unique ID so receivers can deduplicate.
- Latency. Most deliveries start within a few seconds of the event. This is a target, not a guarantee, because a customer's slow endpoint is outside our control.
- Isolation. One customer's failing or slow endpoint must not delay other customers' deliveries.
- Security. Payloads are signed. The service must not be usable to make requests into our own network.
Writing "highly available and scalable" adds nothing. Writing "a slow endpoint must not delay other customers" changes the design.
2. Estimate the load, roughly
Rough numbers decide whether you need one database or many, and whether something fits in memory. The goal is the order of magnitude, with the assumptions written down so they can be challenged.
These are illustrative assumptions for this exercise, not measurements:
| Quantity | Assumption | Derived | | --- | --- | --- | | Events per day | 50 million | ≈ 580 per second on average | | Peak-to-average ratio | 5× | ≈ 2,900 events per second at peak | | Matching endpoints per event | 3 on average | ≈ 8,700 delivery attempts per second at peak, before retries | | Payload size | 2 KB | ≈ 100 GB of event payloads per day | | Retention of delivery history | 30 days | ≈ 3 TB of payloads, plus attempt metadata |
Two things stand out. Write throughput is significant but well within what a single well-tuned relational database can handle with batching. Storage over the retention window is large enough that payloads may belong in cheaper storage, with only metadata in the primary database. If the numbers were a hundred times larger, the design would change. That is why you estimate first.
3. Define the interfaces and the data
Before boxes and arrows, decide what the system exposes and what it stores. Most of the design follows from the data model.
CREATE TABLE endpoints (
id uuid PRIMARY KEY,
customer_id uuid NOT NULL,
url text NOT NULL,
event_types text[] NOT NULL,
secret bytea NOT NULL, -- for HMAC signatures
status text NOT NULL DEFAULT 'active', -- active | disabled | failing
created_at timestamptz NOT NULL DEFAULT now()
);
CREATE TABLE events (
id uuid PRIMARY KEY,
customer_id uuid NOT NULL,
type text NOT NULL,
payload_ref text NOT NULL, -- object storage key
created_at timestamptz NOT NULL DEFAULT now()
);
CREATE TABLE deliveries (
id uuid PRIMARY KEY,
event_id uuid NOT NULL REFERENCES events(id),
endpoint_id uuid NOT NULL REFERENCES endpoints(id),
status text NOT NULL, -- pending | succeeded | failed | dead
attempt_count int NOT NULL DEFAULT 0,
next_attempt_at timestamptz NOT NULL DEFAULT now(),
last_error text,
UNIQUE (event_id, endpoint_id)
);
CREATE INDEX deliveries_due ON deliveries (next_attempt_at)
WHERE status = 'pending';The UNIQUE (event_id, endpoint_id) constraint makes fan-out idempotent: if intake is retried, it cannot create duplicate deliveries. The partial index keeps the "what is due now" query cheap even as history grows.
The interfaces are small: an internal POST /events for product services, and a customer-facing API to manage endpoints, list deliveries, and trigger redelivery.
4. Find the hard part
Every design has one or two problems that dominate it. Spend your time there. For webhooks, the hard part is not throughput. It is untrusted, unreliable receivers:
- Endpoints can be slow, time out, return errors for hours, or disappear permanently.
- A single large customer can have one endpoint receiving a large share of all events.
- The URL is supplied by a customer, so it could point at internal infrastructure, such as a cloud metadata address or a private service. That is a server-side request forgery (SSRF) risk.
Each of these needs an explicit mechanism:
- Per-endpoint concurrency limits and timeouts, so a slow endpoint occupies a bounded number of workers.
- Backoff per delivery and circuit breaking per endpoint. After sustained failures, mark the endpoint
failing, slow its retries sharply, and notify the customer. - Egress controls. Resolve the hostname, reject private, loopback, and link-local addresses, re-check after redirects or disable redirects, and send traffic through an egress proxy that enforces the same rules.
5. Sketch the architecture
With the data model and the hard part understood, the components almost choose themselves:
- Intake validates the event, stores the payload, and inserts the event plus one delivery row per matching endpoint, all in one transaction.
- A scheduler repeatedly claims due deliveries and dispatches them.
- Workers send signed requests, enforce per-endpoint concurrency, and record the outcome of every attempt.
- Failures get a new
next_attempt_atwith exponential backoff and jitter. After the final attempt, the delivery becomesdeadand is visible for manual redelivery.
Claiming due work from Postgres without double-dispatch is a well-known pattern using row locks that skip already-claimed rows:
WITH due AS (
SELECT id
FROM deliveries
WHERE status = 'pending' AND next_attempt_at <= now()
ORDER BY next_attempt_at
LIMIT 500
FOR UPDATE SKIP LOCKED
)
UPDATE deliveries d
SET next_attempt_at = now() + interval '5 minutes' -- lease: retried if the worker dies
FROM due
WHERE d.id = due.id
RETURNING d.id, d.event_id, d.endpoint_id, d.attempt_count;The lease trick matters. Instead of a separate "in progress" status that can get stuck, the claim pushes next_attempt_at forward. If the worker crashes, the delivery simply becomes due again after the lease expires.
6. Walk through the failures
A design is not done until you have traced what happens when each component fails. A table keeps it honest:
| Failure | Behavior | Why it is acceptable |
| --- | --- | --- |
| Intake crashes mid-request | Transaction rolls back; producer retries with the same event ID | Unique event ID prevents duplicates |
| Worker crashes after sending, before recording | Lease expires; delivery is sent again | At-least-once is the contract; receivers dedupe by event ID |
| Endpoint down for hours | Backoff grows; endpoint marked failing; customer notified | Bounded resource use; no impact on others |
| Database unavailable | Intake rejects new events; producers buffer and retry | Durability preserved; no event accepted that cannot be stored |
| Queue unavailable | Scheduler cannot dispatch; deliveries remain pending | The table is the source of truth; nothing is lost |
The last row is a design choice worth noting. Because the database, not the queue, is authoritative, a queue outage delays delivery but cannot lose it.
7. Decide what you will measure
Choose metrics that map to the requirements, not just to the components:
- Time from event to first attempt, as a distribution, against the latency target.
- Age of the oldest due delivery, which is the clearest signal of falling behind.
- Success rate per endpoint, which drives circuit breaking and customer notifications.
- Dead deliveries per day, and how many are manually redelivered.
8. Name the tradeoffs
The last step is to say what you chose and what would make you change it:
- Postgres as the work queue keeps the system simple and consistent. If claim contention or table bloat becomes the bottleneck at higher volumes, move dispatch to a dedicated broker and keep the table as the record of attempts.
- No ordering guarantee across events. Ordered delivery per endpoint would force sequential delivery and let one failing event block all later ones. Including a timestamp and sequence number lets receivers order events themselves if they need to.
- Payloads in object storage trade a second read per delivery for much cheaper retention.
That sequence is the method: behavior, numbers, data, the hard part, the shape, the failures, the measures, the tradeoffs. Components change with every problem. The questions do not.
For the decisions in this process that are hardest to reverse later, see System Architecture Is the Set of Decisions That Are Expensive to Change.