AI Engineering / AI Agents / Agentic Systems / Orchestration
Graph Engineering: Modeling Agent Workflows as Explicit Graphs
Once an agent workflow has branches, parallel steps, approvals, and revision cycles, it is a graph whether or not you draw it. Typed state, nodes, conditional edges, reducers, checkpoints, and how to test and version the result.
On this page
A single agent loop, where the model decides, a tool runs, and the result feeds back, covers a lot of ground. It stops being enough when the workflow develops structure: classify the request first, run two retrievals in parallel, pause for a human to answer a question, review the draft and send it back for revision at most twice.
You can encode that structure inside one long prompt and hope the model follows it. Or you can encode it in application code, where it can be seen, tested, bounded, and resumed. Graph engineering is the second option: modeling the workflow as an explicit graph of steps, with typed state flowing between them.
Three primitives: state, nodes, edges
State is a typed object that represents everything the workflow knows so far. It is the only thing passed between steps.
type ResearchState = {
question: string;
intent?: 'research' | 'clarify';
findings: Finding[];
draft?: string;
review?: { passed: boolean; issues: string[] };
revisions: number;
};Nodes are functions from state to a partial state update. A node might call a model, call a tool, run deterministic code, or wait for a human. A node does not decide where the workflow goes next.
Edges connect nodes. A static edge always goes to the same next node. A conditional edge runs a small router function over the state and returns the name of the next node.
Keeping routing out of nodes is the important separation. A node does one job and reports what it found. The graph decides what happens next, and that decision is visible in one place.
A minimal executor
You do not need a framework to get the benefits of the model. The core executor is small:
type NodeFn<S> = (state: S, ctx: RunContext) => Promise<Partial<S>>;
type Router<S> = (state: S) => string;
type Graph<S> = {
nodes: Record<string, NodeFn<S>>;
edges: Record<string, string | Router<S>>; // node -> next node, or a router
reducers: { [K in keyof S]?: (current: S[K], update: S[K]) => S[K] };
start: string;
};
const END = '__end__';
async function run<S>(graph: Graph<S>, initial: S, ctx: RunContext) {
let state = initial;
let current = graph.start;
for (let step = 0; current !== END; step += 1) {
if (step >= ctx.maxSteps) throw new StepLimitExceeded(current, step);
ctx.signal.throwIfAborted();
const update = await graph.nodes[current](state, ctx);
state = applyUpdate(state, update, graph.reducers);
const edge = graph.edges[current];
const next = typeof edge === 'function' ? edge(state) : edge;
await ctx.checkpoint({ node: current, next, state, step });
ctx.trace.record({ node: current, next, step });
current = next;
}
return state;
}
function applyUpdate<S>(state: S, update: Partial<S>, reducers: Graph<S>['reducers']): S {
const next = { ...state };
for (const key of Object.keys(update) as (keyof S)[]) {
const reduce = reducers[key];
const value = update[key] as S[keyof S];
next[key] = reduce ? reduce(state[key], value) : value;
}
return next;
}This version runs nodes sequentially. Parallel branches add a fan-out step that runs several nodes concurrently and applies their updates through the same reducers. The rest of the model stays the same.
The routers for the workflow above are ordinary functions:
const routeAfterClassify = (s: ResearchState) =>
s.intent === 'clarify' ? 'clarify' : 'gather'; // 'gather' fans out to search + docs
const routeAfterReview = (s: ResearchState) => {
if (s.review?.passed) return END;
if (s.revisions >= 2) return 'escalate'; // bounded cycle
return 'draft';
};Merging parallel branches
When two branches run at the same time and both update state, the merge must be deterministic. Reducers declare, per state key, how concurrent updates combine:
- Lists of findings or messages usually append, with deduplication by source ID.
- Counters add.
- Scalars need a rule: one designated owner branch, or "conflicting writes are an error". Last-write-wins on a scalar written by two concurrent branches makes the result depend on timing.
const reducers: Graph<ResearchState>['reducers'] = {
findings: (current, update) => dedupeBy([...current, ...update], f => f.sourceId),
revisions: (current, update) => current + update,
};If you find yourself needing complex merge logic, the branches probably are not independent and should run sequentially.
Cycles need bounds in the state
Some cycles are the point of the workflow: review sends the draft back for revision, a verifier rejects tool output and the step is retried. Every cycle needs a counter that lives in state, such as revisions above, and a router that checks it. A graph-wide step limit is the backstop, not the primary control. When a bound is hit, route to an explicit node like escalate or finish_with_caveats, not to a generic exception.
Checkpoint after every node
Because the state is explicit and every transition goes through the executor, persisting the state after each node is cheap. It pays for itself several times over:
- Resumption. A crashed or redeployed worker resumes from the last checkpoint instead of rerunning expensive model calls.
- Human-in-the-loop. A
clarifyorapprovenode saves state and ends the current execution. When the human responds, the run loads the checkpoint, adds the answer to state, and continues. Waiting a day costs nothing. - Inspection and replay. The sequence of checkpoints is a complete history of the run. You can see the state that led to a bad decision, and replay from that point with a changed prompt or router.
Checkpointing makes one property essential: nodes with side effects must be idempotent, because a node may run again after a crash between its side effect and the checkpoint. Pass a key derived from (runId, node, step) to any external write.
Make routing inspectable
Routing is where workflows most often go wrong silently. Two practices help:
- Prefer deterministic routing where the decision can be made from state: "has findings", "revisions < 2", "user is on a plan that allows web search". Reserve model-based routing for genuinely semantic decisions.
- When a model routes, constrain and record it. Ask for a value from a fixed enum with structured output, validate it, fall back to a safe default on anything else, and log the chosen route with the model's stated reason. A router that returns free text will eventually return a node name that does not exist.
Testing a graph
An explicit graph decomposes cleanly into testable parts:
- Nodes are functions of state. Test them with fixed inputs, using a fake model client where they call one.
- Routers are pure functions. Table-driven tests cover every branch, including the bounds.
- Whole runs can be tested with recorded model responses: given this question and these canned outputs, assert the sequence of visited nodes and the final state. These tests catch routing regressions without depending on a live model.
- Evaluation runs with a live model measure quality on a fixed set of inputs, and the node trace tells you which step failed when quality drops.
Versioning graphs with runs in flight
A long-running workflow may be checkpointed under one version of the graph and resumed under another. Store the graph version with each checkpoint. Make changes additive where possible, and when node semantics or state shape change incompatibly, either let old runs finish on the old version or write an explicit migration for their state. Silently resuming an old state in a new graph produces failures that are very hard to diagnose.
When not to build a graph
A graph adds structure, and structure has a cost. If the task is one model working with a handful of tools toward one goal, the bounded loop described in An Agent Loop Is a State Machine, Not a Prompt is simpler and often better: the model is good at choosing the next tool, and a graph would only restate that choice in code.
Reach for a graph when the workflow has structure you need to guarantee rather than suggest: mandatory steps, parallel work, approvals, bounded revision cycles, or runs that must survive restarts. Frameworks such as LangGraph implement this state-node-edge model with checkpointing built in. Whether you use one or a fifty-line executor, the engineering decisions are the same: what is in state, who decides each transition, how branches merge, and what bounds every cycle.
The loops that run inside and around these graphs, and how to keep them converging, are the subject of Loop Engineering.