AI Engineering / AI Agents / Agentic Systems / Reliability
An Agent Loop Is a State Machine, Not a Prompt
A practical model for bounded agent execution: make decisions, tool calls, state transitions, and stop conditions explicit in application code.
On this page
An agent is often introduced as a prompt that can call tools. That description is useful for a demo, but it hides the part that determines whether the feature behaves like software: the control loop around the model.
The model proposes a next action. The application decides whether that action is valid, executes it, records the result, and determines what happens next. That is a state machine with an external decision-maker, not an unconstrained conversation.
Make the states visible
A small agent can be described with a few states:
- Ready: collect the user request and the allowed context.
- Deciding: ask the model for a final response or a tool call.
- Executing: validate and run one approved tool.
- Complete: return a final response.
- Stopped: end because of a limit, cancellation, or failure.
Each transition should have an owner. The model can request a tool, but only application code should authorize and execute it. A tool result should become explicit input to the next decision rather than disappearing into an opaque transcript.
Bound every path
At minimum, define a maximum number of model turns, a maximum number of tool executions, a request deadline, and a cancellation path. These limits are not tuning details. They describe the contract of the feature under slow, invalid, or looping behavior.
An intentionally small loop might look like this:
type TurnResult =
| { kind: 'final'; text: string }
| { kind: 'tool'; name: string; arguments: unknown };
async function runAgent(request: Request, signal: AbortSignal) {
const history = createInitialHistory(request);
for (let step = 0; step < MAX_STEPS; step += 1) {
signal.throwIfAborted();
const decision: TurnResult = await model.next(history, { signal });
if (decision.kind === 'final') {
return { state: 'complete' as const, text: decision.text };
}
const tool = tools.get(decision.name);
if (!tool) {
history.push(toolUnavailable(decision.name));
continue;
}
const input = tool.inputSchema.parse(decision.arguments);
const output = await tool.execute(input, { signal });
history.push(toolResult(tool.name, output));
}
return { state: 'stopped' as const, reason: 'step_limit' };
}The example leaves out persistence and transport details, but the boundaries are visible: the tool registry is an allowlist, arguments are parsed before execution, cancellation reaches model and tool calls, and a step limit has a defined result.
Treat tool output as untrusted input
Tool results can be large, malformed, stale, or contain text that resembles instructions. Normalize outputs to the fields the next step needs, cap their size, and keep provenance attached. Do not concatenate arbitrary output into a privileged instruction and assume the model will preserve the original trust boundary.
The same applies to model-produced arguments. A JSON-shaped response is not validated input. Parse it against a schema, enforce authorization with the requesting user’s identity, and reject operations the current request is not allowed to perform.
Decide what each stop means
“The model stopped” is not a product state. Distinguish completion from a step limit, timeout, cancellation, provider failure, invalid tool arguments, and tool failure. Some failures are retryable; others need a clear response or a human decision. Preserve that distinction in logs and in the result returned to the caller.
Retries need their own bounds. Retrying a read is different from retrying a side effect. For writes, use idempotency keys or another operation-specific duplicate-prevention strategy; a generic retry wrapper can turn a transient network failure into repeated work.
Keep the loop observable
Record a trace with a request identifier, model decision type, selected tool name, validation outcome, duration, and final stop reason. Avoid logging secrets or entire private prompts by default. A useful trace answers where time went and why the run ended without requiring someone to reconstruct hidden state from a text transcript.
Once the loop is explicit, the prompt becomes one input to the system rather than the system itself. That makes evaluation more precise: tests can assert which tools are available, how invalid calls are handled, whether limits are respected, and what result each terminal state produces.
The prompt still matters. But dependable behavior comes from the prompt working inside a loop whose permissions, transitions, and stopping rules are ordinary application code.