AI Engineering / AI Agents / Agentic Systems / Reliability / Evaluation
Loop Engineering: Designing the Loops Around a Model
Agents are nested loops running at different time scales. Each needs a termination contract, a progress signal, a verifier outside the model, and a deliberate way to give up. Budgets, no-progress detection, feedback, and the offline improvement loop.
On this page
An earlier article argued that an agent loop is a state machine: explicit states, owned transitions, and defined stop reasons. That is the control structure of a single run. This article is about the engineering of loops as a design element: what makes an iterative process converge, how to tell it has stopped making progress, and how loops at different time scales fit together.
The reason this deserves its own discipline is simple. Almost every useful agent behavior is iteration, such as trying, checking, and adjusting, and iteration without engineering has two failure modes: it stops too early, or it never stops.
Agents are nested loops
A capable agent usually runs several loops, nested inside each other, at different speeds:
- The step loop runs in seconds: decide, call a tool, observe the result.
- The task loop runs in minutes: produce a candidate, check it against something outside the model, revise or escalate.
- The human loop runs on human time: ask for clarification or approval, wait, resume.
- The improvement loop runs over days: collect failed runs, turn them into evaluation cases, change prompts, tools, or the harness, and verify the change against the suite.
Each loop has the same four design questions: when is it done, when does it give up, how does it know it is progressing, and what does it carry into the next iteration? Problems usually come from answering these for the inner loop and forgetting the outer ones.
Every loop needs a termination contract
"Stop when the model says it is finished" is not a contract. A termination contract has three parts:
- A definition of done that is checked, not asserted: tests pass, the output validates against a schema, the checklist is satisfied.
- Budgets across several dimensions, because loops fail in different ways: steps, tokens, wall-clock time, and money.
- A defined result for each exit, so the caller can tell "done" from "gave up after the budget" from "cancelled".
type Budget = { steps: number; tokens: number; ms: number; usd: number };
class LoopBudget {
private used = { steps: 0, tokens: 0, ms: 0, usd: 0 };
private readonly startedAt = Date.now();
constructor(private readonly limit: Budget) {}
charge(usage: { tokens: number; usd: number }) {
this.used.steps += 1;
this.used.tokens += usage.tokens;
this.used.usd += usage.usd;
this.used.ms = Date.now() - this.startedAt;
}
/** The first exhausted dimension, or null if the loop may continue. */
exhausted(): keyof Budget | null {
for (const key of Object.keys(this.limit) as (keyof Budget)[]) {
if (this.used[key] >= this.limit[key]) return key;
}
return null;
}
}Budgets are product decisions. The step limit for a code-fixing agent working in the background and the time limit for an interactive assistant are different numbers for different reasons, and the people who own the product should be able to see and change them.
Detect when the loop has stopped making progress
Budgets end runaway loops eventually. Progress detection ends them early, before they waste the budget. Common stuck patterns:
- Repetition: the same tool called with the same arguments, getting the same result.
- Oscillation: alternating between two states, such as applying a fix and reverting it.
- Stagnation: many steps without any change to the state that matters, like the set of failing tests or the fields still missing.
A cheap, general detector fingerprints the parts of state that represent progress and watches a window of recent iterations:
function progressKey(state: TaskState): string {
// Only what represents progress, not timestamps or transcript text.
return hash({
failingTests: [...state.failingTests].sort(),
filesChanged: [...state.filesChanged].sort(),
lastAction: state.lastAction && {
tool: state.lastAction.tool,
args: state.lastAction.args,
},
});
}
class ProgressMonitor {
private recent: string[] = [];
constructor(private readonly window = 4) {}
/** Returns a reason when the loop looks stuck, otherwise null. */
observe(state: TaskState): 'repeating' | 'oscillating' | null {
const key = progressKey(state);
this.recent = [...this.recent, key].slice(-this.window);
if (this.recent.length === this.window && new Set(this.recent).size === 1) {
return 'repeating';
}
const [a, b, c, d] = this.recent;
if (this.recent.length === 4 && a === c && b === d && a !== b) {
return 'oscillating';
}
return null;
}
}What counts as progress is domain-specific, and choosing it well is most of the work. When the monitor fires, do not just stop. Change something: inject the observation into the context ("the last three attempts produced the same failing test"), switch strategy, or escalate.
Verify with something outside the model
The quality of the task loop is bounded by the quality of its verification signal. Asking the same model whether its output is correct is the weakest signal available. It shares the blind spots that produced the output. Prefer checks that can disagree with the model:
| Verifier | Signal quality | Examples | | --- | --- | --- | | Execution | High | Run the tests, compile the code, execute the query against a fixture | | Validation | High, narrow | Schema validation, type checking, linting, policy rules | | Grounding | Medium | Check that cited facts appear in the retrieved sources | | A separate grader | Medium; needs calibration | A different prompt or model with a rubric, checked against human labels | | Self-review | Low | The same model asked to critique its own output |
Where a strong verifier exists, as it does for code, structured data, and anything with an executable spec, loops can iterate toward correctness. Where it does not, more iterations mostly produce different outputs, not better ones, and the loop should be short with a human checkpoint.
Feed back the error, not just "try again"
A retry is only useful if the next attempt has information the previous one lacked. Verifier output should be shaped into specific, actionable feedback:
- Include the failing test name, the assertion, and the relevant lines of output, not the full log.
- Say what was attempted and what happened, in a sentence: "Added a null check in
parseOrder;test_empty_cartstill fails with the same error." - Cap the size of feedback, and keep the original instructions in the context. Feedback should add to the task, not replace it.
Carrying every failed attempt forward makes the context grow with each iteration and buries the instructions. A practical policy is to keep the latest attempt and its feedback in full, plus a short summary of earlier attempts ("tried X and Y; both failed because Z"). This also helps the model avoid repeating approaches that already failed.
Escalate on purpose
Giving up is an outcome to design, not a failure to hide. A good escalation:
- happens for a reason the loop can name: budget exhausted, no progress, verifier unavailable, or a decision outside the agent's authority;
- hands over the current best candidate, what was tried, and why it stopped;
- asks a specific question when a human can unblock it ("Should the migration drop the unused column, or keep it?").
An agent that escalates clearly at the right moment is more useful than one that keeps going with low confidence.
The improvement loop
The slowest loop is the one that makes all the others better over time:
- Collect traces of runs, especially the ones that hit budgets, escalated, or were corrected by users.
- Categorize failures: wrong tool, bad arguments, missed instruction, verifier gap, insufficient context.
- Convert representative failures into evaluation cases with a checkable expected outcome.
- Change one thing, such as a prompt, a tool description, a verifier, or a budget, and run the full suite, comparing with the baseline.
- Ship only when the target category improves without regressions elsewhere.
This is the same loop as ordinary software regression testing, applied to behavior that is probabilistic. The discipline that matters most is step 4: change one thing at a time, and measure against a fixed suite, not against a handful of examples you remember.
Measuring loops
Per-run success rate is the headline, but loop-specific metrics explain it:
- Iterations to success, as a distribution. A long tail means progress detection or feedback needs work.
- Budget exhaustion rate, by dimension. Which budget ends runs tells you what to optimize.
- First-attempt verifier pass rate, a proxy for how much the loop is compensating for weak first drafts.
- Escalation rate and resolution, meaning how often humans are asked and whether their answers led to success.
- Cost per successful task, not per run, because cheap failed runs are not cheap.
Loops are where agent systems spend their time and money. Engineering them deliberately, with contracts, progress signals, external verification, and honest exits, is most of what separates an agent that demos well from one that can be left running.
The software that hosts these loops and decides what they may do is covered in Harness Engineering.