AI Engineering
Who Decides the Next Step: The Anatomy of an AI Agent
An agent is defined by where control flow lives, not by how capable the model is. State and context, planning, observation, memory, verification, recovery, and permissions, and why reliability comes from the software around the model.
- Author
- Akshay Shinde
- Published
- October 8, 2026
- Read
- 16 min read
- Topics
- AI Agents, Agentic Systems, Reliability, Evaluation
AI Engineering / AI Agents / Agentic Systems / Reliability / Evaluation
Who Decides the Next Step: The Anatomy of an AI Agent
An agent is defined by where control flow lives, not by how capable the model is. State and context, planning, observation, memory, verification, recovery, and permissions, and why reliability comes from the software around the model.
On this page
- State is not context
- Plans are hypotheses
- Observations are an interface you design
- Memory is a write problem
- Verification and recovery
- Deciding what the model should decide
- Why a smarter model does not make a reliable agent
- More agents, more surfaces
- Permissions belong to the action
- Software with a nondeterministic component
Ask ten engineers what an AI agent is and you will get answers about intelligence: it reasons, it plans, it figures things out. Those answers describe the model. They say very little about the system, and the system is what you have to build, test, and operate.
A more useful definition is mechanical. A system becomes an agent when the model, rather than your code, decides what happens next, and keeps deciding until it judges the work finished. Everything interesting about agents, including most of what makes them hard, follows from that one transfer of control.
Take a concrete task: checkout errors spiked after last night's deploy; find out why and propose a fix. It can be built at least four ways, and the differences between them are about who holds control.
A single model call. Paste the error summary and the deploy diff into a prompt and ask what went wrong. For a narrow question it can be enough, and when it fails it fails cheaply: one call, one answer, nothing touched. Its limit is that it only knows what you thought to include.
A workflow. Code fetches the error logs, a model classifies the failure, code pulls the relevant diff, a model summarizes, code posts the summary. Each step is fixed. Anthropic's Building effective agents draws the line here: workflows orchestrate models and tools "through predefined code paths". The model contributes judgment at known points. It never chooses the path.
A tool-using workflow. Same fixed sequence, but at one step the model chooses a tool and its arguments, say which log query to run. Code still owns the order of operations. The model fills in a blank.
An agent. The model sees the alert, decides to query logs, reads the result, decides to look at the diff, notices a changed timeout, decides to check the payment provider's latency, and so on. Nobody wrote that sequence down. Its length was not known in advance. That is the defining property, and it is also why the agent version is the hardest to test: you cannot enumerate its paths, so you cannot write a test for each one.
The usual framing treats this ladder as a maturity model, with agents at the top. It is better read as a cost curve. Each rung buys flexibility with predictability. If the investigation always follows the same five steps, a workflow will be cheaper, faster, and far easier to debug than an agent that rediscovers those steps on every run. Climb only when the path cannot be written down.
Mechanically, an agent is a loop: give the model context, get a proposed action, execute it, feed the result back, repeat. ReAct (Yao et al., 2022) popularized interleaving reasoning with actions this way, and many agent frameworks since are variations on it. Anthropic's guide describes agents as "typically just LLMs using tools based on environmental feedback in a loop."
The loop is ten lines of code. What deserves attention is the anatomy of a single pass through it, because each part fails in its own way and needs its own owner.
Notice how little of a step is the model. The model proposes. Code decides what it saw, whether the proposal may run, where it runs, how the result comes back, whether the step counts as progress, and what to do if it does not. An earlier article treated the loop itself as a state machine; here the interest is in the pieces inside a step.
State is not context
An easy structural mistake in agent code is treating the conversation transcript as the state of the run. Chat APIs take a list of messages and the list is right there, so the shortcut is natural. But a transcript is an append-only log of what was said. It is a poor representation of what is true.
Separate the two:
- Run state is durable and owned by code: the task, the budget spent, which hypotheses have been ruled out, which side effects have happened, what is waiting on a human. It lives in a database or a checkpoint, survives a crash, and can be inspected without reading prose.
- Context is a projection of that state, rebuilt for each model call. It contains what this particular decision needs, in a form the model reads well.
type RunState = {
task: { goal: string; requestedBy: string };
budget: { steps: number; tokens: number; deadline: string };
findings: Finding[]; // facts established by tool results
ruledOut: string[]; // hypotheses already rejected
effects: SideEffect[]; // what this run has changed in the world
pending?: { kind: 'approval' | 'question'; detail: string };
status: 'running' | 'waiting' | 'done' | 'failed';
};
// Context is derived, not accumulated.
function buildContext(state: RunState, step: number): ModelInput {
return {
instructions: SYSTEM_INSTRUCTIONS,
goal: state.task.goal,
established: summarize(state.findings),
doNotRetry: state.ruledOut,
lastObservation: latestObservation(state),
remaining: { steps: state.budget.steps - step },
};
}Once state is explicit, a lot of behavior becomes ordinary code. Resuming after a crash means loading a record, and showing a human what the agent knows means rendering one. Detecting that the agent is retrying a hypothesis it already rejected is a set lookup instead of a hope that the model remembers.
It also changes how you debug. When an agent does something strange, the first question is no longer "what was the model thinking?" but "what was in the context?", and that has a definite answer.
Plans are hypotheses
Agents plan in two broad styles. In the implicit style, the plan exists only in the model's reasoning between actions, re-derived each step. In the explicit style, the model writes a plan as data (a list of steps, each with a goal and a completion condition) and the loop works through it.
Explicit plans are easier to inspect, display, and check. A plan that touches production can be rejected before any step runs, shown to a user for confirmation, or compared with earlier runs to see where behavior changed.
Their weakness is that the world disagrees with them. In the checkout example, a plan might be "read logs, read diff, identify the bad change, write a fix." The logs then show the errors started forty minutes before the deploy. The plan is now wrong, and an agent that keeps executing it will produce a confident, well-formatted fix for the wrong problem.
So plans should be treated as hypotheses that observations can falsify. The practical version:
- Give each plan step an expected observation, and replan when the actual one contradicts it.
- Cap how often replanning can happen, or a confused agent will spend its budget rewriting plans instead of acting.
- Store the plan in run state, not only in the transcript, so a replan is a visible state change.
Planning quality is mostly a model property. Replanning triggers are an engineering property, and they matter more.
Observations are an interface you design
Most attention in tool design goes to the input side: names, descriptions, argument schemas. The output side, what the model sees after a tool runs, tends to get less care, even though it is all the model has to go on for its next decision.
An observation is the only way the world reaches the model. If the log query returns 4,000 lines, the model either sees all of them and loses the thread, or sees a truncated slice and does not know it was truncated. If an API fails with a stack trace, the model may try to interpret the stack trace rather than the failure. If a search returns nothing, "no results" and "the search service timed out" look identical unless you make them different.
A few rules hold up well:
- Shape results for the next decision. Return counts, the top matches, and a marker that more exist, not a raw dump.
- Make failure modes distinct. Empty, error, denied, and timed out are four different observations, and the model should be able to tell them apart.
- Say what was not shown. "Showing 20 of 4,112 matching lines, sorted by time" is information the model can use to refine its next query.
- Treat every observation as untrusted input. Tool output can contain text written by anyone: a log line, a web page, a ticket. That text can contain instructions. More on this below.
The application-side contract for tools, including authorization, validation, and idempotency, is covered in Tool Calling Is an Application Contract. Observations are the other half of that contract.
Memory is a write problem
"Memory" in agent discussions usually means retrieval: find the relevant past facts and put them in context. Retrieval matters, but the harder problem sits on the other side. What gets written to memory, and who decided it was true?
It helps to name three tiers, because they have different lifetimes and different failure modes:
| Tier | Lifetime | Typical contents | Main risk |
|---|---|---|---|
| Working context | One model call | Instructions, current observations, recent turns | Overflow, distraction, missing facts |
| Run state | One task | Findings, plan, effects, budget | Inconsistency after partial failure |
| Long-term memory | Across tasks | User preferences, learned facts, past outcomes | Stale or wrong facts persisting indefinitely |
The third tier is where agents quietly get worse over time. A model that concludes "the payments service is flaky" during one investigation, and writes that to memory, has created a belief that will bias every future investigation. Nobody reviewed it. It may have been true for one afternoon. Worse, if the conclusion came from reading untrusted content, an attacker has a way to plant beliefs.
Treat long-term memory writes like database writes from an untrusted client. Record their provenance (which run, which observation). Give them an expiry or a confidence level. Prefer storing observations over conclusions, since an observation can be re-interpreted later and a conclusion cannot. Where the stakes justify it, require that memory writes be confirmed by a person or by a deterministic check.
Reflexion (Shinn et al., 2023) showed that agents can improve across attempts by writing verbal reflections on their failures into memory. It is a good result, and the same mechanism that lets an agent learn a useful lesson lets it learn a wrong one. The difference is whether something outside the model checks what was learned.
Verification and recovery
An agent decides it is done. That decision is a claim, and the model making it is the same model that might be wrong about everything that came before.
The strength of an agent depends heavily on what can verify its work. In coding, there are tests, compilers, and linters, signals that can flatly disagree with the model. In the checkout example, a proposed fix can be checked by running the failing request against a staging build. Where such a check exists, an agent can iterate toward correctness, because each attempt is measured against something real.
Where no external check exists, as with "write a good summary" or "find the root cause" when the cause is not reproducible, more iterations tend to produce different answers rather than better ones. That is a design signal: keep the loop short, show the evidence, and put a person at the end. Loop Engineering goes further into ranking verifiers and feeding their results back.
Verification also applies to individual steps, not only the final answer. Did the query actually run against the right environment? Did the file the agent claims to have edited change? These are cheap deterministic checks, and they catch a class of failures where the model's narrative and the world have drifted apart.
When a check fails, or a step goes wrong for any other reason, a generic retry is usually the wrong response. Different failures need different handling, and the first job of recovery code is to tell them apart.
type StepFailure =
| { kind: 'transient'; cause: 'timeout' | 'rate_limit' | 'unavailable' }
| { kind: 'invalid_action'; detail: string } // bad args, unknown tool
| { kind: 'denied'; policy: string } // permission said no
| { kind: 'verification'; check: string; evidence: string }
| { kind: 'no_progress'; repeatedFor: number }
| { kind: 'impossible'; reason: string }; // task cannot succeed
function recover(f: StepFailure, state: RunState): Recovery {
switch (f.kind) {
case 'transient':
return { action: 'retry', backoff: true }; // code handles it; model never sees it
case 'invalid_action':
return { action: 'feedback', to: 'model', detail: f.detail };
case 'denied':
return {
action: 'feedback',
to: 'model',
detail: `Not permitted: ${f.policy}`,
};
case 'verification':
return { action: 'replan', evidence: f.evidence };
case 'no_progress':
return f.repeatedFor > 2
? { action: 'escalate' }
: { action: 'replan', evidence: 'repeating' };
case 'impossible':
return { action: 'stop', report: f.reason };
}
}Two details in that sketch matter more than they look. Transient failures are retried by code without consulting the model, since there is nothing for it to decide. And denials go back to the model as observations, so it can choose a different approach instead of trying the same forbidden action again in different words.
Side effects complicate everything. If step six sent an email and step nine fails, retrying from the start sends the email twice. Recovery needs to know which effects have happened (that is what effects in run state is for), which can be compensated, and which cannot. This is the same problem as partial failure in any distributed system, and the same tools apply: idempotency keys, checkpoints after each committed step, and compensating actions written by people rather than improvised by the model.
Deciding what the model should decide
Every agent design contains dozens of small decisions about where a judgment lives. Should the model decide whether to retry? Whether a result is relevant? Whether the task is done? Whether to ask a human?
A useful heuristic weighs three things:
- Can the options be enumerated? If the answer is one of a known set, code or a narrow classifier can usually make the call, and the decision becomes testable.
- Is the decision reversible? Reversible, low-cost decisions are good candidates for the model. Irreversible ones should have code or a person in the path.
- What does a wrong answer cost? A wrong search query wastes a step. A wrong refund costs money and trust.
Retrying a timeout is settled by the first question alone: the options are retry or give up, and a backoff policy chooses between them more reliably than a model would. Choosing which log to read next points the other way on all three: the options are open-ended, a bad choice is easy to reverse, and it costs one step that the next observation will expose. Issuing a refund has enumerable options but is hard to reverse and costly when wrong, so the model can propose it while code and a person decide. Deciding the incident is resolved sits in the middle, so let the model propose it and let a check confirm it.
This is where much of the engineering effort goes. Making the model smarter shrinks the error rate of its decisions. Moving a decision out of the model, when it never needed one, replaces a probabilistic error with code that can be tested.
Why a smarter model does not make a reliable agent
Suppose each step of an agent run is right 95% of the time and the task takes twenty steps. If errors were independent and unrecoverable, the run would succeed about 36% of the time (0.95²⁰). At 99% per step, it is about 82%. This is arithmetic rather than a benchmark. Real errors are correlated, and good recovery fixes many of them. But it shows the shape of the problem: long trajectories amplify small per-step error rates, and recovery design matters as much as per-step accuracy.
There is a second, subtler problem: variance. An agent with 80% average success is not 80% reliable on each task. It may solve some tasks every time and others only some of the time, and the average hides the second group. The τ-bench paper introduced pass^k for exactly this reason: the probability that an agent succeeds on all of k independent trials of the same task. For a task with an 80% success rate, pass^4 is about 41%. For someone who runs the same kind of task repeatedly, pass^k is closer to their experience than a single-trial success rate.
So evaluation has to measure consistency, not just capability. Run each eval case several times. Track the spread, not only the mean. A model upgrade that raises average success while widening variance can be a regression for users.
Measuring this depends on traces. When a traditional service misbehaves, you read logs and metrics and form a theory. When an agent misbehaves, the useful artifact is the trace: every context that was sent, every proposal, every policy decision, every observation, every state change, in order, with costs and timings.
Without traces, agent debugging becomes guessing. With them, most failures sort quickly into a few kinds: the model never saw the relevant fact, it saw the fact and misread it, it chose a reasonable action that a tool executed badly, or verification passed something it should not have. Each points to a different fix, and only one of them is "use a better model."
Evaluation builds on the same traces. Score outcomes (was the root cause right?) and trajectories (did it take a reasonable path, stay within budget, avoid forbidden actions?). Turn every production failure into a regression case. Rerun the suite whenever the model, a prompt, a tool, or the harness changes, because any of them can shift behavior. Harness Engineering covers recording and evaluation infrastructure in more depth.
More agents, more surfaces
Splitting work across several agents (an orchestrator that delegates to workers, or agents that hand off to each other) is often presented as the next step up. Sometimes it is. The real benefit is context isolation: a worker investigating the payment provider gets a clean context focused on that question, rather than a crowded one shared with everything else the run has seen. Parallel workers can also explore independent hypotheses at the same time.
The costs are real and easy to underestimate. Anthropic's write-up of their multi-agent research system reports that in their data, agents used about 4× the tokens of chat interactions, and multi-agent systems about 15×. Every handoff is a lossy compression: the orchestrator sees a worker's summary, not its evidence. Failures become harder to attribute, since a wrong final answer might come from a bad delegation, a bad worker, or a bad merge.
A reasonable default is one agent with good tools until a specific limit, usually context pressure or a need for parallelism, shows up in traces. When splitting, give each worker a narrow task, an explicit output schema, and its own budget, and have it return evidence alongside conclusions. Where the structure of the work is known, it is often better expressed as an explicit graph than as agents negotiating with each other.
Permissions belong to the action
Instructions in a system prompt are not a security boundary. They describe intended behavior to a model that can be persuaded otherwise, by a user or by text it reads along the way.
That second route is the dangerous one. Simon Willison's lethal trifecta names the combination to avoid: an agent that has access to private data, is exposed to untrusted content, and can communicate externally. With all three, an instruction hidden in a web page, an email, or a log line can direct the agent to send private data somewhere. The model cannot reliably tell instructions from its operator apart from instructions embedded in content, because both arrive as tokens in the same context.
The defenses are architectural:
- Permission checks run in code, per action, against the identity of the user the agent acts for, not the agent's own broad credentials.
- Remove one leg of the trifecta where possible. An investigation agent that reads untrusted logs probably does not need to send email or make arbitrary HTTP requests.
- Tie human approval to reversibility. Reading is free. Writing to a staging branch needs no approval. Merging, paying, deleting, or messaging customers does. Approval prompts should show the exact action and its arguments, not the model's description of them.
- Mark context that came from untrusted sources, and restrict what can happen in a step after the agent has read it.
Human-in-the-loop works best as a designed checkpoint at specific actions. Asking for approval on every step trains people to click "approve" without reading, which turns the checkpoint into a formality.
Software with a nondeterministic component
The language around agents invites a particular mental model: an autonomous entity that you instruct, trust, and occasionally correct. That model leads to bad engineering. It suggests that reliability problems are solved by better instructions, and that the right response to failure is a sterner prompt.
The more productive mental model is duller. An agent is a distributed system with one nondeterministic component, and that component's output is a proposal. The engineering question that matters is how far a proposal is allowed to travel, through state, tools, side effects, and memory, before something other than the model has checked it.