AI Engineering / AI Agents / Agentic Systems / AI Infrastructure / Tool Calling
Harness Engineering: The Software Around the Model
Much of what users experience as agent behavior is decided outside the model: what context it sees, which tools exist, what they are allowed to do, where they run, and how runs are recorded and judged. A practical guide to building that harness.
On this page
- The parts of a harness
- Context assembly is a budgeting problem
- Tools are an interface for a non-human caller
- Permissions are tiered, and enforced outside the model
- Execution needs isolation
- State, interruption, and resumption
- Observe every run
- Evaluate with the same harness
- Harness changes are behavior changes
When an agent does something impressive or something alarming, the credit or blame usually goes to the model. Often it belongs elsewhere. The model saw a context someone assembled, chose among tools someone designed, and acted with permissions someone granted, inside an environment someone configured. That surrounding software is the harness.
Harness engineering is the work of building it deliberately. Two teams using the same model can ship agents with very different reliability, cost, and safety, and most of the difference lives here.
The parts of a harness
A harness has a small number of responsibilities, and it helps to name them separately because they fail separately:
- Context assembly: deciding what the model sees on each call.
- Tools and permissions: what actions exist, and which of them a given run may take.
- Execution: where actions actually run, with what isolation and limits.
- State and control: budgets, checkpoints, cancellation, and pausing for humans.
- Observation and evaluation: recording what happened and judging whether it was good.
Context assembly is a budgeting problem
The context window is finite and every token in it competes for the model's attention. Context assembly decides, on every call, what goes in, in what order, and what gets cut.
Treat it as an explicit allocation rather than string concatenation:
type Section = {
name: string;
priority: number; // lower = more important
required: boolean; // never dropped
content: () => Promise<string>;
maxTokens: number; // cap for this section
};
async function assembleContext(sections: Section[], budget: number) {
const ordered = [...sections].sort((a, b) => a.priority - b.priority);
const included: { name: string; text: string }[] = [];
let used = 0;
for (const section of ordered) {
const text = truncateToTokens(await section.content(), section.maxTokens);
const cost = countTokens(text);
if (used + cost > budget) {
if (section.required) throw new ContextBudgetError(section.name);
continue; // optional sections are dropped, and the drop is recorded
}
included.push({ name: section.name, text });
used += cost;
}
return { included, used };
}A few principles hold across most agents:
- Instructions and tool definitions are required and stable. Keep them first and unchanged between calls. Besides clarity, a stable prefix is what lets providers' prompt caching reuse work.
- Retrieved material is ranked and capped. Ten relevant passages beat forty loosely related ones.
- History is compacted, not accumulated. Keep recent turns verbatim and summarize older ones. Keep tool results only as long as they are useful.
- Record what was dropped. When an agent misses something, the first question is whether it was ever in the context.
Tools are an interface for a non-human caller
Tool design is API design for a caller that reads documentation literally and cannot ask follow-up questions. Tool Calling Is an Application Contract covers the application-side contract. The harness-level guidance is about the shape of the tool set:
- Fewer, sharper tools. Overlapping tools (
search_docs,find_docs,query_knowledge) force the model to guess. One tool with clear parameters is easier to use correctly. - Describe when to use a tool, and when not to. The description is the model's only guide.
- Return errors the model can act on.
{"error": "date must be ISO-8601, got '12/10'"}leads to a corrected call. A stack trace or a bare500does not. - Shape and cap outputs. Return the fields the next step needs, with pagination or truncation markers, not a raw dump that crowds out everything else.
Permissions are tiered, and enforced outside the model
Instructions like "do not delete files" are guidance, not enforcement. Every tool call the model proposes should pass through a policy that the model cannot change:
| Tier | Examples | Default policy | | --- | --- | --- | | Read-only | Search, read file, list records | Allow; log | | Reversible writes in scope | Edit a draft, write in a sandbox workspace | Allow; log; make undo possible | | External side effects | Send an email, open a pull request, call a paid API | Require confirmation, or an explicit grant for this run | | Destructive or irreversible | Delete data, deploy to production, move money | Deny by default; require a human with authority |
type Decision = { kind: 'allow' } | { kind: 'confirm'; reason: string } | { kind: 'deny'; reason: string };
function authorize(call: ToolCall, run: RunContext): Decision {
const tool = registry.get(call.name);
if (!tool) return { kind: 'deny', reason: 'unknown tool' };
if (!run.grantedTools.has(tool.name)) return { kind: 'deny', reason: 'not granted for this run' };
switch (tool.tier) {
case 'read':
return { kind: 'allow' };
case 'reversible':
return tool.inScope(call.args, run) ? { kind: 'allow' } : { kind: 'deny', reason: 'out of scope' };
case 'external':
return run.preApproved.has(tool.name)
? { kind: 'allow' }
: { kind: 'confirm', reason: tool.describeEffect(call.args) };
case 'destructive':
return { kind: 'confirm', reason: `Irreversible: ${tool.describeEffect(call.args)}` };
}
}The user's own permissions bound everything. An agent acting for a user should never be able to do what that user could not.
Execution needs isolation
When tools execute code, touch a filesystem, or make network requests, they need an environment that limits the damage of a bad decision or a malicious instruction hidden in content the agent read:
- Sandboxed execution: containers or VMs per run, with no access to host credentials.
- Filesystem scope: a workspace directory, not the whole machine.
- Network egress controls: an allowlist of destinations, which also limits where data can be exfiltrated to.
- Resource limits: CPU, memory, wall-clock time, and output size per tool call.
- Secrets out of context: tools receive credentials from the harness at execution time. The model never sees them, so it cannot repeat them.
Prompt injection, meaning instructions embedded in web pages, documents, or tool output, is best handled here. The harness cannot reliably stop a model from being persuaded. It can make sure a persuaded model still cannot do much harm.
State, interruption, and resumption
Long-running agents need the same operational properties as any long-running job:
- Budgets for steps, tokens, time, and cost, checked by the harness, as described in Loop Engineering.
- Cancellation that reaches in-flight model and tool calls.
- Checkpoints so a run can resume after a crash or a deploy.
- Pauses for approval that persist state and resume when a human responds, without holding a process open.
Observe every run
A trace per run should answer "what did the agent see, decide, and do, and why did it stop?" without anyone reconstructing it from logs:
- the versions of the model, the instructions, and the tool definitions;
- the assembled context per call, or a reference to it, with what was dropped;
- each proposed tool call, the policy decision, arguments (with secrets redacted), duration, and result size;
- token usage and cost per call;
- the stop reason.
Versioning is the part teams most often skip. A change to a tool description is a behavior change. If traces do not record which description was active, a regression cannot be attributed.
Evaluate with the same harness
An evaluation suite that calls the model directly measures the model. Users experience the model inside the harness. Run evaluations through the production harness, with the same context assembly, tools (against fixtures), policies, and budgets, so that a change to any of them is measured.
A workable structure:
- Fixtures: realistic tasks with the files, records, and tool responses they need.
- Graders: deterministic checks first (did the tests pass, is the output valid, was the forbidden tool called), and model-based rubric grading only where needed, checked periodically against human judgment.
- Comparison: every change runs against the baseline, and results are compared per task category, not only in aggregate.
Harness changes are behavior changes
The practical consequence of all of this is organizational. Prompts, tool definitions, context policies, permission tiers, and budgets are code that determines product behavior. Keep them in version control, review them, test them against the evaluation suite, and roll them out with the same care as any other change. Teams that treat the harness as configuration to tweak in a dashboard end up with agents whose behavior nobody can explain.
The model provides capability. The harness decides how much of it is usable, safely and repeatably. That is where most of the engineering is.