AI Engineering / Tool Calling / AI Infrastructure / Backend Engineering
Tool Calling Is an Application Contract
Tool schemas are only one part of the boundary. Authorization, validation, side effects, idempotency, and result shaping belong to the application.
On this page
A model-generated tool call is a request, not a capability. The application owns the decision to run it.
That distinction matters because a tool usually crosses a real system boundary: it reads private data, changes state, sends a message, or invokes another service. A schema can describe the shape of the arguments, but it cannot decide whether this user may perform this action or whether the action is safe to repeat.
Separate proposal from execution
Keep the model-facing tool definition separate from the executor. The definition tells the model which operation exists and what input it expects. The executor performs application checks using trusted context that the model does not control.
async function handleToolCall(call: ModelToolCall, context: RequestContext) {
const tool = toolRegistry.get(call.name);
if (!tool) {
return { ok: false, code: 'UNKNOWN_TOOL' };
}
const input = tool.inputSchema.parse(call.arguments);
await tool.authorize(context.user, input);
return tool.execute(input, {
userId: context.user.id,
requestId: context.requestId,
});
}The user identity, tenant, and request identifier come from the authenticated application request, not from model output. This makes the trust boundary reviewable and testable.
Validate in layers
Useful checks happen at more than one layer:
- Shape: parse the arguments and reject missing or unexpected values.
- Meaning: verify that referenced records exist and that values are valid for the operation.
- Authorization: check access using the authenticated actor and the target resource.
- Policy: apply product rules such as confirmation for destructive actions or limits on bulk operations.
- Execution: use the narrowest service method that performs the requested work.
The model can help choose among allowed actions, but it should not supply authority. For example, a tool should not accept an arbitrary userId just because the schema can represent one. Derive identity from trusted request context wherever possible.
Make side effects deliberate
Read-only tools and side-effecting tools deserve different handling. A search can usually be repeated. Creating a payment, sending an email, or changing a record cannot be assumed safe to repeat after a timeout.
For operations where retries are possible, define idempotency at the business-operation boundary. An idempotency key should represent the intended action and be checked by the system that commits it. Retrying an HTTP request without that guarantee can duplicate the side effect even when the agent loop itself is correctly bounded.
Destructive or externally visible actions may also need an explicit confirmation step. The confirmation should be tied to a specific normalized action and its important parameters, rather than a generic “continue?” after the model has had an opportunity to change its proposal.
Return useful, bounded results
Tool output is another input to the model, so it should be shaped for the next decision. Return only fields needed for the task, cap large collections, include stable identifiers when follow-up actions depend on them, and represent errors with predictable codes. Do not send database rows or upstream response bodies wholesale.
Keep error categories distinct. Invalid arguments, access denied, not found, rate limited, and transient dependency failure imply different next steps. A model may decide how to explain an error, but application code should determine whether retrying is permitted.
Test the contract without the model
Most tool guarantees can be tested as ordinary application behavior:
- unknown tool names never execute;
- invalid input fails before a side effect;
- authorization uses the authenticated actor;
- retries do not duplicate supported writes;
- output is bounded and contains no unintended fields;
- cancellation and timeouts reach downstream work.
Model evaluation can then focus on whether the right tool is selected and whether the result is used correctly. Keeping those questions separate makes failures easier to locate: tool selection is a model behavior; access control and execution are application behavior.
Treating tool calling as a contract keeps the model useful without letting a generated JSON object silently become permission to operate your system.