A model on its own can only produce text. Tool calling is the mechanism that lets it ask your code to do something — query a database, hit an internal API, read a file — and then carry on with the answer. It is the single feature that separates a chat box from software that actually does a job, and it is simpler than the frameworks built on top of it suggest.
A tool is a JSON schema and a function
You describe what the tool does and what arguments it takes. The model never executes anything — it returns a structured request to call it, and your code decides whether to comply.
tools = [
{
"name": "get_disk_usage",
"description": (
"Return disk usage for a mount point on the application server. "
"Use when asked about free space, full disks or storage alerts."
),
"strict": True,
"input_schema": {
"type": "object",
"properties": {
"mount": {
"type": "string",
"description": "Absolute mount point, e.g. /var",
}
},
"required": ["mount"],
"additionalProperties": False,
},
}
]
Two details worth copying. The description is prompt text, not documentation — it is how the model decides whether this tool is the right one, so write when to use it, not just what it is. And strict: True (which requires additionalProperties: false plus a complete required list) guarantees the arguments you get back validate against the schema, so you can stop writing defensive parsing.
The loop, by hand
Write it once manually — it is worth understanding before you hand it to a helper. The pattern is: call, check stop_reason, run any requested tools, append the results, call again.
messages = [{"role": "user", "content": "Is /var nearly full on the app server?"}]
for _ in range(10):
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
tools=tools,
messages=messages,
)
if response.stop_reason != "tool_use":
break
messages.append({"role": "assistant", "content": response.content})
results = []
for block in response.content:
if block.type != "tool_use":
continue
try:
output = run_tool(block.name, block.input)
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": output,
})
except Exception as exc:
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": f"{type(exc).__name__}: {exc}",
"is_error": True,
})
messages.append({"role": "user", "content": results})
else:
raise RuntimeError("tool loop did not terminate after 10 turns")
print(next(b.text for b in response.content if b.type == "text"))
That for ... else is doing real work: the else branch only runs if the loop finished without breaking, which is exactly the runaway case. Never write this as while True.
Letting the SDK drive it
For anything new, use the tool runner. You decorate ordinary Python functions, the schema is generated from the signature and docstring, and the loop is handled for you.
import shutil
import anthropic
from anthropic import beta_tool
client = anthropic.Anthropic()
@beta_tool
def get_disk_usage(mount: str) -> str:
"""Return disk usage for a mount point on the application server.
Args:
mount: Absolute mount point, e.g. /var
"""
total, used, free = shutil.disk_usage(mount)
return f"{used / total:.0%} used, {free // 2**30} GiB free on {mount}"
runner = client.beta.messages.tool_runner(
model="claude-opus-5",
max_tokens=16000,
tools=[get_disk_usage],
messages=[{"role": "user", "content": "Is /var nearly full?"}],
)
for message in runner:
print(message.stop_reason)
TypeScript has the same thing: betaZodTool from the SDK’s Zod helpers to define the tool, then client.beta.messages.toolRunner({...}), which resolves to the final message. If you would rather not depend on Zod there is a raw JSON Schema variant.
The gotchas that bite
Loops that never terminate
The classic failure is a tool that returns something ambiguous — an empty list, a bare “OK”, a stack trace — so the model tries again with slightly different arguments, forever. Fix it at the source: return concrete, self-describing data (“no rows matched status=’open’”, not []), and keep the hard iteration cap as a backstop. Log the tool name and arguments on every iteration; a runaway loop is obvious in the log and invisible in the output.
Answer every block, in one message
A single assistant turn can contain several tool_use blocks — that is parallel tool calling, and it is on by default.
- Return all the
tool_resultblocks in one user message. Splitting them across several messages quietly teaches the model to stop calling tools in parallel. - Every
tool_use_idmust get a matching result. A missing one is a 400, not a warning. - When a tool fails, return the result with
is_error: True— do not drop the block. The model can then recover or report the failure.
Treat arguments as untrusted input
block.input is already a parsed dictionary — never do string matching against a serialised version of it, because escaping varies. More importantly, the arguments came from a model reading text that may have come from a user. Validate the mount point against an allowlist. Parameterise the SQL. A tool is an API endpoint with an unusually persuasive caller.
Designing tools worth calling
Fewer, broader tools beat a sprawl of narrow ones — a model choosing between six tools is far more reliable than one choosing between forty. Name them after the intent rather than the implementation, and keep results small: every tool result is resent on the next request, so a 50KB JSON blob is a recurring cost for the rest of the conversation. If you find yourself writing the same wrappers repeatedly, MCP standardises them so they are reusable across clients — I keep a running set of notes in my MCP hub.
What I’d do
Start with the tool runner and typed functions — it removes the loop bookkeeping without hiding the model’s behaviour. Drop to the manual loop only when you need a request shape the runner will not build. Either way, put a hard iteration cap in, set strict: True on every tool schema, and make sure each tool returns something a human would find unambiguous. Tool calling fails far more often on vague results than on wrong code.

Leave a Reply
You must be logged in to post a comment.