Tools & retrieval
What is Tool Calling?
Also known as: function calling
Tool calling is when a model emits a structured request that a tool executes, then returns a result.
The result is an observation. The model does not magically browse; a tool does. Also called function calling in some APIs.
The model emits `{ "tool": "search", "query": "MCP spec 2024" }`. Your app runs the search and returns snippets as the next observation. The model never touched the network — your code did.
Visual
Tool calling flow
The application sits between the model and the tool, validating every call.
Why it matters
Tool calling is what makes agents capable of acting in the world. Without tools, a model can only generate text. With tools, it can search the web, read files, call APIs, write to databases, and perform any action you expose.
The crucial insight is that the model does not execute tools — your application does. The model emits a structured request; your code validates, executes, and returns the result. This separation is what makes tool calling safe and controllable.
How tool calling works
The process has four steps. First, you declare available tools with their names, descriptions, and parameter schemas. Second, the model generates a tool call request matching one of those schemas. Third, your application validates and executes the request. Fourth, you return the result as an observation in the model's context.
This is fundamentally different from giving the model raw access to execute code. The application layer is the gatekeeper.
Designing good tools
Good tools have clear names, typed parameters, and predictable outputs. They do one thing well. Bad tools have vague descriptions, untyped inputs, and side effects that are not declared.
Each tool should return structured data, not prose. The model needs to parse the result programmatically to decide what to do next.
Key takeaways
- 1The model requests tool calls — your application executes them.
- 2Tools should have typed parameters, clear descriptions, and predictable outputs.
- 3The application layer is the gatekeeper for safety and validation.
- 4Return structured data, not prose, as tool results.
Common mistakes
- ✕Giving the model raw code execution without sandboxing.
- ✕Writing tool descriptions that are vague or misleading.
- ✕Returning unstructured text as tool results, making parsing unreliable.
- ✕Exposing write tools without confirmation gates.
In practice
Most production agents use 2–5 tools. Common tool sets include search + browser for research, repository + test runner for coding, and database + API for business tasks. Start small: each additional tool increases the model's decision space and error surface.
Related terms
Concept neighborhood
Terms linked from Tool Calling in the glossary graph.
- Tool Calling
- Function Calling
- Agent
- Observation
FAQ
- Is tool calling the same as plugins?
- Plugins are a product packaging of the same idea: the model requests a structured action and your code executes it. Tool calling is the underlying mechanism.
- Can the model run tools directly?
- No. The application executes tools and returns results. The model only generates the request.
- How many tools should an agent have?
- Start with 2–5 focused tools. More tools increase the decision space and error rate. Add tools only when the agent demonstrably needs them.