Home / Learn / Tools, skills and protocols

Guide

Tools, Skills and Protocols for AI Agents: A Design Guide

Tools for AI agents, explained: task-shaped tools, schemas, sandboxes, skills and protocols such as MCP. Start with Chapter 2, free to read.

By Enrique Gutiérrez · Last reviewed

A tool is a function the model can ask your code to run; the model never executes anything itself. Tool calling is the request convention, a skill is a procedure loaded when needed, and a protocol is how any agent finds any tool. Designing tools for AI agents well separates a reliable agent from a flaky one.

This guide to tools for AI agents follows Part II of AI Agents, Engineered: Chapter 5, “Tools and the Action Space” and Chapter 6, “Skills, Protocols, and Interoperability”, both in the full book. The exchange itself is in Chapter 2, “The Engine”, free to read.

What is tool calling, and what actually happens in the exchange?

Tool calling, also called function calling, is a convention in which you describe a set of functions alongside the prompt, and the model, when it judges an action useful, emits a structured request naming one function and its arguments instead of prose. Then it stops and waits for your code to answer.

Chapter 2 puts the division of labor in two short sentences: “The model never executes anything. It has no hands; it only writes.” Your harness receives the request, validates it, runs the function, and sends the result back as a new message. The model reads that result and decides what to do next. The book’s summary of the arrangement is the line I would keep if I kept only one: “the model proposes; your code disposes.”

So every side effect happens inside ordinary code you wrote, which is where logging, permissions and refusals belong. And the result goes back to the model, not to the user, because only the model knows why it asked. See the tool call glossary entry and the agent loop explainer.

Whether the request is well formed is a separate question: prompt instructions ask for the shape, constrained decoding enforces it. See structured output for tool calls and what is tool calling in LLMs. Tool calling is the interface, not the connectivity; how it differs from a protocol is the subject of MCP vs function calling.

What is an agent’s action space, and why does it matter so much?

An agent’s action space is the full set of actions available to it, and in practice that means the tools you expose. Chapter 5 opens by asking you to take this literally: “The set of tools you expose is the complete inventory of what your agent can ever do.”

The consequence is that many agent failures are settled before the first token is generated. A model with no way to search your codebase will guess at its contents. A model with no way to ask a clarifying question will charge ahead on a wrong assumption. Neither failure is fixed by better prompting, because the missing capability was never on the list. That is why the book calls choosing and shaping the action space “the core craft of building agents,” upstream of the prompt.

A tool is also a new kind of contract. Anthropic’s tool-writing guidance calls it “a contract between deterministic systems and non-deterministic agents” (Aizawa, 2025): the caller may pick the wrong tool with perfect syntax or invent one that never existed. Design tools for AI agents with that caller in mind, defensively.

How do you choose the set and granularity of tools for AI agents?

Shape each tool to a task, not to an endpoint. Wrapping every API endpoint one-for-one works, which is what makes it a trap: the model ends up doing data plumbing in its head, reading every intermediate result into context, where each extra byte costs tokens and adds a chance to misread.

Chapter 5’s example is a meeting assistant over a calendar API with list_users, list_events and create_event. Wrapped one-for-one, the agent pulls the directory, fetches six calendars, and intersects them by reading. Redesigned around the task, there is one tool, schedule_event(participants, duration, window), which does all of that in ordinary code and returns one line: booked, Thursday, two o’clock. The book’s rule: “The right unit for a tool is a task; the endpoints are plumbing, and plumbing belongs inside.”

The same scheduling job in two action spaces.
Figure 5.2 The same scheduling job in two action spaces. With one endpoint wrapped per tool (top), every intermediate byte crosses the model’s desk and every cross-reference risks a misread. Shaped to the task (bottom), the deterministic work happens inside the tool and the model sees one line back. Chapter 7 prices the difference. Reuse this diagram

The same move generalizes. Expose search_logs, which returns matching lines with a little context, instead of read_logs, which returns everything. Expose one get_customer_context instead of three lookups the model must chain and collate. The model keeps the part that needed judgment, understanding the request, and the computation moves to where it is cheap and exact.

Fit is empirical: expose a tool, run real tasks, read the transcripts, and expect the answer to move as models improve. As the book puts it, “A tool that rescues today’s model can constrain next year’s.” The guide on how to design tools for LLM agents goes further.

How many tools is too many?

There is no magic number, but the direction is clear: past a point, every addition makes the agent worse. Each definition rides in context on every pass and is one more option to weigh, and overlapping tools present a choice where none should exist. The book compresses it: “Past a certain count, additions subtract.”

Three moves add capability without adding tools. Give the model an entry point instead of an inventory, so detail loads only when needed. Delegate self-contained jobs to a subagent. And when a large set is unavoidable, namespace it by service and resource (crm_contacts_search, billing_invoices_search) so kinship shows in the names. When the agent keeps picking the wrong one anyway, the post on why your AI agent calls the wrong tool lists the usual causes.

What belongs in a tool’s interface?

A tool’s interface is its name, description and argument schema, plus two more surfaces people forget: what it returns and how it fails. The definition is everything the model will ever know about the tool, and it is reread on every pass, so each part works as a prompt whether or not you wrote it as one.

The anatomy of a tool across the boundary between your process and the model’s context.
Figure 5.3 The anatomy of a tool across the boundary between your process and the model’s context. Three parts cross to the model’s side and steer it like prompt text—the description above all, in accent, the highest-leverage surface you write. The function itself never leaves; only its results and errors travel back, as still more text the model has to read. Reuse this diagram

Write the description like documentation for a capable new hire, making explicit what you carry implicitly: the query format, the niche term, how one resource relates to another. Name parameters so they admit one reading; Anthropic’s example is to use user_id rather than user (Aizawa, 2025). Pin categorical values with enums and forbid extra properties. The same source reports a web-search tool whose model kept appending the current year to queries, fixed by “improving the tool description,” with no change to the model. Chapter 5 turns that anecdote into a debugging order: “when a model misuses a tool, suspect the contract before you blame the intelligence.”

Returns are context for the next decision: names rather than opaque IDs, the relevant slice rather than the dump, and a note when you truncate. Errors are the highest-signal message a tool sends. Error 422 teaches nothing; severity must be one of: low, med, high (got "urgent") lets the model fix its own call on the next pass.

Why are tool descriptions an attack surface?

Because the model reads a tool description as instructions, and whoever writes the description can write anything into it. If you connect to tools someone else built, their descriptions enter your agent’s context with the authority of documentation, and often without a human ever reading them in full.

Chapter 6 says it plainly: a protocol server “is code you run or trust, its tool descriptions are words your model will read and obey, and the marketplace where you found it vets less than you hope.” Security researchers have demonstrated the attack. Invariant Labs named it the tool poisoning attack: it “occurs when malicious instructions are embedded within MCP tool descriptions that are invisible to users but visible to AI models” (Beurer-Kellner and Fischer, 2025). Their proof of concept hid, inside an innocent-looking addition tool, instructions to read private keys and pass them out through an extra argument. They also described “rug pulls,” where a server changes a description after approval.

This is a form of prompt injection, and the defenses are the ones you apply to any dependency: read descriptions before installing, pin versions, and keep consequential actions behind checks in your harness. The post on the lethal trifecta explains which combination of capabilities to break up.

Why must agent tools be idempotent?

An operation is idempotent when running it twice has the same effect as running it once, and agent tools need the property because agents retry constantly: on timeouts, on ambiguous output, on resume after a crash. A retried create should return the existing record, flagged as pre-existing, instead of minting a second one.

Without it, every retry is a chance to send the email twice or charge the card twice, and the agent will not notice the duplicate the way a human would. Chapter 5 starts the idea at the tool level, and Chapter 18 builds the reliability machinery on it. The companion post on idempotent tools and safe retries covers idempotency keys, receipts and which retry policy fits which tool, with a worked example of a crash mid-send.

How do you retrofit software built for humans into agent tools?

Make explicit what a human used to supply implicitly. Command-line programs and internal services assume a forgiving, resourceful user who answers prompts, skims help once and eyeballs messy output; an agent can do none of that, so every gap human judgment papered over becomes a hang, a burned retry or a duplicate record.

The rules unpack in the order they bite. Never block on an interactive prompt: every input should be passable as a flag, argument or file, and a program with no terminal attached should fail fast instead of waiting. Keep destructive actions confirming by default, with an explicit flag to bypass the question non-interactively. Emit machine-parseable output, with one convention everywhere. Make help text say when to use a command. Offer a dry run on mutations, and validate inputs as strictly as a public API.

One much-cited set of field notes compresses the priority: “Crashes are tolerable; hangs are problematic” (Ronacher, 2025). A hang teaches the agent nothing; a fast, clear error is something it can act on. Ronacher’s examples: logs written to a file the agent can read, and a second launch that errors with “services already running.”

None of these rules hurt your human users, and together they turn the same software into usable tools for AI agents. The post on designing agent tools for software built for humans turns them into a checklist.

Is code execution the universal tool?

Code execution is the closest thing to a universal action: one tool, conceptually run_code(source), lets the model write a program that filters, joins, computes and calls your other tools as functions, while an isolated runtime executes it. It is a meta-tool, and it resolves the moment a task steps off any fixed menu.

The slogan the book adopts is “discretion on the outside, determinism on the inside.” The model decides what to compute; the runtime computes it exactly. Models are also better at this than you might expect. Cloudflare’s engineers, converting tool definitions into a code library the model programs against, concluded that “LLMs are better at writing code to call MCP, than at calling MCP directly,” because models have read far more real code than tool-call transcripts (Varda and Pai, 2025).

The economics follow from where data lives. If ten thousand rows stay in the runtime, the model reads back five lines instead of the whole export. Anthropic reported one rebuilt workflow going “from 150,000 tokens to 2,000 tokens,” a 98.7% reduction (Anthropic, November 2025); that is one vendor’s figure on one workflow. The agent cost per task estimator shows how much of a bill is context re-read every turn.

The honest edges: a sandbox is infrastructure to run and pay for, a short fixed menu is easier to audit, and a wrong program gives the wrong answer with perfect consistency. The book’s rule: “Give the model the judgment, and give the runtime the arithmetic.”

What does the sandbox isolate, and why is it mandatory?

A sandbox is the sealed room the interpreter runs in: its own filesystem, capped processor, memory and time, and no network except destinations you deliberately allow. It is mandatory because the code was written by a probabilistic author that may have been steered by untrusted text it read along the way.

Chapter 5 states the rule without hedging: “never run model-written code with access to anything you would not hand a stranger.” Isolation comes in strengths, from ordinary containers to kernels that intercept system calls to per-session micro virtual machines; those are examples of a category, not a shopping list. Two settings do most of the protecting. Deny the network by default, because a sandbox that can phone out can exfiltrate whatever the code touched. And give each session a fresh sandbox, torn down afterward, so nothing leaks between runs. A proven script can be kept and reused, which is where skills begin. The post on sandboxing agent tool execution goes deeper.

When should an agent drive a screen instead of an API?

Only when nothing better exists. Computer use hands the model a screenshot, a cursor and a keyboard so it can operate software built for human eyes, such as a legacy portal with no API. It reaches everything and is the slowest, most fragile way to reach anything, so treat it as the bottom rung of a ladder.

The rungs, in the order the book recommends: an API wherever one exists, because it is a deterministic contract; then a structured browser tool that targets elements through the page’s accessibility tree, by name and role rather than by position; then pixel-level control, reserved for software that offers no better door. Targeting by meaning survives a layout shift that sends a coordinate click into whatever moved in, and it lets your harness know what is being clicked, so it can refuse purchases while allowing reads.

The three ways to reach another system, drawn as a ladder.
Figure 5.5 The three ways to reach another system, drawn as a ladder. Reliability falls and reach widens as you descend—an API is a deterministic contract, a structured browser targets elements by meaning through the accessibility tree, and raw pixels can drive anything at all but the most cost. Start at the top rung (in accent) and climb down only when the higher door does not exist. Reuse this diagram

Treat the lower rungs as reserve rather than residence: APIs for most steps, the screen only for the gap. The agent sees discrete snapshots, so the harness should confirm each state change before the next action.

The headline risk is that every page the agent reads was written by strangers. Hostile text hidden in a page is indirect prompt injection, which Chapter 17 opens with; until then, keep least privilege and a human confirmation before consequential actions.

What is a skill, and how does progressive disclosure keep it cheap?

A skill is packaged expertise: a named folder holding instructions written in ordinary language, plus optional reference files, templates and scripts, which the agent discovers by a short description and loads only when a task calls for it. Tools give an agent hands; skills give it your procedures.

The failure skills fix is distinctive: every tool call succeeds and the work is still wrong, because the agent skips the release checklist everyone knows. Pasting every procedure into the standing prompt pays for all of them on every call. A skill avoids that through progressive disclosure, loading in three levels: the name and one-line description of every skill, always present; the instruction body of one skill, loaded when the agent judges it relevant; and the bundled references and scripts, fetched piece by piece. Scripts can be executed without being read, so their cost in context is close to zero.

The mechanism lives or dies on one line. At the moment of choosing, the description is all the agent sees, and a vague one means the skill silently never loads. The book’s advice: say what the skill does and when to use it, in the words a user would type, and “Treat the description as API documentation for a caller that cannot read your source.” Triggering is probabilistic, so test it on your actual host. And a skill is files, which cuts both ways: “a skill is a dependency that ships instructions, and you already know better than to install dependencies unread.”

How do skills, tools and protocols differ?

A tool is a single action, a verb. A protocol is connectivity, a standard way to discover and call capabilities. A skill is expertise, the know-how for using the verbs and connections to your standard. They are layers, not alternatives, and most design confusion comes from treating a gap in one layer as a gap in another.

The three layers this chapter keeps apart.
Figure 6.3 The three layers this chapter keeps apart. Tools are the actions an agent can take, protocols are the standard way it connects to them, and skills (in accent) are the expertise that uses both well—each answering a different question. The shared baseline is the point: they are complementary layers of one stack, not rivals, and conflating them is the common beginner error. Reuse this diagram
Tool Skill Protocol
What it is One action the agent can take Packaged procedure, references and scripts Agreed contract for discovery and invocation
Question it answers What can the agent do? How do we do this here? How does the agent reach it?
Example query_db, send_email, run_code Monthly revenue report, house style A tool protocol; an agent-to-agent protocol
Missing it looks like Agent improvises or guesses Every call succeeds, the work is still generic Agent cannot reach the system at all
What to fix Cut a task-shaped tool Write the skill from observed failures Install or implement a server
Main risk Duplicate side effects, blast radius Stale or malicious instructions Untrusted servers and drifting contracts

The diagnostic is one question: what kind of thing is missing? Chapter 6 runs it on a database. A protocol gives reach, the tools itemize that reach, and the agent still produces the wrong revenue report until a skill supplies your fiscal calendar and definitions. “Connectivity without expertise yields generic output; expertise without connectivity cannot act.” The stack, in the book’s compact form: “the loop reasons, the runtime executes, the protocol connects, and the skills guide.” See when a tool should become a skill.

What problem do protocols solve, and what do they add?

A protocol solves the M×N adapter problem. Without a standard, every agent needs a hand-written adapter for every system; with one, each system implements the standard once and each agent speaks it once, and in the book’s words “the count collapses from a multiplication, M×N, to an addition, M+N.”

The multiplication a protocol removes.
Figure 6.4 The multiplication a protocol removes. Three agents and five systems wired directly need fifteen hand-built adapters, and every added agent or system bills the whole other side. With one shared contract in the middle (in accent), the same reach costs eight implementations, and an addition costs one. Reuse this diagram

Two kinds matter. A tool protocol connects an agent to tools and data. A prominent example is the Model Context Protocol (MCP): a server advertises capabilities in machine-readable form, and a client inside the agent asks “what do you offer?” and calls what comes back. An agent-to-agent protocol connects an agent to another agent that reasons on its own; Agent2Agent (A2A) is one widely cited example, and its launch post calls it “an open protocol that complements Anthropic’s Model Context Protocol” (Google, April 2025). Both are labeled examples of categories that will keep evolving, not recommendations. If the far end is a capability, use a tool protocol; if it plans and owns its method, use an agent protocol and a well-written brief.

Protocols have costs: young, moving specs; a far end that drifts without asking; and a generic interface that, for one integration you control, is more ceremony than a direct call. And trust: a standard socket makes connecting to a stranger a one-line change, hence the book’s warning that “every socket is a door.” The post on what MCP is in AI takes the protocol layer further.

Worked example: giving a support agent hands

Suppose a company runs three agents (support, research, coding) that need five systems: ticket tracker, CRM, wiki, code host and data warehouse. The numbers are illustrative.

Connectivity. Hand-written, that is up to 3 × 5 = 15 adapters, and each new system costs three more. Behind a tool protocol it is 3 + 5 = 8 implementations, and the sixth system costs one. With one agent and one system, a direct call wins.

Tools. Instead of the CRM’s forty endpoints, the support agent gets three task-shaped tools: get_customer_context, search_tickets and issue_refund(order_id, amount, idempotency_key). The key derives from the action’s logical identity (run, step, tool, arguments), so a retry returns the existing refund.

Code. For “find all customers charged twice last week,” the agent’s script joins warehouse and ticket data in a network-less sandbox and reads back twelve names, not a day of transactions.

Expertise. Every call succeeds, yet replies ignore the refund policy: a skill gap, not a tool gap, fixed by a refund skill with the policy and a reply template.

Trust. The third-party CRM server’s descriptions are read and pinned before installation, and issue_refund sits behind an approval gate above a set amount. That is not enough: the agent holds all three legs of the lethal trifecta, reading untrusted tickets, holding private CRM data and sending replies out. A hostile ticket can request another customer’s records, so break a leg: scope CRM lookups to the ticket’s own customer and check replies before they leave. The lethal trifecta audit checks your own agent.

If the double-charge investigation takes eight tool calls, each 95% reliable unchecked, the run succeeds about 66% of the time (0.95⁸); the compounding error calculator shows why fewer, larger, verified steps help.

Where does each idea live in the book and on the site?

For each topic covered here, the table gives the chapter that develops it, the related post or tool, and the glossary term, so you can go one level deeper wherever you need to.

Concept Chapter Related post or tool Glossary
The tool-calling exchange Chapter 2 (free) What is tool calling in LLMs? tool call
Action space and granularity Chapter 5 How to design tools for LLM agents action space
The tool interface Chapter 5 Why agents call the wrong tool tool
Idempotency and retries Chapter 5 Idempotent tools and safe retries idempotency
Retrofitting human software Chapter 5 Agent tools for software built for humans harness
Code execution and sandboxes Chapter 5 Agent cost per task estimator sandbox
Skills and progressive disclosure Chapter 6 When a tool should become a skill skill
Protocols and the M×N problem Chapter 6 What is MCP in AI? protocol
Poisoned descriptions, injection Chapter 6 The lethal trifecta prompt injection

Where does this advice stop applying?

Most of this guide assumes a capable model and an action space you control, and most advice about tools for AI agents carries the same assumption. Two limits are worth stating. First, everything here is empirical. Tool granularity, namespacing and skill descriptions behave differently across models and hosts, so the right answer for your agent comes from transcripts and with-and-without comparisons on your tasks, not from this page.

Second, not every integration deserves the full stack. A single lookup does not need a sandbox, and a procedure you type once is not yet a skill: the simplest thing that works, upgraded on evidence. Security is the exception that does not scale down; one third-party server is enough to need the security and operations guide.

The agent now has hands, a shelf of manuals and standard sockets. The question that remains is the book’s thesis: what signal tells you each of them worked? Read Chapter 2 free for the exchange underneath every tool call, and when you want Chapters 5 and 6, see the formats.

The chapters behind this guide

  1. Chapter 5: Tools and the Action Space In the full book
  2. Chapter 6: Skills, Protocols, and Interoperability In the full book

Tools and explainers for this topic

Tool

Agent cost-per-task estimator

Estimate what one agent run costs in tokens: the fixed prompt, the history it re-reads every step, retries and subagents. Free AI agent cost estimator.

Tool

Compounding error calculator

A free compounding error calculator for AI agents: whole-run success from per-step reliability, and the reliability a long task needs.

Explainer · 3 min

The agent loop: four beats and three exits

A three-minute animated explainer of the agent loop: the four beats of every pass, the history that is the agent's only memory, and three exits ranked by trust.

Articles in this cluster

Questions readers ask

Does the model run the code when it calls a tool?
No. The model only writes a structured request naming a function and its arguments, then stops. Your harness decides whether to execute it, runs it in ordinary code, and sends the result back to the model. Every side effect happens on your side of that boundary, where you can log, test, restrict or refuse it.
Do I need a protocol to give an agent tools?
No. Tool calling works with functions you describe directly to the model and run in your own harness. A protocol earns its place when several agents must reach several systems, or when you want to plug in tools other people built; for one agent and one system you control, a direct call is simpler and easier to audit.
How many tools is too many for an AI agent?
There is no fixed number, because it depends on the model and on how distinct the tools are. Every definition sits in context on every pass and is one more option to weigh, and overlapping tools confuse selection. Keep the set small and task-shaped, add tools only when a measured gap demands them, and test on your own tasks.
How do I know a tool description is the problem?
Read the transcripts. If the model picks between two overlapping tools, guesses a parameter’s format, or calls a tool at the wrong moment, the contract is unclear: the description says what the tool does but not when to use it. Tighten names and descriptions before blaming the model; the post “Why Your AI Agent Calls the Wrong Tool” lists six fixes.
Can I give an AI agent a shell?
Yes, inside a sandbox: an isolated runtime with its own filesystem, capped resources, no network except destinations you allow, and a fresh instance per session. Never run model-written code with access to anything you would not hand a stranger, because the code may have been steered by text the agent read.