Home / Blog / Tools, skills and protocols / How to Design Tools for LLM Agents: Shape Them …

Tools, skills and protocols

How to Design Tools for LLM Agents: Shape Them to the Task

How to design tools for LLM agents: shape each tool to a task, not an endpoint. One API wrapped two ways, a symptom table and a template to copy.

By Enrique Gutiérrez · Published · 24 min read

How to design tools for LLM agents comes down to one move: shape each tool to a task the agent performs, not to an endpoint of the API underneath. Then write the description, the argument schema, the result and the error as text a model will read, because each of the four is a prompt.

I hold “fewer, larger tools” as a starting hypothesis to test, not a law. Chapter 5 of the book this site belongs to says that what settles granularity is the model’s current ability, “an empirical matter: found by testing, never by taste.” I found no controlled study that wraps one API both ways and compares them, and no public measurement of actionable against opaque error messages. So the case below rests on reasoning, one example counted by hand, and adjacent measurements, each with its date and caveat.

What does the model actually see of your tool?

The model sees three pieces of text per tool: a name, a description and a schema for the arguments. Your function stays in your process. A tool call is the model writing a name and an argument object; your harness runs the function and hands the result back as more text for the next turn of the agent loop.

Two consequences follow. The first is that the caller improvises. Chapter 5 of AI Agents, Engineered describes a caller “that may call the right tool with wrong arguments, call the wrong tool with perfect syntax, or invent, in a confident flourish, a tool that has never existed.”

The second is cost. Each call to the model is one pass, and “Every tool definition rides on the desk on every single pass.” That list is the agent’s action space, “the complete inventory of what your agent can ever do.”

Should you expose your REST API one tool per endpoint?

Usually not. An endpoint-per-tool wrapper makes the model your glue code: it chains the calls, carries ids from one result to the next, and performs the joins by reading. Each step costs a model pass and context, and each is a chance to misread.

The chapter’s first section opens on this exact exercise, a calendar API of three endpoints wrapped one for one, and calls it “a trap worth studying.” The agent finishes the scheduling job “the way you would compute a spreadsheet in your head. It is possible, expensive, and wrong just often enough.”

The same scheduling job in two action spaces.
Figure 5.2 The same scheduling job in two action spaces. With one endpoint wrapped per tool (top), every intermediate byte crosses the model’s desk and every cross-reference risks a misread. Shaped to the task (bottom), the deterministic work happens inside the tool and the model sees one line back. Chapter 7 prices the difference. Reuse this diagram

A 2025 playbook from a company that runs many tool servers gives the alternative in one line: “Design top-down from workflows, not bottom-up from API endpoints.”

None of this says your API is badly designed. It is correct design aimed at a different caller. The table is my synthesis of where API habits carry over, where they mislead and where they are not enough alone.

API habit For an agent caller Why
Small orthogonal endpoints the client composes Misleads The client is the model; composing costs a pass and context per step
Return the complete resource Misleads Every field is read and paid for on every later pass
Status code plus machine-readable error Not enough alone Keep the code and a retryable flag for your harness; add the words the model needs: what failed, what to try
Flexible generic parameters (filter, options, data) Misleads Flexibility reads as ambiguity; one name, one meaning
Opaque stable ids Partly misleads Keep them for calls, and lead with handles a model can copy
Strict input validation Transfers, more so This caller invents values a compiled client never would
Idempotent writes Transfers, more so Agents retry constantly
Consistent naming Transfers Names are how the model tells neighbors apart

Worked example: one billing API wrapped two ways

Here is one job done through both shapes, with the calls counted. A support agent has a ticket from [email protected]: “I was charged twice for my last order, please refund the duplicate.” The API has customers, orders, charges and refunds.

Everything below is illustrative, counted by hand for this post on a scripted happy path where every call succeeds first time. None of it is a token or accuracy measurement.

Six endpoint-shaped tools

Tool Wraps Used in this job
get_customers(q, page) search customers yes
get_customer(id) one customer no
list_orders(customer, status) a customer’s orders yes
get_order(id) one order, with line items and charges yes
list_refunds(order) refunds on an order yes
create_refund(charge, amount, reason) create a refund yes
Step Call What lands in the context
1 get_customers(q="[email protected]") a page of customer objects
2 list_orders(customer=…) every order object for that customer
3 get_order(id=…) the full order, line items and both charges
4 list_refunds(order=…) refund objects, to check nothing was refunded already
5 create_refund(charge=…, amount=…, reason=…) the refund object

That is five tool calls, so six model passes once you add the pass that writes the reply. The model spots the duplicate by reading amounts and timestamps in step 3, and it copies three identifiers (customer, order, charge) from a result into a later call.

Two task-shaped tools

Tool Kind What happens inside, in code
find_refundable_charges(customer_email) read finds the customer, walks orders, charges and refunds, flags likely duplicates, returns at most ten short rows, newest first
refund_payment(charge_ref, amount_cents, reason, confirm) write checks the charge belongs to the customer on the ticket, which the harness supplies and the model cannot change; checks the charge is refundable; derives the idempotency key from the run, the step and the arguments; issues the refund; returns its reference and what is left to refund
Step Call What lands in the context
1 find_refundable_charges(customer_email="[email protected]") up to ten rows, one flagged looks_like_duplicate
2 refund_payment(charge_ref="CH-2291-B", reason="duplicate", confirm=true) after the harness gate approves: refund reference, amount, amount still refundable

The count, side by side

Counted on the scripted happy path Endpoint-shaped Task-shaped
Tools defined 6 2
Tool calls 5 2
Model passes (calls plus the final reply) 6 3
Definition loads (tools × passes, at equal definition length) 36 6
Results the model reads before the write 4 1
Identifiers the model copies between calls 3 1
Where the duplicate is detected by the model, reading in code

Neither column counts an approval. If your harness gates the refund, the call pauses for a person and no model pass is added; if the model asks in conversation first, add one pass to each column.

Two cautions on the first and fourth rows. The six tools cover the whole API and the two cover this one job; an agent with five jobs carries more than two. And a task-shaped definition is longer: the repaired refund_payment below is about eight times the size of the bare create_refund (1,502 against 184 characters of minified JSON, which is what I measured). Count tokens, not definitions, before claiming a saving.

The pagination, the join across four resources and the already-refunded check have moved into ordinary code. In the chapter’s words, “The right unit for a tool is a task; the endpoints are plumbing, and plumbing belongs inside.”

Why is the refund still its own tool?

Because the refund is the step you want to gate, log by name and make safe to repeat, and a single handle_refund tool would hide that side effect behind a lookup. The read side collapsed from four calls to one. The write did not move.

The rule I use has two parts, and it is this post’s extension of the chapter, which does not state one. First, one risk level per tool: a tool either never changes anything, or it is a write and is named as a write. The same playbook cited above makes that cut from the permissions side: “each tool should stick to a single risk level.” Second, split the lookup from the write when a choice sits between them, meaning a judgment or an approval that depends on what the lookup found.

Two cases show how it cuts. In the refund, someone should confirm which charge is the duplicate, so the lookup and the write are split. The chapter’s own schedule_event reads calendars and books the slot in one tool, and passes: it is named as a write, and any free shared half hour will do, so no choice sits between lookup and booking.

A separate write can also carry an idempotency key, derived in code so the model never touches it. The post on idempotent tools and safe retries covers keys, receipts and retry policy.

How many tools is too many, and why does the agent pick the wrong one?

No fixed number exists. Selection gets worse as the set grows, and how fast depends on the model. “I tried doing the MCP approach with about 100 tools, but the agent picks the wrong tool a lot of the time and it seems to have gotten significantly worse the more tools I added,” one developer wrote in August 2025.

The measurements that exist, all dated and none permanent:

  • Catalog size. A 2025 preprint on long-context function calling (Kate and colleagues) reports “a performance drop of 7% to 85% as the number of tools increases,” depending on the model, as the catalog grew from 8K to 120K tokens of definitions and the right tool’s position in it moved.
  • Loading definitions on demand. One vendor’s internal test (2025) on large tool libraries: accuracy rose from 49% to 74% for one frontier model and from 79.5% to 88.1% for a later model from the same vendor. The stronger model gained less from the fix.
  • What definitions cost. The same vendor’s 2025 example of five connected tool servers: “58 tools consuming approximately 55K tokens before the conversation even starts.” By my arithmetic that is about 950 tokens per definition on every pass, an illustration and not a constant.

Two model providers’ documentation, as read in October 2026, gives soft ceilings. One says to aim for “fewer than 20 functions available at the start of a turn,” adding “this is just a soft suggestion.” A second says to keep the active set to “10-20 tools maximum.”

These are vendor statements at a date, not measurements. The chapter’s rule will outlast them: “keep the tool set small, task-shaped, and unambiguous, and treat every addition as a cost that must argue for itself.”

Which levers reduce wrong picks?

Four, in the order I would try them:

  1. Merge or separate overlapping tools. Two tools with similar names and jobs “present a choice where none should exist,” as the chapter says. Merge them, or state the boundary in both descriptions.
  2. Namespace what remains. Prefixes by service and resource, such as the chapter’s crm_contacts_search and billing_invoices_search, let the model read kinship in the names. One vendor’s engineering guide (Aizawa, 2025) reports that prefix against suffix had “non-trivial effects” on its evaluations, varying by model, so test the scheme.
  3. Load fewer definitions. Give each kind of task only the tools it needs, or delegate a group of tools to a subagent with its own context.
  4. Keep the set fixed within a run. One practitioner report from 2025 (Ji) advises: “unless absolutely necessary, avoid dynamically adding or removing tools mid-iteration.” So choose the set when a run starts; that last step is my judgment, not a result.

A tool protocol changes how definitions reach your agent and not what the model reads, so all of this holds on either side of the MCP versus function calling choice.

How do you write the description and schema so the arguments come out right?

Write the description as onboarding for a capable newcomer, and write the schema so each argument has exactly one reasonable reading. The chapter’s warning is that “Ambiguity a human colleague would resolve with one glance at the codebase is, for the model, permanent.” It is also the cheapest surface to fix.

A flawed definition, linted

Take the write from the endpoint set, as a spec-to-tool generator would produce it. The key that holds the schema is spelled differently from one provider to the next; the shape is the same.

{
  "name": "create_refund",
  "description": "Create a refund.",
  "parameters": {
    "type": "object",
    "properties": {
      "charge": {"type": "string"},
      "amount": {"type": "number"},
      "reason": {"type": "string"}
    }
  }
}

The embed below is this site’s tool schema linter, loaded with that definition. It is pattern matching over the three parts the model sees, and nothing more.

With JavaScript on, the Tool schema linter runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Tool schema linter on its own page to share a result by link.

It returns 12 findings: 1 to fix first, 7 that should be fixed and 4 to consider. The fix-first finding is that the description only restates the name. The seven: no “use when,” no “returns,” nothing marked required, three of three arguments undescribed, nothing that makes a retry safe, a destructive action with no confirmation, and a name that “names an action but not what it acts on.”

That last one is a false alarm. The linter’s word list reads “refund” as a verb, so it sees two verbs and no object. The four to consider: no “do not use,” no account of failure, extra properties allowed, and no word on whether the action can be undone.

The repair, slot by slot

Each finding maps to a sentence or a schema keyword. Three repairs go further than the linter asked: it did not object to charge, amount or a free-text reason, yet charge_ref says which identifier, amount_cents rules out a unit mistake, and an enum on reason closes a set the model would otherwise fill with prose.

{
  "name": "refund_payment",
  "description": "Refunds one charge to the customer's original payment method. Permanent: a refund cannot be undone. Use when the customer has asked for money back on a specific charge and you hold its charge_ref from find_refundable_charges. Do not use to look charges up or to cancel an order; call find_refundable_charges first. Returns the refund reference, the amount refunded in cents and the amount still refundable on the charge. Safe to retry: repeating this call within the same run returns the existing refund, flagged as pre-existing, and moves no more money. On error returns a code, a message, whether a retry can help, and a hint. Example: already_refunded, not retryable, nothing is left to refund on this charge. If you do not know the charge_ref, do not guess. Ask the user which charge they mean.",
  "parameters": {
    "type": "object",
    "properties": {
      "charge_ref": {
        "type": "string",
        "description": "The charge to refund, copied exactly from find_refundable_charges. Example: CH-2291-B."
      },
      "amount_cents": {
        "type": "integer",
        "minimum": 1,
        "description": "Amount to refund, in cents. Example: 4900 for 49.00. Omit to refund everything still refundable."
      },
      "reason": {
        "type": "string",
        "enum": ["duplicate", "customer_request", "fraud", "other"],
        "description": "Why the refund is being issued."
      },
      "confirm": {
        "type": "boolean",
        "description": "Must be true. Set it only after the user has approved this exact refund."
      }
    },
    "required": ["charge_ref", "reason", "confirm"],
    "additionalProperties": false
  }
}

Paste that into the linter and it returns no findings. Keep the old name on the repaired body and one remains, the same false alarm. A model-set confirm flag is a reminder to the model; the approval that counts is enforced in your harness. Letting an omitted amount_cents mean “everything” suits this job, a full duplicate; for most writes the safer default is to require the amount.

Now the limit that matters most for this post. The linter reads the name, the description and the schema of each tool, plus a few checks across the set. One of those looks for one-for-one wrappers, and it needs three operations on one resource before it speaks; this set never gives it that. So it did not see that the tools are endpoint-shaped.

I ran all six endpoint-shaped tools through it as a set, each with a bare one-line description. It returned 60 findings, one of them about the set (an overlap between get_customers and get_customer) and none saying the set mirrors an API, although the linter has a check for that.

A description template to copy

The description has six slots, the same six the linter’s own skeleton prints. Side effects go in the first slot, and the last two lines of the block are for the arguments.

Does: <one sentence: the task this tool performs and what it changes, if
      anything. For a write: can it be undone, and is a repeat safe?>
Use when: <the situations, in the words a user would use>.
Do not use when: <the near-miss cases>; use <neighbor_tool> instead.
Returns: <the fields the next step needs, by name>. For lists: at most <N>
      items, <sort order>. If more exist, the result says so and how to narrow.
On error: returns <code>, <message>, <whether a retry can help>, <hint>.
      Example: <code>, <retryable or not>, <what to try next>.
If unknown: If you do not know <required value>, do not guess. Ask the user,
      or stop and report if no one is there.

Each argument: <what it is>, <format or unit>, <one example>.
      A closed set of values is an enum. What the call needs is required.

Longer is not free. A 2026 preprint (Hasan and colleagues) examined 856 tools across 103 servers of one tool protocol; a model-based scanner found that “97.1% of the analyzed tool descriptions contain at least one smell.” Augmenting every component of the descriptions improved task success “by a median of 5.85 percentage points,” and it also “increases the number of execution steps by 67.46% and regresses performance in 16.67% of cases.” Fuller descriptions help on balance, and each rewrite needs checking.

What does a strict schema buy, and what doesn’t it?

A strict schema removes malformed calls and leaves wrong ones. With constrained decoding, where the model can only emit output that fits the schema, a required field cannot be omitted and an enum cannot receive an invented value; that is the stronger grade of structured output for tool calls, and the weaker one only asks. That makes the schema the cheapest mistake-proofing you have: anything expressible as a type, an enum or a required list stops being a description problem.

It doesn’t make the call true. A conforming call can carry a customer id that was never issued. Chapter 2 has the sentence to remember: “The schema polices the format; you remain the auditor of the content.” Validation of meaning stays inside the tool, and what the tool says when validation fails is the subject of the error section below.

What should a tool return so the result doesn’t eat the context?

Return what the next decision needs, in a bounded size, with identifiers the model can act on. A client picks fields out of a response body. The model reads all of it, and it stays in the context window for every later pass.

An issue filed against one tool server in April 2025 shows the fault at full size: “When list_commits() returns 30 results, the follow-up LLM call will use >64k tokens.”

The 2025 long-context preprint cited above reports “a 7% to 91% degradation in answer retrieval as the tool responses length increases,” depending on the model, as responses grew from 10K to 80K tokens. And in the 2024 paper that named the agent-computer interface (Yang and colleagues), one model on a 300-task coding subset resolved 18.0% of tasks with a 100-line file viewer, 12.7% when shown the whole file and 14.3% with 30 lines. More was worse, and so was too little.

Four rules cover most tools:

  • Return the slice. Ask, with the chapter, “what informs the next action?” and send those fields. In Aizawa’s one example, a concise response format took 72 tokens against 206 for the detailed one.
  • Cap with a default, and say when you cut. The chapter mentions one production agent that caps every response at twenty-five thousand tokens, “a figure I pass along as illustration only.” Whatever your cap, announce it, because “a silent cut reads as a complete answer.”
  • Page, or hand back a handle. Offer a filter and a cursor. For bulky data, return a file path or a reference the next tool accepts.
  • Lead with identifiers a model can copy. “Prefer fields a language model can grip: names over UUIDs, file types over MIME strings, readable labels over hex identifiers.” After a write, return the new record’s handle; a bare “done” strands the next step. A short handle is easier to mis-copy into another real record than a UUID is, so the tool must check that the handle belongs to the customer or the run it was issued for.

What should a tool return when it fails?

An error should be a short instruction the model can act on: what went wrong, what a correct call looks like, what to try next, and whether a retry can help. This is where backend instincts fall shortest. A status code is written for a client that branches on it. This caller branches on words.

The chapter treats the error as a design surface: “An error arrives at the exact moment the model does not know what to do next; it is the highest-signal message your tool will ever send, and most tools waste it.” It adds, “Error 422 teaches nothing.”

Situation What a backend returns by habit What the model can act on
Bad argument 422 Unprocessable Entity invalid_reason: reason must be one of duplicate, customer_request, fraud, other (got “double charge”). Not retryable as sent; resend with a valid reason.
Business rule 409 Conflict already_refunded: charge CH-2291-B was refunded in full on 3 June. Not retryable. Nothing is left to refund; tell the customer.
Unknown identifier 404 Not Found charge_not_found: no charge CH-2291-8 for this customer. Not retryable as sent. Do not guess another reference; call find_refundable_charges and copy one from the result.

Each cell in the right-hand column says what happened, whether a retry can help and what to do next. Where a correct call exists, as in the first and third rows, it says what one looks like. The examples are mine and illustrative. Error text enters the model’s context and may be repeated to the person on the other end, so keep stack traces, queries, hostnames and other customers’ data out of it.

Two plumbing faults undo all of this. The first is a harness that raises on a tool error, which ends the run. An issue filed in October 2026 against one agent library describes the result: “The model never sees the error, so it can’t correct itself.” The second is a failure returned in the success channel with no flag, which the model reads as data.

One tool protocol, taken as an example, writes the principle into its specification: the 2025-11-25 revision of the Model Context Protocol says tool execution errors “contain actionable feedback that language models can use to self-correct and retry with adjusted parameters,” and that clients should pass them to the model. A timeout on a write is its own case, the ambiguous failure, and the idempotency post linked above owns it.

Few powerful tools or many narrow ones?

In my reading both camps are right about different work, and the dividing line is the write; the chapter leaves the crossing point to testing. A general tool, such as a shell or a code interpreter running in a sandbox, suits exploration and bulk data. A named narrow tool suits an action you must permission, log and make repeatable.

One infrastructure team’s 2025 write-up argues that models have read far more code than tool-call transcripts: “LLMs are better at writing code to call MCP, than at calling MCP directly.” Another team removed 80% of its data agent’s tools in 2025 in favor of command execution over well-documented files, and reported 5 of 5 test queries passing against 4 of 5 before. That is five queries on one internal agent.

The general tool breaks in three places. A sandbox is infrastructure you now run. Results depend on what the tool finds; over a messy data layer, the same post warns, “You’ll just get faster bad queries.” The chapter names the third: “a general tool offers no semantic guardrails, demands more skill from the model, and is harder to permission and audit.”

A raw shell command cannot be recognized as a refund. Code execution need not lose that. In the infrastructure team’s design the model’s program calls named tool functions, and the chapter describes the same thing, code “calling your other tools as ordinary functions inside the program”; a refund issued that way is still a call your harness can recognize, gate and log.

The curated set breaks elsewhere. One step off the menu and the model is back to composing by reading.

My reading, which none of these sources states: use one general read-and-compute capability where you can sandbox it, and keep every action that changes something a named tool, whether the model calls it directly or from code. The chapter leaves the crossing point to “measurement on your tasks.”

Which design fault matches the symptom in your transcript?

Find the symptom, then fix the interface before the prompt. The chapter’s debugging order is “when a model misuses a tool, suspect the contract before you blame the intelligence.” The table has fifteen rows under five symptoms; the buttons narrow it to yours.

What you see in the transcript Design fault Fix Symptom
Calls search_x when find_x was right, or alternates between them Two tools overlap; neither description names the other Merge them, or state the boundary in both descriptions wrong tool
Uses a tool on turns where no tool was needed The tool is in the set but irrelevant to most tasks Remove it, or load it only for the tasks that need it wrong tool
Wrong picks rose after tools were added Set too large for this model; names do not group Cut to the task-shaped few; namespace by service and resource wrong tool
Invents a tool name that does not exist Many similar names; the model pattern-completes Fewer, more distinct names; return an error listing the valid names wrong tool
Passes a name where an id was wanted Parameter named user, id or data Rename it to say one thing (user_id); give the format and an example wrong arguments
Sends a value outside the allowed set Free string where a category was meant Enum; reject with the valid set in the error wrong arguments
Makes up a required value No instruction for the unknown case Escape hatch in the description: do not guess, ask the user wrong arguments
Relative path, wrong unit, wrong time zone The argument admits a mistake the tool could rule out Require the unambiguous form: absolute path, ISO date, cents wrong arguments
Five calls and a join done by reading Endpoint-shaped tools; the model is the glue code One task-shaped tool that runs the chain in code too many calls
Same list call repeated with small variations No filter, or the page size is too small Add the filter the task needs; size the page from real use too many calls
One result consumes most of the window The tool returns the whole object or the whole list Return the slice the next step needs; cap it with a default context flood
Treats a truncated result as complete Silent truncation Say it was cut and how to narrow context flood
Retries the same failing call with tiny changes The error says nothing actionable Error as instructions: what failed, valid values, what to try failed calls
Run aborts on a bad argument The harness raises on a tool error Return tool errors to the model as results failed calls
Reports success after a failed call Failure sits in the success channel with no flag Mark errors as errors; put the reason first failed calls

The table is my synthesis; each fix traces to the chapter or to a source cited in this post. If your AI agent calls the wrong tool, start with the first four rows.

How do you test whether the redesign worked?

Run the same small set of tasks before and after, and compare the pass count and four counters. None of the four pages I compared on how to design tools for LLM agents says what to count, yet the chapter’s standard for a tool set is evidence from “the record of the agent actually working.”

  1. Write five to ten realistic tasks from real requests, each needing more than one call. Hold two back that you never tune against.
  2. Run each a few times with the current tools and record per run, first, whether the task ended in the right state, checked against the system and not against the agent’s reply. Then the four counters: calls to a tool the task did not need or to the wrong one of two similar tools; calls rejected for invalid arguments; tool calls per task; tokens per task, definitions and results included.
  3. Read the transcripts whole before changing anything, and tag each bad step with a row of the symptom table.
  4. Change one surface at a time (the set, then descriptions, then results, then errors) and rerun.
  5. Run the held-back tasks once at the end. If they disagree with the tuned set, you fitted the tools to your examples.
  6. Repeat when the model changes, and try removing a tool as well as adding one.

A redesign that lowers the four counters and lowers the pass count is a worse design. With five to ten tasks and a few runs each, a difference of one or two calls is noise; look for a counter that moves on most tasks, and put any pass-rate difference through the eval sample size calculator before believing it.

The worked example shows the limit of a paper count. Its scripted path fixes calls per task at 5 and 2; the pass count and the other three counters appear only when a real model runs the tasks. So I can promise the procedure and not a size of improvement. For the token side, the context window budget planner shows what share of the window your tool definitions take.

Where do task-shaped tools cost you, and what is not measured?

They cost code, flexibility and upkeep, and the evidence for them is thinner than the confidence of the guidance suggests.

You own more code. The join the model did by reading is now a function you write, test and maintain. That suits a job you see every day and not one you see twice a year.

A task tool encodes a guess about the tasks. When the request steps off the menu, the agent has nothing to reach for. Keep one general read path for those cases, or accept the refusal.

A tool can outlive its usefulness. “A tool that rescues today’s model can constrain next year’s,” the chapter warns. “Revisit the set on every model change, and remember that subtraction is a legitimate design move.”

The measurements are narrow. Every percentage here describes particular models on particular tasks in the stated year, and the two gaps named at the top of this post remain open.

How to design tools for LLM agents, in one sentence

If you are about to wrap forty endpoints as forty tools, list the five jobs the agent exists to do instead. Collapse each job’s read chain into one tool. By the rule I use above, name every tool that changes something as a write, and split it from the lookup when a judgment or an approval sits between the two. Then run the tasks and count what changes.

The rest of the craft, from retrofitting command-line programs to code execution, is in Chapter 5, Tools and the Action Space (in the full book). The free guide to tools, skills and protocols collects the related posts and tools, and you can see the formats.

Questions readers ask

How many tools is too many for an LLM agent?
There is no fixed number. Two model providers' documentation, as read in October 2026, suggested fewer than 20 tools at a time, or an active set of 10 to 20, both as soft guidance. Published measurements show accuracy falling as tool catalogs grow. Count wrong-tool calls on your own tasks and merge or remove tools until they stop.
Should I expose my REST API to an agent as one tool per endpoint?
Usually not. Endpoints are shaped for code that composes them for free. A model pays one model pass and some context for every step, and does the joins by reading. Wrap the task in one tool, keep the endpoints inside it, name every tool that changes something as a write, and split it from the lookup when a judgment or an approval sits between the two.
Why does my AI agent call the wrong tool?
Most often because two tools overlap and neither description names the other, because the set is too large for the model, or because a description says what the tool is and not when to use it. Check the tool definitions before blaming the model.
What makes a good tool description?
Six things: what the tool does and changes, when to use it, when to use a neighbor instead, what it returns and how results are capped, how it fails, and what to do when a required value is unknown. Each argument also needs its meaning, its format and one example.
What should a tool return when it fails?
A result the model can read and act on: what went wrong, what a correct call looks like, what to try next and whether retrying can help. A bare status code or a stack trace gives the model nothing to correct, so it guesses or repeats the call.
Does structured output fix wrong tool arguments?
It fixes malformed ones. A strict schema mode can stop a missing required field or an invented enum value. It cannot tell a real customer id from a made-up one, so the tool still has to validate meaning and return an error the model can act on.

Sources

  1. Ken Aizawa (Anthropic) (2025). Writing effective tools for agents, with agents
  2. John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, Ofir Press (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
  3. Kiran Kate, Tejaswini Pedapati, Kinjal Basu, Yara Rizk, Vijil Chenthamarakshan, Subhajit Chaudhury, Mayank Agarwal, Ibrahim Abdelaziz (2025). LongFuncEval: Measuring the effectiveness of long context models for function calling (preprint)
  4. Mohammed Mehedi Hasan, Hao Li, Gopi Krishnan Rajbahadur, Bram Adams, Ahmed E. Hassan (2026). Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions (preprint)
  5. Anthropic (2025). Introducing advanced tool use on the Claude Developer Platform
  6. OpenAI (2026). Function calling (documentation; one example of a model provider's guide, read October 2026)
  7. Google (2026). Function calling, Gemini API (documentation; a second example of a model provider's guide, read October 2026)
  8. Block Engineering (2025). Block's Playbook for Designing MCP Servers
  9. Model Context Protocol (2025). Model Context Protocol specification, revision 2025-11-25: Tools (one example of a tool protocol)
  10. Yichao 'Peak' Ji (2025). Context Engineering for AI Agents: Lessons from Building Manus
  11. Andrew Qu (Vercel) (2025). We removed 80% of our agent's tools
  12. Kenton Varda, Sunil Pai (Cloudflare) (2025). Code Mode: the better way to use MCP
  13. github/github-mcp-server issue tracker (2025). Issue #142: Excessive context size for list_commits (reported April 2025)
  14. openai/openai-agents-python issue tracker (2026). Issue #5304: Sandbox `exec_command` / `view_image` abort the whole run on bad model arguments instead of returning a tool error (reported October 2026)
  15. Hacker News (2025). Hacker News comment on tool selection with about 100 tools (August 2025)