Home / Blog / Tools, skills and protocols / Why Your AI Agent Calls the Wrong Tool (and Six…

Tools, skills and protocols

Why Your AI Agent Calls the Wrong Tool (and Six Fixes)

When your AI agent calls the wrong tool, the record shows which of six causes it is. Classify the call, run five checks, apply one fix and measure it.

By Enrique Gutiérrez · Published · 19 min read

When an AI agent calls the wrong tool, the cause is one of six: the tool definitions drifted, the tool the task needs is missing, the model completed a name nobody defined, stale context pointed the wrong way, the set is too large, or two tools overlap. Each leaves different evidence in the record of the run, and each has a different fix.

That matters because the reflex is to reword the system prompt, and rewording addresses one of the six. This post is for the engineer holding a transcript with a bad call in it. It gives four symptoms to sort the call into, five checks that separate the causes, one fix per cause, and a way to measure that the fix worked.

The book this site belongs to names the caller you are dealing with in the opening of Chapter 5: a tool is a contract with “a non-deterministic caller, one that may call the right tool with wrong arguments, call the wrong tool with perfect syntax, or invent, in a confident flourish, a tool that has never existed.”

The six-way split and the five checks are this post’s own arrangement. The causes trace to Chapter 5, Tools and the Action Space (in the full book) and Chapter 15, Observability and Debugging (in the full book) and to the sources listed at the end.

Which of four symptoms is in your transcript?

A bad tool call is one of four symptoms, and you can tell them apart from one model request and its reply. Compare the name the model wrote with the list of tool definitions sent in that same request, then with what the step needed.

Symptom Test against the record
Wrong tool The name is on the list that was sent. A different tool, or no tool, was right.
Wrong arguments The name is on the list and is the right tool. One or more values are wrong.
Invented tool The name is not on the list that was sent.
No call The reply is prose, and the step needed a tool that was on the list or should have been.

The test needs the list as it was sent, which is not the same as the list in your repository. Chapter 15, in the section “Traces and Spans”, quotes a debugging guide on the difference for prompts: “The rendered prompt is what the model actually saw. The template plus variables is what you intended. Bugs live in the difference.”

The same holds for tool definitions. If your trace does not carry them, the post on agent observability covers what to record.

Research papers use a similar split. A 2024 preprint (Xu and colleagues) sorts “tool hallucinations into two main types, tool selection hallucination and tool usage hallucination”, which correspond to choosing the tool and filling its arguments. An earlier benchmark, MetaTool (Huang and colleagues, 2023), separates deciding whether to use a tool at all from deciding which one.

I use four symptoms because the invented name and the missing call have their own causes. It also helps to treat “answer without a tool” as one more option the model can choose. A missed call is then a confusion between a tool and that option, and it goes through the same checks as any other wrong pick.

Wrong arguments are the symptom this post covers least. Two of the six causes produce them, and the table below shows which. The rest are faults in a single tool’s contract, covered in the posts on how to design tools for LLM agents and structured output for tool calls.

What are the six causes when an AI agent calls the wrong tool?

Six causes account for the runs where an AI agent calls the wrong tool, and they differ in where the fault sits: in the definitions as deployed, in the set of tools, in the context of one step, or in the wording. The order below is the order in which they are cheapest to rule out.

The anatomy of a tool across the boundary between your process and the model’s context.
Figure 5.3 The anatomy of a tool across the boundary between your process and the model’s context. Three parts cross to the model’s side and steer it like prompt text—the description above all, in accent, the highest-leverage surface you write. The function itself never leaves; only its results and errors travel back, as still more text the model has to read. Reuse this diagram

1. Definition drift. The definitions the model read are not the ones your code, your prompts or your skills assume. A tool was renamed, a parameter changed, or a server someone else runs changed its list.

2. A missing tool. No tool on the list does what the step needs. The model picks the nearest one, answers from memory, or writes a name for the tool it wishes it had.

3. A hallucinated name. A tool that does the job exists, and the model wrote a different, plausible name for it. Nothing in the harness checked.

4. Stale context. The definitions are fine, and something earlier in the run points the wrong way: a superseded result, an old error, a pattern of earlier calls.

5. Too many tools. The right definition is present and buried. The same step succeeds when fewer definitions are sent.

6. Overlapping tools. Two options fit the request as described, and the descriptions do not say which one owns it.

Only the last of these is mainly a wording problem. Drift, a missing tool and an unchecked name are faults in the system around the model; stale context and set size are faults in what one step was shown. That agrees with Chapter 15’s reading of its own bug list, in “A Taxonomy of Common Agent Bugs”: “these are systems bugs, not model bugs, and a trace catches them in the act.”

How do you tell the six causes apart?

When an AI agent calls the wrong tool, run five checks in order and stop at the first that finds something. The first three read the record. The last two replay the single failing step, which is cheaper than rerunning the task and isolates one decision.

A replay here means sending the model one request again: the same model, the same system prompt, the same tool definitions, and the input of the step that went wrong. The model is non-deterministic, so one replay proves nothing. Send it ten times and write down a rate, such as “wrong tool in 7 of 10”. The post on how to debug an AI agent covers reproduction in general; this is the narrow version for one tool choice.

  1. Diff the definitions. Compare the tool list sent in the failing request with the list in the last good run and with the handlers deployed. Search your prompts, skills and tool error messages for tool names, and check each one against the list. Any mismatch is cause 1.
  2. Ask whether any listed tool does the job. Write down what the step needed in one sentence, then look for a tool that does it. If none does, it is cause 2, and the model’s pick was the nearest neighbor.
  3. For an invented name, look for the real one. If checks 1 and 2 are clean and a listed tool does the job under a different name, it is cause 3.
  4. Replay the step with a clean history. Keep the definitions; replace the accumulated history with the user’s request and only the facts the step needs, stated correctly. If the wrong-pick rate collapses, the history was the cause: cause 4.
  5. Replay with a short list. Keep the clean history; send only the right tool, the tool that was picked, and up to three others. If the rate collapses, it is cause 5. If the model still splits between the same two options, it is cause 6.

Two notes on reading the result. A rate from ten replays separates 7 of 10 from 1 of 10; it does not separate 4 of 10 from 2 of 10, so run more when the difference is small. And causes stack: a large set makes an overlap worse. Fix the first cause the checks find, then run them again.

Why not ask the model why it chose?

Because the explanation is generated text, written after the fact, and it is weak evidence. One vendor’s engineering guide to writing tools (Aizawa, 2025) puts it in one line: “LLMs don’t always say what they mean.” Read the reasoning for hints about which description misled the model. Let the replay rates decide.

Which evidence points to which fix?

The table maps what the record shows to a cause and its fix. It has ten rows; the buttons narrow it to your symptom. Fix numbers match cause numbers, and each fix is described in the next section.

What the record shows Settled by Cause Fix Symptom
The called name is not on the list sent, and a prompt, a skill or a tool error message mentions that name Check 1 1. Definition drift Version the definitions and test every mentioned name invented tool
The called name is not on the list, and no listed tool does that job Check 2 2. Missing tool Add the tool or state the exit invented tool
The called name is not on the list, and a listed tool does the job under a similar name Check 3 3. Hallucinated name Check names in code; return the valid list invented tool
The expected tool was not in the request, or its description differs from the last good run Check 1 1. Definition drift Version the definitions and test every mentioned name wrong tool, no call
No listed tool does the job; the agent picked the nearest one or answered from memory Check 2 2. Missing tool Add the tool or state the exit wrong tool, no call
The clean replay picks right; the failing step’s input holds a superseded result or an old error Check 4 4. Stale context Date results and retire superseded ones wrong tool, no call, wrong arguments
The clean replay is still wrong; the short-list replay picks right Check 5 5. Too many tools Send each task only the definitions it needs wrong tool, no call
The short-list replay still splits between the same two options Check 5 6. Overlapping tools Merge, or state the boundary in both descriptions wrong tool, no call
The arguments would have been valid under an earlier version of the schema Check 1 1. Definition drift Version the definitions and test every mentioned name wrong arguments
Right tool, current schema, and a value with no source in the request or in any result None of the five Not a selection fault: an invented or malformed argument The argument rows of the tool-design post wrong arguments

Three reported cases show the table at work. In the first, a developer wrote in March 2026: “we had a support agent that would sometimes call the”refund order” tool when the user just wanted to check order status. the tool worked perfectly, the LLM just kept picking the wrong one.”

The name is on the list, so the symptom is a wrong tool. A status tool exists, so check 2 is clean. The comment does not say what a replay showed. If the clean replay and the short-list replay both stay wrong, the table gives cause 6 and the fix is a boundary stated in both descriptions.

In the second, an issue filed in October 2026 against one tool server reports that 26 of its authentication error messages “tell the model to use”the ‘authenticate’ tool”, which no longer exists.” An agent that follows that message writes a name that is not on the list. Check 1 finds it: a name in tool text with no matching definition. That is cause 1, and no description edit would have helped.

In the third, an issue from September 2026 describes an agent that chose a general object-fetching tool over the search tool for “find a person named X”. The reporter calls it “a tool-selection/discoverability problem caused by the MCP tool description, not a bug in the search implementation or MCP transport.” Two listed tools, one request, a description that never mentions people: the evidence points to cause 6, and the short-list replay of check 5 would confirm it.

What are the six fixes, one per cause?

Each cause has one fix that belongs to it and five that do not touch it. The fixes are numbered to match the causes.

Fix 1: version the definitions and test every name that mentions them

Treat the tool definitions as a versioned artifact. Compute a hash of the list as sent, record it on every model call, and add a test that extracts each tool name from your prompts, skills and tool error messages and checks it against the live list.

Chapter 15 applies the same discipline to prompts and models in its account of silent degradation: check the responding model, the prompt version and the prompt hash, because “One of the three almost always differs.” Tool definitions are a fourth thing to pin. They are also the one most likely to be held by someone else, when the list comes from a server reached over a tool protocol; the post on what MCP is in AI covers that arrangement, with MCP as one example of such a protocol.

A second report shows how small the drift can be. One issue from October 2026 describes instruction files that hardcode a tool’s full name. When the same server is installed another way, the name gains a prefix, and the instructions then point “the agent at a tool that is not in its list; the agent searches, guesses or falls back to slower paths.”

Fix 2: add the tool, or state the exit

When no tool does the job, either build the tool or tell the agent what to do when nothing fits: say so, or ask. Which of the two depends on how often the gap appears, so log every hit.

An invented name is useful evidence here. One practitioner wrote in July 2025: “We do however take note of hallucinated tool calls and have had it suggest an implementation we start with and have several such tools in production now.” A name the model keeps writing is a description of the missing capability.

The exit matters as much as the tool. The 2024 preprint cited above proposes widening the action space with what its authors call “indecisive actions, allowing LLMs to defer tool use, seek clarification, or adjust tool selection dynamically”. That is a training method. The part that transfers to any agent is that declining needs to be an available move.

Fix 3: check every name in code and return the valid list

Before dispatch, compare the requested name with the registered tools. On a miss, run nothing, and return a result that says the name is unknown and lists the valid names. This is the tool-error advice of Chapter 5, Tools and the Action Space (in the full book) applied to the name: say what happened and what a correct call looks like.

This fix contains the failure; it does not stop the model from writing the name. Distinct, consistently patterned names lower how often it happens, and the tool contract linter flags naming problems in a pasted set. A check in code is what keeps a made-up name from ever reaching a handler.

Fix 4: date the results and retire the superseded ones

Stale context is repaired where the context is assembled. Have each result say when it was read, re-read before any step that acts on a value that can change, and remove or mark a result once a newer one replaces it.

Chapter 15 describes the mechanism under “Lost and poisoned context”: “an early wrong fact (a stale record fetched at step two, a mislabeled result) gets treated as ground truth forever after”. An old error works the same way. As an illustration: a tool fails once at step two with an authentication error, recovers, and the agent routes around it for the rest of the run because the error is still sitting in the history.

Fix 5: send each task only the definitions it needs

Reduce what is sent, by task type, by retrieval over the catalog, or by handing a group of tools to a subagent with its own context window. Chapter 5’s rule is to “keep the tool set small, task-shaped, and unambiguous, and treat every addition as a cost that must argue for itself.”

One measurement of the size effect: a 2025 preprint, RAG-MCP (Gan and Sun), retrieved the relevant definitions before calling the model and reports that this “more than triples tool selection accuracy (43.13% vs 13.62% baseline) on benchmark tasks”. That is one benchmark, one year and one set of models. The design post lists the other dated measurements and the levers in detail, so I do not repeat them.

Fix 6: merge the pair, or state the boundary in both descriptions

Overlapping tools, in Chapter 5’s words, “present a choice where none should exist.” Remove the choice by merging the two. If they have to stay separate, write the boundary into both descriptions, each naming the other: use this one when, use that one when.

The vendor guide cited earlier agrees on the cause: “When tools overlap in function or have a vague purpose, agents can get confused about which ones to use.” If the confused pair is a tool and the no-tool answer, the same fix applies to one description: state when to call it and when to answer without it.

How do you measure that the fix worked?

Build a tool-selection eval set from the requests where your AI agent calls the wrong tool, add the ones where it chose well, and read the results per pair of tools. An overall accuracy number moves for many reasons; the count for the pair you targeted moves for one.

A selection eval set is a list of cases, each a request with the tool that should be called first. Chapter 15 gives the composition: “for each tool, prompts that should trigger it and near-misses that should not”. Add cases where the right answer is no tool at all. The template is this post’s own.

TOOL-SELECTION EVAL CASE
id:                 sel-017
source:             trace id or ticket this case came from, or "written"
request:            the user's message, verbatim
history:            none | a fixed, recorded history (name the fixture)
definitions:        version or hash of the tool list this case runs against
expected:           one tool name | none | ask-the-user
also acceptable:    names that are not errors for this request, if any
must not call:      tools whose call here would do harm (writes, sends)
kind:               trigger | near-miss for <tool> | no-tool | missing-tool
cause when found:   1-6 from the triage table, or "not from a failure"
held out:           yes | no  (held-out cases are never used for tuning)
runs per case:      3 or more; record the called name for every run

Then tabulate expected against called. Give the table two extra columns: none, for runs that made no call, and invented, for names that were not on the list.

A worked example with illustrative numbers

The numbers below are invented for illustration and describe no real system. A support agent has five tools. The set has 60 cases, run 3 times each, for 180 runs.

Expected order_status order_search refund_create kb_search ticket_escalate none invented
order_status (36 runs) 24 10 0 0 0 2 0
order_search (36) 6 29 0 1 0 0 0
refund_create (24) 0 0 22 0 2 0 0
kb_search (36) 0 0 0 31 0 5 0
ticket_escalate (24) 0 0 1 0 21 0 2
none (24) 0 0 0 5 0 19 0

The diagonal sums to 146, so selection accuracy is 146 of 180 runs, or 81.1%, with 34 errors.

The errors are concentrated. The two order tools are confused with each other in 16 runs (10 plus 6), which is 47.1% of the errors. The knowledge-base search and the no-tool answer are confused in 10 runs (5 plus 5), another 29.4%. Two pairs hold 26 of the 34 errors, or 76.5%.

That reading already says which causes to suspect. Errors piled in one pair point to cause 6. Errors spread thinly across many pairs would point to cause 5. The two invented runs, 1.1% of 180, go to checks 1 to 3.

Suppose the fix is fix 6 on the order pair: one boundary sentence in each description. On the rerun, in this illustration, the two order rows change to 33 correct out of 36 each, with 4 confusions between them, and no other row changes. Accuracy rises to 159 of 180, or 88.3%, a gain of 7.2 points, and errors fall from 34 to 21.

The number tied to the fix is the pair count: 16 of the 72 order runs before (22.2%), 4 of 72 after (5.6%). The largest remaining pair is now the knowledge-base search against no tool, 10 of the 21 errors. That is the next thing to work on, and the overall figure alone would not have shown it.

What can go wrong with the measurement?

Four things. Three runs of one case are not three independent cases, so judge by how many cases changed, and put any difference through the eval sample size calculator or the post on how many eval examples you need before believing it.

Tuning descriptions against the same cases fits the wording to the cases. The vendor guide’s remedy is the standard one: “We relied on held-out test sets to ensure we did not overfit”. Keep cases that are never used for tuning.

A description edit can also make things worse. A 2026 preprint on 856 tools from 103 tool servers (Hasan and colleagues) found that augmenting descriptions “improves task success rates by a median of 5.85 percentage points” and also “increases the number of execution steps by 67.46% and regresses performance in 16.67% of cases.” Rerun the whole set after every edit.

Last, first-call selection is a narrow measure. It says nothing about whether the task ended in the right state. Keep an end-to-end check beside it.

Where does this diagnosis stop working?

The diagnosis stops where the record stops, and the fixes lower how often an AI agent calls the wrong tool without bringing the rate to zero.

It needs the request as sent. Without the tool list and the rendered input of the failing step, none of the five checks can be run as written: the first three read the list as sent, and the last two replay the input. The fix for that is tracing, and it comes before any of the six.

Replays cost money and prove rates. Ten replays of the step as it failed, ten with a clean history and ten with a short list is thirty model calls for one bug. That is cheap for a recurring failure and poor value for a one-off on a harmless read.

The split depends on the model. A set that is too large for one model is fine for another, so a model change can move a case from cause 5 to no fault at all, or the reverse. Rerun the selection set when the model changes.

A lower rate is still a rate. For a tool that writes, sends or pays, a wrong pick that happens rarely still happens. One developer described the exposure in April 2026: “any hallucinated tool call can fire off messages to an entire org while the agent is doing something unrelated.” Selection fixes reduce how often; a check in code decides what a wrong pick can reach. The post on AI agent guardrails covers that side.

Every figure cited here is dated. The percentages come from particular models on particular benchmarks in the stated years. The procedure is what should carry over to your system, and you will have to measure your own numbers.

The debugging order to keep

Chapter 5 gives the rule in “The Tool Interface”: “when a model misuses a tool, suspect the contract before you blame the intelligence.” For an AI agent that calls the wrong tool, I would widen the contract to include what is deployed and what the step was shown. Check the definitions as sent, then the set, then the name, then the history, and only then the wording.

The craft of the tools themselves is in Chapter 5, Tools and the Action Space (in the full book), and traces, replay and the bug taxonomy are in Chapter 15, Observability and Debugging (in the full book); the agent bug bestiary walks the same taxonomy as a short wizard. The free guide to tools, skills and protocols collects the related posts and tools, and you can see the formats.

Questions readers ask

Why does my AI agent call the wrong tool?
For one of six reasons: the tool definitions drifted from what the code or the prompts expect, the tool the task needs does not exist, the model completed a name that was never defined, an old result in the context points the wrong way, the set is too large, or two tools overlap. Read the request record and replay the failing step to see which.
How do I tell a hallucinated tool call from a wrong tool call?
Compare the name the model wrote with the list of tools sent in that same request. A name that is not on the list is an invented tool. A name that is on the list, where another tool or no tool was right, is a wrong tool. The two have different causes and different fixes.
Will a better system prompt stop wrong tool calls?
Sometimes, and for one cause only. Wording helps when two tools overlap or a description does not say when to use the tool. It does nothing for a missing tool, a renamed tool, a stale result in the history or a name the harness never checks. Find the cause first.
How do I test tool selection?
Build a set of cases, each a request with the tool that should be called, including cases where no tool should be called and near-misses for each tool. Run every case several times, tabulate expected against called, and read the errors per pair of tools. Keep some cases that are never used for tuning.
Does a strict schema mode prevent wrong tool calls?
It can keep the name and the arguments inside what the schema allows, where the interface offers it. It cannot make the choice among valid tools correct: a call to the refund tool on a status question fits the schema perfectly. Selection is tested, not constrained.

Sources

  1. Ken Aizawa (Anthropic) (2025). Writing effective tools for agents, with agents (published 11 September 2025)
  2. Tiantian Gan, Qiyao Sun (2025). RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation (preprint)
  3. Yue Huang, Jiawen Shi, Yuan Li, Chenrui Fan, Siyuan Wu, Qihui Zhang, Yixin Liu, Pan Zhou, Yao Wan, Neil Zhenqiang Gong, Lichao Sun (2023). MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
  4. Hongshen Xu, Zichen Zhu, Lei Pan, Zihan Wang, Su Zhu, Da Ma, Ruisheng Cao, Lu Chen, Kai Yu (2024). Reducing Tool Hallucination via Reliability Alignment (preprint)
  5. Mohammed Mehedi Hasan, Hao Li, Gopi Krishnan Rajbahadur, Bram Adams, Ahmed E. Hassan (2026). Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions (preprint)
  6. Hacker News commenter MickeyShmueli (2026). Hacker News comment on a support agent that picked the refund tool for a status question (4 March 2026)
  7. Hacker News commenter garfij (2025). Hacker News comment on keeping note of hallucinated tool calls (8 July 2025)
  8. Hacker News commenter yjcho9317 (2026). Hacker News comment on what a hallucinated tool call can reach (8 April 2026)
  9. gramps-web-mcp-rs issue tracker (2026). Issue #70: search tool description doesn't mention person or name lookups, causing agents to pick the wrong tool (reported 23 September 2026)
  10. outlook-assistant issue tracker (2026). Issue #275: tool errors don't set the error flag, and 26 auth messages point to a non-existent tool (reported 3 October 2026)
  11. agentic-board issue tracker (2026). Issue #758: flag skills that call tools by a name that does not exist when the server comes from a plugin (reported 1 October 2026)