Home / Blog / Tools, skills and protocols / What Is Tool Calling in LLMs? The Exchange, Ste…

Tools, skills and protocols

What Is Tool Calling in LLMs? The Exchange, Step by Step

What is tool calling in LLMs? An exchange: the model writes a request, your code runs the function, the result goes back. See who does each step.

By Enrique Gutiérrez · Published · 26 min read

What is tool calling in LLMs? It is an exchange: a large language model (an LLM) writes a request for your code to run a function, your code runs it or declines, and the result goes back to the model as more input. The model executes nothing. It reads text and writes text, and that is all it does.

The book this site belongs to gives the definition in two sentences: “The model never executes anything. It has no hands; it only writes.” (Chapter 2, section “Structured Output and Function Calling”, free to read online.)

Your own experience may seem to contradict that: a chat assistant searched the web for you, and no code of yours was involved. That was the same exchange with its middle step run on the provider’s servers, an exception covered below.

Every snippet in this post is in my own illustrative notation and follows no provider’s format; a card near the end is for writing down your interface’s names.

What is tool calling in LLMs, in one exchange?

Tool calling is five steps: your code sends a question with a list of tool definitions, the model writes a request to use one, your code runs the function, your code sends the result back, and the model writes an answer or another request. Steps 1 and 4 are calls to the model, and steps 2 and 5 are its replies. Step 3 is the only one in which anything happens.

A tool is a function you are willing to run on the model’s behalf. What you send the model is a tool definition: a name, a description in plain language, and a schema for the arguments (a declaration of which fields exist and what type each holds). Chapter 3 (“A Minimal Agent from Scratch”) counts “the ordinary function that actually runs” as a fourth part. Of the other three it says, “The first three travel to the model with every call”, and of the function: “The fourth never leaves your process.”

Function calling, step by step.
Figure 2.5 Function calling, step by step. Your code sends the question and the list of tools; the model replies not with prose but with a structured request to call one; your code runs the real function (in accent, step 3, the only place a side effect happens); the result goes back to the model, which then writes the final answer. The model proposes; your code disposes. Reuse this diagram

Three hosted providers recur below under fixed labels: the June 2023 provider, whose announcement that month brought function calling into its programming interface; the May 2024 provider, whose tool use left beta that month; and a third hosted provider. The terms table has the sources.

The numbering is the book’s. The June 2023 provider’s guide also lists “five high level steps”, the third being “Execute code on the application side with input from the tool call” (documentation, read 6 October 2026), and the third hosted provider’s shorter list heads its execution step “Execute Function Code (Your Responsibility)” (documentation, last updated 23 September 2026).

Who writes each message, and who reads it?

Your code writes steps 1 and 4 and the model reads them; the model writes steps 2 and 5 and your code reads them. Nothing in the middle is addressed to the user. The table is my arrangement of the book’s five steps.

Step What it is Who writes it What it contains Who reads it What can go wrong here
1 The call to the model Your code The standing instructions, the user’s words, and one definition per tool: name, description, argument schema. No function code. The model A tool is missing from the list, or its description does not say when to use it
2 The tool call The model A request that names a tool and gives arguments, or several such requests. Prose may come with it, or none. Your code No request; an unwanted one; the wrong tool; a tool that does not exist; invented arguments; a request that arrives as plain text
3 The execution (not a message) Your code Nothing is sent. Your code checks the name and the arguments, decides whether the call may run, runs each requested function and catches what it raises. Nobody An unchecked argument; a call that runs and should not have; an exception that ends the run
4 The result, in a second call to the model Your code The function’s output, or its error, usually as text, one per request and paired with the request it answers. The exchange so far goes with it, resent by your code or held by the server. The tool definitions go again too. The model The result is shown to the user and never sent; the model’s step-2 message is left out; the pairing is lost; the text carries someone else’s instructions
5 The answer The model Prose for the user, or another request, in which case the exchange returns to step 3 Your code, which shows prose to the user The answer ignores or misstates the result

A tool call is one request in the step-2 message, and that message is addressed to your code, which is why a first reply can come back with no answer in it. One Stack Overflow asker had got exactly that far: “This gets me to the point where I have LLM describe what function should be invoked – but what’s next?” (question 79253601, 5 December 2024).

On the day the June 2023 announcement was published, one forum commenter asked: “Where do you tell it where the API lives?” (Hacker News, 13 June 2023). Nowhere. The model never calls it, and step 1 has no line for an address.

The step-4 result is for the model. Chapter 2 names the mistake: “the classic first bug is to execute the tool and hand its raw result to the user, when the result must instead go back to the model (step 4)—the model alone knows why it asked and what to do next.”

What does one exchange look like, written out?

Written out, one exchange is a definition the model is shown, a request the model writes, a result your code returns, and whatever the model writes next. The example is a store’s support desk and a customer charged twice. The notation is mine and illustrative, and r1 stands for whatever your interface uses to pair a result with its request.

STEP 1  your code -> model
  instructions: You are a support assistant for a store.
  user:         I was charged twice for order A-1042. Can you fix it?
  tools:
    name:        get_order
    description: Look up one order by its id. Use it before you say
                 anything about what an order contains.
    arguments:   { order_id: text in the form A-0000, required }

    name:        get_payments
    description: List the charges made against one order.
    arguments:   { order_id: text in the form A-0000, required }

STEP 2  model -> your code
  prose:    (none)
  request:  r1  get_order { order_id: "A-1042" }

STEP 3  your code
  get_order is on the list; "A-1042" fits the form; a lookup is allowed
  runs the real get_order("A-1042")

STEP 4  your code -> model
  the exchange so far (resent by you, or held by the server),
  the tool definitions, again, and
  result for r1:  ok · order A-1042, one desk lamp, customer reports
                  being charged twice

STEP 5  model -> your code
  either   prose for the user             (the exchange ends)
  or       another request                (back to step 3)
  here:    request:  r2  get_payments { order_id: "1042" }

Step 2 here has no prose. A model may write a line of prose beside a request, and that line is still not the answer. The argument has a source: “A-1042” is in the user’s message.

Chapter 15, in “A Taxonomy of Common Agent Bugs”, says an argument’s value “should be sourced from the user’s text or from a prior span’s result” (a span is one recorded step of a run), and that “a value with no source was invented.”

What happens when the result is an error?

The error travels the same route as a result: your code sends it to the model at step 4, and the model reads it and writes a corrected request. In this script the model’s second request has dropped the prefix of the order id, so the exchange continues from step 3.

STEP 3  your code
  get_payments is on the list; "1042" does not fit the form A-0000
  the call is refused, and nothing runs

STEP 4  your code -> model
  result for r2:  error · bad_argument: unknown order id format,
                  expected A-0000 · retryable: false

STEP 5  model -> your code
  request:  r3  get_payments { order_id: "A-1042" }

The end of the run, illustrative and not in the interactive's script:

STEP 3, 4  your code -> model
  runs get_payments("A-1042")
  result for r3:  ok · two identical charges, one minute apart

STEP 5  model -> your code
  prose:  Order A-1042 was charged twice, one minute apart. I can't
          issue refunds myself, so the next step is our billing team.

Arguments are generated afresh on every call to the model, as text, which is how an id written correctly at step 2 came out wrong one call later. Nothing ran at step 3, and the result went back anyway; Chapter 2’s advice for a bad call is to “feed the failure back as data”.

“retryable: false” tells your code not to run the same call again; the model is free to ask again with a corrected argument, and does. The error also says what a valid id looks like, and writing errors a model can act on is part of how to design tools for LLM agents.

Play your code’s part

In the interactive below you stand at step 4 of this same run, with one step-3 choice, running the call again. The scripted model asks for get_order("A-1042"), you are shown what came back, and you choose what your code does with it: append the result and continue, retry the call, stop the run, or escalate to a person. The second pass is the malformed get_payments("1042") and its error.

With JavaScript on, the Run the loop runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Run the loop on its own page to share a result by link.

Those first two passes are this post: a result goes back to the model, and so does an error. Passes 3 to 7 are decisions about retrying, escalating and stopping, which belong to the agent loop and to idempotent tools and safe retries. In real code the append is unconditional, as the tool itself says.

Who does what in a tool call?

The model does three things in a tool-using turn: it decides whether to ask for a tool, it writes the request, and it reads the result and writes what comes next. Your code does eight, and the software that serves the model does three. The table is my own sorting of fourteen jobs that people attribute to “the model”; press “the model” in the filter and three rows remain, which is the lesson.

By the serving layer I mean the software between your code and the model itself: the provider’s servers on a hosted interface, or the runtime or library that loads the model on your own machine. Your code is what the book calls the harness.

# Job in one tool-using turn Step Who does it
1 Decide which tools exist for this call to the model. It is the list you send, every time. 1 your code
2 Write each tool’s name, description and argument schema: all the model learns about the tool. 1 your code
3 Turn the definitions into the text the model reads 1 the serving layer
4 Decide whether to ask for a tool or to answer in prose. A judgment, made from the question and the descriptions. 2 the model
5 Pick the tool and write the arguments. Generated text, so either can be wrong or invented. 2 the model
6 Hold the arguments to the schema while they are generated, where a strict mode exists and is on 2 the serving layer
7 Turn the model’s raw output into a structured call. Where a call can be lost or mangled. 2 the serving layer
8 Check that the tool exists and the arguments are valid, and set any argument the model should not write, such as who the user is 3 your code
9 Decide whether this call may run: authorization, confirmation, limits 3 your code
10 Run the function, once or again after a failure, in the order your code chooses. The only side effect in the exchange. 3 your code
11 Turn a failure into a result the model can read. An exception ends the run; a result lets the model correct it. 3 your code
12 Send the result back, paired with its request, in a new call to the model. The model keeps nothing from one call to the next. 4 your code
13 Read the result and write what comes next: an answer for the user, or another request 5 the model
14 Stop the exchange when it has gone on too long after 5 your code

At least two rows depend on your setup. Parsing (job 7) falls to your code with a bare model library and no server: one library’s documentation says that for some models “you’ll need to manually translate the output string into a tool call dict” (library documentation, read 6 October 2026). And the earlier exchange (job 12) is resent by your code on some interfaces and held by the server on others, so where anything remembers yesterday’s conversation, it is the serving layer. In the two server-held examples I read, the tool definitions still went with each call.

What changes when the provider runs the tool?

When the provider runs the tool, steps 3 and 4 move from your code to the provider’s servers, and the model’s three jobs stay where they were. The May 2024 provider’s documentation says so in a parenthesis: the model “emits a structured request, your code (or [the provider’s] servers) runs the operation, and the result flows back into the conversation” (documentation, read 6 October 2026; the bracket replaces the company’s name).

When a turn calls only such tools, that page says, a server-side loop “executes the operation and feeds the output back to the model before the response reaches you”, unless the loop stops early. Your application’s job becomes “to enable the tool and read the final answer rather than to participate in the execution loop”.

In the table’s terms, checking, running, error handling and sending the result back (jobs 8, 10, 11 and 12) pass to the serving layer, and the provider wrote the definition (job 2). Choosing the tool stays yours (job 1), and permission (job 9) shrinks to that one choice. Stopping (job 14) is the provider’s while its loop runs. The result still reaches your code, after the fact: the response shows “what ran and what came back”.

So do not hand an action you would want to confirm to a loop you cannot gate. I did not read how any provider validates or limits the calls it runs.

A tool reached over a tool protocol is a third arrangement. A server on the far side runs the function (job 10), and its author wrote the description (job 2). When your program holds the connection, the call still passes through it on the way out and the result on the way back, so checking, permission, returning the result and stopping (jobs 8, 9, 12 and 14) stay yours. When a provider connects to that server for you, it is the provider-run case above. One protocol’s specification says as much of permission: “there SHOULD always be a human in the loop with the ability to deny tool invocations” (specification, revision of 25 November 2025).

Who did what, when someone says “the model did it”?

The two tables can sort sentences they do not contain, and trying it is the quickest check that you hold the idea. Before reading across each row, decide who did what, and at which step.

What someone says What happened, by step Who did it
“The model searched the web for me.” Step 2: the model wrote a request for a search tool. Step 3: the search ran, in your code if you defined the tool, on the provider’s servers if the provider runs it. Step 5: the model read what came back. The model asked and read (jobs 5, 13). Your code or the serving layer searched (job 10).
“The model retried the API call when it timed out.” A call is made again only at step 3. Either your code retried there and the model never knew, or your code returned the timeout at step 4 and the model wrote the same request again at step 5. Your code called again (job 10). At most, the model asked again (jobs 13 and 5).
“The model refused to delete the account because it was dangerous.” At step 2 the model wrote prose and no request. The decision that binds is at step 3, and the run never reached it. The model declined to ask (job 4). Only your code can refuse a call (job 9).
“The model read the file and saw it was empty.” Step 2: a request to read. Step 3: your code opened the file. Step 4: your code sent back empty text. Step 5: the model read that text. Your code read the file (job 10). The model read your result (job 13), and a failed read returned as an empty string would look the same to it (job 11).
“The model called two tools at the same time.” Step 2: one message holding two requests, where the model and the interface allow it; the model wrote both before reading either result. Step 3: your code ran both, one after the other or together, as you wrote it. Step 4: two results, each paired with its request. The model wrote two requests (job 5). Your code chose the timing (job 10).

How does the model know which tool to use?

Everything the model knows about your tool is the name, the description and the argument schema you sent, which reach it as text in its input; it weighs them against the question. Chapter 2 says of its own example: “The decision of whether and how to call create_ticket is made from this definition alone—there is no source code to consult, no wiki, no colleague to ask.”

The May 2024 provider documents that its interface “constructs a special system prompt from the tool definitions, tool configuration, and any user-specified system prompt” (documentation, read 6 October 2026); a system prompt is the standing instructions at the top of the model’s input.

For open-weight models, one model library’s documentation describes turning a function you pass into a schema, and warns: “The parser will also ignore the actual code inside the function!”

The other interfaces I read do not document this step. Wherever the text is assembled, a tool the model ignores or misuses is first a problem of wording.

What does the model write, and who parses it?

The model writes the request as text, usually in a format it was trained to produce, and the serving layer parses that text into the structured call your code receives.

On training, the June 2023 announcement that brought the feature into one widely used programming interface said its models “have been fine-tuned to both detect when a function needs to be called (depending on the user’s input) and to respond with JSON that adheres to the function signature” (provider announcement, 13 June 2023, as archived on 21 December 2023). Fine-tuned means given further training, and JSON is a common text format for structured data.

The format is not shared. A 2026 essay by Rémi Louf prints one operation in three model families’ formats and concludes: “wire formats are training-time decisions, and nothing constrains them to a shared convention” (Louf, 9 April 2026). Whether a call is marked off by reserved tokens is one of those decisions.

A hosted interface hides that format from you, the essay says; with an open model, when no parser fits, “the output comes back garbled: reasoning tokens in arguments, malformed JSON, missing tool calls.” So whether a model supports tool calling has two halves: was it trained to write calls, and can the software serving it read them.

How is tool calling different from function calling, structured output, plugins and protocols?

Function calling, tool calling and tool use are three names for the exchange in this post. Structured output is the shape of the request, a plugin and a tool protocol are ways for definitions and calls to travel, and an agent is the exchange run in a loop. The table is my arrangement; each date is the one printed on a primary announcement.

Term What it names Who acts on the model’s output What I could date
Function calling, tool calling, tool use The five-step exchange Your code, or the provider’s servers for a tool the provider runs “Function calling” entered the June 2023 provider’s programming interface on 13 June 2023 (as archived). The May 2024 provider’s “tool use” left beta on 30 May 2024.
A tool call One request in the step-2 message; a message can hold several Your code
Structured output Any model output held to a machine-readable shape. A tool call is one use; a shaped reply to the user is another. Your code reads data. Nothing need run. The June 2023 provider announced a mode that returns valid JSON on 6 November 2023, and output matched to a supplied schema on 6 August 2024 (both as archived; the second is linked below).
A plugin A packaging of third-party tools for the June 2023 provider’s chat product, “described by a manifest file” That chat product; the developer supplied “an API with endpoints you’d like a language model to call” Announced 23 March 2023 (as archived)
A tool protocol A standard for how a program discovers tool definitions and forwards calls to the servers that implement them A server on the far side runs the function; your program decides whether to forward the call The May 2024 provider announced one on 25 November 2024 as “a new standard for connecting AI assistants to the systems where data lives”
An agent The exchange run in a loop, under limits, with the model choosing each next step Your code, repeatedly

Which of these words mean the same thing?

The first three do. Three documentation sets I read on 6 October 2026 say so in their opening lines: “Function calling (also known as tool calling)”, “Tool use (also called function calling)” and “tool calling (also known as function calling)”. The book says “interchangeably”.

Structured output and tool calling share machinery and differ in purpose: one shapes a reply for a program to read, the other asks your code to run something.

Plugins and protocols leave the model’s part alone: it still reads definitions and writes a request. They change who wrote the descriptions and where the function runs. And Chapter 2 defines an agent as “this request–execute–return exchange run in a loop.”

Why does a tool call go wrong?

A tool call goes wrong in a handful of ways, and each follows from one fact: the request is generated text, written from descriptions and passed through a parser. Chapter 5, Tools and the Action Space (in the full book) frames the design problem: “A tool is a contract between your deterministic code and a non-deterministic caller”. The grouping below is this post’s own, and it names each fault without fixing it.

The fault Why it is possible Where it is fixed
The model answered in prose and made no call Asking is a judgment (job 4), made from the question and the descriptions Confirm the definitions were sent, then rewrite the description to say when to use the tool
It asked for a tool nobody needed, or for the wrong one The same judgment, erring the other way The size of the tool set and descriptions that name their neighbors: the first two levers for an agent that calls the wrong tool
It invented an argument, or named a tool that does not exist With no value to hand, a model writes a plausible one Job 8: check the name against your list, trace each value to a source, and return the failure as a result
The arguments were well formed and wrong A schema checks shape Validation of meaning in your code; see “Who is responsible for what runs?” below
One request where you expected two, or the reverse How many requests a reply may hold depends on the model and the interface Chapter 2: “your handler should expect zero, one, or many”
The call arrived as plain text The serving layer did not parse the model’s format (job 7) The model’s and the server’s documentation; a locally run model can fail in either half
Your code never sent the result back The result was shown to the user, or the model’s own message was dropped The “classic first bug”; it shows up the first time you build an AI agent from scratch

What does a strict schema fix?

A strict schema mode removes arguments that do not fit the schema and leaves arguments that fit and are wrong. The mechanism is constrained decoding, a filter on generation. What a strict schema buys and what it does not is a design question, and structured output for tool calls is a longer subject than this section.

The June 2023 provider published figures for the difference. Its announcement of 6 August 2024, as archived, reports on “our evals of complex JSON schema following”: an earlier model of its own “scores less than 40%”, training took its newest model to “93% on our benchmark”, and with the constraint switched on that model “scores a perfect 100%” (provider announcement, as archived on 28 December 2024). These are one vendor’s internal numbers, on its own models, in 2024, and the post gives no test-set size.

The same announcement states the limits. It lists two cases in which the output does not match the schema even with the constraint on: the model refuses, or generation is cut off by a length limit. And when the output does match, “the model may still make mistakes within the values of the JSON object”. Chapter 2: “well-formed is a statement about syntax, and safe is a statement about consequences. A flawlessly schema-conformant call to delete_account still deletes the account.”

Who is responsible for what runs?

Your code is responsible, because step 3 is yours: the model’s request is a proposal, and whether it runs is decided in ordinary code you wrote. Chapter 2 says of a call that perhaps should not run: “Whether it should run—authorization, confirmation, limits—is your runtime’s decision to make at step 3”.

The request’s arguments are untrusted input to your function, so validate their meaning as well as their shape; Chapter 2 lists what that takes: “ranges, cross-field rules, “does this customer ID actually exist.”” Identity is not an argument for the model to write: the call runs in your code, and your code already knows which user is signed in.

The result is the other half. Whatever the function returns, the model reads, in the same input as your instructions. The June 2023 announcement carried the warning itself: “a proof-of-concept exploit illustrates how untrusted data from a tool’s output can instruct the model to perform unintended actions” (provider announcement, as archived on 21 December 2023). It recommended “user confirmation steps before performing actions with real-world impact”.

So running the function yourself protects you from the model’s request and does nothing about the function’s reply.

The attack is called prompt injection. How to prevent prompt injection in AI agents starts from this fact, and AI agent guardrails are the checks that sit at step 3. Where the function runs is a third decision, and sandboxing an agent’s tool execution is how you limit what a mistaken call can reach.

What should you see in your first captured exchange?

You should see nine things, in order, in one real exchange with one tool. Use one tool you defined yourself, and no helper that runs tools for you: some client libraries and frameworks do, and then step 3 happens out of your sight.

Print everything you send and everything that comes back, in full, for each call to the model, before you write a loop. If the second reply is another request, as in the example above, repeat items 3 to 7 for it.

  1. My first call to the model carries the tool definitions and none of my function’s code.
  2. The model’s reply contains a request and no result. Whatever prose comes with it is not the answer.
  3. The request names a tool that is on my list.
  4. Every argument value the model wrote traces to the user’s words or to an earlier result.
  5. At this point nothing has happened. My function has not run, and I can point at the line of my code that will run it.
  6. My second call to the model carries the result, paired with the request it answers: the id on my result matches the id on the request, or whatever my interface pairs by.
  7. The model’s own step-2 message is in that second call too, sent by me, or, where the server holds it, my second call carries a reference to the first reply.
  8. The final text is the model’s, written after it read my result. I can find my result’s values in it.
  9. I can say which part of the result’s text came from a source I do not control.

Item 5 is the one to stay on. The site’s three-minute agent loop explainer has the line for it: “Nothing has happened yet; a request is only text until your code runs it.”

A card for your provider’s names

The card below is a template to fill in from your own provider’s documentation, one line per part of the exchange.

My provider's names for the one exchange
----------------------------------------
Tool definitions go in the call as:          ____________
A definition's three parts are called:       name / ________ / ________
Tool definitions go with every call:         yes
The model signals "I want a tool" by:        ____________
A tool request carries:                      tool name / arguments / id? (yes / no)
Arguments arrive as:                         parsed data / text I must parse
One reply can hold several requests:         yes / no / can be switched off
I send a result back as:                     ____________
A result is paired with its request by:      id / tool name / position
I flag a failed tool by:                     ____________
The standing instructions go:                in the message list / in their own field
The conversation is kept by:                 me (I resend it) / the server (I send a reference)
Schema conformance of arguments is:          asked / enforced when I switch on ________
Tools the provider runs for me, if any:      ____________

The third line has one option because every example I read, server-held conversations included, sent the definitions with each call.

Limits: what differs between interfaces, and what is unmeasured

The five steps hold in every interface I read; the details around them differ by category, and the failure rates are unmeasured in public. I read the documentation of three hosted interfaces, one local runtime, one open-model library and one local inference server, all on 6 October 2026. Each page states that provider’s behavior on that date.

What differs. A result is paired with its request by an id in the three hosted interfaces, and by the tool’s name in the local runtime’s example. The hosted pages document several requests in one reply; the open-model library says “in most cases, models only emit a single tool call at a time”, and the local server says parallel calls are “supported on some models but disabled by default” (project documentation, read 6 October 2026).

What is unmeasured. I found no public rate, with a denominator, for wrong-tool or invented-argument calls in real use, and none for how often a serving layer drops a call. The largest group of reader questions I collected is about those faults, which says where I looked: sites where people post what broke.

What this post did not do. The worked exchange is a script and no model wrote it. Four announcements were read as archived copies, because the live pages refused the download. I found no source for who coined “function calling” or “tool use”, and no dated page for the often-repeated rename from “functions” to “tools” at one provider.

The takeaway

Next time someone asks “what is tool calling in LLMs?”, point at step 3. The model wrote a request at step 2 and read a result at step 5. The call itself happened in between, in code that someone wrote and can check, and that someone is you unless you handed the step to a provider. That is why an agent’s actions can be checked at all: each one passes through a line of code that can log it, test it or refuse it.

The mechanism is in Chapter 2, “The Engine”, free to read online, and the craft of the tools themselves is in Chapter 5, Tools and the Action Space (in the full book). The guide to tools, skills and protocols maps the rest, and you can see the formats.

Questions readers ask

Does the model run my function?
No. The model writes a request that names one of the tools you described and supplies arguments, then stops. Your code decides whether to run the function, runs it and sends the result back. The exception is a tool the provider hosts, such as a web search some providers offer: there the provider's servers run it, and the model still only writes the request.
Why do I have to call the model twice for one tool call?
Because the model's first reply is addressed to your code, not to the user. It contains a request and no answer. Your code runs the function and calls the model a second time with the result, and only then can the model write for the user. Chapter 2 of AI Agents, Engineered says the exchange has two model calls at minimum.
How does the model know which tool to use?
From the question and from the name, the description and the argument schema you sent. It knows nothing else about your tool. Those three are turned into text in the model's input. The function's code is never sent, so a tool the model ignores or misuses is first a problem of wording in the definition.
Is tool calling the same as function calling, and how does it differ from structured output?
Tool calling, function calling and tool use are three names for the same exchange; three documentation sets, read in October 2026, treat them as synonyms. Structured output is wider: any model output held to a machine-readable shape. A tool call is structured output that asks your code to run something; a shaped reply to the user is structured output that runs nothing.
Can a model call two tools at once?
It can write two requests in one message where the model and the interface allow it; your code then runs both and returns one result for each. Whether a given model does this, and whether it can be switched off, differs by interface and by model, so a handler should expect zero, one or many requests in a reply.

Sources

  1. OpenAI (2023). Function calling and other API updates (announcement dated 13 June 2023; read as archived, snapshot of 21 December 2023)
  2. OpenAI (2023). Announcement of plugins for a chat product (dated 23 March 2023; read as archived, snapshot of 31 December 2023)
  3. OpenAI (2023). New models and developer products announced at a developer conference (dated 6 November 2023; read as archived, snapshot of 31 December 2023)
  4. OpenAI (2024). Introducing Structured Outputs in the API (dated 6 August 2024; read as archived, snapshot of 28 December 2024)
  5. Anthropic (2024). Announcement that tool use is generally available (30 May 2024)
  6. Anthropic documentation (2024). API release notes, entry of 30 May 2024 (read 6 October 2026)
  7. Anthropic documentation (2026). How tool use works, including where tools run (read 6 October 2026)
  8. Anthropic documentation (2026). Tool-use documentation: defining tools, section on the system prompt built from tool definitions (read 6 October 2026)
  9. Anthropic documentation (2026). Tool use overview (read 6 October 2026)
  10. OpenAI documentation (2026). Function calling guide (read 6 October 2026)
  11. Google documentation (2026). Function calling guide of a third hosted provider (last updated 23 September 2026; read 6 October 2026)
  12. Ollama documentation (2026). Tool calling in a local model runtime (read 6 October 2026)
  13. Hugging Face documentation (2026). Tool use, in the documentation of an open-source model library (read 6 October 2026)
  14. llama.cpp project (2026). Function calling notes of a local inference server (read 6 October 2026)
  15. Rémi Louf (2026). Tool calling, open source, and the M×N problem (9 April 2026)
  16. Anthropic (2024). Introducing the Model Context Protocol (25 November 2024; one example of a tool protocol)
  17. Model Context Protocol (2025). Model Context Protocol specification, revision 2025-11-25: Tools (one example of a tool protocol)
  18. Hacker News commenter MuffinFlavored (2023). Forum comment asking where the model is told where the API lives (13 June 2023)
  19. Stack Overflow (2024). Stack Overflow question 79253601: how to use tools with LLMs (asked 5 December 2024)