In a tool calling LLM agent, the model requests a function by name with arguments, and your code runs it. A tool is one action. A skill is a written procedure, with optional references and scripts, that loads only when a task calls for it. A tool should become a skill when its definition starts carrying steps, order and house rules.
That last sentence is this post’s claim, and a team running a tool calling LLM agent in production needs two qualifications before it is useful. Moving a procedure out of the always-loaded text saves less than the threads promise, and the arithmetic below shows why. It also adds a way to fail that standing text does not have: the model can decline to load the procedure, and nothing errors when it does.
What is tool calling in LLMs, and where does a skill begin?
Tool calling is the mechanism by which a model asks your program to act: it emits a structured request naming a function and its arguments, and your code decides whether to run it. The book’s glossary defines a tool as “A function the model can ask your program to run.” The tool call is the request, and the model executes nothing. Function calling is the same thing under another name; one model provider’s documentation, read in October 2026, writes “Function calling (also known as tool calling).”
A skill is a different kind of object. The glossary calls it “A bundle of instructions, reference files, and scripts an agent loads on demand,” and files it under packaged expertise. One vendor’s 2025 engineering write-up put the pattern on the map, and I cite it as an example. It describes skills as “organized folders of instructions, scripts, and resources that agents can discover and load dynamically to perform better at specific tasks.”
The name misleads in one respect a tech lead should hear early. A skill in this sense is a document. Nothing is trained, and the model that reads it is the same model as before.
Chapter 6 of AI Agents, Engineered (in the full book) compresses the difference into five words: “a tool is a verb.” The skill is the judgment about the verbs. The third layer, a protocol, is “An agreement about how two independently written programs talk,” which is how an agent reaches tools somebody else wrote.
Which layer does a gap in a tool calling LLM agent belong to?
Ask what kind of thing is missing: reach to a system, a safe action, a procedure, a rule every task needs, or a step that must never be skipped. Each answer names a different layer, with its own standing cost and its own way of failing.
The first three bins are the chapter’s. Its diagnostic is “one question: what kind of thing is missing?” No reach means connectivity. Reach without a safe, task-shaped action means a tool. And the third: “If it has the reach and the actions and still does the job like a talented stranger, the gap is expertise; write the skill.”
My addition is the last two bins, arranged from things the book says elsewhere, so the five-bin sorting is this post’s extension of the chapter. Chapter 6 allows that writing a correction into the standing prompt is sometimes correct: “for one paragraph, that is the right fix.” Chapter 5 adopts a slogan it credits to a framework vendor’s blog post: “discretion on the outside, determinism on the inside.”
| What is missing | Layer | What it costs while unused | How it fails |
|---|---|---|---|
| Reach to a system | Protocol server, or a direct integration | Whatever definitions the host loads | The far end changes or breaks |
| A safe, task-shaped action | Tool | Its whole definition, every pass | Wrong tool, wrong arguments, an error |
| A procedure some tasks need | Skill | One short menu line, every pass | Silently not loaded, or loaded and not followed |
| A rule nearly every task needs | Standing instruction | Its whole text, every pass | Diluted as the standing text grows |
| A step that must never be skipped | Code, in a tool or the harness | Nothing in the context | Loudly, with an error |
The chapter explains why the sorting matters in practice: “teams routinely buy another connector when what they are missing is a procedure, and write another procedure when what they are missing is reach.”
The skill row holds when one miss is affordable; the placement table further down covers the rest.
Are tools, skills and protocols rivals?
They are layers that combine, with one real overlap at the edge. A commenter in an April 2026 discussion of this argument put the layering in one sentence: a protocol server “exposes all the things that can be done, and Skills encode a workflow/expertise/perspective on how something should be done given all the capabilities.”
The overlap is this. A script plus a written manual can replace a protocol server when one team owns both ends, the host can run commands and the credentials are already present. David Mohl’s essay I Still Prefer MCP Over Skills, submitted to the same forum that day, states the other side fairly: “Skills are great for pure knowledge and teaching an LLM how to use an existing tool.” Access to a service, he argues, is a different job, and one a protocol does far better.
The book’s ledger agrees with both. “For one tightly coupled integration you fully control, a direct API call is simpler, faster, and easier to debug.” Protocols earn their cost when many agents and third parties share a capability, which the post on what MCP is in AI works through, and the M×N protocol calculator does the multiplication. The Model Context Protocol, a prominent example of a tool protocol as of 2026, describes itself (read October 2026) as “an open-source standard for connecting AI applications to external systems.”
One complaint dissolves once the layers are apart. The complaint that a protocol bloats the context describes a host that loads every advertised definition up front. Loading policy belongs to the host; reach belongs to the protocol.
What are the signs that a tool has grown into a skill?
The clearest sign is a tool description that contains a procedure: numbered steps, ordering words such as first and before you, or house rules. Every task re-reads that description on every pass, including the tasks that never call the tool.
A builder on a February 2026 thread described exactly this in one server’s search tool: “the search tool description is basically a mini tutorial. This is going right into the context window.”
The table is this post’s own synthesis. The book supplies two of the promotion paths in so many words: the repeated correction and the proven script. The description that has become a procedure is my extension; the book supports it only indirectly, through the cost of definitions and its advice on what a description is for. The table has twelve rows, and the buttons narrow it by the layer the fix lands in.
| Sign you can observe | What it means | Move | Layer |
|---|---|---|---|
| A tool description contains numbered steps, ordering words or house rules | The description carries a procedure, re-read on every pass by tasks that never call the tool | Cut the description back to what, when, returns and fails. A rule that must hold on every call goes into the tool’s code. The remaining procedure moves to a skill that names the tool when a minority of tasks need it, and the description keeps one line pointing to it | skill, code |
| You type the same correction into a second session | A procedure lives in your head and is injected by hand | Write it down once, then place it with the placement table further down | skill, standing, code |
| A section of the system prompt is used by a minority of tasks | Paid on every pass, and it competes for attention on unrelated tasks | A skill whose description says what it does and when to use it, if one miss is affordable; if one miss is costly, it stays standing with a check | skill, standing |
| The agent rewrites the same script every session | Proven code is being re-derived at model prices | Bundle the script with a skill when a minority of tasks need it; it runs without being read | skill |
| Every tool call succeeds and the output is still wrong for your house | The gap is expertise: wrong format, wrong fiscal calendar, wrong house style | Write the procedure down; a skill when a minority of tasks need it and one miss is affordable | skill |
| Three or more tools only make sense in a fixed order | One task if the order never varies; a recipe if judgment enters between steps | One task-shaped tool, or a skill with a stop between steps; a stop that must hold is enforced by the tool or the harness | tool, skill |
| A step must happen on every run, no exceptions, and a program can enforce it | Prose is followed most of the time; a skill can be skipped silently | Enforce it in code inside the tool or the harness | code |
| A rule applies to nearly every task | A skill would add a decision point and a miss rate for no saving | Keep it in the standing instructions, short | standing |
| The agent cannot reach the system at all | Connectivity is missing | A direct integration for one system you own; a protocol server for several agents or third parties | protocol |
| The agent reaches the system through endpoint-shaped calls only | The tools follow the API’s shape, and the task needs a different one | Redesign the tool | tool |
| The text records a fact about this user, this run or last Tuesday | This is state | Memory | memory |
| The skill menu runs to dozens of near-identical lines | The too-many-tools problem, one level up | Remove skills that do not earn their line | curate |
Every row that ends in a skill assumes two things: a minority of tasks need the procedure, and one miss is affordable and would be caught. The placement table further down gives the answer when either fails.
How do the rows trace to the book?
Seven rows trace to the book, and the rest are reasoning from it. The reach row and the endpoint-shaped row are the first two bins of the chapter’s diagnostic. The second row is the scene that opens Chapter 6’s first section, a correction such as “always run the migration dry-run before applying it” typed twice: “The first time you typed it, it was a prompt. The second time, it was a chore.” If that dry-run must never be skipped, the migration tool should enforce it; the written text only explains when and why.
The fourth comes from the code-execution section of Chapter 5 (in the full book), where a proven script can be “saved, named, and loaded next time instead of re-derived.” The fifth is the failure that opens Chapter 6: “every tool call succeeds, and the work is still wrong.”
The memory row is a boundary the chapter draws firmly: “writing session facts into a skill folder is a category error that produces stale, misfiring instructions.” And the last row is its third caveat: “Curate the library like you curate the toolset: install what earns its line.”
Two rows hand off to other posts. Endpoint-shaped tools are the subject of how to design tools for LLM agents, which covers the tool contract and which I do not repeat here. Bundled scripts need somewhere safe to run, and sandboxing agent tool execution is that subject.
What does the always-loaded pile cost on every pass?
It costs its full size multiplied by the number of model passes, because the standing text is re-sent each time. Chapter 5 states the mechanism for tools: “Every tool definition rides on the desk on every single pass, and each one is an option the model must weigh every time it decides what to do next.”
Here is an example with the arithmetic shown. Every number in it is illustrative: the inputs are invented and round, and they describe no real system. Take a tool calling LLM agent for operations work whose standing text holds core instructions, six procedures pasted into the system prompt, and 24 tool definitions.
| Always loaded | Arithmetic | Tokens per pass | Share |
|---|---|---|---|
| Core instructions | 1 × 1,000 | 1,000 | 3% |
| Six pasted procedures | 6 × 1,500 | 9,000 | 28% |
| 24 tool definitions | 24 × 900 | 21,600 | 68% |
| Total | 1,000 + 9,000 + 21,600 | 31,600 | 100% (rounded) |
A 20-pass task re-reads that pile 20 times: 31,600 × 20 = 632,000 token-reads, before the conversation or any tool result is counted. People in the threads I read quote the session-start total and stop there. The per-task figure is the one that matters.
Is 900 tokens per definition fair? It sits inside what two dated reports show, and the range is wide. One vendor’s 2025 example counts “58 tools consuming approximately 55K tokens before the conversation even starts,” about 950 each by my division; the tool-design post linked above uses the same example. A user’s measurement from November 2025 put one server’s 27 tools at roughly 18,000 tokens, about 670 each.
What do skills move, and what do they leave?
Skills move the procedures and leave the tool definitions untouched. Give each of the six procedures a 100-token menu line, so the menu costs 6 × 100 = 600 per pass. The task needs one procedure; its 1,500-token body loads on pass 3 and stays for the remaining 18 passes.
| Illustrative, over one 20-pass task | Arithmetic | Token-reads | Change |
|---|---|---|---|
| Procedures, pasted | 9,000 × 20 | 180,000 | |
| Procedures as skills, one used | 600 × 20 + 1,500 × 18 | 39,000 | 78% less |
| Procedures as skills, none used | 600 × 20 | 12,000 | 93% less |
| Whole pile, before | 31,600 × 20 | 632,000 | |
| Whole pile, after, one skill used | (1,000 + 600 + 21,600) × 20 + 27,000 | 491,000 | 22% less |
Both numbers are true, and a reader needs both. Skills removed about four-fifths of the procedure cost and about one-fifth of the total, because two-thirds of the pile was tool definitions. In the same example, dropping 12 of the 24 definitions would save 12 × 900 × 20 = 216,000 token-reads, more than the 141,000 the six skills saved.
So the tool table of a tool calling LLM agent needs its own fix: a smaller set, or a host that loads definitions on demand. The vendor write-up cited above reports its own figures for the second. Loading every definition up front for a library of more than 50 tools consumed about 77K tokens. Fetching definitions through a search tool consumed about 8.7K, in one vendor’s product at one date.
I link the context window budget planner here and do not embed it. It prices one pass, so it reproduces the 31,600; the 632,000 is that figure times 20. Prompt caching lowers the price of the repeated text, while the attention it takes stays the same; the prompt caching savings calculator covers the price side.
How does progressive disclosure keep a skill cheap?
Progressive disclosure keeps a skill cheap by loading it in layers: a short description always, the procedure when the model judges it relevant, and reference material only when a branch of the task reaches it. The glossary’s definition is “Revealing information in layers, each layer pulled in only when the task demonstrates it is needed.”
The open specification of one skill format, as read in October 2026, recommends about 100 tokens of metadata per skill and a body under 5,000 tokens. Those are one format’s recommendations at one date. The three-level shape is the part that should last, and Chapter 5 names it without any format: “give the model an entry point instead of an inventory.”
Two hosts’ documentation, both read in October 2026, shows the same line being drawn. One contrasts skills, “Task-specific, loaded on-demand,” with “Always applied” instructions. A second says its initial skill list “uses at most 2% of the model’s context window,” and that with many skills installed it “shortens skill descriptions first.”
The menu has a budget, and a description you wrote with care may reach the model shortened. The chapter’s warning applies with full force: “A vague description means the skill never loads, and the failure is perfectly silent.”
When should a procedure stay in the standing prompt?
A procedure should stay in the standing prompt when most tasks need it or when a single miss would be costly. The miss rate decides this far more often than token cost.
Take the token side first. A procedure of B tokens costs B on every pass when it is standing. Packaged, it costs a menu line m on every pass, plus B on the fraction f of tasks that load it.
The skill is cheaper whenever f is below 1 − m/B. With the illustrative B = 1,500 and m = 100, that is 1 − 100/1,500 = 93%. A procedure would have to be used on more than nine tasks in ten before tokens argued for keeping it standing.
What standing text has, and a skill lacks, is the absence of a decision. The chapter says so in four words: “First, triggering is probabilistic.”
What did the one published eval measure?
It measured a coding agent learning one framework’s new APIs, with the documentation delivered three ways. This is Jude Gao’s write-up for a framework vendor, dated 27 January 2026. Packaged as a skill with no further prompting, the pass rate was 53%, the same as with no documentation, because “In 56% of eval cases, the skill was never invoked.”
An explicit instruction to use the skill “improved the trigger rate to 95%+ and boosted the pass rate to 79%.” A compressed 8KB index kept in the always-loaded instructions file passed 100%. The author’s explanation is three words long: “No decision point.”
Read it as one eval. It covers one vendor, one framework’s documentation and a coding agent the write-up does not name, at one date. I did not capture the size of its test set. The knowledge was also needed on essentially every task in the suite, which is the case where always-loaded text is supposed to win.
Two smaller reports point the same way. A second team wrote in February 2026 that in its test “Skills were rarely invoked,” with no rate given. A user’s bug report from January 2026 reads: “No error messages. The skill is silently ignored.”
So what is the rule?
Sort by whether a program can enforce the step, how often the procedure is needed, and what one miss costs. This rule is mine, reasoned from the mechanism and that one eval; no controlled study compares the placements. Read the rows from the top and stop at the first that fits.
| The text is… | Where it goes | Why |
|---|---|---|
| 1. A step that must never be skipped, and a program can enforce it | Code in the tool or the harness | Skipping it becomes an error |
| 2. Needed on most tasks | Standing instructions, kept short | A decision point adds a miss rate for almost no saving |
| 3. Needed on a minority of tasks, and one miss is costly | Standing instructions, plus a check on the result | You cannot see a skill that did not load |
| 4. Needed on a minority of tasks; a miss is cheap and would be noticed | A skill | The saving is large and a miss is caught |
| 5. Needed on a minority of tasks; a miss is cheap and would go unseen | A skill, with an activation test | The saving is real; the test is how you find the misses |
Take a release checklist that one task in ten needs and that is costly to skip. Its steps a program can verify go to code under the first row. The remaining text reaches the third row and stays standing, which is the price of having no decision point.
A consequential write should never exist only as a line in a skill. The skill may say when a refund is appropriate; the refund stays a named tool with its own check.
The book’s glossary calls a file of always-loaded project rules a standing project-instructions file. The eval’s winning 8KB artifact was itself an index pointing to files read on demand. The disagreement was about where the menu line lives and how firmly it is worded.
What goes in a skill? A skeleton to copy
A skill needs six things: what it does, when to use it, the tools it names, the steps, the checks, and pointers to anything heavier. The skeleton below is in no vendor’s syntax on purpose; the portable idea is a short description, a procedure loaded on demand and resources loaded later still.
name: <short-name-for-the-job>
description: <What it does, in the third person>. Use when <the words a
user would actually type>. Do not use for <the nearest neighboring task>.
What this is for
<One or two sentences: the job, and the standard the output must meet.>
Tools this procedure uses
<tool_name>: <what it is used for here>
<write_tool_name>: <a write; the tool's own check still applies>
Steps
1. <step, naming the tool it uses>
2. <step>
STOP: do not go on to step 3 until <a condition a person or a check
can confirm>.
3. <step>
Checks before you finish
- <an invariant the output must satisfy>
- <what to report if a check fails>
Reference (open only if the task reaches it)
<document>: <one line on when to open it>
<script>: <what it does; run it, do not read it>
Changelog
<date> <who> <what changed, and which observed failure it fixes>
Three lines carry most of the weight. The description is all the model sees when it chooses, so it says what and when in a user’s words. The STOP line exists because an agent’s default motion is forward; in the chapter’s words, “The gate is the pattern.” And the script line reflects what a bundled script buys: “the model’s judgment decides when to run it, and the code’s determinism guarantees what it does.”
How big should one skill be?
One skill should cover one job, with a body short enough to read like the procedure. The chapter’s rule is “Keep each skill narrow: one job done reliably beats a sprawling do-everything folder that is hard to trigger, hard to test, and hard to debug.”
Its reason for a short body is “every token of the body is a tax on every task that triggers the skill, including the tasks that needed only a third of it.” The open format cited above, as read in October 2026, recommends a body under 5,000 tokens and under 500 lines, with heavier material moved to reference files. Treat those as one format’s working figures.
Narrow can be overdone. One practitioner complained in September 2026 about a set split so finely that a task “needs another skill and have to go fetch it.” The unit I use is the trigger: one job with several branches is one skill with references, and two jobs that share a topic are two skills.
How do you check that the move worked?
Run the same tasks with and without the skill, and test activation in both directions. A skill that never loads looks identical to a tool calling LLM agent working from general knowledge, so the check has to be one you go and make.
- Pick five to ten representative tasks, some that need the procedure and some that must not trigger it.
- Record whether the body loaded on each run, from the trace; the agent’s reply will not tell you.
- Compare the pass count with the skill installed and removed. The chapter’s instruction is “run your representative tasks with and without the skill and take the difference.”
- Phrase the tasks the way people ask, typos and vagueness included.
- Repeat when the model or the host changes. A description that fires on one setup can sit inert on another.
A skill you did not write deserves a read before it is installed. The chapter’s sentence is “a skill is a dependency that ships instructions, and you already know better than to install dependencies unread.” A 2026 survey preprint reports in its abstract that “curated skills can substantially improve agent success rates while self-generated skills may degrade them”; I read the abstract only and record no effect size.
Where does this advice break, and what are the alternatives?
It breaks where the host or the model cannot carry the mechanism, and where the evidence runs out. Four limits are worth stating.
The mechanism assumes a capable host. On-demand loading needs an agent that can read files, and one that can run code if scripts are bundled. It also needs a model that can judge relevance from two sentences.
A bigger window does not make the question moot. Drew Breunig’s 2025 account reports a small model that failed a task with 46 tools in context and succeeded with 19, although all 46 fit. A 2025 preprint (Gan and Sun) reports tool selection accuracy of 43.13% with retrieved candidates against a 13.62% baseline; I verified its abstract only. Chapter 5’s version: “Past a certain count, additions subtract.”
The menu claim rests on analogy. I found no study of selection accuracy against skill count. That a large menu degrades selection is the chapter’s caveat, that it “recreates, one level up, the too-many-tools confusion,” plus the tool evidence.
Every number here is dated. Definition sizes and trigger rates come from 2025 and 2026 hosts and models, and they will move. The shape should hold for any tool calling LLM agent: always-loaded cost multiplies by passes, and a decision point has a miss rate.
The alternatives are few. A smaller tool set chosen per kind of task is one, and a subagent that carries a group of tools in a context of its own is the other. The context engineering checklist treats both as part of a wider audit.
The sorting to do this week
Open your agent’s assembled prompt and its tool table, and sort every block: an action, a procedure some tasks need, a rule every task needs, a step that must never be skipped, or reach. Price the always-loaded pile per pass, then per task. Move a procedure into a skill only when a minority of tasks need it and one miss is affordable, and then go and check that it loads.
The tool side of this argument is in Chapter 5, Tools and the Action Space (in the full book), and skills, progressive disclosure and protocols are in Chapter 6, Skills, Protocols, and Interoperability (in the full book). The free guide to tools, skills and protocols collects the related posts and tools, and you can see the formats.
Questions readers ask
- What is tool calling in LLMs?
- Tool calling is the mechanism by which a language model asks a program to act. The model emits a structured request naming a function and its arguments; the program validates the request, runs the function and returns the result as text for the next pass. The model itself executes nothing. Function calling is another name for the same thing.
- What is the difference between a tool and a skill?
- A tool is one action behind a name, a description and an argument schema, and its definition is sent to the model on every pass. A skill is a written procedure, with optional reference files and scripts, for using actions to your standard. Only a short description of the skill is always loaded; the body loads when the model judges it relevant.
- Do skills replace MCP servers or other tool protocols?
- No. A protocol gives an agent reach to a system across a boundary; a skill gives it procedure. A script plus a written manual can stand in for a protocol server when one team owns both ends, the host can run commands and credentials are already in place. It cannot where several agents from different owners need the same capability.
- Why did my skill not trigger?
- At the moment of choosing, the model sees only the skill's short description. If the description does not say what the skill does and when to use it in the words a user would type, or the host shortened it to fit a budget, the body never loads and nothing errors. Test activation on tasks that should trigger the skill and on tasks that should not.
- When should instructions stay in the system prompt instead of a skill?
- When most tasks need them, or when a single miss would be costly. Standing instructions have no decision point, so they cannot fail to load. Keep them short. A step that must never be skipped, and that a program can enforce, belongs in code inside a tool or the harness, where skipping it produces an error.
Sources
- Zhang, Lazuka, Murag (Anthropic) (2025). Equipping agents for the real world with Agent Skills (one vendor's engineering write-up; the pattern's origin, cited as an example)
- Agent Skills (2026). Agent Skills specification (open format; read October 2026)
- Visual Studio Code documentation (2026). Agent Skills, Visual Studio Code documentation (one example of a host implementing the format; read October 2026)
- OpenAI (2026). Skills, developer documentation (a second example of a host implementing the format; read October 2026)
- Jude Gao (Vercel) (2026). AGENTS.md outperforms skills in our agent evals
- Bertelli and Çelik (LlamaIndex) (2026). Skills vs MCP tools for agents: when to use what
- Anthropic (2025). Introducing advanced tool use on the Claude Developer Platform (one vendor's write-up, cited for its token counts)
- OpenAI (2026). Function calling (documentation; one example of a model provider's guide, read October 2026)
- Model Context Protocol (2026). What is the Model Context Protocol (MCP)? (one example of a tool protocol; read October 2026)
- Gan and Sun (2025). RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation (preprint)
- Drew Breunig (2025). How Contexts Fail and How to Fix Them
- Jiang, Li, Deng, Ma, Wang, Wang, Yu (2026). SoK: Agentic Skills -- Beyond Tool Use in LLM Agents (preprint)
- David Mohl (2026). I Still Prefer MCP Over Skills (submitted to Hacker News April 2026)
- anthropics/claude-code issue tracker (2025). Issue #11364: Lazy-load MCP tool definitions to reduce context usage (a user's measurement, November 2025)
- anthropics/claude-code issue tracker (2026). Issue #21947: Personal skill not auto-triggering despite matching task description (January 2026)
- Hacker News (2026). Hacker News comments on tools, skills and protocol servers (February to September 2026)