Home / Tools / Skill token budget

Free tool · runs in your browser · from Chapter 6

Skill token budget

Budget the context cost of agent skills: the standing menu, what one task loads under progressive disclosure, and a check of whether a description fires.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The three levels, the illustrative sizes (a hundred tokens per menu line, a couple of thousand per body), the description advice with its two examples and the three caveats are Chapter 6's. The arithmetic, the bundled-material default, the window threshold, the description heuristics and the token estimate are the tool's own.

What is an agent skills token budget?

An agent skills token budget is an estimate of how much context a library of skills takes up: what is always present, what one task pulls in, and what stays on disk. The calculator above works it out for your numbers, compares it with loading the whole library into the prompt, and checks whether a skill’s description is written so the skill will load when it should.

A skill, in Chapter 6 of the book, is know-how packaged as files in a named folder: instructions first, then optional reference files, templates and scripts. The budget question arises as soon as there are many of them. The chapter asks the reader to do the arithmetic before reading on, and the tool starts from the same figures.

Progressive disclosure.
Figure 6.2 Progressive disclosure. The library stays on disk; the desk receives the menu of every skill, the body of the one skill a task triggers, and only the slices of bundled material that this task’s branches actually reach. Unused expertise costs a run its menu line and nothing more, which is why the library can grow without the desk feeling it. Reuse this diagram

Why not put every procedure in the system prompt?

Putting every procedure in the system prompt fails because the agent then pays for all of them on every call, although a given task uses one or two. Chapter 6 runs the numbers on a team that has written down a hundred procedures of a few pages each, where “A few pages is on the order of a couple of thousand tokens.”

The chapter states the result in round terms. “Load the library the naive way, all of it in the standing prompt, and you are carrying a few hundred thousand tokens of expertise on every single call, for a task that will use, at most, one or two procedures of the hundred.”

The bodies alone come to about 200,000 tokens on those figures. The calculator opens at 300,000 because it adds its own illustration of 1,000 tokens of bundled material per skill, which a naive load would carry too.

The cost is not only the bill. The chapter lists three problems in one sentence: “The desk cannot hold it; even where it physically fits, Part III will show you that the quality of attention degrades long before the space runs out; and you are billed for the whole pile on every pass of the loop.” A library that fits in the window still crowds it, and a crowded window is where context rot sets in.

What are the three levels of progressive disclosure?

The three levels are the menu, the instruction body and the bundled material: a short entry for every skill that is always in context, the full procedure of a skill that loads only when the agent decides it applies, and supporting files that are read one piece at a time. The chapter’s name for the arrangement is progressive disclosure, “revealing information in layers, each layer pulled in only when the task demonstrates it is needed”.

Level What it holds When it enters the context The chapter’s illustration
1. The menu Name and a short description of every installed skill Always “perhaps a hundred tokens per entry”
2. The body The procedure of one skill “only when the agent judges, from the menu, that the skill applies” A couple of thousand tokens for a few pages
3. Bundled material References, templates, scripts “piece by piece, only when a particular branch of the task requires that particular piece” No figure given

On the chapter’s hundred-skill library the calculator gives a 10,000-token menu. A task that triggers one skill and reads 1,000 tokens of its reference material puts 13,000 tokens on the desk, against 300,000 for the library loaded whole. The figure for bundled material is the tool’s own illustration, since the chapter gives none, and the chapter calls the numbers it quotes “illustration rather than specification”.

The second button loads a smaller case that a footnote in the chapter quotes from one implementation’s guide: “an agent with 10 skills starts each call with roughly 1,000 tokens of L1 metadata instead of 10,000 tokens in a monolithic prompt”. The calculator reproduces the 90% that guide reports. The footnote adds the right caution: “The figures are one implementation’s; the three-level shape is the durable part.”

What the mechanism buys is stated in the caption of the chapter’s figure: “Unused expertise costs a run its menu line and nothing more, which is why the library can grow without the desk feeling it.”

Why do scripts cost close to nothing?

Scripts cost close to nothing because running a program does not require reading it. The chapter gives them “a privileged corner of the third level”, for a reason it states plainly: “they can be executed without ever being read, so their cost on the desk is close to zero no matter how large they grow.”

Its worked example is a skill for handling PDF files that ships a program to extract every form field. The program is written once, and “the agent runs it without reading either the program or the PDF into its context; the source text stays on disk, the document stays on disk, and only the small, structured result lands on the desk.”

The calculator counts each script run at zero tokens. That is the chapter’s “close to zero”, rounded down, and it leaves out the script’s output, which does enter the context and can be large if the script is careless about what it prints. Treat the zero as the cost of the code, not of the result.

What makes a skill description work?

A skill description works when it tells the agent what the skill does and when to use it, in words that match how people ask. The chapter is emphatic that this single line decides whether everything else in the folder is ever read: “At the moment of choosing, the menu line is all it has; the body of your skill, however excellent, is still on disk, invisible by design.”

The failure is quiet. “A vague description means the skill never loads, and the failure is perfectly silent: no error, no complaint, just an agent doing the task from general knowledge, slightly wrong, while the cure sits shelved a few kilobytes away.”

The advice fits in one sentence, which the second half of the tool turns into checks. “Say what the skill does and when to use it, in that order, in the third person, using the concrete words a user would actually type”. The chapter then quotes two examples from vendor guidance. The one to imitate reads “Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when the user mentions PDFs, forms, or document extraction”. The other, “Helps with documents”, it calls “the epitaph of a skill that will never fire.” Both are buttons above.

The six checks are the tool’s reading of that advice, done with word patterns:

  1. Says what it does. It does not open with a trigger or with a verb such as “helps” that fits any skill.
  2. Says when to use it. There is a “use when” clause, and it comes after the what.
  3. Third person. No “I”, “we” or “you” outside quotation marks, so a quoted user phrase is left alone.
  4. Concrete words. Enough specific terms for a request to match.
  5. Not too eager. No catch-all such as “any task” or “always”.
  6. Fits a menu line. A rough token count against the line size you set.

The chapter’s summary of the craft is the right frame for all six: “Treat the description as API documentation for a caller that cannot read your source.”

What does the budget not tell you?

The budget does not tell you whether a skill will fire, how a smaller model will handle it, or how many skills are too many; the chapter raises all three as caveats, and the tool quotes them beside the result.

Triggering is probabilistic. The agent decides to load a skill by judgment, and “a skill that fires reliably on one setup can sit inert on another”. The opposite failure is the one the “Skills that trigger wrongly” field prices: a description that is too eager ends up “firing on unrelated tasks and dragging its body onto desks that had no use for it.” The chapter’s instruction is to test in both directions, with “tasks that should trigger it and tasks that should not.”

The host must be capable. The mechanism assumes an agent that can navigate files and run code, and a model that can judge relevance from two sentences. “Smaller models follow multi-file skills less reliably”.

The menu is not free. “At a hundred tokens a line, a large installed library costs real context before any skill fires”. A long list of similar descriptions also brings back the confusion of an oversized tool list. The rule that follows is short: “Curate the library like you curate the toolset: install what earns its line.” Type your window size and the tool reports the menu’s share of it, with a warning above 5%, a threshold that is the tool’s and not the book’s.

Two limits belong to the tool itself. Its token counts are per call, and an agent re-reads everything on the desk at every step of its loop, so the agent cost-per-task estimator is the place to see what a larger prefix does to a whole run. And the description check sees words, not behavior, and knows common English phrasings only. The context window budget planner covers the rest of what competes for the window, the M×N protocol calculator covers the other half of this chapter, and the guide to tools and protocols puts skills, tools and protocols side by side. The full section is in Chapter 6, Skills, Protocols, and Interoperability (in the full book).

Questions readers ask

Why do scripts cost almost nothing on the desk?
Because the agent runs a bundled script without reading it. The program's source stays on disk, the file it processes stays on disk, and only the small result enters the context. Chapter 6 says scripts can be executed without ever being read, so their cost on the desk is close to zero however large they grow. The calculator counts the script at zero and does not model its output.
What does a good skill description look like?
It says what the skill does and then when to use it, in the third person, with the concrete words a user would type. The chapter quotes one as the shape to imitate: it lists the actions (extract text and tables, fill forms, merge documents) and then the triggers (working with PDF files, or the user mentioning PDFs, forms or document extraction).
How many skills is too many?
There is no fixed number. Every installed skill adds a menu line to every call, so a large library costs context before any skill fires, and many near-identical descriptions make the choice harder for the model. The chapter's rule is to curate: install what earns its line. Type your window size to see what share the menu takes.
What is progressive disclosure in agent skills?
Progressive disclosure is loading information in layers, each pulled in only when the task shows it is needed. For skills there are three: a short name and description that is always present, an instruction body loaded when the agent judges the skill relevant, and bundled references, templates and scripts fetched one piece at a time.
Can the description check tell me whether my skill will trigger?
No. It looks for word patterns that the chapter's advice implies, in common English phrasings only, and a description can pass all six and still sit unused on a given model or host. Triggering is a judgment the model makes under uncertainty. The only real test is to try tasks that should trigger the skill and tasks that should not, on the setup you actually run.

Sources

  1. Anthropic (2025). Equipping agents for the real world with Agent Skills
  2. Agent Skills (2025). Agent Skills specification
  3. Anthropic (2026). Skill authoring best practices
  4. Lavi Nigam and Shubham Saboo (Google Developers Blog) (2026). Developer's Guide to Building ADK Agents with Skills