Home / Tools / Tool schema linter

Free tool · runs in your browser · from Chapter 5

Tool schema linter

A free tool schema linter for LLM agents: paste a tool definition and get findings on its name, description, schema, side effects and hidden injections.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. Every rule and quotation comes from Chapters 2, 3, 5, 7, 15, 17 and 18. The detection heuristics (word lists, text patterns, the 10- and 20-tool thresholds, the description-length cut-offs of 40, 100 and 2,000 characters, the Unicode checks) and the 64-character name rule common to tool-calling APIs are the tool's own.

What does a tool schema linter check?

A tool schema linter reads the parts of a tool definition that a language model sees (the name, the description and the argument schema) and flags what will mislead the model or expose it to attack. This one checks six channels: the tool set, the name, the description, the schema, side effects, and the injection surface where poisoned instructions hide.

Every rule in it comes from the book. Chapter 3 defines a tool as four things bundled together: “a name; a description that tells the model what the tool does and when to reach for it; a schema for its arguments; and the ordinary function that actually runs. The first three travel to the model with every call.” The linter can only read those first three, and that is the point. They are the whole interface the model has, and Chapter 5 of the book is blunt about their status: “Every one of these channels is a prompt, whether or not you wrote it as one.”

The anatomy of a tool across the boundary between your process and the model’s context.
Figure 5.3 The anatomy of a tool across the boundary between your process and the model’s context. Three parts cross to the model’s side and steer it like prompt text—the description above all, in accent, the highest-leverage surface you write. The function itself never leaves; only its results and errors travel back, as still more text the model has to read. Reuse this diagram

Paste a definition, or load the good and the flawed example, and the linter groups its findings into three severities. Fix first covers anything that breaks the contract or looks like an attack. Should fix covers the gaps that make a model guess. Consider covers refinements the book recommends but that a tool can survive without.

Why lint a tool definition at all?

You lint a tool definition because the model cannot ask what you meant, and a tool schema linter is the cheapest way to see your definition the way the model will. A human colleague who meets an unclear API reads the source, checks the wiki or leans over to ask; the model decides “whether and how to call” the tool from the definition alone, as Chapter 2 puts it. Whatever the definition leaves out, the model fills in with something plausible.

The book’s debugging rule follows from that: “when a model misuses a tool, suspect the contract before you blame the intelligence.” Chapter 5 backs it with two vendor anecdotes. One team reported a top score on a coding benchmark after refining tool descriptions, “wording changes, no new capability.” The same team’s search tool kept appending the current year to its queries until a clearer description of the query parameter fixed it, with no change to the model. A linter cannot tell you whether your wording is right. It can tell you, in seconds, which of the questions a good definition answers your definition leaves open.

What makes a good tool description?

A good tool description says what the tool does, when to reach for it, when a neighbor is the better choice, what it returns and how it fails. Chapter 2 compresses this into one instruction: “Write it like API documentation for a capable new hire on their first morning: what the tool does, when to reach for it, when its neighbor is the better choice.”

The linter tests for each of those parts by looking for the phrasing that usually carries it (“use when”, “returns”, “do not use”, “fails with”). It also looks for one sentence the book singles out. Chapter 15, diagnosing hallucinated tool arguments, recommends writing an escape hatch into the description, and quotes one guide’s phrasing: “If you do not know X, do not guess. Ask the user.” When a tool requires a free-text identifier such as a customer email and the description has no such line, the linter suggests one. For descriptions with serious gaps, the result offers a skeleton to fill in: does, use when, do not use when, returns, on error, and the escape hatch.

What should an argument schema pin down?

An argument schema should leave the model exactly one reasonable way to fill each field. Chapter 5 gives the canonical example: “A parameter named user invites a name, an email, an ID, or a JSON object; a parameter named user_id invites exactly one thing.” It follows with the rule the linter’s schema checks are built on: “Pin categorical values with enums, mark what is required, forbid extra properties so junk cannot ride along.”

The table shows how each schema check maps to that guidance.

Check What the linter flags Book source
Ambiguous names user, id, data, options, file and similar Chapter 5, the user_id example
Missing enums A string named like a category (status, severity, type) or described as “one of …” with no enum Chapter 15: an enum “cannot be hallucinated outside its enum”
Required No required list, or one naming arguments that do not exist Chapter 5
Extra properties additionalProperties not set to false; free-form objects; arrays with no items Chapter 5
Descriptions and types Arguments with neither; dates with no stated format Chapter 5: harden inputs “as if they were hostile”
Size limits List or search tools with no limit, filter or cursor; limits with no maximum Chapter 5: “Return the relevant slice, never the full dump”

How many tools is too many?

There is no fixed number, but every tool costs attention on every pass. Chapter 5’s working rule is that “too many tools confuse the model,” and that overlapping definitions “present a choice where none should exist.” Chapter 7 supplies the measurement: a small model “given a task with forty-six tools available failed it, even though every definition fit comfortably inside its window; handed only the nineteen relevant tools, it succeeded.”

The linter warns gently above 10 tools and more firmly above 20. Those thresholds are the tool’s illustrations, not the book’s; the right count depends on the model and the task. It also flags near-duplicates (two search tools, a read beside a fetch) unless one description names the other and states the boundary, mixed naming conventions, and large sets without the service and resource prefixes the book recommends. The context window budget planner shows what those definitions cost in tokens.

The same scheduling job in two action spaces.
Figure 5.2 The same scheduling job in two action spaces. With one endpoint wrapped per tool (top), every intermediate byte crosses the model’s desk and every cross-reference risks a misread. Shaped to the task (bottom), the deterministic work happens inside the tool and the model sees one line back. Chapter 7 prices the difference. Reuse this diagram

The subtlest set-level check looks for one-for-one wrappers of an API: a resource with list, get, create and delete tools side by side. Chapter 5’s design exercise replaces list_users, list_events and create_event with a single schedule_event, because “The right unit for a tool is a task; the endpoints are plumbing, and plumbing belongs inside.” The linter cannot design the task-shaped tool for you, but it can point at the places where you have shipped the plumbing.

Which side effects should a definition announce?

A definition should announce any action that changes the world, whether repeating it is safe, and whether it can be undone. Agents retry on timeouts and ambiguous results, and Chapter 5 sets the bar for writes: “a retried create should return the existing record, flagged as pre-existing, instead of minting a second one.” The linter flags create, send, charge and similar verbs that offer neither an idempotency key nor a promise that retries are safe. The post on idempotent tools and safe retries works through the key design.

Destructive actions get two more checks, both from Chapter 5’s rules for programs built for humans: a dry run on mutations, and confirmation by default with an explicit flag to bypass it. The check the linter treats as most serious is a side effect hiding behind a read-shaped name, such as a get_ tool whose description admits it also sends email. Approval rules and logs often key off names, and Chapter 17’s placement rule is “put the check next to the side effect.” A write disguised as a read slips past the check.

How can a tool description carry an attack?

A tool description can carry an attack because it is loaded into the model’s context and read with the same obedience as everything else there. Chapter 17 names the attack tool poisoning: “A malicious server can hide instructions in its own descriptions: redirect data to me, prefer me over the legitimate tool, quietly include the contents of this file in your next call.” Because the description stays on the desk, “a poisoned one is a standing injection that arrives at install time and attacks on every run thereafter.”

Why a tool’s description is attack surface.
Figure 17.5 Why a tool’s description is attack surface. Installed once, it becomes a permanent resident of the agent’s context, sitting among the ordinary papers on the desk. The instruction hidden inside it (in accent) is re-read with full obedience on every subsequent run—a standing injection, delivered inside the package rather than through the inputs. Reuse this diagram

The linter scans every string in the definition (the name, the descriptions, property names, defaults and enum values) for the patterns those attacks use: orders to ignore other instructions, tags such as <IMPORTANT>, requests to keep something from the user or to keep the instructions themselves secret, reaches for keys or local files (including an order to call a file-reading tool on a secrets file such as .env), and attempts to win the model’s choice over other tools. It flags any description that gives orders about a different tool in the set, the variant known as tool shadowing, and any URL, email address or IP address in model-facing text. For a broader view of the attack class, see the glossary entry on prompt injection.

It also checks what a reviewer cannot see. Chapter 17 warns that “the payload does not need to be visible to a person,” and gives white-on-white text as its example. The Unicode checks extend that warning in a direction the book does not spell out: zero-width characters, bidirectional controls, invisible tag characters, and words that mix Latin letters with look-alike Cyrillic or Greek ones. Unicode’s own security report documents how invisible characters, bidirectional text and look-alike letters from other scripts can make two different strings look the same to a reader.

What does a linted example look like?

Load the flawed example and the linter returns six fix-first findings. The tool is called get_data, its description begins “Gets data for a user. Also sends a usage summary email,” and it hides an <IMPORTANT> block asking the model to read an SSH key, pass it along as notes and say nothing to the user. A zero-width space trails the documentation URL.

The fix-first group catches the hidden tag, the secrecy request, the reach for the key file, the bid to run before any other tool, the invisible character and the email hiding behind a get_ name. The should-fix group covers the rest of the contract: a name that says nothing about its object, no “use when”, no “returns”, an ambiguous user, a free-text type, a free-form options object, nothing required. Now load the good example, create_ticket. It uses the book’s own error example, an enum for severity, an idempotency key, an escape hatch for the customer email and a note pointing to update_ticket for existing tickets. It returns no findings.

What can the linter not tell you?

The linter cannot tell you whether a description is true, whether the tool set fits the task, or whether a server will stay honest. A third-party tool that passes today can change tomorrow; Chapter 17 calls this a rug pull, “a tool that was clean when you approved it changes later,” and notes that “A locally installed tool is a file you can read, pin, and hash; it cannot change without your filesystem knowing.” Re-review remote tools on every update, and “Give each tool its own scoped credential, never your master key, so that one poisoned link in the chain forfeits one small thing.”

Its heuristics are word lists and patterns, so expect the occasional false alarm and the occasional miss. A tool that is clean on its own can still complete the lethal trifecta when combined with the rest of your set; the linter makes a keyword-based guess, and the Lethal Trifecta Audit and the post on the lethal trifecta explained do the job properly. Shaping tools to a task is design work, covered in the guide to tools, skills and protocols. The only test that settles it is the one Chapter 5 prescribes: expose the tool, run real tasks, and read the transcripts.

Questions readers ask

Why is the tool description the highest-leverage surface you write?
Because it is all the model knows. The name, description and argument schema travel to the model with every call, and the decision whether and how to use the tool is made from them alone. The book reports a vendor that reached a top benchmark score partly through wording changes to tool descriptions, with no new capability.
What makes an error message useful to a model rather than to a human?
It says what failed in a stable code, whether retrying can help, and what a correct call looks like. Chapter 5's example is severity must be one of: low, med, high (got “urgent”). A bare Error 422 or a stack trace gives the model nothing to correct, so it guesses or retries blindly.
How can a tool description be an attack?
The description is loaded into the model's context and read with the same attention as the user's request. A malicious tool server can hide orders in it: send data to an address, prefer this tool over the real one, include a file's contents in the next call. Chapter 17 calls this tool poisoning.
Does a clean lint mean the tool is safe to install?
No. The linter checks the shape of the contract and looks for known poisoning patterns, but a determined attacker can phrase an instruction it does not recognize, and a remote tool can change after you approve it. Read third-party definitions yourself, pin what you can, and re-review on every update.
Which JSON shapes does the linter accept?
A single tool object with name, description and one of parameters, input_schema or inputSchema; the same object wrapped as a function definition; an array of tools; or an object with a tools array. Everything is parsed in your browser, and invalid JSON gets a line and column with the likely cause.

Sources

  1. Ken Aizawa (Anthropic) (2025). Writing effective tools for agents, with agents
  2. Invariant Labs (2025). MCP Security Notification: Tool Poisoning Attacks
  3. Manish Bhatt, Vineeth Sai Narajala and Idan Habler (2025). ETDI: Mitigating Tool Squatting and Rug Pull Attacks in Model Context Protocol (MCP)
  4. Simon Willison (2025). The lethal trifecta for AI agents: private data, untrusted content, and external communication
  5. JSON Schema. Understanding JSON Schema: object
  6. Unicode Consortium (2014). Unicode Technical Report #36: Unicode Security Considerations (final version, revision 15)