Home / Blog / Tools, skills and protocols / Structured Output for Tool Calls: Two Grades of…

Tools, skills and protocols

Structured Output for Tool Calls: Two Grades of Guarantee

Structured output for tool calls comes in two grades. See what each guarantees, what neither covers, and copy a boundary check that holds under both.

By Enrique Gutiérrez · Published · 23 min read

Structured output for tool calls comes in two grades. In the weak grade the model is asked to follow a schema, and it usually does. In the strong grade, constrained decoding removes every token that would break the schema. Both grades leave the same things unchecked: a cut-off output, a refusal, and a well-formed value that is wrong.

I wrote this for an engineer who already ships tool calls and has just met a parse error on a path believed safe, or a rejected schema. By the end you can name the grade each tool runs under, choose a grade by what a malformed call costs, and copy one boundary check that holds under both. Every provider detail below is a snapshot read on 6 October 2026, and those details move.

What is structured output for tool calls, and what are the two grades?

Structured output for tool calls is the model’s request to run a tool, written as an object that fits the tool’s argument schema, so your code can read it without parsing prose. The book’s glossary defines a tool call as “The model’s structured request to run a tool: the tool’s name plus arguments fitting its schema, emitted as structured output instead of prose.” The exchange around that request is the subject of what tool calling is in LLMs.

The grades come from Chapter 2 of AI Agents, Engineered, in the section “Structured Output and Function Calling”, which is free to read online. The chapter warns that the method of compliance matters “because there are two grades of guarantee and they sound alike.”

Asking is the weak grade. The schema sits in the prompt or the tool definition, sometimes with a mode that guarantees the reply parses as JSON. The chapter’s verdict on that mode: “Parsing is a low bar. The model can still omit a required field, invent an extra one, or put a string where the number goes—syntax guaranteed, shape merely hoped for.”

The strong grade is constrained decoding. “Constrained decoding compiles your schema into a filter that runs at every single step, striking from the list each token that would violate the schema before the draw is made.” The glossary files it as “The strong grade of structured output; asking nicely in the prompt is the weak one”.

The text so far feeds the model, which prints odds for the next token (mat 0.42, floor 0.18, sofa 0.13, rug 0.09); the decoder draws one ticket, this time floor, and it fills the open slot in the sentence.
Figure 2.3 Sampling as a weighted lottery run once per token. The model prints the odds and the decoder draws a ticket; the favorite usually wins, but not always, and each draw takes its place in the sentence so far. Reuse this diagram

How does constrained decoding work?

Constrained decoding works by deleting tickets from the lottery in the figure before each draw. The model still prints odds for every token; the decoder discards the tokens that no valid output could continue with, then draws from the rest. In the chapter’s words, “The model chooses freely among valid continuations and cannot emit an invalid one—conformance is enforced in the machinery rather than requested in the prompt.”

Willard and Louf’s 2023 preprint, Efficient Guided Generation for Large Language Models, describes the trick that makes this fast. Generation is treated as a walk through the states of a finite-state machine, with “an index over a language model’s vocabulary” built ahead of time. The set of allowed tokens at each state then becomes a lookup.

Three consequences follow from the mechanism. The filter decides which tokens are eligible; it has no view on which eligible token is correct. It knows nothing about your output limit, so a cut-off output is a valid beginning that never became a valid object. And the schema has to be compiled before generation starts, which is why constrained modes accept only part of the schema language.

Chapter 2 prices the strong grade in one clause: “the price is a little decoding overhead and some limits on how elaborate the schema may be.” The sections below report what the providers’ own pages put inside those limits.

What does each grade guarantee, and what does it leave to you?

The weak grade guarantees nothing formal about shape, its parse-only form guarantees valid JSON for an output that finishes, and the strong grade guarantees that a finished output fits the supported part of your schema. The table is this post’s own arrangement. The first and last rows are the book’s two grades, and the middle row is the chapter’s aside about a mode that only parses.

Grade What is guaranteed What is left open How it fails in production What to validate
Weak: asked (schema in the prompt or the tool definition) Nothing formal. Conformance is a rate you measure Syntax, required fields, types, the closed field list, the allowed values Prose around the object; a missing or invented field; a string where a number goes; a tool name you never offered Everything: completion, parse, schema, values
Weak: asked, with a parse-only mode The text parses as JSON, if generation finished Shape: fields, types, allowed values Valid JSON of the wrong shape; a cut-off object Completion, schema, values
Strong: constrained decoding A finished output fits the part of the schema the stack supports An unfinished output; a refusal; schema features the stack rejects or drops; every value A request rejected before generation; arguments cut off at the output limit; a refusal in place of a call; a well-typed invented value Completion, refusal, the full schema, values

Chapter 2 gives the reason in eight words: “a guarantee about shape says nothing about truth.”

How do you tell which grade a call ran under?

You tell from four facts about the request and the response, none of which is visible in the arguments themselves. The list is mine, assembled from the documentation quoted in the next section.

  • Was the schema sent as a constraint, or only as text the model reads?
  • Did the serving stack accept every schema feature, or did it reject some, or drop some without saying?
  • Did generation end normally, or at the output limit, or with a refusal?
  • Is every path constrained? A fallback model, a second provider or a turn with several calls may run under a different grade.

A dated example of the last item: one provider’s 2024 announcement, read as archived, said its constrained mode was “not compatible with parallel function calls.” Check your provider’s current page.

What does every grade of structured output for tool calls leave open?

Four things stay open under every grade: wrong values, an output cut off at the output limit, a refusal, and schema features the constrained mode does not accept. The book covers the first and gives the fourth one clause. The rest comes from what providers and open stacks say about themselves.

Each row below is a labeled example of a category, quoted as its page stood when I read it on 6 October 2026. Nothing here recommends a provider. Subsets, limits and defaults change, so treat the rows as a dated snapshot and read your own provider’s page.

Source, as read 6 October 2026 What its page says is guaranteed What its page says can still go wrong
Hosted provider A, documentation, undated “Reliable type-safety: No need to validate or retry incorrectly formatted responses” “Structured Outputs can still contain mistakes.” Its 2024 announcement adds: “The model can fail to follow the schema if the model chooses to refuse an unsafe request.”
Hosted provider B, documentation, undated “Structured outputs guarantee schema-compliant responses through constrained decoding” On a refusal, “the refusal message takes precedence over schema constraints”; at the token limit, “The output may be incomplete and not match your schema”
Hosted provider C, documentation last updated 23 September 2026 “syntactically correct JSON” “always validate values in your application”; “Not all JSON Schema features are supported.”
Hosted provider D, release note of 27 November 2024 “the tool calls are guaranteed to adhere to the tool names, parameter names, parameter data types, and required parameters, without the risk of hallucinations.” The note lists no exceptions
Open inference server, documentation, undated When a tool is forced: “You are guaranteed a validly-parsable function call - not a high-quality one.” When the model chooses the tool and the strict setting is off, “arguments may occasionally be malformed or violate the function’s parameter schema.”
Open inference stack, grammar guide, undated Grammars “constrain model outputs” “Unsupported features are skipped silently.”

Provider B’s page on strict tools lists “No need to validate and retry tool calls” as a benefit. The same provider’s page in the table has a section headed “Invalid outputs”. Provider C tells you to “always validate values in your application”. The pages agree once you see which layer each sentence is about: the first is about shape and the last is about values.

What about values that fit the schema and are wrong?

Wrong values pass through every grade untouched, because a schema describes types and the filter enforces eligibility. Chapter 2 says what happens when the model lacks the fact a field asks for: “it will make it produce a well-typed number.”

Provider D’s sentence shows how easily this is misread. It lists tool names, parameter names, data types and required parameters. A reader who stops at “without the risk of hallucinations” will take that to cover values, and the list covers none.

The chapter’s instruction therefore stands under both grades: “validation stays in your code regardless”, and it names ranges, cross-field rules and whether an identifier exists. The post on how to design tools for LLM agents covers the contract side, and I do not repeat it.

What happens when the output is cut off at the output limit?

A cut-off output is a valid beginning of an object, and it will not parse. The constraint held for every token generated; the limit stopped generation before the closing brace.

Provider B documents the case in the table above, and provider A’s page handles an incomplete response in its sample code. A person writing as staff of provider A explained on Hacker News on 6 August 2024 why it can be worse under a constraint: “In JSON mode and Structured Outputs, failures are rarer but more catastrophic. If the model gets too confused, it can get stuck in loops where it just prints technically valid output forever without ever closing the object.”

What it changes is the order of your checks: read how generation ended before you parse. A parse error on a cut-off call means the limit was too low or the model looped.

What happens when the model refuses?

A refusal replaces the call, and both hosted providers that document it let the refusal win over the schema. Provider B’s wording is that “the refusal message takes precedence over schema constraints”.

Code that assumes a constrained reply always parses throws at this point, and the exception hides what happened. Check for the refusal first and route it as its own outcome. My default is to surface it and not ask again.

Which schema features does a constrained mode reject?

A constrained mode accepts a subset of the schema language, the subset differs from stack to stack, and what happens to an unsupported feature differs too. These are the “limits on how elaborate the schema may be” from the chapter, read from three pages on 6 October 2026.

Provider A’s documentation says “all fields or function parameters must be specified as required”, so an optional field has to be expressed another way. Provider B’s page lists “Recursive schemas” and “Numerical constraints (such as minimum, maximum, multipleOf)” as unsupported. It adds: “If you use an unsupported feature, you’ll receive a 400 error with details.” On that date its page also capped a request at 20 strict tools and 24 optional parameters across them, for a reason it states: “Each optional parameter roughly doubles a portion of the grammar’s state space.”

The open stack takes the opposite policy. Its guide says “Unsupported features are skipped silently.” Under that policy your request succeeds and the output is guaranteed to fit a weaker schema than the one you wrote.

Two practical points follow. A schema that works under the weak grade may be rejected under the strong one, and a schema one stack accepts may be rejected by another. Under silent dropping, only a check against your full schema shows what was enforced.

JSONSchemaBench, a 2025 preprint that tested several engines on real-world schemas, reported “the best framework supporting twice as many schemas as the worst.”

Does constraining the output hurt answer quality?

The published evidence is contested. For schema-constrained output, the studies I read report differences of a few points in both directions on reasoning benchmarks. One 2024 paper reports much larger drops in two other conditions, and no study scored calls to real tools.

Study What it tested What it found Its limits
Tam and five co-authors, “Let Me Speak Freely?”, EMNLP 2024 Industry Track Reasoning and classification tasks on three hosted and two open models, under three methods the paper names: JSON-mode (filed under constrained decoding; it “ensures the output is valid JSON”), format-restricting instructions, and conversion from free text afterwards. Schema-constrained output appears in one table, on one model “a significant decline in LLMs’ reasoning abilities under format restrictions”; in one classification table, JSON-mode “performs much better than text” The authors: “we were unable to include results from more powerful language models”; the dataset is “limited in scope”
Will Kurt, “Say What You Mean”, a company blog The same three reasoning tasks, re-run on one open model with prompts matched across conditions Structured 0.78, 0.77 and 0.44 against unstructured 0.77, 0.73 and 0.41 The author’s employer builds a constrained-generation library. I captured no publication date from the page, and no test-set sizes or intervals. It was submitted to Hacker News on 24 November 2024
Geng and eight co-authors, JSONSchemaBench, 2025 preprint The same three tasks on one open model, four open constraint engines against unconstrained output Every engine at or above unconstrained; the introduction says “up to 4%” Prompts follow the rebuttal’s setup. Several authors work at the company behind the best-scoring engine. A preprint, with no intervals in the text I read

One table in the paper that says constraints hurt covers schema-constrained output, on a single small hosted model. There, free text scored 94.57, 82.85 and 83.11 on three reasoning tasks, and schema-constrained output scored 91.71, 81.77 and 86.07. By my subtraction, free text led by 2.86 and 1.08 points on two tasks and trailed by 2.96 on the third.

The larger drops in that paper appear in two other conditions. In the same table its JSON-mode, this post’s parse-only form, sat 6.42 to 7.62 points below free text. In another table, one hosted model scored 86.99 on a math benchmark when the prompt asked for JSON and 23.44 when the prompt also spelled out the schema, with a standard deviation of 22.9 across prompt variants. A second hosted model moved from 89.66 to 89.21.

Field order explains one drop, by the authors’ account. On one task, they report, every JSON-mode response from one model placed the answer key before the reason key, “resulting in zero-shot direct answering instead of zero-shot chain-of-thought reasoning.” Elsewhere they write that failing to reason first caused “a large drop in final performance”.

The counter-evidence is modest too. The rebuttal’s gains are 1, 4 and 3 points by my subtraction. The preprint’s table has unconstrained output at 50.7%, 52.6% and 80.1%; the largest gain on each task is 3.3, 3.3 and 3.7 points, and the smallest on one task is zero. Both sources have vendor ties and the second reuses the first’s prompts, so I do not count them as two independent results.

What can and cannot be concluded?

Three things can be said, and one cannot. Field order matters, because generation is sequential: put any reasoning field before the answer field, and confirm your stack keeps that order. Provider B’s page, as read on 6 October 2026, says required properties are emitted first and optional ones after, which can reorder a schema.

The grammar should leave room for the reasoning. An ICML 2025 paper, CRANE, reports that grammars allowing only the final answer reduce reasoning, and that widening the grammar restores it. A NeurIPS 2024 paper argues that constraining “can distort the LLM’s distribution”. I read only the abstracts of both.

And tell the model about the shape as well as enforcing it. The open inference server quoted above recommends stating the format in the prompt, so that “the model’s intended generation is aligned with the schema that it’s being forced to generate”.

What cannot be concluded is anything about calls to real tools. Every study above scored final answers on reasoning or classification benchmarks. The 2024 paper says its JSON-mode on one provider’s models ran through that provider’s function calling interface, yet what it scored was a benchmark answer. Applying these results to a short call whose schema the model was shown is an inference, so run your own evaluation set with the constraint on and off.

Which grade should each tool get?

Choose the grade of structured output for tool calls one tool at a time, by what a malformed call would cost. This rule is mine. It is reasoned from the mechanism, from the chapter’s closing caution about consequences and from one provider’s tip, and nobody has tested it.

A fair objection comes first. If the boundary check below catches every bad shape, the strong grade looks redundant. It earns its place because your validator is code and can have gaps, and because each bounced call costs a model turn. It also has a price: a narrower schema language, a delay on the first request with a new schema on some providers, and per-request budgets on others.

Provider B’s documentation suggests marking only critical tools as strict when there are many, reserving the mode “for tools where schema violations cause real problems”.

The rule, in three questions:

  • First: if a malformed call slipped past your checks, would the worst case be one wasted model turn? Then asking is enough, with every failure validated and returned.
  • Second: could a malformed call be stored, acted on without an error, or reach a side effect? Then constrain it.
  • Third: does the constrained mode reject part of the schema? For a cheap read, stay with asking. For anything else, rewrite the schema into the supported subset and move the rejected rule into the boundary check.

The buttons narrow the table by consequence.

Tool kind What a malformed call costs Grade What the boundary still checks Consequence
Search or lookup, read-only One wasted turn Asking, validated and returned Query length, result limit, whether this caller may read it low
Read with a one-off schema built per request One wasted turn; a constrained mode may add a first-request delay Asking, validated and returned The full schema low
Read whose schema uses a feature the constrained mode rejects One wasted turn Asking, validated and returned The full schema, then values low
Classifier or router verdict A wrong branch taken with no error Constrained, with a label for “none of these” How often the exit label is used; the reasoning field comes before the label medium
Extraction into a record A corrupt row Constrained Ranges, cross-field sums, that the source text contains the value medium
Write to your own system A bad record or a duplicate Constrained Identifiers exist, the caller may act on them, the retry is safe high
Irreversible or external action: pay, send, delete Money, mail or data that cannot be recalled Constrained, plus an approval step All five checks; shape is the smallest of them high
Write whose schema uses a feature the constrained mode rejects As the two rows above Constrained on a rewritten schema; the rejected rule moves to check 3 The full schema you wrote, then values high

The classifier row is where the strong grade can mislead. The 2024 paper above found its JSON-mode competitive or better on classification tasks, yet a forced choice among labels that do not fit still returns a label. Give the list an exit and watch how often it is taken. How to build an LLM classifier is its own subject, with rubrics and test sets that this post leaves alone.

The consequence tier classifier sorts an agent’s actions into tiers and is a quick way to fill in the last column. The tool schema linter flags two things constrained modes commonly require, a closed field list and a correct required list. It knows no provider’s subset, so it cannot tell you a schema will be rejected.

How much does one percent matter over a run?

It matters by compounding, shown here with invented inputs. If each call conforms with probability 0.99 and calls are independent, the chance of at least one non-conforming call in 20 is 1 − 0.9920, or 18.2%. At 0.999 it is 2.0%.

Both inputs are illustrative; I found no verified public rate of malformed tool calls. The 18.2% is the chance of needing the recovery path, since a caught call can still be repaired. The compounding error calculator does the sum for other inputs.

One tool, walked through

Take set_discount, an illustrative write to your own system whose schema limits percent to whole numbers from 0 to 30. Suppose your provider’s constrained mode rejects numeric bounds.

The guarantee table says the strong grade would cover the field’s name and type and leave the bound open. The consequence rule puts a write in the high tier, and its third question applies: send the constrained mode the schema without the bound, and keep the bound in the full schema your code validates against. A call with percent: 45 then finishes, parses, and fails check 3 of the template below.

The model receives: percent must be a whole number from 0 to 30 (got 45); resend with a value in that range. Two more tools, walked the same way: a read-only search_orders lands in the low tier and stays with asking, and an issue_refund lands in the high tier, constrained, with an approval step at check 5.

What should you check at the boundary, whichever grade you have?

Check five things in order: that generation finished, that the arguments parse, that they fit the full schema you wrote, that the values mean something, and that the call is allowed. The template is this post’s own, written as plain text in the book’s pseudocode style, and it is the same under both grades.

ON EVERY MODEL RESPONSE THAT MAY CARRY A TOOL CALL
A response may hold zero, one or many calls: run the checks once per call.
The checks are the same under either grade. Stop at the first that fails.

1. COMPLETE?  Read how generation ended before you parse anything.
     cut off at the output limit -> do not parse, do not execute.
                                    Raise the limit or ask for less,
                                    then re-ask once. If that is also
                                    cut off, treat it as a loop: fail.
                                    count: truncated
     refusal                     -> do not parse, do not re-ask.
                                    Surface it to the user or operator.
                                    count: refused

2. PARSES?    Is the argument text one valid JSON object?
     no -> return to the model: where parsing stopped, and an excerpt.
           count: did_not_parse

3. FITS?      Is the tool name one you offered? Do the arguments validate
              against the FULL schema you wrote, including every rule the
              constrained mode rejected or did not enforce?
     no -> return: the field, what was expected, what arrived,
           the valid set.
           count: did_not_fit

4. MEANS?     Values, checked in your code:
     in range     amount_cents > 0 and amount_cents <= refundable_cents
     units        the unit is in the field name; reject 12.5 for *_cents
     exists       does this id exist, and did it come from the user
                  or from an earlier result in this run?
     consistent   end >= start; line items sum to the total
     no -> return: what failed, the valid set, what a correct call
           looks like.
           count: did_not_mean

5. ALLOWED?   Authorization, confirmation, limits, for this caller
              and this tool.
     no -> do not execute. Say why, and what the model may do next.
           count: not_allowed

RETURNING A FAILURE TO THE MODEL
  Send it as the tool result, flagged as an error.
  Each return, and each re-ask, spends one from a retry cap
  (illustrative: 2 per call). When the cap is spent, fail the run loudly.

READING THE COUNTERS, PER TOOL
  did_not_parse above zero on a tool you believe is constrained, or
  did_not_fit on a rule the constrained mode enforces
      -> the call did not run under the grade you think it did.
  did_not_mean
      -> the rate no grade can lower. Watch it.

Check 3 validates against the full schema for the reason the open stack’s guide gave: a feature can be dropped silently, and a fallback path may run unconstrained. It costs little, and its counter is how you learn which grade of structured output for tool calls you are running.

Check 4 is where Chapter 5’s advice on inputs applies. Chapter 5, Tools and the Action Space (in the full book) tells you to “Pin categorical values with enums, mark what is required, forbid extra properties so junk cannot ride along.” It also describes how model errors differ from typos, and concludes: “Validate as strictly as you would for a public API, because the caller, however capable, is not a trusted operator”.

A command-line program wrapped as a tool needs this boundary most. A person at a terminal supplied checks 4 and 5 by looking at the screen, and the wrapped program has nobody looking. Building agent tools for software built for humans is the retrofit subject of that same chapter.

Check 5 is the Chapter 2 caution in one line. Its list is “authorization, confirmation, limits”, and a call can pass the first four checks and still be one that should not run.

When should you retry, and what do you send back?

Send a failure back to the model when the model can fix it, which covers checks 2, 3 and 4. A cut-off output wants a larger limit, a refusal wants a person, and a call that is not allowed wants a reason with no second attempt.

What you send back is Chapter 5’s subject. Its test for an error message is three items, “what went wrong, the valid set, the fix implied”, drawn from its own example: severity must be one of: low, med, high (got "urgent"). The chapter’s general form: “Write every error as instructions to a capable colleague who cannot see your face: what happened, what a correct call looks like, what to try next.”

Applied to the template, each return has a fixed shape.

Failed check What the model is sent (illustrative)
2, parses arguments did not parse: prose before the opening brace. Resend the object alone.
3, fits reason must be one of: duplicate, customer_request, fraud, other (got "double charge"). Resend with a valid reason.
4, means no charge CH-2291-8 for this customer. Do not guess another id; look the charge up and copy one from the result.

Sending the identical request again, with no error attached, repeats the draw that produced the fault. The return gives the model something new to condition on.

Then cap it. Chapter 2: “Cap those retries with a hard budget, after which the run fails loudly; a stubborn confusion must never be allowed to loop at your expense.” Timeouts, backoff and retries of writes are a different problem with different tools, and the post on idempotent tools and safe retries owns it.

Where does this stop being reliable?

It stops where the evidence does, and the limits of this post are specific.

The provider details are a snapshot. Every quotation from documentation was read on 6 October 2026, and two sources are from 2024. Subsets, limits and defaults move. The four categories should outlast them: what finishes, what is refused, what is enforced, what is true.

There is no public rate. I found no neutral measurement of how often tool calls are malformed under the weak grade. One provider published figures for its own models in 2024, and the tool-calling post linked at the top reports them.

The quality studies are about answers. None scored calls to real tools, the 2024 paper’s JSON-mode used one provider’s function calling interface, and the two sources reporting gains have vendor ties.

The rule and the template are untested. Both are my synthesis. The retry cap of 2 is illustrative, and leaving a refusal alone is a default you may have reason to change.

The check to keep

Structured output for tool calls gives you a statement about shape, and the grade tells you how strong that statement is. Write down the grade each tool runs under today, move the consequential ones to the strong grade, and put the five checks in front of every tool. Then watch did_not_mean, because it is the count no decoder will bring down for you.

The two grades are in Chapter 2, “The Engine”, free to read online, as is the glossary entry for structured output. The tool interface and its error channel are in Chapter 5, Tools and the Action Space (in the full book). The guide to tools, skills and protocols collects the neighboring posts and tools, and you can see the formats.

Questions readers ask

What is structured output for tool calls?
It is the model's request to run a tool, written as an object that fits the tool's argument schema, so code can read it without parsing prose. It comes in two grades. In the weak grade the model is asked to follow the schema. In the strong grade, constrained decoding makes tokens that would break the schema impossible to choose.
Does strict mode guarantee valid JSON?
For an output that finishes, and against the schema features the provider supports. Two hosted providers' documentation, read on 6 October 2026, names two cases where the output may not match the schema even with the constraint on: the model refuses, or generation stops at the output limit. Values inside the object are never covered.
Why was my schema rejected in strict mode?
A constrained mode compiles the schema into a grammar and accepts only part of the schema language. Which part differs by provider and changes over time: as read on 6 October 2026, one hosted provider's page lists recursive schemas and numeric bounds as unsupported, and another requires every field to be marked required. Read your provider's current list.
Does constrained decoding reduce answer quality?
The evidence is contested. A 2024 paper reports a significant decline in reasoning under format restrictions, tested with format instructions, a JSON-mode and conversion from free text afterwards, and ties a large drop to the answer being generated before the reasoning. On schema-constrained output, that paper, a rebuttal and a 2025 preprint differ by four points or fewer, in both directions. None scored calls to real tools.
Do I still need to validate tool arguments if the schema is enforced?
Yes. Check that generation finished, that the reply is a call and no refusal, that the arguments fit the full schema you wrote, and that the values are in range, in the right unit and refer to things that exist. One hosted provider's page, read on 6 October 2026, says to always validate values in your application.

Sources

  1. OpenAI (2026). Structured model outputs (documentation; one example of a hosted provider's constrained mode, read 6 October 2026)
  2. OpenAI (2024). Introducing Structured Outputs in the API (announcement of 6 August 2024, as archived on 28 December 2024)
  3. Anthropic (2026). Structured outputs (documentation; a second example of a hosted provider's constrained mode, read 6 October 2026)
  4. Anthropic (2026). Strict tool use (documentation, read 6 October 2026)
  5. Google (2026). Structured outputs, Gemini API documentation (a third example; page last updated 23 September 2026)
  6. Cohere (2024). Structured Outputs support for tool use (release note of 27 November 2024)
  7. vLLM project (2026). Tool Calling (documentation of an open inference server, read 6 October 2026)
  8. llama.cpp project (2026). GBNF Guide, grammars/README.md (documentation of an open inference stack, read 6 October 2026)
  9. Brandon T. Willard, Rémi Louf (2023). Efficient Guided Generation for Large Language Models (preprint, arXiv:2307.09702)
  10. Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, Yun-Nung Chen (2024). Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance (EMNLP 2024 Industry Track, pages 1218–1236)
  11. Will Kurt (.txt) (2024). Say What You Mean: A Response to 'Let Me Speak Freely' (company blog; no publication date captured from the page; submitted to Hacker News on 24 November 2024)
  12. Saibo Geng, Hudson Cooper, Michał Moskal, Samuel Jenkins, Julian Berman, Nathan Ranchin, Robert West, Eric Horvitz, Harsha Nori (2025). JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models (preprint, arXiv:2501.10868)
  13. Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova, Loris D'Antoni (2024). Grammar-Aligned Decoding (NeurIPS 2024, arXiv:2405.21047; abstract read)
  14. Debangshu Banerjee, Tarun Suresh, Shubham Ugare, Sasa Misailovic, Gagandeep Singh (2025). CRANE: Reasoning with constrained LLM generation (ICML 2025, arXiv:2502.09061; abstract read)
  15. tedsanders (Hacker News) (2024). Hacker News comment on why a strict mode is optional (6 August 2024)