What is a screen-then-detail cascade?
A screen-then-detail cascade is two classifiers in a row: a cheap one that handles the obvious items, and the full one that sees only what the cheap one would not decide. The estimator above takes your volume, your costs per item and three rates, and shows what the arrangement costs a day against running the full classifier on everything.
Chapter 26 of the book builds the full classifier first: a written guide, chosen examples, a rubric answered before the verdict, and a flag for ambiguous items. The chapter calls that “the full instrument”. The cascade comes in the chapter’s last section, as a response to volume: “At fifty thousand, notice what the full instrument spends per item—the rubric answers, the rationale, real output tokens every time—and reach for the standard remedy: a cheap screen in front of the detailed pass.”
The screen, in the chapter’s words, “can be a smaller model running a slimmed guide (definitions only, no rubric), or sometimes plain rules”. It deals with the unmistakable bulk. The rest moves on, and of the items the full instrument sees, a few carry its ambiguity flag and go to a person.
The figure’s caption says what the widths mean: “Most items never touch the expensive call—which is what makes the cascade pay, and the abstain flag is exactly its escalation trigger.” The estimator redraws those bands as three bars sized from your own numbers.
How does the estimator compute the cost?
The estimator charges every item for the screen and only the escalated items for the full instrument, then compares the sum with the full instrument on every item. The arithmetic is the tool’s. The chapters describe the shape and give no formula.
- Cascade cost a day = items × (screen + gate) + items × escalation rate × full instrument.
- Without a cascade = items × full instrument.
- Cleared and wrong a day = items × (1 − escalation rate) × the screen’s error rate on cleared items.
- Person-hours a day = items × escalation rate × flag rate × minutes per item ÷ 60.
The first line follows Chapter 10’s description of a cascade, where “every request pays the cheap call and the escalated fraction pays twice”. With the defaults, 50,000 items a day pay 300 units each for the screen, and the 7,500 that escalate pay 2,000 more. That comes to 30 million units against 100 million for the full instrument alone, so the cascade costs 30% of the alternative.
Costs are in whatever unit you count in: tokens, or tokens weighted by how your provider bills input and output. The page holds no prices and none are needed, because the comparison is a ratio. The defaults are illustrative and are not measurements of any system.
What is the break-even escalation rate?
The break-even escalation rate is the share of escalated items at which the cascade costs exactly as much as having no cascade. The tool computes it as 1 − (screen + gate) ÷ full instrument, and marks it on the chart with a dashed line.
With the defaults the break-even is 85%. A screen that costs half as much as the full instrument breaks even at 50%. A screen that costs as much as the full instrument never breaks even, and the tool says so in its headline.
The chart shows why this number is worth knowing before you build anything. The cascade’s cost is a straight line that rises with the escalation rate, and the full instrument alone is a flat line. Your point sits somewhere on the rising line. The distance below the flat line is your saving, and it shrinks with every item the screen cannot decide.
Why is a high escalation rate a warning?
A high escalation rate is a warning because it means the screen is not doing the one job it was added for. Chapter 19 puts it in a sentence about cost levers in general: “a cascade where half the traffic escalates is a cheap tier that is failing its job”.
The estimator warns at 50% and above. Treating 50% as a hard line is the tool’s reading of that sentence, and the chapter does not state a threshold. A cascade at 50% can still cost less than the full instrument, which is why the tool reports the break-even separately. The two numbers answer different questions: whether the cascade pays, and whether the screen is any good.
The same chapter says what to do about it: “use the cheapest evaluator that works, and watch the escalation rate”. The tool adds a reading of its own: an escalation rate that climbs over the weeks suggests the traffic has drifted away from what the screen’s definitions cover.
What does a separate quality gate cost?
A separate quality gate costs its price on every item, including the easy ones the screen got right. Some cascades let the screen set its own abstain flag. Others run a second model over the screen’s answer to decide whether to escalate. The gate field is for the second kind, and it defaults to zero.
Chapter 19 names this failure “spending the savings on the referee”, and describes it: “if the quality gate in your cascade is itself an expensive model, every request (including the easy majority) pays for the judge”. The estimator adds the gate to every item, lowers the break-even accordingly, and shows what share of the cascade’s cost the gate accounts for.
Chapter 26’s own design avoids the separate gate. The full classifier already has an ambiguity flag, and of that mechanism Chapter 26 says: “The abstain machinery you built in the reasoning section turns out to be exactly the escalation trigger the cascade needs.”
What does the screen let through unchecked?
The screen lets through every item it clears with confidence, and some of those labels are wrong. Nothing downstream looks at a cleared item, so these are the errors the cascade cannot catch. The estimator multiplies the cleared volume by the error rate you enter and reports the count per day.
This is the reason for the screen’s design rule. The chapter says “its one design requirement is that it knows what it may not decide: anything uncertain, low-confidence, or flagged goes onward to the full instrument rather than being guessed at.” A screen tuned to escalate less will look better on the cost table and worse on this count.
Chapter 25 explains why the count matters more than overall accuracy. The two kinds of triage error do not cost the same: “A ticket wrongly escalated costs a specialist a shrug.” A ticket wrongly cleared is found late or never. So, the chapter says, “you must monitor the false-confident rate (how often the system was sure and wrong) rather than the accuracy number alone”.
The tool cannot know your screen’s error rate. You measure it by running the screen over a labeled sample and counting the wrong labels among the items it did not escalate. The three-set splitter prepares such a sample, and the eval sample size calculator says how many rows you need before a small rate means anything.
How much work reaches a person?
The work that reaches a person is the flagged share of the escalated items, which is usually a small number. The figure’s caption describes it: “Of those, only a thin flagged trickle reaches a person.” With the defaults, 750 items a day are flagged, and at two minutes each that is 25 person-hours.
The estimator reports those hours beside the machine cost and does not add them together. Converting hours into the same unit as tokens would need a price for both, and the site carries no prices. The tool also leaves out the comparison case: the full instrument running alone would flag items too, and that workload is not modeled.
The flagged items are worth more than their handling time. The chapter says “the flagged stream doubles as the best telemetry the instrument produces.” Each one is a boundary rule the guide has not written yet, or a kind of input nobody expected.
When is a cascade not worth building?
A cascade is not worth building when the volume is small, when a person is waiting on each answer, or when the screen is barely cheaper than the full instrument. The chapter is direct about the first case: “At fifty tickets a day, run the full instrument on everything and think no more about it.” It gives no rule for the volumes in between, and the tool does not invent one. It shows the absolute saving per day and leaves the judgment to you.
The second case is about time. Chapter 10 says of the cascade that “its price is latency, since every request pays the cheap call and the escalated fraction pays twice, waiting on two models in sequence.” It reports where practitioner guidance draws the line: a cascade is “a natural fit for batch and background work, often unacceptable on a path where a person is waiting.” The estimator counts how many items a day pay that double wait. For the tokens inside a single call, the agent cost-per-task estimator does the arithmetic.
One assumption sits under every number on this page: both stages judge each item alone. Chapter 26’s operating rule is to “classify every item on a fresh desk”, and it explains the reason: “amnesia between items is what makes their verdicts comparable.” A design that batches many items into one call would have a different cost per item and a different error rate, and the estimate would not describe it.
The guide to agent patterns places the cascade beside routing, its close relative, which decides up front where each item goes. The full chapter, with the classifier that both stages are built from, is Chapter 26, Agents as Classifiers and Scorers (in the full book).
Questions readers ask
- What must the cheap screen be able to say?
- It must be able to say that it does not know. Chapter 26 gives the screen one design requirement: it knows what it may not decide, and anything uncertain, low-confidence or flagged goes onward to the full instrument instead of being guessed at. A screen that always answers will always be cheap, and its mistakes will leave the system with no second look.
- Why does the estimator warn at a 50% escalation rate?
- Because the screen exists to keep most items away from the expensive call. Every item pays for the screen, and each escalated item then pays for the full instrument as well. Chapter 19 says a cascade where half the traffic escalates is a cheap tier that is failing its job. The estimator turns that sentence into a warning at 50%, and it also shows the exact rate at which your cascade stops saving anything.
- Why does every item get a fresh context?
- So that each item is judged by the same instrument. Chapter 26 gives two reasons. A long history of earlier verdicts degrades the judgment of later ones, and earlier verdicts in the window nudge the next one toward the same label. A fresh context also keeps the cost per item flat, which the estimate assumes.
- What is the break-even escalation rate?
- It is the escalation rate at which the cascade costs exactly what the full classifier alone would cost. The tool computes it as one minus the per-item cost of the screen and gate divided by the per-item cost of the full instrument. With a screen at 300 units and a full pass at 2,000, the break-even is 85%. Above it, the cascade costs more than having no cascade.
- Does the estimate include the cost of the people who review flagged items?
- It reports their time and does not price it. The tool shows how many items a day reach a person and the hours that takes at the minutes per item you enter. Adding those hours to a token count would need a price for both, and the page holds no prices. It also compares machine cost only: the full classifier alone would flag items too, and that workload is not modeled.
Sources
- Lingjiao Chen, Matei Zaharia, James Zou (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (arXiv:2305.05176)
- Isaac Ong, Amjad Almahairi, Vincent Wu, et al. (2024). RouteLLM: Learning to Route LLMs with Preference Data (arXiv:2406.18665)
- C. K. Chow (1970). On Optimum Recognition Error and Reject Tradeoff (IEEE Transactions on Information Theory 16(1):41–46)
- Paul Viola, Michael Jones (2001). Rapid Object Detection Using a Boosted Cascade of Simple Features (CVPR 2001)