How to monitor AI agents in production comes down to five signals that uptime cannot show: outcome quality on a judged sample of real runs, cost per task, gate and guardrail hits, behavior shape, and drift against your offline eval baseline. Give each an alert rule sized to the runs you score, and watch the monitor too.
The grouping into five is mine. Chapter 15 of AI Agents, Engineered (in the full book) names the ingredients in its section on production monitoring: quality, cost, latency, step counts, tool errors, finish reasons and drift. It does not number them, and the gate signals come from Chapters 12, 17 and 19. Every threshold below is illustrative, a starting point to replace with your own baseline.
I wrote this for the person who decides how to monitor AI agents in production and then asks a team for the dashboard, which is why it ends in a spec you can paste into a ticket.
The post assumes each run already leaves a trace, the book’s “complete, structured record of one run”. What a trace should record is the subject of agent observability and of the post on LLM tracing with OpenTelemetry. Here the unit is the population of runs, read week by week.
Why is uptime not enough to monitor AI agents in production?
Uptime, latency and error rate describe the service, and an agent can keep all three green while it gives wrong answers. Chapter 15 quotes a tooling vendor’s survey of debugging practice on why: “Most agent failures do not trigger visible errors because the system still returns a successful status code even when the result is wrong.”
Keep the classic signals. The Google SRE book’s chapter on monitoring, written by Rob Ewaschuk, opens its list with “The four golden signals of monitoring are latency, traffic, errors, and saturation.” They tell you the agent is reachable and fast. They say nothing about whether the refund it just approved matched the policy.
The chapter puts you inside that gap with a scene. Every dashboard is green, and for six days the agent has been answering a growing share of billing questions with last quarter’s refund policy, because a reorganized document still wins the search. The chapter’s verdict on that dashboard: “Your monitoring is faithfully reporting the health of everything except the product.”
An Ask HN post from March 2026 described monitoring that “tells us if the API is up and how fast it responds, but nothing about whether the outputs are actually good”. Nobody replied.
What did two providers’ postmortems say about passing evals?
Two model providers have published postmortems in which a quality regression reached users while their evaluations looked fine. Both incidents happened inside a model provider, upstream of any agent built on its models, and I read them for one thing: how the drop was detected.
OpenAI’s postmortem of 2 May 2025 covers a model update rolled out on 24 and 25 April 2025 that made responses sycophantic (flattering and agreeable past the point of usefulness); the rollback began on 28 April. The company wrote that “our offline evaluations—especially those testing behavior—generally looked good,” and that “the A/B tests seemed to indicate that the small number of users who tried the model liked it”. It also wrote: “We also didn’t have specific deployment evaluations tracking sycophancy.”
Anthropic’s postmortem of 17 September 2025 covers three infrastructure bugs that intermittently degraded response quality between early August and early September 2025. One routing bug “initially affected 0.8% of requests,” and early reports were “difficult to distinguish from normal variation in user feedback.” The company’s own summary is blunt: “The evaluations we ran simply didn’t capture the degradation users were reporting.” Among its fixes, it committed to “run them continuously on true production systems”.
What they show is the shape of the gap: offline evals passed, a small A/B test looked good, and the regression was found in production traffic and user reports. An online signal you never defined cannot alert you.
Did the model get worse, or did its behavior change?
Hosted models do change between snapshots, and whether a change counts as worse depends on what your eval measures. Chen, Zaharia and Zou compared two snapshots of the same hosted model, three months apart in 2023. Their first version reported accuracy at identifying prime numbers falling from 97.6% to 2.4%; the revised abstract, on a test of primes and composites, reports “84% accuracy” falling to “51% accuracy.”
Narayanan and Kapoor answered within days (19 July 2023), before that revision. The first version’s test numbers were all prime, so, as I read their argument, the test could not separate a model that reasons worse from one that changed which answer it leans toward. Their conclusion: “everything in the paper is consistent with the behavior of the models changing over time. None of it suggests a degradation in capability.”
Both sides agree on the part that matters to you. The rebuttal itself says “Behavior drift makes it hard to build reliable products on top of LLM APIs.” I read the paper’s abstracts only, and the pair is old by this field’s standards; I use it for the argument, which is that a baseline you trust is what turns “it feels different” into a measurement.
How to monitor AI agents in production: what are the five signals?
The five signals are outcome quality, cost per task, gate and guardrail hits, behavior shape, and drift against the eval baseline, plus a coverage line that says whether the monitor is running. Each row names the metrics the book uses and where they live; the grouping, the coverage line and every threshold are mine and illustrative.
The buttons narrow the table by who hears the alert. A page wakes the on-call engineer; a ticket goes to the agent’s owner the same day; daily and weekly rules land in a review.
| Signal | What it catches | How to measure | Alert rule (illustrative) | Chapter | Who hears it |
|---|---|---|---|---|---|
| 1. Outcome quality | Wrong answers behind a green status; tasks reported done that were not | Programmatic checks on every run where one exists; a calibrated judge on a random sample of finished runs; rephrasing and abandonment on every run | Page: a programmatic check fails on more than 5% of an hour’s runs, with at least 200 runs. Weekly: judged pass rate below the trailing four weeks with two-sided p < 0.05 | 15 (task success, output quality); 16 (online layer) | page, weekly |
| 2. Cost per task | Loops, retry storms, swelling context, a model swap that changes token use | Cost per run from token counts and a versioned price table; p50 and p95; cost per successful task | Page: a run’s spend passed its hard cap, which means the cap failed. Ticket: median cost per task at 2× its trailing four weeks. Daily: p95 cost per run above 3× the expected cost per task | 15 (cost per run and per day); 19 (caps) | page, ticket, daily |
| 3. Gate and guardrail hits | A gate nobody reads any more; a guardrail facing an attack or blocking real work; budgets running out | Approvals requested and granted, time to approve; escalations by trigger; guardrail blocks by rule; runs ended by a budget cap | Daily: blocks, escalations or approval requests of one type above 3× their baseline. Weekly: approval rate above 98% for three weeks; any rise in runs ended by a cap | 12, 17, 19 (my grouping; not a Chapter 15 signal) | daily, weekly |
| 4. Behavior shape | Wandering agents, failing tools, empty or truncated tool results, upstream changes of shape | Step-count distribution; error and retry rate per tool; finish-reason mix; result size per tool, with the share of empty results | Page: one tool’s error rate above 5× its baseline over 15 minutes, with at least 50 calls. Daily: step-count p95 up 30% on four weeks; finish-reason mix shifted; a tool’s result sizes jump or its empty share triples | 15 (step counts, tool errors, finish reasons, result size) | page, daily |
| 5. Drift against the eval baseline | An alias rolled forward to a new model, a prompt edit, a moved corpus, a changed user mix | Requested and responding model, prompt version and prompt hash on every run; the offline suite re-run against the production configuration; week-over-week on rows 1 to 4 | Ticket: the responding model or the prompt hash changed with no deploy. Weekly: the scheduled offline re-run scores below its release baseline | 15 (drift, week-over-week rule); 16 (offline suite) | ticket, weekly |
| Coverage (the monitor’s health) | An online evaluator skipping runs, judgments dropped under rate limits, alerts switched off | Scored runs ÷ eligible runs per day; failed judgments; when each alert rule last ran; date of the last judge calibration | Daily: coverage under 80% of the planned rate; failed judgments not retried; any alert rule silent for over 24 hours; calibration older than the last judge or rubric change | Mine, from tool issue trackers; judge calibration from 16 | daily |
Notice how few rows page. The SRE book’s rule is that “Every page should be actionable,” and it names the cost of breaking it: “When pages occur too frequently, employees second-guess, skim, or even ignore incoming alerts”. The pages above fire on frequent, unambiguous events: a cap that failed, a tool error rate, a deterministic check failing at volume. Judged quality, gates and drift are read in a review, because the sample behind them is small and the book’s advice for drift is “it is a trend, so alert on trends”.
How do you score outcome quality on live runs?
Score outcome quality in three layers: a programmatic check on every run where one exists, a calibrated judge on a random sample of finished runs, and implicit user signals on every run. The book’s wording for the first two is that quality is “scored programmatically where a check exists (the answer parses, the citation resolves, the refund matches the policy) and by a calibrated model-graded judge where it does not”.
The judge is an LLM-as-a-judge: a model with a written rubric, scoring outputs the way a reviewer would. Chapter 15 warns that “judge scores on live traffic inherit every bias that chapter dissects, and they cost latency and money besides”. That cost is why the judge samples, and why the sample size decides what it can see (worked through below).
The implicit signals are free. Chapter 16 (in the full book) puts it in one line: “A user who immediately rephrases the same request has graded the first answer, for free.” Explicit thumbs are sparser; one engineer called them “a weak signal” in May 2025, and OpenAI’s postmortem above reported that a user-feedback reward “can sometimes favor more agreeable responses”.
Deciding what counts as success for a task comes before any of this. That question, how you know if an AI agent is working at all, has to be answered task by task; the dashboard only counts the answers.
Why track cost per task as well as the monthly bill?
Track cost per task because the invoice arrives weeks after the run that caused it, and a loop or a retry storm shows up in the per-task distribution the same day. Chapter 15’s rule for choosing tiles covers both halves: “The signals worth wiring are the ones a user would recognize as the product succeeding or failing, plus the ones that predict your invoice.”
The page rule sounds odd at first: alert when a run spent past its hard budget. A cap should stop the run, so a run that passed it means the cap failed, and that is worth waking someone for. Chapter 19 (in the full book) states why caps exist: “a capped run has a bounded worst-case bill and an uncapped one does not”.
The daily rule compares p95 cost per run with the expected cost per task, which you can estimate before launch with the agent cost per task estimator. A doubling of the median cost per task opens a ticket the same day; like every multiple here, the 2× is illustrative. Report cost per successful task beside it; a cheaper agent that fails more often can raise that number while the average falls. Bringing the number down is its own job, the subject of reducing AI agent costs; the monitor’s job is to notice when it moves.
Why count gate and guardrail hits?
Count them because a gate or a guardrail is also a sensor, and its counts reveal problems that no quality score shows. This row is my grouping: Chapter 15 does not list gate hits as a monitoring signal, and the material comes from the chapters on oversight, security and cost.
An approval gate is, in the book’s glossary, “A point where the run halts until a person approves, rejects, or edits the proposed action.” An approval rate that creeps toward always-yes means the person has stopped reading. Chapter 12 (in the full book) says it in two sentences: “Over-gating does more than annoy. It defeats the gate.” Chapter 17 adds: “Prompting on everything trains the click that defeats the prompt.”
A surge in approval requests counts too: it says the agent now wants a signature for work it used to do alone. The other counts read the same way. A spike in guardrail blocks of one type is either an attack or a rule blocking real work, and both deserve a look the same day. A rise in runs ended by a budget cap says the agent needs more steps than you planned, or that it has started to wander.
What does behavior shape catch?
Behavior shape catches the agent doing its work differently: more steps, a failing tool, a result that comes back empty or cut short. Chapter 15 names “the step-count distribution, which is your loop detector”, and beside it the error and retry rate per tool.
Finish reasons (the reason each generation stopped) get their own tile in the chapter: “A drift toward ‘ran out of tokens’ or toward content filtering is an early, cheap symptom that something upstream changed shape.” The chapter’s bug taxonomy adds result size: “record every tool result’s size as a span attribute, alert when a tool’s result distribution jumps”.
I add the share of empty results per tool to that rule. A tool that returns an empty list with a success status raises no error, so the error-rate page never sees it; the size distribution does.
How do you see drift against the eval baseline?
You see drift by comparing this week with earlier weeks and with the offline baseline, and by recording on every run the facts that change silently. The chapter names three routes: users change, the model behind an alias changes, the knowledge the agent retrieves changes. “Each route produces the same maddening symptom, a quality score sagging a few points with no commit to blame.”
The chapter’s rule for that slide: “A week-over-week regression rule catches a slow slide that no absolute threshold will, because the numbers pass through every threshold slowly, one un-alarming day at a time.” The fast half of drift detection is bookkeeping. Record the requested model, the responding model, the prompt version and a hash of the rendered prompt; when any of them changes without a deploy, open a ticket that day. The slow half is re-running the offline suite against the production configuration on a schedule, so the baseline is measured on what users get.
How do you evaluate AI agents in production with online evals?
You evaluate AI agents in production by scoring a sample of live runs, which is online evaluation, and feeding what it finds back into the offline suite you run before every release. Chapter 16 draws the line: offline evaluation answers “did this change help?”, and it cannot answer “what is happening out there?”
The two layers form one loop, in the chapter’s words: “Online surfaces the failures you never imagined; each becomes an offline task; the offline suite guards it forever.” An eval set grown this way stays close to real traffic. Chapter 15 gives the same loop from the other end: “a flagged trace is an eval case that has not been written down yet”.
Neither layer is enough alone. A passing suite before release still leaves the gap the two postmortems describe, and Chapter 16 says so plainly: “a green suite is evidence, never proof”.
What should the sampling rule be?
Run cheap programmatic checks on every run, a judge on a random sample sized by the arithmetic below, and the judge on every run only where the action is consequential and the volume is low. That last clause is my rule. For capture, Chapter 15 advises keeping everything in the first weeks, then to “sample the successes down to a fraction while keeping every errored run”.
Hamel Husain and Shreya Shankar’s evals FAQ (modified September 2026) adds a guard against sampling only what you already suspect: “Keep some random traces in every batch.” A sample chosen by your current signals can only find the failures those signals describe.
Products in this category expose much the same controls on an online evaluator. Two examples, read in October 2026: one product’s docs describe a filter, a sampling rate and a weekly spend limit on the judge. A second product scores matching traces “after the trace has gone idle,” so a long run is judged once it has finished. The concept is the four controls: which runs, what share, when, and at what cost.
How many live runs does the judge need to score?
The judge needs enough scored runs that the smallest drop you would act on stands out from the noise, and that is usually more than a week of sampled traffic on a small agent. One engineer named the problem on Hacker News in 2024: an A/B test needs “a big % difference” or “a big number of samples,” and “LLM apps usually don’t meet either of those two criteria.”
Here is the worked read, with illustrative inputs. The baseline is the trailing four weeks: 360 of 400 judged runs passed, 90%. This week the judge passed 41 of 50, which is 82%, eight points down, and someone has already said “quality dropped” in a meeting.
Run a two-proportion test (the standard check of whether two pass rates differ by more than chance) and the two-sided p is 0.087. The 95% Wilson interval for this week runs from 69.2% to 90.2%, and the baseline’s 90% sits inside it. At fifty judged runs, an eight-point drop is still within noise.
The same 82% on 200 judged runs gives p = 0.0055, and this week’s interval (76.1% to 86.7%) no longer contains the baseline’s 90%. The calculator below opens on the 50-run case; Version A is the baseline window and Version B is this week, and the tool reports about 295 runs per period to detect this gap reliably.
With JavaScript on, the Eval sample-size calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Eval sample-size calculator on its own page to share a result by link.
Pick the smallest drop you would act on, then compute the judged runs that make it visible. With equal periods, 80% power and a 5% false-alarm rate, a fall from 90% to 80% needs about 199 judged runs per period and 90% to 85% needs about 686. Why the count grows with the inverse square of the gap, and which tricks shrink it, is explained in how many eval examples you need; the arithmetic is the same for live runs.
If your agent cannot produce that many judged runs in a week, widen the window and keep the bar where it is. The SRE Workbook’s chapter on alerting on SLOs makes the point for low-traffic services: “if a system receives 10 requests per hour, then a single failed request results in an hourly error rate of 10%”. An hourly judge average on a small agent is that hourly error rate. Read the failing runs anyway; Husain and Shankar’s advice is to “track confidence intervals for production metrics. If the lower bound crosses your threshold, investigate further.”
Does the judge itself drift?
Yes: the judge is a model with a rubric, so it can drift with its model, its prompt and the people who wrote the criteria. The book’s glossary entry says “A judge is an instrument that must itself be calibrated against human judgment before it is trusted”.
Calibration decays. Shankar and colleagues named one reason in a 2024 paper on validating LLM graders: “criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria”. The rubric you wrote at launch is a guess about failures you had not yet seen.
Treat the judge like any instrument on the dashboard. Pin its model and version its rubric, record both on every score, and re-check agreement with human labels whenever either changes and on a fixed schedule in between. The judge agreement calculator computes agreement beyond chance with an interval, and the post on whether an LLM judge is reliable covers the method; the coverage line carries the date of the last check.
Is the monitor itself running?
Not always: online evaluators and alert rules can stop working without raising an error, and a quality line that stops moving looks the same as a stable one. I found this failure in the issue trackers of widely used open-source observability tools, read in October 2026; none of the five top search results I read mention it.
Three reports from two projects, as examples of the category. In July 2026, a user of one tool reported that when the judge’s endpoint rate-limits, “Assessments fail silently, the trace is dropped until the next scheduler tick”. In March 2026, another reported that any trace running longer than the scheduler’s scan interval “is systematically ignored and never evaluated”.
The third is the one I would show a CTO. In August 2026, a user of a second tool reported alert monitors that switch themselves off, adding that “from the outside a disabled monitor and a monitor with nothing to report look identical”. The same report says: “Three separate monitors across three projects have died this way for us over five days.”
These are reports at a date; some were fixed or closed later, and the shapes will recur in other tools. The defense is the coverage line in the table: scored runs divided by eligible runs, failed judgments, and when each alert last ran. The check on “last ran” has to live outside the tool it watches, the way a heartbeat does, or it can go quiet the same way.
What should the dashboard spec say?
The spec should fit on one screen and name, for each signal, how it is measured, the alert rule, who hears it and the baseline it compares against. The template below follows the table above; the bracketed values are yours to set, and the defaults in the table are illustrative.
AGENT DASHBOARD: [agent name] · owner: [name] · on-call: [rota] · reviewed: [date]
Eval baseline: [offline suite] at [version] · pass [k/n] · re-run on prod config: [weekly]
Judge: rubric [version] · judge model [pinned id] · agreement with humans [value] on [n] labels · calibrated [date]
SERVICE: uptime · latency p50/p95 · error rate (keep the existing alerts)
1 OUTCOME QUALITY checks on 100% of runs · judge on [r]% random sample of finished runs
· rephrase and abandonment on 100%
page: a programmatic check fails on > [5]% of an hour's runs (min [200] runs) -> on-call
weekly: judged pass rate below trailing 4 weeks, two-sided p < 0.05 -> owner
(judged runs this week: [n]; smallest drop visible at this n: [d] points)
2 COST PER TASK cost per run p50/p95 · cost per successful task · price table [version]
page: a run spent past its hard cap (the cap failed) -> on-call
ticket: median cost per task >= [2]x its 4-week level -> owner, same day
daily: p95 cost per run > [3]x expected cost per task -> owner
3 GATES & GUARDS approvals requested/granted · time to approve · escalations by trigger
· guardrail blocks by rule · runs ended by a budget cap
daily: one block, escalation or approval-request type > [3]x its baseline -> owner
weekly: approval rate > [98]% for [3] weeks; any rise in cap-ended runs -> owner
4 BEHAVIOR SHAPE steps p50/p95 · errors and retries per tool · finish reasons
· result size per tool, incl. share of empty results
page: one tool's error rate > [5]x baseline over [15] min (min [50] calls) -> on-call
daily: steps p95 up [30]%; finish-reason mix shifted;
a tool's result sizes jump or its empty share triples -> owner
5 DRIFT requested vs responding model · prompt version · rendered-prompt hash
· week-over-week on 1-4
ticket: responding model or prompt hash changed with no deploy -> owner, same day
weekly: scheduled offline re-run below its release baseline -> owner
COVERAGE scored / eligible runs · failed judgments · last run of each alert rule
daily: coverage < [80]% of planned rate; failed judgments not retried;
any alert rule silent > [24] h (checked from outside the tool);
judge calibration older than the last judge or rubric change -> owner
REVIEW (weekly, 30 min): read the 10 worst runs by cost, steps and judge score;
every flagged run becomes an offline eval case
Two lines in it carry the most weight. The “smallest drop visible at this n” line forces the team to say what the quality number can and cannot see. The review line keeps a person reading runs, which Chapter 15 defends with a sentence I would put on the wall: “The tile you glance at every morning should be capable of ruining your morning; if it can only ever go up, it is decoration.”
Would this dashboard catch three quiet failures?
It catches each of the three with a different rule, and the walk-through below shows which rule fires and which stay silent. Try each one before reading the answer: which row of the table would you expect to fire?
A model update lowers answer quality while latency improves. The provider rolls the alias you call forward to a new snapshot that answers faster and slightly worse, so latency improves, errors stay flat and no page fires. The drift ticket fires the same day, because the responding model changed with no deploy of yours.
The weekly quality test stays quiet if the judge scores fifty runs a week, exactly as in the worked read. So the ticket’s first action is the offline re-run against the production configuration, where the suite is large enough to measure the drop.
A tool starts returning empty results. A search backend loses its index and answers every query with an empty list and a success status. The tool’s error rate does not move, so the error-rate page stays silent. The next morning’s daily rule fires on that tool’s share of empty results, and the step-count rule may follow as the agent retries its searches.
The online judge silently stops scoring. The judge endpoint starts rate-limiting and judgments are dropped, the shape in the July 2026 report. The weekly quality test cannot fire, because a handful of scored runs never reaches significance, and the few that were scored may look fine. The coverage rule fires the next morning: scored runs fall below 80% of the planned rate and failed judgments sit unretried. Had the alert rule itself been switched off, the outside check on when each rule last ran would fire instead.
Where does this approach break?
Any answer to how to monitor AI agents in production breaks first on low-volume agents. With tens of runs a week, no statistical alert on quality is available, and the honest monitor is a person reading every failing run and every flagged one.
The arithmetic assumes independent runs and a stable judge. Runs from one user’s session are correlated and count for less than their number suggests, and a judge whose rubric changed mid-window makes the two periods incomparable. The thresholds in the table are placeholders; the right multiple for a cost alert depends on how spiky your traffic is, and you learn it from your first month of data.
The two provider postmortems show a detection gap; they do not tell you how often your agent will regress. The five-signal grouping is my synthesis of the book’s material, and other groupings work as well.
Monitoring tells you something changed. Finding out why is a different job, done by comparing a good trace with a bad one, which the post on AI agent failure modes walks through. The fix belongs in the eval suite, so the same failure fails a test next time.
The dashboard to ask for this week
The short answer to how to monitor AI agents in production is one screen. Ask your team for five signals and a coverage line, an alert rule and an owner for each, and the number of judged runs printed next to the quality figure. Then ask which of the three quiet failures above it would have caught last month.
The production-monitoring material behind this post is in Chapter 15, Observability and Debugging (in the full book), with the online and offline layers in Chapter 16, Evaluating Agents (in the full book). The guide to evaluating and observing agents collects the related posts and tools, and you can see the formats.
Questions readers ask
- How do you know if an AI agent is working in production?
- Run programmatic checks on every run where a check exists, score a random sample of finished runs with a calibrated LLM judge, and track implicit signals such as a user rephrasing the same request or abandoning the session. Compare the weekly judged pass rate, with its confidence interval, against your offline eval baseline, and read the worst runs by hand every week.
- What is the difference between offline and online evals for LLM apps?
- Offline evals run a fixed, curated suite before release and answer whether a change helped, because the dataset holds still. Online evals score a sample of live runs and answer what is happening with real users. Each failure that online evaluation finds should become a new offline case, so the suite keeps growing from real traffic.
- How many production runs should an LLM judge score?
- Enough that the smallest drop you would act on stands out from the noise. With independent runs and illustrative numbers, seeing a fall from 90% to 80% at 80% power takes about 200 judged runs in each period compared, and seeing 90% to 85% takes about 686. A team that scores fewer should review quality weekly or monthly, never hourly.
- How do you detect drift after a model update?
- Record the model you requested and the model that responded, plus the prompt version and a hash of the rendered prompt, on every run. Raise a ticket the same day when the responding model or the prompt hash changes without a deploy, and re-run the offline eval suite against the production configuration on a schedule.
- How do you avoid alert fatigue when monitoring AI agents?
- Page only on signals that are frequent, unambiguous and actionable, such as a breached spending cap or one tool's error rate jumping. Send noisy, low-volume signals such as judged quality, approval rates and drift to a daily or weekly review with a named owner, and remove any alert that nobody acts on.
Sources
- OpenAI (2025). Expanding on what we missed with sycophancy (a model provider's postmortem)
- Anthropic (2025). A postmortem of three recent issues (a model provider's postmortem)
- Lingjiao Chen, Matei Zaharia, James Zou (2023). How is ChatGPT's behavior changing over time? (preprint; v1 and v3 abstracts read)
- Arvind Narayanan, Sayash Kapoor (2023). Is GPT-4 getting worse over time?
- Rob Ewaschuk (2016). Monitoring Distributed Systems, in Site Reliability Engineering, ch. 6 (ed. Betsy Beyer and others)
- Steven Thurgood (2018). Alerting on SLOs, in The Site Reliability Workbook, ch. 5
- Hamel Husain, Shreya Shankar (2026). AI Evals: Everything You Need to Know (evals FAQ, modified September 2026)
- Shreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran, Ian Arawjo (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences (preprint; abstract read)
- LangChain (2026). Set up LLM-as-a-judge online evaluators (one example of a product's online-evaluation docs; modified 6 October 2026)
- Braintrust (2026). Score production traces (a second example of a product's online-evaluation docs; modified 23 September 2026)
- rahul-oss-sap, mlflow/mlflow issue tracker (2026). Issue #24404: online judge assessments dropped under rate limits (July 2026)
- Datashift-RobinDillen, mlflow/mlflow issue tracker (2026). Issue #21870: long-running traces never evaluated by the online scheduler (March 2026)
- pikonha, langfuse/langfuse issue tracker (2026). Issue #16610: alert monitors that switch themselves off (August 2026)
- llmskeptic, Hacker News (2026). Ask HN: How do you monitor AI features in production? (March 2026)
- navaed01, Hacker News (2025). Ask HN: How are you checking if your LLM is giving customers the right answer? (May 2025)
- jneagu, Hacker News (2024). Comment on the statistical power of A/B tests for LLM apps (October 2024)
- imviky, Hacker News (2026). Ask HN: How do you find out if the LLM API is giving degraded responses? (June 2026)