Home / Blog / Agent fundamentals / Why Do AI Agents Hallucinate? The Engine Explains It

Agent fundamentals

Why Do AI Agents Hallucinate? The Engine Explains It

Why do AI agents hallucinate? The model writes the likeliest text where a fact is missing. Sort your case into seven causes and pick the fix that fits.

By Enrique Gutiérrez · Published · 22 min read

Why do AI agents hallucinate? Because the language model inside predicts the most plausible next piece of text, and where it lacks a fact it writes a plausible one in the same confident tone. An agent then acts on that text and reads it back as context on every later step.

That answer comes from Chapter 2 of the book, which is free to read online. The rest of this post turns it into something you can run on a failing transcript: seven causes, five questions and a report template. The seven-way split is my own arrangement of things the book says in separate places. The book prints no such table and gives no rate.

Why do AI agents hallucinate, mechanically?

AI agents hallucinate because a language model does one job: given the text so far, it scores every possible next token, and the surrounding system picks one and repeats. A token is a chunk of text, often a word or a piece of one. There is no second mode for facts and no step where the model looks anything up.

Chapter 2, The Engine: How Language Models Work is blunt about what follows from that loop: “the model never drafts its answer in advance”. Each token is a local guess, conditioned on everything before it. The guessing skill sits in the model’s weights, the billions of numbers fixed during training.

Autoregressive generation, the engine’s whole method.
Figure 2.1 Autoregressive generation, the engine’s whole method. The text so far enters the model, which returns a ranked list of next-token guesses; one is chosen, appended to the text, and the model runs again on the now-longer input. The chosen token and its return path are in accent—there is no plan held in reserve, only this local guess repeated. Reuse this diagram

The book calls those weights a lossy compression of what the model read. “The model absorbed the shape of its reading and kept no archive of the text itself”, so recalling a precise figure “is a gamble”.

Hallucination is, in the book’s words, “the field’s term for output that is fluent, specific, confident, and wrong”. The glossary entry adds the part that surprises newcomers: “The mechanism is ordinary generation doing its job where a fact is missing”. Nothing broke. The engine ran exactly as built.

What is the difference between the odds and the draw?

The odds are the probabilities the model assigns to every candidate token. The draw is a separate step: a small piece of code, the decoder, picks one candidate, weighted by those odds. Most confusion about hallucination comes from merging the two.

Chapter 2 has an image for this, a weighted lottery, “run once per token: the model prints the odds, and the decoder draws a ticket.” The book uses the lottery to explain why two runs of the same prompt differ.

Using the same picture to locate hallucination is my own framing for this post. The draw is where run-to-run differences come from. The odds are where a missing fact shows. When the model lacks the fact, the odds still favor something that looks right, and one wrong, specific token can hold most of the probability.

A Hacker News commenter described the pattern in 2023 after a model recommended a function that did not exist. In TeMPOraL’s reading, the model was “guessing at some reasonable patterns - like, upon seeing function calls like InitializeFoo() or AddBar(), it would assume existence of functions UninitializeFoo() and RemoveBar()”. The invented name scored well because it had the shape of its neighbors.

Why does a missing fact come out sounding certain?

A missing fact sounds certain because fluent text is the only thing the engine produces, and it produces it whether or not the fact is there. Chapter 2 states the relation in five words: “Plausible and true are correlated.”

Then comes the clause I would ask a student to memorize: “but the two part company exactly where the model lacks the fact, and at that moment the machinery produces the most plausible-looking string instead of stopping. Fluent first, correct second.”

The danger is in the delivery. In the book’s phrasing, “the fabricated answer arrives in exactly the same assured tone as the correct one”, and “Tone carries no information about truth”. You cannot hear which sentence was a guess. The polished assistant manner comes from a later training stage, and the book is plain about that stage: “It adds no new knowledge, and it does not remove the tendency to fabricate.”

The book’s summary of why a model guesses instead of declining is that “we graded these systems into guessing”. It cites Kalai, Nachum, Vempala and Zhang (2025), whose abstract argues that “language models hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty”. That paper is a theoretical argument with bounds. It measures no model’s rate.

Why is an AI agent hallucination worse than a chatbot’s?

An AI agent hallucination costs more for two reasons: the agent acts on the invented text, and the invented text goes back into the context that every later step reads. A chatbot’s fabrication waits for a person to act on it. An agent’s fabrication becomes the next action.

An agent here means the engineering sense of what an AI agent is: a model that chooses its own next step, in a loop, within limits set by code. That code is the harness, the ordinary program around the model that runs tools, keeps the transcript and decides when to stop. The context is the text in front of the model on one call, held in its context window.

A loop gives fabrication more surfaces. A chatbot can invent a sentence. An agent can also invent a tool name, an argument inside a tool call, a summary of a result it never received, or a report about its own actions.

The second reason is the transcript. Chapter 2 puts it briefly: “Hallucination snowballing was this pattern applied to facts; it holds for actions too.” The arithmetic of that decay belongs to the post on how compounding errors in AI agents make small mistakes snowball.

Public bug reports show the action version. In a 2026 issue on a coding agent’s tracker, a user wrote that a subagent’s report “looked indistinguishable from a verified result”. The same user added: “The tool-use metadata for the run showed tool_uses: 0.” In a second report, the assistant “presented non-existent tool output as if it were a genuine tool result”.

Both are user reports, closed without a maintainer’s diagnosis. I found no study that measures how often an agent reports an action it never took. Of the seven causes below, this one has the least measurement behind it.

Why doesn’t a prompt line remove it?

A prompt line rarely removes hallucination because it asks the model to withhold the facts it lacks, and the model has no reading of which facts those are. People who ask why do AI agents hallucinate often mean which setting or sentence turns it off. The book keeps its claim at “rarely”: the engine’s failure modes “respond to engineering, rarely to sterner wording in the prompt.”

Two pieces of Chapter 2 carry the reasoning. First, “an instruction is a request, not a guarantee”. Second, the tone problem cuts both ways. You cannot hear the difference between a known fact and a guess, “and in any useful sense neither can the model.”

A prompt can do one useful thing here: offer an exit. The book’s example is a fallback line: “if the answer is not in the provided text, say you do not have enough information.” Its comment follows: “Without that last line, a model that lacks the answer will do what it was trained to do, which is answer.”

The forum record matches. A Hacker News reader asked in 2025 “is adding ‘do not hallucinate’ to prompts effective in preventing hallucinations?” Five replies came back and none carried a measurement.

In an informal 2024 test by another commenter, the instruction changed the behavior of two chat assistants: one ran a web search and the other declined. Both had an exit available. That is one person’s test, cited only as an illustration.

Two papers speak to removing it for good. Xu, Jain and Kankanhalli (2024) state in their abstract that “it is impossible to eliminate hallucination in LLMs”, for models “used as general problem solvers”. I read that paper at abstract level, so I report the claim and its stated scope and nothing more.

Kalai and colleagues answer under a heading of their own: “Hallucinations are inevitable only for base models”. A base model is the model as it leaves its first training stage, before it is tuned into an assistant. A system that may look a fact up or answer “I don’t know” is outside that claim.

My reading of the pair, in my own words: a prompt sentence cannot put a fact into the model. A system around the model can supply facts, allow declining and check the result.

The book adds its own hedge, and I would keep it. These are properties of current models: “some will soften as training methods improve, and none, at the time of writing, is solved”.

Does temperature zero stop hallucination?

Temperature zero does not stop hallucination, because temperature changes the draw and leaves the odds untouched. Temperature is a setting in the decoder. It widens or narrows the gaps between the model’s scores on their way into the draw, and it never changes which token ranks first. At zero the top-scoring token wins every time.

Chapter 2 is exact about where these settings act: “the ranked list is fixed, and the knobs govern only how you pick from it.” The consequence comes a few lines later: “correctness sits at neither end—a wrong answer can be drawn confidently at any setting.” If the top-scoring token is the invented one, temperature zero selects it every time.

One study puts numbers on this, with limits worth reading first. Spracklen and colleagues (2025 revision of a 2024 preprint) sent single code-generation prompts to 16 models and collected 576,000 Python and JavaScript samples. No agent was involved, nothing could run an install, and the models date from 2023 and 2024. The rates will have moved since.

In that setup, the average share of invented package names was “at least 5.2% for commercial models and 21.7% for open-source models”. Temperature mattered in the direction you would expect: “All models exhibited a clear increase in hallucination rate as temperature value increases”. For the commercial models the authors saw “only a slight increase in hallucination rate between temperatures 0 and 1”. Lower settings shaved the problem and left it standing.

The repeat test in the same paper says more about the mechanism: “43% of hallucinated packages were repeated in all 10 queries, while 39% did not repeat at all across the 10 queries.” A name that returns ten times out of ten reads like a property of the odds. Noise in the draw would scatter it.

A 2023 Hacker News comment claimed that temperature zero mostly cures the problem, with a telling condition attached: “especially with a good prompt including the information you are asking about.” Putting the information in the prompt changes the odds.

Which of seven causes is in your trace?

This post sorts an AI agent hallucination into seven causes, each with its own evidence in the trace and its own response. A trace is the recorded sequence of a run: every model call with the exact input it received, and every tool call with its raw result.

The table is a synthesis. Chapter 2 gives one mechanism, and the symptoms appear across other chapters, some of them in the full book. For a wider research map, a 2025 survey of hallucinations in LLM-based agents counts “eighteen triggering causes” in its abstract.

Symptom you see Cause Confirm it from the trace Response Where the book says it Where the fix lives
An invented package, function, citation, ID or URL Missing fact. The value was never in the model’s input, and the weights hold no reliable copy Search the model’s input at that step for the value or a source for it. Nothing is there, no tool fetched it, and a source could have supplied the true value. The invented one was never true: the thing does not exist, or it exists and is the wrong one here Put the fact in the context by prompt, document or tool, then check the value against that source in code Chapter 2 (free): lossy weights, plausible versus true context
A deprecated call, an old version number, a policy or price that has since changed Stale fact. The value was true before the model’s training cutoff The claim matches an earlier state of the world, and no dated source was in the model’s input Deliver time-sensitive facts in the prompt or by tool, with their date Chapter 2 (free): the training cutoff context
A confident summary of data the agent never received Hole made by the harness. An error swallowed, an empty result passed on, a result cut without notice A tool call just before the claim failed, came back empty or was cut, and the model’s input for the next step does not say so Return the failure to the model as a readable error, and say when a result was cut Chapter 2 (free): silent truncation. Chapter 18 (in the full book): the swallowed error harness
A tool name that is not on the list, or an argument from nowhere Interface gap. A free-text field, an ambiguous description, a name nobody checks Compare the called name with the tool list that was sent. Compare each argument with the user’s text and earlier results Close the choice with a fixed list of allowed values, validate in code, return the valid names as an error Chapter 2 (free): closed schemas, failures fed back. Chapter 15 (in the full book): the argument test tool contract
“Tests pass”, “file updated”, a pasted “output” or a quote from a tool result, with no record behind it Unverified action claim. Nothing compared the sentence with the record Match the claim to a tool call id, its exit status and its raw output. No such call exists, or the record says otherwise Compare the sentence with the record in code. The harness decides “done” from the exit status, and a quoted result must appear in the raw output Chapter 1 (free): confidence carries no information. Chapter 18 (in the full book) verification
Later steps add detail to a claim nobody sourced Snowball. An earlier fabrication now sits in the transcript Scroll back to the first appearance of the claim. The agent wrote it with no source, and later steps cite it Check the claim in a fresh context or with a check that involves no model, then restart from a clean summary Chapter 2 (free): snowballing. Chapter 7 (in the full book): context poisoning verification
An answer where “I don’t know” was right, or a correct answer dropped after pushback No exit. The agent had no allowed way to decline or to ask No source in reach held the fact and the instructions offered no fallback, or the claim flipped after an objection that carried no evidence State the fallback in the instructions and add an “ask” or “cannot do” action. After an objection, re-run the check that settled the answer, and change it only on new evidence Chapter 2 (free): the fallback line, and sycophancy (giving way under pushback) context, tool contract

The table reduces to four rules. If the fact was never in the model’s input, supply it. If a tool failed before the claim, repair the harness first. If the claim is about something the agent did or about what a tool returned, check the record and ignore the sentence.

If the model had a free choice of names, close the choice. Only the last row has a fix that is an instruction to the model, and even there the instruction only offers an exit. Rows 1 and 2 use the prompt to carry a fact.

Which five trace questions sort the causes?

Five questions, asked in order, assign one cause to one step, and the first question that stops you is your answer. The first three look for causes outside the model’s knowledge, where a prompt edit or a retrieval layer would be wasted effort.

  1. Where does the claim first appear? Scroll back from where you noticed it. If the agent wrote it at an earlier step and this step only builds on it, this step is a snowball. Go to the earlier step and ask questions 2 to 5 there.
  2. What did the model receive just before the claim? Find the last tool call before it. If that call failed, returned nothing or was cut, and the model’s input for the next step does not say so, stop: a hole made by the harness.
  3. Is the claim about something the agent did, or about what a tool returned? Look for the call id, the exit status and the raw output. If no matching call exists, or the record contradicts the sentence, stop: an unverified action claim.
  4. Is the invented thing a tool name or a parameter of a call? Compare the name with the tool list sent, and each parameter (an ID, a date, a path) with the user’s text and earlier results. If a value has no source, stop: an interface gap. Text the agent authors, such as code or prose, goes to question 5.
  5. Was the fact in the model’s input, and what should the step have produced? Search the input for it. If the true value sat in the prompt or the user’s text and the answer still got it wrong, you have a reading error, which is outside this table (see the limits below). If the value was absent and a prompt, document or tool could have supplied it, the cause is a missing fact when the invented value was never true. It is a stale fact when the invented value used to be true. If no source in reach held the answer, or a right answer was dropped after an objection, the cause is no exit.

Question 2 needs a trace that records what the model received, which can differ from what the tool returned. If your harness logs only the tool’s side, add the model’s side first. The agent bug bestiary catalogs neighboring specimens, such as the swallowed error and hallucinated arguments.

Worked example: “all tests pass” after a failing run

In this case a coding agent announces that all tests pass after the test tool returned an error. The excerpt below is plain pseudocode, invented for illustration.

step 14
  model   -> harness : call run_tests(path="tests/")        id=c14
  harness            : executed c14, exit_status=1
  harness -> model   : result c14: "2 failed, 31 passed.
                       FAILED test_refund_window: AssertionError"
step 15
  model   -> user    : "All tests pass. The refund fix is complete."
  harness            : final answer received, run marked done

Question 1: the claim first appears at step 15, and nothing earlier says the tests pass. No snowball. Question 2: the last tool call, c14, failed. The line harness -> model shows the model received the failure text, so the harness made no hole.

Question 3: the claim is about the result of the agent’s own action. The matching call exists, and its record, exit_status=1, contradicts the sentence. Stop here. The cause is an unverified action claim, and the fix lives in verification.

The response follows from the row. The harness already held the exit status when it marked the run done. The repair is one rule in ordinary code: a run that claims passing tests closes only if the last test call exited with status zero.

Change one line and the diagnosis changes. Had the third line read result c14: "", question 2 would have stopped you at a hole made by the harness. The repair would be to return the error text. Chapter 18, Reliability, State, and the Harness (in the full book) has a sentence for that outcome: “The failure has been laundered into an answer.”

Two more cases from my test set run through the same questions and land in different rows:

Case Question that stops it Cause Response
The agent writes an import for a package that exists in no registry 5, the value was never true Missing fact Put the project’s real dependency list in the context, then check every import against it in code before accepting the file
The agent quotes a refund policy that changed after its training data ended 5, the value used to be true Stale fact Fetch the current policy by tool or place it in the prompt, with its date

Does retrieval fix it?

Retrieval reduces hallucination and does not remove it. Retrieval means fetching relevant text and placing it in the context before the model answers, the pattern known as retrieval-augmented generation, or RAG. It is the main response to the first two rows of the table.

The book’s support for it is strong and specific: “The model is markedly more reliable when the facts it needs are in its context than when it must dredge them from its weights”. Chapter 3, The Agent Loop (in the full book) gives the reason in a clause, “a model that can look a fact up no longer needs to invent it”, and names the idea grounding.

Now the measured residue. Magesh and colleagues (2024) ran what their abstract calls a preregistered evaluation (the method was fixed before the results were in) of retrieval-based legal research tools from two large providers. The abstract reports that “hallucinations are reduced relative to general-purpose chatbots”, and that the tools “each hallucinate between 17% and 33% of the time.”

Limits on that number matter. It covers one domain, legal research, in commercial tools as they stood in 2024. I read the preprint at abstract level, so I cannot tell you exactly what the authors counted as a hallucination. It shows that retrieval left a residue in one field, and it supports no general rate.

Two things produce a residue. The fetch can miss or return the wrong passage, and in a fixed pipeline, Chapter 8, Retrieval and Knowledge (in the full book) says, “no step exists at which anything could notice”. The model can also write past what the passage says, which is my own addition and follows from the mechanism above.

Kalai and colleagues make a related point about incentives: “the binary grading system itself still rewards guessing whenever search fails to yield a confident answer.”

Does structured output stop a hallucinated tool call?

Structured output answers one cause, the interface gap, and leaves the other six where they were. A schema is a declaration of which fields an output may have and which values each may hold. Constrained decoding enforces it by striking invalid tokens from the list before the draw.

That closes a real door. In the book’s words, “if the field list is closed, no invented field can begin”. When an agent hallucinated a tool call by inventing a name, the book’s recovery is to “feed the failure back as data”. Return an error listing the valid tools, and cap the retries.

The limit is stated just as plainly in Chapter 2’s section on structured output: “a guarantee about shape says nothing about truth”. If the model lacks the order total, the schema “will make it produce a well-typed number.” The verdict follows: “A schema, misused, makes fabrication easier to consume, not rarer.”

A wrong call has more causes than an invented name, and they deserve their own diagnosis. The six causes of why agents call the wrong tool are a separate sorting; here the wrong call is one row of seven.

A harness can add a fault of its own on top. A 2026 issue on an agent framework reports that a call to a name outside the registered tools was dropped in silence: “The step executes nothing, records nothing to memory”. The model asked for something and got no reply. That is a hole made by the harness, stacked on an interface gap.

Why can’t you ask the model whether it made something up?

Asking the model whether it made something up is a weak check because the model tends to agree with the frame of the question. Chapter 2 says “Are you sure?” “is weak verification, because the model tends to ratify whatever frame the question hands it”. Chapter 1, What Is an Agent? gives the general rule: “The model’s own confidence carries no information about whether it is right”.

One court record shows the cost. In Mata v. Avianca, a 2023 sanctions order in the Southern District of New York, the court found that lawyers “submitted non-existent judicial opinions with fake quotes and citations” produced by a chat assistant. The order is one matter, dated 22 June 2023, and I cite the order itself.

The detail that matters here is the self-check. According to the order’s findings, a lawyer asked the tool “Is Varghese a real case” and “Are the other cases you provided fake”. The tool, the order records, “responded that it had supplied ‘real’ authorities”. The court imposed a penalty of $5,000.

Two cautions. The tool was a general chat assistant used on its own, so this is the model’s behavior without an agent around it. The court also declined to condemn the technology: “there is nothing inherently improper about using a reliable artificial intelligence tool for assistance.”

A check earns the name when it comes from outside the thing being checked. The book’s word for such a source is an oracle: a test, a registry lookup, a file you can open, an exit status. The post on compounding errors covers one useful exception, a model asked in a fresh context, with the 2023 study behind it.

What should a hallucination report contain?

A hallucination report should contain four things from the trace: the symptom, the evidence line, the cause and the response you chose. Writing them down forces the step most people skip, which is copying the line that proves the cause.

The template below is plain text. Fill one per incident and keep them.

HALLUCINATION REPORT

Run or trace id:
Step where the claim first appears:
Symptom (the claim, copied word for word):
What is true instead, and how I know (file, command, source, date):

Evidence line (copied from the trace, never paraphrased):

Trace questions, in order. The first one that stops you is the cause.
1. Written by the agent at an earlier step?            yes / no   step:
2. Tool call before it failed, empty or cut,
   and the model's input does not say so?              yes / no
3. Claim about the agent's own action or a tool's
   result, with no matching record or a record that
   says otherwise?                                     yes / no
4. Tool name or call parameter that no list or
   source supplied?                                    yes / no
5. True value in the prompt or the user's text?
   yes -> reading error, outside the table
   no  -> never true / used to be true / no source in reach
   right answer dropped after an objection -> no exit

Cause (one row of the table):
Where the fix lives: context / harness / tool contract / verification
Response chosen:
What this response does not cover:
Check that shows it worked (same input, N runs, count of failures):

The last two lines are the ones I would refuse to leave blank. Every response in the table covers one cause, and a fix you never re-ran is still a guess.

Where does this diagnosis stop working?

This diagnosis stops working when a run mixes causes across steps, when the trace is incomplete, or when the error falls outside the seven rows. A full answer to why do AI agents hallucinate is wider than any seven-row table.

First, the questions classify one step. A long run can hold a stale fact at step 3 and a snowball built on it at step 12. Classify each step separately and fix the earliest first.

Second, the table leaves one kind of error out. A fact that sat in the prompt or the user’s own text, with no tool call behind it, and still came out wrong is a reading error, which the full book treats in its chapter on managing context. A misquoted tool result is different: question 3 catches it, because the record holds the raw output. Wrong arithmetic is a separate entry in Chapter 2’s list, with its own remedy: route exact work to a calculator or to executed code.

Third, the numbers in this post are dated and narrow. The package study measured single prompts to models from 2023 and 2024, and the retrieval study covered one domain. The court order is one matter. None of them gives a general hallucination rate.

None of this is exotic engineering. Most of what is new about AI agents for backend developers is this habit of treating a model’s sentence as unverified input. The wider set of basics sits in the agent fundamentals guide.

The check to keep

The answer to why do AI agents hallucinate fits in one line: the engine writes plausible text where a fact is missing, and the loop acts on it. The working question for any run is where the fact was supposed to come from: the context, a tool result, or the record of an action. Find the step, copy the evidence line, name one cause, change one layer.

The mechanism is in Chapter 2, “The Engine”, free to read online along with the glossary. The trace-reading craft and the error-handling patterns are in Chapters 15 and 18, in the full book; see the formats.

Questions readers ask

Why do AI agents hallucinate?
The language model inside an agent predicts a plausible next token whether or not it holds the fact. Where the fact is missing, the most plausible text is an invented one, and it arrives in the same confident tone as a true one. The agent then acts on that text and reads it back as context on every later step.
Does setting temperature to zero stop hallucination?
No. Temperature changes how a token is drawn from the model's odds and leaves the odds themselves alone. Chapter 2 of the book says a wrong answer can be drawn confidently at any setting. In one study of code-generation prompts, invented package names rose with temperature and were still produced at low settings.
Does RAG stop AI agent hallucination?
It reduces it and does not remove it. A model is more reliable when the fact is in its context, but a fetch can miss, and the model can still write past what the passage says. A preregistered 2024 study of retrieval-based legal research tools, in one domain, found that each hallucinated between 17% and 33% of the time.
Why does my agent say it ran something it did not run?
Nothing in the loop compared the sentence with the record. A report such as 'tests pass' is generated text like any other. Match each claimed action to a tool call id, an exit status and the raw output, and let the harness decide that a task is done from that record.
Can hallucination be eliminated?
From a model that must always answer, apparently no: one formal paper argues elimination is impossible for a model used as a general problem solver. A system is a different matter. Kalai and colleagues (2025) write that hallucinations are inevitable only for base models, since a system may look a fact up or answer 'I don't know'.

Sources

  1. Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang (2025). Why Language Models Hallucinate
  2. Ziwei Xu, Sanjay Jain, Mohan Kankanhalli (2024). Hallucination is Inevitable: An Innate Limitation of Large Language Models
  3. Joseph Spracklen, Raveen Wijewickrama, A H M Nazmus Sakib, Anindya Maiti, Bimal Viswanath, Murtuza Jadliwala (2025). We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
  4. Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, Daniel E. Ho (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (preprint)
  5. Muru Zhang, Ofir Press, William Merrill, Alisa Liu, Noah A. Smith (2023). How Language Model Hallucinations Can Snowball
  6. Xixun Lin and colleagues (2025). LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
  7. U.S. District Court, Southern District of New York (Castel, J.) (2023). Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), Opinion and Order on Sanctions, Doc. 54 (S.D.N.Y., 22 June 2023)
  8. Isamu (Hacker News) (2025). Hacker News comment asking whether 'do not hallucinate' works in a prompt
  9. tkgally (Hacker News) (2024). Hacker News comment reporting a hands-on test of a no-hallucination instruction
  10. ilaksh (Hacker News) (2023). Hacker News comment claiming temperature zero mostly prevents hallucination
  11. TeMPOraL (Hacker News) (2023). Hacker News comment on why a model invents a function name
  12. kunal-heart (GitHub) (2026). GitHub issue: a subagent returned a full report with zero tool uses
  13. likebear1968, with a comment by Necmttn (GitHub) (2026). GitHub issue: non-existent tool output presented as a genuine tool result
  14. BlueX888 (GitHub) (2026). GitHub issue: a streamed call to an unregistered tool name is silently dropped