Home / Blog / Context engineering and memory / Why Your AI Agent Ignores Instructions (Is It t…

Context engineering and memory

Why Your AI Agent Ignores Instructions (Is It the Context?)

When an AI agent ignores instructions, one of four causes is at work. See what 14 studies measured, then run four checks on the failing request.

By Enrique Gutiérrez · Published · 25 min read

When an AI agent ignores instructions, one of four things happened. The rule never reached the model, other text contradicted it, it was present and still broken deep in a long transcript, or prose could never guarantee it. Published studies measure parts of each, mostly on single prompts. Four checks on one saved request tell you which.

The title asks whether it’s the context, and the honest answer is “sometimes”. Context is one of the four causes. The popular story, in which the rule lost an attention contest to newer, longer tool output, is a hypothesis where agents are concerned.

Of the fourteen studies tabled below, seven tested a single prompt and only two ran an agent. One of those two reports that token distance “is not a significant factor” in the drift it measured.

No study I found places a standing rule at different positions in an agent transcript and measures compliance. None tests capital letters, and none measures how often a summary pass drops a rule. So this post sets out the evidence with its limits, then a diagnostic that turns the hypothesis into a measurement on your own transcript.

What do engineers report when an AI agent ignores instructions?

They report four experiences under one complaint. A rule held and then drifted, a rule broke on the first step, louder wording changed nothing, or a rule never reached the model. I collected 21 reports on October 6, 2026, from 13 Hacker News comments and 8 GitHub issues.

One commenter (buschleague, February 2026) wrote that “agents follow them for the first few steps, then gradually drift.” Another (chickensong, December 2025) saw rules followed “somewhat reliably at the beginning and end of the conversation” and ignored in the middle, then added: “but I have no real proof.”

A third commenter (pessimizer, October 2025) ended up typing a rule in capitals: “THE LAST RULE IS THE MOST IMPORTANT RULE”. The commenter’s verdict on magic words: the hunt “is just an illusion of control.”

Some rules never arrive. A GitHub issue from September 2026 (franjorub) describes project rules that one client loaded and another skipped: “The rules read as binding to anyone reviewing the repo, but never reach the model.” Two reporters independently invented a probe for this, an instruction with a visible, harmless effect. One calls it a “behavioral canary” (lyonsno, March 2026).

Every one of these people would say their AI agent ignores instructions, and they need different things. With three reports reused from a neighboring post’s research, 9 of the 24 ask whether the rule reached the model at all. Six describe drift, 6 describe something outranking the rule, 5 ask about louder wording, and 4 need a rule that cannot fail; some count twice. The sample is skewed: all but two of the 21 come from coding-agent harnesses, and nobody in the set cites a study.

What are the four causes?

The four causes are: the rule was not in the request, the rule was contradicted, the rule was present and lost, and the rule was never enforceable in prose. Each names a different thing to change, which is why sorting comes before fixing.

A language model keeps nothing between calls, so “the agent forgot” is always a statement about one request. The system prompt and every other standing instruction are resent by the surrounding software on each call, with the whole transcript so far. The figure, from Chapter 2 (free to read), shows the stack that is rebuilt every time.

The anatomy of a chat prompt.
Figure 2.4 The anatomy of a chat prompt. Standing instructions go in the system message, the current request in the user message, and the model’s own answers accumulate as assistant messages. Because the model retains nothing between calls, the surrounding software resends the entire stack on every turn (the accent loop)—what feels like memory is only the transcript being re-read. Reuse this diagram
Cause What happened How well it is evidenced What you change
Not in the request The rule was never loaded, or a trim, a cap or a summary removed or weakened it An engineering fact. How often a summary drops a rule is unmeasured The code that assembles the request
Contradicted The rule is present, and text of equal or higher standing says the opposite Measured on single prompts with formatting rules (Geng, 2026; Zhang, 2025) The conflicting text
Present and lost The rule is present, nothing contradicts it, and a long transcript still ends in a violation Length, position and crowding are measured on single prompts and single tool calls. For agent transcripts, a hypothesis Placement, the pile, or the run’s earlier steps
Never enforceable in prose The rule fails at some rate even on a short, clean request Measured as a rate on agent-style benchmarks (Yao, 2024; Qi, 2025) A check in code

The third row is where the title’s question lives. If your question is what is context rot in general, read the post on what context rot is, with its symptoms and four fixes.

What has been measured, and on what?

Fourteen studies bear on the question, and the setting of each limits what it can say. Seven tested a single prompt, three a scripted chat, two a tool call in one step, and two an agent run. Only one ran long, tool-using trajectories.

I name no models; each result belongs to the models of its year. The last column is the setting, and picking “agent run” shows how little is left.

Study What was varied What was measured Result, in the paper’s numbers What it does not show Setting
Liu et al., 2023, “Lost in the Middle” (TACL) Position of the answer passage among 10 to 30 passages QA accuracy, 4 models A U-shaped curve; one model “can drop by more than 20%” and, mid-input, falls below its 56.1% no-document score Instructions or agents. Restating the query “minimally changes trends” on this task single prompt
Hong, Troynikov and Huber, 2025, “Context Rot” (Chroma technical report) Input length, distractors, needle position Retrieval and QA accuracy, 18 models “model performance consistently degrades with increasing input length”; “no notable variation” across 11 needle positions Agents or standing rules. Vendor research single prompt
Levy, Jacoby and Goldberg, 2024 (ACL) Irrelevant padding around one question, about 250 to 3,000 tokens Reasoning accuracy Average accuracy falls “from 0.92 to 0.68” at 3,000 tokens Instructions; long inputs; later models single prompt
Du et al., 2025 (Findings of EMNLP) Input length, with retrieval verified perfect Math, QA and coding accuracy, 5 models Performance “still degrades substantially (13.9%–85%)”, also “when all relevant evidence is placed immediately before the question” Agents; that placement cures length single prompt
Jaroslawicz et al., 2025 (preprint) 10 to 500 keyword-inclusion rules in one report-writing prompt Share of rules satisfied, 20 models Best models: “68% accuracy at the max density of 500 instructions”; top two near-perfect “through 150 or more instructions”; a bias toward earlier rules Harder rules or other tasks; results “may not generalize to other task types” single prompt
Mu et al., 2025 (preprint) 1 to 20 if-then guardrails added to one real system prompt Pass rate, 5 models, n=100; a pass needs every guardrail answered correctly in one chat Performance “uniformly approaches zero as the number of guardrails increases”; I recorded no rate per count Tools, conflicts or long history: the test has none. How often a single rule fails: the score is all-or-nothing multi-turn chat
Zhang et al., 2025, IHEval (NAACL) Aligned or conflicting instructions across system message, user message, history and tool output; 3,538 examples Task score, 13 models The best model shows “a 22-point drop from its aligned setting”; the best open model resolves 48% of conflicts Multi-step runs; each example is one or two turns single prompt
Geng et al., 2026, “Control Illusion” (AAAI; preprint 2025) Two mutually exclusive formatting rules, one marked as the priority; 1,200 data points Responses satisfying only the priority rule, 6 models 74.8% to 90.8% follow one rule alone; 9.6% to 45.8% under conflict; the conflict is acknowledged in 0.1% to 20.3% Past “single-turn interactions and simple, verifiable constraints”; agents are outside it single prompt
Li et al., 2024 (COLM) Rounds of chat between two copies of a model, one holding a system prompt Whether the system prompt still holds each round “a significant instruction drift within eight rounds of conversations”; I recorded no headline rate Tools; models after 2023 multi-turn chat
Laban et al., 2025 (preprint; I read the abstract only) One message, against the same information over several turns Task performance “an average drop of 39% across six generation tasks” Standing rules; tools multi-turn chat
Kate et al., 2025, LongFuncEval (preprint) Tool catalog size; tool response length; tool conversation length Tool-call accuracy; answer retrieval Drops of “7% to 85%”, “7% to 91%”, and 13% to 40%, in that order Obedience to a standing rule tool calling
Qi et al., 2025, AgentIF (preprint) 707 instructions from 50 agentic applications, averaging 11.9 constraints Constraints met; instructions with every constraint met Best scores 59.8 and 27.2; past 6,000 words of instruction the second is “nearly 0” for all models A loop; accumulating tool output tool calling
Yao et al., 2024, τ-bench Repeated trials under a written policy, with tools and a simulated user Task success; pass8, the share of tasks passed on all eight tries The best agent of 2024 passes 61.2% of retail tasks and 35.2% of airline tasks on one try; pass8 is under 25% in retail Why a policy line was missed agent run
Arike et al., 2025 (technical report) Competing-phase length, pressure, emphasis, filler for history; 4 models, one simulated trading environment Drift from the goal in the system prompt “all evaluated models exhibit some degree of goal drift”; “token distance is not a significant factor in the goal drift patterns we observe” Rule lists; a second environment agent run

My reading is thinner for two rows: I read only the abstract of Laban, and for Mu I have the sentence without the pass rate at each count. Six rows are preprints or technical reports (Jaroslawicz, Mu, Laban, Kate, Qi, Arike), and one is vendor research (Hong).

What do the studies agree on?

They agree on four things, each inside its own setting. More input is used less well, more rules are followed less completely, conflicts are resolved unreliably, and a written policy is followed at a rate below certainty.

Length. Irrelevant padding had cut accuracy by the 3,000-token mark (Levy), perfect retrieval did not prevent the drop (Du), and long tool responses hurt answer retrieval (Kate). Each is a single call, and none measured obedience to a rule.

Rule count. Two models stayed near-perfect through 150 or more keyword rules (Jaroslawicz). Performance approached zero as if-then guardrails rose from 1 to 20, on an all-or-nothing score (Mu). The count a model can hold depends on the kind of rule, so I quote none.

Conflict. When two formatting rules clashed and one was marked as the priority, that rule alone was satisfied in 9.6% to 45.8% of responses (Geng). An explicit “You must always follow this constraint” left obedience “far from reliable priority control”.

Rate. Agents given tools and a written policy passed the same retail task on all eight tries in under a quarter of cases (Yao). It is the only one of the four measured on an agent run.

Where do the studies disagree?

They disagree on whether distance or content drives drift, and on whether the start or the end of an input is the stronger position.

Distance or content? Li and colleagues found drift within eight rounds of chat. In a small open model they also saw attention to the system prompt’s tokens drop between turns. That supports the distance story, without tools, on models of 2023.

Arike and colleagues ran the only long, tool-using trajectories in the table. Filler of the same length in place of the history produced far less drift. What tracked drift was a context full of examples of the other behavior, a result the authors call “only suggestive evidence”. Both findings can hold, since the settings differ.

Start or end? Inside a list of rules, earlier ones were followed more (Jaroslawicz). Inside a long input, a fact at either end was used better than one in the middle (Liu). Inside an 80,000-token tool response, later records were answered better, by 5% to 75% depending on the model (Kate). And on one 2025 task, position made “no notable variation” (Hong).

Order inside a rule list and distance from the end of an input are different quantities. My reconciliation is practice: order the rule block by importance, restate one rule at the tail, then measure.

What has nobody measured?

Three things the popular advice depends on have no published measurement that I could find. They are the position of a standing rule in an agent transcript, the effect of capital letters, and rules lost to compaction.

Position in a transcript. No study varies where a standing rule sits in a long, tool-heavy run and measures compliance with it. The placement rules further down are an inference from neighboring measurements and the book’s reasoning.

Capitals. No controlled study tests capital letters, “IMPORTANT” or a repeated “never”. The next section has the two nearest measurements.

Compaction. Readers report rules that vanished after a summary pass, and nobody has counted how often compaction does it. Keeping a rule alive across summaries and restarts is a question of agent memory architecture: which tier holds the rule, and what reloads it into every request.

My search was one person’s search on one day. Treat “none found” as exactly that.

Does writing the rule louder help?

Sometimes it moves the rate, and the rule remains a request. With no controlled study of capitals, the nearest evidence is two neighbors that point in different directions.

In Geng’s single-prompt study, an explicit priority sentence did not make the prioritized rule win reliably. In Arike’s agent study, an emphatic statement of the goal “significantly reduces goal drift” under pressure and “proves insufficient in goal switching scenarios”.

Vendor notes are dated examples of a category. One provider’s prompting documentation (Anthropic, undated, read October 6, 2026) tells users of its newer models to “dial back any aggressive language”. A second provider’s guide (OpenAI, April 14, 2025, one model family) says that between conflicting instructions the model “tends to follow the one closer to the end of the prompt”.

Chapter 7 (in the full book) tells of a chat product whose builders had to instruct it “nine times, some in capital letters” before it complied. The book files that under fighting the weights, a mode it admits comes “without their weight of published measurement yet”, and advises: “stop pushing and route around”. Treat emphasis like any other edit, and measure the violation rate before and after.

How do you find which cause you have?

When an AI agent ignores instructions, find the first step that broke the rule and save the exact request the model received there. Then run four checks in order: present, contradicted, lost, enforceable. Each verdict names one artifact to change.

Knowing how to debug an AI agent starts with the same move for every bug. Chapter 15, Observability and Debugging (in the full book) says to read the trace forward to the first wrong step, because “that first step is the bug, and everything after it is consequence”. Then put the chapter’s question to that step: “what exactly did the model see at the moment it made this call?”

Save the request as sent, with the model identifier and sampling settings. Every replay starts from it, because starting from the user’s query would rerun your assembly code, which is a suspect. The book: “No user in the loop, no context-construction code in the way; that code is part of what may have produced the bug, so you deliberately bypass it.”

The inspection step and the clean desk are the book’s, from Chapters 15 and 7. The split into contradicted and lost, the edits in check 3, the replay count and the thresholds are mine, and every number is illustrative.

What does each check ask?

Each check asks one question of the saved request, and the first verdict ends the triage.

Check 1: was the rule present? Search by meaning as well as by string, since a summary may have reworded it. A paraphrase that weakens the rule counts as cut: the fix is a summary that keeps the rule word for word. Chapter 2’s advice is to “check for overflow before you blame the model”. The post that explains what context engineering is sorts absent text into never sent and cut off, with probes for when you cannot print the request.

Check 2: was it contradicted by something of equal or higher standing? Standing means text the model is meant to obey. That covers another standing instruction, a tool description, a line the harness wraps around your rules, and a later user turn. Chapter 7 calls the result clash, rival texts “all sharing one desk with nothing to mark which supersedes which”. Remove or reconcile one side before adding any emphasis.

Instruction-shaped text inside a tool result has lower standing and goes to check 3. It is also a security matter, prompt injection.

Check 3: was it present, uncontradicted and still broken? Replay the saved request and record the violation rate, then change one thing per arm, starting with a clean desk. Restating the rule last tests position (Liu, Kate). Clearing acted-on tool results tests length (Levy, Du). Removing the run’s earlier violations tests precedent (Arike, Laban).

Check 4: was it ever enforceable in prose? A rule that still fails on a clean desk points away from the pile. Cut the standing text to that one rule and replay. If the rate falls, the other rules were crowding it (Qi; Mu, where a pass needs every rule to hold). If it stays, ask what one violation costs.

IGNORED-INSTRUCTION TRIAGE (one rule, one failing run)

Rule, as one sentence:   ____
Run or trace id:         ____
First violating step:    ____  (read forward; stop at the first one)
Saved request:           ____  (system text, messages, tool definitions,
                               tool results, model id, sampling settings)

CHECK 1  PRESENT?
  [ ] Absent; no code tried to include it        -> NOT IN THE REQUEST (never loaded)
  [ ] Absent; a trim, cap or summary removed it  -> NOT IN THE REQUEST (cut)
  [ ] Present only as a weaker paraphrase
      (never became prefer; scope dropped)       -> NOT IN THE REQUEST (cut)
      Fix: the summary keeps the rule verbatim.
  [ ] Present in its own words, at position ____ -> Check 2
  Fix for all three: assembly code. The step-1 request shows which.
  After the fix, replay; if the violation persists, go to Check 2.

CHECK 2  CONTRADICTED?
  Text of equal or higher standing that says the opposite: a standing
  instruction, a tool description, a harness wrapper, a later user turn.
  [ ] Found: ____  -> CONTRADICTED. Remove or reconcile one side, replay,
                      record the rate. Add no emphasis first.
  [ ] None         -> Check 3
  Lower-standing text that pulls the other way (test it in Check 3):
      instructions inside a tool result:       ____
      earlier steps that broke the same rule:  ____

CHECK 3  PRESENT, UNCONTRADICTED AND STILL BROKEN?
  Replay the saved request N times at production sampling settings.
  N = 10 is illustrative.                      Baseline:  ____ / N
  Change ONE thing per arm:
  a. clean desk (standing text, tools, this step's input)  ____ / N
  b. full request, rule restated as the last line          ____ / N
  c. full request, acted-on tool results cleared           ____ / N
  d. full request, earlier violations and stray
     instructions removed                                  ____ / N
  Illustrative rule for N = 10: a gap under 5 runs is nothing, 6 or
  more is a result, exactly 5 gets 20 more runs per arm.
  [ ] a falls, and b, c or d falls   -> LOST
      b = placement, c = length, d = precedent. Fix the one that moved.
  [ ] a falls; none of b, c, d does  -> change two at a time, else unresolved
  [ ] a stays near the baseline      -> Check 4 (not shown is weaker than
                                        cleared; rerun at 30 per arm if
                                        the rule matters)
  [ ] baseline under 5 / N           -> too rare to compare arms at this N;
                                        run a and e, then the cost question

CHECK 4  ENFORCEABLE IN PROSE?
  e. clean desk, standing text cut to this one rule        ____ / N
  [ ] e falls  -> the rule list is crowding it; shorten the list
  [ ] e stays  -> prose does not hold this rule on this model
  What does ONE violation cost? ____
  [ ] Cheap   -> accept the measured rate, or change the wording or the
                 format and measure again
  [ ] Costly  -> BELONGS IN CODE: a check in the tool, a gate on the
                 action, a validator on the output. Keep the sentence
                 in the prompt as a hint.

Ask the cost question after every verdict.

What can a replay tell you, and what can it not?

A replay tells you how often this exact request produces the violation, and whether one edit moves that rate by a wide margin. It cannot tell you why the model did it, and ten runs cannot resolve small differences.

Three violations in ten against two in ten is no difference: Fisher’s exact test gives p = 1.0. Seven against two gives p = 0.07, and eight against two gives p = 0.02.

Ten runs also miss real effects often. By my arithmetic, an edit that truly cuts the violation rate from 70% to 30% shows a gap of six or more in about 24% of trials at ten replays per arm. So an arm that did not move is “not shown”, and the pile is not cleared. When the gap is small and the rule matters, rerun at 30 per arm.

A clean streak proves less than it feels like. A rule broken 26% of the time still passes ten replays in a row about once in twenty tries, because 0.74 multiplied by itself ten times is 0.049. The eval sample size calculator does this arithmetic for other counts.

Chapter 15’s single replay “at temperature zero” asks whether the failure reproduces at all, and Chapter 7 warns that “a single pass proves little”. Repeated replays at production settings ask how often.

On the why, the book is plain: replay shows you “what happened, never why the model said it”. One commenter (reedlaw, May 2026) got a fluent account from the model and wrote: “I don’t trust [the model’s] own explanation”.

What does one failing transcript look like under the four checks?

Here is a constructed run that ends at check 2. An operations agent has a standing rule to ask before it writes to the staging database, and at step 9 it runs a migration unasked. The run and every count in this section are illustrative.

Request for step 9, as sent (excerpt; constructed for this post)

[system, line 14]    Ask the user before running any command that
                     writes to the staging database.
[tool run_sql,       Runs SQL against staging. Safe to use directly;
 description]        no confirmation needed for staging.
[user, turn 1]       Add a nullable column "region" to "orders" on staging.
[steps 2-8]          schema reads and one dry run
                     (6,200 tokens of tool results)
[assistant, step 9]  run_sql("ALTER TABLE orders ADD COLUMN region text")
                     <- first violation; nobody was asked

Step 0. Steps 2 to 8 only read, so step 9 is the first violation. Its request is saved.

Check 1. Line 14 of the system text holds the rule word for word, so the rule is present.

Check 2. The description of the SQL tool, written by another team, says “no confirmation needed for staging”. A tool description is sent on every call and carries the same standing as the system text. The verdict is contradicted, and the checks stop here.

The tempting move was to rewrite line 14 in capitals. Geng’s numbers argue against it: a rule marked as the priority won alone in 9.6% to 45.8% of responses. The repair is to delete the sentence from the tool description.

In this illustration the saved request violates the rule in 7 of 10 replays, and in 0 of 10 once the sentence is gone. The cost question still applies: an unasked write to shared data is expensive, so the tool now also demands a confirmation for statements that write.

What does a run that ends in code look like?

It looks like a low violation rate that no edit moves, on a rule that one miss makes expensive. An alert-triage agent has a paging tool and the rule “Never page the on-call engineer for a severity-3 alert”. At step 2 of a short session it pages for one.

Check 1 finds the rule present. Check 2 finds nothing of any standing that says otherwise. Check 3 gives 2 violations in 10 at baseline, too few for any edit to open a gap of five, and 1 in 10 on a clean desk. Check 4 cuts the standing text to the single rule and gets 1 in 10 again.

Ten replays cannot separate one in ten from two in ten, and here they do not need to. One violation wakes a person at night for no reason, so the verdict is “belongs in code”. The paging tool now rejects alerts below severity 2 unless a human has set an override. The sentence stays in the prompt as a hint.

Where should a standing rule sit in a long request?

Put it in the slot that is resent on every call, restate the one load-bearing rule last when long material follows, and keep the list around it short. These rules rest on single-prompt measurements and the book’s reasoning. For agent transcripts they are hypotheses that check 3 tests.

Chapter 2 gives the order rule: when a prompt opens with pages of material, “put the instruction after the pile, and restate the one load-bearing rule near the end”. It adds: “The last thing the model reads should be the thing you most need it to do.”

Rule Rests on Strength of the evidence Limit
Keep a standing rule in the slot resent on every call; confirm by printing a late request Chapter 2; the reports of rules that never arrived The book’s reasoning; an engineering fact A summary pass that rewrites instructions needs the rule put back
When long material follows, restate the one load-bearing rule as the last line Chapter 2; Liu, 2023; Kate, 2025 Measured for facts and questions in single prompts and tool output. Untested for a standing rule in an agent run Liu, 2023: repeating the query before and after the documents “minimally changes trends” on that QA task. Du, 2025: length still hurt evidence placed last. Hong, 2025: no position effect on one task
Keep the always-loaded rule list short, most important first Mu, 2025; Qi, 2025; Jaroslawicz, 2025 Measured on single prompts, single responses and one scripted chat No count transfers across kinds of rule or models. Mu’s score is all-or-nothing across the rules
Clear acted-on tool results before you reword anything Levy, 2024; Du, 2025; Kate, 2025; Chapter 7 Measured on single prompts and single tool calls Trim too much and a later step lacks a fact
Delete the text that contradicts the rule Geng, 2026; Zhang, 2025; Chapter 7 Measured on single-turn formatting conflicts Some text is not yours to delete; then use a check in code
After the first violation, repair and restart from notes Arike, 2025; Laban, 2025; Chapter 7 One agent study its authors call suggestive; one chat study read in abstract A restart costs the history; write the plan to a file first
If you re-inject a rule on a schedule, measure what it buys Li, 2024 Measured on chat models of 2023, without tools. Practice from outside the book Repetition “consumes a substantial portion of the context window”; one report below saw no effect
Enforce in code any rule that one miss makes expensive Chapter 17; the glossary; Yao, 2024 The book’s reasoning; a rate on one agent benchmark Rules of taste have no cheap checker

One commenter (klardotsh, September 2026) supplies the counter-report on re-injection: with style reminders injected on a timer, the agent “still largely ignores the request”.

Both vendor guides cited earlier cover one long document plus a question in a single call, and they disagree. The undated one puts long documents above the instructions; the April 2025 one says a single copy of the instructions “above the provided context works better than below”.

For the tickable form of these lines, see the context engineering checklist for long agent runs. Procedures that only some tasks need are the subject of when a tool should become a skill.

How small does the rule get as the window fills?

The rule keeps its size while its share of the request shrinks. In an illustrative run, 1,500 tokens of standing instructions are 21.4% of a 7,000-token first request. About 23 steps later, at 2,400 new tokens per step, they are 2.4% of 63,000, and 42,000 tokens of tool results are 66.7%.

The context window budget planner does this count for your own run. Share of tokens is a different quantity from share of attention, and no study in the table derives the second from the first. With an AI agent context window full of old tool output, the count tells you what to cut first when you manage an agent’s context window.

When does the rule have to move into code?

Move it when one violation costs more than you can accept, whichever check caught the failure. A prompt rule is followed at a rate. Better placement may raise that rate, which is what check 3 measures, and it cannot make it a guarantee.

The book’s glossary (free to read) defines the system prompt as “a request rather than an enforcement mechanism”. A guardrail, it says, is “A check that lives outside the model’s reasoning and runs deterministically, whatever the model has decided”. Chapter 2 gives the picture: a model can honor a format rule “a thousand times and then, on run one thousand and one, wrap the word in a courteous sentence”.

One vendor says as much of its own product: a coding agent’s documentation (read October 6, 2026) describes its rules file as “context, not enforced configuration”. Practitioners arrive at the same place. One commenter (bayganyo, August 2026) added a check that blocks any reply over 150 words: “It’s then forced to redo its output to comply, and it’s like night and day.”

Chapter 17, Security, Safety, and Guardrails (in the full book) has the sentence I would pin above the desk: “A refund cap enforced in the payment tool’s own code cannot be exceeded by any sequence of tokens whatsoever.” The check is a test inside the tool, a gate on the action, or a validator on the output. For rules of taste, such as tone, you accept a measured rate and watch it.

What can this post not tell you?

It cannot tell you that your AI agent ignores instructions because of its context, and it cannot give you a placement rule proven on agents. It gives you a way to find out on one transcript, with small samples.

The studies used models from 2023 to mid-2025, and thresholds move with each generation.

A system that makes one model call per request needs only checks 1, 2 and 4. A rule the model resists alone on a clean desk may be fighting the weights, which the book itself holds loosely. A broken tool result is an ordinary bug, and the agent bug bestiary sorts those.

One request, four checks

The next time an AI agent ignores instructions, leave the wording alone and open the request of the first step that broke the rule. Is the rule there, and does anything of equal standing say otherwise? Does one edit move the rate, and what does one miss cost?

The order rule is in Chapter 2, free to read, as is the glossary. The clean desk is in Chapter 7, “Managing the Context Window”, and the replay method in Chapter 15, “Observability and Debugging”; both are in the full book. For the rest of this cluster, start at the context engineering guide, or see the formats.

Questions readers ask

Why does my AI agent ignore instructions?
For one of four reasons. The instruction was not in the request the model received; it was there and other text of equal or higher standing contradicted it; it was there, uncontradicted, and still broken in a long transcript; or it is a rule that prose cannot guarantee. Searching the saved request of the first violating step and replaying it a few times tells them apart.
Why does an agent follow a rule at first and then drift?
Each step adds text between the rule and the next output, and may add examples of the opposite behavior. A 2024 study of scripted chats (Li and colleagues) found significant drift within eight rounds on chat models of 2023. A 2025 agent study (Arike and colleagues) found token distance was not a significant factor and, as suggestive evidence, that examples in the context mattered. Each result belongs to the models and the setting of its year.
Does writing an instruction in capital letters make an agent follow it?
No controlled study found for this post tests capital letters. In one study an explicit priority sentence left conflicting formatting rules unreliable; in another an emphatic goal statement reduced drift under pressure and was insufficient once the context held a long stretch of the other behavior. Measure the violation rate before and after the edit on the saved request.
How many instructions can a model follow at once?
There is no general number. A 2025 benchmark of simple keyword-inclusion rules saw its two best models stay near-perfect through 150 or more rules. Another 2025 study saw performance approach zero as if-then guardrails rose from 1 to 20, where a pass meant every guardrail in the chat was answered correctly. The count depends on the kind of rule and on the model.
Where should I put an important instruction in a long prompt?
In the slot that is resent on every call, with the one load-bearing rule restated as the last line when long material follows it. That is the book's rule from Chapter 2, and it is consistent with single-prompt studies of facts and questions. No published study tests it for a standing rule in an agent transcript, so replay your own failing request with and without the restatement.

Sources

  1. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang (2023). Lost in the Middle: How Language Models Use Long Contexts (TACL)
  2. Kelly Hong, Anton Troynikov, Jeff Huber (Chroma) (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance (technical report)
  3. Mosh Levy, Alon Jacoby, Yoav Goldberg (2024). Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models (ACL 2024)
  4. Yufeng Du, Minyang Tian, Srikanth Ronanki, Subendhu Rongali, Sravan Bodapati, Aram Galstyan, Azton Wells, Roy Schwartz, Eliu A. Huerta, Hao Peng (2025). Context Length Alone Hurts LLM Performance Despite Perfect Retrieval (Findings of EMNLP 2025)
  5. Daniel Jaroslawicz, Brendan Whiting, Parth Shah, Karime Maamari (2025). How Many Instructions Can LLMs Follow at Once? (preprint)
  6. Norman Mu, Jonathan Lu, Michael Lavery, David Wagner (2025). A Closer Look at System Prompt Robustness (preprint)
  7. Kenneth Li, Tianle Liu, Naomi Bashkansky, David Bau, Fernanda Viégas, Hanspeter Pfister, Martin Wattenberg (2024). Measuring and Controlling Instruction (In)Stability in Language Model Dialogs (COLM 2024)
  8. Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer Neville (2025). LLMs Get Lost In Multi-Turn Conversation (preprint; abstract read)
  9. Zhihan Zhang, Shiyang Li, Zixuan Zhang, Xin Liu, Haoming Jiang, Xianfeng Tang, Yifan Gao, Zheng Li, Haodong Wang, Zhaoxuan Tan, Yichuan Li, Qingyu Yin, Bing Yin, Meng Jiang (2025). IHEval: Evaluating Language Models on Following the Instruction Hierarchy (NAACL 2025)
  10. Yilin Geng, Haonan Li, Honglin Mu, Xudong Han, Timothy Baldwin, Omri Abend, Eduard Hovy, Lea Frermann (2026). Control Illusion: The Failure of Instruction Hierarchies in Large Language Models (AAAI 2026)
  11. Kiran Kate, Tejaswini Pedapati, Kinjal Basu, Yara Rizk, Vijil Chenthamarakshan, Subhajit Chaudhury, Mayank Agarwal, Ibrahim Abdelaziz (2025). LongFuncEval: Measuring the effectiveness of long context models for function calling (preprint)
  12. Yunjia Qi, Hao Peng, Xiaozhi Wang, Amy Xin, Youfeng Liu, Bin Xu, Lei Hou, Juanzi Li (2025). AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios (preprint)
  13. Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
  14. Rauno Arike, Elizabeth Donoway, Henning Bartsch, Marius Hobbhahn (2025). Technical Report: Evaluating Goal Drift in Language Model Agents
  15. Anthropic (2026). Prompting best practices (vendor documentation, undated; read 2026-10-06)
  16. OpenAI Cookbook (2025). GPT-4.1 Prompting Guide (vendor guide, dated 2025-04-14)
  17. Anthropic (2026). How Claude remembers your project (product documentation, undated; read 2026-10-06)
  18. buschleague (Hacker News) (2026). Hacker News comment: rules followed for the first few steps, then drift
  19. chickensong (Hacker News) (2025). Hacker News comment: followed at the beginning and end, ignored in the middle
  20. pessimizer (Hacker News) (2025). Hacker News comment: escalating a rule to capitals
  21. reedlaw (Hacker News) (2026). Hacker News comment: distrusting the model's own explanation
  22. klardotsh (Hacker News) (2026). Hacker News comment: injected reminders that changed nothing
  23. bayganyo (Hacker News) (2026). Hacker News comment: a blocking check on reply length
  24. franjorub (GitHub) (2026). GitHub issue: project rules loaded by one client and skipped by another
  25. lyonsno (GitHub) (2026). GitHub issue: an instructions file above the repository root is not honored