Home / Blog / Agents at work / Build vs Buy AI Agents: Let the Durable Layers Decide

Agents at work

Build vs Buy AI Agents: Let the Durable Layers Decide

Build vs buy AI agents, decided layer by layer: own the evals, tool contracts and traces, rent the model and framework, and test every exit. See the table.

By Enrique Gutiérrez · Published · 17 min read

Build vs buy AI agents is one decision per layer. The parts that say what good work is for your business (task definitions, test cases, tool contracts, permission rules, traces) outlast any vendor, so you own them. The model, the framework and the runtime change often and are the same for everyone, so you rent them and keep a way out.

I wrote this for a CTO or an engineering manager with three proposals on the table: a finished agent product that demoed well, a prototype on a framework, and an engineer who says forty lines against the model interface would do. The usual build vs buy AI agents comparison treats each proposal as a whole. This one splits the system into layers, gives each layer a verdict, and then asks of every proposal what you could take with you on the day you leave.

Why is “build vs buy AI agents” the wrong size of question?

Because an agent system is several things with different lifespans, and a single verdict for all of them is wrong for some of them. The classic rule for software sits at the level of the business function. Martin Fowler put it this way in 2010: “it’s all about whether the underlying business function is a differentiator or not.” Payroll is his example of the utility side.

The pages that rank for this query apply that rule to the agent as a unit. One services firm’s guide (January 2026) says “build what differentiates you, buy what doesn’t.” A platform vendor’s guide (September 2025) says “For most organizations, the reality isn’t build or buy. It’s both.” Each of the three I read comes from a company that sells one side of the choice, and each stops before saying which part is which.

The same rule works one level down. Chapter 27 of the book (in the full book), in the section “The Ecosystem as Durable Layers,” says of any product that “a product that spans two boxes is two products, and should be judged twice.” A finished agent product spans nearly all the boxes. Judging it once is how a buyer ends up with a good loop and no copy of their own test cases.

People who have built these systems ask the question in that finer grain. An Ask HN thread from May 2026 puts it this way: “For those who built this stuff in-house: was it ever a build-vs-buy conversation? What would a tool have had to do for you to buy instead of build?” The rest of this post is an answer to the second sentence.

Which layers of an agent system last, and which churn?

The layers themselves last and their occupants churn, and inside your own system the things that last are mostly words and data. Chapter 27 strips the brand names off an agent stack and finds six layers, each answering one question: foundation models, frameworks and SDKs, protocols, skills, runtimes and sandboxes, and evaluation and observability. Its verdict on the catalog of products is short: “The catalog is weather.”

The ecosystem as six durable layers with two cross-cutting rails.
Figure 27.3 The ecosystem as six durable layers with two cross-cutting rails. Strip the brand names off any serious agent stack and these six remain, each answering one question the earlier chapters already answered. The two rails—a surface where a person steers, a rail that enforces and records permission—run alongside every layer rather than living in one. The skills layer is picked out (in accent) because it is where most builders actually add value; products churn, but the shelves stay. Reuse this diagram

The section that follows, “Build versus Adopt,” walks the layers and comes back with an inventory. On one side is “what is durably yours: the prompts and skills, the tool contracts, the eval sets and the traces they grew from, the memory files, the goal functions, the guardrail policies”. On the other is “what is rented: the model, the framework, the runtime, the platform.” The chapter’s summary of what that does to the decision: “it converts a shopping question into a custody question, and custody has a discipline: never store what you own inside what you rent.”

To turn the inventory into something a team can fill in, I ask two questions of each layer. The questions and the three verdicts are this post’s own; the placements follow the chapter, except the prompt-wording row, explained below.

  1. Does it outlast a change of model or vendor? A test case with its correct outcome does. A workaround for one model’s habits does not.
  2. Is it about your business, or the same for every company? Your refund limits are yours. Retrying a rate-limited call is the same everywhere.

Anything that is yours gets the verdict own: you hold the master copy in a plain format, in a repository or a store you control. Durability changes how much to invest in it. Anything generic that churns gets rent, with an exit. Anything generic that lasts gets standard: implement the shared one.

What does the table say for each layer?

Twelve rows cover the layers of a typical system. Pick a verdict to narrow the table to the rows that share it.

Layer Outlasts a model or vendor change? Yours or generic? What to do Verdict
Task definitions and acceptance criteria Yes Yours Write them as specs in your repository Own
Instructions, policies and skills Yes Yours Keep them as files; a console holds a copy, never the master Own
Eval set and rubrics Yes Yours Grow the set from your own traces; make it runnable against any option Own
Tool contracts for your systems Yes Yours Names, schemas, descriptions and error text live with your code Own
Data, memory and permission rules Yes Yours Decide who may do what in your own systems and enforce it in code Own
Traces and transcripts Yes Yours Export step-level records to a store you control Own
Prompt wording tuned to one model No Yours Keep it in cheap files and expect to rewrite it on a model change Own
The model No Generic Call it through one seam so that changing it is one edit Rent
The loop and its plumbing No Generic Use a framework for a named need, or write the small version Rent
Runtime and sandbox No Generic Adopt isolation built by specialists; the policy stays in your rows Rent
Scoring machinery and dashboards No Generic Adopt the run-and-score and display machinery; the cases stay yours Rent
Tool protocol and trace format Yes Generic Implement the shared standard and avoid a private variant Standard

How do you read the two rows that look odd?

Prompt wording is owned and still churns, and protocols are neither built nor bought. The book lists prompts on the durably owned side and does not divide them; the division here is mine. A prompt holds two things. The rules it states are durable. The phrasing that makes one model follow them is not, and Chapter 18 gives the habit that goes with it: “on every model change, reread your standing machinery and delete what the new model has made redundant.” Own both, and spend effort on the first.

The standard row has its own verb in the book: “a protocol is valuable in proportion to how boringly standard it is, so you implement the standard and resist the bespoke.” Whether a tool protocol pays off for your number of tools is the subject of MCP vs function calling, and the post on LLM tracing with OpenTelemetry covers one example of a shared trace format.

The eval rows split one book layer in two, as the chapter does: “Adopt the plumbing (trace collection, dashboards, the run-and-score machinery) and own the substance.” An eval set is a list of cases with the outcomes you count as correct. No vendor has that list for your task.

What do the three options hand you, layer by layer?

A finished product rents you everything and stores your owned layers on its side unless you arrange otherwise; a framework rents you the loop; the bare model interface rents you only the model. The table sets the three against the same layers. It describes categories, and any single product can differ from its row, which is why the exit questions further down exist.

Layer Buy a finished agent product Build on a framework Build on the bare model interface
The model Vendor picks it and may change it You pick, through the framework’s adapter You pick, through your own seam
The loop and plumbing Vendor’s, closed The framework’s, readable when the source is open Yours to write and maintain
Runtime and sandbox Vendor’s You choose one You choose one
Task definitions and instructions Typed into the vendor’s console Your files, in the framework’s shapes Your files
Tool contracts Vendor’s connectors, or your own tools if a standard protocol is accepted Your functions, with the framework’s wrappers Your functions
Eval set Yours only if you write one and can run it against the product Yours, with any runner Yours, and you write the runner
Permission rules Vendor’s admin settings, plus the credentials you grant Your code Your code
Traces Vendor’s store; export is the question Hooks into a store you pick The logging you write
What you write first Configuration The agent’s logic The agent’s logic and its plumbing

Two rows deserve a second look before a meeting. The model row for a finished product means a change you did not make can alter the agent’s behavior; the question about release notice is number 15 in the post on AI agent reliability for enterprise buyers. And the eval row says that buying removes none of the work of defining good output. Chapter 27 is plain that the substance of evaluation is something “no product can sell you.”

When is a framework the right middle?

When you can name the plumbing it saves you from writing. Chapter 3, The Agent Loop (in the full book) lists what a framework adds to a loop: state and persistence, retries, tracing, memory management and multi-agent orchestration. Its rule is “Adopt for a named need, never for the feeling that serious systems use frameworks,” and its limit is that “a framework can make an agent more dependable and cannot make it smarter.”

Schluntz and Zhang’s essay, written from work with teams building agents, reports that “the most successful implementations weren’t using complex frameworks or specialized libraries” and adds a condition for those who adopt one: “If you do use a framework, ensure you understand the underlying code” (Anthropic, 2024). The post on what an agent harness is has the full table of framework parts, the need that justifies each, and an audit list. I won’t repeat it here.

What does leaving cost under each option?

Leaving costs the rework of every owned layer that stayed behind, plus whatever history cannot be recovered at all. I found no public measurement of that cost by option, so this section gives a way to count it for your own case. Chapter 27 names three kinds of lock-in, and each maps to rows of the first table.

Model lock-in “lives in code that depends on one provider’s raw response shape or calling conventions”. The chapter calls it the pettiest and the most common, and its insurance is “one seam in your codebase through which all model calls pass,” which the glossary calls a model client.

Orchestration lock-in “lives in workflows that exist only in a vendor’s configuration language”. The insurance is logic in ordinary code and prompts in ordinary files.

Data lock-in is “whether your transcripts, memories, and eval results can leave.” The chapter says teams often discover it last, and adds: “The question costs nothing asked before adoption and a migration afterward; ask it before.”

A forum comment from June 2026 objects that lock-in can’t be real: “How can you lock in when the harnesses are basically thin clients around the APIs and you can replicate them using agents in a short period of time?” For the rented rows I think that is right, and it is the argument of this post: the loop is cheap to replace. The comment says nothing about two years of traces and an eval set that exist only in a vendor’s database.

What is the exit test?

The exit test asks, for each of five durable layers, which of three things is true. The five are task definitions and instructions, the eval set and rubrics, tool contracts, data and permission rules, and traces. I fold prompt wording into the first, since it gets rewritten on a move anyway.

  • Leaves with you: the master copy is in a store you control, in a format another tool can read.
  • Mirrored: the working copy sits with the vendor and you keep the master, so leaving means re-entering it.
  • Stays behind: the only copy is on the vendor’s side, or the export loses what you need.

Before testing anything, mark each layer must keep or would rebuild. An audit trail in a regulated process is a must-keep. Five connectors to systems with documented interfaces may be a rebuild you accept. An option that leaves a must-keep layer behind is out until the vendor’s answer changes.

These are the questions that produce the answers. Send them to a vendor before a contract, or ask them of a framework before the second sprint.

  • Where is the master copy of our instructions and policies, and can we export it as plain text at any time?
  • Can we run our own test cases against the agent, unattended, and get per-case results out?
  • Do our tools connect through a standard protocol or through connectors that exist only here?
  • Can we export a step-level trace of every run, with tool calls and results, in a documented format?
  • Does trace export still work after the contract ends, and for how many days?
  • Are the permission rules for what the agent may do readable outside your admin screen?
  • Can we pin the model that runs our tasks for a period we choose, and move our instructions unchanged if you switch it?
  • If we stop paying, what do we hold on that day, and in which file formats?
  • Which parts of our logic are written in a language or format that only this product reads?
  • For each part we rent, what is the second supplier, and what would switching require us to rewrite?

Worked example: an invoice-matching agent

In this worked example one contract clause decides between buying and building. The numbers are illustrative. A finance team wants an agent that matches supplier invoices to purchase orders, posts the clean ones and flags the rest. The owned layers come to 4 instruction files (a matching policy and three skills), 120 past invoices with their correct outcome, 5 tool contracts for the accounting system, and 8 permission rules such as posting limits. The agent pauses for a person on flagged invoices, so of Chapter 3’s five plumbing parts it needs three: persistence, retries and tracing.

The team marks four layers must keep and marks tool contracts would rebuild, since the accounting system’s interface is documented. Then it runs the exit test on what each option’s supplier says.

Durable layer Finished product Framework Bare interface
Task definitions and instructions Mirrored (console; master kept in the repository) Leaves Leaves
Eval set and rubrics Leaves (cases run through a test tenant) Leaves Leaves
Tool contracts Stays (built-in connectors) Leaves Leaves
Data and permission rules Mirrored (admin settings; rules kept as a document) Leaves Leaves
Traces Stays (outcomes export, steps do not) Leaves Leaves
Count 1 leaves, 2 mirrored, 2 stay 5 leave 5 leave
Plumbing parts the team writes 0 of 3 0 of 3 (rented) 3 of 3

The product leaves one must-keep layer behind: traces. Finance needs to reconstruct why an invoice was posted, and an outcome log does not show the steps. As it stands, the product is out. If the vendor adds step-level trace export to the contract, only tool contracts stay behind, the team already accepted that rebuild, and the product is back in.

The count also prices the exit. Leaving the product later means redoing 17 items (4 instruction files re-entered elsewhere, 5 connectors rebuilt as tools, 8 rules re-implemented). Leaving the framework means redoing 6 (5 tool wrappers and 1 loop). The product’s 17 counts owned items only; whatever replaces it also needs a loop and the three plumbing parts, rented or written. Neither number is large, and the gap between them is not what decides this case. The trace clause is.

With the clause, the team buys: the product and the framework both leave no plumbing to write, and the product also leaves nothing to operate. Without it, the team takes the framework, because persistence for the human pause is a named need and writing it is the bare option’s cost. In both outcomes the 120 cases sit in the team’s repository and run against whatever is chosen. Whether the agent is worth running at all is a separate sum, covered in the post on AI agent ROI.

Which questions settle build vs buy?

Four steps settle it, and they are the same for a product, a framework or a plan to write everything. The rule is this post’s own, built on the book’s custody discipline.

  1. Mark the five durable layers must keep or would rebuild for this agent.
  2. Run the exit test on each option. Drop any option that leaves a must-keep layer behind.
  3. Among the options still in, take the one that leaves you the least to write and operate, and drop any rented part whose need you cannot name.
  4. Record the exit for each rented part and the event that reopens the decision.

Step 3 has one precondition from Chapter 3: a team that has never written a loop should write one small one first, because “the bare version teaches you what every part is for.” Otherwise nobody on the team can judge a vendor’s answers. Step 1 goes faster when the task is already written down as a spec, which is the subject of spec-driven development with AI agents. And a product manager is often the person who can say which layers are must-keep, a point the post on AI agents for product managers takes further.

The memo below holds the result on one page.

BUILD VS BUY DECISION: <agent name>            Date: <date>   Owner: <name>

1. What the agent does, in one sentence:
2. Durable layers (must keep / would rebuild) and where the master copy lives:
   - Task definitions and instructions:
   - Eval set and rubrics (number of cases):
   - Tool contracts:
   - Data and permission rules:
   - Traces (retention we need):
3. Options considered and the exit test for each
   (per layer: leaves / mirrored / stays behind):
   - Finished product:
   - Framework:
   - Bare model interface:
4. Options dropped, and the must-keep layer each left behind:
5. Decision, and the named need each rented part serves:
   - Model:
   - Loop and plumbing:
   - Runtime and sandbox:
   - Scoring machinery and dashboards:
6. Exit for each rented part (second supplier, what we rewrite):
7. Reopen this decision when: <a model or price change, a contract renewal,
   a must-keep layer we can no longer export, a plumbing need we can now name>

Where does AI agent technical debt pile up?

It piles up where your code or your records took a vendor’s shape, and in plumbing nobody budgeted for. Chapter 27 locates the first kind exactly: “The lock-in is rarely the framework’s name on the box; it is the hundred small places where your code absorbed a vendor’s shape because absorbing it was, that afternoon, the path of least resistance.”

The second kind is older than agents. Sculley and colleagues wrote of machine learning systems in general, in 2015, that “it is common to incur massive ongoing maintenance costs in real-world ML systems.” That paper predates LLM agents, and I cite it for the framing only. A reply in the May 2026 thread gives the agent-era version: “building the ‘boring infra stuff’ around it is a time-sink no one prepares you for…” (Hacker News).

Both kinds argue for the same split. Rent the boring infrastructure when a need is named, and keep the vendor’s shape out of the owned rows. A tool definition that passes the tool contract linter in plain form will survive a change of framework with a new wrapper around it.

Is it an agent at all, or plain automation?

Check that first, because a fixed sequence of steps needs none of this. If the path through the task is known in advance, a workflow with model calls at fixed points is cheaper to run and easier to test, and build versus buy becomes an ordinary software question. The post on when not to use AI agents gives the tests, and the should this be an agent? tool walks through them for one task.

Where does this rule not apply?

It does not apply to throwaway work, and it has three other limits. First, on prototypes the book is direct: “A weekend prototype deserves none of this; wrap nothing, hard-code freely, let it be disposable.” It also warns against the opposite error: “Portability has a price, and buying it everywhere is its own failure mode.”

Second, the split assumes you can write your own test cases. A team with no examples of correct output cannot run the exit test on the eval layer, and should collect cases before choosing anything.

Third, a generic function may have no owned layer worth protecting. A meeting-notes summarizer for internal use has no tool contracts and a trace nobody audits. Mark everything would rebuild, and the rule returns “buy” at step 3 without ceremony.

Fourth, the table leaves out the screen where people meet the agent and the organization’s governance process. Chapter 27 draws both as rails beside the layers. A finished product’s review screen can be the main thing you are paying for, and nothing here measures it.

The line to hold in the meeting

The choice among a product, a framework and the bare interface is a choice about the rented rows, and rented rows are cheap to change if the owned rows stayed with you. So the question to ask of every proposal is which of your five layers it would keep. Chapter 27 closes the section with the same instruction: “Rent the engine. Own the judgment. Keep the receipts in your own drawer.”

The layers and the custody rule are in Chapter 27, “The Frontier and How to Keep Learning”, in the full book; the framework question is settled at the scale of one loop in Chapter 3, The Agent Loop (in the full book). The Agents at work guide collects the related posts and tools, and you can see the formats.

Questions readers ask

Should you build or buy an AI agent?
Decide it per layer. Own the parts that are about your business: task definitions, the eval set, tool contracts, permission rules and traces. Rent the parts that are generic and change often, such as the model, a framework and the runtime. Buy a finished product only if those owned parts can leave with you or you would accept rebuilding them.
Should you use an agent framework or build your own?
The book's rule in Chapter 3 is to adopt for a named need, such as runs that must survive a restart, traces you need, or several agents to coordinate. With no such need, a loop written against the model interface is small and fully visible. Either way, keep instructions, tool definitions and test cases in your own files.
What is vendor lock-in for AI agents?
Chapter 27 of the book names three kinds. Model lock-in is code that depends on one provider's response shape. Orchestration lock-in is workflows that exist only in a vendor's configuration language. Data lock-in is transcripts, memories and eval results that cannot leave. The third is the one the chapter says teams often discover last.
What should you own when you buy an AI agent product?
Own the written definition of the task, a set of test cases with correct outcomes that you can run against the product, the rules for what the agent may do, and a copy of the step-level record of what it did. Those four are about your business, and a later vendor or an in-house build will need all of them.
Can you build AI agents without a framework?
Yes. An agent is a model call in a loop with tools, a message history and stop rules, and the minimal version is short. What a framework adds is plumbing: persistence, retries, tracing, memory management and orchestration. You write those parts yourself when you need them, or rent them when writing them costs more than the dependency.

Sources

  1. Erik Schluntz and Barry Zhang, Anthropic (2024). Building Effective Agents
  2. Martin Fowler (2010). Utility Vs Strategic Dichotomy
  3. D. Sculley et al., NeurIPS (2015). Hidden Technical Debt in Machine Learning Systems
  4. Retool (2025). Build vs buy AI agents: why custom solutions win long-term
  5. Turing (2026). Build vs. Buy AI Agents: Everything You Need to Know
  6. Catie Grasso and Catalina Herrera, Dataiku (2025). Build vs. buy for AI agents: A practical guide
  7. Hacker News thread (2026). Ask HN: What are your worst war stories bringing agentic applications into prod
  8. Hacker News thread (2025). Ask HN: Why Is Every Company Building Own Agent Framework? Isn't One Enough?