AI agent pricing models come in four kinds: per seat, usage, per task and per outcome. An agent’s cost per task rises with the square of its step count, so a single flat price quietly subsidizes the hardest work, and per-outcome pricing, the most attractive on a slide, is the most dangerous to start with.
That claim rests on arithmetic you can check, and this post does the arithmetic in the open. It prices one illustrative agent by task difficulty, runs the same customer through each model twice (once with an easy mix of work, once with a harder one), and shows which models break. Then it argues for an order of moves: start hybrid, bill outcomes in shadow, and switch only when you can verify what you are billing for.
What are the main AI agent pricing models?
The main AI agent pricing models are per seat (a flat fee per user or per agent), usage (tokens, credits or compute units), per task (a flat fee per action, conversation or workflow) and per outcome (a fee only when a defined result happens). Most real price pages mix them, and the mix decides who carries the cost variance.
The taxonomy is not settled. Kyle Poyar and Manny Medina, after analyzing “60+ AI agent companies” in April 2025, cut it as per agent, per action, per workflow and per outcome; other investors and vendors draw the lines slightly differently. I use the four AI agent pricing models below because each one puts the variance in a different place, and that is the property this post is about. The table describes each one; the buttons narrow it by who absorbs a hard month.
| Model | What the buyer pays for | Predictable for the buyer? | What it needs to work | Dated examples of the category | Who carries the variance |
|---|---|---|---|---|---|
| Per seat or per agent | Access, per person or per deployed agent | Yes | Usage per seat that stays within a narrow band | GitHub Copilot seat plans, which by October 2026 came with a monthly allowance of included AI credits; Cursor, which in June 2025 moved its individual Pro plan from request counts to included usage | vendor |
| Usage (tokens, credits, compute units) | Consumption, converted at a published rate | No: the bill tracks the work | A buyer who can forecast volume, and alerts before surprises | Salesforce Agentforce Flex Credits, announced May 2025 at “20 Flex Credits ($0.10 per action)”; the included credits under GitHub’s seats | buyer |
| Per task (action, conversation, workflow) | A unit of work, at one flat price | Mostly | Tasks of similar difficulty | Salesforce’s $2 per conversation (2024); Replit, which in June 2025 replaced a flat price per checkpoint with effort-based pricing | vendor |
| Per outcome | A defined result, such as a resolved ticket | Yes, per result | An outcome both sides can verify and attribute | Intercom Fin at “$0.99 per outcome” (pricing page, October 2026); Zendesk, which in August 2024 said customers pay only for issues “resolved autonomously by AI” | vendor |
| Hybrid | A base fee with included usage, overage and a cap | Within the cap | Cost per task known by class | Sierra, which in December 2024 offered blended outcome and consumption pricing; most seat plans with included credits | shared |
Two of those dated examples tell the same story from different companies. Both Cursor and Replit started with a flat unit over work whose difficulty varied enormously, and both moved to a price that follows effort. Cursor’s July 2025 explanation is the clearest primary statement of the problem I have found: “the hardest requests cost an order of magnitude more than simple ones. API-based pricing is the best way to reflect that.” It also apologized and refunded “unexpected charges,” which is what a pricing change discovered late costs in public.
Why does a hard task cost so much more than an easy one?
A hard task costs disproportionately more because an agent re-sends its whole transcript on every step, so each step pays for all the steps before it. The input bill therefore has a term that grows with the square of the run’s length, and long runs are where most of the money goes.
The mechanism is in Chapter 19 of the book (in the full book), which describes a run’s input bill as “the fixed equipment times the number of steps, plus the per-step growth times a term that rises with the square of the run’s length.” In symbols, with a fixed prefix F (system prompt, tool definitions, task), a growth g per step (what the model writes plus what the tool returns) and n steps, input tokens come to F·n + g·n(n−1)/2. Output is roughly the model’s own writing per step times n, and it is the small part: in an agent, the chapter notes, “the agent’s economics are mostly input economics.”
Here is the chapter’s napkin applied to three task classes. Everything in this table is illustrative: F = 3,200 tokens, 300 tokens written and 700 returned per step (so g = 1,000), and round prices of $3 and $15 for each million input and output tokens, chosen for easy arithmetic and matching no vendor’s rate. No retries yet.
| Task class | Steps n | Input tokens (F·n + history) | Output tokens | Cost per run |
|---|---|---|---|---|
| Easy | 5 | 16,000 + 10,000 = 26,000 | 1,500 | $0.078 + $0.0225 = $0.1005 |
| Typical | 10 | 32,000 + 45,000 = 77,000 | 3,000 | $0.231 + $0.045 = $0.276 |
| Hard | 30 | 96,000 + 435,000 = 531,000 | 9,000 | $1.593 + $0.135 = $1.728 |
The hard task has six times the steps of the easy one and costs about seventeen times as much. At thirty steps, history being re-read is 82% of the input, and the naive estimate (prefix times steps, 96,000 tokens) is low by a factor of 5.5. The estimator below opens on the hard row; move the step slider to 10 or 5 to reproduce the other two, and the retry slider to see what failures add.
With JavaScript on, the Agent cost-per-task estimator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Agent cost-per-task estimator on its own page to share a result by link.
Retries multiply the whole bill. The chapter cites one cost-governance framework that puts them at ten to twenty percent of consumption in poorly instrumented systems; at 10% the hard run becomes about $1.90.
The lever that most often cuts the fixed part is prompt caching (billing a repeated prefix at a discount), and the prompt caching savings calculator shows how much it removes for your own prefix. Neither changes the shape: the hard tail still dominates. A coding agent working in a legacy codebase is a good picture of that tail, long runs that read a lot of unfamiliar code before they change any of it.
How does each pricing model hold up when the work gets harder?
Each model holds up differently when the mix of tasks gets harder: flat prices (per seat, per task) hand the extra cost to the vendor, usage pricing hands it to the buyer, and per-outcome pricing hands it to the vendor twice, because hard tasks both cost more and succeed less often. The worked example makes the size of that difference visible.
Take one customer running 1,000 tasks a month. In mix A the work is 60% easy, 30% typical and 10% hard, which costs 600 × $0.1005 + 300 × $0.276 + 100 × $1.728 = $315.90. The hard tenth of the volume is 55% of that cost. In mix B the customer has learned what the agent can do and hands it harder work: 40% easy, 30% typical, 30% hard, for $641.40, with the hard share now 81%.
Nothing about the agent changed. The AI agent cost per month doubled because the customer got better at using it, and this is the test every one of the AI agent pricing models has to pass.
I calibrated every model to earn the same $500 a month on mix A, then ran mix B through it. The prices are illustrative; the success rates for the outcome row (95% easy, 85% typical, 60% hard) are assumptions, chosen to show the direction.
| Model (illustrative price) | Revenue, mix A | Margin, mix A | Revenue, mix B | Margin, mix B | Who carries the variance |
|---|---|---|---|---|---|
| Per seat: 5 seats × $100 | $500.00 | 36.8% | $500.00 | −28.3% | Vendor |
| Per task: $0.50 a task | $500.00 | 36.8% | $500.00 | −28.3% | Vendor |
| Usage: cost × 1.583 (500 ÷ 315.90) | $500.00 | 36.8% | $1,015.19 | 36.8% | Buyer (the bill doubles) |
| Per outcome: $0.565 a success | 885 successes → $500.03 | 36.8% | 815 successes → $460.48 | −39.3% | Vendor, who also pays for failed runs |
| Hybrid: $400 base incl. $250 of compute at cost, overage at cost × 1.583 (500 ÷ 315.90) | $504.31 | 37.4% | $1,019.50 | 37.1% | Shared, bounded by the buyer’s cap |
Per seat and per task land on the same numbers here because both are flat; the per-task row hides a sharper story inside it. At $0.50 a task the easy class earns an 80% margin, the typical class 45%, and every hard task loses 246%. That is the subsidy in plain view: the easy work pays for the hard work, until the mix shifts. Under per seat, a successful agent can hurt you a second way. Sierra put it bluntly in December 2024: “the more effective their AI becomes, the fewer contact center seats their clients need.”
Usage pricing keeps the margin and moves the shock to the buyer, which is honest and also the reason buyers fear meters. The hybrid row only looks safe. Its mix B bill is $1,019.50, almost the same as pure usage; what the hybrid adds is a floor the buyer understands and a cap the buyer controls, so the doubling arrives as a conversation about a budget before it arrives on an invoice. In the buyer’s own unit, the hybrid bill is about $0.50 a task on mix A and $1.02 a task on mix B, a number the buyer can set against what a task costs them today.
Why is per-outcome pricing the riskiest place to start?
Per-outcome pricing is the riskiest place to start because the meter underneath it never stops: the vendor pays for every run, successful or not, while the buyer pays only for successes. Hard tasks fail more and cost more, so a single outcome price is the deepest subsidy of the four models.
Chapter 19’s image for this is a taxi meter, because the token bill “charges for the distance traveled, with complete indifference to whether the ride arrived anywhere useful. A run that fails after forty steps is billed for forty steps.” Under outcome pricing, the vendor owns that meter. Divide each class’s cost by its success rate and you get the number that matters, the cost per successful outcome: about $0.11 for easy tasks, $0.32 for typical ones and $2.88 for hard ones in the example. One price across all three cannot be right for more than one of them.
Bessemer’s February 2026 pricing playbook names the same risk from the investor’s side. “The risk is real: a difficult customer issue could consume far more compute than anticipated.” Tiering outcomes by difficulty helps with the cost side. Some outcome-priced products already price outcome types differently: in October 2026, Fin’s page listed qualifications at $9.99 each against $0.99 per resolution. It does not help with the harder problem, which is what counts as an outcome at all.
Can you verify the outcome you are billing for?
You can bill for an outcome safely only when both sides can verify it from records the agent does not control. Otherwise the outcome is a proxy, and outcome pricing pays the vendor on a signal the vendor’s own system partly produces, which is the trust problem the book treats in two chapters.
Look at a real definition, as an example of the category. Intercom’s Fin pricing page (read October 2026) bills a resolution when “No further help is requested after Fin’s last answer.” That is a measurable event, and the page also states that you are not charged “when a conversation is simply passed to your team without an outcome,” which is a good rule. (The same page does bill a handoff when Fin completes a Procedure you configured to end in one, so “handoff” needs its own line in any contract.) But the event is the customer’s silence.
Silence can mean the problem was fixed, or that the customer gave up. The definition is workable precisely because support has an outside check the vendor does not control: a repeat contact, a refund, a satisfaction score, a baseline cost per ticket.
Chapter 23 (in the full book) explains why support is the friendly case: “A deflection rate (the share of contacts the agent resolves with no human touching them) against a known cost per contact is a business case that computes itself.”
Research and advisory work sits at the other end. It lives behind what the chapter calls the verification gap: the distance between how convincing an output looks and how cheaply it can be checked. The chapter’s cost line is blunt. “An agent whose output costs as much to verify as to produce by hand has a return of approximately nothing, however impressive the demo.” The same holds for an outcome price: if checking the outcome costs as much as doing the work, the price is fiction.
There is also a quieter risk in how outcome pricing feels to the buyer. Chapter 24 (in the full book) warns that an interface built to make users trust an agent reaches for confident tone and polished summaries. “Every one of those moves raises trust without touching reliability, which is to say every one of them manufactures over-trust.”
An outcome counter on an invoice is the same kind of move. The goal is appropriate reliance, so the outcome definition should come with the evidence a buyer can audit. Poyar and Medina reach the same conclusion from market patterns. “If you don’t have a clear path to attribute results to your agent either via A/B testing or with a POC, this may not be the best route for you.”
What should an outcome contract include before you charge for it?
An outcome contract should define a result both sides can observe, say what does not count, set rules for partial and late outcomes, and give the buyer a way to dispute. It should also rest on your own measured cost per success by task class. Work through the list against your product before outcome pricing goes on a quote.
- The outcome is observable by both parties from records the agent cannot write by itself (a closed order, a paid invoice, no repeat contact within an agreed window). The agent’s own report never counts.
- The definition names what is not billable: escalations to a human, handoffs, and repeat contacts on the same issue.
- Partial and lagging outcomes have a written rule: when they are credited, at what fraction, and how long you wait.
- There is a dispute window, and a sample of billed outcomes is checked by a person on a fixed schedule.
- You know cost per successful outcome for each task class, measured from real traces.
- Outcome prices are tiered by difficulty, or the hard class is priced or capped separately.
- At least one full billing cycle of outcome invoices has been computed in shadow beside the real bill, with the disputes counted.
- The buyer has a monthly cap, and you have a minimum that covers your fixed costs.
- The agent cannot produce the billable signal by ending the interaction early, closing a ticket, or declaring success.
How do you start pricing an agent without betting the company?
Start hybrid, shadow-bill outcomes, then move. A base fee with included usage, overage at a published rate and a buyer-set cap keeps your margin through a mix shift; computing outcome invoices in parallel shows you, at no risk, what outcome pricing would have earned and how often it would have been disputed.
The order matters because each step produces the evidence the next one needs. The hybrid period gives you real traces, and from them your cost per task by class, the distribution that the napkin only guesses at. The shadow period gives you cost per success and a dispute rate, and it lets the buyer see outcome invoices before they are real. If the shadow invoices are disputed often, you have learned that cheaply.
Set the base fee from the costs you can see. In the example, $250 of included compute covers most of mix A ($315.90) and roughly 40% of mix B, so the base fee works as a commitment from the buyer and a budget on the vendor’s side. Do the sizing with your own numbers, task class by task class; the companion post on how much an AI agent costs to run walks through gathering them from traces. If you are still deciding which work to give an agent at all, the post on AI agent use cases for startups is the earlier question.
Then keep measuring the right quantity. The book’s warning holds for pricing as much as for engineering: “a stable invoice can conceal token volume growing as fast as prices fall, so watch consumption, never just the bill.” Token prices tend to fall, and customers who trust the agent tend to hand it longer tasks. Both forces act on the same price page.
What margins should you plan for and explain to investors?
Plan for gross margins well below software-as-a-service benchmarks, and explain them with your cost per task by class. Every outside figure here is a dated snapshot from one firm’s sample, so use them as context for a conversation with a board.
In February 2020, before the current generation of models, Martin Casado and Matt Bornstein of a16z wrote that AI companies showed “gross margins often in the 50-60% range – well below the 60-80%+ benchmark for comparable SaaS businesses.” Bessemer’s February 2026 playbook gives a similar range: “Companies see 50-60% gross margins vs. 80-90% for SaaS.”
ICONIQ’s July 2026 report, from surveys of over 300 executives, says “Gross margins are improving, jumping from 45% in 2025 to a projected 53% in 2026, and 59% in 2027.” It adds that as products scale, “talent’s share of cost falls while inference rises.” The worked example’s 36.8% is lower than all of these only because I chose a simple markup, so read it as arithmetic about one invented product.
What an investor or a board actually needs is the shape behind the margin: what share of cost comes from the hard tail, what happens to margin when the mix moves, and what the cap does to the worst month. For leaders weighing agentic AI as a business decision, the companion piece on framing AI agent ROI against risk adds the other two columns of the ledger, what an error costs and what checking costs.
Where does this advice not apply?
The advice to start hybrid and move to outcomes last fits products where task difficulty varies widely and outcomes are hard to verify. Where tasks are uniform, a flat per-task price is simpler and fine; where the outcome is cheap to verify and the vendor already has cost-per-success data, outcome pricing can come first.
The example has real limits.
Its prices, step counts and success rates are illustrative, and real cost per task is a distribution with a long tail that three tidy classes only approximate; the agent cost per task estimator handles one run at a time, and only your traces show the spread. A commenter in a July 2026 Indie Hackers thread on token costs and margin risk put it in one line: “the workflow cost isn’t a number, it’s a distribution.” (By their own disclosure, the commenter had found a competing tool.) Pricing is also a market decision. A buyer who already pays per resolution elsewhere may refuse anything else, and a competitor’s per-outcome page may force your hand before your data is ready. In that case, the checklist above becomes the negotiation: tier by difficulty, cap the hard class, and write the dispute rule into the contract.
There are alternatives the four models leave out. A share of measured savings is one; Zuora’s August 2025 survey of the category cites a vendor charging a percentage of the cloud costs it saves. Another is pricing the human-equivalent work in the buyer’s own units. Both are outcome pricing in a different coat, and both carry the same verification question.
The one thing to keep
AI agent pricing models differ mainly in who carries the cost of a hard task, and an agent’s hard tasks cost far more than its average suggests, because cost grows with the square of the steps. Price on the curve by task class, start hybrid with a cap, run outcomes in shadow, and bill for outcomes only once you and the buyer can both check them. As Chapter 19 puts it, teams that skip the napkin “discover their unit economics from the first month’s invoice, which is the most expensive way to learn arithmetic.”
The cost model behind every number here is in Chapter 19, “Cost, Latency, and Performance”, in the full book; the ROI frame is in Chapter 23 and trust calibration in Chapter 24, both also in the full book. The Agents at work guide collects the related posts and tools, and you can see the formats.
Questions readers ask
- What is the most common AI agent pricing model?
- Subscriptions remain the most common, and pure models are rare. ICONIQ's July 2026 survey of over 300 executives found consumption-based pricing rising from 35% to 42% of companies in six months and outcome-based from 18% to 23%. Treat those shares as a dated snapshot.
- How much does an AI agent cost to run per month?
- Multiply your task count by the cost per task of each difficulty class, because an average hides the hard tail. In this post's illustrative example (round prices, not any vendor's rate), 1,000 tasks a month cost $315.90 with 10% hard tasks and $641.40 with 30%, before retries, which one cost-governance framework puts at 10 to 20% more in poorly instrumented systems.
- Should I charge per seat for an AI agent?
- Only with included usage and a cap underneath. A seat price is predictable for the buyer, but the vendor carries all the cost variance, and a working agent can reduce the number of seats a customer needs. Several coding-agent products that began with flat seat plans added usage meters under them.
- When does outcome-based pricing work?
- When the outcome is cheap to verify by both sides, attributable to the agent, and priced by difficulty. Support resolutions with a measurable baseline are the clearest case. Where the outcome is a judgment about words, such as a research brief, the outcome becomes a proxy the vendor partly controls.
- How do I protect margins from heavy users?
- Price on cost per task by class; include a usage allowance in the base fee; charge overage at a known rate; and give the buyer a monthly cap they set. Then watch token consumption per task, because a stable invoice can hide volume rising as prices fall.
Sources
- Kyle Poyar and Manny Medina (Growth Unhinged) (2025). A new framework for AI agent pricing
- Cursor (2025). Clarifying our pricing (June 2025 pricing update)
- Replit (2025). Introducing Effort-Based Pricing for Replit Agent
- Indie Hackers thread (2026). How are AI product founders keeping token costs from turning into margin risk?
- Sierra (2024). Outcome-based pricing for AI agents
- Intercom (retrieved 2026). Fin pricing
- Bessemer Venture Partners (2026). The AI pricing and monetization playbook
- ICONIQ (2026). 2026 State of AI Report: The Builder's Economy
- Martin Casado and Matt Bornstein (a16z) (2020). The New Business of AI (and How It's Different From Traditional Software)
- Salesforce (2025). Agentforce flexible pricing (press release)
- GitHub (retrieved 2026). GitHub Copilot plans
- Tien Tzuo (Zuora) (2025). The four kinds of agentic AI pricing models
- Zendesk (2024). Zendesk outcome-based pricing (newsroom)