Guide
What an AI Agent Run Actually Costs
The honest answer is that you cannot get it from a per-request average, and the reason is measurable. A Stanford Digital Economy Lab study of eight frontier models on SWE-bench Verified, published in April 2026, found that runs on the same task can differ by up to 30x in total tokens. Same task, same prompt, thirty times the bill.
So the unit that matters is the run, not the request. A worked example below puts one twelve-step coding run at about $0.85, and about $0.47 once the stable part of its context is cached. The spread around those numbers is the harder problem.
Why a run is a different animal from a request
A chat turn is one call with a roughly knowable size. An agent run is a session: the model calls a tool, reads the result, decides what to do next, and repeats until it is finished or stopped. Nobody knows in advance how many steps that takes, including the model. The Stanford authors found that frontier models fail to predict their own token usage, with correlations reaching only 0.39.
The multiples are large and come from the vendors' own measurements. Anthropic's engineering team reports that agents use about 4x the tokens of a chat interaction and multi-agent systems about 15x, and that token usage alone explains 80% of the variance in performance. The Stanford figure for agentic coding tasks against ordinary code chat is around 1000x.
This is what caught Uber. TechCrunch reported in June 2026 that the company burned its entire 2026 AI budget in four months as coding agent use spread, and responded with a monthly cap of $1,500 per employee per tool. A company with a $3.4 billion research budget did not have a forecasting problem. It had a metering problem, and the fix it reached for was a cap.
The average is the wrong number
Divide a monthly bill by request count and you get a figure that looks reassuring. Our worked run below costs $0.85 across twelve calls, which is about seven cents a call. Seven cents sounds like nothing.
Now apply the 30x spread. The same run, on a bad day, is $25.63. A thousand runs a day at the mean is a $25,600 month, which is the number that goes in the plan; the month you actually get depends on how heavy the tail was, and nothing in the average tells you. Averages describe a distribution that agents do not have.
The practical consequence: forecast from percentiles, and enforce on the run.
What actually drives the cost
Five mechanisms account for most of it:
- Context replay. Every step re-sends the conversation so far, so input grows as the run proceeds.
- Fan-out. A supervisor that spawns subagents multiplies the run rather than adding to it.
- Retries. A failed tool call or a malformed response repeats work already paid for.
- Reasoning tokens, billed as output, invisible in the response body.
- Tool results re-entering context. A large file read once is then carried for the rest of the run.
Most cost advice tells you output is the expensive one, and for chat it is: output runs about five times the input rate on most tiers. The Stanford study found the opposite for agent runs, with input tokens rather than output tokens driving overall cost. Context replay is why.
That flips what you should optimise. Trimming responses saves little. Caching the stable prefix saves a lot, because on Claude models a cache read costs roughly a tenth of the base input rate.
One run, costed out
A twelve-step coding run on Claude Sonnet 5, at the published rates of $2 per million input tokens, $10 per million output, $0.20 per million cache reads and $2.50 per million cache writes.
Stable prefix (system prompt + repo context): 20,000 tokens, unchanged all run
Each step adds ~2,200 tokens to context (tool output + prior response)
Each step produces 700 output tokens
Steps: 12
Input, step k = 20,000 + (k-1) x 2,200
Total input = 385,200 tokens
Total output = 8,400 tokens
Without caching:
input 385,200 x $2 / 1M = $0.7704
output 8,400 x $10 / 1M = $0.0840
run cost = $0.854
Input is 90% of that. Now cache the 20,000-token prefix: write it once, read it on the other eleven steps.
cache write 20,000 x $2.50 / 1M = $0.0500
cache reads 220,000 x $0.20 / 1M = $0.0440
other input 145,200 x $2.00 / 1M = $0.2904
output 8,400 x $10.00/ 1M = $0.0840
run cost = $0.468
A 45% cut, and every cent of it came from the input side. Had you instead squeezed the responses to nothing, the ceiling on your saving was eight cents.
Per-run accounting, in practice
The mechanics are the same wherever you meter. Give each user-initiated action an identifier, set server-side, and attach it to every model call that action causes. With an in-path setup the integration is a base URL change plus metadata injection for full attribution:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://www.spendline.ai/v1",
apiKey: process.env.OPENAI_API_KEY,
defaultHeaders: {
"x-spendline-key": process.env.SPENDLINE_API_KEY,
},
});
const response = await client.chat.completions.create(
{ model: "claude-sonnet-5", messages },
{
headers: {
"x-agent-id": agentRun.id, // every call in this run shares it
"x-customer-id": account.id, // who the run was for
"x-workflow-id": "repo_refactor",
},
}
);
Cost per run then falls out of the ledger, and so does the shape you actually need:
SELECT agent_id,
SUM(cost_usd) AS run_cost,
COUNT(*) AS calls
FROM ai_calls
WHERE created_at >= '2026-09-01'
GROUP BY agent_id
ORDER BY run_cost DESC
LIMIT 20;
That ordering matters more than the average it replaces. The twenty most expensive runs of the month are where the money went, and they are usually a handful of loops rather than broad growth.
What each control actually stops
| Control | Stops a loop | Stops an expensive single run | Knows about dollars | Where it fails |
|---|---|---|---|---|
| Monthly provider spend alert | No | No | Yes | Arrives after the money is spent |
| Per-user monthly cap | Eventually | No | Yes | A single run can exhaust it in minutes |
| Step or depth limit | Yes | No | No | Counts steps; a 5-step run on a huge context is still expensive |
| Per-run token ceiling | Yes | Partly | No | Token counts differ in price by model |
| Per-run dollar cap, checked before each call | Yes | Yes | Yes | Needs in-path metering |
Two rows do the work. A control that cannot see dollars cannot protect a budget denominated in dollars, and a control evaluated after the run is a report. Step limits remain worth having as a cheap backstop, but they are not the primary control. The wider set of per-run controls is covered in how to stop runaway AI agent spend.
Common failure modes
- Reporting the mean. It hides the tail, which for agents is where the budget goes.
- No run identifier. A $25 run then looks like 300 unrelated small calls and nothing flags it.
- Optimising output for agent workloads. The lever is on the input side.
- Trusting the model's own estimate of what a run will cost. Measured correlation with actual usage was 0.39.
- Assuming an expensive run was at least a good one. Accuracy peaked at intermediate cost in the Stanford data and saturated above it.
- Per-user caps standing in for per-run caps. They bound the month, not the incident.
Where this leads
Once every run is a priced, attributed record, two things follow that are not otherwise available. Runs can be stopped at a dollar ceiling before the next call is forwarded rather than reported on afterwards. And the cost of a run can be rolled up to the customer it was for, which is the number that tells you whether an account is worth having. That rollup is the subject of tracking LLM costs per customer, and what it does to margin is worked through in AI gross margin by customer.
FAQ
How much does one AI agent run cost? There is no single figure. Runs on the same task varied by up to 30x in the Stanford data. The worked twelve-step run here is $0.85 uncached and $0.47 cached, and the spread around it matters more than either number.
Why are agent runs so much more expensive than chat? A run is many calls, each replaying the context so far. Anthropic reports roughly 4x chat tokens for agents and 15x for multi-agent systems; Stanford puts agentic coding tasks near 1000x ordinary code chat.
Is output or input the expensive part of an agent run? Input, which reverses the usual advice. Context replay makes input volume grow with run length, and the Stanford study found input driving overall cost.
Does spending more tokens produce better agent results? Not reliably. Accuracy peaked at intermediate cost and saturated above it, so the priciest run of a task is often not the successful one.
How do you cap the cost of an agent run? Attach a server-set run identifier to every call, accumulate spend against it, and check the total before each call is forwarded. Step limits are a backstop because they count steps rather than dollars.
Sources and method
Written from building and operating Spendline's AI spend governance proxy, with public primary sources checked on 29 September 2026: token multiples and the 80% variance finding from Anthropic's multi-agent engineering post; the 30x same-task variance, the 1000x comparison, the input-dominance finding, the accuracy saturation result and the 0.39 self-prediction correlation from the Stanford Digital Economy Lab study (Bai et al., 14 April 2026, eight frontier models on SWE-bench Verified); the Uber budget and the $1,500 cap from TechCrunch, 2 June 2026; model rates from Anthropic's published price list. Prices change without notice, so check them before relying on a figure. The worked example is our own arithmetic on stated assumptions, and the control comparison is our categorisation judgment. Last updated: September 2026.
Want to know what one run of your agent costs for your heaviest customer? Book a free 30-minute call. No integration is needed: we go through where your AI cost lands by customer and workflow, and how you would find the accounts that cost more than they pay. If the call turns up a real gap, we offer a free 60-day pilot on your own traffic. Book the 30-minute call
Not ready to talk? Take the 5-minute assessment instead.