AI Agent Cost: What It Takes to Build One and What It Costs to Run

Every guide to AI agent cost quotes a build range between $10,000 and $500,000 and cites nothing. The number that actually matters is what one run costs and how that scales. Here is the arithmetic, from published rate cards.

Rahul Ramakrishnan

Key takeaways

  • Every build-cost range published on this topic, including the ones reported below, is asserted by the vendor publishing it, with no methodology, no sample, and no linked source.
  • The number that decides whether an agent is viable is cost per run at your expected volume, and almost nobody publishes the arithmetic for it.
  • A six-step agent on Claude Sonnet 5 at published rates works out to roughly $0.15 per run once retries are included, which is $1,450 a month at 10,000 runs.
  • Token spend is only the floor. InfoWorld's David Linthicum puts all-in operating cost at two to five times raw token cost, higher under regulation.
  • Run cost tracks how much reasoning happens per execution, so an agent that re-derives the same steps every time has a bill that grows in proportion to how well it works.

The two costs people confuse

Build cost is one-off. Engineering time, integration work, guardrails, evaluation, and deployment plumbing. You pay it once, then again in smaller increments as you maintain it.

Run cost is recurring and mostly tokens. Cost per run multiplied by runs per month. This is the number that scales with adoption, and the one that gets left out of the quote.

Nearly every cost guide on this subject names the distinction in a sentence and then spends the rest of the page on the build number. That is backwards. The build number is a project you can cancel. The run number is a liability that compounds.

What it costs to build an AI agent

Here are the ranges the market publishes. Read the third column before you use any of them.

| Scope | Published range | Sourced or asserted | |---|---|---| | Single-purpose agent, one system | $10,000 to $30,000 | Asserted. Codessavvy gives $15,000 to $30,000 as its own fixed price; Branchnode gives $10,000 to $30,000. No methodology. | | Multi-agent or multi-system workflow | $30,000 to $100,000 | Asserted. Codessavvy $30,000 to $80,000, Branchnode $30,000 to $100,000. No sample, no rate survey. | | Regulated or heavily integrated build | $100,000 to $500,000 | Asserted. Branchnode gives $100,000 to $250,000 for regulated work, says complex builds cross $500,000. Nothing linked. | | Contractor and agency labour | $25 to $300 per hour | Asserted. Branchnode lists US agencies at $125 to $175, nearshore $40 to $90, Southeast Asia $25 to $50. No survey named. | | In-house team, fully loaded | $400,000 to $1.2M per year | Asserted, and the only figure you can check against your own salary bands. |

Those are the market's claims about itself, reported as claims. I read the pages ranking for this query and none publishes a methodology, a rate survey, or a sample. Branchnode's page carries roughly 38 dollar amounts across about 29 distinct values and not one outbound citation in the body; its single analyst reference, a 2025 Gartner estimate that 60 percent of AI projects would be abandoned over data readiness, sits in plain text with no link to the report.

None of that makes the ranges wrong. It makes them unfalsifiable, which is a worse problem. If you have been quoted $65,000, no page on the internet will tell you whether that is reasonable, because none of them shows how any of these numbers were derived.

What actually drives the build number

The model call is the cheap part of a build. What costs money:

  • Integrations. Every system of record the agent touches needs auth, error handling, rate-limit behaviour, and a story for what happens when a write half-succeeds.
  • Guardrails and permissions. Scoping what the agent may do, under whose credentials, and where it must stop and ask a person.
  • Evaluation. A test set, a way to run the agent against it, and a defensible answer to "how do you know it works."
  • The loop itself. How agentic workflows are structured is most of the design work, because step count drives both reliability and cost.

Two of those four exist only because the thing is an agent. That is why an agent costs more than a script.

What it costs to run an AI agent

Where the tokens actually go

An agent is a loop, and the loop is what makes it expensive. Four mechanisms account for most of the bill.

Re-sent context. Each step sends the whole conversation so far back to the model. System prompt, tool definitions, every prior response, every tool result. A six-step agent does not send its context once. It sends a growing version of it six times.

Reasoning tokens. Internal reasoning is billed at the output rate, five times the input rate on Claude Sonnet 5 and six times on Gemini 3.1 Pro. Google states plainly that output prices include thinking tokens. Reasoning is usually the most expensive line in a run.

Tool results. Anthropic's own figures put an average 10 kB web page at about 2,500 tokens and a 100 kB documentation page at about 25,000. Each one lands in context and stays there for the rest of the run.

Retries and failures. A run that fails at step five still bills for steps one through five. Leaving failed runs out understates the bill by a tenth or more.

Working out your cost per run

Take a collections agent handling one overdue invoice. Pull the invoice, check the contract terms, reconcile against payments received, draft a collections note, write the result back to the system of record. Six model steps.

Assumptions, stated so you can replace them:

  • System prompt plus tool definitions: 2,000 tokens, re-sent on every step.
  • Task input: 500 tokens.
  • Tool results: between 400 and 1,500 tokens per step.
  • Output including reasoning: 5,200 tokens across the run.
  • Model: Claude Sonnet 5 at $2 per million input tokens and $10 per million output, per the published rate card.

Input accumulates as the loop runs, because each step carries everything before it:

| Step | Input tokens sent | |---|---| | 1 | 2,500 | | 2 | 4,000 | | 3 | 5,800 | | 4 | 7,100 | | 5 | 8,200 | | 6 | 9,200 | | Total | 36,800 |

The arithmetic:

  • Input: 36,800 × $2 / 1,000,000 = $0.0736
  • Output: 5,200 × $10 / 1,000,000 = $0.052
  • Successful run: $0.126
  • With a 15 percent retry rate, billed in full: $0.126 × 1.15 = $0.145

About fifteen cents a run. Note what the loop did to the input side. The agent's context is roughly 9,200 tokens at its largest and it billed for 36,800, because four times over it paid again for context it had already sent.

Swap the model and the number moves. The same profile on Claude Haiku 4.5 at $1 and $5 per million is $0.072 a run. On Claude Opus 5 at $5 and $25 it is $0.361. That spread is why routing to cheaper models per step is the first lever worth pulling, along with prompt caching, where a cache read is priced at 0.1x base input.

One caveat. Token usage on the same task varies substantially between runs, because step count, tool-result size, and reasoning length all move. Fifteen cents is a central estimate with a real spread around it, not a fixed unit price. Measure your own distribution before committing to a forecast.

What happens at 10x volume

| Runs per month | Estimated monthly run cost | Notes | |---|---|---| | 500 | $73 | Pilot. A rounding error that tells you nothing about viability. | | 2,000 | $290 | Still trivially affordable. Where most forecasts stop. | | 10,000 | $1,450 | Production volume. Worth a line in the budget. | | 50,000 | $7,250 | Non-token costs start to dominate and need their own model. | | 200,000 | $29,000 | Any per-run inefficiency is a five-figure monthly problem. |

Token cost is linear in volume, which sounds benign and is the trap. A pilot at 500 runs a month costs $73 and proves nothing about economics. The same agent succeeding at 200,000 runs costs $29,000 a month in tokens alone.

Build cost versus run cost: where they cross

Take a $30,000 build, the middle of the asserted band, running 10,000 times a month at $0.145.

| Line | Amount | |---|---| | Build, one-off | $30,000 | | Token cost, 12 months | $17,400 | | All-in run cost at 2x tokens, 12 months | $34,800 | | All-in run cost at 5x tokens, 12 months | $87,000 | | First-year total | $64,800 to $117,000 |

The two-to-five multiplier is not ours. It comes from David Linthicum's analysis in InfoWorld, which puts all-in operating cost at two to five times raw token spend once orchestration, vector stores, observability, evaluation, and security controls are counted, and higher under regulation. His line is worth keeping: "The model call may be inexpensive, but the system around the model is not."

Cumulative run cost overtakes the build between month four and month ten depending on where you land in that multiplier, or around month twenty-one on tokens alone. On any of those paths, the quote you spent three weeks negotiating is the smaller half of year one.

The costs nobody quotes

Four things sit inside Linthicum's multiplier.

Evaluation runs. Every regression test executes the agent and bills at production rates. A 200-case eval suite on the agent above costs $29 a pass, which is fine weekly and expensive in CI on every commit.

Human review. Early on, someone checks a meaningful share of outputs. At 10,000 runs a month and two minutes of review on a tenth of them, that is 33 hours of someone's month.

Observability. Attributing spend to the step that caused it needs instrumentation that is usually bought rather than built. Attributing cost to what the agent did is the difference between a bill you can act on and a bill you can only pay.

Rework. Prompt and tool changes invalidate your evaluation baseline, so you re-run it. This one surprises people most, because it appears in no month's forecast and in every month's invoice.

How to estimate this for your own case

  1. Write down the agent's steps. Count model calls, not tasks. Step count is the primary driver of per-run cost.
  2. Measure context growth per step, including tool results. Assume the largest realistic tool output, not the average.
  3. Estimate output tokens per step including reasoning, and price them at the output rate. Do not price reasoning as input.
  4. Compute cost per run from a provider's published rate card, step by step, as above. Use OpenAI's, Anthropic's, or Google's directly.
  5. Add your observed retry and failure rate, billed in full. If you have not measured it, use 15 percent and flag it as a guess.
  6. Multiply by expected monthly volume at launch and at ten times launch. Only the second number tells you anything.
  7. Apply a multiplier for everything around the model. Two is optimistic under regulation. Then decide which agent use cases pay back against that total, not against the token line.

Why agent costs grow, and what to do about it

Here is the structural problem underneath the arithmetic. Most agents are billed for thinking, and most of that thinking is repetition. The same collections workflow, reasoned through from scratch on Monday and reasoned through again on Tuesday, at full price both times. That produces a cost curve that rises with adoption, which is the wrong shape for anything you want to succeed. The better the agent works, the more it runs, and the more it costs. No amount of model-shopping fixes a cost structure that charges you again for a decision you already made.

Major's answer is to change what gets billed. When an agent on Major works out how to handle a repeatable part of a task, it builds an app for that part, deterministic code with its own managed database, storage, and logs, and from then on it runs the app rather than reasoning through the step again. The model stays for the judgment calls. Reason once, run forever.

The consequence for the bill is structural rather than a discount. Spend front-loads into the runs where the agent is still working the problem out, then flattens as the app library grows, because the repeated work has stopped passing through the model. Because the apps are reusable across the organisation, later work starts from a lower baseline than earlier work did. It is the same move that makes the work auditable, since an app has permissions and logs where a re-reasoned prompt has neither, and using fewer model calls is how both properties arrive together.

This does not make agent cost free or fully predictable. The judgment steps still cost tokens and still vary run to run, the build still has to happen, and the honest version of the claim is that repeatable work moves out of a metered, probabilistic layer into one where it runs the same way every time. It also does not help you if your workflow is simple enough to live in one prompt. If it is long, branches, touches four systems of record, and someone will eventually ask what it did in March, the shape of the cost curve is the whole decision, and reducing what each run costs is a smaller lever than reducing how many runs need reasoning at all.

If you have just calculated your cost per run and did not like the slope, look at which of your agent's steps never needed the model twice. That is the part Major turns into an app, and you can see how agents on Major build the apps they then run.

Related articles

Frequently asked questions

How much does it cost to build an AI agent?
Published ranges run from about $10,000 for a single-purpose agent to $500,000 for a regulated, heavily integrated build. Treat those figures as asserted rather than measured: the vendor pages publishing them state no methodology, no sample, and no rate survey. The variable that decides viability is cost per run at your expected volume, which the build quote does not tell you.
What are the ongoing costs of running an AI agent?
Token spend on every step of the loop, including re-sent context and reasoning tokens billed at the output rate. Then tool and infrastructure costs, evaluation runs that bill at production rates, human review of outputs, and the observability needed to attribute spend. InfoWorld puts all-in operating cost at two to five times raw token spend, higher under regulation.
Why is an AI agent more expensive than just calling an API?
A single API call sends context once. An agent runs a loop, and each step re-sends the whole conversation so far, so context is billed repeatedly. Reasoning tokens are charged at the output rate, tool results accumulate in context, and a run that fails at step five still bills for the first five steps.
How much does an AI agent cost per run?
A six-step collections agent on Claude Sonnet 5, sending 36,800 accumulated input tokens and producing 5,200 output tokens including reasoning, costs $0.126 per successful run at the published rates of $2 and $10 per million tokens. With a 15 percent retry rate it is about $0.145. Token usage varies run to run, so treat that as a central estimate.
When should you not build an AI agent?
When the task is deterministic enough that code is cheaper and more reliable. If the steps are fixed and the decisions are rule-based, a script runs the same way every time at no token cost. Agents earn their cost where genuine judgment is required at some point in a long workflow crossing several systems.