LLM Cost Management: Control Spend by Running Less AI
LLM cost management starts with visibility, budgets, routing, and attribution. The durable fix is architectural: move repeatable work into deterministic apps so token demand grows more slowly as usage grows.

Key takeaways
- Metering tells you where the money went and reduces nothing. Measurement and reduction are separate projects with separate owners.
- Every model call should emit one event carrying tenant, workflow, model, token counts, retries, latency, and outcome, including failures.
- A real budget has four fields: soft threshold, hard stop, escalation path, named owner. Anything less is a chart.
- Caching, batching, and routing lower the price of a call. None of them remove the call.
- The structural lever is workload placement. Repeatable steps belong in deterministic code, and the model belongs on judgment.
The actual definition: what is LLM cost management?
LLM cost management is the practice of metering, attributing, budgeting, and reducing what a system spends on model inference and everything attached to it: input and output tokens, retries and re-runs, tool calls the model triggers, retrieval infrastructure, and the human review time spent catching what the model got wrong.
Infracost's glossary calls it "the systematic approach of controlling, monitoring, and optimizing expenses associated with large language model operations" (Infracost). Accurate, and it describes a finance function. It says nothing about where the work runs, which is the variable with the largest coefficient.
Write the spend out symbolically and the coefficient becomes visible:
cost ≈ calls × (1 + r) × (t_in × p_in + t_out × p_out) + tool_fees + retrieval_infra + review_time
where r is the retry and re-run rate, t_in and t_out are tokens per call, and p_in and p_out are the per-token prices. Most cost work targets p_in and p_out. Some targets t_in. Almost none targets calls, the only term that compounds with adoption.
Measurement and reduction get conflated constantly. Instrumenting every request is a prerequisite for reduction and is not itself a reduction. A team can ship perfect per-tenant cost attribution and spend exactly as much the following month.
Why this matters now
Per-token prices keep falling and per-token discounts keep improving. Anthropic bills cache reads at "0.1 times the base input tokens price" against a 5-minute default lifetime, with 5-minute writes at "1.25 times" and 1-hour writes at "2 times" the base input price (Anthropic prompt caching docs). OpenAI describes reused tokens as "discounted up to 90%" above a "1,024 tokens" minimum cacheable prefix (OpenAI prompt caching docs).
Worth taking, all of it. These discounts also apply per call. An agent that re-derives the same customer-tier classification on every run gets a cheaper re-derivation, then does it again tomorrow. Discounting a repeated operation leaves it repeated.
Agentic workloads changed the shape of the bill. One user request now fans out into planning steps, tool selections, retries on malformed output, and verification passes. The call count per unit of work went up, inside code most FinOps owners cannot see into. We cover the per-call levers and what each one treats in LLM cost optimization.
The cost controls that matter
Meter every request
One event per model call, emitted whether the call succeeded or not. Failed calls that consumed input tokens still cost money, and a collector that records only successes under-reports spend exactly when a system is misbehaving.
| Field | Purpose | | --- | --- | | event_id, timestamp | Dedup and time-series joins | | tenant_id | Per-customer attribution and isolation | | workflow_id, step_id | Locates spend inside the agent, per step | | model, provider | Price lookup and routing analysis | | input_tokens, output_tokens | The billable quantity | | cached_input_tokens | Separates discounted reads from writes | | attempt, retry_of | Makes the (1 + r) term measurable | | latency_ms | Cost-to-latency tradeoff on routing changes | | outcome, error_code | Distinguishes paid failures from free ones | | cost_usd, price_version | Reprices history when a price sheet changes | | prompt_hash | Finds repeated work without storing prompts | | actor, credential_id | Ties spend to the authorizing identity |
OneUptime's LLMOps write-up specifies a close variant of this schema and works as an implementation reference, with the caveat that it is vendor material and flags its own price tables as placeholders: "Replace these illustrative values with current provider price sheets" (OneUptime). Hardcoded prices are the most common reason a cost pipeline goes quietly wrong.
Two fields carry privacy weight. Store prompt_hash rather than prompt text, so repeated-work analysis does not put a second copy of customer data in your telemetry store. And scope the cost store by tenant_id the way you scope the application database. A shared analytics table is how multi-tenant isolation gets broken by the observability layer rather than the product. The decision-context side of this is a separate problem, covered in agent observability.
Attribute spend to owners
Attribution turns a number into a decision. Infracost lists the usual dimensions: project, department, user, application. For agents, add workflow and step, because "the support agent cost $40k last quarter" is not actionable and "the ticket-classification step cost $40k last quarter" is.
The FinOps Foundation notes that accurate forecasting depends on "the ability to fully categorize and allocate technology costs" (FinOps Framework). An unallocated pool of model spend cannot be forecast, defended, or cut. Name a human owner per workflow, not a team. A person who gets the alert.
Enforce budgets and hard stops
A budget with no enforcement action is a chart. Give every scope a graduated ladder. OneUptime models one as ALERT, THROTTLE, DOWNGRADE, BLOCK across thresholds at 50%, 80%, 95%, and 100%.
| Scope | Soft threshold | Hard stop | Escalation | Owner | | --- | --- | --- | --- | --- | | Per-run ceiling | 60% of expected tokens | Terminate, return partial state | Page the on-call | Workflow owner | | Per-tenant daily | 80% of daily allowance | Queue non-urgent work | Notify account owner | Support lead | | Per-team monthly | 80% of budget | Block non-production runs | Finance and engineering review | Team manager | | Per-agent task type | 3x the 30-day median | Refuse the task | Incident channel | Platform lead |
The per-run ceiling is the control most teams skip and need most. Runaway retry loops do their damage inside hours, and a monthly budget check will not catch a loop that burns a quarter's allowance overnight.
Downgrade needs a defined floor, because a downgrade path with no cheaper model left resolves to a block, and that block should be intentional. Every hard stop needs an exception path with an audit record, or the first production incident gets resolved by someone turning the limit off for good.
Route by task and quality
Routing sends cheap tasks to cheap models. Cast AI describes an "LLM proxy [that] intelligently selects the most optimal LLM model for user queries," framed as a fix for teams defaulting to oversized models across a catalog of "at least 20 different models" (Cast AI). That vendor page offers no savings figures, and any page that does should be read against your own traffic.
Routing helps when task difficulty varies widely and you can classify it cheaply. It stops helping in three cases. A classifier that is itself a model call adds a call. A downgrade that raises the retry rate lets the (1 + r) term eat the per-token gain. And a task that needs a specific model's tool-use behavior has no cheaper substitute at any price. More on the tradeoffs in model routing.
Cache and batch safely
Caching and batching are the highest-return per-call moves. Both have sharp edges.
Prompt caching pays off when a long stable prefix precedes a short variable suffix, which describes most agent system prompts. Order context accordingly: stable instructions and schemas first, volatile user input last. Watch the write multiplier. A 1-hour Anthropic write costs "2 times the base input tokens price," so a prefix written and never re-read costs more than skipping the cache.
Exact-match caches must key on model and sampling settings. A cache that ignores temperature returns one sampled output as though it were the distribution, which is a correctness bug wearing a cost-saving costume. Semantic caches answer a new question with an old answer based on embedding proximity, which can produce a fluent, confident, wrong response. Use them only where a wrong answer is cheap.
Batching trades latency for price. Anthropic's Message Batches API describes "reducing costs by 50%," with most batches "completing within 1 hour," a hard 24-hour expiry, and a ceiling of "100,000 Message requests or 256 MB in size" (Anthropic batch processing docs). Expired requests are not billed, so batching needs a re-submission path rather than a silent drop.
Account for failures honestly across all three. A discarded cache read, an expired batch, and a routed call retried at a higher tier all show up as spend with nothing to show for it. That category deserves its own line in the report.
What dashboards miss
A dashboard cannot fix an architecture that re-reasons the same task. That is the part the implementation-heavy guides leave out, and it is where the money is.
Take a collections workflow. The agent pulls an invoice, checks it against contract terms, decides whether the payment is late, drafts a note, updates the system of record, and escalates the disputed ones. Five of those six steps produce a stable output for a given input: lookups, comparisons, format changes, writes. One of them, the judgment about whether a disputed case should escalate, needs a model.
Run all six through the model and you pay for all six on every invoice, at whatever the current discounted rate happens to be. You also get a system with no durable state, because the intermediate results lived in a context window that ended when the run did. The state problem and the cost problem are the same problem. Work that is not stored gets redone, and redone work is billed again.
Here is the checklist we use to decide what leaves the model.
- Does this step produce the same output for the same input? If yes, it is code.
- Are the rules still changing week to week? If yes, keep it in the model until they settle.
- Does the step need judgment about ambiguous input, or only a lookup, a comparison, a transformation, or a write?
- Is the result paid for repeatedly because it was never persisted anywhere?
- Can you name the owner, the permissions, and the audit record for this step today? If not, moving it into code gets you all three.
Steps that leave the model stop appearing in the calls term, the only change to the cost identity that does not degrade as usage grows.
What this article does not cover: self-hosted GPU economics, fine-tuning and training spend, and evaluation cost. Evaluation is the harder problem, because the quality bar you enforce sets the retry rate that sets the bill, and it deserves separate treatment.
What we're building at Major in response
Cost governance that stops at metering and budgets is a reporting function rather than a control. It tells you accurately how much you spent re-deriving things you already knew.
Major is the enterprise platform where agents build the software they run on. When one of our agents works out how to handle a repeatable part of a task, it builds an app for that part and runs the app from then on. The app is deterministic code, holds its state in a managed database, keeps its files, and writes its own logs under scoped credentials with role-based access. The model stays in the loop for the judgment calls and stops being billed for the lookups. Reason once. Run forever.
The economic shape is front-loaded and then flatter. Building the app costs reasoning up front. Running it costs what code costs. We will not put a percentage on that, because the honest version of the claim is structural rather than numeric and depends on how much of your workflow is genuinely repeatable.
Two tradeoffs are real. An app built from a wrong understanding runs the wrong logic reliably, which is worse than an inconsistent model getting it right half the time, so the permissions, logs, and review path around the app matter as much as the app. And a workflow whose rules change weekly should stay in the model, because you will pay to rebuild the app more often than you save by running it. The case for compiling a step into software gets stronger as the rules settle and volume climbs, the same condition under which governance at the point of action starts to matter.
If you want to flatten the cost curve rather than re-price it, ask your platform which steps it can compile into deterministic code and what state those steps keep afterward. You can see how Major's agents build the apps they run on.
Related articles
Frequently asked questions
- What is LLM cost management?
- LLM cost management is the practice of metering, attributing, budgeting, and reducing what a system spends on model inference and its dependencies. That covers input and output tokens, retries and re-runs, tool calls the model triggers, retrieval infrastructure, and human review time. Metering and attribution make the spend visible. Reducing it means changing where the work runs.
- How do companies control LLM costs?
- They emit one cost event per model call, attribute spend to a named owner per workflow and step, and enforce graduated budgets that end in a hard stop. Then they lower the price of each call through prompt caching, batch APIs, and routing cheap tasks to cheap models. The larger lever is removing repeatable steps from the model entirely and running them as deterministic code.
- How do you set an LLM budget?
- Baseline 30 days of per-workflow token spend, then set a soft alert around 80 percent of that baseline and a hard stop at the ceiling. Add a per-run token ceiling, because retry loops burn a monthly allowance overnight. Name one person as owner rather than a team, and define an exception path that writes an audit record when someone raises a limit.
- Does model routing reduce LLM costs?
- Routing reduces cost when task difficulty varies widely and you can classify it cheaply. It stops helping when the classifier is itself a model call, when the cheaper model raises the retry rate enough to cancel the per-token saving, or when the task depends on a specific model tool-use behavior. Routing lowers the price of a call and does not remove the call.