LLM API Pricing in 2026: The Rate Card Is Not the Bill
Every LLM pricing table shows two columns: input and output. Your invoice has a dozen more, cached reads, cache writes, context tiers, batch discounts, service tiers, tokenizer density. Here is what the rate cards actually charge you for.

Key takeaways
- LLM API rates are quoted per million tokens, and output costs four to five times input on the same model.
- Two of the three major providers price the same model differently depending on how long your prompt is, so a single "input" column is factually wrong for most of the market.
- Cached input reads cost about a tenth of standard input, which makes cache hit rate a bigger cost variable than the choice between two similarly priced models.
- Batch processing is half price at all three providers, and it is the easiest discount to claim on any workload that does not need an answer in the same second.
- Reasoning tokens bill as output, so an agent that thinks longer pays at the most expensive rate on the card.
How LLM API pricing works
You are billed per token, in and out, quoted per million. Output is priced several times higher than input across every provider: Claude Sonnet 5 charges $2 per million input tokens and $10 per million output, and OpenAI's gpt-5.4 charges $2.50 and $15. The ratio is consistent enough that you can treat it as a rule. Generating a token requires a full forward pass through the model; processing an input token happens in parallel with all the others in the prompt. That asymmetry shows up directly in the price.
This is the part everyone already knows, and it is the part every comparison table shows. The rest of the bill comes from six or seven dimensions that most tables do not have a column for.
Current rates from the three major providers
Prices verified on 17 September 2026 against each provider's own pricing page. These change frequently. Check the linked page before you commit a forecast to anyone.
| Provider | Model | Input $/1M | Cached input $/1M | Output $/1M | Context tier | Verified | |---|---|---|---|---|---|---| | Anthropic | Claude Opus 5 | $5.00 | $0.50 | $25.00 | None; 1M window at standard rate | 2026-09-17 | | Anthropic | Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | None; 1M window at standard rate | 2026-09-17 | | Anthropic | Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | None published | 2026-09-17 | | OpenAI | gpt-6-astra | $10.00 | $1.00 | $50.00 | Long context: $20.00 / $2.00 / $75.00 | 2026-09-17 | | OpenAI | gpt-5.6-terra | $2.00 | $0.20 | $12.00 | Long context: $4.00 / $0.40 / $18.00 | 2026-09-17 | | OpenAI | gpt-5.4 | $2.50 | $0.25 | $15.00 | Above 272k: $5.00 / $0.50 / $22.50 | 2026-09-17 | | OpenAI | gpt-5-nano | $0.05 | $0.005 | $0.40 | None | 2026-09-17 | | Google | Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 | Above 200k: $4.00 / $0.40 / $18.00 | 2026-09-17 | | Google | Gemini 2.5 Pro | $1.25 | $0.125 | $10.00 | Above 200k: $2.50 / $0.25 / $15.00 | 2026-09-17 | | Google | Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | None | 2026-09-17 |
Read the context tier column before you read anything else. gpt-5.6-terra and Gemini 3.1 Pro Preview both list $2.00 input. On a 300k-token prompt, one of them is charging $4.00 and the other is still charging $2.00, because the boundaries sit in different places. Claude Sonnet 5 stays at $2.00 all the way to a million tokens. Three identical-looking rows, three different bills.
The dimensions the pricing tables leave out
| Dimension | How it is charged | Why it changes your bill | |---|---|---| | Input tokens | Per million, at the listed rate | The baseline everyone models | | Output tokens | Per million, 4x to 5x input | Dominates chat and drafting workloads | | Cached input reads | 0.1x base input at Anthropic and OpenAI; roughly 0.1x at Google | The single largest lever on repeated-context workloads | | Cache writes | 1.25x base input for a 5-minute cache, 2x for a 1-hour cache at Anthropic | A cache you never read from costs more than no cache | | Cache storage | Google charges per million tokens per hour, from $0.50 to $8.10 depending on model and tier | Rent on idle context, invisible in per-token math | | Context-length tier | OpenAI and Google roughly double input above a threshold; Anthropic does not tier | Long-context RAG can cost 2x the quoted rate | | Batch tier | 50% off input and output at all three | Free money on anything asynchronous | | Fast or priority tier | 2x standard at OpenAI and Anthropic; 1.8x at Google | Latency is a line item | | Reasoning tokens | Billed as output | Where agent bills actually land | | Server-side tools | Web search is $10 per 1,000 calls at both OpenAI and Anthropic; Google charges $14 per 1,000 after a free allowance on Gemini 3 models | Per-call fees stack on top of tokens | | Data residency | 1.1x multiplier at Anthropic for US-pinned inference; 10% uplift at OpenAI on models released after 5 March 2026 | A compliance decision with a price attached | | Tokenizer density | Not charged; changes how many tokens the same text becomes | Moves effective price without moving the rate card |
Cached input, and why it is the biggest lever
A cache read costs 10% of the standard input price at both Anthropic and OpenAI. On Claude Sonnet 5, that is $0.20 instead of $2.00. Any workload with a stable prefix, meaning a system prompt, a policy document, a schema, a tool manifest, gets most of its input bill cut by an order of magnitude. If you are choosing between two models forty cents apart on paper, cache hit rate will swamp the difference.
Cache writes are not free
Anthropic charges 1.25x base input to write a 5-minute cache and 2x for a one-hour cache. The break-even is easy arithmetic: at 1.25x write and 0.1x read, the 5-minute cache pays for itself after one read. The one-hour cache needs two. Bursty traffic that writes the cache and then goes quiet pays the premium and collects nothing. Google prices this differently again, renting storage by the hour, so a cached context sitting idle overnight accrues charges whether or not anyone queries it.
Context-length tiering
This is the most under-covered fact in this category. OpenAI's gpt-5.4 charges $2.50 per million input under 272k tokens and $5.00 above it. Google draws its line at 200k: Gemini 2.5 Pro is $1.25 below and $2.50 above. Anthropic states plainly that Claude 4.6 and later include the full 1M token context window at standard pricing, and that a 900k-token request bills at the same per-token rate as a 9k one.
So the correct answer to "what does this model cost per input token" depends on how long the prompt is. Every two-column table ranking for this keyword prints a single figure and calls it the price.
Batch and service tiers
All three providers discount batch by 50% on both input and output. If your workload tolerates minutes instead of seconds, that is the cheapest optimization available and it requires no model change. In the other direction, premium tiers exist for latency. OpenAI's Fast mode runs 2x standard. Anthropic's Fast mode prices Claude Opus 5 at $10 input and $50 output against a $5 and $25 standard rate. Google's Priority tier runs about 1.8x. These stack multiplicatively with cache multipliers and residency uplifts, which is how a bill ends up three times a forecast built on the base rate.
Reasoning tokens are billed as output
Extended thinking produces tokens you never see, and they bill at the output rate. Google states directly that Gemini Flash output pricing includes thinking tokens. This is where agent economics get decided: an agent that reasons for 4,000 tokens before emitting a 200-token answer is paying for 4,200 output tokens. Routing requests by cost matters here more than anywhere else, because reasoning-heavy calls hit the most expensive rate on the card.
Tokenizer density
Anthropic's pricing page notes that Claude 4.7 and later use a newer tokenizer producing roughly 30% more tokens for the same text than Claude Sonnet 4.6 and earlier. Nothing on the rate card changes. Your effective price per unit of English goes up by about a third. Comparing a $2.00 rate on one tokenizer to a $2.00 rate on another is not comparing like with like, and no aggregator table has a column for it. This is also why picking a model per task beats picking one on headline rate.
The extras
Web search runs $10 per 1,000 calls at OpenAI and Anthropic. OpenAI bills code interpreter containers per 20-minute session, from $0.03 to $1.92 depending on memory, and file search storage at $0.10 per GB per day. Anthropic gives 1,550 free container hours a month then charges $0.05 per hour. None of this appears in a per-token comparison, and for a tool-using agent it is not a rounding error.
What this costs on a real workload
A support triage agent handling 2,000 tickets a day. Each run sends a 12,000-token stable prefix (triage policy, product taxonomy, response templates) plus 3,000 tokens of ticket and customer history, and generates 700 tokens. Claude Sonnet 5 at the rates above, 60,000 runs a month.
Standard. Input: 15,000 × $2 / 1M = $0.030. Output: 700 × $10 / 1M = $0.007. Per run, $0.037. Monthly: $2,220.
With caching. Cache read: 12,000 × $0.20 / 1M = $0.0024. Uncached input: 3,000 × $2 / 1M = $0.006. Output: $0.007. Per run, $0.0154, plus roughly 720 cache writes a month at $0.03 each. Monthly: about $946.
Batch plus caching. Every figure halves. Per run, $0.0077. Monthly: about $473.
Same workload, same model, same output quality. $2,220 or $473 depending on two configuration choices. That spread is wider than the gap between most of the models in the table above. Configuration decides more of this bill than model selection does.
How to actually estimate your bill
- Count tokens on ten real requests with the provider's token counting endpoint. Do not estimate from character counts, and do not reuse a count taken on a different model's tokenizer.
- Split input into the stable prefix and the variable part. The ratio between them tells you what caching is worth before you build it.
- Check whether your typical prompt crosses a context tier boundary. If it sits near 200k or 272k, model both sides.
- Add reasoning tokens to your output estimate. Measure them; do not assume the visible answer is the whole output.
- Add per-call tool fees separately, multiplied by calls per run rather than runs per month.
- Multiply by real volume, including retries and failed runs. A 5% retry rate is 5% on the bill.
- Re-verify the rates on the provider page before you present the number. At least one scheduled change is already published: Google lists Gemini 3.8, 3.7 and 3.6 Flash doubling on 1 January 2027, from $0.75 to $1.50 input and $3.75 to $7.50 output.
This article does not cover self-hosting economics or fine-tuning costs. Both change the shape of the calculation enough to need their own treatment.
The variable that dominates the bill
Every lever on this page reduces the price of a model call. Caching, batching, tier selection, the levers that actually reduce spend and the orchestration layer where you implement them: all of it makes each call cheaper. None of it makes fewer calls.
Look at the triage workload again. 60,000 runs a month, and most of those tickets are four or five shapes the agent has already worked out how to handle. Optimized perfectly, the system still pays full freight every time it reasons its way to a conclusion it reached last Tuesday. Rate-card optimization has a floor and you hit it fast.
Major moves the floor. When an agent on Major works out how to handle a repeatable part of a task, it builds an app for that part: deterministic code with its own managed database, storage and audit log. From then on it runs the app instead of re-reasoning the step, so that work leaves the token bill rather than moving to a cheaper tier. Because the app holds its own state, the context that would otherwise be re-sent on every run lives in a database instead of a prompt. The model still handles the judgment calls, the escalations, the cases the app was not built for. What changes is the cost curve: front-loaded while the app library is being built, then flat as it grows, rather than climbing with usage. Reason once, run forever. That is a different shape from anything on a pricing page, and it is the part of the bill you control after routing and caching are done. If you are also weighing quality against cost, how the models compare on capability and coordinating models and the app layer are the next two questions.
Start with the workload you priced above. Point an agent at your triage queue, let it build the app for the parts that repeat, and watch which calls stop happening. Build your triage agent on Major and price the workload again in a month.
Related articles
Frequently asked questions
- How are LLM APIs priced?
- Per million tokens, counted separately for input and output, with output priced roughly four to five times input on the same model. Two further dimensions carry real weight. Cached input reads bill at about a tenth of the standard input rate, and OpenAI and Google both charge more for input once a prompt crosses a context-length threshold. Anthropic charges one rate across its full context window.
- Why is output more expensive than input?
- Input tokens are processed together in a single pass over the prompt. Output tokens are produced one at a time, each requiring its own pass through the model, so generation consumes far more compute per token than reading does. Providers price that difference directly, which is why output sits at four to five times the input rate across Anthropic, OpenAI and Google.
- What is the cheapest LLM API?
- The lowest per-token tier is not the cheapest per finished job. gpt-5-nano lists at $0.05 per million input and $0.40 output as of 17 September 2026, but a small model that needs two attempts, longer reasoning, or a retry on a malformed response can cost more per completed task than a mid-tier model that answers correctly the first time. Price the job, not the token.
- Does prompt caching actually save money?
- Yes, once you read the cache. Anthropic charges 1.25x the base input rate to write a five-minute cache and 2x for a one-hour cache, against a read price of 0.1x. That means the five-minute cache pays for itself after a single read and the one-hour cache after two. Bursty traffic that writes a cache and then goes quiet pays the write premium for nothing.
- Are LLM API prices going up or down?
- Both, depending on the model. Per-token rates on comparable capability have generally fallen as newer tiers launch, and Anthropic cancelled a scheduled increase on Claude Sonnet 5, keeping it at $2 and $10. Google has published the opposite: Gemini 3.8, 3.7 and 3.6 Flash double on 1 January 2027, from $0.75 to $1.50 input and $3.75 to $7.50 output.