Cheapest LLM API: Why the Lowest Token Price Rarely Wins
The cheapest LLM API is rarely the cheapest system. Compare token price with retries, routing, failure cost, and the deterministic app layer that makes spend predictable.

The short answer
The cheapest LLM API is the model endpoint that produces an accepted result for your specific task at the lowest all-in cost per success. That cost covers input and output tokens, failed attempts that get rerun, and the human minutes spent fixing outputs that returned a clean 200 but still got the task wrong. Token price is one input to that number.
The formula:
Cost per successful task = (token cost per attempt ÷ success rate) + scoped overhead
Scope the overhead or it swallows everything. For most teams it means human review and correction minutes per task at a loaded hourly rate. One-time engineering setup goes on a separate line.
What API pricing actually measures
Providers bill per token, split into input and output. Input tokens are everything you send, including the system prompt, retrieved documents, history, and tool definitions. Output tokens are what the model generates. Output is priced higher almost everywhere, because each generated token needs its own pass through the model while the prompt is processed in parallel. CostGoat's comparison explains the same asymmetry.
Live comparators such as CostGoat and Price Per Token list input price, output price, cached-input price, and context window for hundreds of models. Prices move, so use these comparators to understand which fields to compare, then confirm any number you budget against on the provider's official pricing page on the day you commit.
Price tables measure the cost of a token. You pay for tasks, and the ratio between the two depends on prompt shape. An extraction job that sends a 6,000-token contract and gets back a 200-token JSON object is input-heavy, so input price dominates. A drafting job that turns a 300-token brief into a 2,000-token memo is output-heavy. The same two models can swap rank on the same price table depending on which job you run.
Why per-token price is not task cost
Take an invoice-extraction workflow. The numbers below are illustrative, chosen to show the arithmetic, and do not describe any provider's pricing.
Model A costs $0.002 per attempt. Model B costs $0.010. A schema validator catches malformed output and triggers a retry. Model A passes on the first try 70% of the time and Model B passes 97% of the time. Per 1,000 invoices, token cost per success works out to about $2.86 for A and $10.31 for B. Model A wins by about seven dollars.
Then the validator's blind spot shows up. It catches broken JSON, but it can't catch a well-formed invoice total that is simply wrong. Suppose 8% of A's accepted outputs carry a wrong value against 1% of B's, and each one costs a finance reviewer four minutes to find and fix. That is 320 reviewer minutes for A and 40 for B. At a $60 loaded hourly rate, A's overhead is $320 and B's is $40. The "expensive" model is cheaper by a wide margin.
Each retry also adds a full round trip of latency. Verbose models spend more output tokens per answer, so a lower rate can still produce a larger bill. A model that needs a longer prompt to behave pays for that prompt on every call.
Free, budget, mid-tier, and frontier workload fits
| Workload | Model tier | Retry risk | What to measure | |---|---|---|---| | Ticket classification into fixed labels | Budget | Low, output is constrained | Accuracy on a labeled set of 200 tickets | | Structured extraction from documents | Budget to mid-tier | Medium, valid but wrong values | Field-level accuracy against a gold set | | Internal digest summaries | Budget | Low to medium | Output length and a sampled human rating | | Prototypes, evals, personal tools | Free | High, rate limits and outages | Throughput under the provider's limits | | Customer-facing drafts | Mid-tier | Medium, tone and fact errors | Approval rate and edit distance | | Contract review, code changes, multi-step reasoning | Frontier | High cost per failure | Acceptance-test pass rate and review minutes | | Agent tool calls across several systems | Mid-tier to frontier | High, errors compound per step | End-to-end task completion rate |
Free tiers deserve their own warning. They throttle requests per minute and per day, and the limits vary by provider and change without notice. OpenRouter's free API comparison puts the math plainly: 20 requests per minute means one request every three seconds, and 1,000 per day is roughly 40 an hour. It also notes that some free endpoints shrink context windows or serve lower-precision weights, and some providers train on free-tier prompts. Use free tiers for evals. Keep them away from anything customer-facing or confidential.
Model routing and measurement
The cheapest capable model should handle each task. "Capable" has a strict meaning here: it passes your acceptance test. A model that fails the test is expensive at any price.
LLM routing applies that rule per request, sending easy queries to small models and escalating hard ones. Routing lowers the price of each call. It does nothing about how many calls you make.
To pick on evidence rather than a price column:
- Build an acceptance test for each workload. Collect 100 to 300 real examples with known correct outputs and a pass/fail rule a script can check.
- Run every candidate model across the set and log input tokens, output tokens, retries, pass rate, and latency for each example.
- Compute cost per successful task with scoped overhead, choose the cheapest model that clears the bar, and rerun the test whenever a provider changes a price or a model version.
How deterministic apps flatten repeatable spend
An agent that re-reasons a workflow on every run pays for that reasoning every time. It re-parses the same input format and re-derives a formatting rule it worked out last Tuesday. A cheaper model makes each call cheaper. The calls keep happening.
Most steps in a repeatable workflow don't need a model. Fetching a record. Validating a schema. Writing a row. Once an agent has figured out those steps, they can run as code. Code returns the same output for the same input, never needs a probabilistic retry, and spends zero tokens per run. The Google Sheets automation pattern shows this on a single connector, and AI agents for project management shows it across a multi-step operational workflow.
The model still does the judgment work: the ambiguous line item, the exception, the call a rule can't encode. It just stops re-deriving the plumbing. Cost becomes front-loaded, because the agent spends tokens figuring the workflow out once, then flatter, because later runs execute code and only call the model where judgment is needed.
The Major take
Price tables rank the visible unit, the token. Actual spend is set by how often a workflow reasons from scratch and how often that reasoning fails. No provider switch fixes the second part.
Major is the enterprise platform where agents build the software they run on. When a Major agent works out a repeatable workflow, it builds an app for it, with a managed database, permissions, and audit logs, and runs that app on later executions. Take vendor invoices. A new invoice email arrives as the trigger. The app parses it, writes an invoice record to its database, and checks the total against the stored contract terms in code. Only when a line item doesn't match a known rule does the agent call the model to decide whether to approve, hold, or escalate. That decision lands as an audit row with the input, the model's reasoning, and the outcome. Extraction you would have paid a model to repeat 1,000 times a month runs as code, and model spend concentrates on the real exceptions. Reason once, run forever.
If you are measuring cost per successful task on an invoice, ticket, or extraction workflow right now, the next step is moving its repeatable steps out of the prompt and into an app with its own state and audit trail. Build your first cost-predictable workflow app on Major.
Related articles
- LLM Routing: How It Works and What the Benchmarks Show
- LLM Cost Management: Control Spend by Running Less AI
- AI Model Selection: Choose Per Task, Then Govern the Workflow
Related articles
Frequently asked questions
- What is the cheapest LLM API right now?
- It changes month to month, so check a live comparator such as CostGoat or Price Per Token and confirm on the provider's official pricing page before you commit. The lowest per-token rate is only a starting point. The cheapest API for you is the one with the lowest cost per successful task on your own workload, counting retries and human correction.
- Are free LLM APIs safe for production?
- Rarely. Free tiers cap requests per minute and per day, can tighten those limits without notice, and come with no SLA or compensation for outages. Some shrink context windows, and some providers use free-tier prompts for training. They work for prototypes and evals. Customer-facing or confidential workloads belong on a paid tier with a clear data policy.
- Is a cheaper model always lower quality?
- No. Quality only matters relative to the task. A small model can classify tickets into fixed labels as well as a frontier model, at a fraction of the cost. On multi-step reasoning, the same small model may fail often enough that retries and human fixes make it the more expensive choice per success.
- How do I cut LLM costs without switching models?
- Trim prompts and retrieved context, cap output length, and use cached-input pricing for stable system prompts. Route easy requests to smaller models. The biggest structural cut comes from moving repeatable workflow steps into deterministic apps, so the model stops re-reasoning the same plumbing on every run and only handles the judgment calls.