Grok Pricing in 2026: Plans, API Costs, and Agent Spend

Every Grok pricing page lists the subscription tiers and the per-token rate. None of them explains the 200k context cliff that reprices your whole request, or what a recurring agent workflow actually costs on its thousandth run. Here is the rate card, and the part that comes afte

Jason Bao
grok-pricing-hero.png

The short answer

All figures on this page were retrieved from official xAI documentation on 16 September 2026. Two things to know before the tables.

First, Grok is sold twice. There is a consumer subscription you buy to talk to Grok in an app, and there is an API you pay for per token to build on. They are separate products with separate rate cards, and the subscription does not include API credit.

Second, the subscription rate card is harder to pin down than the API one. xAI's pricing page names more tiers than it prices, its team and API tabs render client-side, and the page rejects automated retrieval outright. The API rates, by contrast, are fully published in the docs, down to the cached-input rate per model. So this page is precise where xAI is precise and says "not published" where it is not. No figure here comes from a third-party pricing roundup, and several of those roundups are currently quoting rates for models xAI no longer serves.

Grok subscription plans

xAI's pricing page presents a comparison matrix across several named tiers but attaches a dollar figure to only some of them. The Individual tab is the only one that renders server-side; Team and API pricing load in the browser, so those numbers are not in the page source. On 16 September 2026, x.ai/pricing returned HTTP 403 to automated retrieval, so no subscription dollar figure below is asserted from a retrieved official source.

| Tier | Price | What it is for | | --- | --- | --- | | Free | $0, with usage limits | Casual use in the Grok app and on X. Rate-limited on messages and on the newest models. | | SuperGrok Lite | Not published in retrievable source | Entry paid individual tier. Appears in xAI's feature matrix without a price. | | SuperGrok | Not published in retrievable source | The mainstream paid individual plan. Higher limits and access to current flagship models. | | SuperGrok Heavy | Not published in retrievable source | Top individual tier, positioned around multi-agent and heavier reasoning workloads. | | Business / Team | Not published; tab loads client-side | Seat-based team plan with admin controls, team management, and compliance features. | | Enterprise | Quoted on request | Contracted deployment, security and compliance review, dedicated support. |

A figure circulating widely on AI-tools blogs puts SuperGrok Heavy at roughly $300 a month. Treat that as unverified. It is not in the retrievable official page, and repeating an unsourced price on a page whose job is accuracy would be a poor trade. If you need the exact number, read it from a logged-in view of x.ai/pricing, which will also show the current promotional state.

If you are still deciding whether Grok is the right model before you decide what it costs, how Grok compares against ChatGPT and Grok against Claude are the better starting points.

Grok API pricing per token

This is the rate card that matters if you are building. Every model carries three separate rates: input, cached input, and output. Each also carries two tiers, one for prompts under 200,000 tokens and one for prompts at or above it.

All rates per 1M tokens, USD. Retrieved from docs.x.ai/docs/models and docs.x.ai/docs/pricing on 16 September 2026.

| Model | Context window | Input | Cached input | Output | | --- | --- | --- | --- | --- | | grok-4.6 (prompt under 200k) | 500k | $2.00 | $0.50 | $6.00 | | grok-4.6 (prompt 200k and above) | 500k | $4.00 | $1.00 | $12.00 | | grok-4.5 (prompt under 200k) | 500k | $2.00 | $0.30 | $6.00 | | grok-4.5 (prompt 200k and above) | 500k | $4.00 | $0.60 | $12.00 | | grok-4.3 (prompt under 200k) | 1M | $1.25 | $0.20 | $2.50 | | grok-4.3 (prompt 200k and above) | 1M | $2.50 | $0.40 | $5.00 | | grok-4.20-0309-reasoning (under 200k) | 1M | $1.25 | $0.20 | $2.50 | | grok-4.20-0309-non-reasoning (under 200k) | 1M | $1.25 | $0.20 | $2.50 | | grok-4.20-multi-agent-0309 (under 200k) | 1M | $1.25 | $0.20 | $2.50 | | grok-build-0.1 (prompt under 200k) | 256k | $1.00 | $0.20 | $2.00 | | grok-build-0.1 (prompt 200k and above) | 256k | $2.00 | $0.40 | $4.00 |

The three grok-4.20 variants each double to $2.50 input, $0.40 cached, $5.00 output above the threshold, the same pattern as grok-4.3.

There is no free API tier. Credits are prepaid, and xAI's billing docs state that requests are rejected once prepaid credits deplete. A claim doing the rounds that xAI grants roughly $150 a month in free API credits through a data-sharing programme does not appear anywhere in current documentation. Treat it as unverified.

Image, video, and voice models are priced per image, per second, and per minute or character respectively, not per token, so they sit outside this table.

The 200k context cliff

This is the single most consequential rule on the rate card and no page ranking for Grok pricing mentions it.

The docs are unambiguous: "requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request." Every token, not the overage. A prompt of 199,000 tokens on grok-4.6 bills input at $2.00 per million. A prompt of 201,000 tokens bills all 201,000 at $4.00 per million, and the completion at $12.00 instead of $6.00.

For a chat product this is a rounding error, because almost no human conversation gets there. For an agent it is a structural risk, because an agent that carries its own history forward walks toward the threshold on every run. The workload does not announce that it crossed. The rate just doubles, quietly, on a Tuesday, and the per-token price you budgeted against is no longer the price you are paying.

The costs that are not in the token table

A real bill and a rate-card estimate differ by three things.

Server-side tool invocations are billed per call, on top of tokens. Web Search and X Search are $5 per 1,000 calls, Code Execution $5 per 1,000, File Attachments $10 per 1,000, and Collections Search $2.50 per 1,000. xAI has noted that from 21 September 2026, X Search moves to $5 per 1,000 posts fetched plus $10 per 1,000 user profiles fetched. Image, video-understanding, and remote MCP tools are token-billed with no invocation fee.

The Batch API takes 20% off standard rates across all token types. Flat 20%, not a range, and it applies to four models only: grok-4.3 and the three grok-4.20 variants. Other models get no batch discount. A ranking page currently states the discount is "20 to 50%", which the official pricing page contradicts.

Priority processing bills at a 2x multiplier across all token types, with caching discounts applied first, and is charged only when the response confirms the priority service tier. The US regional endpoint adds a 1.1x multiplier, currently on grok-4.6 only.

One more thing that behaves like a cost even though it is not priced: rate limits are tiered by cumulative spend. Tier 0 starts at $0 and the ladder runs $50, $250, $1,000, and $5,000, with Enterprise on request. Tiers unlock automatically and never downgrade. A new account is capacity-limited until it has spent enough to not be.

What Grok costs behind a recurring agent

Here is a model, not a benchmark. Every input is an assumption, labelled, so you can substitute your own.

Assume a support-triage agent on grok-4.6 at standard rates, no batch, no priority. Per ticket: 12,000 input tokens, of which 9,000 are a stable prefix that hits the cached rate and 3,000 are the new ticket text, plus 800 output tokens and one Web Search call.

  1. Uncached input: 3,000 tokens at $2.00 per 1M = $0.0060
  2. Cached input: 9,000 tokens at $0.50 per 1M = $0.0045
  3. Output: 800 tokens at $6.00 per 1M = $0.0048
  4. Web Search: one call at $5 per 1,000 = $0.0050
  5. Total per ticket: about $0.0203

At 400 tickets a day that is roughly $8.12 a day, or about $244 over a 30-day month. Cheap. This is the number that gets into the business case.

Now run the same agent after someone lets conversation history accumulate. Assume the prompt has grown to 210,000 tokens, 180,000 of it cached, 800 output, same single search. The request is now above the threshold, so it bills at $4.00 input, $1.00 cached, $12.00 output. That comes to about $0.3146 per ticket, roughly $126 a day, about $3,775 a month.

Be precise about what caused that. Most of the jump is token volume: the prompt got 17 times bigger. The threshold accounts for a doubling on top. Priced at the under-200k rates, that same fat prompt would run about $0.16 a ticket instead of $0.31. Both effects are real and they compound, which is why a growing-context agent is the one workload where per-token rate genuinely stops predicting the bill.

The position worth stating plainly: per-token rate is the least important number in an agent's bill. The two figures that determine spend are how many times the same reasoning gets repeated and whether the prompt crosses the long-context threshold. Most teams optimise the rate card and ignore both. And comparing Grok's headline input rate against another provider's is close to meaningless without the cached-input rate, the tool fees, and the context tier sitting next to it.

The spend controls xAI actually ships

xAI has built four mechanisms for exactly this problem, and none of the pages ranking for Grok pricing mentions that any of them exist.

| Mechanism | What it does | Where it is documented | | --- | --- | --- | | Per-response cost reporting | Every inference response returns cost_in_usd_ticks in the usage object: the actual amount billed for that request, after discounts and inclusive of server-side tool fees. Ticks are integers at 1 USD = 10^10 ticks, so summing thousands of requests does not accumulate float error. | docs.x.ai/docs/cost-tracking | | Per-key ACL scoping | API keys are scoped by ACL to specific endpoints (api-key:endpoint:chat, api-key:endpoint:image) or single models (api-key:model:grok-4.6), and carry their own qps, qpm, and tpm ceilings. Requests are rejected when a key's limit trips. | docs.x.ai/docs/management-api-guide | | Hard billing ceiling | The invoiced billing limit defaults to $0, which means requests are automatically rejected once prepaid credits deplete. A stop, not an alert. Auto top-up is opt-in, with a configurable trigger balance, a minimum $5 top-up, a monthly cap, and a limit of five top-ups per 24 hours. | docs.x.ai/console/billing | | Usage attribution by key | Usage Explorer groups spend by API key, and filters by key, model, request IP, cluster, or token type. xAI's own worked example is comparing consumption across test and production keys. | docs.x.ai/console/usage |

Two of these deserve emphasis. cost_in_usd_ticks makes per-feature attribution a fact rather than an estimate, which is rare: most providers make you reconstruct it from token counts. Note one caveat, that the Vercel AI SDK does not currently surface the field in its response metadata, so use the OpenAI SDK or raw REST if you want it. And per-key ACL scoping is the credential story for a production agent, since a key that can only reach one model on one endpoint at a capped token rate is a much smaller blast radius than a team-wide key. That reasoning generalises well beyond xAI, and scoping what an agent can do covers the wider pattern. If you are assessing the governance side rather than the arithmetic, what enterprise-grade actually means is the better read.

How to bring the number down

The cached-input rate is the first lever, and it is the one most agent code accidentally forfeits. On grok-4.6 the cached rate is a quarter of the standard input rate, and xAI applies caching discounts before the priority multiplier. A recurring agent with a stable system prompt, a fixed policy document, and the variable part of the request at the end will hit that rate repeatedly. An agent that rebuilds its prompt in a different order each run, or interleaves fresh content into the middle of an otherwise stable block, will not. Same model, same task, four times the input bill.

The second lever is the model. Not everything needs the flagship: grok-4.3 runs at $1.25 input against grok-4.6's $2.00, and classification or extraction steps rarely need the top model. Routing cheaper models per task is a real saving and a well-understood pattern.

The third lever is larger than both and it is not a pricing decision. The largest line item in a mature agent's bill is usually work that stopped needing judgment weeks ago. A triage agent that has classified 40,000 tickets is not discovering anything new on ticket 40,001, and it is paying full freight to re-derive a conclusion it has reached thousands of times. Caching makes re-reading the prompt cheaper. It does not stop the model from doing the thinking again. The only way that cost goes to zero is for the settled part of the work to stop being a model call and become code. That argument generalises past any single provider, and what actually reduces LLM cost is where to follow it.

The Major take

You can now measure Grok spend to the request. xAI returns the exact billed cost on every response, caps it at a hard ceiling, and attributes it by key. That is genuinely good instrumentation and better than most providers ship.

What none of it changes is why the number is what it is. A recurring agent pays the model to re-derive the same conclusion on every run, and as its context grows it eventually crosses the long-context threshold and pays a higher rate on every token in the request. Instrumentation makes a climbing bill legible. It does not flatten it.

Major is the enterprise platform where agents build the software they run on. When one of its agents works out how to handle a repeatable part of a task, it builds an app for that part and runs the app instead of reasoning through it again, so that work is code on the next run and on the thousandth. State lives in the app's managed database rather than accumulating in a prompt, so context does not grow run over run and the workload does not drift into the higher-priced tier. The model still costs tokens for the judgment calls. Everything settled stops asking. Front-loaded, then flat, rather than climbing with usage. Reason once, run forever.

If the arithmetic above described your workload, the thing to build is the narrow one: take the triage step your agent has already settled and let it become an app with its own database, its own logs, and its own permissions, then keep Grok for the calls that still need a model. Get started on Major and build the app that replaces your agent's repeated reasoning.

Related articles

Frequently asked questions

How much is the Grok subscription?
Grok has a free tier with usage limits, paid individual plans under the SuperGrok name, a seat-based Business or Team plan, and Enterprise pricing quoted on request. As of 16 September 2026, xAI's pricing page does not publish a dollar figure for SuperGrok Lite, SuperGrok, or SuperGrok Heavy in its retrievable page source, and the Team and API tabs load client-side. Check a logged-in view for exact figures.
Is paid Grok worth it?
It splits by who is asking. For a consumer, the paid tiers buy higher message limits and access to current flagship models, so the value depends on how often you hit the free tier's caps against comparable assistants. For anyone building software, the subscription is the wrong product: API access is billed separately per token, and a subscription includes no API credit.
Is there a free version of Grok?
Yes for consumer use. Grok has a free tier in the app and on X, rate-limited on messages and on access to the newest models. There is no free API tier. API usage is billed per token against prepaid credits, and xAI's billing docs state requests are rejected once those credits deplete.
How much does the Grok API cost?
It is billed per token, with separate input, cached-input, and output rates for every model. As of 16 September 2026, grok-4.6 runs $2.00 per 1M input tokens, $0.50 cached, and $6.00 output for prompts under 200,000 tokens. Above that threshold the rates double to $4.00, $1.00, and $12.00, and the higher rate applies to every token in the request.
How do you control Grok API spending?
Four documented mechanisms. The invoiced billing limit defaults to $0, so requests are rejected once prepaid credits run out. API keys are scoped by ACL to specific models and endpoints with their own rate ceilings. Every response returns cost_in_usd_ticks, the exact billed cost. Usage Explorer attributes spend by key. The largest lever, though, is not spending the tokens at all.