Gemini API Pricing: Every Tier That Changes What You Pay
Google's pricing page runs 9,500 words across three dozen near-identical tables and answers every question except the one you came with. Here is the same rate card compressed: which model, which tier, what the 200k threshold does, and what changes on January 1.

Key takeaways
- Four variables set your Gemini bill: model, service tier, prompt length, and the surface you call from.
- Batch and Flex bill at exactly half of Standard on every model that offers them. Priority bills at 1.8 times Standard.
- Pro-class models charge more above a 200,000-token prompt. Flash and Flash-Lite are flat at any prompt length.
- Thinking tokens are billed inside the output price. Google's row label reads "Output price (including thinking tokens)."
- Gemini 3.8, 3.7, and 3.6 Flash double on January 1, 2027: $0.75/$3.75 becomes $1.50/$7.50 per million.
The four things that determine your Gemini bill
Google's Developer API pricing page runs to roughly 9,500 words because it prints one table per model and there are three dozen models. Read it as thirty-six rate cards and you will be reading it forever. Read it as one rate card with four inputs and you can price any model on the list, including the ones Google ships after this article is published.
The model
Three classes matter for most text workloads. Pro is the reasoning tier and carries the highest per-token rate along with the prompt-length threshold described below. Flash is the working tier, priced roughly one-third to one-half of Pro on output. Flash-Lite is the cheap classifier tier, and at $0.10 input and $0.40 output per million on Gemini 2.5 Flash-Lite it is an order of magnitude below Pro.
The class names are stable across generations, but the prices inside a class are not. Gemini 3.5 Flash costs $1.50 input and $9.00 output. Gemini 3.8 Flash, a newer model, costs $0.75 and $3.75. Newer does not mean more expensive here, and assuming it does is how teams end up paying twice for the same tier of capability. If you are picking between classes rather than generations, routing work to the right model is the decision underneath the price.
The service tier
This is the most under-used lever on the card, and the arithmetic is unusually clean. Across every model that offers them, Batch and Flex bill at exactly half the Standard input and output price. Gemini 2.5 Pro Standard input is $1.25 per million, and Batch and Flex are $0.625. Gemini 3.8 Flash Standard output is $3.75, and Batch and Flex are $1.875. The ratio holds on Flash-Lite too.
Priority runs the other direction at exactly 1.8 times Standard. Gemini 3.5 Flash-Lite input is $0.30 on Standard and $0.54 on Priority. So the spread between the cheapest and most expensive way to run the same model is 3.6x, with no change to the model or its weights. You are buying a place in the queue.
Take the position seriously: for most production workloads the tier decision matters more than the model decision. Downgrading a model to save money means accepting worse output and hoping the quality holds. Moving a non-interactive job from Standard to Batch halves the bill at identical output quality. That is the safer of the levers that actually move an LLM bill, and almost nobody pulls it first.
| Tier | Cost vs Standard | Latency trade-off | Use when | |---|---|---|---| | Standard | 1x | Synchronous, default responsiveness | A person or a live system is waiting on the response | | Batch | 0.5x | Asynchronous, results returned when the job completes | Overnight jobs, backfills, bulk classification, evaluation runs | | Flex | 0.5x | Served at reduced priority, variable latency | Background work that can absorb a slow response | | Priority | 1.8x | Faster and more consistent service | Latency directly affects revenue or a contractual SLA |
Batch and Flex are marked "Not available" on several models, including Gemini 3.8 Flash and Gemini 3.5 Flash-Lite on the Free tier. Check availability per model before you design a workload around the half-price rate.
The 200,000-token prompt threshold
Pro-class models are priced in two bands. Gemini 2.5 Pro on Standard bills $1.25 input and $10.00 output for prompts up to 200,000 tokens, and $2.50 input and $15.00 output above it. Gemini 3.1 Pro Preview goes from $2.00/$12.00 to $4.00/$18.00 at the same boundary. Input doubles. Output rises by half.
The threshold applies to the whole prompt, so a long system prompt plus a large retrieved context can cross it without any single document looking big. Flash and Flash-Lite models have no such band. They charge one rate regardless of prompt length, which is a reason to prefer them for long-context work even when Pro would handle the reasoning better.
Any single number quoted for a Pro model without stating which side of 200k it sits on is incomplete.
Where you call it from
The Developer API and Vertex on Google Cloud are separate surfaces with separate rate cards, and blending them is the most common factual error in Gemini pricing writeups. The Vertex pricing page adds two dimensions the Developer API page does not show.
First, a regional split. Prices are listed for the Global endpoint and for regional endpoints separately, with regional carrying roughly a 10% premium. Gemini 2.5 Pro Standard input up to 200K tokens shows $1.25 per million on Global and $1.375 on regional. If data residency requires a specific region, that constraint has a price attached to it.
Second, cached input is a separate column, priced at roughly a tenth of the normal input rate. On a workload that re-sends the same system prompt and reference material on every call, context caching is a larger lever than the tier choice. The Developer API page does not surface it the same way.
How much does the Gemini API cost? The current rates
Gemini is billed per million tokens. On the paid Developer API, Standard rates run from $0.10 input and $0.40 output per million on Gemini 2.5 Flash-Lite up to $2.00 input and $12.00 output on Gemini 3.1 Pro Preview, with Pro-class rates roughly doubling above a 200,000-token prompt. Batch and Flex cost half of Standard. Priority costs 1.8 times Standard.
Figures below checked on September 8, 2026 against https://ai.google.dev/gemini-api/docs/pricing, which was last updated September 4, 2026. All figures are Developer API, paid Standard tier, USD per 1M tokens. Halve them for Batch or Flex. Multiply by 1.8 for Priority.
| Model | Input per 1M | Output per 1M | 200k threshold applies? | Best for | |---|---|---|---|---| | Gemini 3.1 Pro Preview | $2.00 / $4.00 | $12.00 / $18.00 | Yes | Hardest reasoning, agent planning, complex code | | Gemini 2.5 Pro | $1.25 / $2.50 | $10.00 / $15.00 | Yes | Reasoning at a lower rate, and free on the Free tier | | Gemini 3.5 Flash | $1.50 | $9.00 | No | Quality-sensitive throughput work | | Gemini 3.8 Flash | $0.75 | $3.75 | No | General production workhorse, rising Jan 1 2027 | | Gemini 3 Flash Preview | $0.50 text, $1.00 audio | $3.00 | No | Multimodal throughput including audio input | | Gemini 3.5 Flash-Lite | $0.30 | $2.50 | No | High-volume classification and extraction | | Gemini 3.1 Flash-Lite | $0.25 text, $0.50 audio | $1.50 | No | Cheap multimodal tagging at scale | | Gemini 2.5 Flash-Lite | $0.10 text, $0.30 audio | $0.40 | No | The cheapest usable text tier on the card |
Where two figures appear in a Pro row, the first is for prompts up to 200,000 tokens and the second is above it. This table covers the text and multimodal models most teams choose between. Google also prices image generation, Veo video by the second, Lyria per song, and embeddings on their own scales, and those are outside the scope here. Price is one axis. Comparing models beyond price is the other half of the decision, and choosing a model per task usually beats standardising on one.
The free tier: what it covers and what it costs you
The free tier is real and it is broader than most summaries suggest. Gemini 2.5 Pro is listed as "Free of charge" on the Standard and Priority tiers. So are Gemini 3.8, 3.7, 3.6, and 3.5 Flash, and the Flash-Lite line. Batch and Flex are marked "Not available" on most free-tier rows, and the image, video, and Omni models are paid only. Any writeup claiming Pro models went paid-only is contradicted by Google's live page.
What the free tier costs you is data terms. Google's pricing tables carry a row labelled "Used to improve our products," and the value is Yes for the Free Tier and No for the Paid Tier. The billing documentation puts it the other way round: paid-service prompts and responses are covered by Google's paid-services data terms and are not used to improve Google products, and no equivalent assurance is given for free-tier traffic. For a prototype on synthetic data that is a fair trade. For anything touching customer records it decides the question before price does.
Rate limits are the other boundary. Gemini enforces requests per minute, input tokens per minute, and requests per day, applied per project rather than per API key, with daily quotas resetting at midnight Pacific. Linking a billing account moves you to Tier 1 more or less immediately. Tier 2 needs $100 paid and three days since the first payment. Tier 3 needs $1,000 and thirty days. Monthly caps rise with the tier, from $250 on Tier 1 to $20,000 and above on Tier 3.
What changes on January 1, 2027
Google's page lists Gemini 3.8 Flash, 3.7 Flash, and 3.6 Flash at $0.75 input and $3.75 output per million through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. The Batch and Flex rates follow, moving from $0.375/$1.875 to $0.75/$3.75. Priority goes from $1.35/$6.75 to $2.70/$13.50.
That is a doubling, published in advance, on the models most likely to be carrying a production workload. If you are budgeting a twelve-month plan on any of the three, the second half of the year costs twice the first half. Gemini 3.5 Flash at $1.50/$9.00 has no scheduled increase listed, which as of January will make it cheaper on input than the model that currently undercuts it. Worth re-checking the card in December rather than assuming today's ordering survives.
What this costs at your volume
Take a nightly job that classifies 40,000 support tickets. Assume 1,500 input tokens per ticket, including the system prompt and the ticket body, and 200 output tokens back. That is 60M input and 8M output tokens per run. Model: Gemini 3.8 Flash.
Standard: 60 × $0.75 = $45.00 input, plus 8 × $3.75 = $30.00 output. $75.00 per run, $2,250 across a 30-day month.
Batch: 60 × $0.375 = $22.50 input, plus 8 × $1.875 = $15.00 output. $37.50 per run, $1,125 per month.
The job runs overnight and nobody is waiting on it, so the Standard tier was buying responsiveness that had no value. Now compare the other reflex. Running the same workload on Gemini 3.5 Flash-Lite at Standard gives 60 × $0.30 = $18.00 plus 8 × $2.50 = $20.00, or $38.00 per run. Almost exactly the same saving as Batch, except you also downgraded the model and now have to prove classification accuracy held.
After January 1, 2027, the Standard run becomes $150.00 and the Batch run becomes $75.00. The tier lever keeps its value. It just applies to a bigger number. Assumptions here are round and yours will differ, but the ratios do not.
The Major take
Google's tier system is an admission worth reading closely. Defer a job and pay exactly half. Demand it now and pay 1.8 times. The spread between Batch and Priority is 3.6x on identical model output, which is Google saying plainly that when and how work runs is worth as much as what runs it.
Every tier on that card still assumes the model performs the work. Tier selection sets the price of a request and does nothing about the number of requests. A workflow that re-derives the same result on every run keeps buying the same answer at whatever tier it chose, and the bill still tracks usage. Batch at half price on a workload that triples is a line that goes up.
On Major, when an agent works out how to handle a repeatable part of a task, it builds an app for that part and runs the app instead of reasoning through the step again. That step leaves the rate card entirely, at every tier, because no request gets sent. The app carries its own database, storage, and logs, so the account history and reference context that used to be re-sent as input tokens on every run live in the app rather than being re-bought each time. That is the difference between coordinating models and the app layer and pointing more traffic at a cheaper row.
The model still reasons. It makes the judgment calls, handles the cases the app was not built for, and writes the apps in the first place. The economics are front-loaded and then flat: building the app spends real reasoning up front, and after that repeated execution stops scaling token spend with usage. Reason once, run forever.
Batch is the cheapest tier Google sells. The cheaper one is the request that was never sent.
If the nightly classification job above looked familiar and the curve looked wrong, the next move is building the part of it that never needs re-reasoning. On Major, an agent works out the classification rules once, builds a governed app that applies them with its own database and audit log, and calls Gemini only for the tickets that genuinely need judgment. Get started on Major and build the classification agent that stops re-buying the same answer.
Related articles
Frequently asked questions
- How much does the Gemini API cost?
- Gemini API pricing is per million tokens and depends on model, service tier, and prompt length. On the paid Developer API, Standard rates run from $0.10 input and $0.40 output per million on Gemini 2.5 Flash-Lite to $2.00 input and $12.00 output on Gemini 3.1 Pro Preview. Pro-class rates roughly double above 200,000-token prompts. Batch and Flex cost half of Standard. Checked September 8, 2026.
- How do you use the Gemini API for free?
- Create a project and API key in Google AI Studio without linking billing. The free tier covers Gemini 2.5 Pro, the Flash line, and Flash-Lite on Standard, with lower rate limits and Batch and Flex mostly unavailable. Google's pricing tables mark free-tier content as used to improve its products, while paid-tier content is not. That data distinction rules the free tier out for customer records.
- Are there rate limits on the Gemini API?
- Yes. Google enforces requests per minute, input tokens per minute, and requests per day, applied per project rather than per API key, with daily quotas resetting at midnight Pacific. Limits vary by model and rise with your billing tier. Linking a billing account reaches Tier 1 almost immediately. Tier 2 requires $100 paid and three days.
- Are thinking tokens billed separately?
- No. Thinking tokens are included in the output price. Google's own row label on the pricing page reads "Output price (including thinking tokens)." So a reasoning-heavy call costs more because it generates more output tokens, not because a separate line item exists for the reasoning itself.
- What is the difference between the Standard, Batch, Flex, and Priority tiers?
- They price the same model output at different service levels. Standard is the synchronous default. Batch and Flex both bill at exactly half of Standard, with results returned asynchronously or at reduced priority. Priority bills at 1.8 times Standard for faster, more consistent service. The spread from Batch to Priority is 3.6x on identical model quality, which makes tier choice the largest safe cost lever.