gemini4.ai

Cost

Gemini 4 API pricing

No price exists, because no model exists. What you can do today is budget against the model Gemini 4 would replace, and understand which of the leaked specifications would actually move your bill.

Gemini 4 published price

None

No model ID, no pricing entry

Circulating figure

$2.25 / 1M

Unsourced. Treat as noise.

Odds input ≤ $3 / 1M

55%

Forecast, not a quote

Budget against

$2 / $12

Gemini 3.1 Pro standard tier

The only real numbers: Gemini 3.1 Pro today

Gemini 3.1 Pro has been the Pro-tier flagship since 19 February 2026, and its rates are the correct baseline for any Gemini 4 planning. All figures are per million tokens on the standard paid tier of the Gemini Developer API.

TierInputOutput
Standard (≤ 200k prompt)$2.00$12.00
Long context (> 200k prompt)$4.00$18.00
Cached input$0.20
Batch (50% off standard)$1.00$6.00

Rates as published for September 2026; verify against theofficial Gemini API pricing pagebefore committing a budget. Vertex AI and enterprise agreements price separately.

About the $2.25 figure

A price of $2.25 per million input tokens appears across aggregator articles about Gemini 4. It has no traceable origin: no Google document, no leaked screenshot, no named source. Its proximity to the current $2.00 standard rate suggests it was constructed rather than discovered — a plausible-looking increment that then propagated by citation.

Forecasters give roughly 55% odds that standard text input launches at $3.00 per million tokens or less. That is a wide band stated honestly, and it is more informative than a precise number with nothing behind it.

What the leaked specs would do to your bill

The 256k output ceiling matters most

Output tokens cost six times input on the current Pro tier. A model that can emit 256,000 tokens in one response makes it far easier to accidentally emit a very expensive one. If the leaked ceiling holds, output budgeting — not context budgeting — is the thing to instrument before you migrate.

A multi-million context window mostly changes the threshold

Gemini already charges a long-context premium above 200k tokens: double on input, 50% more on output. A larger maximum window does not remove that structure, it extends the range over which the premium applies. Watch where the threshold lands, not how big the window is.

High-effort compute is likely a separate price

Inference-time compute costs real money to serve, and the industry convention is to expose it as a distinct tier or a per-request parameter with its own rate. Expect a second number rather than a bump to the base rate — and expect it to be the expensive one.


Frequently asked questions

How much does the Gemini 4 API cost?

Nothing is published. There is no Gemini 4 model ID, so there is no price, no quota, no region list and no service tier. Any pricing page listing Gemini 4 today is publishing a guess.

Where does the $2.25 per million tokens figure come from?

It circulates across aggregator sites with no traceable origin — no Google source, no leaked document, no named insider. It sits slightly above the current Gemini 3.1 Pro standard input rate of $2.00, which is probably where it was reverse-engineered from.

What does Gemini 3.1 Pro cost right now?

On the standard paid tier: $2.00 per million input tokens and $12.00 per million output tokens for prompts up to 200,000 tokens, rising to $4.00 and $18.00 above that threshold. Cached input is $0.20 per million, and batch processing runs at half the standard rate.

Will Gemini 4 be more expensive than Gemini 3.1 Pro?

Unknown, and the leaked specs cut both ways. A larger context window and a 256k output ceiling push costs up, while every Gemini generation so far has held or lowered the headline Pro rate as serving efficiency improved. A reported high-effort compute mode would most likely be priced as a separate, higher tier rather than folded into the base rate.

Can I budget for Gemini 4 today?

Use the current Gemini 3.1 Pro rates as your planning baseline and treat anything beyond that as scenario work. If your workload is output-heavy, note that output tokens cost six times input on the current Pro tier — a larger output ceiling changes your bill far more than a larger context window does.

Reviewed 20 September 2026. Related: leaked specifications,release-date tracking.