gemini4.ai

Performance

Gemini 4 benchmarks

There is no Gemini 4 benchmark. Not a low one, not a disputed one — none. What circulates instead are two unsourced comparative claims, and this page exists to state precisely how thin they are.

What is actually claimed

Gemini 4 performance claims, reviewed 20 September 2026
AttributeWhat is claimedConfidenceSource
Complex coding / SWE tasksReported to outpace Anthropic Fable 5.1.No scores, no benchmark name, no evaluation harness has been published.LEAKNokiaPowerUserCometAPI
Multi-step reasoningReported to hold an edge over OpenAI GPT-6 Astra.LEAKCometAPIGeeky Gadgets
Official benchmark reportNone. Google has published no Gemini 4 evaluation of any kind.CONFGoogle AI for Developers

Why these claims carry so little weight

A benchmark claim is only useful if you can locate four things: the benchmark, the score, the harness and the comparison model's version. The circulating Gemini 4 claims supply none of them.

On the Arena.ai sightings

Anonymous pre-release testing on public arenas is real and routine — labs stage unlabelled checkpoints to collect preference data before launch. Two caveats apply to reading anything into it.

First, attribution is guesswork: observers infer which lab an unlabelled model belongs to from formatting habits and refusal style. Second, arena scores measure human preference on short interactions, which correlates loosely with capability on long-horizon tasks. A strong arena showing is a signal about response style at least as much as about intelligence.

Where the first real numbers will appear

When Gemini 4 launches, verifiable performance data arrives in a predictable order:

  1. Google's announcement post — a first-party table. Vendor-selected benchmarks, vendor-selected baselines, but real numbers with a named methodology.
  2. The model card — evaluation details, safety results and the caveats the blog post omits.
  3. LMArena — within hours, once the model is routable.
  4. Independent harnesses — SWE-bench Verified, ARC-AGI, Terminal-Bench and similar, over the following days, run by third parties with published scaffolds.

This page will be rewritten against those numbers the moment they exist. Until then, the accurate answer to "how good is Gemini 4" is that nobody outside Google knows.

For a grounded reference point, the current Pro flagship — Gemini 3.1 Pro, February 2026 — is the model Gemini 4 would have to beat. Itsdocumented capabilities are here, and they are what any leaked comparison should be measured against.


Frequently asked questions

What are the Gemini 4 benchmark scores?

There are none. Google has published no Gemini 4 evaluation, and no independent lab has tested the model because no public endpoint exists. Any table of "Gemini 4 benchmark scores" circulating today is fabricated or extrapolated.

Does Gemini 4 beat Claude Fable 5.1 or GPT-6 Astra?

Leaked evaluations are reported to show an advantage in complex coding against Fable 5.1 and in multi-step reasoning against GPT-6 Astra. No scores, benchmark names or harness details accompany those reports, so the claims cannot be checked, reproduced or ranked.

Was Gemini 4 tested on Arena.ai?

Observers reported side-by-side comparisons on Arena.ai against an unlabelled model believed to be a Gemini 4 Pro checkpoint. Anonymous-model testing is a normal pre-release practice, but attribution is inferred from output style rather than confirmed, and arena results measure human preference rather than capability.

When will real Gemini 4 benchmarks appear?

Google publishes a benchmark table in its announcement post, typically on launch day. Independent numbers from LMArena, SWE-bench, ARC-AGI and similar follow within days of API access. Until an API model ID exists, nobody outside Google can run an evaluation.

Reviewed 20 September 2026. Confidence labels are defined on themethodology page. Confirmed rows are checkable against a Google source today; nothing else is.