Performance
Gemini 4 benchmarks
There is no Gemini 4 benchmark. Not a low one, not a disputed one — none. What circulates instead are two unsourced comparative claims, and this page exists to state precisely how thin they are.
What is actually claimed
| Attribute | What is claimed | Confidence | Source |
|---|---|---|---|
| Complex coding / SWE tasks | Reported to outpace Anthropic Fable 5.1.No scores, no benchmark name, no evaluation harness has been published. | LEAK | NokiaPowerUserCometAPI |
| Multi-step reasoning | Reported to hold an edge over OpenAI GPT-6 Astra. | LEAK | CometAPIGeeky Gadgets |
| Official benchmark report | None. Google has published no Gemini 4 evaluation of any kind. | CONF | Google AI for Developers |
Why these claims carry so little weight
A benchmark claim is only useful if you can locate four things: the benchmark, the score, the harness and the comparison model's version. The circulating Gemini 4 claims supply none of them.
- No benchmark named. "Complex coding tests" is a category, not a measurement. SWE-bench Verified, Terminal-Bench and LiveCodeBench produce very different orderings.
- No score. Without a number there is no margin, and a claim with no margin cannot be distinguished from noise.
- No harness. Agentic coding scores swing by double-digit percentages on scaffolding alone. The same weights can look state-of-the-art or mediocre.
- No baseline version. "Fable 5.1" and "GPT-6 Astra" both have configuration and effort settings that materially change results.
- An unreleased model. Even a perfectly measured checkpoint score need not describe the model that eventually ships.
On the Arena.ai sightings
Anonymous pre-release testing on public arenas is real and routine — labs stage unlabelled checkpoints to collect preference data before launch. Two caveats apply to reading anything into it.
First, attribution is guesswork: observers infer which lab an unlabelled model belongs to from formatting habits and refusal style. Second, arena scores measure human preference on short interactions, which correlates loosely with capability on long-horizon tasks. A strong arena showing is a signal about response style at least as much as about intelligence.
Where the first real numbers will appear
When Gemini 4 launches, verifiable performance data arrives in a predictable order:
- Google's announcement post — a first-party table. Vendor-selected benchmarks, vendor-selected baselines, but real numbers with a named methodology.
- The model card — evaluation details, safety results and the caveats the blog post omits.
- LMArena — within hours, once the model is routable.
- Independent harnesses — SWE-bench Verified, ARC-AGI, Terminal-Bench and similar, over the following days, run by third parties with published scaffolds.
This page will be rewritten against those numbers the moment they exist. Until then, the accurate answer to "how good is Gemini 4" is that nobody outside Google knows.
For a grounded reference point, the current Pro flagship — Gemini 3.1 Pro, February 2026 — is the model Gemini 4 would have to beat. Itsdocumented capabilities are here, and they are what any leaked comparison should be measured against.
Frequently asked questions
What are the Gemini 4 benchmark scores?
There are none. Google has published no Gemini 4 evaluation, and no independent lab has tested the model because no public endpoint exists. Any table of "Gemini 4 benchmark scores" circulating today is fabricated or extrapolated.
Does Gemini 4 beat Claude Fable 5.1 or GPT-6 Astra?
Leaked evaluations are reported to show an advantage in complex coding against Fable 5.1 and in multi-step reasoning against GPT-6 Astra. No scores, benchmark names or harness details accompany those reports, so the claims cannot be checked, reproduced or ranked.
Was Gemini 4 tested on Arena.ai?
Observers reported side-by-side comparisons on Arena.ai against an unlabelled model believed to be a Gemini 4 Pro checkpoint. Anonymous-model testing is a normal pre-release practice, but attribution is inferred from output style rather than confirmed, and arena results measure human preference rather than capability.
When will real Gemini 4 benchmarks appear?
Google publishes a benchmark table in its announcement post, typically on launch day. Independent numbers from LMArena, SWE-bench, ARC-AGI and similar follow within days of API access. Until an API model ID exists, nobody outside Google can run an evaluation.