← comparisons · head to head

Gemini 3.5 Flash vs Gemini 3.1 Pro

Of the 5 public leaderboards that rate both, Gemini 3.5 Flash scores higher on 3 and Gemini 3.1 Pro on 2. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Gemini 3.5 Flash is at 27.3 and Gemini 3.1 Pro at 22.7 (out of 100).

Scores as published by each board; table composed September 24, 2026.

Get the next launch alert, free →

Leaderboard by leaderboard

BoardGemini 3.5 FlashGemini 3.1 ProHigher
LMArena
Elo (community votes)
1,5001,486.8Gemini 3.5 Flash
LiveBench
global average (0-100)
74.677Gemini 3.1 Pro
SimpleBench
AVG@5 % (private set)
76.779.6Gemini 3.1 Pro
Vals AI
Vals Index accuracy % · weighted finance + coding tasks
53.141.9Gemini 3.5 Flash
Artificial Analysis
intelligence index · measured runs only
32.629.7Gemini 3.5 Flash

Raw scores are each board's own scale, so compare within a row, not across rows. Full table: the bench index.

Gemini 3.5 Flash at a glance

Developer
Google
Status
Live since May 19, 2026
API price
$1.50 input / $9 output per 1M tokens · Google, gemini-3.5-flash
Context window
1.05M tokens

Live — preview since I/O, on OpenRouter.

Full Gemini 3.5 Flash tracker →

Gemini 3.1 Pro at a glance

Developer
Google
Status
Live

Live.

Full Gemini 3.1 Pro tracker →

Gemini 3.5 Flash vs Gemini 3.1 Pro: quick answers

Which is better, Gemini 3.5 Flash or Gemini 3.1 Pro?

It depends on what you measure. Of the 5 public leaderboards that rate both, Gemini 3.5 Flash scores higher on 3 and Gemini 3.1 Pro on 2. Gemini 3.5 Flash's widest lead is on Vals AI; Gemini 3.1 Pro's widest lead is on LiveBench. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Gemini 3.5 Flash is at 27.3 and Gemini 3.1 Pro at 22.7 (out of 100).

Who makes Gemini 3.5 Flash and Gemini 3.1 Pro?

Both are from Google.

Where do these numbers come from?

Each score is the leaderboard's own published figure, read by ArtificialWatch (table composed September 24, 2026). Boards measure different things, so raw scores are only comparable within a row.

More comparisons

GPT-6 Astra vs Claude Fable 5.1GPT-6 Astra vs GPT-5.6GPT-6 Astra vs Claude Opus 5Claude Fable 5.1 vs Claude Fable 5Claude Fable 5.1 vs Claude Opus 5Claude Opus 5 vs GPT-5.6Claude Opus 5 vs Claude Opus 4.8Claude Opus 5 vs Claude Fable 5Claude Opus 5 vs Kimi K3Kimi K3 vs Claude Fable 5Kimi K3 vs GLM-5.2Kimi K3 vs GPT-5.6DeepSeek V4 Pro vs GLM-5.2DeepSeek V4 Pro vs Kimi K3Grok 4.7 vs GPT-6 AstraGrok 4.7 vs Claude Opus 5Grok 4.7 vs Claude Fable 5.1Qwen 3.8 Max vs Claude Opus 5Qwen 3.8 Max vs Claude Fable 5Qwen 3.8 Max vs Kimi K3Qwen 3.8 Max vs DeepSeek V4 ProQwen 3.8 Max vs GLM-5.2Qwen 3.8 Max vs GPT-5.6GPT-6 Astra vs Claude Fable 5Claude Fable 5 vs Claude Opus 4.8Claude Fable 5 vs Claude Sonnet 5Claude Sonnet 5 vs Claude Opus 4.8Claude Sonnet 5 vs Claude Opus 5GLM-5.2 vs Claude Opus 4.8GLM-5.2 vs Claude Fable 5GLM-5.2 vs GPT-5.5Gemini 3.5 Flash vs Claude Sonnet 5GPT-6 Astra vs Claude Opus 4.8Claude Fable 5.1 vs Claude Opus 4.8Kimi K3 vs Claude Opus 4.8
Hear the minute the next one drops.

Watch free: Chrome push the moment a new model answers on a public API, an email about 15 minutes later. Also: the bench index · API pricing · the wire