← comparisons · head to head
Gemini 3.5 Flash vs Gemini 3.1 Pro
Of the 5 public leaderboards that rate both, Gemini 3.5 Flash scores higher on 3 and Gemini 3.1 Pro on 2. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Gemini 3.5 Flash is at 27.3 and Gemini 3.1 Pro at 22.7 (out of 100).
Scores as published by each board; table composed September 24, 2026.
Get the next launch alert, free →
Leaderboard by leaderboard
| Board | Gemini 3.5 Flash | Gemini 3.1 Pro | Higher |
|---|---|---|---|
| LMArena Elo (community votes) | 1,500 | 1,486.8 | Gemini 3.5 Flash |
| LiveBench global average (0-100) | 74.6 | 77 | Gemini 3.1 Pro |
| SimpleBench AVG@5 % (private set) | 76.7 | 79.6 | Gemini 3.1 Pro |
| Vals AI Vals Index accuracy % · weighted finance + coding tasks | 53.1 | 41.9 | Gemini 3.5 Flash |
| Artificial Analysis intelligence index · measured runs only | 32.6 | 29.7 | Gemini 3.5 Flash |
Raw scores are each board's own scale, so compare within a row, not across rows. Full table: the bench index.
Gemini 3.5 Flash at a glance
- Developer
- Status
- Live since May 19, 2026
- API price
- $1.50 input / $9 output per 1M tokens · Google, gemini-3.5-flash
- Context window
- 1.05M tokens
Live — preview since I/O, on OpenRouter.
Gemini 3.1 Pro at a glance
- Developer
- Status
- Live
Live.
Gemini 3.5 Flash vs Gemini 3.1 Pro: quick answers
Which is better, Gemini 3.5 Flash or Gemini 3.1 Pro?
It depends on what you measure. Of the 5 public leaderboards that rate both, Gemini 3.5 Flash scores higher on 3 and Gemini 3.1 Pro on 2. Gemini 3.5 Flash's widest lead is on Vals AI; Gemini 3.1 Pro's widest lead is on LiveBench. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Gemini 3.5 Flash is at 27.3 and Gemini 3.1 Pro at 22.7 (out of 100).
Who makes Gemini 3.5 Flash and Gemini 3.1 Pro?
Both are from Google.
Where do these numbers come from?
Each score is the leaderboard's own published figure, read by ArtificialWatch (table composed September 24, 2026). Boards measure different things, so raw scores are only comparable within a row.
More comparisons
Watch free: Chrome push the moment a new model answers on a public API, an email about 15 minutes later. Also: the bench index · API pricing · the wire