✳ the bench index
Every leaderboard, one table.
The public boards each tell part of the story — community votes here, live task suites there, lab-reported launch numbers somewhere else. This table reads them together: each model's best-variant score on every board we track, normalized to a percentile rank, and combined into one composite index. Refreshed daily, free, no signup.
Table last refreshed 2026-07-25· each board's own as-of date is listed with the sources below.
| # | Model | AW Index | LMArena | LiveBench | SimpleBench | LLM Stats | Vals AI | AA |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 94.6 | — | — | 80.6 93 | 58 96 | 74.8 94 | 60.7 100 |
| 2 | Claude Fable 5Anthropic | 94.1 | 1,635.8 94 | 80.8 94 | 81.9 100 | 57.5 82 | 75.1 100 | 59.9 94 |
| 3 | GPT-5.6OpenAI | 88.2 | 1,625 88 | 82.4 100 | 71.7 67 | 58 96 | 73.1 81 | 58.9 88 |
| 4 | Kimi K3Moonshot | 78.7 | 1,682 100 | 78.5 75 | 60.7 40 | 55.7 73 | 74.7 88 | 57.1 82 |
| 5 | Claude Opus 4.8Anthropic | 75.7 | 1,567.5 77 | 78.9 81 | 64.8 47 | 52.6 64 | 70.4 75 | 55.7 77 |
| 6 | GPT-5.5OpenAI | 63.4 | 1,525.3 47 | 79.9 88 | 76.9 80 | 49.1 27 | 68 56 | 54.8 71 |
| 7 | Grok 4.5xAI | 57.9 | 1,550.2 71 | 76.2 63 | 70 53 | 49.5 36 | 65.3 50 | 53.8 65 |
| 8 | Muse Spark 1.1Meta | 56.3 | 1,536.3 59 | 76.2 56 | — | 52.1 55 | 68.4 63 | 50.6 47 |
| 9 | Claude Sonnet 5Anthropic | 54.4 | 1,541.1 65 | 74.8 50 | 60.6 33 | 50.9 46 | 68.6 69 | 53.4 59 |
| 10 | Gemini 3.5 FlashGoogle | 41.2 | 1,492.6 24 | 74.6 44 | 76.7 73 | — | 62.7 38 | 50.2 41 |
| 11 | GLM-5.2Zhipu AI | 40.6 | 1,587.7 82 | 73.2 38 | 58.8 27 | 47.1 9 | 65 44 | 51.1 53 |
| 12 | Seed 2.1 ProByteDance | 35.6 | 1,528.3 53 | — | — | 47.3 18 | — | — |
| 13 | Gemini 3.1 ProGoogle | 35.3 | 1,485.8 12 | 77.1 69 | 79.6 87 | — | 53.8 6 | 46.5 35 |
| 14 | Qwen 3.7 MaxAlibaba | 30.3 | 1,518.7 41 | 73.1 31 | 70.4 60 | 46.7 0 | 57.5 25 | 46 29 |
| 15 | MiniMax M3MiniMax | 17.6 | 1,492.5 18 | 67.3 0 | 45.8 0 | — | 58.9 31 | 44.4 24 |
| 16 | DeepSeek V4 ProDeepSeek | 17.6 | 1,464.2 6 | 71.6 19 | 50.9 13 | — | 55.6 19 | 44.3 18 |
| 17 | Kimi K2.7 CodeMoonshot | 13.1 | 1,517.4 29 | 68.4 6 | 57.9 20 | — | — | 41.9 6 |
| 18 | Kimi K2.6Moonshot | 12.5 | 1,518.7 35 | 70.5 13 | — | — | 55.2 13 | 44.2 12 |
| 19 | InklingThinking Machines | 0 | 1,444.9 0 | 71.7 25 | 50 7 | — | 49.3 0 | 40.7 0 |
Cell format: raw board score + percentile chip. Hover any cell for the exact variant behind the number.
Labs ship many configs — thinking modes, effort levels, harnesses. Per board, each model is scored by its best-performing variant, identically for everyone. Search, grounding, image and video product modes are excluded.
Elo ≈1500 and 0–100 scores can't be averaged. Each board's scores become percentile ranks (0–100) within the set of models on this table that the board actually scores.
A model's AW Index is the median of its percentiles across boards — so no single board's methodology can dominate. A model needs 2+ boards to be ranked; boards scoring fewer than 5 of these models display but don't vote.
Every raw number comes from the source board, links back to it, and carries its fetch date. Percentiles are relative to this table's model set, not the boards' full lists.
Benchmarks tell you who leads today; the radar tells you the minute that changes. Watch the models · the wire · track record