✳ the bench index

Every leaderboard, one table.

The public boards each tell part of the story — community votes here, live task suites there, lab-reported launch numbers somewhere else. This table reads them together: each model's best-variant score on every board we track, normalized to a percentile rank, and combined into one composite index. Refreshed daily, free, no signup.

Table last refreshed 2026-09-12· each board's own as-of date is listed with the sources below.

The index at a glancetrailingleading
Claude Opus 596.7
Claude Fable 593.8
GPT-5.687.5
Kimi K372.6
GPT-5.562.5
Claude Opus 4.860
Claude Sonnet 555.6
Grok 4.553.3
Muse Spark 1.146.9
DeepSeek V4 Pro43.8
Gemini 3.5 Flash37.5
Gemini 3.1 Pro31.3
GLM-5.229.4
Qwen 3.7 Max25
Kimi K2.618.8
Kimi K2.7 Code13.1
MiniMax M35.6
#ModelAW IndexLMArenaLiveBenchSimpleBenchLLM StatsVals AIAA
1Claude Opus 5Anthropic96.71,687 10080.1 8280.6 9354.9 8067.2 10050.7 100
2Claude Fable 5Anthropic93.81,628 8983 10081.9 10054.3 6066 9449.7 94
3GPT-5.6OpenAI87.51,617.3 8381.1 9471.7 6755 10063.7 8847.1 88
4Kimi K3Moonshot72.61,674.3 9479.2 7760.7 4053.1 4057.8 6943.8 81
5GPT-5.5OpenAI62.51,529.7 4480.2 8876.9 8057.4 6338.6 63
6Claude Opus 4.8Anthropic601,558.8 6176.2 5964.8 4750.9 060.9 8142 75
7Claude Sonnet 5Anthropic55.61,542.6 5676 5360.6 3359.6 7538.4 56
8Grok 4.5xAI53.31,574.3 6775.8 4770 5351.5 3139.1 69
9Muse Spark 1.1Meta46.91,541.5 5075.3 4154.8 5634.3 44
10DeepSeek V4 ProDeepSeek43.81,580.5 7277.4 7150.9 1352.1 2052.4 3836.3 50
11Gemini 3.5 FlashGoogle37.51,500 1774.6 3576.7 7353.1 4433 38
12Gemini 3.1 ProGoogle31.31,487 1177 6579.6 8741.9 630.4 31
13GLM-5.2Zhipu AI29.41,591.8 7873.2 2958.8 2753.1 5027.1 13
14Qwen 3.7 MaxAlibaba251,517.1 2873.1 2470.4 6044.8 2529.9 25
15Kimi K2.6Moonshot18.81,521 3970.5 1243.5 19
16Kimi K2.7 CodeMoonshot13.11,510.6 2268.4 657.9 2026.3 6
17MiniMax M3MiniMax5.61,486.7 667.3 045.8 042.7 1329.6 19
18InklingThinking Machines01,439.2 071.9 1850 734.1 025.5 0
Seed 2.1 ProByteDanceneeds 2+ boards1,518.7 33

Cell format: raw board score + percentile chip. Hover any cell for the exact variant behind the number.

How the index is computed
01Best variant wins

Labs ship many configs — thinking modes, effort levels, harnesses. Per board, each model is scored by its best-performing variant, identically for everyone. Search, grounding, image and video product modes are excluded.

02Percentiles, not raw averages

Elo ≈1500 and 0–100 scores can't be averaged. Each board's scores become percentile ranks (0–100) within the set of models on this table that the board actually scores.

03The index is a median

A model's AW Index is the median of its percentiles across boards — so no single board's methodology can dominate. A model needs 2+ boards to be ranked; boards scoring fewer than 5 of these models display but don't vote.

04Nothing is estimated

Every raw number comes from the source board, links back to it, and carries its fetch date. Percentiles are relative to this table's model set, not the boards' full lists.

The boards
LMArenacommunity votes — millions of blind head-to-headsElo (community votes) · 19 models mapped · as of 2026-09-12LiveBenchindependent — contamination-free live task suiteglobal average (0-100) · 18 models mapped · as of 2026-09-12SimpleBenchindependent — private reasoning set, human-baselinedAVG@5 % (private set) · 16 models mapped · as of 2026-09-12LLM StatsLLM Stats' headline composite — verified benchmarks + live perfLLM Stats Score · composite (0-100), their rendered top 15 · 6 models mapped · as of 2026-09-12Vals AIindependent — expert-crafted private tasks, measured runsVals Index accuracy % · weighted finance + coding tasks · 17 models mapped · as of 2026-09-12Artificial Analysisindependent — standardized eval suite, measured runs onlyintelligence index · measured runs only · 17 models mapped · as of 2026-09-08

Benchmarks tell you who leads today; the radar tells you the minute that changes. Watch the models · the wire · track record