✳ the bench index

Every leaderboard, one table.

The public boards each tell part of the story — community votes here, live task suites there, lab-reported launch numbers somewhere else. This table reads them together: each model's best-variant score on every board we track, normalized to a percentile rank, and combined into one composite index. Refreshed daily, free, no signup.

Table last refreshed 2026-07-25· each board's own as-of date is listed with the sources below.

The index at a glancetrailingleading
Claude Opus 594.6
Claude Fable 594.1
GPT-5.688.2
Kimi K378.7
Claude Opus 4.875.7
GPT-5.563.4
Grok 4.557.9
Muse Spark 1.156.3
Claude Sonnet 554.4
Gemini 3.5 Flash41.2
GLM-5.240.6
Seed 2.1 Pro35.6
Gemini 3.1 Pro35.3
Qwen 3.7 Max30.3
MiniMax M317.6
DeepSeek V4 Pro17.6
Kimi K2.7 Code13.1
Kimi K2.612.5
#ModelAW IndexLMArenaLiveBenchSimpleBenchLLM StatsVals AIAA
1Claude Opus 5Anthropic94.680.6 9358 9674.8 9460.7 100
2Claude Fable 5Anthropic94.11,635.8 9480.8 9481.9 10057.5 8275.1 10059.9 94
3GPT-5.6OpenAI88.21,625 8882.4 10071.7 6758 9673.1 8158.9 88
4Kimi K3Moonshot78.71,682 10078.5 7560.7 4055.7 7374.7 8857.1 82
5Claude Opus 4.8Anthropic75.71,567.5 7778.9 8164.8 4752.6 6470.4 7555.7 77
6GPT-5.5OpenAI63.41,525.3 4779.9 8876.9 8049.1 2768 5654.8 71
7Grok 4.5xAI57.91,550.2 7176.2 6370 5349.5 3665.3 5053.8 65
8Muse Spark 1.1Meta56.31,536.3 5976.2 5652.1 5568.4 6350.6 47
9Claude Sonnet 5Anthropic54.41,541.1 6574.8 5060.6 3350.9 4668.6 6953.4 59
10Gemini 3.5 FlashGoogle41.21,492.6 2474.6 4476.7 7362.7 3850.2 41
11GLM-5.2Zhipu AI40.61,587.7 8273.2 3858.8 2747.1 965 4451.1 53
12Seed 2.1 ProByteDance35.61,528.3 5347.3 18
13Gemini 3.1 ProGoogle35.31,485.8 1277.1 6979.6 8753.8 646.5 35
14Qwen 3.7 MaxAlibaba30.31,518.7 4173.1 3170.4 6046.7 057.5 2546 29
15MiniMax M3MiniMax17.61,492.5 1867.3 045.8 058.9 3144.4 24
16DeepSeek V4 ProDeepSeek17.61,464.2 671.6 1950.9 1355.6 1944.3 18
17Kimi K2.7 CodeMoonshot13.11,517.4 2968.4 657.9 2041.9 6
18Kimi K2.6Moonshot12.51,518.7 3570.5 1355.2 1344.2 12
19InklingThinking Machines01,444.9 071.7 2550 749.3 040.7 0

Cell format: raw board score + percentile chip. Hover any cell for the exact variant behind the number.

How the index is computed
01Best variant wins

Labs ship many configs — thinking modes, effort levels, harnesses. Per board, each model is scored by its best-performing variant, identically for everyone. Search, grounding, image and video product modes are excluded.

02Percentiles, not raw averages

Elo ≈1500 and 0–100 scores can't be averaged. Each board's scores become percentile ranks (0–100) within the set of models on this table that the board actually scores.

03The index is a median

A model's AW Index is the median of its percentiles across boards — so no single board's methodology can dominate. A model needs 2+ boards to be ranked; boards scoring fewer than 5 of these models display but don't vote.

04Nothing is estimated

Every raw number comes from the source board, links back to it, and carries its fetch date. Percentiles are relative to this table's model set, not the boards' full lists.

The boards
LMArenacommunity votes — millions of blind head-to-headsElo (community votes) · 18 models mapped · as of 2026-07-25LiveBenchindependent — contamination-free live task suiteglobal average (0-100) · 17 models mapped · as of 2026-07-21SimpleBenchindependent — private reasoning set, human-baselinedAVG@5 % (private set) · 16 models mapped · as of 2026-07-25LLM StatsLLM Stats' headline composite — verified benchmarks + live perfLLM Stats Score · composite (0-100), their rendered top 15 · 12 models mapped · as of 2026-07-25Vals AIindependent — expert-crafted private tasks, measured runsVals Index accuracy % · weighted finance + coding tasks · 17 models mapped · as of 2026-07-25Artificial Analysisindependent — standardized eval suite, measured runs onlyintelligence index · measured runs only · 18 models mapped · as of 2026-07-25

Benchmarks tell you who leads today; the radar tells you the minute that changes. Watch the models · the wire · track record