← comparisons · head to head

Kimi K3 vs Claude Opus 4.8

Of the 5 public leaderboards that rate both, Kimi K3 scores higher on 3 and Claude Opus 4.8 on 2. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Kimi K3 is at 60.4 and Claude Opus 4.8 at 47.8 (out of 100).

Scores as published by each board; table composed September 24, 2026.

Get the next launch alert, free →

Leaderboard by leaderboard

BoardKimi K3Claude Opus 4.8Higher
LMArena
Elo (community votes)
1,659.61,555.2Kimi K3
LiveBench
global average (0-100)
79.276.2Kimi K3
SimpleBench
AVG@5 % (private set)
60.764.8Claude Opus 4.8
Vals AI
Vals Index accuracy % · weighted finance + coding tasks
57.860.9Claude Opus 4.8
Artificial Analysis
intelligence index · measured runs only
43.641.8Kimi K3

Raw scores are each board's own scale, so compare within a row, not across rows. Full table: the bench index.

Kimi K3 at a glance

Developer
Moonshot
Status
Live since July 15, 2026
API price
$3 input / $15 output per 1M tokens · Moonshot AI, kimi-k3
Context window
1.05M tokens

Live — dropped Jul 15 · 2.8T · open weights landed Jul 28 (1.56 TB, MXFP4).

Kimi K3, day one: impressive, expensive, and a little slow

2026-07-16 · Kimi K3 has been live for a day, and the picture forming is more interesting than the launch-night hype. Here's everything we're seeing so far.

Full Kimi K3 tracker →

Claude Opus 4.8 at a glance

Developer
Anthropic
Status
Live since May 27, 2026
API price
$5 input / $25 output per 1M tokens · Anthropic, claude-opus-4-8
Context window
1M tokens

Live.

Full Claude Opus 4.8 tracker →

Kimi K3 vs Claude Opus 4.8: quick answers

Which is better, Kimi K3 or Claude Opus 4.8?

It depends on what you measure. Of the 5 public leaderboards that rate both, Kimi K3 scores higher on 3 and Claude Opus 4.8 on 2. Kimi K3's widest lead is on LMArena; Claude Opus 4.8's widest lead is on Vals AI. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Kimi K3 is at 60.4 and Claude Opus 4.8 at 47.8 (out of 100).

Which is newer, Kimi K3 or Claude Opus 4.8?

Kimi K3. It went live on July 15, 2026, 49 days after Claude Opus 4.8 (May 27, 2026).

Which is cheaper, Kimi K3 or Claude Opus 4.8?

Kimi K3 is cheaper on both input and output. On each vendor's own API, Kimi K3 is $3 input / $15 output per million tokens and Claude Opus 4.8 is $5 input / $25 output per million tokens. Prices as published in the models.dev catalog, as of 2026-09-24.

Which has the bigger context window, Kimi K3 or Claude Opus 4.8?

Kimi K3, at 1.05M tokens against Claude Opus 4.8's 1M, as each vendor lists it.

Who makes Kimi K3 and Claude Opus 4.8?

Kimi K3 is from Moonshot; Claude Opus 4.8 is from Anthropic.

Where do these numbers come from?

Each score is the leaderboard's own published figure, read by ArtificialWatch (table composed September 24, 2026). Boards measure different things, so raw scores are only comparable within a row.

More comparisons

GPT-6 Astra vs Claude Fable 5.1GPT-6 Astra vs GPT-5.6GPT-6 Astra vs Claude Opus 5Claude Fable 5.1 vs Claude Fable 5Claude Fable 5.1 vs Claude Opus 5Claude Opus 5 vs GPT-5.6Claude Opus 5 vs Claude Opus 4.8Claude Opus 5 vs Claude Fable 5Claude Opus 5 vs Kimi K3Kimi K3 vs Claude Fable 5Kimi K3 vs GLM-5.2Kimi K3 vs GPT-5.6DeepSeek V4 Pro vs GLM-5.2DeepSeek V4 Pro vs Kimi K3Grok 4.7 vs GPT-6 AstraGrok 4.7 vs Claude Opus 5Grok 4.7 vs Claude Fable 5.1Qwen 3.8 Max vs Claude Opus 5Qwen 3.8 Max vs Claude Fable 5Qwen 3.8 Max vs Kimi K3Qwen 3.8 Max vs DeepSeek V4 ProQwen 3.8 Max vs GLM-5.2Qwen 3.8 Max vs GPT-5.6GPT-6 Astra vs Claude Fable 5Claude Fable 5 vs Claude Opus 4.8Claude Fable 5 vs Claude Sonnet 5Claude Sonnet 5 vs Claude Opus 4.8Claude Sonnet 5 vs Claude Opus 5GLM-5.2 vs Claude Opus 4.8GLM-5.2 vs Claude Fable 5GLM-5.2 vs GPT-5.5Gemini 3.5 Flash vs Gemini 3.1 ProGemini 3.5 Flash vs Claude Sonnet 5GPT-6 Astra vs Claude Opus 4.8Claude Fable 5.1 vs Claude Opus 4.8
Hear the minute the next one drops.

Watch free: Chrome push the moment a new model answers on a public API, an email about 15 minutes later. Also: the bench index · API pricing · the wire