← comparisons · head to head

Kimi K3 vs DeepSeek V4.1 Flash

Of the 6 public leaderboards that rate both, Kimi K3 scores higher on 3 and DeepSeek V4.1 Flash on 3. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Kimi K3 is at 54.1 and DeepSeek V4.1 Flash at 49.2 (out of 100).

Scores as published by each board; table composed September 28, 2026.

Get the next launch alert, free →

Leaderboard by leaderboard

BoardKimi K3DeepSeek V4.1 FlashHigher
LMArena
Elo (community votes)
1,659.91,620.7Kimi K3
LiveBench
global average (0-100)
79.281.1DeepSeek V4.1 Flash
SimpleBench
AVG@5 % (private set)
60.766.7DeepSeek V4.1 Flash
LLM Stats
LLM Stats Score · composite (0-100), their rendered top 15
52.551.4Kimi K3
Vals AI
Vals Index accuracy % · weighted finance + coding tasks
57.857.9DeepSeek V4.1 Flash
Artificial Analysis
intelligence index · measured runs only
43.639.5Kimi K3

Raw scores are each board's own scale, so compare within a row, not across rows. Full table: the bench index.

For coding

On the ArtificialWatch coding index, Kimi K3 is at 80 (#7 of 31) and DeepSeek V4.1 Flash at 72.4 (#8). Kimi K3 scores higher on 2 of the 3 coding boards that rate both and DeepSeek V4.1 Flash on 1.

Coding boardKimi K3DeepSeek V4.1 Flash
LMArena · Coding1,540.51,532.2
LiveBench · Coding71.878.7
Vals Vibe Code Bench8584.7

The full AI coding leaderboard →

Kimi K3 at a glance

Developer
Moonshot
Status
Live since July 15, 2026
API price
$3 input / $15 output per 1M tokens · Moonshot AI, kimi-k3
Context window
1.05M tokens

Live — dropped Jul 15 · 2.8T · open weights landed Jul 27 (1.56 TB, MXFP4).

Kimi K3, day one: impressive, expensive, and a little slow

2026-07-16 · Kimi K3 has been live for a day, and the picture forming is more interesting than the launch-night hype. Here's everything we're seeing so far.

Full Kimi K3 tracker →

DeepSeek V4.1 Flash at a glance

Developer
DeepSeek
Status
Live since September 10, 2026

Live — Sep 10 · 1M ctx · $0.15/$0.60 per 1M as our sweep read it · DeepSeek serves it under the unversioned id deepseek-flash, so the id itself never carries the version.

Full DeepSeek V4.1 Flash tracker →

Kimi K3 vs DeepSeek V4.1 Flash: quick answers

Which is better, Kimi K3 or DeepSeek V4.1 Flash?

It depends on what you measure. Of the 6 public leaderboards that rate both, Kimi K3 scores higher on 3 and DeepSeek V4.1 Flash on 3. Kimi K3's widest lead is on LLM Stats; DeepSeek V4.1 Flash's widest lead is on LiveBench. On the ArtificialWatch Index, which takes the median of each model's percentile across the boards that score it, Kimi K3 is at 54.1 and DeepSeek V4.1 Flash at 49.2 (out of 100).

Which is newer, Kimi K3 or DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash. It went live on September 10, 2026, 57 days after Kimi K3 (July 15, 2026).

How do Kimi K3 and DeepSeek V4.1 Flash compare for coding?

On the ArtificialWatch coding index, Kimi K3 is at 80 (#7 of 31) and DeepSeek V4.1 Flash at 72.4 (#8). Kimi K3 scores higher on 2 of the 3 coding boards that rate both and DeepSeek V4.1 Flash on 1.

Who makes Kimi K3 and DeepSeek V4.1 Flash?

Kimi K3 is from Moonshot; DeepSeek V4.1 Flash is from DeepSeek.

Where do these numbers come from?

Each score is the leaderboard's own published figure, read by ArtificialWatch (table composed September 28, 2026). Boards measure different things, so raw scores are only comparable within a row.

More comparisons

GPT-6 Astra vs Claude Fable 5.1GPT-6 Astra vs GPT-5.6GPT-6 Astra vs Claude Opus 5Claude Fable 5.1 vs Claude Fable 5Claude Fable 5.1 vs Claude Opus 5Claude Opus 5 vs GPT-5.6Claude Opus 5 vs Claude Opus 4.8Claude Opus 5 vs Claude Fable 5Claude Opus 5 vs Kimi K3Kimi K3 vs Claude Fable 5Kimi K3 vs GLM-5.2Kimi K3 vs GPT-5.6DeepSeek V4 Pro vs GLM-5.2DeepSeek V4 Pro vs Kimi K3Grok 4.7 vs GPT-6 AstraGrok 4.7 vs Claude Opus 5Grok 4.7 vs Claude Fable 5.1Qwen 3.8 Max vs Claude Opus 5Qwen 3.8 Max vs Claude Fable 5Qwen 3.8 Max vs Kimi K3Qwen 3.8 Max vs DeepSeek V4 ProQwen 3.8 Max vs GLM-5.2Qwen 3.8 Max vs GPT-5.6GPT-6 Astra vs Claude Fable 5Claude Fable 5 vs Claude Opus 4.8Claude Fable 5 vs Claude Sonnet 5Claude Sonnet 5 vs Claude Opus 4.8Claude Sonnet 5 vs Claude Opus 5GLM-5.2 vs Claude Opus 4.8GLM-5.2 vs Claude Fable 5GLM-5.2 vs GPT-5.5Gemini 3.5 Flash vs Gemini 3.1 ProGemini 3.5 Flash vs Claude Sonnet 5GPT-6 Astra vs Claude Opus 4.8Claude Fable 5.1 vs Claude Opus 4.8Kimi K3 vs Claude Opus 4.8DeepSeek V4 Pro vs Claude Opus 4.8DeepSeek V4 Pro vs GPT-5.5DeepSeek V4 Pro vs Claude Fable 5GLM-5.3 vs Kimi K3GLM-5.3 vs GLM-5.2GLM-5.3 vs GPT-6 AstraGrok 4.6 vs Claude Opus 5Grok 4.6 vs Claude Fable 5Grok 4.6 vs GPT-5.6Grok 4.6 vs Grok 4.5Grok 4.6 vs Claude Sonnet 5Gemini 3.8 Flash vs Gemini 3.1 ProGemini 3.8 Flash vs Claude Opus 5Muse Spark 1.3 vs Gemini 3.8 FlashMuse Spark 1.3 vs GLM-5.3GPT-6 Sol vs GPT-6 AstraGPT-6 Sol vs Claude Opus 5GPT-6 Sol vs Claude Fable 5GPT-6 Luna vs GPT-6 AstraGPT-6 Sol vs GPT-6 LunaClaude Fable 5 vs GPT-5.6Claude Opus 4.8 vs GPT-5.5GPT-5.6 vs Claude Opus 4.8GPT-5.6 vs Claude Sonnet 5Claude Sonnet 5 vs GPT-5.5Claude Fable 5 vs GPT-5.5DeepSeek V4 Pro vs Claude Opus 5Qwen 3.8 Max vs GLM-5.3Qwen 3.8 Max vs Claude Opus 4.8GLM-5.3 Flash vs GLM-5.3GLM-5.3 Flash vs GLM-5.2MiniMax M3 vs Kimi K3MiniMax M3 vs DeepSeek V4 ProMiniMax M3 vs GLM-5.2MiniMax M3 vs GPT-5.5GLM-5.3 vs DeepSeek V4 ProGLM-5.3 vs Claude Opus 5GLM-5.3 vs Claude Fable 5Grok 4.7 vs Grok 4.6Muse Spark 1.3 vs Claude Fable 5.1Muse Spark 1.3 vs Claude Opus 5Muse Spark 1.3 vs DeepSeek V4.1 FlashMuse Spark 1.3 vs Grok 4.6Muse Spark 1.3 vs GPT-6 AstraGemini 3.8 Flash vs GPT-6 AstraGemini 3.8 Flash vs Claude Fable 5.1Gemini 3.8 Flash vs Claude Sonnet 5Gemini 3.8 Flash vs Grok 4.6Gemini 3.8 Flash vs GPT-5.6GPT-5.6 vs GPT-5.5
sweeping every 60 seconds

Hear the minute the next one drops. Free.

A free Chrome push the moment a new model answers on a public API, an email about 15 minutes later. Texts or a phone call on paid plans.

  • Free forever
  • No card
  • Unsubscribe in one click

Also: the bench index · API pricing · the wire