✳ the wire · opinion

Kimi K3, day one: impressive, expensive, and a little slow

2026-07-16 · Kimi K3 · by ArtificialWatch

Kimi K3, day one: impressive, expensive, and a little slow

Kimi K3 has been live for a day, and the picture forming is more interesting than the launch-night hype. Here's everything we're seeing so far.

First, the price. K3 costs $3/M input and $15/M output — verified against the live API listings. What made Kimi the community favorite through the K2 era was affordability, and this isn't that: Grok 4.5 sits at $2/M in and $6/M out, which makes K3 two and a half times more expensive on output tokens than a model it's competing with head-to-head. The affordability era appears to be over; the question is whether the quality justifies the new tier.

Second, the subsidy era looks over too. Moonshot isn't cushioning consumer plans the way it used to, and early adopters are reporting they're burning through usage allowances fast. The 10-30% launch top-up bonus softens that for now — but it expires.

Third, quality. On Artificial Analysis' independent runs, K3 lands just below the frontier: 1668 Elo on GDPval v2 knowledge work versus Claude Fable 5's 1760 — but comfortably above Claude Opus 4.8's 1600 and GPT-5.5's 1494. On agentic browsing it posts a 91.2% BrowseComp, beating Fable 5's 88.0% though short of Sol's 92.2% (the 'SOTA BrowseComp' claim circulating on X doesn't survive contact with the numbers). Slightly worse than Fable 5 and GPT-5.6 overall is a fair one-line read — remarkable for an open-weights-bound model, but you're paying closed-model prices for it.

Fourth, read the terms before you send it anything sensitive. Moonshot's TOS states: "We may share your personal information with our corporate affiliates..." and "We may disclose personal information to public authorities... to comply with legal obligations or lawful requests by public authorities, courts, or regulators." For a company operating under Chinese jurisdiction, draw your own conclusions about what 'lawful requests by public authorities' can mean. Enterprises with data-residency requirements should wait for the open weights and self-host.

Fifth, it's not fast. Watching @_MaxBlade build a Subway Surfers-style game with it on stream: over 25 minutes and roughly $30 of tokens for the build. But — and this is the part that matters — the results were genuinely impressive. This model is really good at game development, and the early artifact quality in long agentic runs is the strongest argument for it.

The bottom line, one day in: K3 is a real frontier-adjacent model with a genuinely new architecture (2.8T parameters, Kimi Delta Attention, native vision, 1M context), priced like it knows it. If the full open weights land within days as promised, self-hosters get the best deal in AI. If you're paying the API rates, run your own evals first — that's exactly what our prompt packs are for.

Update, July 16: Artificial Analysis published its full assessment — 57 on the Intelligence Index (Opus 4.8/GPT-5.5-class, behind Fable 5 and Sol), #1 on AutomationBench, and $0.94 cost per task, which is actually cheaper per task than Opus 4.8 despite the sticker shock. And the game-dev instinct above proved out fast: K3 took #1 on Frontend Code Arena at 1679 points, first in six of seven domains — second only to Fable 5 in, fittingly, Gaming. The weights now have a date: July 27.

Related: Artificial Analysis (benchmark runs) · Kimi K3 tracker · the bench index

← back to the wire