✳ the wire · intel

Kimi K3 inference: Fireworks Fast is 4.5× the throughput at 1.5× the price, and uptime is the real spread

Kimi K3confirmedby ArtificialWatch
Fireworks AI — announcement art
Source imagery · verified against a primary source

The provider table for Kimi K3 is unusually wide right now, and the cheapest row is not the interesting one.

Most providers — Modal, Baseten, Together, Moonshot itself, DigitalOcean, Fireworks standard — sit at the same $3.00 / $15.00 per 1M with $0.30 cache reads. Within that identical price, P50 latency runs from 3.32s (Modal) to 7.29s (Moonshot's own endpoint), and throughput from 17 tps (Baseten) to 36 tps (Modal). Same model, same price, better than 2× the speed depending on who you route to.

The rest of this analysis — 3 more paragraphs — plus the source, is on the paid plans.

Every wire item is source-verified before it posts and graded: confirmed · strong · reported · rumor. Subscribers see the primary source on each one, the reasoning behind its grade, and the full analysis — plus the launch alerts themselves, which fire the minute a model answers rather than whenever the feed catches up.

See the plans →

Don't read about it here first.

The wire is curated after the fact. The radar sweeps every provider API every 60 seconds and emails you the minute a new model actually answers — Opus 5 took 61 seconds from first sighting to inbox. Free forever, no card.

Email + Chrome push on the free tier. Paid plans add instant SMS and an automated phone call. Unsubscribe in one click, always.

← back to the wire