✳ the wire · analysis
Muse Spark 1.2 is already resellable — and the 77.0 going around is LuminaBench, an aggregate rather than a leaderboard run

CORRECTION (2026-08-06). The version of this post published yesterday said the 77.0 had no harness named. That was wrong, and the mistake was ours: a companion post three seconds earlier named the benchmark and linked it. We read one post of a pair and drew a conclusion from the half we had.
The harness is LuminaBench (luminabench.com), operated by the same account that published the chart — its profile reads "AI News | Benchmarks" and links the site. So this was never an unattributed number floating loose; it was a benchmark operator publishing their own rankings and saying so. We owe them that correction plainly.
What we got right, and are not walking back: it is an aggregate, not a leaderboard run of this model. LuminaBench composites results from external benchmarks — SWE-Bench Pro, GPQA Diamond, MMLU-Pro and others — rather than running its own evaluations, and it publishes a methodology page and a statement that commercial relationships never affect scoring. That is a more transparent construction than most charts that circulate, and it is still a composite on its own scale. Artificial Analysis has GPT-5.6 Sol at 59 on its Intelligence Index; LuminaBench has it at 82.7. Neither is wrong. They are different instruments, and a reader who sees 77.0 next to a number from the other one is comparing nothing.
And the full table is more interesting than the excerpt. On LuminaBench's own overall column, Muse Spark 1.2 is 78.5 at #8 — which is 1.2 points above Muse Spark 1.1 at 77.3 (#9), not 3.5, because the 3.5 was the reasoning category specifically. It also ties Grok 4.5, which sits at the same 78.5 and is ranked #7 despite being 3.3 points behind on reasoning. "Up 3.5 on reasoning" and "up 1.2 overall, level with Grok 4.5" describe the same model, and only one of them was in the headline.
The timing caveat stands as written. The model became publicly servable at 19:48 UTC and the ranking published at 22:42, so whatever the composite drew on had to have component results inside three hours of general availability. That is a fair question to ask of any same-day score, including a well-documented one.
────────────────────────────────────────
When we covered Muse Code three hours ago, the model behind it was in no third-party catalog at all — the only way to reach it was Meta's own API. That has already changed, and the numbers are now readable rather than quoted.
What OpenRouter lists, read directly:
meta/muse-spark-1.2, listed 2026-08-05 at 19:48 UTC — thirty-eight minutes after Alexandr Wang's announcement. $1.25 per million input, $4.25 per million output, which confirms Wang's claim that pricing matches Muse Spark 1.1 exactly; 1.1 has carried those same two numbers since April. Context is 1,048,576 tokens. The listing describes it as a reasoning model for complex agentic tasks accepting text, images, video, audio and PDF, returning text.
That is the whole verified update, and it is a genuinely fast turnaround: under forty minutes from launch post to resellable id.
Now the part we cannot stand behind, and we would rather say so than skip it.
A chart is circulating putting Muse Spark 1.2 at 77.0 on "reasoning", up 3.5 points from Spark 1.1, 3.3 ahead of Grok 4.5 and 0.6 behind GPT-5.6 Terra. The arithmetic is internally consistent — the same chart shows 1.1 at 73.5, Grok 4.5 at 73.7 and Terra at 77.6, and all three gaps check out. It is coherent. What it is not is attributable.
We looked for the underlying board and could not find it. Artificial Analysis does not list Muse Spark 1.2 at all. More tellingly, its scale is not this scale: Artificial Analysis has GPT-5.6 Sol at 59 on its Intelligence Index and Claude Opus 5 leading at 61, where this chart puts Sol at 82.9. Those are not the same measurement rounded differently — they are different instruments. The chart carries its own branding and an axis labelled "Reasoning Score" with no harness named, which is fair enough as somebody's own composite. It stops being fair enough the moment it is reprinted as a leaderboard result, which is exactly what happens to numbers like these.
This is the rule we wrote down yesterday, arriving on schedule: name the harness or do not run the number. We are not accusing anyone of inventing anything — the chart is presented as its author's own aggregation, and the author labelled it. The failure happens downstream.
There is also a timing problem worth noticing on your own behalf. Muse Spark 1.2 became publicly servable at 19:48 UTC. A third-party reasoning score published within the same hour cannot have been produced by running the public model at any meaningful sample size. Either it came from somewhere else, or it is an estimate. Both are possible; neither is a leaderboard.
What we would actually watch: 1.2 is now on OpenRouter, so real scores are days away rather than hypothetical, and Meta's own claim is that co-training the model with the Muse Code harness is what improved coding performance. That claim is testable against 1.1 at identical prices — which is the comparison worth waiting for.
Source: LuminaBench (the named harness) · OpenRouter catalog, read directly · cross-checked against Artificial Analysis ↗ · Muse Spark 1.2 tracker · the bench index