✳ the wire · analysis

GLM-5.3 tops CyberGym and trails the frontier on exploitation — and we missed it because the sweep had no eyes on models.dev

GLM-5.3confirmedby ArtificialWatch
GLM-5.3 tops CyberGym and trails the frontier on exploitation — and we missed it because the sweep had no eyes on models.dev
Source imagery · verified against a primary source

Z.ai shipped GLM-5.3 today, and we did not catch it. The entry was armed and the pattern was right — the sweep simply had no source that carried the id. That is now fixed, and the fix is the more useful half of this post.

What launched, from Z.ai's own blog. GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training, on the same 743B base. Z.ai reports a 50% improvement over 5.2 on its in-house Code Bench and open-source state of the art on Terminal Bench 3.0 and Agents' Last Exam.

The headline nobody should skim past is the cyber one, and Z.ai's framing is unusually candid: cyber capability "developed faster than we expected". GLM-5.3 tops the CyberGym vulnerability-discovery column on its own comparison table at 84.5, above Fable 5 at 83.8, GPT-5.6 Sol at 83.6 and DeepSeek-V4 Pro at 83.3.

Read the exploitation rows before concluding it leads on cyber generally, because that is where the same table says something different. On ExploitGym at 2h/6h, GLM-5.3 scores 105/130 against GPT-5.6 Sol's 216/293 and Fable 5's 181/247. On ExploitBench it is 54.4 against Sol's 76.5 and Fable 5's 78.0. So GLM-5.3 is the strongest open model at finding vulnerabilities and still well behind the closed frontier at exploiting them. Z.ai's line that "gains are largest further up the exploitation chain" is a claim about its improvement over 5.2 — 24.4 to 54.4 on ExploitBench — not about where it stands. Both readings are on the same page; only one is in the headline.

The open-weights claim needs a date attached. Z.ai calls this the most capable open-weights model for coding, and the Hugging Face link on the page says "coming soon": the weights land roughly two weeks after launch, once safety evaluation and hardening are complete. So the claim is currently about a model nobody outside can download, and the stated reason for the delay is the capability the post is advertising.

Against Kimi K3, the open model it has to beat, its own table is mixed rather than decisive. GLM-5.3 wins Terminal Bench 3.0 (28.3 to 17.4), ProgramBench (19.0 to 17.5), AutomationBench (48.2 to 46.7), PostTrainBench (39.8 to 32.0) and GDPval-AA v2 (1769 to 1682). It loses Terminal Bench 2.1 (88.2 to 88.3), DeepSWE v1.1 (66.9 to 67.5) and SWE-Marathon (42.5 to 48.1). "Most capable" is defensible on the balance of rows; it is not a sweep.

One week apart, two opposite answers to the same capability. On August 7 OpenAI said it could not rule out Critical cyber capability for Astra and slowed the model down, pausing internal work that did not meet new security controls. Today Z.ai ships a model whose cyber capability it says outran expectations, publishes the benchmarks, and gates only the weights for a fortnight. We are not equating them — a Critical designation under a published preparedness framework is a different object from a two-week hold — but the contrast is the story of the month in one line: the same emergent capability produced a delay at one lab and a launch post at another. ExploitGym, the benchmark family in the middle of both, is the same one OpenAI told us in July its escaped research prototype had been running in.

Now the part about us, because the question we got was fair.

Our GLM entry was armed and its pattern matched glm-5.3 correctly. It did not fire because the launch sweep reads Anthropic's native list, OpenRouter, seven key-gated first-party APIs and one public chat surface — and GLM-5.3 is on none of them. Z.ai launched on its own blog; the id reached models.dev under three coding-plan providers within hours, and OpenRouter still has nothing. models.dev was already in the codebase, used for price drift and for the armed audit, and was never wired into the launch sweep itself. A correct entry, a correct pattern, and no eyes on the shelf the model was sitting on.

That is fixed: models.dev is now a launch source. Worth stating the safety case, because adding 6,300 ids to a system that fires paid alerts deserves one. Running today's models.dev catalog through the registry produces exactly one detection — GLM-5.3 — and nothing else. That is not luck: the armed audit has been checking every armed pattern against this exact list for weeks, so the patterns arrived pre-vetted rather than newly exposed.

And the second failure, which was worse. The autopilot has been red since yesterday, so no sweep published anything at all. The cause was a test asserting that 75% of the whole wire carries real cover art. Radar items are machine-written and use a drawn cover by design, so every sweep pushed that ratio down; at 92 items it crossed the line, the test failed, and the failing test blocked the same autopilot's publish gate. The radar broke itself by working. The test now asserts what its own comment always said — every hand-written item carries real art, machine-written ones may use the fallback — which is a stricter rule that sweep volume cannot erode.

Finally, the registry split. GLM-5.3 and 5.5 shared one armed entry, so firing on 5.3 would have spent the single one-shot and left GLM-5.5 — a launch we have covered three times — permanently silent. 5.3 is now live, 5.5 stays armed on its own. That is the same mistake we found in the Grok entry yesterday, which suggests the rule worth writing down is: a hunt that succeeds is not finished until the entry that caught it has been retired.

Source: Z.ai's own tech blog, read directly · models.dev catalog · our launch-sweep source list · GLM-5.3 tracker · the bench index

sweeping every 60 seconds

Know the minute it drops — not the minute we write it up.

Claude Opus 5 went live at 16:51 UTC. The alert was in subscribers' inboxes at 16:52.

  • Free forever
  • No card
  • Unsubscribe in one click

← back to the wire