✳ the wire · intel
Opus 5's benchmarks went up and the developer reviews went down — which is a question our drift harness exists to answer

Five days after Opus 5 shipped, the reaction has split. Anthropic priced it at $5/$25 against Fable 5's $10/$50 with higher coding and reasoning scores, and on Terminal-Bench 2.1 v1.1 it leads the field at 74% pass@1. Meanwhile a visible share of developers report messy outputs, degradation on long-context coding, and a preference for complex solutions where a direct one would do — with some staying on Opus 4.8 or moving to GPT-5.6.
We are labelling this reported, not confirmed, and we want to be exact about why. 'Feels worse' is not measurable from the outside. It has at least four ordinary explanations that have nothing to do with a weaker model — thinking-mode mismatch, long-thread context handling, shared capacity pressure at peak, and routing differences between Claude surfaces — and every one of them produces the same complaint.
Every wire item is source-verified before it posts and graded: confirmed · strong · reported · rumor. Subscribers see the primary source on each one, the reasoning behind its grade, and the full analysis — plus the launch alerts themselves, which fire the minute a model answers rather than whenever the feed catches up.


