Anthropic shipped Claude Opus 5 this afternoon at five dollars per million input tokens and twenty-five out, identical to Opus 4.8, with the claim that it comes close to Fable 5 at half the price. Fable now bills at ten and fifty, so the halving is arithmetic rather than marketing.

On Frontier-Bench v0.1, an agentic terminal coding test, Opus 5 scores 43.3% against Opus 4.8's 18.7% and Fable 5's 33.7%. That gap is wide enough to survive a lot of scepticism, and some should be applied, because the footnote says the run was internal, on the mini-SWE-agent harness, mean reward over five attempts, with Opus 4.8 substituted whenever the safety classifier refused a task. A refusal scores zero otherwise, so 43.3% describes Opus 5 with Opus 4.8 covering its refusals, not Opus 5 alone. The same substitution holds up Fable's 33.7%. Opus 4.8's own 18.7% had nothing to fall back on.

Read the announcement looking for GPT-5.6 Sol or Gemini 3.1 Pro and they are not in it. Every named comparison is Claude against Claude: Opus 4.8, Fable 5, and Mythos 5, which still leads on cybersecurity work. The charts plot cost per task along one axis, so the claim isn't "we score higher", it's "we score higher for the money." Even the boldest line, a score three times the next-best model on ARC-AGI 3, which tests unfamiliar problems rather than variations on training data, never says who next-best is. OfficeChai reads that field as including Sol. Anthropic's own labels don't.

Two ways to take this, and I can't pick between them. Cost per task is the honest axis for anyone running agents in production, where finishing a job in fewer tokens beats a prettier percentage, and the shared leaderboards are rotting anyway, with SWE-bench Verified showing signs of contamination that let models pattern-match issue text instead of reasoning through it. Against that, a curve generated on your own harness, against your own back catalogue, cannot be checked by anyone outside the building.

Where the rivals actually stand is easier to state than to compare. GPT-5.6 Sol has been out since July 9 at $5/$30, and its max configuration scored 59 on the Artificial Analysis Intelligence Index against Opus 4.8's 55.7. Anthropic published nothing today that updates that composite, so on general reasoning the gap is unmeasured rather than closed. Gemini 3.1 Pro runs at roughly $2 and $12, a third-party estimate Google has never confirmed, well under half what Opus 5 charges, and it takes audio and video natively, which today's announcement never claims for Opus 5.

I went looking for one clean head-to-head reasoning score to settle it and found the same tracker quoting 94.6% and 80.0% for Sol on the same benchmark, on the same page. So no ranking from me, and I'd distrust anyone selling one this week. What today's sheet supports is narrow: Opus 5 costs what Opus 4.8 cost, which makes trying it cheap. Everything past that waits on numbers Anthropic didn't publish.

Sources: