Tuesday, Not Thursday
Anthropic shipped Claude Fable 5.1 this morning, a Tuesday, and I owe the Polymarket traders an apology. Two days ago I wrote that I'd bet against their median of 1 September, on the grounds that nobody had shipped anything to justify a fortnight of repricing. They were right and I wasn't. The 36kr story that had Fable 5.1 held back to land within hours of GPT-6 was wrong too, in the other direction: Anthropic went first.
The oddest line in the launch post is a footnote saying every score was produced with the production safeguards switched on, and that where a safeguard intervened the task was either marked zero or handed to an Opus to finish. Anthropic is telling you its own table is depressed, and I believe the table more for it. Terminal-Bench-Science goes from 24.7 percent on Fable 5 to 52.6. The stated error of up to four and a half points doesn't threaten a 28-point gap, but the same footnote admits the public leaderboard puts Opus 5 ahead of the old Fable on this test, so the base being doubled was a weak one. Terminal-Bench 4.0 rises from 42.0 to 55.8, and the unrestricted Mythos 5.1 scores 60.9 on the same run, so five points of coding ability still sit behind the cyber classifier. The comparison column throughout is GPT-5.6 Sol, not the model everyone is waiting for.
The price on the box hasn't moved: $10 in, $50 out, the same figures that sent Fable upstairs in July. What moved is cache reads, which drop to $0.25 per million, a quarter of what Fable 5 charged. Anthropic's arithmetic says typical workloads come out around 25 percent cheaper and long agentic runs up to 45 percent, because a multi-hour session re-reads its own transcript on every turn and that line item swamps the rest. Cognition says it's moving its Opus 5 traffic in Devin over on launch day because the new cache price makes a Fable-class model "finally economical" for work it had kept on Opus, which is the endorsement that matters, since the whole worry about Fable 5 was that nobody could afford to run it. A Hacker News commenter linked the FT's report of sluggish corporate demand for Fable 5, and the cache cut reads like an answer to that report rather than to any benchmark.
Reception on Hacker News, a few hours in, splits in two. One camp says the price cut is the only real change. The larger camp is still arguing about the fallback classifier. On one side are people who can't get Fable to write an auth endpoint, review unsafe Rust it had just written, or look at anything mentioning seccomp. On the other is Simon Willison, who gets punted to Opus rarely and one-shots most of what he tries. Anthropic's claim of 60 percent fewer cyber false positives will be tested by exactly those people this week, and letting the model find vulnerabilities without writing exploits for them is a more usable line than the one it replaced.
As for OpenAI, its response is Astra, and the question is whether competition can move a date that a safety framework set. On 7 August the company said it couldn't rule out critical cyber capability, and a White House official told Axios it had volunteered its plans to delay. That framework doesn't run on a two-week timer, and Anthropic taking its customers doesn't change what the evaluations say. What has changed is the evidence that the gate is being cleared: TestingCatalog had the first outputs circulating on the 29th, and a leaker says partners got a build called ultima-alpha over the weekend with a wider launch aimed at Thursday the 3rd. The same leaker called Fable 5.1 for last week, and the reply thread under the post says so, so I'd take the partner build and leave the date. I have chased that Thursday before and found nothing at the end of it. Partners holding a build is further than any previous rumour got, so my guess is within a fortnight, with the framework still able to hold the door.
Sources:
-
Introducing Claude Fable 5.1 and Claude Mythos 5.1 — Anthropic
-
Claude Fable 5.1 model overview — Claude Platform Docs
-
Claude Fable 5.1 and Claude Mythos 5.1 discussion — Hacker News
-
OpenAI slows release of Astra model citing cyber capabilities — Axios
-
First outputs from GPT-6 "Astra" model from OpenAI — TestingCatalog
-
ChatGPT 6.0 Astra & Image 2.0 update on the horizon, September 3rd? — Reddit
Filed under AI & machine learning
This post is timestamped using Blockchain technology. Verify