
GPT-6 Astra and Claude Fable 5.1: What Actually Changed
Two large model releases landed within four days of each other in September 2026: Claude Fable 5.1 from Anthropic and GPT-6 Astra from OpenAI. If you are trying to work out which one to put into a workflow, the honest answer today is that the two makers have published very different amounts of information. This post lays out what is actually known about each, marks clearly where the record is empty, and gives you a way to test the difference yourself instead of waiting for a scoreboard.
What launched, and when
Anthropic released Claude Fable 5.1 on 1 September 2026. The API name is claude-fable-5-1. Alongside it sits Claude Mythos 5.1, which Anthropic describes as the same model with fewer restrictions, available only through a trusted-access programme aimed at cybersecurity and life sciences work.
OpenAI put GPT-6 Astra into limited preview on 3 September 2026 and released it publicly on 4 September 2026. That public release is for paying users only, and it is a restricted version of the model rather than the full thing.
The release patterns already differ. One is a general availability launch with a separately gated high-capability variant. The other is a staged rollout where even the public version is restricted.
What each maker claims
OpenAI describes Astra as a "generational leap" for cybersecurity, professional work, software engineering and science, and states that it is faster and handles more tasks than earlier versions. The examples OpenAI gives are practical rather than academic: filling in a tax return, building a game scene, ordering food, searching for jobs. OpenAI also says Astra was trained on more than 100,000 GPUs at its Stargate facility in Texas, and calls it "by far" its largest training run to date.
Anthropic's claims for Fable 5.1 are narrower and more numeric. According to Anthropic, the model scores higher than Fable 5 across a set of named benchmarks, and it costs less to run on typical workloads because of a change to cache pricing. Anthropic also states two safety changes: cybersecurity safeguards now permit vulnerability research, and biology safeguards produce 85% fewer false positives on harmless requests.
Note the shape of the difference. OpenAI's claims are about breadth of capability and scale of training. Anthropic's are about measured deltas and unit economics. Neither shape is inherently more trustworthy, but the two are not comparable statements.
The benchmark situation, which is the real story
Anthropic published benchmark numbers for Fable 5.1 against its own predecessor, Fable 5:
- Terminal-Bench-Science 0.1: 52.6% versus 24.7%
- Terminal-Bench 4.0 (coding): 55.8% versus 42.0%
- Humanity's Last Exam, without tools: 60.9% versus 57.8%
- CursorBench 3.2.0: 73.4% versus 70.5%
- OSWorld 2.0, strict: 41.7% versus 36.1%
OpenAI has published no benchmark scores for GPT-6 Astra in the material available at the time of writing.
That means a head-to-head benchmark comparison between these two models is not possible today. Not difficult, not contested: impossible, because half the numbers do not exist publicly. Anyone publishing a table with both models scored side by side is either using a third-party evaluation they should be naming, or filling gaps with guesses.
Read the Anthropic numbers for what they are, too. They compare Fable 5.1 to Fable 5, not to any competitor. A 52.6% against a previous 24.7% tells you the direction of travel inside one product line. It says nothing about where a different maker's model would land on the same test.
Pricing transparency
Anthropic published prices for Fable 5.1: $10 per million input tokens, $50 per million output tokens, and $0.25 per million cache-read tokens. The cache-read figure is a reduction from $1.00, which Anthropic describes as a 75% cut. On the back of that change, Anthropic says it measures roughly 25% lower costs on typical workloads and up to roughly 45% on agentic workloads.
OpenAI has not published pricing for GPT-6 Astra in the source material. We are not going to estimate it. If you need a cost projection for Astra, you need OpenAI's own rate card, and until that is out you cannot budget for it with any confidence.
This asymmetry has a practical consequence that outranks most capability arguments. You can model a monthly bill for Fable 5.1 from your own token volumes right now. You cannot do that for Astra yet.
Availability and access restrictions
Fable 5.1 is reachable through Claude.ai, the Claude API, Amazon Web Services, Google Cloud and Microsoft Azure, and it is in Claude Code and Claude Enterprise. Three major clouds matters if your procurement already runs through one of them, because it often means no new vendor contract. The less restricted Mythos 5.1 variant sits behind a trusted-access programme.
Astra's public release is limited to paying users and is a restricted version. Its advanced cybersecurity capabilities are held back, with access initially for testers only and possible later expansion through a programme OpenAI refers to as "Daybreak Blue".
Both makers, then, are gating their most capable security-related behaviour behind a vetting process. The mechanisms differ in name and detail, but the pattern is the same on both sides.
The chain-of-thought point on Astra, in plain words
Astra uses a technique OpenAI calls "recurrent depth", also described as looped transformers. The part that matters to a non-researcher is the consequence, not the architecture.
Most current reasoning models produce a visible trail of intermediate steps before they answer. That trail is useful in two ways. It lets you check whether the model reached a right answer for a right reason, and it gives safety researchers something to inspect when they want to know what a model was doing. According to the available material, recurrent depth hides some or all of that chain of thought, because part of the reasoning happens in repeated internal passes rather than as generated text you can read.
The concern raised about this is monitorability: if the reasoning is not expressed in readable form, it is harder for anyone, including the maker, to follow how the model got somewhere. For an ordinary buyer this is not an abstract debate. If you work in a regulated setting, or you need to explain to a client or an auditor why an automated decision came out the way it did, a model that shows less of its working is harder to defend. If you are doing casual drafting work, it may not matter to you at all.
What this means if you are choosing today
Set aside which model is "better", because nothing published supports an answer to that. Ask instead which unknowns you can live with.
If you need a cost model before you commit, Fable 5.1 is the only one of the two you can price today. If you need proof of capability against a named test, Anthropic has published a self-comparison and OpenAI has published nothing, so you would be buying Astra on the strength of the maker's description. If you need explainable reasoning, the chain-of-thought question on Astra is worth resolving before you build anything on it. If you specifically need the tasks OpenAI highlighted, and you are already a paying OpenAI customer, trying the restricted release directly costs you little.
Both models are days old. Independent evaluations and real deployment reports will arrive over the coming weeks, and they will be worth more than either launch announcement. Deciding now means deciding on incomplete information, which is a legitimate choice as long as you know that is what you are doing.
How to test this for yourself
The fastest route past a stalled comparison is your own evidence. Build a small fixed set of tasks that look like your actual work, not like a benchmark: one long document to summarise, one piece of code to fix with a real error attached, one messy dataset to restructure, one piece of writing in your house voice, one multi-step task that needs the model to use a tool and then act on the result.
Run the identical prompt on each model and save both outputs verbatim. Score them on things you can check: did it get the facts right, did it follow the format you asked for, how much editing did the output need. Repeat the set a week later, because both models are new and behaviour can shift as the makers adjust them. Keep notes on token usage too, since on the model with published prices that converts straight into a cost per task.
Our comparison guide walks through a 30-minute version of this self-test with a scoring sheet you can reuse. It works for any two models, including ones that launch after this post.
Where to go next
- AI tool comparisons for the 30-minute self-test method and the scoring sheet
- Prompt engineering for writing the fixed prompts your test set depends on
- AI for professionals for putting a chosen model into day-to-day work
Stuck on a step, or unsure how to score a test run against your own work? Write to the desk and describe what you are trying to compare.
Comments
No comments yet. Be the first to share your thoughts.


