Lavine Web & AI Solutions Logo
Industry News
September 2026 · 5 min read

Two frontier models, three days apart. Here's what actually changed.

Anthropic shipped Claude Fable 5.1 on 1 September. OpenAI shipped GPT-6 Astra on the 3rd. Both list at about £7 per million input tokens and £37 per million output - and that identical price tag hides the one number that will decide your bill.

NewsAI StrategyAgentic AISmall Business

The releases: Two launches, one week

Anthropic released Claude Fable 5.1 on 1 September, alongside a restricted sibling called Mythos 5.1 that only goes to vetted cybersecurity and life sciences teams. Two days later OpenAI announced GPT-6 Astra, trained on more than 100,000 GPUs at its Stargate site in Texas.

Both models run on the big three clouds. Both carry a context window of roughly one million tokens and cap output at 128,000. Both charge the same headline price per token. If you were waiting for one of these launches to make your choice obvious, it did the opposite.

We use Claude every day for client work, so Fable 5.1 landed on our desk immediately. Astra rolled out far more slowly, and the reason for that is more interesting than the benchmarks.

The benchmarks: A genuine split decision

Each company published a table where it wins. That is normal. What is unusual is how cleanly the wins divide by task type.

Astra takes science and computer use. It scores 97.6% on FrontierMath Tier 4 against Fable's 87.8%, 64.6% on Terminal-Bench Science against 52.6%, and 92.7% on ScreenSpot-Pro against 87.3%. OpenAI also claims it drives a computer roughly twice as fast as the previous generation, and Codex now keeps searchable notes across context windows instead of compressing old work into a summary.

Fable 5.1 wins on reasoning depth. It reaches 65.0% on Humanity's Last Exam with tools against Astra's 57.2%, and 63% on SciCode against 56%. Independent evaluator Artificial Analysis puts it ahead on coding agent work.

The more useful comparison is against each model's own predecessor. Fable 5.1 more than doubled its Terminal-Bench Science score, from 24.7% to 52.6%, and lifted AutomationBench from 17.1% to 31.4%. Anthropic also published results from real scientific work: protein binder designs hitting around a 50% success rate where 10-15% is typical, a Venus elevation map at 2-3km resolution instead of 10-20km, and GPU kernel rewrites running up to 2.5 times faster.

The pricing: Where the real difference hides

The headline token prices match, so ignore them. The number that moves is cache reads: about 18p per million on Fable 5.1 against 74p on Astra. Anthropic cut that price by 75% at launch. Both vendors price in dollars, so every figure here is converted at about $1.35 to the pound and will drift with the rate.

A coding agent or a long research loop re-reads the same context on every single turn, so cache reads end up dominating the invoice. Anthropic estimates around 25% lower cost on typical workloads and up to 45% on heavily agentic ones. Astra also doubles its input price above 272,000 tokens, where Anthropic charges one flat rate all the way to a million.

Token efficiency pushes back the other way. Artificial Analysis measured the cost of running its full evaluation suite at roughly £3,900 on Astra against £9,700 on Fable 5.1, because Astra reaches its score using far fewer tokens. Fable wins the jobs that loop, Astra wins the jobs that produce one expensive answer.

Price an agent against Fable. Price high-volume one-shot analysis against Astra. Neither answer survives if you assume the list price tells you the cost.

The catch: Both vendors got more careful

Astra is the first model OpenAI has classified at the Critical cybersecurity threshold under its Preparedness Framework, meaning it can find unknown security flaws and build new ways to exploit them with very little human direction. The public version refuses a set of cyber prompts outright, full capability sits behind a vetted programme called Daybreak, and enterprise administrators have to turn it on deliberately.

Astra also uses a technique OpenAI calls recurrent depth, which obscures its chain of thought. Safety researchers have flagged that as a monitorability problem, because you can no longer read the reasoning to check what the model is doing.

Anthropic attacked the same problem from the opposite end. Fable 5.1 reports 60% fewer false positives on cybersecurity work and 85% fewer on benign biology requests, and it now allows defensive vulnerability discovery while still refusing exploit development. It is the same risk handled two ways: OpenAI gated the capability, Anthropic tuned the refusals.

The practical effect on a small business is identical. Check what your plan actually enables before you promise a client anything, because the model in the marketing page and the model on your account may not be the same thing.

The bottom line: Pick per job, not per vendor

Two frontier models, three days apart, at an identical list price with a split benchmark result is the clearest signal yet that brand loyalty makes a poor procurement strategy. The question worth asking is which model suits the job in front of you, and what that job costs when you run it a thousand times.

For most of the SMEs and startups we work with, none of this changes the plan. Both models sailed past the capability needed for the work that pays the bills - drafting, summarising, customer triage, internal search, tidying up messy data. What changed is the cost of running those things at scale, and the cache read line is where you will feel it.

If you are weighing up which model to build on, or what your current AI spend is actually buying, get in touch and we can work through it with you.

Run your own workload against both and compare the invoices. The benchmark tables are marketing; your bill is not.


Steve Lavine is a full-stack developer and founder of Lavine Web & AI Solutions, working with SMEs and startups across the UK. lavine.dev

Leave a comment
Claude Fable 5.1 vs GPT-6 Astra: Same Price, Real Gap | Lavine Web and AI Solutions