Astra's Catwalk and the Workbench
Launch season hit like a fashion week where everyone shows the same coat. Three labs, three flagships, all photographed on the same Eames chair, and all claiming to own the future. OpenAI is pitching GPT-6 Astra as the world’s most intelligent and aligned model.1 Its launch pack puts the new system at the front of computer use, scientific terminal work, long-context retrieval and several cyber evaluations.2 Meta says Muse Spark 1.3 has frontier performance, but even sympathetic coverage notes its best results come from a model developers can’t broadly use yet.3 Anthropic has quietly let Fable 5.1 lap them both on the main index. The way they are all fighting over the same ultra heavy benchmark crown while the people actually paying the bills just want a reliable workhorse is wild, and worth pulling apart.
On Artificial Analysis, Astra at its max reasoning effort gives the same headline Intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high).4 In the same index it scores exactly equal to Sol on general intelligence at 61 but significantly lower on the agentic cut at 51 vs 58, and lower than Fable 5.1, Opus 5 and even Muse Spark 1.3 in both.5 Dig a layer deeper and Muse Spark 1.3 (max) beats GPT-6 Astra on the coding agent staples that actually decide whether an engineering run finishes, with DeepSWE at 75.4 percent to 74.1 percent and a bigger gap on AutomationBench.6 One community note calls Muse Spark 1.3 Max the first Meta model to surpass OpenAI’s best on Artificial Analysis’ Intelligence Index.7 I have been running the sort of high but not heroic effort settings an engineer would actually live in, Opus High, Sol xHigh, that sort of band. In that band Muse at xhigh matches Opus on the intelligence score for a fraction of the cost per task, while Astra at high sits a point lower and costs materially more even before you crank it to max. That is the workhorse test, and Astra flunks the value part. It is not a rounding error, it is a third of the budget on a big run.
There is a decent counter to all this bean counting, and I do not want to handwave it. Latent Space argues Astra looks written to do the automated engineering OpenAI needs internally, the sort of long horizon, terminal heavy, tool using work that benchmarks barely capture. If that is true the disconnect I am pointing to loses a lot of force. The problem is OpenAI did not make that case. It put the cherry picked terminal table front and centre and left everyone else to infer the internal usefulness. On Hacker News the mood is already sour on these tables. One commenter calls the ARC-AGI-3 scorecard extremely misleading because the harness flatters the new model.8 Another asks whether anyone really believes the index reflects reality when it places Muse above Astra, Sol and Fable.9 I get the scepticism. When the same top line can be shared by three very different systems it stops feeling like a ruler and starts feeling like a mood board.
So where does that leave us. I reckon the labs are marketing to the wrong room. Anthropic is at least consistent, it sells Fable as a top shelf reasoning engine and prices it accordingly, and the index backs it. Meta has finally built a frontier model that is also a good deal, which is why engineers are dusting off harnesses to give Muse a proper crack. OpenAI still owns the workhorse tier that pays the bills, Sol and Luna on subscription are fast, capable, and cheap enough to run all day, but Astra’s PR is aimed at the finance crowd that wants to hear world’s most intelligent, not the builder who wants to know cost per completed ticket. If Astra really is a specialist for OpenAI’s own automation loop, say so, show it, price it for that job. If the next press release has fewer superlatives and more dollars per finished task, I will listen. Until then I am with Mat on this one: Astra is a pass at scale, not because it is a bad model, but because the top shelf flex is not what the people writing the cheques are actually buying. The next move that matters will not be another saturated TerminalBench chart, it will be a cheaper, faster, more reliable workhorse, and whoever ships that without the theatre will clean up.
Sources
- 1 GPT-6 Astra: A new generation of intelligence | OpenAI — openai.com
- 2 GPT-6 Astra: Price, Access and What the Benchmarks Show — digitalapplied.com
- 3 Meta says Muse Spark 1.3 has frontier performance — venturebeat.com
- 4 @aniviacat on GPT-6 Astra — ycombinator.com
- 5 @eis on GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index — ycombinator.com
- 6 @kzrdude on OpenAI begins rolling out GPT-6 Astra — ycombinator.com
- 7 @vinhnx on Muse Spark 1.3 — ycombinator.com
- 8 @intenex on GPT-6 Astra — ycombinator.com
- 9 @modeless on GPT-6 Astra — ycombinator.com
How this was made
- 01-research z-ai/glm-5.3 $0.398
- 04-nominate z-ai/glm-5.3 $0.003
- 06-write meta/muse-spark-1.2 $0.069
- 08-visualise anthropic/claude-sonnet-5 $0.043
total $0.513