Four labs shipped inside three days. Anthropic went first with Claude Fable 5.1 and Mythos 5.1 on the 1st, Google put out Gemini 3.8 Flash and a gated Cyber variant on the 2nd, Meta released Muse Spark 1.3 the same day, and OpenAI closed the week on the 3rd with GPT-6 Astra and a press briefing that ended with "Welcome to the AGI era." Two of those four are being discussed. The other two are being used as evidence that the race now has two horses in it.

Astra's reception was not the coronation the launch copy asked for. Epoch AI rolls more than fifty benchmarks into a single score and puts Astra first at 169, ahead of Fable 5.1 at 163. Artificial Analysis runs its own aggregate and put Astra at 61, exactly level with GPT-5.6 Sol and five points behind Fable 5.1. Two respectable independent shops looked at the same model in the same week and disagreed about whether anything much had happened. The split runs through the sub-scores too: AA has Fable 5.1 leading its Coding Agent Index at 70 against Astra's 67, and Astra costs about 75 percent more per task than GPT-5.6 Sol did at maximum effort. Whatever Astra is, it isn't a clean generational step past the model Anthropic shipped two days earlier.

The vendor numbers are worse than merely contradictory. Astra scored 99.9 percent on ARC-AGI-3 using OpenAI's own Provider Adapter harness and 62.7 percent on the standardised one: same model, same benchmark, thirty-seven points of scaffolding. That gap is why I'll take the independent aggregates seriously and leave the launch-page figures alone, even when the aggregates contradict each other, which they plainly do.

Fable 5.1's reception was quieter and stranger. It led the Artificial Analysis index at 66 on release, the highest any model had posted. Within days AA rebuilt the index, and the new version has Fable 5.1 first at 57, Astra at 55, Opus 5 at 54, and Muse Spark 1.3 down at 53 alongside Fable 5. Zvi Mowshowitz called the retroactive adjustment "more than a little suspicious" while allowing that the new ordering is more plausible, and he's right on both counts. The interesting casualty is Meta. AA's own launch writeup had Astra trailing Muse Spark 1.3 at maximum effort; after the rebuild, Muse Spark sits two points below it. Nothing about either model changed in between.

Google's problem is real and it isn't a benchmark problem. The company has shipped four Gemini Flash models in 106 days, 3.5 through 3.8 since May, and the flagship Gemini 3.5 Pro that Sundar Pichai promised in June still hasn't appeared. That cadence is genuinely impressive and it has a hole in the middle. The 3.8 Flash release is a good product at $0.75 and $3.75 per million tokens through the end of December, and by Google's own positioning it's the budget tier. Fortune counts Google's best model slipping to tenth on the Artificial Analysis ranking after Meta went past it, which is a strange sentence to write about the company that published the transformer paper.

Meta is the case the two-horse framing handles worst. Muse Spark 1.3 claims about 20 percent fewer tool calls and 25 percent fewer tokens than 1.2, and it lands within a couple of points of Astra on either version of the AA index while costing a fraction of Astra's $10 and $50. It also arrived without the thing that used to make Meta distinctive. Muse Spark was the company's first proprietary model since Superintelligence Labs formed, a deliberate walk away from Llama's open weights, and the reward so far has been a week of write-ups filing it among the also-rans.

So the gap is a category rather than a capability, and the clean way to see it is price rather than safety paperwork. OpenAI and Anthropic both shipped at $10 and $50 per million tokens. Google and Meta both shipped under $5 on output. Only two labs put a flagship-tier model on sale this week, which is not the same claim as only two labs being able to build one. On the gating question all four behaved identically: Google's Fairwind programme holds back the Cyber variant of 3.8 Flash in the same shape that Daybreak holds back Astra, and Anthropic restricts Mythos 5.1 to vetted organisations rather than selling it. Four labs, one playbook, two price tiers.

Whether the flagship tier is the race worth winning is a separate question, and the money doesn't obviously say yes. Anthropic's own flagship took six percent of Anthropic's tokens in its first month on sale, and Astra shipped to a limited set of organisations with no date for everyone else. The cheapest capable models in the field are the two that Google and Meta just built.

Sources: