Two weeks have gone by since GPT-6 Astra landed on 3 September, and neither of the labs most people watch has shipped anything since. Read as a lull, that is an artefact of watching two companies. Across the frontier the median gap between model releases this year is eleven days, down from 37.5 days in 2023, so a week with nothing shipping anywhere has become the exception. Whether it comes from a lab you have heard of is a different matter.

OpenAI's median interval has compressed from 170.5 days in 2023 to 49 this year, which puts two weeks after Astra nearer the middle of its own gap than the end of it. Anthropic went on the 1st, Google on the 2nd, OpenAI on the 3rd, and all three paired a public model with a restricted cyber tier built on the same weights. That week cost three launch budgets and three safety reviews, and nobody runs it twice in a fortnight. Something will still ship: the Chinese labs Moonshot and Z.ai have emerged as serious contenders with open-weight models at or near closed-model performance, and a release from either counts in the eleven-day median while registering with almost nobody.

The fatigue is real and gets described from the wrong side. CNBC framed it as an attention problem, too many launches and not enough excitement to go round. The venture investor Trace Cohen relocated the cost to where it actually lands, on engineering teams absorbing two to three engineer-weeks per migration: rebuild the eval set, re-tune the prompts, re-measure cost per task, re-certify anything regulated. That bill appears on no pricing page.

A different complaint, and the one I would weight highest, came in July, when more than a thousand employees across the major labs signed a petition asking for a slower pace. Eval rebuilds are not what that petition is about. It points at burnout and at safety work being compressed into the gaps between launches, and it comes from inside the buildings doing the launching, which is the part that should worry people more than any consumer boredom.

On the aggregate numbers, nothing happened. Artificial Analysis put Astra at 61 on its Intelligence Index, exactly level with GPT-5.6 Sol and five points behind Claude Fable 5.1 at 66; a second reading came in at 61.2 against Sol's 60.9, inside the noise. Ten days ago I wrote about two independent shops reaching opposite verdicts on the same model in the same week, and the flat composite is what stuck.

Terminal-Bench 4.0 went from 37.3% to 57.9%. ScreenSpot Pro went from 76.9 to 92.7, AutomationBench more than doubled, and ExploitBench went from 78.5% to 100%, which mostly means ExploitBench has stopped measuring anything. Fable 5.1 moved from 42.0 to 55.8 on Terminal-Bench and roughly doubled its predecessor on the science variant. None of that reaches you through a single index score. A composite exists to return one number for an entire capability profile, and a profile this uneven is the case it represents worst.

OpenAI then undercut its own case. Astra's 99.9% on ARC-AGI-3 came from OpenAI's provider adapter, while ARC Prize's neutral harness measured 62.7% on the same model. Teach a reader to discount a number like that and the discount lands on the honest ones too, so a genuine doubling on AutomationBench arrives already devalued.

Spikiness runs downhill as well. Astra scores 57.2% on Humanity's Last Exam with tools against Fable 5.1's 65%, and it lost ground on economically-weighted work. A model can be a step change at operating a computer and flat or worse at reasoning across a long document.

One change from that week never appeared on a benchmark table at all. Cache reads fell from $1.00 to $0.25 per million tokens on the 1st while input and output stayed at $10 and $50. For an agent carrying the same prefix through hundreds of steps, cache reads are most of the invoice, so a footnote in the pricing section does more for a monthly bill than any score in the launch post.

I asked Claude last week what Anthropic would do next, and three of its four predictions had already happened before it made them. The date-guessing has the same record: every named release Thursday this summer turned back into a game of telephone when I chased it down. Eleven days is the only number here I'd put money on, and it tells you nothing about who.

Sources: