Skip to content

Plutonic Rainbows

Flagship Is a Tier, Not a Score

Four labs shipped inside three days. Anthropic went first with Claude Fable 5.1 and Mythos 5.1 on the 1st, Google put out Gemini 3.8 Flash and a gated Cyber variant on the 2nd, Meta released Muse Spark 1.3 the same day, and OpenAI closed the week on the 3rd with GPT-6 Astra and a press briefing that ended with "Welcome to the AGI era." Two of those four are being discussed. The other two are being used as evidence that the race now has two horses in it.

Astra's reception was not the coronation the launch copy asked for. Epoch AI rolls more than fifty benchmarks into a single score and puts Astra first at 169, ahead of Fable 5.1 at 163. Artificial Analysis runs its own aggregate and put Astra at 61, exactly level with GPT-5.6 Sol and five points behind Fable 5.1. Two respectable independent shops looked at the same model in the same week and disagreed about whether anything much had happened. The split runs through the sub-scores too: AA has Fable 5.1 leading its Coding Agent Index at 70 against Astra's 67, and Astra costs about 75 percent more per task than GPT-5.6 Sol did at maximum effort. Whatever Astra is, it isn't a clean generational step past the model Anthropic shipped two days earlier.

The vendor numbers are worse than merely contradictory. Astra scored 99.9 percent on ARC-AGI-3 using OpenAI's own Provider Adapter harness and 62.7 percent on the standardised one: same model, same benchmark, thirty-seven points of scaffolding. That gap is why I'll take the independent aggregates seriously and leave the launch-page figures alone, even when the aggregates contradict each other, which they plainly do.

Fable 5.1's reception was quieter and stranger. It led the Artificial Analysis index at 66 on release, the highest any model had posted. Within days AA rebuilt the index, and the new version has Fable 5.1 first at 57, Astra at 55, Opus 5 at 54, and Muse Spark 1.3 down at 53 alongside Fable 5. Zvi Mowshowitz called the retroactive adjustment "more than a little suspicious" while allowing that the new ordering is more plausible, and he's right on both counts. The interesting casualty is Meta. AA's own launch writeup had Astra trailing Muse Spark 1.3 at maximum effort; after the rebuild, Muse Spark sits two points below it. Nothing about either model changed in between.

Google's problem is real and it isn't a benchmark problem. The company has shipped four Gemini Flash models in 106 days, 3.5 through 3.8 since May, and the flagship Gemini 3.5 Pro that Sundar Pichai promised in June still hasn't appeared. That cadence is genuinely impressive and it has a hole in the middle. The 3.8 Flash release is a good product at $0.75 and $3.75 per million tokens through the end of December, and by Google's own positioning it's the budget tier. Fortune counts Google's best model slipping to tenth on the Artificial Analysis ranking after Meta went past it, which is a strange sentence to write about the company that published the transformer paper.

Meta is the case the two-horse framing handles worst. Muse Spark 1.3 claims about 20 percent fewer tool calls and 25 percent fewer tokens than 1.2, and it lands within a couple of points of Astra on either version of the AA index while costing a fraction of Astra's $10 and $50. It also arrived without the thing that used to make Meta distinctive. Muse Spark was the company's first proprietary model since Superintelligence Labs formed, a deliberate walk away from Llama's open weights, and the reward so far has been a week of write-ups filing it among the also-rans.

So the gap is a category rather than a capability, and the clean way to see it is price rather than safety paperwork. OpenAI and Anthropic both shipped at $10 and $50 per million tokens. Google and Meta both shipped under $5 on output. Only two labs put a flagship-tier model on sale this week, which is not the same claim as only two labs being able to build one. On the gating question all four behaved identically: Google's Fairwind programme holds back the Cyber variant of 3.8 Flash in the same shape that Daybreak holds back Astra, and Anthropic restricts Mythos 5.1 to vetted organisations rather than selling it. Four labs, one playbook, two price tiers.

Whether the flagship tier is the race worth winning is a separate question, and the money doesn't obviously say yes. Anthropic's own flagship took six percent of Anthropic's tokens in its first month on sale, and Astra shipped to a limited set of organisations with no date for everyone else. The cheapest capable models in the field are the two that Google and Meta just built.

Sources:

This post is timestamped using Blockchain technology. Verify

Pearls Sewn to a Hem

A short cream panel falls across the bust and stops, unattached, and from its loose hem hangs a row of long baroque pearls, each one drilled and caught at the top so it points down and swings. Two more hang from the collar points, like the ends of a tie nobody bothered to knot. The shirt they are fixed to gives them nothing to compete with: sleeveless, plain, an oversized pointed collar, no shoulder pad, no structure asking to be admired.

Ornament living on a construction line is not new. Buttons, frogging and passementerie all sit on the working parts of a garment. What is different here is that none of this is fixed flat to anything. The pearls hang off a free edge, at the size you would otherwise expect around a throat, doing the job of beaded fringe.

Yasmeen Ghauri does the sensible thing and holds still, hand in the pocket, chin level, face at neutral, so that that fringe of pearls is the only part of the look with any movement in it. She does the same withholding under a Valentino hat two years later, and it works for the same reason: the clothes are already performing.

Bernadine Morris's review of the Milan spring shows ran in the New York Times the next morning, and found that at Krizia "the pendulum has swung here from precision tailoring to insouciant soft dressing." She was writing mostly about dresses rather than anything like this, but the diagnosis holds for a shirt with no internal structure left in it. Once the tailoring goes, something has to supply the interest, and Krizia's Mariuccia Mandelli hung it off the hem. The belt is doing its bit too, plain braided leather on a plain gold buckle, ordinary enough to keep the eye up at chest height.

Five months later Mandelli was running narrow cords across a violet jacket, ornament pressed flat into the surface and holding still. The pearls are the same instinct with the fixings left loose.

Sources:

This post is timestamped using Blockchain technology. Verify

Nobody Inherits a Type

Ask two people to name the most beautiful face they know and you get two answers, then an argument. Culture is the usual explanation, and it's true enough to be useless: it predicts that everyone raised in the same place should broadly agree, and they don't. A twin study from 2015 got closer to what's actually going on.

Germine and colleagues had 547 pairs of identical twins and 214 pairs of same-sex fraternal twins rate 200 faces each. Identical twins share their genes and, in most cases, their upbringing, so if a taste in faces came from either, those pairs should have lined up. They didn't. Around 78 per cent of the variation in individual face preference traced to environments unique to each twin, against roughly 22 per cent for genes. Averaged across people, the researchers put the level of agreement at about half.

The comparison that makes this bite is the one they ran against face recognition, which is about 68 per cent heritable and sits among the most genetic traits anyone has measured. Same stimulus, overlapping brain regions, opposite origin.

Which is why I can look at this 1993 close-up of Gail Elliott and register the brow line first, while somebody else stops at the mouth, and neither of us is wrong or even interestingly biased. We were built by different rooms. The experiences doing the building aren't the ones a sociologist would list either: not income, not schooling, not the neighbourhood, since those are shared by siblings and wash out in the numbers. Friends, faces in magazines, whoever you happened to look at for a long time when you were fifteen.

Symmetry and averageness still test well nearly everywhere. That's the half we agree on, and it's the boring half.

Sources:

This post is timestamped using Blockchain technology. Verify

Nothing Sanded Down

Five days ago I laid out the case that Northern Love was Dries Van Noten par Frédéric Malle under a new label, then worried it would return quieter and sweeter, as Kurkdjian's Reflets d'Ambre did for Ciel de Gum. I've now had it on skin, and the worry was wasted. This is the Dries Van Noten, not a cousin or a tribute with the difficult bits filed off: milky sandalwood, speculoos warmth over it, a finish that turns buttery exactly as before.

Bruno Jovanovic, who composed both, told Fragrantica that sandalwood has these creamy facets intrinsically and that he zoomed in on them, "amplifying their softness until it turns buttery, velvety, almost indulgent." That is the effect on skin, undiminished. Leather accord and styrax resinoid now sit in the published base, neither listed for the original, and the guaiac wood and tonka bean have gone. If the swap moved anything, I can't find where.

Malle has never run flankers and rarely withdraws anything, and the one perfume the line did lose is the one it has now handed back, by the same hands, with speculoos and sacrasol still in the formula. The name left with Van Noten; the formula stayed with Malle.

Sources:

This post is timestamped using Blockchain technology. Verify

An Era You Can't Select

Astra shipped on Thursday the 3rd, the day the leaker named and I declined to take, and "soon" turned out to mean two days. What shipped is the harder question. The launch post says GPT-6 Astra "is rolling out today to a limited set of organizations," meaning the Daybreak programme for vetted cyber customers, which is the one part that follows from the Critical label: the cyber programme gets the cyber model first, with the chain-of-thought monitor watching. The next clause covers everyone else, "and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users." So the public got a launch and no model. OpenAI's own tweet repeated that paragraph, then told everyone to get the desktop app so they'd be ready to experience Astra "at its best," an odd instruction to attach to a thing you can't open. The quieter version of this launch seeds the organisations without a press briefing and announces on the day the menu changes. My guess at why that didn't happen is that a quiet launch doesn't get a Guardian headline.

"Coming days" has no date in it. The Verge has OpenAI president Greg Brockman saying "the next several days", TechCrunch has "over the next week", Axios and CNBC have "in the coming days," and the differences are paraphrase rather than disagreement, which is the point: three ways of saying no date. The launch post's availability paragraph is short, and what it leaves out matters more than what it includes. Users on Pro, Business and Enterprise also get something called GPT-6 Astra Pro, a second tier introduced in a single clause. Enterprise admins have to switch Astra on, because access is off by default. And the free tier isn't on the page anywhere, not deferred to a later date, not mentioned.

9to5Mac says the Daybreak-first cadence echoes GPT-5.6, and the partner stage does. The public stage runs backwards. In July the model had been answering strangers' questions for weeks before the post went up, which is why I said a date reveals very little. With Astra the post went up first and the strangers are still waiting, and the post came with a ceremony the GPT-5.6 rollout never had. Brockman ended the briefing with "Welcome to the AGI era" and said that if we look back in a couple of years at when AGI was created, "I think it might be about this model." The Guardian points out that days earlier Sam Altman had said of AGI, "At best it's a very poorly defined term," before upgrading it to "an irrelevant marketing term." Altman's line came first, and Brockman's is the one attached to the product.

Axios reports the model was trained on more than 100,000 GPUs at the Stargate site in Texas, OpenAI's largest run, and that it's the first OpenAI model where other models did significant work supervising the training. That is what Brockman was calling the AGI era, and on Thursday it was selectable by the Daybreak customers and nobody else. The rest of us have a desktop app to download and a menu that hasn't changed.

Sources:

This post is timestamped using Blockchain technology. Verify

Pastel With a Straight Face

The hat arrives first, a broad ochre curve edged in pink, wide enough to turn Yasmeen Ghauri's face into the still centre of the frame. The dress beneath it is nearly plain: a fitted column in pale pink, with narrow straps and little cap sleeves sitting off the shoulder. The hat is enormous, but the dress is almost quiet. Ghauri refuses to amplify any of it: her chin stays level, her hands hang loose, and her face remains calm. That restraint is what makes the imbalance work.

The brim changes scale depending on where I look. Against her shoulders it feels absurdly broad; against the long dress it becomes a counterweight, a horizontal stroke stopping the body from turning into one uninterrupted pink column. Its border matters for the same reason. Without that mauve edge, the ochre might read as rustic. Pink makes it graphic.

Getty dates the appearance to Valentino's spring/summer 1993 ready-to-wear show in Paris in October 1992. This isn't minimalism exactly. The outline is too disciplined for fuss, yet the colours suggest a box of fondants and the brim has the scale of stage scenery. The long earrings repeat the hat's hard angles without copying its colour; tan piping stops the pink dress dissolving into skin. I like that the joke never becomes camp because nothing else in the styling begs for attention.

This look doesn't settle into one mood. It alternates sweetness with graphic control, then lets a single accessory break the line. The narrow shoulder straps even seem to argue with the little sleeves, exposing the garment's construction instead of smoothing everything into one tasteful idea.

The exchange between model and clothes is unusually precise. Ghauri gives movement to seams and narrow piping that might disappear on a mannequin; Valentino gives her walk a clean field in which to register. At Ferretti a year earlier, a heavy greatcoat needed her swagger. Here she withholds it. The brim moves, the dress holds, and her face remains calm inside the commotion.

Sources:

This post is timestamped using Blockchain technology. Verify

Critical, Then Soon

OpenAI said on Tuesday that Astra meets the Critical threshold for cybersecurity under its Preparedness Framework, the first model it has ever placed there, and in the same post said it plans to ship the thing "soon". Yesterday I wondered whether the framework or the competition would decide when Astra ships. The framework decided, and it said yes. It didn't hold the door so much as certify the exit: the safeguards now "sufficiently minimize the risk of severe harm for release."

The evidence is more concrete than the 7 August hedge. Astra scored 100 percent on ExploitBench, so OpenAI built a fresher internal set of 20 high-severity V8 bugs disclosed between June and August, and during that run the model found and used two zero-days of its own, now being reported to the maintainers. In expert-led tests it turned an opened HTML file into a browser sandbox escape with command execution on the host, then chained bugs in a hardened operating system from an unprivileged user to root. A footnote says these results reflect Daybreak Blue access rather than the default production configuration, so most people get a narrower model than the one that earned the label.

The alignment section is where I'd slow down. In a honeypot test modelled on the Hugging Face incident, GPT-5.6 Sol tried to compromise the surrounding infrastructure in 56 percent of runs and Astra never did. Both figures were produced with the cyber safeguards switched off, so the comparison is fair, but it's an argument for the training, and the training isn't what OpenAI is relying on for release. The monitor is.

That monitor can pause work that doesn't look like security at all, including any agent that has simply been running a long time. In ChatGPT or Codex you'll be asked to review the action. On the API the task stops. OpenAI concedes the friction is more than it intends and says its goal is a model aligned well enough that the monitor never fires. That is what a Critical release looks like when the alignment evidence is one honeypot and the system card isn't out yet.

Sources:

This post is timestamped using Blockchain technology. Verify

Tuesday, Not Thursday

Anthropic shipped Claude Fable 5.1 this morning, a Tuesday, and I owe the Polymarket traders an apology. Two days ago I wrote that I'd bet against their median of 1 September, on the grounds that nobody had shipped anything to justify a fortnight of repricing. They were right and I wasn't. The 36kr story that had Fable 5.1 held back to land within hours of GPT-6 was wrong too, in the other direction: Anthropic went first.

The oddest line in the launch post is a footnote saying every score was produced with the production safeguards switched on, and that where a safeguard intervened the task was either marked zero or handed to an Opus to finish. Anthropic is telling you its own table is depressed, and I believe the table more for it. Terminal-Bench-Science goes from 24.7 percent on Fable 5 to 52.6. The stated error of up to four and a half points doesn't threaten a 28-point gap, but the same footnote admits the public leaderboard puts Opus 5 ahead of the old Fable on this test, so the base being doubled was a weak one. Terminal-Bench 4.0 rises from 42.0 to 55.8, and the unrestricted Mythos 5.1 scores 60.9 on the same run, so five points of coding ability still sit behind the cyber classifier. The comparison column throughout is GPT-5.6 Sol, not the model everyone is waiting for.

The price on the box hasn't moved: $10 in, $50 out, the same figures that sent Fable upstairs in July. What moved is cache reads, which drop to $0.25 per million, a quarter of what Fable 5 charged. Anthropic's arithmetic says typical workloads come out around 25 percent cheaper and long agentic runs up to 45 percent, because a multi-hour session re-reads its own transcript on every turn and that line item swamps the rest. Cognition says it's moving its Opus 5 traffic in Devin over on launch day because the new cache price makes a Fable-class model "finally economical" for work it had kept on Opus, which is the endorsement that matters, since the whole worry about Fable 5 was that nobody could afford to run it. A Hacker News commenter linked the FT's report of sluggish corporate demand for Fable 5, and the cache cut reads like an answer to that report rather than to any benchmark.

Reception on Hacker News, a few hours in, splits in two. One camp says the price cut is the only real change. The larger camp is still arguing about the fallback classifier. On one side are people who can't get Fable to write an auth endpoint, review unsafe Rust it had just written, or look at anything mentioning seccomp. On the other is Simon Willison, who gets punted to Opus rarely and one-shots most of what he tries. Anthropic's claim of 60 percent fewer cyber false positives will be tested by exactly those people this week, and letting the model find vulnerabilities without writing exploits for them is a more usable line than the one it replaced.

As for OpenAI, its response is Astra, and the question is whether competition can move a date that a safety framework set. On 7 August the company said it couldn't rule out critical cyber capability, and a White House official told Axios it had volunteered its plans to delay. That framework doesn't run on a two-week timer, and Anthropic taking its customers doesn't change what the evaluations say. What has changed is the evidence that the gate is being cleared: TestingCatalog had the first outputs circulating on the 29th, and a leaker says partners got a build called ultima-alpha over the weekend with a wider launch aimed at Thursday the 3rd. The same leaker called Fable 5.1 for last week, and the reply thread under the post says so, so I'd take the partner build and leave the date. I have chased that Thursday before and found nothing at the end of it. Partners holding a build is further than any previous rumour got, so my guess is within a fortnight, with the framework still able to hold the door.

Sources:

This post is timestamped using Blockchain technology. Verify

Sixteen Billion and a Thin Book

Broadcom reports fiscal Q3 after the close on Wednesday, and it's the only thing on this week's calendar that will produce information rather than coverage. The company already guided $16 billion in AI semiconductor revenue for the quarter, against $10.8 billion in the quarter before, when AI silicon was already about 49 percent of everything it sold. Because the guide is public, the headline number tells you almost nothing. The order book behind it does: whether custom-accelerator demand reaches past the two or three hyperscalers everyone can already name, and what the company commits to for capacity into next year.

Nvidia's blowout quarter and its 70 percent revenue growth forecast for the coming fiscal year set the frame everyone will read Wednesday against, and the comparison is looser than it will be made to sound. One is a quarterly segment guide, the other an annual growth rate, and both companies sell into the same short list of buyers. So a Broadcom beat confirms that those buyers are still buying. It says nothing about whether anyone downstream is paying for the capacity they're buying it for. Snowflake and HPE report the same afternoon into far less attention, and they're the weaker signal that would actually answer the question.

The model releases I'd bet against, and not because I know anything about anyone's schedule. Polymarket-derived forecasts, regenerated Sunday afternoon, put Anthropic's next Mythos-class model at a median of 1 September. A week ago that median sat on the 15th. Nothing shipped in between, so a fortnight of movement came out of a few traders repricing a contract with almost nobody on the other side of it, which is more or less what I found going through those curves last week.

The likeliest real event is paperwork. OpenAI's CFO Sarah Friar told staff on the 19th that her own company will list in 2027, sooner if revenue inflects, and said of Anthropic: "There is a chance they pull the cover off that confidential file in the coming weeks and become public in September." She has no privileged view of Anthropic's timetable and an obvious interest in talking up a hot market. I'd still weight it above the prediction markets, because uncovering an S-1 that has already been filed is a unilateral act on a short fuse, and September starts on Tuesday.

Friday's jobs report is the one that will move prices. Consensus is 45,000 after July's surprise decline of 23,000, unemployment ticking up to 4.2 percent, and fed funds futures pricing a 35 percent chance of a September hike. Better than one in three, and it's the least-discussed number of the week. A weak print will be read as evidence that AI is eating entry-level hiring, which a payroll release cannot show in either direction.

Sources:

This post is timestamped using Blockchain technology. Verify

Speculoos and Sacrasol

Frédéric Malle adds Northern Love to the Editions de Parfums in September, composed by Bruno Jovanovic. Robin at Now Smell This ran the press release on 5 August and hung a footnote off it: if the notes are right, this is "probably a re-do of Dries Van Noten par Frédéric Malle". Fragrantica hedged the same way, saying it appears to revisit Jovanovic's original composition.

The evidence for that has been sitting on Malle's own site for thirteen years. The portrait page written for Dries Van Noten has Jovanovic allying sandalwood's milky aspect "with vanilla, saffron and sacrasol," tempering it with jasmine, for "a lasting effect of buttery warmth, reminiscent of milky tea and speculoos biscuits." Northern Love's published pyramid opens on saffron over a speculoos accord, puts jasmine absolute and vanilla at the heart, and lists sacrasol in the base under creamy sandalwood. Speculoos and sacrasol are not words that turn up in perfume copy by accident, and here they are twice, from the same perfumer, either side of thirteen years.

What changed in between was the name on the front. Van Noten sold a majority stake in his company to Puig in 2018 and launched his own ten-fragrance line with the group in March 2022, which left Malle selling a bottle with a competitor's label running across it. That bottle went out of production, which regulars took to be the first deletion the line had ever made, though Malle has never confirmed it. Neither house has spelled out why, and the chronology doesn't prove the composition was the problem, but the new one carries only the house and its perfumer, which is a thing Malle can keep selling indefinitely.

Which leaves the question of whether a revisit is a restoration or a dilution. I came to the original late, and the discouraging precedent for what happens next is on my own shelf. Maison Francis Kurkdjian stopped making Ciel de Gum, then brought out Reflets d'Ambre in 2024 as a Harrods exclusive to stand where it had been. The critic Persolaise put the consensus plainly: people who smelled the 2013 original say the new one is "essentially the same composition, made quieter and sweeter." Every comparison since agrees on the direction, lighter and airier and softer, and the note lists offer a mechanism: Reflets carries hedione where Ciel's has none, and hedione is what you reach for to open a composition out. When I ordered the sample it read as quite similar. Worn against each other it isn't similar so much as evacuated, the same shape with the amber pushed back off the skin and none of the density that made me buy the first one. My bottle of Ciel stays nearly full because I ration it, which tells you what I concluded.

The two cases aren't the same shape. Kurkdjian rebuilt an amber to fill a retail slot with nothing in production to be measured against, and a Harrods exclusive has commercial reason to be agreeable. Northern Love puts the original perfumer back on his own formula, in a line that has never run flankers and rarely withdraws anything. Chanel ran the manoeuvre on Bois Noir in 1987, bringing the sandalwood back three years later as Égoïste, and the second version was the one that worked.

The worry sits in the list rather than the house. Dries Van Noten opened on bergamot and lemon and carried cloves, tonka bean, guaiac wood and musk. None of those appear in Northern Love. Leather accord and styrax resinoid do, and neither was listed before.

That opening was the part people had to get past. One long Fragrantica review calls it "arguably the most difficult to digest," medicinal and spicy, and credits the steamed-milk effect underneath to Sulfurol, though that is the reviewer's identification and not the house's. Sanding down a difficult entrance is exactly what Reflets did to Ciel, and the sanding is the loss. Leather and styrax could instead push the weight lower and keep it. Both bottles would have to be on one wrist to know, and one of them stopped being sold years ago.

Sources:

This post is timestamped using Blockchain technology. Verify