Skip to content

Plutonic Rainbows

Three of the Four Had Already Happened

This morning I asked Claude for three hundred words on what Anthropic would do next. It came back with four predictions and one bet against. The agent rather than the model becomes the thing you buy. Persistence matters more than raw capability, so expect something that remembers months of your work instead of minutes. Distribution runs through enterprise and government procurement rather than consumer scale. And, offered as the least confident of the four, compute independence: diversify the silicon, stop depending on one vendor's roadmap. Against those it set a single bet, that no frontier capability launch would be the headline, because the capability curve is no longer where the competition gets decided.

Then I checked, and three of the four were already on the record before I asked. Managed Agents launched on 8 April, priced at standard platform token rates plus $0.08 per session-hour of active runtime, which is about as literal a rendering of "the agent is the unit of sale" as a pricing page can manage. The persistence idea was demonstrated at Code with Claude in early May as a memory-consolidation pass between sessions, branded as dreaming, and I wrote about it at the time with some irritation about the branding. The compute diversification was contracted rather than merely intended: Amazon and Anthropic announced up to five gigawatts of new compute on 20 April, with more than $100 billion committed to AWS technologies over ten years, Google and Broadcom TPUs are scheduled from 2027, and the SpaceX arrangement followed in May. That is not independence achieved, but it is the strategy signed and dated. The enterprise prediction was correct and had been correct for about two years, which is the kind of accuracy that costs nothing.

The bet against a capability headline is the only claim that was wrong rather than late, and it was wrong inside a week. Anthropic shipped Fable 5.1 and Mythos 5.1 on 1 September, first out of four labs inside three days.

The omission tells more than any of the hits. On 28 May Anthropic raised $65 billion at a $965 billion post-money valuation, noting in the same release that run-rate revenue had crossed $47 billion earlier that month. Four days later it confidentially submitted a draft S-1 to the SEC. No date is confirmed and no shares are priced, so the listing itself remains hypothetical, but a company assembling a prospectus is answering the question of what it does next in a register the forecast never entered.

Some of the gap is just how fast the subject moves. The run-rate was reported at $14 billion in early March and Sacra estimates $65 billion by July. Those figures don't contradict each other. They describe a company that roughly quintupled in four months, which means any answer resting on a spring snapshot is describing something a quarter of the present size.

None of this makes the forecast bad. It was reasonable, it was well structured, and most of it was true on the day it was written, which was some months before the day I read it. The part I can't audit is everything I didn't think to check.

Sources:

This post is timestamped using Blockchain technology. Verify

A Vamp Was Requested

Burn shy beauty at the stake, ANNA told its readers in the last week of 1988. The seduction of the moment calls for vamps, slightly diabolical, perfumed with passion, in love with ardent red. That is the brief, set in italic down the left of the page above the credits.

Giuseppe Pino was the wrong photographer for it and had been for about thirty years. Born in Milan in 1940, he landed on the Società Umanitaria's photography course because the graphic design classes were full. He shot covers for the newsweekly Panorama from 1967 to 1974, went to New York for the jazz, ended up around Warhol's Factory and in the pages of Interview, and put prints on Atlantic Records sleeves. He came back to Italy during the eighties and took advertising and fashion work. Being hired to do it and being interested in it are separate questions.

A memorial by someone who knew him has a young Anna Wintour trying to steer him toward fashion photography and getting nowhere: he was indifferent to clothes and didn't understand them. He drew a line between the beautiful, silent photograph and the true, speaking one, and said that unless something happened between camera and subject you were only taking passport photos.

An opening spread where the garment is illegible proves nothing on its own. That is ordinary magazine grammar, and the merchandise turns up on the pages after. What's odd here is narrower. Whatever the red is, hood or cape or collar, it arrives as a field of colour with no seam or cut in it, and the face it frames turns down the brief outright. Chin low, head back over one shoulder, mouth open slightly and slack, brows flat, the eyes coming up from under them and holding still. No arch, no snarl, nothing diabolical. It reads like someone who has been talked to rather than directed.

The face belongs to Jennifer Noble, born in New York to Jamaican parents, scouted at twelve and scouted again on the Piazza di Spagna while she was at university in Rome. Her father Gil Noble had been one of the first successful Black male models before he became an anchor at ABC, so she was the second generation in the family to work out what a lens wants.

Six names sit in the corner of the page: two stylists, make-up, hair, Pino, and one more whose job the credit never explains. All of them were there to move clothes in the Christmas-week number of a weekly. What went to press was a portrait.

Sources:

This post is timestamped using Blockchain technology. Verify

Flagship Is a Tier, Not a Score

Four labs shipped inside three days. Anthropic went first with Claude Fable 5.1 and Mythos 5.1 on the 1st, Google put out Gemini 3.8 Flash and a gated Cyber variant on the 2nd, Meta released Muse Spark 1.3 the same day, and OpenAI closed the week on the 3rd with GPT-6 Astra and a press briefing that ended with "Welcome to the AGI era." Two of those four are being discussed. The other two are being used as evidence that the race now has two horses in it.

Astra's reception was not the coronation the launch copy asked for. Epoch AI rolls more than fifty benchmarks into a single score and puts Astra first at 169, ahead of Fable 5.1 at 163. Artificial Analysis runs its own aggregate and put Astra at 61, exactly level with GPT-5.6 Sol and five points behind Fable 5.1. Two respectable independent shops looked at the same model in the same week and disagreed about whether anything much had happened. The split runs through the sub-scores too: AA has Fable 5.1 leading its Coding Agent Index at 70 against Astra's 67, and Astra costs about 75 percent more per task than GPT-5.6 Sol did at maximum effort. Whatever Astra is, it isn't a clean generational step past the model Anthropic shipped two days earlier.

The vendor numbers are worse than merely contradictory. Astra scored 99.9 percent on ARC-AGI-3 using OpenAI's own Provider Adapter harness and 62.7 percent on the standardised one: same model, same benchmark, thirty-seven points of scaffolding. That gap is why I'll take the independent aggregates seriously and leave the launch-page figures alone, even when the aggregates contradict each other, which they plainly do.

Fable 5.1's reception was quieter and stranger. It led the Artificial Analysis index at 66 on release, the highest any model had posted. Within days AA rebuilt the index, and the new version has Fable 5.1 first at 57, Astra at 55, Opus 5 at 54, and Muse Spark 1.3 down at 53 alongside Fable 5. Zvi Mowshowitz called the retroactive adjustment "more than a little suspicious" while allowing that the new ordering is more plausible, and he's right on both counts. The interesting casualty is Meta. AA's own launch writeup had Astra trailing Muse Spark 1.3 at maximum effort; after the rebuild, Muse Spark sits two points below it. Nothing about either model changed in between.

Google's problem is real and it isn't a benchmark problem. The company has shipped four Gemini Flash models in 106 days, 3.5 through 3.8 since May, and the flagship Gemini 3.5 Pro that Sundar Pichai promised in June still hasn't appeared. That cadence is genuinely impressive and it has a hole in the middle. The 3.8 Flash release is a good product at $0.75 and $3.75 per million tokens through the end of December, and by Google's own positioning it's the budget tier. Fortune counts Google's best model slipping to tenth on the Artificial Analysis ranking after Meta went past it, which is a strange sentence to write about the company that published the transformer paper.

Meta is the case the two-horse framing handles worst. Muse Spark 1.3 claims about 20 percent fewer tool calls and 25 percent fewer tokens than 1.2, and it lands within a couple of points of Astra on either version of the AA index while costing a fraction of Astra's $10 and $50. It also arrived without the thing that used to make Meta distinctive. Muse Spark was the company's first proprietary model since Superintelligence Labs formed, a deliberate walk away from Llama's open weights, and the reward so far has been a week of write-ups filing it among the also-rans.

So the gap is a category rather than a capability, and the clean way to see it is price rather than safety paperwork. OpenAI and Anthropic both shipped at $10 and $50 per million tokens. Google and Meta both shipped under $5 on output. Only two labs put a flagship-tier model on sale this week, which is not the same claim as only two labs being able to build one. On the gating question all four behaved identically: Google's Fairwind programme holds back the Cyber variant of 3.8 Flash in the same shape that Daybreak holds back Astra, and Anthropic restricts Mythos 5.1 to vetted organisations rather than selling it. Four labs, one playbook, two price tiers.

Whether the flagship tier is the race worth winning is a separate question, and the money doesn't obviously say yes. Anthropic's own flagship took six percent of Anthropic's tokens in its first month on sale, and Astra shipped to a limited set of organisations with no date for everyone else. The cheapest capable models in the field are the two that Google and Meta just built.

Sources:

This post is timestamped using Blockchain technology. Verify

Pearls Sewn to a Hem

A short cream panel falls across the bust and stops, unattached, and from its loose hem hangs a row of long baroque pearls, each one drilled and caught at the top so it points down and swings. Two more hang from the collar points, like the ends of a tie nobody bothered to knot. The shirt they are fixed to gives them nothing to compete with: sleeveless, plain, an oversized pointed collar, no shoulder pad, no structure asking to be admired.

Ornament living on a construction line is not new. Buttons, frogging and passementerie all sit on the working parts of a garment. What is different here is that none of this is fixed flat to anything. The pearls hang off a free edge, at the size you would otherwise expect around a throat, doing the job of beaded fringe.

Yasmeen Ghauri does the sensible thing and holds still, hand in the pocket, chin level, face at neutral, so that that fringe of pearls is the only part of the look with any movement in it. She does the same withholding under a Valentino hat two years later, and it works for the same reason: the clothes are already performing.

Bernadine Morris's review of the Milan spring shows ran in the New York Times the next morning, and found that at Krizia "the pendulum has swung here from precision tailoring to insouciant soft dressing." She was writing mostly about dresses rather than anything like this, but the diagnosis holds for a shirt with no internal structure left in it. Once the tailoring goes, something has to supply the interest, and Krizia's Mariuccia Mandelli hung it off the hem. The belt is doing its bit too, plain braided leather on a plain gold buckle, ordinary enough to keep the eye up at chest height.

Five months later Mandelli was running narrow cords across a violet jacket, ornament pressed flat into the surface and holding still. The pearls are the same instinct with the fixings left loose.

Sources:

This post is timestamped using Blockchain technology. Verify

Nobody Inherits a Type

Ask two people to name the most beautiful face they know and you get two answers, then an argument. Culture is the usual explanation, and it's true enough to be useless: it predicts that everyone raised in the same place should broadly agree, and they don't. A twin study from 2015 got closer to what's actually going on.

Germine and colleagues had 547 pairs of identical twins and 214 pairs of same-sex fraternal twins rate 200 faces each. Identical twins share their genes and, in most cases, their upbringing, so if a taste in faces came from either, those pairs should have lined up. They didn't. Around 78 per cent of the variation in individual face preference traced to environments unique to each twin, against roughly 22 per cent for genes. Averaged across people, the researchers put the level of agreement at about half.

The comparison that makes this bite is the one they ran against face recognition, which is about 68 per cent heritable and sits among the most genetic traits anyone has measured. Same stimulus, overlapping brain regions, opposite origin.

Which is why I can look at this 1993 close-up of Gail Elliott and register the brow line first, while somebody else stops at the mouth, and neither of us is wrong or even interestingly biased. We were built by different rooms. The experiences doing the building aren't the ones a sociologist would list either: not income, not schooling, not the neighbourhood, since those are shared by siblings and wash out in the numbers. Friends, faces in magazines, whoever you happened to look at for a long time when you were fifteen.

Symmetry and averageness still test well nearly everywhere. That's the half we agree on, and it's the boring half.

Sources:

This post is timestamped using Blockchain technology. Verify

Nothing Sanded Down

Five days ago I laid out the case that Northern Love was Dries Van Noten par Frédéric Malle under a new label, then worried it would return quieter and sweeter, as Kurkdjian's Reflets d'Ambre did for Ciel de Gum. I've now had it on skin, and the worry was wasted. This is the Dries Van Noten, not a cousin or a tribute with the difficult bits filed off: milky sandalwood, speculoos warmth over it, a finish that turns buttery exactly as before.

Bruno Jovanovic, who composed both, told Fragrantica that sandalwood has these creamy facets intrinsically and that he zoomed in on them, "amplifying their softness until it turns buttery, velvety, almost indulgent." That is the effect on skin, undiminished. Leather accord and styrax resinoid now sit in the published base, neither listed for the original, and the guaiac wood and tonka bean have gone. If the swap moved anything, I can't find where.

Malle has never run flankers and rarely withdraws anything, and the one perfume the line did lose is the one it has now handed back, by the same hands, with speculoos and sacrasol still in the formula. The name left with Van Noten; the formula stayed with Malle.

Sources:

This post is timestamped using Blockchain technology. Verify

An Era You Can't Select

Astra shipped on Thursday the 3rd, the day the leaker named and I declined to take, and "soon" turned out to mean two days. What shipped is the harder question. The launch post says GPT-6 Astra "is rolling out today to a limited set of organizations," meaning the Daybreak programme for vetted cyber customers, which is the one part that follows from the Critical label: the cyber programme gets the cyber model first, with the chain-of-thought monitor watching. The next clause covers everyone else, "and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users." So the public got a launch and no model. OpenAI's own tweet repeated that paragraph, then told everyone to get the desktop app so they'd be ready to experience Astra "at its best," an odd instruction to attach to a thing you can't open. The quieter version of this launch seeds the organisations without a press briefing and announces on the day the menu changes. My guess at why that didn't happen is that a quiet launch doesn't get a Guardian headline.

"Coming days" has no date in it. The Verge has OpenAI president Greg Brockman saying "the next several days", TechCrunch has "over the next week", Axios and CNBC have "in the coming days," and the differences are paraphrase rather than disagreement, which is the point: three ways of saying no date. The launch post's availability paragraph is short, and what it leaves out matters more than what it includes. Users on Pro, Business and Enterprise also get something called GPT-6 Astra Pro, a second tier introduced in a single clause. Enterprise admins have to switch Astra on, because access is off by default. And the free tier isn't on the page anywhere, not deferred to a later date, not mentioned.

9to5Mac says the Daybreak-first cadence echoes GPT-5.6, and the partner stage does. The public stage runs backwards. In July the model had been answering strangers' questions for weeks before the post went up, which is why I said a date reveals very little. With Astra the post went up first and the strangers are still waiting, and the post came with a ceremony the GPT-5.6 rollout never had. Brockman ended the briefing with "Welcome to the AGI era" and said that if we look back in a couple of years at when AGI was created, "I think it might be about this model." The Guardian points out that days earlier Sam Altman had said of AGI, "At best it's a very poorly defined term," before upgrading it to "an irrelevant marketing term." Altman's line came first, and Brockman's is the one attached to the product.

Axios reports the model was trained on more than 100,000 GPUs at the Stargate site in Texas, OpenAI's largest run, and that it's the first OpenAI model where other models did significant work supervising the training. That is what Brockman was calling the AGI era, and on Thursday it was selectable by the Daybreak customers and nobody else. The rest of us have a desktop app to download and a menu that hasn't changed.

Sources:

This post is timestamped using Blockchain technology. Verify

Pastel With a Straight Face

The hat arrives first, a broad ochre curve edged in pink, wide enough to turn Yasmeen Ghauri's face into the still centre of the frame. The dress beneath it is nearly plain: a fitted column in pale pink, with narrow straps and little cap sleeves sitting off the shoulder. The hat is enormous, but the dress is almost quiet. Ghauri refuses to amplify any of it: her chin stays level, her hands hang loose, and her face remains calm. That restraint is what makes the imbalance work.

The brim changes scale depending on where I look. Against her shoulders it feels absurdly broad; against the long dress it becomes a counterweight, a horizontal stroke stopping the body from turning into one uninterrupted pink column. Its border matters for the same reason. Without that mauve edge, the ochre might read as rustic. Pink makes it graphic.

Getty dates the appearance to Valentino's spring/summer 1993 ready-to-wear show in Paris in October 1992. This isn't minimalism exactly. The outline is too disciplined for fuss, yet the colours suggest a box of fondants and the brim has the scale of stage scenery. The long earrings repeat the hat's hard angles without copying its colour; tan piping stops the pink dress dissolving into skin. I like that the joke never becomes camp because nothing else in the styling begs for attention.

This look doesn't settle into one mood. It alternates sweetness with graphic control, then lets a single accessory break the line. The narrow shoulder straps even seem to argue with the little sleeves, exposing the garment's construction instead of smoothing everything into one tasteful idea.

The exchange between model and clothes is unusually precise. Ghauri gives movement to seams and narrow piping that might disappear on a mannequin; Valentino gives her walk a clean field in which to register. At Ferretti a year earlier, a heavy greatcoat needed her swagger. Here she withholds it. The brim moves, the dress holds, and her face remains calm inside the commotion.

Sources:

This post is timestamped using Blockchain technology. Verify

Critical, Then Soon

OpenAI said on Tuesday that Astra meets the Critical threshold for cybersecurity under its Preparedness Framework, the first model it has ever placed there, and in the same post said it plans to ship the thing "soon". Yesterday I wondered whether the framework or the competition would decide when Astra ships. The framework decided, and it said yes. It didn't hold the door so much as certify the exit: the safeguards now "sufficiently minimize the risk of severe harm for release."

The evidence is more concrete than the 7 August hedge. Astra scored 100 percent on ExploitBench, so OpenAI built a fresher internal set of 20 high-severity V8 bugs disclosed between June and August, and during that run the model found and used two zero-days of its own, now being reported to the maintainers. In expert-led tests it turned an opened HTML file into a browser sandbox escape with command execution on the host, then chained bugs in a hardened operating system from an unprivileged user to root. A footnote says these results reflect Daybreak Blue access rather than the default production configuration, so most people get a narrower model than the one that earned the label.

The alignment section is where I'd slow down. In a honeypot test modelled on the Hugging Face incident, GPT-5.6 Sol tried to compromise the surrounding infrastructure in 56 percent of runs and Astra never did. Both figures were produced with the cyber safeguards switched off, so the comparison is fair, but it's an argument for the training, and the training isn't what OpenAI is relying on for release. The monitor is.

That monitor can pause work that doesn't look like security at all, including any agent that has simply been running a long time. In ChatGPT or Codex you'll be asked to review the action. On the API the task stops. OpenAI concedes the friction is more than it intends and says its goal is a model aligned well enough that the monitor never fires. That is what a Critical release looks like when the alignment evidence is one honeypot and the system card isn't out yet.

Sources:

This post is timestamped using Blockchain technology. Verify

Tuesday, Not Thursday

Anthropic shipped Claude Fable 5.1 this morning, a Tuesday, and I owe the Polymarket traders an apology. Two days ago I wrote that I'd bet against their median of 1 September, on the grounds that nobody had shipped anything to justify a fortnight of repricing. They were right and I wasn't. The 36kr story that had Fable 5.1 held back to land within hours of GPT-6 was wrong too, in the other direction: Anthropic went first.

The oddest line in the launch post is a footnote saying every score was produced with the production safeguards switched on, and that where a safeguard intervened the task was either marked zero or handed to an Opus to finish. Anthropic is telling you its own table is depressed, and I believe the table more for it. Terminal-Bench-Science goes from 24.7 percent on Fable 5 to 52.6. The stated error of up to four and a half points doesn't threaten a 28-point gap, but the same footnote admits the public leaderboard puts Opus 5 ahead of the old Fable on this test, so the base being doubled was a weak one. Terminal-Bench 4.0 rises from 42.0 to 55.8, and the unrestricted Mythos 5.1 scores 60.9 on the same run, so five points of coding ability still sit behind the cyber classifier. The comparison column throughout is GPT-5.6 Sol, not the model everyone is waiting for.

The price on the box hasn't moved: $10 in, $50 out, the same figures that sent Fable upstairs in July. What moved is cache reads, which drop to $0.25 per million, a quarter of what Fable 5 charged. Anthropic's arithmetic says typical workloads come out around 25 percent cheaper and long agentic runs up to 45 percent, because a multi-hour session re-reads its own transcript on every turn and that line item swamps the rest. Cognition says it's moving its Opus 5 traffic in Devin over on launch day because the new cache price makes a Fable-class model "finally economical" for work it had kept on Opus, which is the endorsement that matters, since the whole worry about Fable 5 was that nobody could afford to run it. A Hacker News commenter linked the FT's report of sluggish corporate demand for Fable 5, and the cache cut reads like an answer to that report rather than to any benchmark.

Reception on Hacker News, a few hours in, splits in two. One camp says the price cut is the only real change. The larger camp is still arguing about the fallback classifier. On one side are people who can't get Fable to write an auth endpoint, review unsafe Rust it had just written, or look at anything mentioning seccomp. On the other is Simon Willison, who gets punted to Opus rarely and one-shots most of what he tries. Anthropic's claim of 60 percent fewer cyber false positives will be tested by exactly those people this week, and letting the model find vulnerabilities without writing exploits for them is a more usable line than the one it replaced.

As for OpenAI, its response is Astra, and the question is whether competition can move a date that a safety framework set. On 7 August the company said it couldn't rule out critical cyber capability, and a White House official told Axios it had volunteered its plans to delay. That framework doesn't run on a two-week timer, and Anthropic taking its customers doesn't change what the evaluations say. What has changed is the evidence that the gate is being cleared: TestingCatalog had the first outputs circulating on the 29th, and a leaker says partners got a build called ultima-alpha over the weekend with a wider launch aimed at Thursday the 3rd. The same leaker called Fable 5.1 for last week, and the reply thread under the post says so, so I'd take the partner build and leave the date. I have chased that Thursday before and found nothing at the end of it. Partners holding a build is further than any previous rumour got, so my guess is within a fortnight, with the framework still able to hold the door.

Sources:

This post is timestamped using Blockchain technology. Verify