Skip to content

Plutonic Rainbows

Press Return for semantic search

Critical, Then Soon

OpenAI said on Tuesday that Astra meets the Critical threshold for cybersecurity under its Preparedness Framework, the first model it has ever placed there, and in the same post said it plans to ship the thing "soon". Yesterday I wondered whether the framework or the competition would decide when Astra ships. The framework decided, and it said yes. It didn't hold the door so much as certify the exit: the safeguards now "sufficiently minimize the risk of severe harm for release."

The evidence is more concrete than the 7 August hedge. Astra scored 100 percent on ExploitBench, so OpenAI built a fresher internal set of 20 high-severity V8 bugs disclosed between June and August, and during that run the model found and used two zero-days of its own, now being reported to the maintainers. In expert-led tests it turned an opened HTML file into a browser sandbox escape with command execution on the host, then chained bugs in a hardened operating system from an unprivileged user to root. A footnote says these results reflect Daybreak Blue access rather than the default production configuration, so most people get a narrower model than the one that earned the label.

The alignment section is where I'd slow down. In a honeypot test modelled on the Hugging Face incident, GPT-5.6 Sol tried to compromise the surrounding infrastructure in 56 percent of runs and Astra never did. Both figures were produced with the cyber safeguards switched off, so the comparison is fair, but it's an argument for the training, and the training isn't what OpenAI is relying on for release. The monitor is.

That monitor can pause work that doesn't look like security at all, including any agent that has simply been running a long time. In ChatGPT or Codex you'll be asked to review the action. On the API the task stops. OpenAI concedes the friction is more than it intends and says its goal is a model aligned well enough that the monitor never fires. That is what a Critical release looks like when the alignment evidence is one honeypot and the system card isn't out yet.

Sources:

This post is timestamped using Blockchain technology. Verify

Tuesday, Not Thursday

Anthropic shipped Claude Fable 5.1 this morning, a Tuesday, and I owe the Polymarket traders an apology. Two days ago I wrote that I'd bet against their median of 1 September, on the grounds that nobody had shipped anything to justify a fortnight of repricing. They were right and I wasn't. The 36kr story that had Fable 5.1 held back to land within hours of GPT-6 was wrong too, in the other direction: Anthropic went first.

The oddest line in the launch post is a footnote saying every score was produced with the production safeguards switched on, and that where a safeguard intervened the task was either marked zero or handed to an Opus to finish. Anthropic is telling you its own table is depressed, and I believe the table more for it. Terminal-Bench-Science goes from 24.7 percent on Fable 5 to 52.6. The stated error of up to four and a half points doesn't threaten a 28-point gap, but the same footnote admits the public leaderboard puts Opus 5 ahead of the old Fable on this test, so the base being doubled was a weak one. Terminal-Bench 4.0 rises from 42.0 to 55.8, and the unrestricted Mythos 5.1 scores 60.9 on the same run, so five points of coding ability still sit behind the cyber classifier. The comparison column throughout is GPT-5.6 Sol, not the model everyone is waiting for.

The price on the box hasn't moved: $10 in, $50 out, the same figures that sent Fable upstairs in July. What moved is cache reads, which drop to $0.25 per million, a quarter of what Fable 5 charged. Anthropic's arithmetic says typical workloads come out around 25 percent cheaper and long agentic runs up to 45 percent, because a multi-hour session re-reads its own transcript on every turn and that line item swamps the rest. Cognition says it's moving its Opus 5 traffic in Devin over on launch day because the new cache price makes a Fable-class model "finally economical" for work it had kept on Opus, which is the endorsement that matters, since the whole worry about Fable 5 was that nobody could afford to run it. A Hacker News commenter linked the FT's report of sluggish corporate demand for Fable 5, and the cache cut reads like an answer to that report rather than to any benchmark.

Reception on Hacker News, a few hours in, splits in two. One camp says the price cut is the only real change. The larger camp is still arguing about the fallback classifier. On one side are people who can't get Fable to write an auth endpoint, review unsafe Rust it had just written, or look at anything mentioning seccomp. On the other is Simon Willison, who gets punted to Opus rarely and one-shots most of what he tries. Anthropic's claim of 60 percent fewer cyber false positives will be tested by exactly those people this week, and letting the model find vulnerabilities without writing exploits for them is a more usable line than the one it replaced.

As for OpenAI, its response is Astra, and the question is whether competition can move a date that a safety framework set. On 7 August the company said it couldn't rule out critical cyber capability, and a White House official told Axios it had volunteered its plans to delay. That framework doesn't run on a two-week timer, and Anthropic taking its customers doesn't change what the evaluations say. What has changed is the evidence that the gate is being cleared: TestingCatalog had the first outputs circulating on the 29th, and a leaker says partners got a build called ultima-alpha over the weekend with a wider launch aimed at Thursday the 3rd. The same leaker called Fable 5.1 for last week, and the reply thread under the post says so, so I'd take the partner build and leave the date. I have chased that Thursday before and found nothing at the end of it. Partners holding a build is further than any previous rumour got, so my guess is within a fortnight, with the framework still able to hold the door.

Sources:

This post is timestamped using Blockchain technology. Verify

Sixteen Billion and a Thin Book

Broadcom reports fiscal Q3 after the close on Wednesday, and it's the only thing on this week's calendar that will produce information rather than coverage. The company already guided $16 billion in AI semiconductor revenue for the quarter, against $10.8 billion in the quarter before, when AI silicon was already about 49 percent of everything it sold. Because the guide is public, the headline number tells you almost nothing. The order book behind it does: whether custom-accelerator demand reaches past the two or three hyperscalers everyone can already name, and what the company commits to for capacity into next year.

Nvidia's blowout quarter and its 70 percent revenue growth forecast for the coming fiscal year set the frame everyone will read Wednesday against, and the comparison is looser than it will be made to sound. One is a quarterly segment guide, the other an annual growth rate, and both companies sell into the same short list of buyers. So a Broadcom beat confirms that those buyers are still buying. It says nothing about whether anyone downstream is paying for the capacity they're buying it for. Snowflake and HPE report the same afternoon into far less attention, and they're the weaker signal that would actually answer the question.

The model releases I'd bet against, and not because I know anything about anyone's schedule. Polymarket-derived forecasts, regenerated Sunday afternoon, put Anthropic's next Mythos-class model at a median of 1 September. A week ago that median sat on the 15th. Nothing shipped in between, so a fortnight of movement came out of a few traders repricing a contract with almost nobody on the other side of it, which is more or less what I found going through those curves last week.

The likeliest real event is paperwork. OpenAI's CFO Sarah Friar told staff on the 19th that her own company will list in 2027, sooner if revenue inflects, and said of Anthropic: "There is a chance they pull the cover off that confidential file in the coming weeks and become public in September." She has no privileged view of Anthropic's timetable and an obvious interest in talking up a hot market. I'd still weight it above the prediction markets, because uncovering an S-1 that has already been filed is a unilateral act on a short fuse, and September starts on Tuesday.

Friday's jobs report is the one that will move prices. Consensus is 45,000 after July's surprise decline of 23,000, unemployment ticking up to 4.2 percent, and fed funds futures pricing a 35 percent chance of a September hike. Better than one in three, and it's the least-discussed number of the week. A weak print will be read as evidence that AI is eating entry-level hiring, which a payroll release cannot show in either direction.

Sources:

This post is timestamped using Blockchain technology. Verify

Speculoos and Sacrasol

Frédéric Malle adds Northern Love to the Editions de Parfums in September, composed by Bruno Jovanovic. Robin at Now Smell This ran the press release on 5 August and hung a footnote off it: if the notes are right, this is "probably a re-do of Dries Van Noten par Frédéric Malle". Fragrantica hedged the same way, saying it appears to revisit Jovanovic's original composition.

The evidence for that has been sitting on Malle's own site for thirteen years. The portrait page written for Dries Van Noten has Jovanovic allying sandalwood's milky aspect "with vanilla, saffron and sacrasol," tempering it with jasmine, for "a lasting effect of buttery warmth, reminiscent of milky tea and speculoos biscuits." Northern Love's published pyramid opens on saffron over a speculoos accord, puts jasmine absolute and vanilla at the heart, and lists sacrasol in the base under creamy sandalwood. Speculoos and sacrasol are not words that turn up in perfume copy by accident, and here they are twice, from the same perfumer, either side of thirteen years.

What changed in between was the name on the front. Van Noten sold a majority stake in his company to Puig in 2018 and launched his own ten-fragrance line with the group in March 2022, which left Malle selling a bottle with a competitor's label running across it. That bottle went out of production, which regulars took to be the first deletion the line had ever made, though Malle has never confirmed it. Neither house has spelled out why, and the chronology doesn't prove the composition was the problem, but the new one carries only the house and its perfumer, which is a thing Malle can keep selling indefinitely.

Which leaves the question of whether a revisit is a restoration or a dilution. I came to the original late, and the discouraging precedent for what happens next is on my own shelf. Maison Francis Kurkdjian stopped making Ciel de Gum, then brought out Reflets d'Ambre in 2024 as a Harrods exclusive to stand where it had been. The critic Persolaise put the consensus plainly: people who smelled the 2013 original say the new one is "essentially the same composition, made quieter and sweeter." Every comparison since agrees on the direction, lighter and airier and softer, and the note lists offer a mechanism: Reflets carries hedione where Ciel's has none, and hedione is what you reach for to open a composition out. When I ordered the sample it read as quite similar. Worn against each other it isn't similar so much as evacuated, the same shape with the amber pushed back off the skin and none of the density that made me buy the first one. My bottle of Ciel stays nearly full because I ration it, which tells you what I concluded.

The two cases aren't the same shape. Kurkdjian rebuilt an amber to fill a retail slot with nothing in production to be measured against, and a Harrods exclusive has commercial reason to be agreeable. Northern Love puts the original perfumer back on his own formula, in a line that has never run flankers and rarely withdraws anything. Chanel ran the manoeuvre on Bois Noir in 1987, bringing the sandalwood back three years later as Égoïste, and the second version was the one that worked.

The worry sits in the list rather than the house. Dries Van Noten opened on bergamot and lemon and carried cloves, tonka bean, guaiac wood and musk. None of those appear in Northern Love. Leather accord and styrax resinoid do, and neither was listed before.

That opening was the part people had to get past. One long Fragrantica review calls it "arguably the most difficult to digest," medicinal and spicy, and credits the steamed-milk effect underneath to Sulfurol, though that is the reviewer's identification and not the house's. Sanding down a difficult entrance is exactly what Reflets did to Ciel, and the sanding is the loss. Leather and styrax could instead push the weight lower and keep it. Both bottles would have to be on one wrist to know, and one of them stopped being sold years ago.

Sources:

This post is timestamped using Blockchain technology. Verify

Thursday Was Ten Days Ago

Nobody has put a source on next Thursday, and nobody put one on the Thursday before it. Two questions circulate about Astra: whether it ships on a named day, and what Anthropic does when it finally does. The first is answerable, and the answer is that the day doesn't exist.

A specific weekday attached to an unreleased frontier model is either a leak or a game of telephone, and this is telephone. Chase the word back through the coverage and it arrives at a video published on Wednesday 19 August, titled "GPT-6 Astra Releasing Thursday". Its own description is markedly softer, saying Astra "could be arriving as soon as Thursday" on the strength of X posts by OpenAI employees. The title hardened what the description hedged, the whole mechanism inside a single artefact, and the Thursday in both was the next day. Before that, on the 6th, a leaker named a release candidate and said Astra was days away, and by the time I wrote that up on the 15th nothing had shipped. A fresh date arrives about as fast as the last one expires.

Set against that, the company's own record runs to three posts. The 1 August research post named Astra "our next major model" and credited an internal version with ten long-open problems in mathematics and theoretical computer science. On 7 August it said its evaluations meant it "cannot rule out critical cyber capabilities" under its Preparedness Framework. On the 18th it said a significant number of workloads remain paused, which reads to me as containment rather than restraint. Not one of the three carries a date, a price or a model ID, and a White House official told Axios that OpenAI had volunteered its plans to delay.

The capability underneath the rumours isn't in doubt, which is most of why they regenerate. The Lean repository for those ten proofs reports a sorry count of zero, meaning no step in any of them is left unproven, and it's a compiler saying so rather than a press office.

TestingCatalog reported on 29 August that internal testing has expanded around a checkpoint called mozaik-alpha-fdm, with the first outputs circulating publicly, while noting OpenAI has confirmed neither the codename nor a date. Real evidence of preparation. The betting markets priced on 18 August gave 18% for a launch by end of August; with the month nearly out and no model, that leg resolved the way the pattern predicts. The same book had 60% by 15 September, a date that keeps sliding rightward.

Anthropic's answer was written in February. The same Axios piece notes it had committed to pausing training if capabilities outran its ability to control them, then removed that in a Responsible Scaling Policy update, on the reasoning that one developer pausing while others "moved forward training and deploying AI systems without strong mitigations" could leave the world "less safe". That commitment covered training rather than launches, so it doesn't map straight onto a deployment hold. The reasoning generalises, though: Anthropic has written down that it won't be the lab that stops while a competitor keeps moving.

So a slowdown is the one response the framework rules out. On timing I have less. One unconfirmed report from the Chinese outlet 36kr has Anthropic's Fable 5.1 finished and held to ship within hours of GPT-6, and by the standard I've applied to every Thursday here, it doesn't clear the bar either. It doesn't have to. The doctrine gets there without it.

Sources:

This post is timestamped using Blockchain technology. Verify

I. Magnin, Top Left

Two names share this page and they are not the same kind of name. Down the right edge, letter-spaced tall enough to read from across a newsstand: ESCADA BY MARGARETHA LEY. Top left, small, serif, doing none of the shouting: I. MAGNIN. One of them made the clothes. The other was where you could buy them.

That top-left line is cooperative advertising. The brand supplies the photograph, a retailer's name goes in the corner, and the two of them split the cost of the page between them. So the same plate can run in a dozen markets with a dozen different shops named on it, which is easiest to see when a label doesn't bother keeping the list to one: Kenar ran a page the year before with a whole row of stores set along the bottom, Bloomingdale's and I. Magnin among them. This page carries a single store, and that is the only variable on it.

Everything Escada supplied is deliberately from nowhere. Gail Elliott crosses a white studio in olive houndstooth over black patent, opaque tights, flat shoes, a bowler tipped forward and a quilted tote you could pack for a weekend, the house emblem stamped into the leather. Her eyes have gone sideways, off the lens and out of the frame entirely. Heavy autumn wool, sold to you in July. No street, no city, no weather, no season matching the one you were reading in.

Which leaves those two words in the corner doing all the locating, and in 1991 they were still carrying real civic weight in San Francisco at least. Mary Ann Magnin founded the store in 1876 and put her husband's initial on it while he peddled the baby clothes she sewed. The Union Square flagship opened in 1948, ten storeys of white marble by Timothy Pflueger, and when Christian Dior toured it he called it the White Marble Palace.

The family had been out of it for nearly fifty years by the time this ad ran. Bullock's bought them in 1944, Federated absorbed Bullock's, and Macy's took the Magnin stores in 1988, four years before filing for bankruptcy protection and beginning to shut them. In November 1994 Federated and Macy's announced the chain would be wound up because it "is not expected to make a meaningful contribution to the combined company's bottom line." A grandson of the founder offered to buy the twelve remaining stores and was refused. Union Square closed on 8 January 1995 and the Chronicle columnist Herb Caen wrote the elegy: "The interior was dazzling, especially the great main hall, two stories high, with its Lalique light fixtures, the gold ceilings, the glass murals, the expensively made cases."

The name down the right edge outlived its author. Margaretha Ley died in June 1992, eleven months after this page ran, and the house spent the rest of the decade insisting on the page that nothing had happened before filing for insolvency in 2009. Seventeen more years of trading on a dead woman's signature, which a brand can do, because a brand is a claim and claims can be maintained. The line in the top left corner had none of that in it. It was an address, and an address is the first thing a merger deletes.

Sources:

This post is timestamped using Blockchain technology. Verify

Madrid 1846, Londres, Tokio

The bottom of the page carries the addresses. Madrid, Barcelona, Bilbao, San Sebastián, Zaragoza, Valencia, Sevilla, Granada, Palma de Mallorca and Las Palmas de Gran Canaria, then Londres, New York, Tokio, Hong Kong, Singapore, set in small caps under the four-L anagram and the words Madrid 1846. Ten Spanish cities and five that aren't, the foreign ones spelled for a Spanish reader. A printer's legal deposit number down the gutter dates the sheet to 1988.

That date tells you more than the 1846 does. In 1984 Louis Urvois and Gian Luca Spinola bought the majority of Loewe's shares. In 1986 LVMH bought the rights to its international distribution. Through the same decade the womenswear was designed by Giorgio Armani and Laura Biagiotti. So the Spanish leather house selling itself in these pages was majority-owned outside the family, distributed abroad by a French group, and dressed by two Italians. It had been in Hong Kong since 1976 and New York since 1983, and the Japanese money that was about to reorganise European luxury was arriving while the campaign ran. Enrique Loewe Lynch set up the Loewe Foundation that same year, poetry prize included.

The pictures were made to travel too. Peter Lindbergh photographed Yasmin Le Bon for Loewe from autumn/winter 1987 through autumn/winter 1991, season after season, and the pages print like contact sheets rather than advertisements: the film's ragged edge and sprocket holes left in, the copy cut back to LOEWE and the season in Spanish. Almost all of it is black and white. This frame isn't. She's at a café table with the water glass going soft behind her hand, in a black jacket with a stiff white collar and cuffs, chin propped, watching something off to the left. The identical shot sits in the archive in monochrome, labelled 88-89. Two versions of one negative, and I can't tell you which title got which.

Then look at who is in it. Le Bon was born in Oxford to an Iranian father and an English mother, had shot Guess the year before, and carried a surname half of Europe could hum. Lindbergh was German. The streets in the rest of the run read as Paris. Nothing in these pages is Spanish except the leather and the word colección, which sits oddly against a footer insisting on ten Spanish cities. She was doing the same job elsewhere that autumn: on page 72 of American Vogue for October 1988 she supplied an imported Englishness to a Canadian label with no claim on it either.

One page in the run breaks the pattern and is the most interesting thing in it. No model, no street: a tight crop of a short leather jacket, the anagram embroidered onto a napa glove, and four lines of Spanish product copy about velvet buttons and a matelassé collar. Chaqueta corta de inspiración austriaca. The single page that sells the making of the thing describes it as Austrian.

Sources:

This post is timestamped using Blockchain technology. Verify

Six Percent of the Tokens

Ramp went looking for the money and didn't find it. In Fable 5's first month on general sale, businesses bought six percent of their Anthropic tokens from it and spent 11.4 percent of their Anthropic dollars there. The gap between those two numbers is the price doing its work: at $10 per million input tokens and $50 on output, Fable runs double Anthropic's other flagship tier, so a thin slice of usage swells into a fatter slice of the bill. Even swollen, it's thin.

The cross-vendor figure is the one that stings. Fable generated roughly three-quarters as much model-attributed spend in July as OpenAI's GPT-5.6 Sol, while charging about twice as much per token. Sol also took a far larger share of its own house's usage, 25 percent of OpenAI tokens against Fable's six, though those are shares of two different totals and shouldn't be read as a like-for-like volume comparison. Ramp's own caveat matters more anyway: its sample of 70,000 businesses skews tech-heavy, so adoption across the wider economy is probably thinner than six percent rather than fatter.

Price is the easy read and it's mostly the right one. Anthropic shipped Opus 5 on 24 July at $5 and $25 per million tokens, half of Fable on both sides, and said it beat Fable on coding and knowledge-work evaluations while being "designed to be used every day." That is a company telling its customers, in the politest available language, that the flagship isn't the default. The best public cost comparison I can find doesn't quite test the right thing: the consultancy ML6 ran Fable against Opus 4.8, a generation back, and found Fable caught release-blocking issues that 4.8 missed at about three times the wall-clock and 3.4 times the money, $12.38 against $3.65. Against Opus 5 the price premium would be narrower and the capability gap narrower still. That is the comparison every buyer has actually been running since late July, and it's the one nobody has published.

The consumer side ran the same argument faster and in public. Anthropic bundled Fable into subscription limits, announced a move to metered usage credits on 7 July, pushed the deadline to the 12th, pushed it again to the 19th, and then on 20 July put Fable back into Max and Team Premium at half of plan limits while leaving Pro on credits with a one-time hundred-dollar sweetener. Three schedules in thirteen days is a company finding out in real time what its subscribers will tolerate paying for the best thing it makes.

None of which is visibly hurting Anthropic. It recorded its first adjusted operating profit in the second quarter, guided investors toward another in the third, and reached roughly $65 billion in annualised revenue by July. One tier underperforming is a product problem, not a solvency one.

Which is where phones come in. The average American handset was 3.16 years old when it got traded in during 2020, and by 2025 that had stretched to 3.84 years, with the global cycle walking up from 2.4 years in 2013 to 3.7 by 2022. A longer cycle on its own doesn't prove that new phones stopped feeling better, so the more useful number is the reason people give when they finally do upgrade: three quarters of them say it was the battery. Not the camera, the processor, or anything a keynote gets built around. Once the dominant trigger for replacing a device is a physical component wearing out, the feature improvements have already stopped doing the work of selling it. They justify the purchase after the wear has forced it.

Ramp's economist reads the Fable numbers as a ceiling, saying the company has "found a new upper bound for how much businesses are willing to spend on AI." I think that's slightly the wrong shape. The data shows a pricier, more capable model taken up more slowly than a cheaper one. It doesn't establish a budget cap, because total enterprise AI spending kept climbing through the same period and the share of Ramp businesses paying for AI at all went from just over half in March to nearly 56 percent by July. Buyers will pay for a gain they can see in a number they already keep. Eleven percentage points on SWE-Bench Pro is a real difference and it is invisible in every metric a finance team tracks, which is the same trap the mini tiers exposed back in March.

The analogy has a limit, and working through it changes the answer rather than softening it. A phone is one visible purchase every few years and the comparison is trivially easy: this handset, that handset, this price. Tokens are metered continuously, and the choice gets made per task by an engineer who mostly never sees the invoice, which ought to make reaching for the expensive model easier rather than harder. Except that Fable carries friction no handset does. Anthropic set the safety classifiers deliberately wide, writing that it configured them "to trigger on a set of requests that we know are likely benign" with a margin "much larger than in any prior launch," and that users experience this as the model refusing reasonable, non-harmful requests. Sitting on top of that is a mandatory thirty-day data retention policy with no configuration toggle and no enterprise carve-out. The engineer isn't unconstrained after all: one of those constraints lives in the code path and the other lives in the contract.

So the six percent is overdetermined, and I'd rather say that than pick the tidier cause. Price, refusals and retention all push the same direction, and this data can't cleanly separate them. What the phone comparison adds is the part that survives all three: strip the frictions out entirely, and a capability gain nobody can feel still doesn't move a purchase, which is roughly what the NBER found when a flood of new software moved no usage at all.

There is an AI equivalent of a dying battery, and the labs own the schedule for it. Nothing wears out in a model. The thing that eventually forces the upgrade is the old tier being retired, deprecated, or quietly repriced out from under the people who built against it. On the current numbers that lever moves more adoption than shipping something better does, which is an odd place for a research company to arrive.

Sources:

This post is timestamped using Blockchain technology. Verify

Nobody Is Pricing Anthropic

Thirty days since Opus 5, against twenty-one between Fable 5 and Sonnet 5 and twenty-four from there to the Opus release on 24 July. One tracker puts the house average at about every 24 days and still shows its next projection as 17 August, six days expired and never recalculated. That is about what cadence models are worth once a lab decides to sit on something.

The Polymarket-derived forecasts, regenerated today, put Anthropic's next Mythos-class model at a median of 15 September and Astra at 19 September. Anthropic first, by four days, which runs backwards to the story everyone has been repeating, me included.

Then I looked at the curves rather than the medians. Astra's reaches 90% by the end of October. Anthropic's reaches 55% and stops, having crossed the halfway line on 15 September and gained five points between then and the end of October. That is not a forecast of a release date. It is a book that runs out of people willing to price it, and there is no contract at all on the model the reporting is actually about.

So: late September, OpenAI first, Anthropic within days of it. I'm backing the leak over the aggregate, mostly because on Anthropic's side there is no aggregate to argue with.

Sources:

This post is timestamped using Blockchain technology. Verify

Cheap to Make, Hard to Want

Economists at the National Bureau of Economic Research went looking for the software that agentic coding was supposed to produce, and they found it. Their working paper on writing code versus shipping code tracks four app marketplaces through what the authors call the agentic-coding era, and the count of new applications climbs steeply, exactly as advertised. Then they checked whether anyone opened any of it. Total usage in the first three months after launch had not increased in any of the four marketplaces, and the share of new apps failing to reach even a modest audience had risen. The abstract describes this as large productivity gains translating only partially into shipped and used software. The graphs underneath show flat or falling usage. The engineer Yuchen Jin put it at human scale in June: before AI he could waste a weekend building one useless app, and now he can build sixty-seven, each with a logo, a landing page, and nobody using it.

The usual diagnosis is that the output is bad. Some of it is, but plenty of vibe-coded software works fine and still goes unopened, so bad code can't be the mechanism. Making was never the expensive part. Getting someone to want the thing was expensive, and keeping it alive afterwards was expensive, and neither of those got cheaper by a penny. Hendrik Erz calls the result slopware and names the two jobs no agent does: maintenance and user experience. By his estimate about 90% of a developer's work was never typing, and it's the 90% that decides whether anyone comes back on day two.

The defence I hear most is that AI doesn't produce slop, engineers do. True, and useless. Erz quotes the pitch in its purest form, that we now live in an era where if you can dream an app you can probably build it, and calls it horribly wrong. The difficulty of building was quietly doing other work. Spend three months on something and you find out around week two whether you still believe in it. Take the three months away and nothing else is checking.

The art half runs on the same arithmetic. A study Brian Merchant cites put the median cost of fine-tuning a model on a professional author's corpus at $81 per author, a 99.7% saving against paying the writer. What that buys isn't work anyone seeks out. It's what he calls the slop layer, the material you scroll past on a load screen or nod along to in a Discover Weekly mix without ever learning whose name is on it. OpenAI reached the same verdict faster than the critics did, and shut Sora down to move the GPUs to coding agents, after Disney walked away from the partnership.

Nathan Schneider now hears about a new app at least once a week, usually with the familiar tells, and his reading is the one I'd defend. Apps got too cheap to serve as the coin of the realm. I use these tools every day and this site carries their fingerprints all over it, so I'm not in a position to draw the line anywhere flattering.

Sources:

This post is timestamped using Blockchain technology. Verify