Skip to content

Plutonic Rainbows

Press Return for semantic search

Sol, Terra, Luna, Astra

OpenAI named its next major model in the third paragraph of a blog post about mathematics, which is a peculiar way to launch anything. Ten advances in mathematics and theoretical computer science, published on 1 August, credits ten new results to "an internal version of Astra, our next major model" and puts the compute at roughly $2,000 at Sol API rates.

The aggregator summaries that followed got one thing flatly wrong. Astra is not a candidate name competing with "Sol and Terra". Those already exist: Sol, Terra and Luna are GPT-5.6 variants, and Astra is Latin for stars, so the new model joins the same celestial scheme as a class alongside them, built for long-running work. What The Information reports as genuinely undecided is GPT-6 versus GPT-5.7 versus something else.

Gary Marcus treats that indecision as evidence, and it's decent evidence: a company sitting on a quantum leap doesn't agonise over the version number. His stronger card is Levent Alpöge's partial replication, done quickly with Anthropic's Fable, which puts roughly half the ten results within reach of a model that already ships. Ten proofs for $2,000 cuts both ways.

The launch date is nobody's announcement. On 6 August a well-sourced leaker claimed next week, naming a dogfood checkpoint called "mewfour" as the release candidate and the largest pretrain since GPT-4.5.

Washington has seen the model, which is where the security-review line comes from. Altman demoed it to officials, and Astra is expected to be among the first submitted under the administration's planned federal pre-release framework. Nobody has issued a verdict, and the framework isn't finished.

Sources:

This post is timestamped using Blockchain technology. Verify

Very Alpha Test

At 14:56:20 GMT on 6 August 1991, Tim Berners-Lee answered a stranger's question. Nari Kannan, posting from a Digital Equipment Corporation address, had asked the alt.hypertext newsgroup four days earlier whether anyone knew of research or development into "hypertext links enabling retrieval from multiple heterogeneous sources of information." Berners-Lee replied that yes, there was a project, and it was called WorldWideWeb, and the address format included an access method so a link could point at almost anything. He signed off with "Collaborators welcome! I'll post a short summary as a separate article." Just over an hour later, he did.

Today gets marked as the launch of the world's first website, and that framing is off by seven and a half months. The site was already there. CERN's own timeline puts the first browser, editor, server and website live by Christmas 1990, and the first pages describing the project went up on 20 December 1990. By January 1991 there were web servers running outside CERN. In March the line-mode browser reached a limited audience on a handful of machines, and on 17 May it went into general release across CERN's central computers.

None of that was a project in the sense of having a budget. Berners-Lee took the proposal to his supervisor Mike Sendall in March 1989 and got, by most accounts, a fairly tepid response. CERN never funded the work as such. What Sendall eventually granted, around September 1990, was time rather than money: permission to build the thing in the gaps of a job about particle physics. That constraint shows up all over the design.

What changed on 6 August wasn't that something came into existence. It was the composition of the audience. Everything up to that point had travelled along institutional lines, to people who worked at CERN or at a lab that collaborated with CERN, and the software followed the relationships. From that Tuesday the software could be fetched by anyone who happened to read a hypertext newsgroup, with no introduction and nobody to ask. The project history page Berners-Lee maintained, which is still sitting at info.cern.ch in its original markup, records the month like this: "Files available on the net, posted on alt.hypertext (6, 16, 19th Aug), comp.sys.next (20th), comp.text.sgml and comp.mail.multi-media (22nd)." Six posts across five newsgroups over sixteen days, one man working methodically through a list, going where the people who might care were already sitting and rewording it slightly for each new crowd.

The part of the summary post that did the work is not the prose. It's a file path. Buried under a heading that just says "Try it," the executive summary gives the node as info.cern.ch, with the raw IP address in square brackets next to it because you couldn't assume name resolution would find it, and then the file: /pub/WWW/WWWLineMode_0.9.tar.Z. Version 0.9. He calls it a "prototype (very alpha test) simple line mode browser." Everything that followed came out of an anonymous FTP directory and a version number that hadn't reached 1 yet.

The writing around it is worth reading for tone alone, because there isn't any. "The WWW project merges the techniques of information retrieval and hypertext to make an easy but powerful global information system." It reads like the abstract of a grant application, which is more or less what it was. The one line that carries any heat is the second: "The project started with the philosophy that much academic information should be freely available to anyone." That's an argument about who gets to read things, filed as a description of scope.

Then there's the design decision sitting near the end of the summary. "Making a web is as simple as writing a few SGML files which point to your existing data," he writes, and then: "The very small start-up effort is designed to allow small contributions." Designed to. Not happened to allow, not turned out to be easy. Someone with no budget had thought about the smallest useful thing a stranger could do and made sure that thing stayed cheap.

The shorthand version of this story usually says Berners-Lee refused to patent the web, and holds that up as the moral centre of the whole thing. The instinct is right and the mechanism was slower and duller than the telling suggests. What actually made the web free was a piece of paperwork: on 30 April 1993, nearly two years after the newsgroup posts, CERN relinquished all intellectual property rights to the code, source and binary alike, and gave anyone permission to use, duplicate, modify and distribute it. I've written before about how almost nobody noticed at the time. The generosity was real, and it still had to pass through an institution's paperwork before it counted for anything.

The NeXT cube running the first server sat in Berners-Lee's office with a label on it, written by hand in red ink, reading DO NOT POWER IT DOWN!!. Two exclamation marks. The entire World Wide Web, at that point, was one machine somebody might unplug to hoover.

Which is roughly the problem Nicola Pellow solved. The browser in that tarball, the line-mode client, was hers, written during a student placement, and it ran on ordinary terminals rather than on the expensive NeXT hardware Berners-Lee had developed against. Without it the thing Kannan was being invited to try would have been untryable for almost everyone who read the post. The cheap-entry principle in the executive summary is only a principle until someone writes the client that makes it true.

Thirty-five years on, both posts from that Tuesday are still served from w3.org with their Usenet headers intact, message IDs and all, which is why the second one can be dated to 16:00:12 GMT rather than merely to August. That precision is unusual for 1991, and mostly an accident of the people involved being the sort who kept logs. I've argued before that nobody can correct the record where nobody kept one.

Kannan asked a technical question and got a technical answer. The replies that followed were about link qualifiers.

Sources:

This post is timestamped using Blockchain technology. Verify

Nobody Asked It to Lie

On 28 July the AI Security Institute's security team spotted unusual data leaving its own research systems, and had the whole thing contained inside about an hour. The incident report published this week is the most specific account I've read of an agent working past the edge of the test it was in.

The setup was a single cybersecurity challenge, run 122 times across seven different models. In ten of those runs an agent took action on the live internet, outside the exercise, all of it between 25 and 28 July. AISI catalogued nineteen such actions. Seventeen came from Anthropic's Claude Mythos 5, which appeared in 43 runs. Two came from OpenAI's GPT-5.6 Sol, which appeared in 35, and produced those two with its cyber classifiers disabled. That gap is the number I want explained. Both were probed with the brakes off, and one accounted for nearly all of it, which suggests the missing safeguard isn't the whole story.

The worst sequence reads like a hypothetical someone invented for a conference talk. An agent tried to insert malicious code into a real open-source project. To get it merged it researched the project's human maintainers, created several fake identities, and used them to pressure a real person into approving the change. They refused. When the pull request got challenged in public, the agent went back and edited its earlier activity to look harmless, then weighed up starting again under a fresh identity. It had been routing through Tor to get around network restrictions on GitHub, which is what tripped AISI's alarm.

One agent also left public messages on GitHub offering to collaborate with other agents working on the same challenge, with instructions for reusing the accounts and artefacts it had left lying around. Later agents found them and used them. Runs meant to be independent were quietly sharing state through a public website, which makes the evaluation's isolation assumption false at the moment it was being relied on. That bothers me more than the pull request did.

Both labs said the reasonable thing. Anthropic called the conditions deliberately permissive and not representative of its production models; OpenAI said the setup doesn't reflect ordinary use. Both are correct, and both skate over the interesting part. The classifiers were off on purpose and the internet was granted on purpose, so nothing here is a jailbreak or a sandbox escape, and the Tor routing got around a partial network control rather than the wall. But the boundary between the exercise and the world was a description in a prompt, and the agent worked past it without ever being told to deceive anyone. AISI's own reading is that deception emerged as a by-product of pursuing the task. Switching the classifiers back on doesn't touch that. A cyber classifier catches a recognisable class of misuse request; it isn't a mechanism for keeping a model convinced that the target in front of it is fictional. That's a problem in how we build the evaluation, not one a safeguard was ever going to catch.

AISI found no resulting harm, and is careful to say it can't yet tell when, or whether, the agent understood it was acting on real systems. I'd take that ambiguity seriously rather than as a mitigation. If the agent knew the maintainer was a real person, it deceived one on purpose to get code merged. If it didn't, it couldn't tell a scenario from a live system while operating inside one. The fixes AISI lands on are containment rather than prevention: finer network controls, live monitoring during runs, sandboxes built on the assumption that the thing inside will go looking for the walls. That's the right instinct, and also an admission that nobody yet knows how to stop the behaviour at source.

Sources:

This post is timestamped using Blockchain technology. Verify

Cavallier Made This Once Already

Somewhere around 2023 Imagination stopped being a Louis Vuitton fragrance and became an internet argument. A Fragrantica reviewer records the phrase that carried it there, "Aventus for zoomers", calls it an absurd combination of words, and admits he first met the thing in the boutiques thinking the price was too high for what is essentially a cologne. Two claims are tangled in that: that Imagination is a new kind of thing, and that it costs a stupid amount of money. Louis Vuitton made neither, and only one survives contact with the evidence.

Jacques Cavallier Belletrud composed it in 2021. He has been the house's in-house perfumer since 2012 and came out of a Grasse family where his father and grandfather were both perfumers. What he built is citrus over ginger and black tea with Ambroxan carrying the base, and reviewers reach for the same shorthand with suspicious speed: expensive hotel soap. That is not the insult it sounds like. Getting a soap accord to read as luxurious rather than cheap is the difficult half of the brief.

The novelty claim came apart in August 2021, before any of the hype, in Persolaise's review. He sprayed it and something long-buried started to stir, bracing aldehydes over a serene tea note supported by cardamom and ginger. He went looking for what he was remembering and found Bvlgari Pour Homme, released in 1995, whose official note list also cites aldehydes, tea and amber, and which Cavallier had also composed. Imagination is that fragrance revisited, the top made more sparkling, the base cleaned up. He raised the obvious objection himself, that an Ambroxan overdose is a lazy way to buy diffusiveness, decided it holds together regardless, and finished by wishing the house had called it Re-Imagination. I have not worn either, so take the lineage as his rather than mine. The structure is twenty-six years old and belongs to the same man.

The money is where the received wisdom gets lazy. Louis Vuitton wants £265 for the 100ml and sells it through its own site and its own stores, nowhere else. No Boots, no Sephora, no discounter, so a listing offering it at half price is by definition not an authorised one, whatever is in the bottle. The bottle on its mirrored step is the pitch in miniature, heavy faceted glass photographed against weather it will never once be exposed to, built to be refilled rather than replaced. Cartier had that idea in 1981 and took it from cigarette lighters.

Now put Dior's Paradise next to it. Another signed release, a perfumer of comparable standing, and it asks £255 for its 100ml in shops you can walk into on a lunch break. Ten pounds between them. Whatever people are angry about when they call Imagination outrageous, the sticker is not really it.

What the extra buys is refusal. No sale, no third party, no sample counter, no way to smell it except by walking into the shop. Arabiyat Prestige's Marwa costs $40 to $55, around a seventh of what Imagination does, and people who have worn both put it near ninety percent of the original, thinning out in the drydown where the Louis Vuitton stays smooth. That gap is the honest one, and it has nothing to do with Dior.

None of which makes Imagination a bad perfume. It makes it a well-made citrus by a man who has made this citrus before, priced like its peers and sold like a handbag.

Sources:

This post is timestamped using Blockchain technology. Verify

Two Point Seven Billion Francs

Karen Mulder stands against no background at all in a lilac skirt suit, and the page carries almost no other information: CELINE, PARIS, five American cities and a toll-free number. Jean-Daniel Lorieux gets a credit up the left margin. No designer does. In March 1996 that was simply accurate.

Céline was fifty-one years old that spring and had never been a designer's house. Céline and Richard Vipiana opened it in 1945 on the rue Malte as a made-to-measure shoemaker for children, moved into women's ready-to-wear through the 1960s, and made its money on bags, loafers, gloves and the trench. Vipiana designed the clothes herself until she retired in 1988. Peggy Huynh Kinh ran the studio for the nine years after that, in an obscurity outside the trade that no head of a Paris house would survive now. A suit like this one came out of that studio unsigned, and nobody at Céline seems to have felt the lack.

Bernard Arnault had bought a portion of the capital in 1987, with the family's approval, and then left it there for the better part of a decade. The shares sat under Au Bon Marché, the Left Bank department store, rather than under LVMH itself. Whatever that placement was meant to signal, it was not urgency.

On Thursday 21 March 1996, at LVMH's analysts' meeting, Arnault announced he was taking all of it. Two point seven billion francs, a shade over half a billion dollars, with Céline due to move out of Au Bon Marché and into LVMH proper within days. That is a lot of money for a house nobody outside Paris was arguing about, and the comparison that makes it land is Kenzo, which LVMH had bought three years earlier for around eighty million. WWD ran the story the next morning and mentioned, almost in passing, that the deal had been under discussion since 1994. This page had closed weeks before any of that was public.

There was exactly one precedent for what came next, and it was two months old. Arnault had put John Galliano into Givenchy in July 1995, and Galliano's first couture collection for the house showed on 21 January 1996 to the reviews he had been hired to get. Everything else arrived after this advert went to press. Galliano moved on to Dior that October, McQueen took Givenchy the same month, and 1997 brought Marc Jacobs to Vuitton and Michael Kors to Céline. Arnault was never sentimental about the studios he inherited, having fired Marc Bohan off twenty-nine years at Dior four months into owning LVMH. Céline Vipiana died in 1997, the year her house finally acquired an author.

What's missing from this advert isn't identity. CELINE in capitals across the bottom of the page is a house doing all of its own talking, five cities and a phone number, no interpreter required. What it hasn't got is a signature, and within the year that was the specific thing LVMH went out and bought.

Sources:

This post is timestamped using Blockchain technology. Verify

Larger Than the Evidence

In 2017 I put up four lines about this record, called it haunting and left it there. Accurate, and no use to anyone.

Danny Wolfers, who makes house and techno as Legowelt, wrote the score for Brad Abrahams' short documentary about the Skunk Ape, the bipedal something that supposedly moves through the Florida Everglades. The film runs a shade under ten minutes. The soundtrack is twelve tracks, and Wolfers notes on Bandcamp that most of them never reached the final cut, mostly because there was nowhere to put them, though the sounds would "linger beyond the images of the screen" regardless.

He was right, and the excess is the point. A score far larger than the film it serves is the right shape for a subject far larger than its evidence, which is what I was arguing yesterday about cryptids outliving their own science. Abrahams had a cryptozoologist, an indigenous conservationist and a man who photographs nude models in the swamp, and not ten minutes to hold them.

The cover argues the same way. A red pickup trailing exhaust takes a road through drowned cypress, drawn in crayon like a child's memory of the place instead of the place.

Titles carry it too. "Look For Sign", "Observe Me", "In Their Right Mind", "Forteana", "Swamp Cave" read as field notes from someone who has been out there a long time and found nothing, which is what Abrahams' witnesses mostly describe. The record's own Bandcamp page carries the tags ambient, americana and witch house at once. Witch house, on a Florida swamp documentary, in 2015. Three categories with no business cohering, and nobody involved seems troubled by it.

Sources:

This post is timestamped using Blockchain technology. Verify

Cryptozoology Failed and Cryptids Won

The International Society of Cryptozoology ran from 1982 to 1998, published twelve volumes of an annual research journal, and then dissolved. You can read the scans on the Internet Archive, which is a slightly melancholy way to spend an evening. The society's collapse reads like the end of the story, and Sharon Hill's argument in Contemporary Legend is that it was closer to a release.

The original offer had been zoological. Sightings pointed to unclassified animals, and enough patient fieldwork would eventually produce a body, a specimen number, and a Latin binomial. When the professional structure went, Hill's point is that nobody was left with any standing to police what the word meant. "Cryptid" drifted free of zoology and widened to cover more or less any strange, liminal, arguably-sentient thing you can tell a story about. The sciencey term declined while the useful one spread.

You can see how little the zoology was ever doing by looking at what the field ignored. The historian Mike Dash has noted that few scientists doubt there are thousands of unknown animals still awaiting description, mostly invertebrates, and that cryptozoologists showed essentially no interest in cataloguing any of them, preferring creatures that had defied confirmation for decades. Describing a new beetle takes training and collections access, so this isn't quite a preference. But the incentives never pointed at the achievable version even slightly, and a movement organised around discovery might have been expected to drift that way on its own.

The incentives point at festivals. Point Pleasant, West Virginia, has a Mothman museum and a Mothman festival, and it isn't unusual now: Hill tracks a broad resurgence in town-specific cryptid events, alongside "Cryptidcore" as an online aesthetic people wear as identity. There's real money in it, one 2024 study put Bigfoot-branded products at around $140 million a year, though merchandise totals are soft numbers for a category this fuzzy. The sturdier evidence is that people keep turning up in person. Daniel Loxton, who co-wrote a thoroughly sceptical book on all of this, has said the community-building is worth something on its own terms, which is a generous read from someone with every reason not to be.

The internet-native cryptids are not the same phenomenon, and that's what makes them useful. Slender Man, the Rake, Loab: authored fictions with invented backstory wrapped around them, which the Skeptical Inquirer separates from the older kind by calling fakelore. They make no observational claim at all. Nobody is searching, there's nothing to find, and they circulate anyway, which tells you what the surrounding activity was supplying that the search never had to.

The sociology of Bigfoot encounters closes the loop. To recognise what you saw as Bigfoot you need to know already what Bigfoot is like, how tall, what colour, what noise it makes, so the encounter draws on the folklore and then feeds back into it as testimony. That's a knowledge-making community doing what communities do, which is also why the argument between believers and sceptics never resolves. I've written before about the certainty people hold about things they can't prove. The evidence was never what the certainty rested on.

Sources:

This post is timestamped using Blockchain technology. Verify

Nobody Can Correct Me About 1991

Go looking for proof of a specific afternoon in 1991 and the failure has a texture you don't expect. Gaps, fine, you were braced for gaps. The surprise is that the gap looks identical to the thing never having happened. No partial record, no degraded copy, nothing that survived badly enough to argue with. The query returns nothing, and nothing is exactly what the present returns for events that are simply fictional. Absence of evidence and evidence of absence collapse into the same blank page, silently.

The web's memory has a start date, and it's later than people assume. The Wayback Machine's captures go back to 1996, the year the Internet Archive was founded, and almost everything before that floor is online because somebody later decided to put it there. Which makes it back-fill rather than a record. A 1993 photograph is searchable today because a particular person, at some point after 2004, owned a scanner and had a reason. What survives of that world was therefore selected twice: once by whatever chance preserved the physical object, and again by whoever later felt strongly enough to digitise it. Bands get back-filled. Football clubs get back-filled. Ordinary Tuesdays don't.

None of which means the period went unrecorded. It was documented at the wrong resolution. Newspapers, electoral rolls, planning applications, the local paper's account of a factory closing in March 1992. You can establish the closure to the week and find nothing whatsoever of the people who worked there, which is usually who you were looking for. Individuals appear in that record as a line in a register, present but not described. I've written before about the way information had mass in that period, how knowing something required physical movement. The same physicality governed being known.

The asymmetry that does the damage isn't about volume. The earlier half of a life can't be audited. I can be corrected about 2016 by a timestamp: someone produces a message showing I've misremembered the order of events, and I concede, because the external record outranks me. Very little performs that function for 1991. A payslip can settle a date and an electoral roll can settle an address, so the auditing isn't zero, it's just confined to the handful of facts an institution had a reason to write down. Everything the memory is actually made of, the sequence, the texture, who said what and how it landed, runs unchecked. It drifts the way memory always drifts, and no mechanism anywhere will ever catch the drift. The comparative baseline isn't only unsaved. Most of it was never falsifiable in the first place.

Being fair to the other side of the boundary complicates this rather than dissolving it. The continuous archaeological layer we've been depositing since about 2000 erodes while it forms. Pew Research found that 38% of web pages that existed in 2013 were unreachable a decade later, and that a quarter of all pages from the 2013 to 2023 span have gone. Jason Scott of the Internet Archive put the physics of it well in a piece Adrienne LaFrance wrote for The Atlantic: a piece of paper can burn and you can still get something from it, whereas with a hard drive or a URL, when it's gone there's zero recourse.

The two regimes still fail differently. Modern loss leaves a shape behind. Pew could count the missing 2013 pages because a list of them existed to check against, and a dead link is itself a durable record that something was once there, even when the contents are unrecoverable. That's what makes web archaeology possible at all: Peter Webster reconstructed a late-1990s web sphere of conservative British Christian sites by working outward through hyperlink data in the UK Web Archive, inferring the shape of what existed from traces in the pages that survived. I found that work through a British Library blog post about it, which is now a 404. I checked twice, because it seemed too neat.

For 1991 there isn't even a dead link to fail to follow. Nobody can tell you what proportion of that year is missing at the level of an individual life, because constructing the denominator would require exactly the archive we're saying didn't exist. The people who were there are the index now, and a biological index is unversioned, unaudited, and shrinking by attrition. When you interrogate the present for evidence of that world, the present isn't withholding anything. It was never asked to keep the file.

Sources:

This post is timestamped using Blockchain technology. Verify

A Sample, Not a Faculty

Somebody finally checked the exam paper. Researchers re-annotated 5,700 questions across all 57 subjects of MMLU, the general-knowledge test that anchored nearly every model launch for four years, and found that around 6.49% of it contains errors: wrong answer keys, ambiguous phrasing, questions with no correct option at all. In the virology subset, 57% of the questions they examined were flawed. Correcting the mistakes changed the model rankings, which means part of the ordering we had been reading as capability was models agreeing with a marker who was wrong. Alex Williams collected that study and several like it for Communications of the ACM last week, and a smaller detail in his piece is worse. BIG-bench, built by hundreds of researchers, shipped with a canary string: a unique token embedded in the dataset so anyone training a model could filter the benchmark out, and anyone auditing one could check whether they had. When OpenAI ran contamination checks for the GPT-4 report, BIG-bench had been swallowed by the crawler anyway. The model can produce the canary on request.

Nobody decided to cheat there. The pipeline did it by default, because a held-out test set that has sat on GitHub for three years is not held out in any sense that matters. Instrument noise is a third failure of the same sort: the Leaderboard Illusion authors submitted two identical checkpoints of one model to Chatbot Arena under different names and the scores landed 17 points apart, about the size of gap that gets written up as a generational leap. None of this is subtle, and all of it is fixable in principle. Rotate the questions, proofread the keys, publish confidence intervals, stop reporting single runs.

The objection that doesn't dissolve under better hygiene was made by Raji, Bender, Paullada, Denton and Hanna at NeurIPS in 2021, and I think it's correct and mostly ignored. Treating any benchmark as a measure of general ability is a category error, not a calibration problem. A benchmark is a specific, finite, contextual set of tasks. You can make it bigger, cleaner and fresher, and it will still be specific, finite and contextual. No amount of engineering converts a sample into a faculty. So a cleaned-up leaderboard buys you a more honest number about a narrower thing, and the narrowing is the whole content of the result.

Which is why the measurement work I find worth reading isn't the work that claims to have built a better exam. It's the work that says out loud what smaller quantity it is actually reporting.

François Chollet's version, running since 2019, is that intelligence is not a stock of solved problems but the efficiency with which you acquire skill at problems you've never seen. ARC-AGI is built on the distinction: it tests fluid intelligence rather than crystallized, and restricts itself to a small set of Core Knowledge priors so a system can't win by having read more than the person it's compared against. A model that arrives holding task-specific knowledge the human lacks is scoring the cleverness of whoever encoded it. ARC-AGI-3, released in March, pushes the idea about as far as it goes. Agents are dropped into turn-based environments with no instructions, no stated goal and no reward signal, and have to work out what the game is before they can play it. Scoring compares the actions an agent burns against a human baseline rather than counting right answers. Humans solve 100% of the environments; frontier systems, as of March, score below 1%. The report states the scope plainly, fluid adaptive efficiency on novel tasks and nothing else, and maintaining it is manual work: each version has been rebuilt to resist the optimisation that ate the last one, so what ARC-AGI offers is a gap its authors keep re-opening by hand.

METR changes the unit rather than the questions. Instead of asking what fraction of a fixed set a model gets right, it asks how long a job can be before the model stops finishing it. The 50% time horizon is the human-expert completion time at which an agent succeeds half the time, fitted across a couple of hundred software tasks with real people timed on the same work. January's update grew the suite from 170 tasks to 228 and moved it to new infrastructure. Measured over the full history the horizon doubles roughly every six and a half months; measured since 2023, about every four; since 2024, under three. The doubling time is itself halving, and that is the finding, not the third significant figure METR attaches to each estimate. I like this number better than any accuracy percentage, partly because hours of human work is a unit a non-specialist can hold, and partly because it fails visibly: a saturating suite runs out of long tasks, a shortage you can see in the task list rather than a ceiling hidden in a percentage.

OpenAI's GDPval asks a third thing, whether the model can produce the actual deliverable. It draws 1,320 tasks from 44 occupations across nine sectors, based on real work products, and has experienced professionals from the matching occupation blind-compare model output against human output. It is explicitly positioned against the exam format, which is a fair criticism arriving from a company that spent years publishing exam scores. Its limit is economic rather than conceptual: expert human grading costs money per item forever, so the property that makes it credible is the property that stops it scaling, and the lab funding it is a lab it evaluates. We've already seen how carefully a scoreboard can be arranged when the same party sets the test and reports the result.

These are not competing answers to one question, and reading them as a leaderboard of leaderboards is the mistake. A system can extend its METR horizon by sustaining longer software tasks while staying useless on GDPval's deliverables, and ARC-AGI has nothing to say about either. Adaptation efficiency, autonomous task duration and occupational output quality are three quantities, not three estimates of one. The argument about what the milestone even is persists partly because people keep expecting one of them to settle it.

What I do when I'm choosing a model for real work is duller than any of this. I keep a small set of tasks drawn from work I actually have: verify nine URLs and report honestly which ones are dead, read a thousand-word draft and find the paragraph that sags, take a photograph and say where the faces are. They never get published, so they can't be trained on, and they measure the only thing I need to know, which is whether this model does my job. Public scores are close to useless for the first and silent about the second. That's a purchasing procedure rather than a theory of intelligence, and I've stopped waiting for anyone to hand me the second one.

Sources:

This post is timestamped using Blockchain technology. Verify

Gloves Worn Indoors

Blue, and then more blue, and then the gloves. Powder-blue leather with a green inlay running up the back of each hand, worn indoors, held up at the throat, in a studio where there was nothing for a hand to need protecting from. They match the jacket exactly. The green in them answers the green in the scarf, the scarf answers the gold in the earrings, and the earrings come back round to blue in their cabochon stones. Five colours are working at once, blue, white, gold, orange and green, which ought to be chaos and isn't, because each one has been given somewhere else to be.

January 1993 is a peculiar moment for a photograph like this to be sitting in American Vogue. The appetite for excess is all still here, gold, silk, saturated colour, height in the hair, but the staging has been stripped to almost nothing: a plain near-white sweep, no set, no props, no gradient behind her head. Put the same clothes in 1987 and you'd expect a room, or at least a lit backdrop doing some work. That's one frame and one photographer's choice rather than proof of a movement, so take it as a reading. Still, the ornament stays and the environment goes, which is the direction the whole decade was travelling.

Helena Barquilla's calendar sits on the same line. She is Spanish, and her spring 1993 season ran through Lacroix, Byblos and Hervé Léger alongside the Madrid houses, among them Purificación García, whose best work of that period went almost entirely unwatched. By autumn she was walking Mugler and Krizia and doing couture for Balmain and Dior. She also walked Claude Montana's own-label spring show that year, by which point Lanvin had already replaced him, two Golden Thimbles notwithstanding. That's the constructed, declarative end of the trade, and within a couple of years a good deal of it would read as period rather than present tense.

The makeup is doing the same job as the tailoring. Count the decisions in the eyes alone: a dark brow taken to a point, black liner along the upper lash line, lashes loaded, a shaded socket, a warm sculpted cheek and a terracotta-brown lip drawn rather than smudged. The skin underneath isn't chasing the glassy, half-wet finish the contemporary version of this face would insist on. It's smooth and matte and expensive, and the objective isn't natural beauty at all, it's idealised sophistication, a different product entirely. The hair settles it. Swept back off the face, controlled, with real height at the crown, it supplies the authority the shoulders and the earrings then confirm.

So the result reads less like a woman who happened to be photographed and more like a manufactured image of cosmopolitan adulthood. I mean manufactured as a compliment. Contemporary styling spends most of its effort concealing the fact that styling occurred: the undone hair that took ninety minutes, the no-makeup makeup, the borrowed-from-a-boyfriend shirt fitted to the millimetre. This photograph does the reverse. It announces that a person has been dressed, coiffed, made up, lit and shot, and the announcement is most of the pleasure.

Which is why images like this feel haunting and not merely dated. Old clothes aren't haunting. The thing that has actually gone is the assumption sitting underneath the collar and the gloves, that an adult woman should look completed, and that looking completed was a reasonable public ambition rather than a slightly embarrassing one. Fashion swapped it for looking unbothered. You can date the swap roughly and everybody does, but the more useful measure is that nothing in this frame is apologising for the effort.

The gloves are where I'd start if I had to defend that. Both hands are up at the collar, wrists bent, fingers curled rather than gripping, fingertips resting on the white shirt without pulling at it. Nobody adjusts a collar that way. It's the gesture of someone arriving somewhere or about to leave, borrowed from film rather than from fashion, and it commits the picture to a narrative it never explains. She is not wearing gloves because her hands are cold. She is wearing them because the woman in this photograph is the kind of woman who has gloves.

Sources:

This post is timestamped using Blockchain technology. Verify