Skip to content

Plutonic Rainbows

Press Return for semantic search

Bias Doesn't Hold Still

Someone at iTutorGroup configured the company's application software to reject women aged 55 and over and men aged 60 and over the moment their birthdate hit the form. The EEOC sued, and in September 2023 the company paid $365,000 to settle. No neural network, no training data, nothing anyone would call artificial intelligence, and that is exactly why it remains the clearest case in the field: the discrimination was legible. You could read it off a config screen and point at it in court. Nothing since has been that easy to read, which is the actual problem with handing hiring to models.

The story most people carry around is the Amazon one. Train a screener on a decade of your own hiring, find your own historical preferences reflected back, quietly kill the project. It fits the intuition that machines launder the past, and for a while the research agreed. Brookings researchers simulated resume screening with large language models and found significant gender and racial discrimination, concentrated on Black men: in their setup, resumes with Black male names were preferred over white male names zero percent of the time. Five leading models tested through VoxDev showed the same intersectional pattern, where Black women, Black men and white women each get sorted differently, so a fairness test built on race alone or gender alone reports nothing.

Then the sign changed. A paired-resume audit of fourteen models, modelled on the classic correspondence experiments and running 24,024 matched job postings per model, found that GPT-3.5-turbo reproduced the human pro-white callback gap at 2.12 percentage points, and that every model released in 2024 or later showed either no gap or a statistically significant reversal favouring Black-coded names, by as much as three points. The same flip appeared on the gender axis. A separate study across twenty-nine models found that language models now advantage female and Black candidates relative to comparable white and male ones while disadvantaging disabled candidates, the effect of demographic identity worth somewhere between six months and a year of extra education. Post-training alignment, not the pretraining corpus, turned out to be the main driver.

I don't think that's good news. A screener that favours one group over another fails the same test whichever way it leans, and a bias that reverses between model generations is harder to govern than one that doesn't, because it can't be certified. It moves with vintage, with provider, with whatever the alignment team shipped last quarter. New York City's Local Law 144 requires an annual bias audit of automated employment tools. The EU AI Act classifies hiring systems as high-risk. Both regimes were designed around the Amazon story, where bias is durable, inherited, and points where you'd predict. An annual audit of something whose sign changes between releases is a photograph of a moving object, and the auditors are mostly reading figures the vendor produced, a weakness familiar from self-reported benchmarks.

The comparison that matters, though, isn't against a fair process. It's against human recruiters, and the human baseline is dreadful. It has been measured since the 2004 correspondence study in which Bertrand and Mullainathan sent fake CVs to Boston and Chicago employers and watched white-sounding names collect fifty percent more callbacks. In a field experiment covering seventy thousand applicants, AI-led interviews produced twelve percent more job offers, eighteen percent more people actually starting work, and better thirty-day retention. Those are outcomes, on real hires, and anyone defending the status quo has to account for them. The fairness number from the same study needs more care than it usually gets: reported gender-based discrimination fell from 5.98 percent to 3.30 percent, and "reported" is carrying the sentence. That is what candidates said about their experience, not a count of who got selected, and a machine interviewer can feel less prejudiced while sorting people exactly as badly. The hiring and retention gains I'd take at face value. The fairness gain I'd want measured a different way.

The catch is what happens when you put a person back in the loop, which is what every compliance framework asks for. In a screening experiment with 528 participants, people deciding alone, or alongside an unbiased model, selected candidates from all racial groups at roughly equal rates. Point them at a biased model and their choices tracked its preferences up to ninety percent of the time. The humans were fine until the recommendation showed up. Human oversight, the phrase carrying most of the load in AI hiring regulation, describes a person whose judgement the system has already colonised.

Against that, a study on Denmark's largest job portal compared recruiters searching manually, algorithmic matching, and the two together, and found the hybrid produced the fairest candidate lists of the three, better than either alone. Oversight worked there. The difference I'd bet on is accountability: the Danish recruiters were professionals on their own platform filling real vacancies with their reputations riding on the shortlist, while the experiment's participants were completing a task for a stranger and had no stake in being right. If that's the mechanism, then "human oversight" in a statute means nothing unless the human carries consequences, and none of the current rules require that. They require a person to be present.

A separate finding sits outside the demographic argument entirely. Maryland researchers ran twenty-two hundred resumes through commercial and open-source models and found self-preference rates of 67 to 82 percent: the models rank resumes generated by themselves above human-written ones of equivalent quality. Whatever that is measuring, it isn't the candidate. It rewards knowing which vendor screens your application, which is knowledge distributed exactly as unevenly as you'd expect.

Applicants adjust too, and that quietly undermines everything above. An IZA experiment found application rates falling 4.6 points for women and 3.2 points for men as AI involvement in the evaluation increased. If women withdraw from AI-screened roles at a higher rate than men, every callback-parity statistic in every study I've cited is computed on a pool the screener already reshaped before it read a single CV. A tool can post clean demographic parity across the applications it receives and still have skewed the workforce, because the skew happened upstream, in who bothered to apply. No audit regime I'm aware of looks there.

None of this makes the tools indefensible, and the scale argument does more work for me than any individual study. A bad human recruiter damages a few hundred careers across a working life, and eventually somebody notices the pattern. A model licensed across the Fortune 500 applies one idiosyncratic preference to millions of applications at once, and the only people positioned to notice are the ones running it. Mobley v. Workday is grinding through the Northern District of California on an age-discrimination theory, and it matters more than its facts because of the position it puts everyone in. The applicant can't see the screen. The employer that licensed the tool often can't either. The vendor owes an explanation to neither. A judgment against a vendor is the only lever anyone has found that reaches inside the model, and it's a slow, expensive, ten-year lever, aimed at an industry that is simultaneously removing the entry-level rungs those applications were pointed at.

Sources:

This post is timestamped using Blockchain technology. Verify

That Cab Never Took a Fare

Red coat over red jacket over red skirt, red gloves against black tights, and a brass rail she's leaning back on with both arms spread wide, as if she's holding the empty plaza still. The only words anywhere in the picture sit in the bottom corner: the Woolmark, and beneath it, Pure new wool. Everything else is colour doing the talking. Looked at cold it's an odd bit of staging, a woman braced against a handrail in a concrete forecourt, smiling at nobody, dressed for a season the light doesn't quite support.

Then the cab. There's a yellow taxi with checkerboard trim parked on what looks like a pier, rounded 1930s fenders in deep red, whitewall tyres, a TAXI dome on the roof, and a woman in a black knit coat and yellow turtleneck resting one gloved hand on the wing like she's just bought it. No taxi in New York looked remotely like that in 1992. The fleet was Caprices and boxy Ford sedans, square and dented and the colour of weak mustard. This car is a costume. Somebody built it, or found it and painted it, so the photograph could say New York faster than a street sign ever would.

The third frame lines three of them up against the water: black, yellow, red, each one a single column of colour from throat to shoe, hands on hips, the skyline behind them dissolved into pale blue mush. I spent longer on that mush than I meant to. Off to the right there's a tower with a steep pyramid crown and I can't turn that shape into anything except 40 Wall Street, with the older Financial District piling up to the left of it. Which puts the camera on the Brooklyn side looking west, on a cold bright morning, the two towers standing just off the right edge of the picture. I'm confident about the building and a good deal less confident about the pier.

Then the front plate on the cab, which reads RJ, Rio de Janeiro. I noticed it late and it irritated me for a good hour, because a number plate is usually the most honest object in a photograph, the sort of small hard evidence I've used to date a picture before. If that one is genuine, the paragraph above is in serious trouble. I've come down on the side of the skyline, on the grounds that a pyramid-crowned Art Deco tower is difficult to fake and a plate screwed to a costume car is not, but I can't get both facts to sit still at once. Either way the cab is a fabrication, down to a licence from the wrong hemisphere.

None of which touches the clothes, and the clothes are the part that hasn't moved an inch. Black and white harlequin diamonds running down one sleeve, red and gold scrollwork down the other, a four-strand gold choker sitting high on the throat, the whole thing knitted rather than printed. That's the part people forget about Escada: the pattern is structural, worked into the cloth on the machine, which is where the density in their late-eighties knitwear comes from. Shoulders still huge, skirts still short, but the colour blocking has been disciplined into flat graphic slabs. It's 1988's appetite wearing 1992's haircut.

Margaretha Ley died in 1992. Whether she saw these frames, approved them, sat in a room in Munich and said yes to the taxi, I have no way of knowing, and the catalogue went out regardless. The women in it are uncredited, which was house practice: Escada gave the shops a byline and gave the models nothing.

So the pictures have quietly changed jobs. The towers went, Ley was gone inside the year, the cab was scrapped or repainted into something else, and the three women at the rail are somewhere in their sixties, still unnamed. The door on that morning is shut. What got out was the light, which came off a red wool coat, went into a lens, and also went straight past it and kept going, and is thirty-four years out from here tonight, well beyond Vega, carrying an arrangement of yellow gloves and cold water that nobody is going to assemble again.

Amendment, 26 July 2026. They have names after all: the three at the water are Gail Elliott, Yasmin Le Bon and Niki Taylor. That also undoes the arithmetic in the sentence above, because Taylor was born in 1975, which made her a teenager when this was shot and makes her fifty-one now rather than sixty.

Sources:

This post is timestamped using Blockchain technology. Verify

Fifteen Percent Went to Shalimar

The parfum flacon floats over a bare shoulder and the woman selling it is caught in a gilt mirror, cropped dark hair, a red mouth, her eyes going sideways to her own reflection rather than out to you. Other bottles crowd the dressing table at the bottom of the frame. It's a picture about a private ritual, not an offer, and it was made for a fragrance that had been on sale for about seventy-five years by then.

I wanted to name her and couldn't. Five search routes turned up nothing. The closest dated match is Bruno Bisang, who shot Mitsouko for Guerlain in Paris in 1997 and kept two Polaroids from the job, and the entries that follow them in his archive are labelled Monica Bellucci and Claudia Schiffer while the Mitsouko pair carry only the campaign name. That's a photographer's filing habit rather than evidence of anything Guerlain decided, and I can't confirm it's even the same shoot. She stays unnamed and I've stopped pretending that counts as a finding.

The commercial picture behind the ad is less romantic. Guerlain went into the 1990s turning over around two billion francs a year, profits slipping, its catalogue visibly ageing, the last real hit being Samsara in 1989. By mid-decade Shalimar and Samsara were each running at roughly 15 percent of the house's perfume sales, and perfume was about 60 percent of the business. Mitsouko isn't named in that breakdown. Two flagships at 15 percent leaves seventy percent unaccounted for and the chypre could be sitting anywhere inside it, so this is an argument from silence and not a sales figure. What it does establish is which two names the house reached for when it had to describe itself, and neither of them was the one the critics loved.

That was the business the family sold. LVMH took 58.9 percent of Guerlain in 1994 for close to two billion francs, bought out the remainder by 1996, and valued the house at over four billion in the process. Champs-Élysées arrived that same year, built from a marketing concept rather than composed by a Guerlain nose.

The pressure from outside was worse. Thierry Mugler's Angel launched in 1992 on patchouli, praline, red berries and vanilla, invented the modern gourmand, and went on to outsell Chanel No. 5 in France; CK One followed in 1994. Neither smelled anything like a fruity chypre, and the mass of the market was already running toward clean and aquatic anyway.

Then the material itself was restricted. Oakmoss is the spine of a chypre, and also a skin sensitiser, and IFRA, the industry's own standards body, pushed its permitted level down to 0.1 percent of the finished product by 1997. That forced reformulation rather than deletion, but wearers noticed. A commenter on Perfume Shrine who bought her first bottle in the early 1990s described going back for a replacement and finding it simply wasn't the same, and only learned about the reformulation years afterward.

Which is one way to read that mirror. A house carrying a scent it can't grow and won't discontinue has little reason to pay for a famous face, and every reason to buy a beautiful anonymous one who is looking at herself rather than at you.

Amendment, 26 July 2026. She has a name after all: the woman in the mirror is Tara Westwood, the Canadian model born in Manitoba. The argument above about why Guerlain wanted an anonymous face is mine, not the house's, and it was always weaker than the sales figures.

Sources:

This post is timestamped using Blockchain technology. Verify

Sam Neill Was in the Hold

Phillip Noyce put three actors on two boats in the middle of the Pacific and gave himself ninety-five minutes. John and Rae Ingram, played by Sam Neill and Nicole Kidman, are sailing out the death of their small son when they take aboard Hughie Warriner, the sole survivor of a sinking schooner, played by Billy Zane as charm with something wrong under it. John rows across to inspect the wreck. Hughie sails off with Rae. The rest is two people trying to out-think each other on a yacht, cut against a man watching the water rise in a hold he cannot open. Terry Hayes stripped Charles Williams's 1963 novel to that, and the shape is close to airtight.

The reviews on release were not what its standing now suggests. Caryn James's New York Times notice on 7 April 1989 called it "an unsettling hybrid of escapist suspense and the kind of pure trash that depends on dead babies and murdered dogs for effect," and noted the novel contained neither. Roger Ebert granted the tension, then spent much of his review on the scaffolding, the killer who talks when he should shoot, the body any viewer can see will get up again. He is right on every count, and I would still say it barely matters. The machinery only becomes visible once the film has stopped, which at ninety-five minutes is most of what a thriller is for. It was made for A$10 million and took A$10.2 million worldwide.

Warner Bros. sold the situation rather than the people, which was right when the cast was an unbankable Neill, an American best known for a few minutes of the first Back to the Future, and a twenty-one-year-old Australian nobody abroad could name. The one-sheet sets Kidman's face above the flat water with the ketch becalmed under her chin, and lets the copy do the rest: a couple alone at sea, a stranger who called for help, and the fatal mistake of answering.

The Australian Film Institute gave it four awards from eight nominations, every one technical: score, cinematography, editing, sound. Nothing for Kidman, Noyce, or the picture itself. The industry could hear the machine working and declined to call it art. Video spent the next decade disagreeing, the same slow correction that turned Jacob's Ladder into a cult film a year later, and so in the end did the New York Times, which placed Dead Calm among the thousand best films ever made. Kidman was cast in Days of Thunder and left for America. Noyce went to Patriot Games. Dean Semler, who shot it, won an Academy Award the next year for Dances With Wolves. The material had already beaten a better director: Orson Welles shot his own version between 1966 and 1969 and never finished it.

Sam Neill died in Sydney on 13 July, aged seventy-eight, of pneumonia, three months after a clinical trial cleared the lymphoma he had been treating since 2022. He has the least showy part in the film by a distance, the sensible husband written offscreen into a flooding hold for most of the second half, and he plays it without one bid for sympathy. The picture belongs to Kidman and Zane, which is exactly how a performance that good goes thirty-seven years without being mentioned.

Sources:

This post is timestamped using Blockchain technology. Verify

Twelve Hundred Guests and a Bad Headline

On 23 February 1981, Yves Saint Laurent launched a men's fragrance at the Opéra-Comique in Paris and had Rudolf Nureyev dance for twelve hundred guests. The tagline was un parfum pour dieux vivants, a perfume for living gods. A spokesman at the press conference called it the most expensive male perfume on the market. Ten days after that party, the Palm Beach Daily News ran a piece under the headline "Kouros Debut Lacked Sweet Smell of Success." I can't get past that headline, all that survives in the citation trail, sitting oddly against a scent that then ran forty-five years without a break.

The name came out of a holiday. Saint Laurent had been to Greece and came back fixed on the archaic statues of young men, the kouroi. He described the blue of the sea and sky, the intense freshness of a world given over to beauty, and then seeing those figures again, and said he had his perfume and its name in one go. Nothing about that account predicts what turned up in the bottle.

Pierre Bourdon composed it. He had joined Roure Bertrand in 1971 to study under Jean Carles, and spent his apprenticeship being corrected by Edmond Roudnitska on bicycle rides. What he delivered was not freshness and marble. Under a bright opening of bergamot and bitter artemisia sits civet and oakmoss, held so that soap and animal never quite resolve into one or the other. Which of the two you notice has been a matter of temperament ever since: when Chandler Burr covered it for the New York Times in 2007, he ran the piece under the title "What the Cat Dragged In."

The advertising made no attempt to soften that. The bottle, a white flask with a chrome collar built to read as a Greek column, appears in the campaign held against a sunburnt shoulder, against a blue lifted straight out of Saint Laurent's account of the Aegean. The face is turned away and cropped at the jaw, there is no setting, and no line of selling copy anywhere beyond the logotype across the foot.

Nineteen eighty-one was a crowded year for masculine scent. Calvin Klein put out a men's fragrance in the same twelve months that almost nobody remembers. Kouros proved that a men's fragrance could be enormous, strange and openly carnal and still sell at scale for decades, and a proof like that changes what a marketing department will sign off. Dior's Fahrenheit arrived in 1988 smelling of petrol, by which point a difficult masculine was a known quantity rather than a gamble.

The vintage bottle question follows directly from that longevity. Kouros has run continuously since 1981 under five owners, and collectors date bottles by that chain: Charles of the Ritz to 1986, the American Parfums Corp. period from late 1986, Sanofi from 1993, PPR and Gucci from 1999, L'Oréal from 2008. Boxes and bases moved with each handover. The useful tell for the earliest stock is the silver metal base, where between 1981 and 1985 the batch code could be painted on rather than stamped, and some bases from the first year carry no inscription at all. Collectors rate that opening period the highest of the five.

Whether the juice changed as often as the packaging is less settled. A first reformulation in 1986, softening the animalic notes for a wider international audience, is disputed to the point that specialists call it the ghost reformulation, with some tasters finding no difference at all. The later changes are not disputed. Oakmoss came under tightening restriction and civet went the way of most animal materials, and by the 2000s people who had worn it since the eighties were reporting that the floor had dropped out. Build a scent on materials the industry later withdraws and this is what you get, which is why old bottles carry the weight they do.

Kouros is still on the shelf, which is the odd part. The flankers came and mostly went, Kouros Fraîcheur in 1993 and Body Kouros in 2000 among them, while the original survives as a quieter, drier version of itself at ordinary designer prices. So there are two objects now wearing the same name. One turns up on the discount shelf, and the other is a sealed box with a painted code on the base, and only the second one is what Bourdon actually signed off in 1981.

Sources:

This post is timestamped using Blockchain technology. Verify

Escada Kept Drawing Carriages

Margaretha Ley's first success was a small horse and carriage embroidered on a sweater. She put it in a shop window and it sold. The needlework came from her training as a seamstress at the Swedish royal court's tailor in Stockholm, and she named the house after a racehorse that had impressed her. By the early 1990s that one motif had become an entire heraldic vocabulary, printed across silk as state coaches and palace gates flanked by guardsmen, worn over tartan with gold chain belts.

The same coach reappears in seed beads on a navy velvet bomber, alongside a crown and a jewelled star, and once more as a Rolls-Royce grille stitched onto turquoise knit with black paw prints walking over the shoulder. "Hearts and stars and animals and flowers, these are things I always like to play with," Ley told the Chicago Tribune in 1989. The motif from that first shop window was still running through the house's pages more than a decade on.

The garments went in the most complete way available. Look at the shoulder on the magenta wool jacket, built out past the arm on a pad wide enough to stand a glass on, or the sleeve on the sweater-dress in the same abstract print, dropped so low it turns the body into a bell. Those proportions have no route back, and the ornament went with them.

The women are a different matter, though not because they look modern. The styling dates them to the month: lacquered volume in the hair, gold button earrings, a lip drawn hard at the edge. What refuses to date is the address. The bob and the red mouth at Bergdorf Goodman meets the lens with a composure that current fashion photography has largely stopped asking for, and the woman in the pink quilted jacket holds the same line with yellow trousers arguing against her. None of them is credited. The advertisements name Daniel Foxx of Palm Desert, Jacobson's, Leigh's of Grand Rapids, so the shops get a byline and the women get nothing. A named model ages alongside her career, the way Tatjana Patitz does in the 1991 campaigns. These faces were never attached to a biography that could run on without them.

Ley died in 1992. The line had already passed 1,200 pieces by the early nineties, taking in jewellery, handbags, gloves, scarves and shoes, so what Michael Stolzenburg held until 1994 and Todd Oldham took on the following year was a full inventory of codes to be kept up by people who had inherited them rather than meant them. The pages themselves sell inheritance hardest of all. Crests, state coaches, royal-court needlework, a Woolmark tag in the corner reading pure new wool: every signal insists this is the thing you hand on. The clothes went to the charity shops and the archive resellers. The face above the collar came through untouched.

Sources:

This post is timestamped using Blockchain technology. Verify

Selling Leather With His Own Face

North Beach Leather was never a house with the designer tucked out of sight. Michael Hoban opened it in 1967 in the San Francisco district that gave it its name, selling three garments: leather jeans, a Levi's-style jacket, and a double-breasted Edwardian coat that, as he later told the Los Angeles Times, "it seemed like everybody in the world wanted." Hells Angels bought first, then Huey Newton, then the Fillmore crowd once Bill Graham started sending his acts round. From 1969 Hoban was cutting custom white leather for Elvis. By 1988 UPI found ten stores and $25 million in wholesale volume behind him.

Then the 8-Ball jacket, dated fall 1989 in the Museum at FIT's collection. Salt-N-Pepa carried it into hip-hop, and within two years cheap copies of North Beach's jacket graphics were everywhere. Hoban sued. Eight manufacturers settled out of court by 1992.

Which is roughly the confidence on display in the 1993 campaign, where he stands dead centre in his own advert, grey and grinning in a black jacket with an Old English B, Daniela Peštová on his left in blocked red and cream, Yasmeen Ghauri on his right in navy. Two of the most bookable faces of the year, hired to flank the man who cut the clothes.

The ending was slower than the rise. SFGate reported the luxury trade falling away after September 2001; by July 2003 North Beach had filed general assignment, closed New York, Las Vegas, Houston and Chicago, and kept only the Grant Avenue shop. Hoban lost the business and moved to Hawaii. Skip Pas, his general manager, took the licence and the name.

Sources:

This post is timestamped using Blockchain technology. Verify

Dated Off a Number Plate

Works undertaken by R. C. Demolition (Bristol) Ltd, telephone 0272 231320, and underneath in smaller red letters, SAFTEY HELMETS MUST BE WORN ON SITE. Someone painted that misspelling, someone else hung the board on the hoarding, and nobody along the way cared enough to fix it.

That number was renumbered out of existence on 16 April 1995, when PhONEday moved Bristol from 0272 to 0117 and stretched local subscriber numbers from six digits to seven. So 231320 became 923 1320 overnight. The digits survive the translation perfectly well. What doesn't survive is any sense of what they were printed to reach.

Behind the hoarding stands a hotel that opened in 1898 as the Grand Clifton Spa and Hydropathic Institution, back when Clifton was selling the water rather than the view; in August 2018 it reopened as Avon Gorge by Hotel du Vin, refurbished and rebranded. Somewhere between those dates sits the long low range on the gorge edge, roof already stripped back to bare joists, cones out, orange tape strung across the pavement, and I can find no account anywhere of what came down or why.

Which leaves the cars to do the dating. The Mercedes on the left carries a D-plate, issued between August 1986 and July 1987. Further down the hill, the concrete Bristol of that same decade got hated, then argued over, then written about at length. This range got a board with a phone number and a spelling mistake on it, and that board is now the most complete document of the works I've been able to find.

Sources:

This post is timestamped using Blockchain technology. Verify

New York Dressed in Black

Three versions of black define this July 1985 American Vogue editorial: a bodysuit under a cashmere skirt, a broad-shouldered coat, then the bodysuit with patterned tights and bracelets stacked to the wrist. Andrea Blanch photographs Ashley Richardson against undecorated grey, while John Sahag pushes her blonde hair into an improbable halo. The title was “New York: A Revealing New Approach.”

Vogue’s “new” clothes were not casual. Donna Karan turned underwear into dinner dressing; Ralph Lauren put a square-shouldered cashmere coat over a turtleneck and silk-velvet stirrup trousers. In the opening page’s long black column, Richardson looks less dressed than engineered, with Carlos Falchi gloves and Wendy Gell cuffs supplying the decorative noise. The prices, $450 for the body and $1,290 for the coat, belong to a Manhattan conducted through department-store counters.

Blanch changes very little between frames, so adjustments count. Ralph Lauren’s coat makes her broad and sealed-off; the close portrait does the reverse, yet Richardson’s expression barely moves. The hair grows larger as the wardrobe gets barer, an exchange of one kind of volume for another.

Kevyn Aucoin’s makeup is quieter than the celebrity faces attached to his name. Vogue specifies pale rose on the eyes, apricot on the cheeks and gloss on the lips. None of it tries to soften the clothes. Richardson’s face remains present while the colour lets Sahag’s hair and the hard geometry below it do the talking.

Seen now, the pages feel less like a general idea of New York than a precise 1985 fantasy: expensive, controlled, ready to move from an office into dinner without changing clothes. I am not sure anyone lives that way, but Richardson makes the proposition look practical. In the final frame, the bracelets are excessive, the tights almost comic, and her calm expression refuses to acknowledge either.

Sources:

This post is timestamped using Blockchain technology. Verify

Opus 5 and the Missing Scoreboard

Anthropic shipped Claude Opus 5 this afternoon at five dollars per million input tokens and twenty-five out, identical to Opus 4.8, with the claim that it comes close to Fable 5 at half the price. Fable now bills at ten and fifty, so the halving is arithmetic rather than marketing.

On Frontier-Bench v0.1, an agentic terminal coding test, Opus 5 scores 43.3% against Opus 4.8's 18.7% and Fable 5's 33.7%. That gap is wide enough to survive a lot of scepticism, and some should be applied, because the footnote says the run was internal, on the mini-SWE-agent harness, mean reward over five attempts, with Opus 4.8 substituted whenever the safety classifier refused a task. A refusal scores zero otherwise, so 43.3% describes Opus 5 with Opus 4.8 covering its refusals, not Opus 5 alone. The same substitution holds up Fable's 33.7%. Opus 4.8's own 18.7% had nothing to fall back on.

Read the announcement looking for GPT-5.6 Sol or Gemini 3.1 Pro and they are not in it. Every named comparison is Claude against Claude: Opus 4.8, Fable 5, and Mythos 5, which still leads on cybersecurity work. The charts plot cost per task along one axis, so the claim isn't "we score higher", it's "we score higher for the money." Even the boldest line, a score three times the next-best model on ARC-AGI 3, which tests unfamiliar problems rather than variations on training data, never says who next-best is. OfficeChai reads that field as including Sol. Anthropic's own labels don't.

Two ways to take this, and I can't pick between them. Cost per task is the honest axis for anyone running agents in production, where finishing a job in fewer tokens beats a prettier percentage, and the shared leaderboards are rotting anyway, with SWE-bench Verified showing signs of contamination that let models pattern-match issue text instead of reasoning through it. Against that, a curve generated on your own harness, against your own back catalogue, cannot be checked by anyone outside the building.

Where the rivals actually stand is easier to state than to compare. GPT-5.6 Sol has been out since July 9 at $5/$30, and its max configuration scored 59 on the Artificial Analysis Intelligence Index against Opus 4.8's 55.7. Anthropic published nothing today that updates that composite, so on general reasoning the gap is unmeasured rather than closed. Gemini 3.1 Pro runs at roughly $2 and $12, a third-party estimate Google has never confirmed, well under half what Opus 5 charges, and it takes audio and video natively, which today's announcement never claims for Opus 5.

I went looking for one clean head-to-head reasoning score to settle it and found the same tracker quoting 94.6% and 80.0% for Sol on the same benchmark, on the same page. So no ranking from me, and I'd distrust anyone selling one this week. What today's sheet supports is narrow: Opus 5 costs what Opus 4.8 cost, which makes trying it cheap. Everything past that waits on numbers Anthropic didn't publish.

Sources:

This post is timestamped using Blockchain technology. Verify