On Tuesday OpenAI posted 722 manuscripts to GitHub, grouped into 372 problem families, all produced by a model nobody outside the company can use. According to TechSpot, the model was handed about 4,000 problems, so what we're seeing is the fraction that came out well enough to publish. Some of it touches the Riemann hypothesis and two other Millennium Prize problems, though OpenAI only claims progress on those, not solutions.

Good or bad depends on which half you mean. The mathematics is probably good, much of it. Back in May the unit-distance disproof held up under outside checking and carried a genuinely new idea. This time OpenAI says around 300 of the main results have been formalised in Lean, the proof assistant that checks every logical step mechanically, and Rutgers' Alex Kontorovich said of the Riemann result that if a human had done it, it would be an instant Fields Medal. Daniel Litt passed on to Fortune a collaborator's line about having crawled their whole life and now being able to fly. I don't doubt that's how it feels from inside.

The release is the bad half, and it's why I wouldn't trust the batch yet. Fortune reports roughly three hours of computing time per solution. A long technical paper can take human referees years. Within days OpenAI had retracted three papers over an elementary error and amended several others, which tells me the checking started after publication, not before. Lean doesn't close that gap on its own, either. A paper covered by TechCrunch found at least two places where the Lean code for OpenAI's Navier–Stokes result didn't match the written proof. Lean will confirm that you proved something. It can't tell you it's the thing you said you proved.

An advisory group of mathematicians asked labs to stop testing hard problems on proprietary models; OpenAI declined. The group asked for the model's reasoning, which is where anyone would learn how a result was found, and TechCrunch counts ten manuscripts out of more than 700 that include it. A company spokesperson also told Scientific American that many of the results aren't yet understood by OpenAI's own mathematicians. So the company published proofs it hadn't absorbed, from a system nobody else can inspect, and handed the reading to the rest of the field for nothing. Harvard's Melanie Wood put the situation exactly: there's no human understanding at the point of release, "and now the work begins." That work is real, slow and unpaid, and OpenAI got to announce the results before any of it was done.

Sources: