OpenAI published ten results on longstanding open problems in mathematics - achieved, by the company's own account, by an internal version of Astra, their next big model. Every proof is formalized in Lean, so a machine can check it. The tokens for all ten cost about $2000.
- The ten problems span fields from sphere packing in high dimensions to group theory and lattice-based cryptography; three close Erdős problems (146, 180 and 183).
- The results are from an internal version of Astra; people prepared the manuscripts with the same model, and it formalized every proof into a Lean certificate.
- OpenAI also published descriptions of the reasoning trail; the tokens for all the solutions cost about $2000 on Sol API.
There's one detail that lifts this release off the marketing shelf: each of the ten proofs comes with a Lean certificate - a formal record a machine checks line by line. In May, OpenAI showed a disproof of an Erdős conjecture, found while testing an unreleased model. Today they show ten results at once, and drop the name of their next big model along the way.
Why the certificate is what matters. A company's claim about its own model normally gets read with a grain of salt - benchmarks get tuned, wording gets stretched. A formalized proof leaves no room for that argument: it either passes the check or it doesn't. That's why this release carries more weight than any percentage table we've seen this year.
And the staging is worth noting too. OpenAI could have unveiled Astra on stage, with a demo and applause. Instead the name slips out in passing, inside a math paper. The move is calculated: the message isn't 'look how fast it is,' it's 'it already does a researcher's job.' There's a stronger signal than the staging, too: after the May result, outside mathematicians published their own follow-ups. The community isn't watching from the stands. It's in the game.
The price is what stays with me - about $2000 in tokens against problems generations of mathematicians broke their teeth on. Community verification is still ahead, and some of the shine may fade by then. But the ratio is new. The sharp researcher today asks the model early, before his colleague has asked.