OpenAI says its next model solved ten problems mathematicians couldn't. The proof you can check isn't the part that was ever in doubt.
Astra's ten results ship with machine-checkable Lean certificates — the strongest verification any AI-generated mathematics has cleared. What a certificate can't tell you is whether the theorem it verifies is the open problem anyone was trying to solve, or how many human hands were on the wheel.

Image: HaeB / Wikimedia Commons (CC BY-SA 4.0)
A machine-checkable proof certificate is one of the few genuinely trustworthy objects in a field full of overclaims. Feed a formal proof to the Lean theorem prover and it will tell you, without opinion or mercy, whether every step follows from the axioms. It cannot be lobbied. It does not care who wrote the proof, or how, or what press release is attached. So when OpenAI published ten new mathematical results on August 1, each one carrying a Lean 4 certificate anyone can download and re-run, it cleared a bar that almost no AI-generated mathematics had reached before. The certificates are real. What they certify is narrower than the headline, and the distance between the two is the entire story.
The results are attributed to Astra, which OpenAI describes as its next major model family — a system built not to answer a prompt in seconds but to coordinate several agents on a single hard problem for hours or days at a stretch. The company has not said whether it will ship as GPT-6, as a point release, or as a separate class alongside its existing Sol, Terra and Luna models, and it has given no release date. What it has given instead is a demonstration: ten problems that mathematicians had left open for at least a decade, and in several cases far longer, that OpenAI says Astra solved for roughly $2,000 in tokens at current API rates.
Taken at face value, the list is not modest. The headline result is the first explicit construction of a non-sofic group — a question that had stood open since Mikhail Gromov introduced the notion of soficity in 1999, and that a generation of group theorists had failed to settle in either direction. OpenAI says Astra also disproved Connes's rigidity conjecture in the theory of von Neumann algebras, proved Ehrhart's volume conjecture, produced the first improvement to the general upper bound on high-dimensional sphere-packing density since 1978, proved a parallel-repetition theorem for two-player quantum games, and resolved three problems from Paul Erdős's famous catalogue. Any one of these, done by a person, would be a strong year's work. The claim is that a single system produced all ten, and that the compute to find them cost less than a mid-range laptop.
The machine that runs for days
It helps to be clear about what kind of system is being described, because the architecture is the reason the demonstration takes the shape it does. Astra is not, in OpenAI's telling, a bigger chatbot. It is a long-horizon system: multiple model instances working as coordinated agents, decomposing a problem, pursuing separate lines, and — this is the hard part — correcting each other as the shared context grows past the point where any single pass could hold it. The engineering problem such a system exists to beat is error compounding, the way a small mistake early in a long chain of reasoning quietly poisons everything downstream. OpenAI concedes the trade-off candidly: the coordination overhead that lets many agents work in parallel can also degrade performance on tightly coupled tasks, where every step depends on the last and there is nothing to parallelize.
Mathematics is a shrewd choice of proving ground for exactly that architecture. A hard proof is long-horizon by nature — days of dead ends and backtracking — and, unlike most tasks a model could run for days, its output can be checked cold. You do not have to trust the model's account of its own reasoning; you can hand the final proof to a verifier that has no stake in the outcome. That checkability is what makes the $2,000 figure land the way OpenAI intends. It reframes original mathematics, tacitly, as a thing you can now buy by the token — priced, repeatable, and cheap. That reframing is the argument the report is really making. It is worth examining before accepting it.
Lean can tell you a proof is correct. It cannot tell you the proof is of the problem you meant to ask.
What the certificate checks, and what it doesn't
This is the point where the technical detail stops being decoration and becomes the argument, so it is worth slowing down. A Lean certificate verifies one thing with total rigor: that a given formal statement follows from the stated axioms by valid steps. It is a guarantee about the proof. It is not a guarantee about the statement. Between the informal conjecture a mathematician cares about — do non-sofic groups exist? — and the formal theorem Lean actually checks, there is a translation step, called formalization, and that step is a human judgment about whether the symbols faithfully capture the question. A formal theorem can be flawlessly proved and still encode a subtly weaker claim than the one the field had been trying to crack. Lean will certify it either way, because checking the faithfulness of the translation is not Lean's job. It is nobody's job — except the reviewers'.
On this run, the reviewers and the translators are, for now, OpenAI. The company is admirably candid about its role: it says its researchers took the model's arguments, refined them into manuscripts, and formalized the proofs in Lean, and that it takes responsibility for their correctness while crediting the mathematical arguments themselves to Astra. Read that division of labour closely. The entity that stands to gain from the sentence "our model solved ten open problems" is also the entity that decided which formal statement counts as each open problem, prepared the write-up, and published the certificate. None of that makes the results wrong. It means the one judgment a certificate cannot automate — is this theorem the problem we cared about? — is currently being made by the interested party.
Three things you can't check from outside
I want to be precise about my confidence here, because the temptation to overshoot exists in both directions. The mathematics may well hold; the early reactions from mathematicians who have looked include Thomas Bloom's flat "this is big," and OpenAI's own Noam Brown calling the work "a major step for scientific reasoning." But "may well hold" is as far as the public evidence licenses anyone to go, for three reasons the announcement itself makes unavoidable.
- The model is internal. Astra has not been released, which means no outside group can put the same system in front of a fresh open problem and see whether the feat reproduces. A result you cannot reproduce is a report, not yet a finding.
- The human contribution is unquantified. The write-up does not spell out how much problem-shaping — choosing the angle of attack, supplying intermediate lemmas, steering away from dead ends — preceded each successful run. The gap between "a model solved this" and "a model solved this with expert hands on the wheel" is exactly the gap between the two stories OpenAI could be telling, and the report leaves it unmeasured.
- The grader is the author. The faithfulness of each formalization, the framing of each manuscript, and the decision that a given result answers a given famous question are, at this moment, OpenAI's calls about OpenAI's model.
OpenAI is not blind to any of this. It cites the Leiden declaration on AI and Mathematics — a statement by mathematicians worried about precisely how machine-produced results get credited, absorbed, and trusted — and says it wants the community to engage with the ten results rather than simply accept them. That is the right instinct, and it concedes the thing worth noticing out loud: the verification that matters most is the one that has not happened yet. The slow kind, done by people with no stake in the answer, at the pace these conjectures usually command. Publishing the .lean files on a public repository is what makes that scrutiny possible. It is not the same as its having been done.
A proof, and a political document
There is a reason a company would publish this particular report in this particular week, and it is not only mathematical. Astra surfaced in public the same week Sam Altman was in Washington demonstrating it to policymakers and regulators behind closed doors, and the same week it is positioned to become the first model submitted under the federal government's new pre-release review framework for frontier systems. A report organized around ten solved conjectures is, among other things, a capability argument — evidence, offered to the people writing the rules, that the model in the room is doing science, not manufacturing risk. The mathematics and the politics are not separable here. They were released together, and on purpose.
OpenAI has said its longer goal is an autonomous AI researcher: a system that does original science without a person shaping each step, on a timeline it has floated as early as 2028. The ten proofs are being offered as the first hard evidence along that path. Whether they actually are depends entirely on the unmeasured variable above — how much of each discovery was the model and how much was the humans around it — which is the one number the report does not give, and the one a regulator asked to assess "advanced capabilities" would most need to see. A demonstration that answers the easy question loudly while leaving the hard one blank is not a neutral document. It is a persuasive one.
So the honest summary is a split decision. The certificates are the most credible thing anyone has yet attached to AI-generated mathematics, and if even a few of these results survive the slow scrutiny of the field, it is a genuine event in a discipline that does not hand those out lightly. But a certificate answers the question "is this proof valid?" with a rigor that can quietly stand in for the question it does not touch: did a machine, more or less on its own, solve problems humans could not — and who gets to say so? Right now the same company owns the model, the manuscripts, the formalization, and the announcement, and is showing the result to its regulators as it goes. The number to watch is not ten, and it is not $2,000. It is how many of these proofs are still standing a year from now, checked by people who do not work there. Until then, the strongest verification in AI mathematics has certified everything about these results except the part that was ever actually in doubt.
References
- The Decoder — OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions
- The Next Web — OpenAI says its next model, Astra, has solved ten open problems in mathematics
- The Information — Exclusive: OpenAI Previews 'Astra' AI Model in DC
- OfficeChai — OpenAI Says It Has Solved 10 Open Math Problems Using Astra, Its New Model
- AI Weekly — OpenAI releases ten Astra math proofs with Lean certificates


