OpenAI Astra Math Model Solves 10 Unsolved Problems

OpenAI Astra Math Proofs Tackle Unsolved Conjectures

OpenAI Astra math capabilities have just rewritten the boundaries of machine-driven discovery by solving ten open problems that resisted human progress for over a decade. The company formally announced the results on August 1, 2026, publishing a PDF of the proofs and walkthroughs of the model’s reasoning process. This marks the first official confirmation of the Astra name and what OpenAI describes as its “next major model family.”

The ten results span pure mathematics and theoretical computer science, from high-dimensional geometry to quantum complexity. The arguments came entirely from an internal version of Astra, while human researchers helped prepare manuscripts and formal verification with the same model. OpenAI says this is a deliberate demonstration that AI can contribute genuine new mathematical knowledge, not merely solve contest problems already known to be solvable.

The release follows a quieter milestone in May, when an earlier unreleased model produced a counterexample to the unit distance conjecture. That result, while preliminary, gave OpenAI the confidence to test Astra against much harder targets. This time, the company selected problems that had seen absolutely no main progress for at least ten years, and far longer in most cases.

CEO Sam Altman had already previewed Astra in Washington, D.C., framing it as a system built to handle long-running, multi-agent tasks. The model family will be the first to go through a planned U.S. government review process before any public launch. For now, OpenAI is releasing these mathematical proofs as a proof of concept for the model’s reasoning stamina and scientific ambition.

Ten Problems With No Progress in a Decade

The list includes several notorious open problems. Astra established the existence of non-sofic groups, a question that had been open since Mikhail Gromov introduced the concept around 1999. A group is sofic if its multiplication table can be approximated by permutations of finite sets. Gromov asked whether all countable discrete groups are sofic, and the field has been stuck for 27 years. Sebastien Bubeck, a researcher at OpenAI, called it “one of many new beautiful results” in a post on X.

Astra also resolved the Connes rigidity conjecture within the theory of von Neumann algebras. The model found a counterexample showing that the map from groups to their associated II₁ factors is not injective even for a rigid class, which refines decades of work on how much algebraic information survives. Other proofs settled problems in binary and spherical coding, high-dimensional sphere packing, arithmetic circuit complexity and the closest vector problem, each with direct implications for coding theory, cryptography and lattice-based security.

In combinatorial number theory and graph theory, Astra delivered new multicolor Ramsey numbers and tackled external number conjectures. As Brian Wang noted, many of these results are of the Erdős-type, feeding into discrete geometry and analytic number theory. The breadth is unusual even for a specialized human mathematician; a single general model generating insights across such disparate domains signals a new kind of tool.

Thomas Bloom, a mathematician at the University of Manchester who curates erdosproblems.com, wrote on X that the findings are “big news” and more significant than the May counterexample. He said the constructions themselves represent a genuine leap, not just a one-off guess. At the same time, Bloom pushed back on the notion that AI is replacing mathematicians, arguing that the model draws on a century of mathematical theory and was built by mathematicians.

The $2000 API Bill and Lean Verification

Generating all ten solutions needed roughly $2,000 worth of compute at the company’s Sol API rates, OpenAI disclosed. That figure covers the raw token cost for the final successful runs. The model produced arguments that humans, with model assistance, then turned into formal research papers. Each proof was also translated into a Lean certificate, creating a machine-checkable guarantee that the reasoning is correct. The code and walkthroughs are available on GitHub.

The use of Lean is significant. The reasoning walkthroughs reveal how Astra moved from high-level strategy to step-by-step formalization. Unlike traditional peer review, a Lean proof can be verified instantly and automatically, removing ambiguity. OpenAI is effectively betting that the only way to verify complex AI-generated mathematics is through such formal systems, a position that echoes earlier calls from fields like experimental mathematics.

Some researchers on Hacker News pointed out that the $2,000 cost is easy to misread. Commenter aabhay noted that the figure might represent only the final successful attempts, not the full experimental budget including discarded runs, prompts and hyperparameter sweeps. Without knowing how many unsolved problems were attempted and how many tries each problem received, the headline cost could be similar to P-value hacking. OpenAI has not disclosed those details.

Yet even with those caveats, the speed and cost are striking. A human mathematician might spend a career on a single conjecture. Astra’s ability to chew through ten problems, across fields, at a price comparable to a conference trip, will force publication norms to evolve. Some, like c7b on Hacker News, argued that AI-powered mathematics needs new standards: open weights, exact prompts and full reproducibility logs. The company’s current release, while detailed, stops short of that level of disclosure.

Experts React While Transparency Questions Linger

Noam Brown, the researcher who helped develop the test-time reasoning technology inside Astra, said on X that the team tried and failed to crack several other major problems. “Sadly, no Millennium Prize Problems (yet),” he wrote, adding that OpenAI did not spend heavily on each problem and that it is “possible to push test-time compute much further.” His message treats the current crop as an early sample, not a ceiling.

The reaction from mathematicians balances excitement with caution. Bloom praised the constructions but also noted that AI-assisted results still rest on decades of human-built theory. The Leiden Declaration on AI and Mathematics, which OpenAI cited, offers guidance on how credit should be assigned when a proof is generated entirely by a machine. The company argued that claiming human authorship for such results would misrepresent both the system’s contribution and the nature of intellectual work.

A separate concern involves reproducibility. As jsenn and c7b discussed on Hacker News, the math itself matters more than the prompt, but the prompt still defines the experiment. If OpenAI had slanted the problem selection or cherry-picked successes without reporting failures, the portfolio could look more impressive than it is. Independent verification will take months, but the Lean certificates give reviewers a concrete starting point.

The ten proofs now enter a more traditional phase: human peer scrutiny and, if they hold, integration into the wider mathematical landscape. For a company racing to build autonomous AI researchers, solving problems that humans could not touch for decades sends an unmistakable signal. The raw results are here; what the mathematics community does with them will determine whether August 1, 2026 becomes a footnote or a turning point.

FAQ

What is OpenAI Astra?

OpenAI describes Astra as its next major model family, built to coordinate multiple agents over hours or days on complex, long-running tasks. It is the first model slated for a planned U.S. government review process before any public release.

Have the Astra math proofs been peer reviewed?

Not yet. The company posted the arguments, Lean certificates and reasoning walkthroughs, but independent mathematicians are only beginning to assess them. Full validation will take time, though the Lean code provides a machine-checkable correctness guarantee.

How does Lean verify mathematical proofs?

Lean is a proof assistant that checks each logical step against a set of foundational axioms. If the system accepts the code, the proof is logically flawless. OpenAI used Lean to certify every solution, so the arguments are already verified at the formal level, even before human review.

SAVE WHILE SHOPPING 1 - How to make money from home online part time jobs

Recommended For You