OpenAI says an internal version of its next major model, called Astra, has solved or substantially advanced ten open problems in mathematics and theoretical computer science. The subjects range from sphere packing and lattice cryptography to group theory, quantum games and Ramsey numbers. The company has released a 253-page manuscript, reasoning walkthroughs and certificates formalized with the Lean proof assistant.

The announcement matters, but it should not be reduced to “AI proved ten theorems.” A derivation can be logically correct while relying on an incomplete formulation, a mistranslated assumption or an overstated interpretation. Lean makes mechanical checking considerably stronger. It does not replace expert reading, comparison with prior literature or the time required to establish scientific consensus.

The short answer

QuestionAnswer
What is OpenAI announcing?Ten results that resolve or improve open problems across several areas of mathematics.
Which model produced them?An internal version of Astra, described as OpenAI's next major model.
Are the proofs public?Yes. A manuscript, explanations and a repository of Lean certificates are available.
Does Lean guarantee everything is true?It checks that a formalized argument follows its rules and assumptions, not that the formalization perfectly captures the scientific question.
Are the results peer reviewed?Open publication enables scrutiny, but it is not by itself independent validation or acceptance by an academic venue.
What cost does OpenAI report?Roughly $2,000 in tokens at Sol API rates to search for the ten solutions.

Ten problems, not one difficult calculation

The set is not limited to hard exercises with known answers. OpenAI presents new bounds for high-dimensional sphere packing and binary and spherical codes. The manuscript also proposes a construction of non-sofic groups, a central question in group theory, and a disproof of a rigidity conjecture involving von Neumann algebras.

In complexity theory, the authors report new lower bounds for computing the permanent with arithmetic circuits. Another result concerns parallel repetition for general two-player quantum games. The list also includes approximation hardness for the closest vector problem, a foundational lattice problem relevant to understanding post-quantum cryptography.

The remaining groups cover Ehrhart's volume conjecture, multicolor Ramsey numbers and several Erdős problems in extremal graph theory. That breadth is part of the claim: the work is meant to demonstrate movement across different vocabularies and methods, rather than optimization for one uniform benchmark.

How Astra reportedly worked

According to OpenAI, Astra explored the problems with substantial inference compute. Humans then prepared the arguments as manuscripts with help from the same model. The workflow is therefore not a finished paper emerging untouched from a black box. It combines automated search, selection, human rewriting and formalization.

OpenAI estimates that the tokens used to find the solutions would cost about $2,000 at its Sol API rate. That figure is striking but incomplete. It likely excludes model training, experimental infrastructure, researcher time, discarded attempts and final formalization. It is a marginal generation estimate, not the full scientific budget.

The model itself is not publicly available at the time of the announcement. An outside team cannot reproduce the complete protocol simply by rerunning the same experiment. It can, however, examine the manuscripts and certificates, which are the most useful artifacts for immediate evaluation.

What a Lean certificate checks

Lean is a proof assistant. A demonstration is translated into formal objects that its kernel verifies step by step. When a certificate is accepted, it becomes extremely difficult to hide a logical leap, an omitted case or an invalid algebraic manipulation inside the formalized portion.

That guarantee is stronger than a cursory reading of hundreds of pages. A public repository also lets other researchers compile the proofs, inspect dependencies and identify which axioms are used. The small trusted kernel is one of the main reasons these systems are valuable.

Lean still answers a bounded question: “Is this term a proof of the theorem as encoded, under these definitions and axioms?” It does not decide whether the formal theorem exactly matches the problem the community intended. An error in translating the statement, an overly strong assumption or a shifted definition can yield a valid proof of a less important result.

Why human review remains necessary

Specialists first need to compare each statement with the literature. An idea may be new, rediscover a scattered result or improve only a case that was already understood. That judgment requires historical knowledge and engagement with the relevant community.

They must also assess explanatory value. A formal proof can be correct while being hard to understand, brittle to generalize or uninformative about the structure of the problem. Mathematics is not only about certifying propositions; it also creates methods that other people can reuse.

Authorship and responsibility matter as well. The manuscripts were produced through a system mixing a model and human researchers. Crediting contributions, documenting choices and allowing criticism is necessary if a corporate announcement is not to become the only available account.

What can be checked now

The first step is to clone openai/ten-proofs, record the Lean version and compile the certificates in the documented environment. A successful build confirms the technical integrity of the received package. It is not yet a mathematical review.

The second is to connect every formal theorem with its readable statement in the manuscript. Definitions, dimensional restrictions, constants and quantifiers deserve particular attention. A tiny difference can dramatically alter the reach of a result.

The third is to follow responses from researchers in each field, corrections to the repository and submission of manuscripts to journals or conferences. The claims will become stronger over time if the arguments survive adversarial, independent reading.

A methodological shift bigger than a leaderboard

Even if individual results require revision, the published workflow could remain significant. A model proposes avenues, a human selects those worth developing and a proof assistant closes part of the verification loop. That chain is more credible than a chatbot claiming a result without inspectable artifacts.

It also changes researchers' work. The bottleneck may move from producing ideas to formulating useful questions, auditing assumptions and organizing thousands of candidates. Formal tools then become trust infrastructure, not merely educational aids.

The Leiden Declaration nevertheless points out that AI adoption in mathematics affects attribution, access to tools, concentration of resources and the way early-career researchers learn. Releasing proofs helps, but does not settle those concerns.

What to remember

OpenAI has provided more than a list of claims in a blog post. The manuscripts and Lean certificates offer concrete material to audit, which is the right direction for scientific results produced with AI assistance.

Careful wording remains essential. A formally checked proof is a powerful milestone, not a shortcut to consensus. If the ten contributions are confirmed by their respective communities, the larger event will be the arrival of a research workflow in which automated generation, human expertise and formal verification operate together.