OpenAI announces the results of ten mathematical research problems using its Astra artificial intelligence model

OpenAI has stated in a press release that the internal version of its Astra AI model has found solutions to ten research problems in mathematics and computing, in areas such as geometry, group theory, operator algebras, quantum complexity, cryptography and combinatorics, amongst others. According to the company, the total number of tokens — basic units of information — required to solve these problems would cost around $2,000. OpenAI expresses “deep respect and understanding” for those concerned about the impact of AI on these disciplines, including the signatories of the Leiden Declaration on AI and Mathematics; and states that the attribution of results must honestly reflect how each one was obtained, whilst encouraging the mathematical community to review them. The company is publishing these results openly in a paper, accompanied by the model’s account of its reasoning process.

 

Expert reactions

260804_Senén_openai

Senén Barro Ameneiro

Director of CiTIUS – Centre of Excellence for Research into Intelligent Technologies, University of Santiago de Compostela
Science Media Centre Spain

Are the conclusions backed up by solid data?

“Yes, insofar as the results have been verified. In other words, although the results were generated by AI, they were subsequently scrutinised by computer tools designed to verify mathematical proofs, as well as by experts, who confirmed the rigour of the solutions provided. Therefore, everything points to these being correct solutions to the ten published problems.

All the problems are interesting, although they vary in difficulty and apparent practical value. Some are highly significant within their respective mathematical fields. However, it is still very difficult to achieve groundbreaking results, although we must not forget that this is equally true for experts. Machines generally solve those problems that are well-defined and whose proof, however complex, may share common elements—in substance or form—with others that are already known and solved, as they have surely learnt from these. For a machine to invent concepts and new theoretical frameworks is quite another matter.”

Are there any significant limitations?

“As is almost always the case, it is the success that is reported, not the failures or the actual costs of achieving that success. This is the case here too, as OpenAI tells us this incredible story that, with around $2,000 in computational costs, it has solved the problems, but it does not mention what the previous unsuccessful attempts entailed in terms of time, people involved, computational and energy costs — other problems attempted but not solved, and problems solved only on the umpteenth attempt. It’s as if I won the Christmas lottery and said that, having spent 20 euros on a ticket, I’d won 20,000, for example. The cost—and therefore the total profit—is not the same if I simply bought the winning ticket as it is if I’d spent 5,000 euros on tickets. Admittedly, when it comes to mathematical proofs, the value of the result is not diminished, but the cost of obtaining it certainly is.

Furthermore, I am not aware that the prompts used in these proofs have been made public. This in no way calls the result into question—which can be verified regardless of the procedure followed—but it would be interesting to know them to assess their potential usefulness in other problems and the relative difficulty of tackling them automatically.”

Is this announcement more promotional than scientific, or is it genuinely revolutionary?

“Any announcement of this kind has a promotional purpose, and this is legitimate provided it does not mislead or conceal relevant information in order to lend credibility to the achievements.

As far as I can see in this case, the scientific content is sound and represents a qualitative advance over previous problems. It is one thing to solve complex problems that are already part of the established body of mathematical knowledge, and quite another to obtain proofs of the truth or falsity of problems that have long been tackled without success. The progress is spectacular, too: in 2024–2025, AI systems reached the level of the International Mathematical Olympiad, where the problems are very difficult but are known to have solutions; now they have moved into the realm of research, where no one knew whether a solution existed. In this case, the frontiers of knowledge are being pushed back, and this is not only usually more complex but also of far greater value.

We might say that the result is not yet revolutionary, but it points to results that will come in the future—perhaps soon—which could well be revolutionary.”

What implications does this have for the mathematical profession?

“Undoubtedly many. Not because AI will replace mathematicians, as some people are already suggesting, but rather as a complement to human work and creativity. Mathematicians will continue to be needed to ask the most pertinent questions and explore uncharted territory, far removed from existing knowledge. Machines can provide a capacity to explore potential solutions that is impossible for us. It remains to be seen, of course, how this will affect mathematical vocations, training and employment in the field of mathematics, but we will certainly have to rethink mathematics education—particularly that of specialists—and be alert to significant changes in the job market and the nature of the work itself, where AI will undoubtedly be present. In university education in particular, greater emphasis will need to be placed on problem formulation, critical verification and formal knowledge, and less on the practical aspects of completing proofs.”

The author has not responded to our request to declare conflicts of interest
EN

260804_Javier_openai

Javier Aramayona

Director of the Institute of Mathematical Sciences (CSIC-UAM-UCM-UC3M)
Science Media Centre Spain

In recent times, we have witnessed a growing presence of artificial intelligence in mathematical research. It has recently been announced that well-known problems in various fields of mathematics have been solved, or substantial progress has been made towards their solution, both by researchers using AI tools and directly by companies in the sector. In many cases, the validity of the proofs is guaranteed by their formalisation in Lean, a proof assistant that enables the automatic verification of arguments.

Combined with formal verification, these tools will enable us to tackle more ambitious problems, explore ideas more quickly and detect structures and patterns that were previously beyond our reach. It is, without doubt, a promising prospect: they have the potential to raise the already excellent standard of current mathematical research even further.

 

At the same time, these advances raise questions that directly challenge the mathematical community: how to evaluate and attribute results obtained with the aid of AI, how to adapt publication models, or how to train the next generations in this context. Answering these questions will require an active and measured dialogue, which we must undertake with a broad perspective, steering clear of both fatalism and triumphalist proclamations. The community has already begun to organise itself in this regard: the recent Leiden Declaration on Artificial Intelligence and Mathematics, endorsed by the International Mathematical Union, proposes precisely such shared standards for the use of these tools.

It is worth remembering, in any case, that mathematics goes far beyond the solving of specific problems, however interesting and difficult they may be. There is a profound interplay between problem-solving, the development of theories and the formulation of conjectures, which in turn drives — and sometimes creates — entire fields of mathematics. Deciding which questions are worth asking, and understanding what their answers tell us, will remain essentially human tasks.

The author has not responded to our request to declare conflicts of interest
EN
Publications
Study types:
  • Research article
  • Non-peer-reviewed
The 5Ws +1
Publish it
FAQ
Contact