AlphaProof Nexus: A new AI tool designed to tackle mathematical proofs

A Google DeepMind team has developed a new artificial intelligence (AI) system based on large language models capable of autonomously tackling select mathematical problems. The tool, called AlphaProof Nexus, searches for proofs and employs verification processes to ensure the resulting solutions are logically sound. According to the study published in Science, AlphaProof Nexus has successfully solved several previously unsolved problems, including nine of the 353 Erdős problems.

 

Expert reactions

Haya - IA mates

Pablo Haya Coll

Researcher at the Computer Linguistics Laboratory of the Autonomous University of Madrid (UAM) and director of Business & Language Analytics (BLA) of the Institute of Knowledge Engineering (IIC)

Science Media Centre Spain

This study demonstrates that an open architecture based on artificial intelligence (AI) agents has successfully and autonomously solved nine open Erdős problems and 44 conjectures from the On-Line Encyclopedia of Integer Sequences. Beyond the specific results, the value of the authors' work lies in a methodology that enables the systematic generation of powerful mathematical discoveries—a capability previously almost exclusively held by cutting-edge research labs.

Since the beginning of the year, we have witnessed a succession of increasingly impressive AI-driven results in mathematics. A watershed moment arrived in September with the—albeit controversial—announcement that one of the "Millennium Prize Problems" (the Navier-Stokes equations) had been solved using specialized mathematical models and thousands of collaborative agents. There is little doubt left that we are witnessing the birth of a new way of exploring and doing mathematics.

The mathematical capabilities of LLMs (Large Language Models) have improved dramatically. However, as Terence Tao has noted on several occasions, the goals of AI labs and the scientific community do not always align. Labs aim to solve as many problems as possible to showcase the value of their technology, whereas the scientific community seeks to deepen its understanding of mathematics. Tao points out that automating problem-solving can come at the expense of professional practice. The value lies not merely in proving a theorem, but in the entire process leading to that proof. According to Tao, AI eliminates some of the knowledge a mathematician typically gains while attempting to solve a problem—knowledge that is just as valuable to the practice of mathematics as the proof itself.

The author has not responded to our request to declare conflicts of interest
EN

Curto - IA mates

Josep Curto

Academic Director of the Master's Degree in Business Intelligence and Big Data at the Open University of Catalonia (UOC) and Adjunct Professor at IE Business School

Science Media Centre Spain

When we read headlines about artificial intelligence (AI) systems "solving open mathematical problems that had been stalled for decades," it is easy to get swept up in the narrative of inevitable progress. However, a close examination of the technical details in the paper on AlphaProof Nexus—published in Science by researchers at Google DeepMind—reveals a pattern that has become all too common in the current AI landscape: a dazzling display of technical capability marred by significant methodological inconsistencies and a severe case of over-engineering.

The paper presents AlphaProof Nexus as an advanced framework for searching for formal proofs in Lean. The authors built a "complete" agent (Agent D) that combines a basic generate-verify loop, integration with the AlphaProof prover via reinforcement learning, and an evolutionary algorithm guided by a voting system and Elo ratings among sub-agents. On paper, the figures are impressive: the autonomous solution of nine open problems from the famous Erdős list, 44 OEIS conjectures, and applied discoveries in optimization and graph theory.

But when one delves deeper into the methodology and the study's own post hoc analysis, the conceptual structure begins to crack.

First. The authors' own ablation analysis reveals a damning fact: Agent A—a basic, simple "Ralph loop" alternating between LLM generation and Lean compiler verification—solved the same nine Erdős problems as the sophisticated Agent D. What does this tell us? That the framework's complexity is not what drives discovery; what really moves the needle is the base capability of the LLM (in this case, Gemini 3.1 Pro) combined with the direct, uncompromising feedback of the formal compiler.

Second. From a data and resource management perspective, the study’s financial assessment is, at best, opaque. Regarding inference cost metrics (in USD), the study shows that simplified agents are more cost-effective for most problems, whereas Agent D justifies its cost only in specific instances, such as Erdős #125. The article omits two relevant details: the calculations account only for inference during successful runs, and the reported costs for Agents B and D exclude the inference cost associated with AlphaProof.

Third, there is a misconception that using a formal language like Lean eliminates "hallucination." While Lean guarantees the logical validity of a proof relative to the statement written in Lean, it does not guarantee that said statement faithfully translates the actual mathematical problem. The article acknowledges several scenarios in which agents hallucinate or fall prey to errors in formalization. This demonstrates that the system still requires intensive expert human oversight to verify that what the agent has proven is indeed what was intended to be proven.

The true lesson of this work—even if the authors attempt to qualify it—is that both the market and the research community should lean toward simplicity: clean agentic loops, powerful language models, and environments with strict feedback mechanisms (compilers, code execution, unit tests, and expert human oversight)—though such setups are likely beyond the reach of many universities and research institutes. Everything else, for the time being, amounts to evolutionary fireworks that add architectural complexity rather than tangible results.

The author has not responded to our request to declare conflicts of interest
EN

Aramayona - IA mates

Science Media Centre Spain

On one hand, integrating AI into mathematical research holds great potential—potential that will only grow as these new tools are refined. On the other hand, it raises crucial questions regarding many research-related processes, such as the evaluation of scientific results, the use of data fed into Large Language Models (LLMs), and the training of early-career researchers in the AI ​​era.

Above all, there is a fundamental issue concerning the understanding of the research results produced. The primary (and perhaps sole) objective of mathematical research is the advancement of mathematical knowledge: understanding objects and structures, developing new theories, and so forth. Problem-solving—however significant the problems may be—serves as a measure of this progress but is never an end in itself. We are currently at a historic juncture marked by a curious paradigm shift: the emergence of a vast array of results that we know to be true—because they come with a certification of validity—yet which no one fully understands.

As a community, we must continue to insist on the need to understand the advances being made, while also respecting and fostering the spaces, structures, and processes that make such understanding possible.

The author has not responded to our request to declare conflicts of interest
EN

Calonge - IA mates

Teodoro Calonge

Professor of the Department of Computer Science at the University of Valladolid

Science Media Centre Spain

This is a high-quality article. It deals with feeding Large Language Models (LLMs) using a technique known as RAG (Retrieval-Augmented Generation).

RAG is the factor that has recently improved the performance of generative AI tools like ChatGPT, which initially yielded lackluster results. Essentially, specialized models were developed by training them with RAG tailored to specific fields (such as law or medicine). It was at this point that significantly better results began to emerge.

The approach here is similar: a model is created or fine-tuned using mathematical proofs, with the aim of generating—via generative AI—satisfactory results for new problems or ones the system has not previously encountered. According to the authors, the method has proven useful for some challenges but has failed in many others.

For this reason, I believe the article serves well as an early attempt to apply AI to mathematical proof, though the results indicate there is still considerable room for improvement in this field.

The author has not responded to our request to declare conflicts of interest
EN
Publications
Journal
Science
Publication date
Authors

Tsoukalas et al.

Study types:
  • Research article
  • Peer reviewed
The 5Ws +1
Publish it
FAQ
Contact