Autor/es reacciones

Teodoro Calonge

Professor of the Department of Computer Science at the University of Valladolid

This is a high-quality article. It deals with feeding Large Language Models (LLMs) using a technique known as RAG (Retrieval-Augmented Generation).

RAG is the factor that has recently improved the performance of generative AI tools like ChatGPT, which initially yielded lackluster results. Essentially, specialized models were developed by training them with RAG tailored to specific fields (such as law or medicine). It was at this point that significantly better results began to emerge.

The approach here is similar: a model is created or fine-tuned using mathematical proofs, with the aim of generating—via generative AI—satisfactory results for new problems or ones the system has not previously encountered. According to the authors, the method has proven useful for some challenges but has failed in many others.

For this reason, I believe the article serves well as an early attempt to apply AI to mathematical proof, though the results indicate there is still considerable room for improvement in this field.

EN