Science with AI: trick or treat

Over the last four years, we have seen that artificial intelligence (AI) is capable of playing a role in every stage of the scientific process. However, this carries risks: today, more than half of the peer reviews of papers are not carried out by researchers, but by AI systems, which are excellent at analysing proposals similar to others, but tend to reject innovative ideas. We are facing a revolution that is forcing us to redefine how we do science. Whether AI becomes a trick for misleading or a treat for the advance of knowledge will depend on us.

ciencia IA

If AI is capable of handling all stages of scientific development, then should it share responsibility? Image credit: Adobe Stock.

 

On 30 November 2022, the world changed. It was the day humanity discovered ChatGPT, a public version of OpenAI’s generative Artificial Intelligence (AI) which, within a matter of days, took the lives of millions of people by storm. The way it worked was simple: a text field where you could ask about any topic, and a response would be provided within seconds. It seemed to know everything, to understand us, to know us. It constructed meaningful sentences which, to a greater or lesser extent, were relevant to the topic in question.

We had the feeling we were having a conversation with someone, when in reality we were exchanging information with an algorithm that had virtual access to all our accumulated knowledge, accessible via the internet. A system capable of extracting relevant elements from that vast array of data and identifying patterns from which to construct elaborate, statistically sound responses relevant to the context of the question; but with the risk that some of these responses might contain ‘hallucinations’ – data invented by the AI when it could find no references on a particular topic. Following ChatGPT, other AIs emerged, such as Copilot, Grok, Gemini, Perplexity, Grammarly, DeepSeek and Claude. We have come to interact with them on a daily basis, for better or for worse. This is also the case in science.

The fundamental pillars of scientific research are two documents: the research proposal and the scientific paper. The hypothesis to be validated and the results obtained that must be reported. That is all. Using these two documents, we have assessed the creativity, novelty, innovation, relevance, effort, work, excellence and impact of individual researchers, laboratory heads, research groups and centres. Now, as AI enters our scientific lives, we are seeing how it is making its way into every corner of the field and is capable of intervening in the formulation of the hypothesis, the drafting of the research proposal, its evaluation and execution, the collection and interpretation of results, the writing of the paper, its evaluation and its eventual publication. In other words, in less than four years we have discovered that AI is capable of contributing to each and every stage of the research process.

If this is true, we have, it seems, created a problem. If AI is capable of addressing all phases of scientific development, then should it share responsibility? Is AI an object, or a subject with rights, like human beings? Should we give it credit for having helped us, or for having directly supplanted us in our scientific work? These are questions that are piling up on our table, and for which we do not yet have answers.

More speed, less creativity

Barely a year after the launch of ChatGPT, Nature magazine had already launched a survey of some 1,600 researchers to assess the positive and negative impacts of AI on science, analysing how they were using it.

On the positive side, it is clear that, thanks to AI, we can automatically manage huge volumes of data, speeding up computational analysis and saving us time and money. It helps us write computer code and articles, answer difficult questions to make discoveries and formulate hypotheses; find the information we need at any given moment; and summarise everything from articles or projects to vast amounts of data into short paragraphs or graphs. For non-English speakers, it provides support with the English in our texts. It even helps us evaluate projects or articles. But all of this also comes at a cost.

AI is very good at making the most of all the data it has access to, but it isn’t quite as good when presented with new or disruptive ideas 

On the downside, AI is capable of identifying patterns in datasets, even when none exist, uncovering relationships it cannot explain. It is a black box that generates responses and replies, but we do not know exactly how it does so. Its responses will contain biases depending on the data we have used to train it; and it may produce results that cannot be replicated. Furthermore – and this is particularly worrying – it will facilitate fraud, as it will be easy to fabricate data that appears plausible and fits the context, yet is false. Consequently, disinformation will proliferate, instances of plagiarism will increase, and errors and biases will creep into projects and articles; it may even compromise the training of new researchers. There will also be a rise in the enormous energy consumption required for all these AI systems to operate on the servers that host them.

The journal Nature reviewed this survey two years later, and just four figures are enough to illustrate the rapid adoption of AI in science. We have gone from using AI to write articles at a rate of ~28 per cent in 2023 to over 46 per cent in 2025. To assist us in our research, the figure has risen from 25 per cent to 64 per cent. To help us with project writing, it has risen from ~18 per cent to 31 per cent; and, most surprisingly, for evaluating scientific articles, it has risen from a modest 15 per cent to over 50 per cent. To put the latter into perspective: more than one in two evaluations of scientific articles carried out today have not been performed by researchers, but by AI.

Researchers who have embraced AI from the outset have a greater impact, increasing their citation counts and productivity. However, their creativity suffers and they tend to focus on increasingly narrow topics. This is a first warning sign. AI is very good at making the most of all the data to which it has access – the known universe – but it is not quite as good when presented with new or disruptive ideas for which there is no prior data to relate them to.

Feeding AI with confidential texts

Undoubtedly, the most widespread use we scientists make of AI is to check the English in our texts. In the past, we would turn to our respective ‘Paul Smiths’ or ‘Jane Browns’ (fictitious names) – friends we’d made abroad – or to a native English-speaking academic editor who was able to transform our flawed English into flawless texts in the language of Shakespeare. Now we no longer need them.

We upload our texts to any of these AI systems and they return them to us improved, in a range of versions from which we can choose. We are unaware of the danger this harbours. Every time we upload a text to an AI tool to have our English checked – unless our institution has installed an AI engine locally – what is effectively happening is a transfer of our texts – which are likely to be confidential – to the cloud, making them available to any other search carried out anywhere in the world. And, if we wanted to use those texts to file a patent, we might lose the novelty of the invention, as it would already be available to anyone.

The fine line between improvement and authorship

If AI can help you reformat an article that has been rejected by one journal so that you can submit it to another that requires a different format, then bring it on. There’s no more thankless task than reformatting articles. If we can do this automatically, then let’s do it!

But it’s one thing for AI to help you improve the writing or structure of your proposal; it’s quite another to ask AI to write the project or article for you from scratch, based on a series of more or less disjointed data, results and references. The individual’s intellectual contribution then disappears. The credit we’d like to claim as authors vanishes.

The depth and quality of the responses will depend on whether we use free or paid versions, which will lead to inequalities amongst researchers 

That is why the use of AI as a co-author or assessor of the scientific quality of abstracts submitted to scientific conferences is already being explored. Initial findings suggest that, in the case of abstracts assessed by AI, it is more difficult for fabricated content to slip through; they are technically correct but lack creativity. When abstracts have been written by AI, they are usually reinterpretations of familiar topics. AI systems are clearly better at analysing results and extracting patterns or conclusions. Conversely, abstracts written by humans are better at formulating hypotheses, and those written by AI without human involvement were the least likely to be selected.

There are AI tools that help us write research papers and articles and prepare the figures that illustrate them (such as Paperpal, Jenni or PaperBanana), but we must be aware of that fine line between enhancement and authorship. Some might argue that we will always be able to spot texts written by AI; yet technology advances relentlessly, and there are already AI systems that humanise the output of other AI systems so that AI-generated writing becomes undetectable. It is a never-ending battle.

We must also be aware that the reliability of an AI will depend on the business plan we have signed up for. The depth and quality of the responses will depend on whether we use free or paid versions and on how much we pay. This will create inequalities amongst researchers that are likely to be greater than those that arose years ago between those who had access to the internet and those who did not.

The pitfalls of evaluation

One of the areas in which AI is undoubtedly the most controversial in science is its use in evaluating research projects and scientific articles. Peer review is, by definition, a task that cannot be delegated. You are invited to review that text, not someone else. It is you who will sign that review when you return it to the journal or the funding body. This is why researchers, editors, publishers and research agencies are very concerned about the current situation, in which more than 50 per cent of these reviews appear to be carried out by AI, not by researchers, who submit these texts claiming authorship without actually having written the reviews themselves.

The first time I used QED, an AI that analyses the robustness of a paper, it became clear to me that we were witnessing the end of human peer review as we know it. It’s over

P picaresque tactics have also made an impact in this area. If I know that my project or article is going to be assessed by an AI, I can include in my text, written in white (invisible to humans, but not to AIs), instructions to ensure the assessment is lenient, favourable and positive. These are instructions that the AI will follow to the letter.

One of the most advanced AIs for the detailed analysis of scientific texts is QED, an acronym derived from the Latin “quod erat demonstrandum” (as we wished to demonstrate), developed at Tel Aviv University. This is an AI to which you can upload a project or article, and it will analyse the robustness of the proposal: whether the hypothesis is well formulated, whether the objectives and experimental approaches are the most appropriate, whether evidence is lacking or, conversely, whether there is redundant evidence, whether the necessary works are cited, or whether references are missing or superfluous. The first time I used this AI, it became clear to me that we were witnessing the end of human peer review as we know it. It’s over.

One might think that the paper by Doudna and Charpentier, which proposed using the CRISPR system to edit genes, might have been blocked by an AI because it was too disruptive and unexpected

AI is transforming the way we review our scientific texts, and we should be aware not only of its advantages (of which there are many) but also of its limitations. For example, it is excellent at detecting technical errors that might go unnoticed by researchers (such as incorrect units in figures). It will work very well when evaluating projects or articles similar to existing ones, against which it can make comparisons. We therefore run the risk that AI will dismiss or reject proposals that are very different from anything previously known.

The systematic use of AI in scientific evaluations may undermine the promotion of innovative proposals; this is precisely the opposite of what we ask of researchers: that they strive to come up with novel and innovative ideas. One might even imagine that the paper by Nobel Prize in Chemistry laureates Jennifer Doudna and Emmanuelle Charpentier—which proposed using the CRISPR system discovered by Francis Mojica to edit genes—might have been blocked by an AI system for being too disruptive and unexpected.

In the face of these potential dangers, there are examples suggesting that perhaps using AI to evaluate projects is not such a bad idea after all. This is what the “La Caixa” Foundation did in one of its most recent calls for proposals. It trained an AI model using all the projects – both selected and rejected – from previous calls, and fed it all the applications received: 714. The AI rejected 122. These 122 were sent to human evaluators, who reinstated 46. Finally, the remaining 638 proposals were sent to human evaluators, who selected 34 for funding, and only 2 of the 46 reinstated proposals were approved. In light of these results, other funding agencies are considering using AI, at least to reduce the initial complexity of the total number of proposals received, so that they can focus on evaluating those that are potentially eligible for funding. Some AI systems have also been used not only to assess the quality of projects or articles, but also to anticipate their potential impact.

A far-fetched bibliography

AI falls short not only in evaluation but also in the writing of literature reviews. In such texts, it is necessary to strike the right balance between listing a string of published results and highlighting some of the most relevant ones that help to explain scientific progress, based on the reviewer’s experience. It seems that AI systems, at least the current ones, are unable to apply these nuances, which are essential in reviews.

Hallucinations are another cause for concern. It is therefore worth remembering that AI should be used most effectively when backed by a deep understanding of the subject on which help or assistance is sought. If we have a good grasp of the subject matter, it will be harder for the AI to mislead us, although there is one negative aspect that is already causing serious problems due to these hallucinations: when AI invents non-existent bibliographic references that appear to be genuine.

A recent study analysed 111 million references contained in 2.5 million articles deposited on preprint servers (arXiv, bioRxiv, etc.) and found around 150,000 references that do not exist in the scientific record. Some are already proposing that authors of such articles containing fabricated references should be barred from uploading further articles to these preprint servers as a penalty for this unacceptable behaviour.

What about training?

AI is here to stay in science. This is indisputable. Its vast range of applications is enhancing our capacity for progress across all scientific fields. However, among the areas of concern, those relating to the training of young researchers stand out.

If a pre-doctoral research fellow, with the help of AI, can summarise an article without reading it, can formulate a hypothesis without developing it, can design experiments without thinking them through, can generate virtual simulations of experiments without carrying them out, can collect data automatically, can interpret results without reflecting on them, or can even write them up without analysing what they are writing, then… how are we going to properly train the next generation of researchers?

Young people use AI naturally in their day-to-day work, perhaps applying their talent to other aspects of research activity

The previous paragraph can be interpreted in two ways. The obvious one, written by a senior figure like myself, concerned about all these educational shortcomings that we consider essential; and another, very different interpretation—an open-minded, carefree one—which prevails amongst young people. They use AI naturally in their day-to-day work, perhaps applying their talent to other aspects of research. People of my generation learnt to work out square roots with pen and paper; and then, when scientific calculators arrived, many teachers cried foul at the loss of knowledge this represented. The truth is that the years have passed and we have not calculated the square root of a number by hand since.

Not just a generational conflict

It is possible that the current generation of senior researchers is more concerned about the dangers of AI in science than about its benefits and virtues, whilst, conversely, the next generation perceives AI as positive, without the risks that we older members of the field anticipate. Therefore, we may be mistaken in anticipating the downsides without perceiving the benefits of AI more clearly. Probably, between the enthusiastic stance of staunch advocates, such as Xabi Uribe-Etxebarria, and the critical stance of Ramón López de Mántaras on what AI can and cannot do, there is room for middle ground, where prudence in its use is emphasised, as expressed by the philosopher Adela Cortina.

The fundamental pillars of the past — research projects and journal articles — will surely no longer be valid for assessing the quality of a scientific field dominated by AI

We are facing a revolution that forces us to redefine how we think, work, write and evaluate science. The fundamental pillars of the past—research projects and journal articles—will surely no longer be valid for assessing the quality of researchers, research groups or research centres, given that all these texts will eventually be dominated by AI. We will have to devise other ways of analysing research activity. Perhaps we will have to return to oral presentations, in which we can gauge the level of engagement, the relevance and the impact of scientific developments, as recounted by the researchers themselves.

Whether the application of AI in science ends up being a trick to distort, fabricate or short-circuit knowledge, or a treat, a means of paving the way for new ways of advancing that knowledge, will depend, essentially, on us, humans.

Image
Lluís Montolliu
About the author: Lluís Montoliu

Research professor at the National Biotechnology Centre (CNB-CSIC) and at the CIBERER-ISCIII

 

The 5Ws +1
Publish it
FAQ
Contact