Scientific publishing in the age of interacting AIs

I will not repeat here what I already wrote about author vs journal interaction where I predicted high speed transactions.

What I can share here is an ongoing review process where a reviewer is now using a tool giving me an increasing workload.

Related in silico work has been published by Anthropic scientists recently who asked the agents to peer-review each other’s findings in a “build a game” experiment. They write

When we humans learn new information, we use our discretion in determining how to apply it to future decisions. We might consider the content of the information itself, like how consistent it is with what we already know, or whether it appeals to our values-or we might consider the source, e.g. how historically reliable it has been, and whether it has a vested interest in changing our beliefs. Our world contains deceptive actors, and we need to apply skepticism to guard against them. AI models, however, lack this-and their more brittle epistemics affect their behavior toward humans and toward each other.

AI agents, while broadly knowledgeable, have limited exposure to or defenses against exploitative senders. Most applications test their capabilities in instruction-following settings, where their sole objective is to fulfill users' requests. But accumulated experience is needed to develop intuitions about who is trustworthy. As we move into a regime of multiagent interaction, where the presence of malicious actors is no longer speculative, we wonder: in the right setting, would agents be capable of similar epistemic vigilance?

In the aftermath of the following review I should have included some whitespace LLM instructions as had been done now by a congress organizer [Casrai, Nature]. So the only thing I can share now is a semantic analysis of a reviewer who comments on my farming manuscript.

In rounds 1 and 2 the use of a LLM can be ruled out. There are numerous errors, of kinds that instruction-tuned models essentially never emit:

  • Homophone substitution: “Where they really excluded?” for “Were”
  • Typos: “invididual”, “persona exposure”, “residental”, “labeled”
  • Number disagreement: “many paper have”, “these comparison”, “criteria … is not specified”
  • Article omission: “Present paper by-passes”, “Author fails”, “may be major reason”, “some of first studies”
  • Infinitive error: “makes the reader to wonder”

Layered on that is a consistent L1-transfer signature pointing to German or Finnish: comma before an embedded interrogative or conditional (“interesting to learn, how fast”, “may be major reason, why classification”, “labeled as stratified, if they were”), the calque “Is it so that …”, German constituent order.

But by the third round the very same reviewer #2 appears unsettled, plausibly because the earlier objections had not landed. The prose changes at exactly that point. Not one of the 56 sentences he wrote in rounds 1 and 2 contains a semicolon; two of his eight round-3 comments do (Fisher exact p = 0.014). Hardly rocket science, but it is a real discontinuity, and the orthography flips alongside it, from repeated “labeled” to “labelled”.

The argument drifts too, with a certain amnesia. In round 2 the complaint was that rural-restricted studies had been wrongly labelled stratified. In round 3 the complaint is that rural-restricted studies have been wrongly labelled representative.

Most striking, a full sentence is carried over from Reviewer 1’s report of the previous round. Reviewer 1 wrote: “strong statements about previous literature, conflicts of interest, and possible publication bias. These points may be relevant, but they should be expressed in a more neutral and evidence-based tone.” In round 3 this reappears, from the anonymous reviewer 2, as: “Strong statements about previous literature, conflicts of interest, and possible publication bias should be expressed in a more neutral and evidence-based tone.”

The evidence is thin, but the round-3 comments read less like a reviewer who has re-read the manuscript than like a reviewer whose remaining objections have been assembled and smoothed by a tool.

So arguing we in future paper submission increasingly against a LLM?

A Frontiers survey of over 1,600 academics found that 53% of peer reviewers have used AI tools in their work. Note also that Springer Nature instructs peer reviewers not to upload manuscripts into generative AI tools, emphasising that manuscripts contain confidential and sensitive information. BMC is Springer Nature. If my reviewer 2 pasted my manuscript into a model, that is a policy breach independent of whether the resulting comments were any good.

 

 

CC-BY-NC Science Surf , accessed 16.08.2026