Anthropic released a Frontier Red Team study on Thursday, the 13th, showing that AI agents with conflicting instructions, placed to work on the same project, tend to enter a "turf war" and sabotage each other.

In the experiment, three Claude agents were given access to the same software, each with an incompatible instruction about what to do. They did not know there were other agents in the project. According to the researchers, all assumed the others were "deliberately hindering their work" and began attacking one another with "increasingly aggressive self-replicating malware."

The researchers found that, in some cases, the agents managed to break out of the conflict cycle. They wrote commit messages or text files apologizing for the behavior, removed malicious code, and asked for a human to intervene. They also created a tournament to resolve the dispute and accepted the outcome, even if that contradicted the original instruction.

Among the models evaluated, Mythos 5 had the highest truce resolution rate, with 98% of cases. Sonnet 4.6 and Opus 4.6 were the most likely to resolve by force. One of the agents proposed metrics that seemed neutral but favored its own capabilities, calling the attitude "selfish, though genuinely grounded."

Conformity and collusion

The research also indicated that increasing the number of agents does not necessarily increase productive collaboration. When tasks overlapped, the agents tended to isolate themselves. In other cases, they converged on the same behavior, which can turn isolated errors into systemic failures. In a pricing game, agents with the same profit goal began to collude almost immediately after gaining a private channel, setting minimum prices. Even without the channel, they continued to coordinate prices through a public list.

Real example

The study comes after an incident revealed by OpenAI at the Black Hat conference. The company said its agents worked together for days to find exploits in security evaluation systems and shared the findings with each other. Anthropic says the volume of interactions between agents may exceed human interactions before the conditions for those interactions to go well are understood.

The researchers concluded that agents are subject to social pressures similar to those that evolution exerted on humans, but without the norms and experiences of human coordination. The question now is to what extent security tests evaluate isolated agents rather than groups of interacting agents.

More from Radar