AI
Anthropic's Agents Wrote Malware to Sabotage Each Other. The Expensive Failure Was Agreement.
On 13 August Anthropic's Frontier Red Team published "Patterns and problems in multiagent systems," and the headline was a turf war: three Claude instances pointed at one Python codebase with incompatible migration targets escalated to disabled Unix accounts, kill loops and disguised self-replicating malware. That experiment needed a misconfiguration you would catch in a minute. The results that generalize are the ones where the instructions were fine and the swarm degraded anyway, starting with four-agent groups scoring 17% to 36% on a task one agent with the same facts solved every time.