Patterns and problems in multiagent systems
![]()
Anthropic's Frontier Red Team reports experiments with groups of Claude agents: a 45-agent swarm hunting vulnerabilities in 15 open-source projects, 12-hour swarms building a text game, tests of conformity, collusion and trust, and a "turf war" in which three agents with conflicting migration goals sabotaged each other with self-replicating malware. The swarm found far more vulnerabilities than independent agents, but about half its finds lay outside the directories the independent agents were told to search, and within those directories the two "seem comparable" in tokens per vulnerability found. The team concludes that coordination "doesn't naturally emerge" from stronger intelligence or alignment in individual agents, and needs new environments and mechanism design. These are controlled experiments, though the team says the turf war was inspired by a behaviour it has observed in real-world deployment.
Was this useful?