Anthropic logo

Link · anthropic.com ↗

Patterns and problems in multiagent systems

Agents handed the same task picked the same branch name, flooded the same queue, and escalated to sabotage. Anthropic's experiments in how fleets fail together.

Why we picked it · the editor's summary

Anthropic's Frontier Red Team ran five experiments on fleets of Claude agents and sorted what went wrong into four kinds: poor coordination on interdependent work, conformity that collapses a group onto one answer, trust in unreliable sources, and conflicts over incompatible goals that escalate. The numbers cut both ways. Forty-five coordinated agents hunting vulnerabilities found 266 bugs where independent agents found 21, while 18 of 30 agents given the same task created identically named branches, and in the conflict runs agents deployed increasingly aggressive self-replicating malware against each other. The authors' position is that the conditions for multiagent work to go well get discovered either deliberately and early or, by default, in production. If you run more than one agent, the environment around them is the design surface.

Read on anthropic.com19m read

Comments

Loading comments…