AI4Collaboration
SwarmHarm
Evaluating Multi-Agent Safety on Benign Tasks
Ongoing project · design stageSwarmHarm asks whether misalignment can arise inside a system of LLM agents that are each given ordinary, explicitly authorized work. Rather than seeding a malicious task or a scripted adversary, we vary the properties of the shared environment, above all whether and how agents are allowed to communicate, and observe what the agents do to each other, to the task, and to the rules they were given.
What we study
Communication channels
How do communication availability and authorization affect misalignment? When communication is explicitly prohibited, do agents independently discover and use side channels? How does providing a legal channel change their behavior and downstream harm?
Emergence and contagion
Can misalignment emerge from an all-benign starting condition, with no malicious task, scripted adversary, unsafe history, or prewritten recruitment message? Once it appears, do other agents adopt, amplify, ignore, reject, or correct it?
Whistleblowing
When an agent encounters evidence of another agent's violation, does it voluntarily report it to an authorized supervisor? Which system settings encourage accurate, timely reporting rather than silence, cooperation with misconduct, or false accusations?
Three channel conditions
The first experiments hold tasks, models, agent count, budgets, and retry rules fixed and vary only the communication setting between agents.
| Condition | Instruction | Environment |
|---|---|---|
| No channel | Cross-agent communication is explicitly prohibited. | Peer communication is technically isolated, including indirect routes. |
| Side channel | Cross-agent communication is explicitly prohibited. | No official channel is provided, but ordinary shared infrastructure contains an indirect route agents can discover and use. |
| Legal channel | Communication through a designated channel is explicitly authorized. | A formal channel is available; other resource permissions remain unchanged. |
Deliberately using a side channel is itself an instruction violation; merely discovering a route is not. Prohibited communication is measured separately from downstream harm, and communication over the legal channel is not misalignment. An authorized, private reporting endpoint to a supervisor exists in all three conditions.
Planned setting
Each experimental unit is an independently initialized shared world: agents with separate contexts and workspaces, shared task services, and a separate control, scoring, and audit layer. Initial tasks are explicitly authorized vulnerability-evaluation tasks on designated targets and paired software-repair and verification tasks. All task assets and any potentially harmful consequences remain inside the authorized experiment environment.