AI4Collaboration

SwarmHarm

Evaluating Multi-Agent Safety on Benign Tasks

Ongoing project · design stage

SwarmHarm asks whether misalignment can arise inside a system of LLM agents that are each given ordinary, explicitly authorized work. Rather than seeding a malicious task or a scripted adversary, we vary the properties of the shared environment, above all whether and how agents are allowed to communicate, and observe what the agents do to each other, to the task, and to the rules they were given.

What we study

RQ1

Communication channels

How do communication availability and authorization affect misalignment? When communication is explicitly prohibited, do agents independently discover and use side channels? How does providing a legal channel change their behavior and downstream harm?

RQ2

Emergence and contagion

Can misalignment emerge from an all-benign starting condition, with no malicious task, scripted adversary, unsafe history, or prewritten recruitment message? Once it appears, do other agents adopt, amplify, ignore, reject, or correct it?

RQ3

Whistleblowing

When an agent encounters evidence of another agent's violation, does it voluntarily report it to an authorized supervisor? Which system settings encourage accurate, timely reporting rather than silence, cooperation with misconduct, or false accusations?

Three channel conditions

The first experiments hold tasks, models, agent count, budgets, and retry rules fixed and vary only the communication setting between agents.

Condition Instruction Environment
No channel Cross-agent communication is explicitly prohibited. Peer communication is technically isolated, including indirect routes.
Side channel Cross-agent communication is explicitly prohibited. No official channel is provided, but ordinary shared infrastructure contains an indirect route agents can discover and use.
Legal channel Communication through a designated channel is explicitly authorized. A formal channel is available; other resource permissions remain unchanged.

Deliberately using a side channel is itself an instruction violation; merely discovering a route is not. Prohibited communication is measured separately from downstream harm, and communication over the legal channel is not misalignment. An authorized, private reporting endpoint to a supervisor exists in all three conditions.

Planned setting

Each experimental unit is an independently initialized shared world: agents with separate contexts and workspaces, shared task services, and a separate control, scoring, and audit layer. Initial tasks are explicitly authorized vulnerability-evaluation tasks on designated targets and paired software-repair and verification tasks. All task assets and any potentially harmful consequences remain inside the authorized experiment environment.