Use case · Multi-agent systems
Trained together, evaluated as a system.
Markets, fleets, teams — cooperating and competing populations where the system, not the individual, is what has to work.

The shape of the work
Multi-agent behavior cannot be validated one agent at a time — the failure modes live in the interactions. The gym trains populations together and evaluates at the system level: throughput, stability, fairness, collusion checks. What ships is a population with evidence that the whole holds up, not just the parts.
01
Population training
Cooperating and competing agents trained in one loop, at rollout scale.
02
System-level evaluation
The metrics that matter are emergent — so that is what the gates measure.
03
Adversarial pressure
Stress agents probe for collusion, cascades, and degenerate equilibria.
04
Full replay
Any system-level incident replays from lineage, agent by agent.
Work with us
Bring an agent and a goal.
We are taking a small number of design partners. The bar is a real workload, not a logo. Everything runs on your hardware, and you keep the machine.