Use cases
One loop, three kinds of buyer.
The gym is the same everywhere: an agent and a goal go in, an evaluated policy comes out. What changes is the workload.
Language models
For teams improving the model itself. Rubrics, checkers, and promotion gates instead of preference guesswork — on your hardware.
Read →LLM agents
Browsers, terminals, codebases, ticket queues — trained against stubs of your real systems and scored on whether the work got done.
Read →Robots & embodied AI
Humanoids, arms, quadrupeds, drones — trained in simulation at scale. A robot model and a task go in; evaluated policies come out.
Read →Game NPCs
Behavior inside engines and open worlds — single characters or whole populations that learn instead of being scripted.
Read →Multimodal models
Image, voice, and video models trained in loops where quality is a measurement, not a vibe — on the hardware where your data lives.
Read →Bio & science agents
Agents that drive simulators, lab protocols, and discovery loops — domains where success is read off an instrument.
Read →Multi-agent systems
Markets, fleets, teams — cooperating and competing populations where the system, not the individual, is what has to work.
Read →Enterprise agents
Agents put into real work — healthcare, finance, logistics, operations. Trained on your systems, proven against your definition of success.
Read →Work with us
Bring an agent and a goal.
We are taking a small number of design partners. The bar is a real workload, not a logo. Everything runs on your hardware, and you keep the machine.