Use case · LLM agents
Agents that operate software, proven on yours.
Browsers, terminals, codebases, ticket queues — trained against stubs of your real systems and scored on whether the work got done.

The shape of the work
Public agent benchmarks top out at tasks measured in minutes; your workload is measured in hours and days. The gym builds the environment from your actual stack — the tool layer stubbed from real APIs — and trains against long-horizon tasks with outcomes as the score. What ships is an agent with evidence, not a demo with luck.
01
Environments from your stack
The agent practices the actual job against stubs of the real interfaces.
02
Long-horizon tasks
Multi-step work with checkpoints and resets — the tasks no public benchmark contains.
03
Scored on outcomes
Did the ticket close, did the code merge, did the number reconcile. Not plausibility.
04
Reviewable everything
Every trajectory, tool call, and change is on disk, diffable, and auditable.
Work with us
Bring an agent and a goal.
We are taking a small number of design partners. The bar is a real workload, not a logo. Everything runs on your hardware, and you keep the machine.