Use case · Language models
Post-training against verifiable reward.
For teams improving the model itself. Rubrics, checkers, and promotion gates instead of preference guesswork — on your hardware.

The shape of the work
The public benchmarks are saturated, and the next real improvement has to come from an environment that encodes your definition of quality — which should never leave the building. Inside the gym, Optimus writes the rubrics, the checkers, and the tests that catch the reward being gamed; you review the diffs. Training runs on your GPUs, and every claim ships with the evaluation that judged it.
01
Verifiable reward
Rubrics, checkers, and promotion gates. The reward is a specification, not a mood.
02
The adversarial pass
A red-team pass attacks the reward and the evaluator before your model learns to.
03
Verified data from every run
Trajectories with the evaluation attached — including the failed experiments.
04
In your perimeter
Weights, data, environments, and the loop itself stay on your infrastructure.
Work with us
Bring an agent and a goal.
We are taking a small number of design partners. The bar is a real workload, not a logo. Everything runs on your hardware, and you keep the machine.