Reward function and success criteria partners.
What an agent learns depends on how it is scored. A reward that is easy to exploit produces an agent that exploits it, and a success criterion that is too loose hides failures. Designing rewards and checks is one of the hardest parts of building an environment.
If you specialise in reward design, verifiers, or task specification for a domain, we would like to work together so Optimus can use your rewards and success checks when it builds environments in that domain. Customers get environments that measure what they mean to measure.

How the reward function and success criteria partnership works
Steps
- 01
Share access
You give us access to the product, its technical documentation, and a contact on your engineering team.
- 02
We integrate
We build the integration and write the documentation Optimus follows, so it knows when to use your product and how.
- 03
Test together
We run it in a real training setup and check the results with you before anything reaches a customer.
- 04
Go live
Optimus can select your product in customer runs, on commercial terms we agree with you.
For researchers
If you research environment design, generalisation, or how agents should be tested, we would like to collaborate. Use Optimus to build and run the environments your experiments need, and bring us the questions you want answered. We are open to publishing results together.
Academic labs, independent researchers, and open-source teams can apply for up to $30,000 in Iacon credits through our research grants.
Talk to us
If a training run consumes what you sell, we should talk. Book a time directly, or write to dev@iaconautonomics.com and tell us which layer you supply.
Book a call