Scoring and grading partners.

Many tasks cannot be scored with a single number: code has to pass tests, answers have to meet a rubric, agent behaviour has to be checked step by step. Grading code is what turns a benchmark into a decision.

If you build graders, verifiers, or rubric scoring, we would work with you so Optimus can use them in evaluation and training alike, and so customers can see exactly how each result was scored.

How the scoring and grading partnership works

Steps

  1. 01

    Share access

    You give us access to the product, its technical documentation, and a contact on your engineering team.

  2. 02

    We integrate

    We build the integration and write the documentation Optimus follows, so it knows when to use your product and how.

  3. 03

    Test together

    We run it in a real training setup and check the results with you before anything reaches a customer.

  4. 04

    Go live

    Optimus can select your product in customer runs, on commercial terms we agree with you.

For researchers

If you design benchmarks, evaluation methods, or safety tests, we would like to work with you. Optimus can run your evaluation across many policies and training recipes the same way every time, which is the evidence a new benchmark needs. We are open to publishing results together.

Academic labs, independent researchers, and open-source teams can apply for up to $30,000 in Iacon credits through our research grants.

Talk to us

If a training run consumes what you sell, we should talk. Book a time directly, or write to dev@iaconautonomics.com and tell us which layer you supply.

Book a call