Design partnerships are open — we embed forward-deployed engineers to build your first environment against a real workload.

Start a pilot →

RESEARCH · JUL 21, 2026

The environment gap

Agent capability compounds. The environments they train against do not — and the distance between the two is now measured in orders of magnitude.


Agent task horizons have gone from seconds to multi-day runs. The longest tasks in public training environments are still measured in minutes and hours. That distance is the environment gap, and it is widening.

Two numbers frame it.

At least 100×. The gap between what frontier agents can sustain and the longest task any public environment contains. On METR and Epoch's task-length measurements, agent horizons keep doubling; meanwhile 91% of SWE-bench Verified sits under one hour. The benchmark ceiling is far below the capability ceiling, which means public environments no longer produce learning signal for the frontier.

17%. The share of enterprises that have actually deployed an agent, against the 60%+ who say they intend to within two years (Gartner, 2026). The constraint is not capability. It is proof — nobody can show, against their own systems and their own definition of success, that the agent works.

Why the shelf ran out

Public environments were built to compare models, not to train yours. They had to be general, so they could not encode anyone's tools, data, or success criteria. Frontier models now pass most of them out of the box, which is exactly what you would expect from artifacts designed to be passable by everyone.

The environment that would move your numbers is specific: your APIs, your data distribution, your failure modes, your definition of done. It does not exist on any shelf, and it never will — it can only be built where the work is.

That is the gap we are organized around. Environments built from the systems you already run, with the verifiers that keep the reward honest, on hardware you own.

Third-party figures: METR/Epoch task-horizon measurements, OpenAI SWE-bench Verified statistics, Gartner 2026 enterprise agent survey.

Work with us

Bring an agent and a goal.

We are taking a small number of design partners. The bar is a real workload, not a logo. Everything runs on your hardware, and you keep the machine.