case study · AUG 27, 2026

Training insurance agents against the systems they actually operate

Production insurance agents, trained and evaluated against the systems they actually operate.

Harper

An insurance agent that works in a demo is not the same artefact as one that works in production. The difference is everything the demo left out: the policy administration system it has to read, the underwriting rules it cannot break, and the cases that only appear at the tail of a real book.

Training against a simplified stand-in produces an agent that is good at the stand-in.

The environment is the real stack

The work is an environment built from the systems Harper already operates, rather than a simulation of them. The agent meets the same interfaces, the same constraints, and the same failure cases it will meet in production, and it is scored on those.

What is recorded

ArkenOS retains the task definition, the runs, the evaluator revision, and the artifacts each candidate produced, so a change that improves one class of case can be checked against the ones it might have cost.

What is not published yet

This study describes the shape of the work and no more. The engagement detail, the evaluation boundary, and any figures follow once Harper has reviewed them and agreed to what is said here — and nothing above should be read as a performance claim in the meantime.

Read next