[blog]

What we learned taking our first agent through assurance

Notes from a regulated client where an operations agent passed internal review before scaling.

Published 2 December 2025 · LoxiLabs

The task

A high-volume operations process: receive a document, extract what matters, check it against policy, prepare the case, route it. People did it in a few minutes each, thousands of times a week. An agent could do most of it — and the client's assurance team had never reviewed one before.

What the reviewers asked for

Not accuracy numbers. They asked: what can it do on its own, what needs a human, what happens when it's wrong, how would we know, and how do we switch it off? Every one of those questions is answered by operations design, not by the model.

What we built to answer them

A golden set of several hundred cases run on every change. Traces for every run with cost, latency, and outcome. Hard spend limits per run. An approval gate for any action with external effect. A one-step rollback. And a runbook owned by the operations lead, not by us.

The result

The review passed. The agent now handles most cases end to end; the rest go to people with the preparation already done. The client reports the outcome in hours returned and cycle time, which is what they asked for at scoping.

Those controls became the first version of our AI Agent Ops Framework.

[related]

[share]

Copy the address bar link, or send this article to a colleague who owns the metric.

Ready to move from pilots to P&L?

Tell us about the decision or workflow you want to change. We'll come back with an honest view on whether it's worth proving, and what it would take.

Start a conversation