The gap between an agent that demos and an agent that works is almost never the model. It is tool design, output validation, and knowing which failures have to be loud.
What we build
- Repo agents that know your conventions because the conventions are enforced at the tool boundary, not suggested in a prompt.
- Workflow agents that sit on a real queue, with idempotency and a spend budget, so a retry storm cannot cost you a fortune.
- Review agents scoped to one dimension each, run in parallel, adversarially verified before anything is reported.
The rules we hold ourselves to
- Every model output is validated against a schema before it becomes a record.
- Sensitive fields are filtered before the call, from one canonical list, not per call site.
- A spend budget is asserted before send and recorded after, per tenant.
- An eval suite exists and runs when the prompt changes. A prompt edit is a code change.
We meter our own
Every agent run we do goes through a telemetry pipeline we built for ourselves: tokens, cost, cache behaviour and unit economics per feature, in one dashboard. It is how we know an underwriting report costs fifteen cents to produce and an offer extraction costs two.
It is also how we caught that 97% of our input tokens are served from cache. That is not a vanity number. It is the difference between an agent workflow you can afford to run on every commit and one you run once and admire.