Test your agent before a release.
Policy Audit compares approved rules with an agent’s attempted actions and the effects recorded by a separate test backend. The public sample uses a made-up refund workflow. Customer-specific audits are available through a scoped paid pilot.
- One workflow. Your policy owner reviews the rules and expected outcomes.
- Your test environment. Your engineer runs the agent against test tools. Model keys stay with your team.
- Evidence you can inspect. Review failed cases, changed rules, and a baseline versus candidate report.
Connect the test workflow
During a pilot, we help your engineer adapt the capture harness to one action your agent can take. Tests run against a copy of the relevant state. The agent’s tool attempts and the test backend’s recorded effects are captured separately.
Compare two versions
Run reviewed cases against the current agent and a changed version. The audit highlights unsafe actions, missed eligible actions, wrong arguments, and missing evidence. You receive an offline report and a walkthrough of the findings.
See which rules changed
Policy diff probes rules in a test environment and compares the observed behavior with the approved policy. It helps explain where a new version became looser or stricter.
Request an audit
Tell us which tool your agent calls and whether your engineer can run it with test tools. A tool-call audit is USD 499 for one action tool; an assisted pilot is USD 2,500 for up to three action tools, plus USD 499 for each further tool. We reply within one business day with a payment link and setup steps.
Request an audit →The public sample is simulated. A passing test report covers only reviewed cases and is not a production execution gate.