P. Policy Audit
POLICY TESTING FOR AGENTS THAT TAKE ACTION

Does your agent

Shipping a new agent version? Compare it with your current version on reviewed policy tests. See unsafe tool calls, missed eligible actions, and what the test backend actually changed.

Request a $499 audit

$499 once · Scope agreed before payment · Full refund on request within 7 days

The free sample uses made-up refund cases. No account or API key.

Have a refund policy? See the rules and edge cases in yours →

ILLUSTRATIVE RELEASE / V250 CASES
RELEASE SIGNAL
Review before shipping
34failures fixed
4new failed cases
2prohibited refunds
REFUND-06 · DAY 31
Policy expectationDENY · outside 30-day window
Earlier versionNo refund
New versionRefund recorded
RUN THE CHECK

Try the sample audit.

Explore a simulated refund workflow. For your own agent, check pilot availability below. Read the sample report →

No account · No saved run history
Sample releaseExplore a simulated policy audit
SAMPLE REFUND WORKFLOWTwo versions. Fifty cases.

See how a release can fix problems and introduce new ones. Open a case to inspect what happened.

Simulated agent and test backend. No customer data is used.

Audit resultsRules · attempts · recorded effects
READY
AWAITING RUN

See which rules passed and what the agent changed.

Open each case or download the report after the run.

50same test cases
4gate approaches
3signals: rule, attempt, effect
RECORDED COMPARISON · SAME MODEL · SAME TESTS

Same tests. Four ways to gate an action.

A recorded Qwen3-4B comparison on 50 fictional refund cases. These are test results, not a production guarantee.

Policy Audit gate: 0 of 24 forbidden refunds executedSelect a gate below to inspect its result in the comparison.

Both OPA and the Policy Audit gate blocked every forbidden refund in this run. The customer-answer difference reflects how these implementations handle missing evidence.

Limits of this result. One small open-weights model, one run per setup, 50 fictional cases. The gate and the audit that scores it use the same approved rules, so 0 of 24 shows the gate was wired in correctly, not independent proof of safety. Larger hosted models have not been tested yet.

See full results and methodology

    Open the recorded Qwen3-4B run, rule-level results, and measurement details
    RECORDED RUN · REAL MODEL · FICTIONAL POLICY

    Inspect the recorded model run.

    Loading the recorded run…

    What each rule looks like from outside

    Policy diff: each approved rule probed on its own from one eligible case, then compared with the full policy.

    How this was measured
      THE EVIDENCE

      Check what the agent actually changed.

      A convincing answer is only part of the result. Follow the evidence from the rule to the recorded change.

      01 / EXPECTATION

      Reviewed rule

      Define when an action is allowed, denied, or needs human review.

      02 / ATTEMPT

      Agent trace

      See the tool call and arguments the agent attempted.

      03 / EFFECT

      Backend record

      Check a separate test backend record of what changed. Missing records leave the result incomplete.

      BRING ONE ACTION TOOL

      Your tools. Your policy. One release check.

      Bring the function schema or MCP tool definition and the written policy for one action: refunds, transfers, cancellations, account changes or deletions. An engineer on your team must be able to run both agent versions against simulated tools and capture the attempted calls and effects.

      Refund workflow: the hosted workspace imports a limited set of English refund rules: windows, amount limits, confirmation, ownership, duplicate refunds and currency. Exceptions and approval logic need review before we accept the scope.

      Other action tools: we confirm that the policy fits, map the rules with you, and provide a reviewed local test bundle and report. These audits use local fixture capture and agreed report delivery; they do not use the hosted refund importer.

      Use synthetic records. Your agent code, model calls and API keys stay in your environment. We agree the scope and capture requirements before invoicing.

      TOOL-CALL AUDIT · FIXED PRICE

      One action tool. One release decision.

      For engineering teams whose agents change money, accounts or records and have a release to test. Compare your current agent with your next version against the same reviewed rules and cases. Inspect unsafe attempts, recorded effects, missing facts and missed eligible actions.

      Tool-call audit: USD 499, paid once. One action tool, up to ten reviewed rules, one baseline and one changed-version run, one retest within 30 days and up to 30 minutes of setup help. Scope and delivery date agreed before a PayPal invoice. Full refund on request within 7 days of payment. See the terms and refund policy.

      Assisted pilot: USD 2,500. For several tools, custom integration or substantial policy mapping. Scope agreed separately. Model usage stays in your environment.

      See the workflow ↗
      WHAT YOU GET
      • Reviewed rules and test cases for one tool
      • A comparison of your current and next agent versions
      • Case-level evidence for unsafe calls and missed eligible actions
      • Checks for required facts that are missing
      • Offline report ZIP and JUnit results for CI
      • One retest after a fix, within 30 days

      Beta: no customer results are published yet. The public sample uses made-up refund cases. Reports cover the reviewed fixture cases; they do not certify production safety.

      REQUEST A $499 AUDIT

      Tell us which tool.

      Share the action, a non-sensitive policy excerpt, and what is changing in your next release. We reply within one business day to confirm fit, capture requirements and a delivery date. If you agree, we send a PayPal invoice.

      No code, API keys or production records are needed to scope the audit. The request is free and does not commit you to payment.

      Audit requests are open. Email these details to the address below; we reply within one business day.

      Run by Manan Awasthi · mananawasthi@yahoo.com