About the project.
Policy Audit tests whether AI agents’ tool calls follow their written policy. It is run by Manan Awasthi, a sole proprietor in India, and grew out of independent research.
This is an early beta. There are no customer pilots yet, the public sample uses made-up refund cases, and the recorded tests use fictional cases: one with a small open-weights model, one with two Claude models. Features, wording and results on this site may change.
Why it exists
Agents that issue refunds, change accounts or approve requests are usually checked on what they say. Policy Audit checks what they do: it compares the approved rules with the tool calls the agent attempted and with a separate record of what actually changed.
The project grew out of research on whether tool-using language models apply a written policy to the facts in front of them. In the recorded experiments they often do not do this reliably, which is why the product checks behavior from the outside instead of trusting the model’s explanation.
What is ready and what is not
- Ready: the public sample audit, the recorded Qwen3-4B comparison and its downloadable run data, and a real-model evidence report with two Claude models and a verification pack.
- Available: tool-call audits for one action tool (USD 499) and assisted pilots for up to three action tools (USD 2,500, plus USD 499 for each further tool). Refund policies import automatically; other tools are mapped with you during setup. The execution gate is not publicly available.
- Not yet: any customer pilot result.
Pricing and requests
Audits are open: USD 499 for one action tool, or USD 2,500 for an assisted pilot covering up to three action tools, plus USD 499 for each further tool. See the terms, and use the request form or the contact page.
Please do not send customer data, API keys or production logs by email.
Policy Audit is run by Manan Awasthi. Contact: mananawasthi@yahoo.com.