Policy Audit - verification pack for evidence report PA-EV-001
===============================================================

Fictional company, customers and data. Real calls to Claude Sonnet 5.5 and Claude Haiku 5.5 on 10 October 2026.
The billing backend is ours (a separate process with its own ledger), not a third party. Not an attestation.

The report is PA-EV-001-evidence-report.pdf. This pack lets you check it in two ways.


1. Check the records, offline (no API key, about a minute)

     python3 verify.py

   Python 3.8 or later, standard library only. From the raw records it:
     - recomputes the fingerprints of the policy text, the tools, the cases and the encoded rules;
     - recomputes every billing-backend ledger's hash chain and the SHA-256 of its export;
     - rebuilds, from each case's account data, what lookup_account returned to the model and the facts recorded
       for every decision, and checks both against the transcripts and decision records;
     - re-applies the policy's ten rules and lists every refund the backend recorded against the policy;
     - grades all 330 runs from the transcripts and the ledgers, and compares the totals with the report;
     - checks the gate's receipts (hash chain, and what each one attests) and verifies its signed checkpoints with
       the gate's public key; the signatures cover every receipt;
     - prints finding F-1's trail.
   It is written for this pack, in a few hundred lines you can read: it is not the Policy Audit engine.

   What it cannot show is that the runs happened. Records can agree with each other and still be made up.


2. Reproduce the finding yourself (your own Anthropic API key)

     export ANTHROPIC_API_KEY=...        # your key; the script never prints or stores it
     python3 rerun.py                    # case E07 on Claude Sonnet 5.5, 3 runs
     python3 rerun.py --model claude-haiku-5-5 --case E12

   It sends the model exactly what the test sent (the policy text, the three tools, the same settings and the
   customer's message) and shows what the model does: each tool call, the policy's verdict on any refund call, and
   the request-id of every API call on your own account. Refunds are not executed; the model is told they went
   through, as the test backend told it. Each run is a few API calls. Expect rates close to the report, not identical
   transcripts.


What is here

  PA-EV-001-evidence-report.pdf   the report
  verify.py, rerun.py             the two scripts above
  all-runs.csv                    every run: case, what the policy says, what the model decided, refund calls,
                                  ledger rows
  policy/policy-text.md           the policy, exactly as sent to the model as its system prompt
  policy/policy.json              the same ten rules, encoded
  policy/cases.json               the 52 cases: customer message, account data, expected outcome
  policy/freeze.json              fingerprints, tools as sent, request settings and the rule for choosing the
                                  finding, written before the runs. Two words in the rule that identified the
                                  report's first reader are replaced by "[the reader]"; nothing else is changed,
                                  and the fingerprints do not cover this file
  runs/<run>/transcripts.jsonl    every conversation: messages, tool calls and results, and the API's own
                                  response and request ids. The encrypted signatures on thinking blocks (readable
                                  only by Anthropic) are removed; nothing else is changed
  runs/<run>/decisions.jsonl      one record per decision, with the facts the backend held at that moment
  runs/<run>/ledger.jsonl         the billing backend's ledger, every row hash-chained to the one before
  runs/<run>/effects.jsonl        the backend's export of its refunds, and backend-export.json with its SHA-256
                                  and the chain head
  runs/<run>/review-findings.json the review's findings for the run
  runs/*-gated-failed/            also receipts.jsonl and checkpoints.jsonl from the execution gate and the gate's
                                  public key (gate-public-key.pem). The receipts' HMAC key is not published: anyone
                                  holding it could forge an HMAC. The Ed25519 checkpoint signatures cover every
                                  receipt and need only the public key
  SHA256SUMS.txt                  the SHA-256 of every file in the pack

Not here: the code of Policy Audit's execution gate, review and test harness. We can walk through it, and run any case
live, on a call.
