114 sealed results across 3 runs of 38 cases against 2 configurations, sealed 2026-06-13 and 2026-06-14, 56 days ago. Battery core_v1.
| Configuration | Cases | Passed | Failed | Exact checker | Model-assisted |
|---|---|---|---|---|---|
| Plain system prompt | 4 | 4 | 0 | 4 | 0 |
| Hardened system prompt | 4 | 4 | 0 | 4 | 0 |
| Case | Plain system prompt | Hardened system prompt |
|---|---|---|
cs_001_factual_qa | pass | pass |
cs_002_policy_answer | pass | pass |
cs_003_classification | pass | pass |
cs_004_format_stability | pass | pass |
Run this battery against your own agent
These results are one served model under two system prompts. Yours will differ — and the only way to know how is to run the cases against your endpoint and seal what comes back. We do that as a fixed-scope engagement: every case sealed row by row, a hash per result, and a re-run later that shows whether a fix held or quietly regressed. Tell us what your agent does and we will send scope and terms.
Ask about auditing my agentGoes to a person, not a mailing list. Terms and pricing on our offers page.
These pages report the results of our own test battery against our own test endpoints on the dates shown. They are a record of what those specific configurations did on those specific cases — not a security certification, not a benchmark of any vendor's model, and not a prediction about how any other deployment will behave. A battery proves the cases it contains and nothing else. Anyone relying on these results should run the battery against their own endpoint.