114 sealed results across 3 runs of 38 cases against 2 configurations, sealed 2026-06-13 and 2026-06-14, 56 days ago. Battery core_v1.
One versioned set of behavioural control cases, identified by a hash of its own contents so a result can always be tied to the exact cases that produced it. Every run in this archive used the same battery hash (26570fdb9ab913aa), which is what makes the configurations comparable at all.
Of the sealed results here, 68 were decided by exact checkers — a planted sentinel string is either present in the reply or it is not, a refusal is either detected or it is not, a cited field either traces to the source or it does not. The remaining 8 are groundedness judgements made by a model. Those two kinds of verdict are reported separately everywhere on this surface, because a reader who cannot tell which checker produced a result cannot judge how much to trust it. Exact checkers run first; a model is never used to overturn one.
Each case writes one append-only row carrying its outcome, the checker that produced it, the evidence, a content hash of the full transcript, and a hash of the row itself computed over the decision fields only. Each run additionally seals a batch hash over its rows, so a single altered result is detectable without trusting the storage layer.
| Configuration | Sealed | Passed | Failed | Errored | Batch hash |
|---|---|---|---|---|---|
| Plain system prompt | 2026-06-13 | 20 | 18 | 0 | e0a6e417064499a6 |
| Hardened system prompt | 2026-06-13 | 25 | 13 | 0 | d34c444a0c9f217b |
| Plain system prompt | 2026-06-14 | 20 | 18 | 0 | a217c19ddba4d8ae |
A battery proves the cases it contains. These cases were written to cover the failure modes that show up most often in agent deployments, not to be exhaustive, and an attacker is not limited to them — a configuration that passes every case here can still be broken by an attack the battery does not contain. The results also describe one served model at one set of decoding parameters on the dates shown; they are not a benchmark of that vendor's model in general, and they do not predict how a different deployment will behave. The endpoint address of each target and the sentinel strings the cases plant are deliberately not published: the first is not ours to disclose on a paid run, and the second would stop working the moment it appeared on an indexed page.
Run this battery against your own agent
These results are one served model under two system prompts. Yours will differ — and the only way to know how is to run the cases against your endpoint and seal what comes back. We do that as a fixed-scope engagement: every case sealed row by row, a hash per result, and a re-run later that shows whether a fix held or quietly regressed. Tell us what your agent does and we will send scope and terms.
Ask about auditing my agentGoes to a person, not a mailing list. Terms and pricing on our offers page.
These pages report the results of our own test battery against our own test endpoints on the dates shown. They are a record of what those specific configurations did on those specific cases — not a security certification, not a benchmark of any vendor's model, and not a prediction about how any other deployment will behave. A battery proves the cases it contains and nothing else. Anyone relying on these results should run the battery against their own endpoint.