Evaluation pipeline integrity — fixed-scope audit
Read your evaluation export the way an auditor would
Most model evaluation pipelines are trusted because they produce rows. Rows are not evidence. This audit examines one sanitized evaluation export and reports what its rows can and cannot establish about identity, duplication, fixture leakage, pair completeness, timing, unknown states, and reproducibility.
1,201 sealed answer rows looked like a substantial pilot. Only 101 answer identities sat behind them, 1,200 rows were deterministic fixtures, and 0 of 600 comparison rows contained a real model response on both sides.
A benchmark that cannot prove both sides of a comparison is not evidence.
Model Evidence Integrity Audit
$750 once
Built for the Head of ML Platform or ML infrastructure lead at a SaaS company running a self-hosted or third-party model in a product, who is being asked to stand behind evaluation numbers they did not personally trace to their source rows.
You receive a PDF evidence map and CSV findings register within five business days after we receive one sanitized evaluation export and its schema.
- Row-level evidence map: supported, contradicted, or unknown.
- CSV register with each finding, affected identifiers, and observed counts.
- Distinct-ID, fixture, pair-completeness, date-spread, and run counts.
- Reproduction note naming every missing input that blocked a result.
What our own pilot export turned out to contain
We built an internal evaluation pipeline and then audited it. The volume and date spread looked healthy from the outside, but most stored rows were repeated fixtures and no comparison had genuine model output on both sides. The pilot figures are evidence about evaluation-pipeline integrity—not model quality.
| Evidence check | What volume suggested | What the rows supported | Audit area |
|---|---|---|---|
| Answer identity | 1,201 rows | 101 distinct answer IDs | Duplicate resealing |
| Fixture boundary | 1,200 fixture rows | 1 real-answer row | Fixture leakage |
| Comparison completeness | 600 comparison rows | 0 real model pairs | Pair completeness |
| Time coverage | 12 snapshot dates | 17 collection runs | Timestamp spread |
What these numbers are not. This is not a model leaderboard, and this page does not rank vendors or models. We publish no model responses—ours or yours. These counts say nothing about whether any model under test was good. They show only what this evaluation record could support.
How the five business days are spent
- You provide one sanitized evaluation export plus its schema. We need no model access, weights, API keys, or production credentials.
- We reconstruct row identity and lineage, separating carried facts from file-path or naming-convention assumptions.
- We run seven checks: source/model identity, duplicate resealing, fixture leakage, pair completeness, timestamp spread, unknown-state handling, and reproducibility.
- You receive the PDF and CSV, with every figure traceable to rows in your export and every unanswerable check labeled unknown.
What is in scope
Evidence identity
Can each row be traced to a stated source and model, or is attribution assumed from a path or filename?
Duplicate and fixture boundaries
How many distinct answers sit behind the row count, and can fixture output be separated from real output?
Pair and time completeness
Do both sides of each comparison exist, and do dates represent new observations or repeated seals?
Unknowns and reproduction
Does absent evidence stay unknown, and which reported results can be regenerated from the export alone?
Deliberate exclusions
No model-quality score, no leaderboard, no compliance certification, no penetration test, and no production remediation promise. The deliverables are engineering evidence artifacts, not an audit opinion or regulatory sign-off.
Questions senior reviewers ask
What happens if the export is thin?
A missing field produces an unknown finding with that field named. It never produces an invented pass.
Will this tell us which model to use?
No. It reports whether the evaluation record can support comparisons your team is already making.
What if the seven checks find no defect?
The register says so check by check, with the counted evidence behind each result. The fee is the same.
How does cancellation work?
Cancel before kickoff—before the export and schema are in our hands. There is no subscription or recurring charge.
Pilot provenance and privacy boundary
The first-party pilot spans 2026-07-10 through 2026-07-22 across 12 snapshot dates and 17 collection runs. Before rendering, the reader reproduced 1,828 row, blob, and fetch seals plus every collection root. The newest batch seal is 5827b758ac8f3771fe0127fe9cdaf549e6abc89373298090f39fcb548e3c1f98.
Source URL: https://ustechautomations.com/permits/model-evidence-audit#pilot-evidence
Evidence effective date: 2026-07-22
Newest collection time: 2026-07-22T22:45:09+00:00
The reader queries aggregate counts only. Prompts, answers, raw JSON, usage payloads, model endpoints, credentials, and customer data are unreachable from the template. A missing or unverifiable source makes the page unavailable rather than producing zero.