Public data index › Model Evidence Integrity Audit

Evaluation pipeline integrity — fixed-scope audit

Read your evaluation export the way an auditor would

Most model evaluation pipelines are trusted because they produce rows. Rows are not evidence. This audit examines one sanitized evaluation export and reports what its rows can and cannot establish about identity, duplication, fixture leakage, pair completeness, timing, unknown states, and reproducibility.

1,201 sealed answer rows looked like a substantial pilot. Only 101 answer identities sat behind them, 1,200 rows were deterministic fixtures, and 0 of 600 comparison rows contained a real model response on both sides.

A benchmark that cannot prove both sides of a comparison is not evidence.

Model Evidence Integrity Audit

$750 once

Built for the Head of ML Platform or ML infrastructure lead at a SaaS company running a self-hosted or third-party model in a product, who is being asked to stand behind evaluation numbers they did not personally trace to their source rows.

You receive a PDF evidence map and CSV findings register within five business days after we receive one sanitized evaluation export and its schema.

  • Row-level evidence map: supported, contradicted, or unknown.
  • CSV register with each finding, affected identifiers, and observed counts.
  • Distinct-ID, fixture, pair-completeness, date-spread, and run counts.
  • Reproduction note naming every missing input that blocked a result.

Start an audit — $750, once →

No subscription. Cancel before kickoff. The five business days begin when the sanitized export and schema arrive. This page is a demand test; no checkout is active.

What our own pilot export turned out to contain

We built an internal evaluation pipeline and then audited it. The volume and date spread looked healthy from the outside, but most stored rows were repeated fixtures and no comparison had genuine model output on both sides. The pilot figures are evidence about evaluation-pipeline integrity—not model quality.

1,201sealed answer rows
101distinct answer identities
1,200deterministic fixture rows
0real model pairs
Evidence checkWhat volume suggestedWhat the rows supportedAudit area
Answer identity1,201 rows101 distinct answer IDsDuplicate resealing
Fixture boundary1,200 fixture rows1 real-answer rowFixture leakage
Comparison completeness600 comparison rows0 real model pairsPair completeness
Time coverage12 snapshot dates17 collection runsTimestamp spread

What these numbers are not. This is not a model leaderboard, and this page does not rank vendors or models. We publish no model responses—ours or yours. These counts say nothing about whether any model under test was good. They show only what this evaluation record could support.

How the five business days are spent

  1. You provide one sanitized evaluation export plus its schema. We need no model access, weights, API keys, or production credentials.
  2. We reconstruct row identity and lineage, separating carried facts from file-path or naming-convention assumptions.
  3. We run seven checks: source/model identity, duplicate resealing, fixture leakage, pair completeness, timestamp spread, unknown-state handling, and reproducibility.
  4. You receive the PDF and CSV, with every figure traceable to rows in your export and every unanswerable check labeled unknown.

What is in scope

Evidence identity

Can each row be traced to a stated source and model, or is attribution assumed from a path or filename?

Duplicate and fixture boundaries

How many distinct answers sit behind the row count, and can fixture output be separated from real output?

Pair and time completeness

Do both sides of each comparison exist, and do dates represent new observations or repeated seals?

Unknowns and reproduction

Does absent evidence stay unknown, and which reported results can be regenerated from the export alone?

Deliberate exclusions

No model-quality score, no leaderboard, no compliance certification, no penetration test, and no production remediation promise. The deliverables are engineering evidence artifacts, not an audit opinion or regulatory sign-off.

Questions senior reviewers ask

What happens if the export is thin?

A missing field produces an unknown finding with that field named. It never produces an invented pass.

Will this tell us which model to use?

No. It reports whether the evaluation record can support comparisons your team is already making.

What if the seven checks find no defect?

The register says so check by check, with the counted evidence behind each result. The fee is the same.

How does cancellation work?

Cancel before kickoff—before the export and schema are in our hands. There is no subscription or recurring charge.

Pilot provenance and privacy boundary

The first-party pilot spans 2026-07-10 through 2026-07-22 across 12 snapshot dates and 17 collection runs. Before rendering, the reader reproduced 1,828 row, blob, and fetch seals plus every collection root. The newest batch seal is 5827b758ac8f3771fe0127fe9cdaf549e6abc89373298090f39fcb548e3c1f98.

Source URL: https://ustechautomations.com/permits/model-evidence-audit#pilot-evidence
Evidence effective date: 2026-07-22
Newest collection time: 2026-07-22T22:45:09+00:00

The reader queries aggregate counts only. Prompts, answers, raw JSON, usage payloads, model endpoints, credentials, and customer data are unreachable from the template. A missing or unverifiable source makes the page unavailable rather than producing zero.

Browse other data services · Agent behavior audit evidence · US Tech Automations