{
  "family": "model-cards",
  "generated": "2026-08-30",
  "note": "Real rows out of dated copies we sealed ourselves. Nothing here is made up.",
  "where_these_rows_came_from": "Where these rows came from, and anything their publisher requires to be printed alongside them, is set out on the page this file came from: https://ustechautomations.com/feeds/model-cards",
  "rows_published": 5,
  "columns": 2,
  "headers": [
    "Claim id",
    "Claim as written"
  ],
  "rows": [
    [
      "3db5b5b96714",
      "We evaluated with a 200k token context window and up to 10 context resets with auto-summarization of previous context if 180k tokens of the windo"
    ],
    [
      "42927b630fe8",
      "Model Malicious - overt (refusal rate) Malicious - covert (refusal rate) Dual use (success rate) Claude Sonnet 4.5 95.56% 52.42% 87.00% Claude Sonnet 4 80.00% 77.00% 54.33% Table 4.1.A Claude Code evaluation results without mitigations."
    ],
    [
      "858314f7d049",
      "Our checkpoint evaluations showed that the model reached 45.3% performance on the hard subset of SWE-bench Verified, still below the 50% checkpoint."
    ],
    [
      "e380e1ecd478",
      "Refusals or callouts along these lines appeared in about 13% of transcripts generated by the automated auditor, especially when the auditor was implementing scenario ideas that were, by design, unusual or implausible."
    ],
    [
      "b3f762f36dd0",
      "Refusal rate: extended thinking Claude Sonnet 4.5 0.02% (\u00b1 0.04%) 0.05% (\u00b1 0.08%) 0.00% (\u00b1 0.00%) Claude Opus 4.1 0.08% (\u00b1 0.09%) 0.13% (\u00b1 0.15%"
    ]
  ]
}
