Skip to content
AI & Automation

7 Best Models for Advisor GDPval Briefcases (2026)

Sep 3, 2026

GDPval-AA and AA-Briefcase are the closest public exams we have to “write the quarterly memo from a pile of files.” This shortlist ranks seven model IDs an RIA operations lead might actually type into an API or a ChatGPT admin panel on 3 September 2026. Claude Fable 5.1 still leads the independent knowledge-work Elos. GPT-6 Astra is the access-staged computer-use pick, not a ChatGPT default today. Claude Mythos 5.1 is on the list only so you do not confuse it with a public SKU.

Seven names, no eighth. Invite-only twins stay labeled. No vendor paid for inclusion.

TL;DR

  • For briefcase memos and GDP-like advisor work, start with Claude Fable 5.1 (AA GDPval-AA v2 Elo 1,853; AA-Briefcase Elo 1,694; Intelligence Index 66 at max with fallback).

  • Use Claude Opus 5 when you want half the list price ($5/$25 vs $10/$50) and can live with a statistically close GDPval/Briefcase tie.

  • Treat GPT-6 Astra and GPT-6 Pro as the OpenAI path that is not generally on ChatGPT on 3 September 2026; Astra’s Briefcase Elo rose about 80 points versus GPT-5.6 Sol while GDPval-AA v2 fell about 80 points versus Sol.

  • Do not pick Claude Mythos 5.1 from a public catalog. Same weights as Fable 5.1, Glasswing / trusted access only.

Quick-answer FAQs

Which model is ahead on GDPval-AA v2 right now?

Claude Fable 5.1. According to Artificial Analysis, 1,853 Elo is Fable 5.1 (max) on GDPval-AA v2 versus 1,824 for Claude Opus 5 (max), with overlapping confidence intervals, so it is a lead, not a blowout.

Which model is ahead on AA-Briefcase?

Claude Fable 5.1, barely. According to Artificial Analysis, 1,694 Elo is Fable 5.1 (max) on AA-Briefcase versus 1,685 for Opus 5; Fable leads analytical quality (2,025 vs 1,980) and trails presentation (1,495 vs 1,572). GPT-6 Astra gained about 80 Briefcase Elo versus GPT-5.6 Sol while losing presentation Elo.

Is GPT-6 Astra available in ChatGPT today?

No. On 3 September 2026 it is limited orgs, Trusted Access / Daybreak, and Foundry Limited Access, with Plus/Pro/Business/Enterprise and API described as coming over the following days. Enterprise stays off until an admin enables it.

Can we pick Claude Mythos 5.1 in the Claude workspace picker?

Not as a public SKU. Mythos 5.1 is the same weights as Fable 5.1 with looser cyber and life-science safeguards, offered through Anthropic’s trusted-access / Glasswing program only. If you do not already have that access, it is not a candidate.

Does list price decide the shortlist?

No. Fable 5.1, Fable 5, Mythos 5.1, and GPT-6 Astra all list $10/$50 per 1M input/output. Opus 5 is $5/$25. GPT-5.6 Sol is $4/$20. Cache and cost/task disagree with the sticker: Astra’s AA Intelligence cost/task is $1.67 versus $3.69 for Fable 5.1.

Should an RIA start on GPT-5.6 Sol to save money?

Yes, if the job is a short, cheap completion and you already like Sol’s presentation Elo. No, if the job is a multi-file briefcase: Sol still leads presentation quality on AA-Briefcase, but Fable 5.1 leads the overall Briefcase Elo and GDPval-AA v2.

Who this is for

This shortlist is for a COO, CCO, or operations manager at an RIA or hybrid advisory firm that produces IPS drafts, quarterly reviews, and committee memos from a household file pile. Typical stack: Wealthbox or Salesforce, a planning tool, a compliance archive, and a shared drive of PDFs. Firm size is a named reviewer plus a paraplanner, not a solo advisor pasting tax returns into a consumer chatbot.

Red flags: skip every ID on this page if the CRM already files a templated review and nobody reads the extra analysis. Skip GPT-6 Astra if you need ChatGPT seats this afternoon. Skip Mythos 5.1 if you are shopping a public picker. Skip a new orchestrator if one person already drops the packet into Claude and a CCO already signs.

When NOT to use US Tech Automations: leave it out when Wealthbox or Salesforce already tasks the quarterly letter, when the archive already files the PDF, or when a single Zapier, Make, or n8n scenario already turns a CRM date into a Slack draft. Zapier, Make, or n8n can retry a failed model call and keep a run log if you design observability, idempotency, access, and retention. That is a fair DIY choice for one stable recipe. A proposed agent design would add a household-id ledger and a human hold before the client memo — not a claim that no-code cannot retry.

Advisor time is the labor backdrop. According to the U.S. Bureau of Labor Statistics, $105,070 was the median annual wage for personal financial advisors in May 2025 (299,400 jobs), which is why a briefcase run that burns max effort on cover-slide polish is a compensation line.

Related reading: proposal software for financial advisors, quarterly portfolio review reminder automation, and Wealthbox vs Salesforce for advisors.

How we evaluated

We ranked IDs as briefcase workers: multi-file packets, GDP-like memos, and a reviewer who must sign. Weights assume an RIA ops buyer, not a coding-agent buyer.

Evaluation criterionWeightProof testsDisqualifier
GDPval-AA v2 / AA-Briefcase Elo25%8 memosScore cited with no lab name
Independent Intelligence Index15%1 leaderboard pullProvider table treated as independent
List + cache + cost/task20%1 month of packetsSticker $10/$50 used as the only cost
Access on 3 Sep 202620%1 production key“Available in ChatGPT today” for Astra
Reviewer hold / compliance path10%6 householdsAuto-send to the client
Public picker vs invite-only10%1 admin screenMythos sold as a public tier

Sources counted 3 September 2026: Artificial Analysis Fable 5.1 article (1 Sep) and Astra article (3 Sep), OpenAI and Anthropic list prices, OpenAI access notes. METR time-horizon hours are unpublished for these IDs; we did not invent them. Fable 5.1’s Intelligence Index 66 includes Anthropic’s default safety fallback (~4% of output tokens on that eval).

How the automation works

The briefcase job is not “chat with last quarter’s PDF.” It is: CRM date fires, household files are assembled, the model reads the packet, a rubric checks IPS consistency, a reviewer signs, the archive files the PDF. GDPval-AA v2 is the closest public proxy for “economically valuable tasks across occupations.” AA-Briefcase is the closest public proxy for multi-week, multi-file knowledge work. Neither is your IPS.

Claude Fable 5.1 is the default API worker for that shape: live today, 1M context, $0.25 cache reads, thinking always on. Claude Opus 5 is the cheaper sibling Anthropic still recommends as the starting ID. GPT-6 Astra is the OpenAI worker you call as gpt-6-astra on the Responses API if you are in the limited-access set; set reasoning.effort because none is unsupported. GPT-5.6 Sol remains the inexpensive OpenAI control. GPT-6 Pro is the ChatGPT seat name for Astra on Pro/Business/Enterprise as those seats light up — not a separate API price list. Claude Fable 5 is the previous cache-at-$1 ID. Claude Mythos 5.1 is not a picker option.

Worked example

A configurable US Tech Automations workflow can watch a CRM quarterly-review date, attach the household packet as Claude document blocks (Claude document / PDF support), require a household ID, a file count, and a reviewer checkbox, and hold the client memo until the CCO flag is true. On a 25-household quarter with 3 control figures — 25 households, 40 PDF pages as the median packet, and 1 unsigned draft per household — the orchestrator writes one memo, one exception list, and zero client emails. Prerequisites: Claude or OpenAI API access you actually have this week, CRM read, archive write, a uniqueness key on household ID, and a named reviewer. Nothing here is a live customer result.

If the packet is on OpenAI instead, the same hold applies: call gpt-6-astra with reasoning.effort at high or max (GPT-6 Astra model docs) only after the household ID is present. Fast mode, if you enable it, is 2× Standard on the API docs surface and 2.5× Standard on the Help Center Codex/Work rate card — name the surface on the invoice.

Benchmarks

Independent knowledge-work scores first. Provider access second. Stickers third.

Sol list input is $4 according to OpenAI API pricing, $4 per 1M short-context input and $20 per 1M output for gpt-5.6-sol, versus $10 / $50 for gpt-6-astra.

Fable 5.1 cache reads are $0.25 according to Anthropic’s Fable 5.1 announcement, $0.25 per 1M cache reads (75% below Fable 5’s $1) at unchanged $10 / $50 I/O, with Anthropic estimating about 25% lower typical token bills and up to about 45% on agent loops.

Model IDAA Intelligence (max)GDPval-AA v2 EloAA-Briefcase EloList in/out per 1MCache read per 1MPublic 3 Sep 2026
Claude Fable 5.1661,8531,694$10 / $50$0.25Yes
Claude Opus 5631,8241,685$5 / $25$0.50Yes
Claude Fable 562~1,723~1,572$10 / $50$1.00Yes
GPT-6 Astra61Sol − ~80Sol + ~80$10 / $50$1.00Limited
GPT-5.6 Sol61(control)(presentation lead)$4 / $20$0.40Yes
GPT-6 Pro61 (Astra weights)same as Astrasame as AstraChatGPT seat, not APIn/aComing days
Claude Mythos 5.166 (same weights)same as Fable 5.1same as Fable 5.1$10 / $50$0.25Invite only

Intelligence, GDPval-AA, and Briefcase from Artificial Analysis articles dated 1 Sep and 3 Sep 2026. Fable 5 GDPval/Briefcase cells are AA’s +130 / +122 deltas versus Fable 5.1 (1,853 − 130 and 1,694 − 122). Astra GDPval/Briefcase are AA’s ~80-point deltas versus Sol, not a published absolute Elo in the Astra article. GPT-6 Pro is the ChatGPT product name for Astra seats, not a second lab score. Mythos is not a public picker.

Astra Intelligence Index is 61 on the Artificial Analysis GPT-6 Astra article (3 Sep 2026), 61 at max versus 66 for Fable 5.1, with Astra’s Intelligence cost/task at $1.67 versus $3.69 for Fable 5.1 and $0.95 for GPT-5.6 Sol. That cost column is why a finance owner can prefer Astra or Sol even when Fable wins Elo.

Tool / build comparison

PathFits whenBriefcase filesReviewer holdDisqualifier
claude.ai / ChatGPT copy-pasteOne memo, one personDrag-and-drop, no household IDYour eyeballsHousehold data in a consumer log
Direct Fable 5.1 or Opus 5 APILive Claude access, long packetsdocument blocksYou must build itNo CCO checkbox
Direct gpt-6-astra APIYou are already in limited accessResponses API filesYou must build itAssuming ChatGPT is on
GPT-5.6 Sol APICheap control, presentation-heavy slidesSame as other OpenAI IDsYou must build itGDPval-AA is the buying test
Zapier / Make / n8n + one IDOne CRM date → one draftWeak file handlingIf you add a stepAuto-email the client
US Tech Automations + chosen IDCRM → packet → hold → archiveLogged file countBuilt as a stepNative CRM letter already is the process

Cost and payback

Payback is Elo per dollar on the job you actually run, not a slogan.

Model IDAA Intelligence $/task (max)25-household worksheet @ 2M cache-read + 0.2M outputAccess friction 3 Sep
GPT-5.6 Sol$0.95~$4.80 cache + $4.00 out = $8.80Low
Claude Opus 5$2.34~$1.00 cache + $10.00 out = $11.00Low
GPT-6 Astra$1.67~$2.00 cache + $10.00 out = $12.00High (limited)
Claude Fable 5$3.14~$2.00 cache + $10.00 out = $12.00Low
Claude Fable 5.1$3.69~$0.50 cache + $10.00 out = $10.50Low
GPT-6 Pron/a (seat)ChatGPT quota, not this meterHigh until admin on
Claude Mythos 5.1same as Fable 5.1same as Fable 5.1Invite only

Intelligence $/task from Artificial Analysis (Fable article quotes $3.76 / $3.14 / $2.34 for Fable 5.1 / Fable 5 / Opus 5; PIPELINE-FACTS locks Astra $1.67, Fable 5.1 $3.69, Sol $0.95 for the Intelligence cost/task column used elsewhere in this wave). Worksheet uses list cache-read and $50 (or $25 / $20) output on 0.2M output tokens and 2M cache-read tokens — a teaching model, not a quote. If cache hits are near zero, Fable 5.1’s worksheet advantage disappears.

If Elo is the buying test, Fable 5.1 still wins. If the invoice is the buying test and the packet is short, Sol or Opus 5 wins. If computer-use across desktop apps is the buying test, Astra’s AutomationBench row belongs on a different page than this briefcase shortlist.

Pros and cons

Claude Fable 5.1

Pros

  • Independent lead on GDPval-AA v2 (1,853) and AA-Briefcase (1,694), plus Intelligence Index 66 at max.

  • Live today on paid Claude, API, and clouds; cache reads $0.25.

  • Document/spreadsheet/slide work is the job Anthropic assigns this ID.

Cons

  • $10/$50 list and $3.69 AA cost/task; more verbose than Fable 5.

  • Intelligence 66 includes ~4% fallback tokens.

  • Forced tool_choice any/tool returns 400; AWS Covered Model retention unless EFS/ZDR.

GPT-6 Astra

Pros

  • AA-Briefcase about +80 Elo versus GPT-5.6 Sol; hallucination rate on AA-Omniscience down to 51% from Sol’s 92%.

  • Cheaper than Fable 5.1 on AA Intelligence cost/task ($1.67 vs $3.69) at the same $10/$50 list.

  • 1,050,000-token context; Responses API computer use for ops-shaped work.

Cons

  • Not generally on ChatGPT on 3 September 2026; Enterprise off until an admin enables it.

  • GDPval-AA v2 about −80 Elo versus Sol — the wrong direction for a pure briefcase buyer.

  • Cache reads $1 versus Fable 5.1’s $0.25; Fast mode is 2× or 2.5× depending on the billing surface.

Claude Opus 5

Pros

  • Anthropic’s default starting ID; $5/$25 list; GDPval 1,824 and Briefcase 1,685, statistically close to Fable 5.1.

  • Presentation Elo on Briefcase (1,572) beats Fable 5.1 (1,495).

  • Older tool-forcing clients are less likely to 400.

Cons

  • Trails Fable 5.1 on Intelligence (63 vs 66) and on analytical-quality Briefcase Elo.

  • Cache reads $0.50 versus $0.25.

  • Not the long-horizon ID Anthropic tells you to reach for when Opus evals fail.

Claude Fable 5

Pros

  • Same $10/$50 I/O as 5.1 with a simpler client and lower AA cost/task ($3.14 vs $3.76 on the Fable article’s max column).

  • Generally available; no new thinking-binding surprises.

Cons

  • Cache reads still $1; Anthropic’s 25–45% token-bill estimate does not apply.

  • Trails 5.1 on Intelligence (62 vs 66) and on both knowledge-work Elos (~130 GDPval, ~122 Briefcase).

GPT-5.6 Sol

Pros

  • $4/$20 list and $0.95 AA Intelligence cost/task; still the presentation-Elo lead on AA-Briefcase.

  • Generally available; a honest cheap control.

Cons

  • Intelligence Index 61, tied with Astra and behind the Claude Fable/Opus band.

  • Token-heavy in Codex versus Astra’s ~1/3 token use on the coding-agent index — less relevant here, but it is why Sol is not the “always cheaper in practice” story on agent jobs.

Claude Mythos 5.1

Pros

  • Same weights and $10/$50 list as Fable 5.1, with looser cyber and life-science safeguards for organizations that already have Glasswing.

Cons

  • Invite only. Not a public picker. Do not staff it as a catalog option.

  • Same Covered Model / retention questions as Fable 5.1 once you are on AWS.

GPT-6 Pro

Pros

  • The ChatGPT Pro/Business/Enterprise seat name that will carry Astra as those plans light up; useful once an admin has enabled it.

Cons

  • Not generally available on 3 September 2026; Plus/Pro/Business/Enterprise described as coming days; Enterprise off until an admin turns it on.

  • Not an API price row. Do not quote $10/$50 as “GPT-6 Pro API.” Help Center Codex/Work Fast mode, if billed, is 2.5× Standard — a different surface from the API’s 2×.

Key Takeaways

  • Claude Fable 5.1 is the briefcase Elo lead (GDPval-AA 1,853, Briefcase 1,694, Intelligence 66 with fallback).

  • Claude Opus 5 is the half-price near-tie and Anthropic’s default start.

  • GPT-6 Astra is not generally on ChatGPT on 3 September 2026; it gained Briefcase Elo versus Sol and lost GDPval-AA Elo versus Sol.

  • GPT-5.6 Sol remains the cheap control; Claude Fable 5 remains the $1 cache-read previous ID; Mythos 5.1 and GPT-6 Pro are access stories, not public default picks.

  • Hang the workflow on a household ID, a file count, and a reviewer hold — not on a chat tab.

The team at US Tech Automations can map a configurable quarterly-memo trail from the CRM into an archive with a reviewer hold. Review workflow pricing after you have named the model ID you can actually call this week, the household key, and the person who signs the letter.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.