7 Best Models for Advisor GDPval Briefcases (2026)
GDPval-AA and AA-Briefcase are the closest public exams we have to “write the quarterly memo from a pile of files.” This shortlist ranks seven model IDs an RIA operations lead might actually type into an API or a ChatGPT admin panel on 3 September 2026. Claude Fable 5.1 still leads the independent knowledge-work Elos. GPT-6 Astra is the access-staged computer-use pick, not a ChatGPT default today. Claude Mythos 5.1 is on the list only so you do not confuse it with a public SKU.
Seven names, no eighth. Invite-only twins stay labeled. No vendor paid for inclusion.
TL;DR
For briefcase memos and GDP-like advisor work, start with Claude Fable 5.1 (AA GDPval-AA v2 Elo 1,853; AA-Briefcase Elo 1,694; Intelligence Index 66 at max with fallback).
Use Claude Opus 5 when you want half the list price ($5/$25 vs $10/$50) and can live with a statistically close GDPval/Briefcase tie.
Treat GPT-6 Astra and GPT-6 Pro as the OpenAI path that is not generally on ChatGPT on 3 September 2026; Astra’s Briefcase Elo rose about 80 points versus GPT-5.6 Sol while GDPval-AA v2 fell about 80 points versus Sol.
Do not pick Claude Mythos 5.1 from a public catalog. Same weights as Fable 5.1, Glasswing / trusted access only.
Quick-answer FAQs
Which model is ahead on GDPval-AA v2 right now?
Claude Fable 5.1. According to Artificial Analysis, 1,853 Elo is Fable 5.1 (max) on GDPval-AA v2 versus 1,824 for Claude Opus 5 (max), with overlapping confidence intervals, so it is a lead, not a blowout.
Which model is ahead on AA-Briefcase?
Claude Fable 5.1, barely. According to Artificial Analysis, 1,694 Elo is Fable 5.1 (max) on AA-Briefcase versus 1,685 for Opus 5; Fable leads analytical quality (2,025 vs 1,980) and trails presentation (1,495 vs 1,572). GPT-6 Astra gained about 80 Briefcase Elo versus GPT-5.6 Sol while losing presentation Elo.
Is GPT-6 Astra available in ChatGPT today?
No. On 3 September 2026 it is limited orgs, Trusted Access / Daybreak, and Foundry Limited Access, with Plus/Pro/Business/Enterprise and API described as coming over the following days. Enterprise stays off until an admin enables it.
Can we pick Claude Mythos 5.1 in the Claude workspace picker?
Not as a public SKU. Mythos 5.1 is the same weights as Fable 5.1 with looser cyber and life-science safeguards, offered through Anthropic’s trusted-access / Glasswing program only. If you do not already have that access, it is not a candidate.
Does list price decide the shortlist?
No. Fable 5.1, Fable 5, Mythos 5.1, and GPT-6 Astra all list $10/$50 per 1M input/output. Opus 5 is $5/$25. GPT-5.6 Sol is $4/$20. Cache and cost/task disagree with the sticker: Astra’s AA Intelligence cost/task is $1.67 versus $3.69 for Fable 5.1.
Should an RIA start on GPT-5.6 Sol to save money?
Yes, if the job is a short, cheap completion and you already like Sol’s presentation Elo. No, if the job is a multi-file briefcase: Sol still leads presentation quality on AA-Briefcase, but Fable 5.1 leads the overall Briefcase Elo and GDPval-AA v2.
Who this is for
This shortlist is for a COO, CCO, or operations manager at an RIA or hybrid advisory firm that produces IPS drafts, quarterly reviews, and committee memos from a household file pile. Typical stack: Wealthbox or Salesforce, a planning tool, a compliance archive, and a shared drive of PDFs. Firm size is a named reviewer plus a paraplanner, not a solo advisor pasting tax returns into a consumer chatbot.
Red flags: skip every ID on this page if the CRM already files a templated review and nobody reads the extra analysis. Skip GPT-6 Astra if you need ChatGPT seats this afternoon. Skip Mythos 5.1 if you are shopping a public picker. Skip a new orchestrator if one person already drops the packet into Claude and a CCO already signs.
When NOT to use US Tech Automations: leave it out when Wealthbox or Salesforce already tasks the quarterly letter, when the archive already files the PDF, or when a single Zapier, Make, or n8n scenario already turns a CRM date into a Slack draft. Zapier, Make, or n8n can retry a failed model call and keep a run log if you design observability, idempotency, access, and retention. That is a fair DIY choice for one stable recipe. A proposed agent design would add a household-id ledger and a human hold before the client memo — not a claim that no-code cannot retry.
Advisor time is the labor backdrop. According to the U.S. Bureau of Labor Statistics, $105,070 was the median annual wage for personal financial advisors in May 2025 (299,400 jobs), which is why a briefcase run that burns max effort on cover-slide polish is a compensation line.
Related reading: proposal software for financial advisors, quarterly portfolio review reminder automation, and Wealthbox vs Salesforce for advisors.
How we evaluated
We ranked IDs as briefcase workers: multi-file packets, GDP-like memos, and a reviewer who must sign. Weights assume an RIA ops buyer, not a coding-agent buyer.
| Evaluation criterion | Weight | Proof tests | Disqualifier |
|---|---|---|---|
| GDPval-AA v2 / AA-Briefcase Elo | 25% | 8 memos | Score cited with no lab name |
| Independent Intelligence Index | 15% | 1 leaderboard pull | Provider table treated as independent |
| List + cache + cost/task | 20% | 1 month of packets | Sticker $10/$50 used as the only cost |
| Access on 3 Sep 2026 | 20% | 1 production key | “Available in ChatGPT today” for Astra |
| Reviewer hold / compliance path | 10% | 6 households | Auto-send to the client |
| Public picker vs invite-only | 10% | 1 admin screen | Mythos sold as a public tier |
Sources counted 3 September 2026: Artificial Analysis Fable 5.1 article (1 Sep) and Astra article (3 Sep), OpenAI and Anthropic list prices, OpenAI access notes. METR time-horizon hours are unpublished for these IDs; we did not invent them. Fable 5.1’s Intelligence Index 66 includes Anthropic’s default safety fallback (~4% of output tokens on that eval).
How the automation works
The briefcase job is not “chat with last quarter’s PDF.” It is: CRM date fires, household files are assembled, the model reads the packet, a rubric checks IPS consistency, a reviewer signs, the archive files the PDF. GDPval-AA v2 is the closest public proxy for “economically valuable tasks across occupations.” AA-Briefcase is the closest public proxy for multi-week, multi-file knowledge work. Neither is your IPS.
Claude Fable 5.1 is the default API worker for that shape: live today, 1M context, $0.25 cache reads, thinking always on. Claude Opus 5 is the cheaper sibling Anthropic still recommends as the starting ID. GPT-6 Astra is the OpenAI worker you call as gpt-6-astra on the Responses API if you are in the limited-access set; set reasoning.effort because none is unsupported. GPT-5.6 Sol remains the inexpensive OpenAI control. GPT-6 Pro is the ChatGPT seat name for Astra on Pro/Business/Enterprise as those seats light up — not a separate API price list. Claude Fable 5 is the previous cache-at-$1 ID. Claude Mythos 5.1 is not a picker option.
Worked example
A configurable US Tech Automations workflow can watch a CRM quarterly-review date, attach the household packet as Claude document blocks (Claude document / PDF support), require a household ID, a file count, and a reviewer checkbox, and hold the client memo until the CCO flag is true. On a 25-household quarter with 3 control figures — 25 households, 40 PDF pages as the median packet, and 1 unsigned draft per household — the orchestrator writes one memo, one exception list, and zero client emails. Prerequisites: Claude or OpenAI API access you actually have this week, CRM read, archive write, a uniqueness key on household ID, and a named reviewer. Nothing here is a live customer result.
If the packet is on OpenAI instead, the same hold applies: call gpt-6-astra with reasoning.effort at high or max (GPT-6 Astra model docs) only after the household ID is present. Fast mode, if you enable it, is 2× Standard on the API docs surface and 2.5× Standard on the Help Center Codex/Work rate card — name the surface on the invoice.
Benchmarks
Independent knowledge-work scores first. Provider access second. Stickers third.
Sol list input is $4 according to OpenAI API pricing, $4 per 1M short-context input and $20 per 1M output for gpt-5.6-sol, versus $10 / $50 for gpt-6-astra.
Fable 5.1 cache reads are $0.25 according to Anthropic’s Fable 5.1 announcement, $0.25 per 1M cache reads (75% below Fable 5’s $1) at unchanged $10 / $50 I/O, with Anthropic estimating about 25% lower typical token bills and up to about 45% on agent loops.
| Model ID | AA Intelligence (max) | GDPval-AA v2 Elo | AA-Briefcase Elo | List in/out per 1M | Cache read per 1M | Public 3 Sep 2026 |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 66 | 1,853 | 1,694 | $10 / $50 | $0.25 | Yes |
| Claude Opus 5 | 63 | 1,824 | 1,685 | $5 / $25 | $0.50 | Yes |
| Claude Fable 5 | 62 | ~1,723 | ~1,572 | $10 / $50 | $1.00 | Yes |
| GPT-6 Astra | 61 | Sol − ~80 | Sol + ~80 | $10 / $50 | $1.00 | Limited |
| GPT-5.6 Sol | 61 | (control) | (presentation lead) | $4 / $20 | $0.40 | Yes |
| GPT-6 Pro | 61 (Astra weights) | same as Astra | same as Astra | ChatGPT seat, not API | n/a | Coming days |
| Claude Mythos 5.1 | 66 (same weights) | same as Fable 5.1 | same as Fable 5.1 | $10 / $50 | $0.25 | Invite only |
Intelligence, GDPval-AA, and Briefcase from Artificial Analysis articles dated 1 Sep and 3 Sep 2026. Fable 5 GDPval/Briefcase cells are AA’s +130 / +122 deltas versus Fable 5.1 (1,853 − 130 and 1,694 − 122). Astra GDPval/Briefcase are AA’s ~80-point deltas versus Sol, not a published absolute Elo in the Astra article. GPT-6 Pro is the ChatGPT product name for Astra seats, not a second lab score. Mythos is not a public picker.
Astra Intelligence Index is 61 on the Artificial Analysis GPT-6 Astra article (3 Sep 2026), 61 at max versus 66 for Fable 5.1, with Astra’s Intelligence cost/task at $1.67 versus $3.69 for Fable 5.1 and $0.95 for GPT-5.6 Sol. That cost column is why a finance owner can prefer Astra or Sol even when Fable wins Elo.
Tool / build comparison
| Path | Fits when | Briefcase files | Reviewer hold | Disqualifier |
|---|---|---|---|---|
| claude.ai / ChatGPT copy-paste | One memo, one person | Drag-and-drop, no household ID | Your eyeballs | Household data in a consumer log |
| Direct Fable 5.1 or Opus 5 API | Live Claude access, long packets | document blocks | You must build it | No CCO checkbox |
Direct gpt-6-astra API | You are already in limited access | Responses API files | You must build it | Assuming ChatGPT is on |
| GPT-5.6 Sol API | Cheap control, presentation-heavy slides | Same as other OpenAI IDs | You must build it | GDPval-AA is the buying test |
| Zapier / Make / n8n + one ID | One CRM date → one draft | Weak file handling | If you add a step | Auto-email the client |
| US Tech Automations + chosen ID | CRM → packet → hold → archive | Logged file count | Built as a step | Native CRM letter already is the process |
Cost and payback
Payback is Elo per dollar on the job you actually run, not a slogan.
| Model ID | AA Intelligence $/task (max) | 25-household worksheet @ 2M cache-read + 0.2M output | Access friction 3 Sep |
|---|---|---|---|
| GPT-5.6 Sol | $0.95 | ~$4.80 cache + $4.00 out = $8.80 | Low |
| Claude Opus 5 | $2.34 | ~$1.00 cache + $10.00 out = $11.00 | Low |
| GPT-6 Astra | $1.67 | ~$2.00 cache + $10.00 out = $12.00 | High (limited) |
| Claude Fable 5 | $3.14 | ~$2.00 cache + $10.00 out = $12.00 | Low |
| Claude Fable 5.1 | $3.69 | ~$0.50 cache + $10.00 out = $10.50 | Low |
| GPT-6 Pro | n/a (seat) | ChatGPT quota, not this meter | High until admin on |
| Claude Mythos 5.1 | same as Fable 5.1 | same as Fable 5.1 | Invite only |
Intelligence $/task from Artificial Analysis (Fable article quotes $3.76 / $3.14 / $2.34 for Fable 5.1 / Fable 5 / Opus 5; PIPELINE-FACTS locks Astra $1.67, Fable 5.1 $3.69, Sol $0.95 for the Intelligence cost/task column used elsewhere in this wave). Worksheet uses list cache-read and $50 (or $25 / $20) output on 0.2M output tokens and 2M cache-read tokens — a teaching model, not a quote. If cache hits are near zero, Fable 5.1’s worksheet advantage disappears.
If Elo is the buying test, Fable 5.1 still wins. If the invoice is the buying test and the packet is short, Sol or Opus 5 wins. If computer-use across desktop apps is the buying test, Astra’s AutomationBench row belongs on a different page than this briefcase shortlist.
Pros and cons
Claude Fable 5.1
Pros
Independent lead on GDPval-AA v2 (1,853) and AA-Briefcase (1,694), plus Intelligence Index 66 at max.
Live today on paid Claude, API, and clouds; cache reads $0.25.
Document/spreadsheet/slide work is the job Anthropic assigns this ID.
Cons
$10/$50 list and $3.69 AA cost/task; more verbose than Fable 5.
Intelligence 66 includes ~4% fallback tokens.
Forced
tool_choiceany/tool returns 400; AWS Covered Model retention unless EFS/ZDR.
GPT-6 Astra
Pros
AA-Briefcase about +80 Elo versus GPT-5.6 Sol; hallucination rate on AA-Omniscience down to 51% from Sol’s 92%.
Cheaper than Fable 5.1 on AA Intelligence cost/task ($1.67 vs $3.69) at the same $10/$50 list.
1,050,000-token context; Responses API computer use for ops-shaped work.
Cons
Not generally on ChatGPT on 3 September 2026; Enterprise off until an admin enables it.
GDPval-AA v2 about −80 Elo versus Sol — the wrong direction for a pure briefcase buyer.
Cache reads $1 versus Fable 5.1’s $0.25; Fast mode is 2× or 2.5× depending on the billing surface.
Claude Opus 5
Pros
Anthropic’s default starting ID; $5/$25 list; GDPval 1,824 and Briefcase 1,685, statistically close to Fable 5.1.
Presentation Elo on Briefcase (1,572) beats Fable 5.1 (1,495).
Older tool-forcing clients are less likely to 400.
Cons
Trails Fable 5.1 on Intelligence (63 vs 66) and on analytical-quality Briefcase Elo.
Cache reads $0.50 versus $0.25.
Not the long-horizon ID Anthropic tells you to reach for when Opus evals fail.
Claude Fable 5
Pros
Same $10/$50 I/O as 5.1 with a simpler client and lower AA cost/task ($3.14 vs $3.76 on the Fable article’s max column).
Generally available; no new thinking-binding surprises.
Cons
Cache reads still $1; Anthropic’s 25–45% token-bill estimate does not apply.
Trails 5.1 on Intelligence (62 vs 66) and on both knowledge-work Elos (~130 GDPval, ~122 Briefcase).
GPT-5.6 Sol
Pros
$4/$20 list and $0.95 AA Intelligence cost/task; still the presentation-Elo lead on AA-Briefcase.
Generally available; a honest cheap control.
Cons
Intelligence Index 61, tied with Astra and behind the Claude Fable/Opus band.
Token-heavy in Codex versus Astra’s ~1/3 token use on the coding-agent index — less relevant here, but it is why Sol is not the “always cheaper in practice” story on agent jobs.
Claude Mythos 5.1
Pros
Same weights and $10/$50 list as Fable 5.1, with looser cyber and life-science safeguards for organizations that already have Glasswing.
Cons
Invite only. Not a public picker. Do not staff it as a catalog option.
Same Covered Model / retention questions as Fable 5.1 once you are on AWS.
GPT-6 Pro
Pros
The ChatGPT Pro/Business/Enterprise seat name that will carry Astra as those plans light up; useful once an admin has enabled it.
Cons
Not generally available on 3 September 2026; Plus/Pro/Business/Enterprise described as coming days; Enterprise off until an admin turns it on.
Not an API price row. Do not quote $10/$50 as “GPT-6 Pro API.” Help Center Codex/Work Fast mode, if billed, is 2.5× Standard — a different surface from the API’s 2×.
Key Takeaways
Claude Fable 5.1 is the briefcase Elo lead (GDPval-AA 1,853, Briefcase 1,694, Intelligence 66 with fallback).
Claude Opus 5 is the half-price near-tie and Anthropic’s default start.
GPT-6 Astra is not generally on ChatGPT on 3 September 2026; it gained Briefcase Elo versus Sol and lost GDPval-AA Elo versus Sol.
GPT-5.6 Sol remains the cheap control; Claude Fable 5 remains the $1 cache-read previous ID; Mythos 5.1 and GPT-6 Pro are access stories, not public default picks.
Hang the workflow on a household ID, a file count, and a reviewer hold — not on a chat tab.
The team at US Tech Automations can map a configurable quarterly-memo trail from the CRM into an archive with a reviewer hold. Review workflow pricing after you have named the model ID you can actually call this week, the household key, and the person who signs the letter.
About the Author

Helping businesses leverage automation for operational efficiency.