Skip to content
AI & Automation

Claude Fable 5.1 vs Fable 5: Output Bills (2026)

Sep 3, 2026

Claude Fable 5.1 vs Claude Fable 5 is an output-token-bill decision for firms that already live in workpapers, not a beauty contest between two Anthropic names. List input and output prices did not move: $10 / $50 per million tokens on both. Cache reads did move: $1.00 on Fable 5, $0.25 on Fable 5.1. Anthropic’s own estimate is about 25% cheaper typical token bills, up to about 45% on agent loops. Independent evals still show Fable 5.1 emitting more output, which is why a CAS team can “upgrade” and watch the invoice grow.

The job on this page is memo-and-workpaper generation: PBC lists, variance narratives, payroll-to-GL stories, engagement-letter handoffs. If the model writes a novel every time Invoice.TotalAmt changes, the cache cut will not save you. If the model rereads the same 40-page manual on every client, the $0.25 cache read is the whole upgrade.

TL;DR

  • Stay on Claude Fable 5 when your prompts are short, cache hit rates are already low, and you do not want Fable 5.1’s extra verbosity or its breaking tool-choice rules.

  • Move to Claude Fable 5.1 when the same policy pack is reused (cache reads at $0.25) and you will cap max_tokens, set effort, and stop dumping whole binders into one turn.

  • Do not read “25% cheaper typical bills” as “cheaper than GPT-6 Astra on AA cost per task”; that Astra comparison is $1.67 vs $3.69 and is a different page.

  • Meter the write. Chat windows are not a general-ledger control.

Who this is for

This page is for partners, CAS managers, and IT at firms of about 8 to 80 people who already call Claude from a workpaper process — API, Bedrock, or a tightly used paid-chat seat — and who just saw Fable 5.1 ship on 1 September 2026. You post payroll journals, chase PBC, and invoice on a cadence. You care what output_tokens does to the month.

Red flags: you have no token dashboard; you plan to paste every binder into claude.ai and call that “the 25% savings”; you need forced tool_choice any/tool (Fable 5.1 returns 400); you want Mythos 5.1 as a public SKU; nobody reviews a draft before it hits the GL.

When NOT to use US Tech Automations: if Fable 5 or Fable 5.1 already drafts the only memo and a senior already types it into the workpaper, stay in that seat. If a single Zapier, Make, or n8n scenario already posts Gusto-to-sheet with retries you configured, keep it. Those tools can run a one-event path with the history you turn on; you still own the journal and the reviewer. US Tech Automations belongs when a Fable call must sit on a payroll or invoice event with a hold, which is the shape on agentic workflows.

Onboarding still has to finish even when the model is quiet; see CAS client onboarding in 30 days. PBC tracking that still lives in a shared inbox will also burn tokens without improving the file; fix the request log first with audit PBC request tracking rather than asking Fable 5.1 to narrate a missing PBC for the third time.

The hidden cost of manual output-token review

The hidden cost is not the $10 / $50 sticker. It is partners rereading 1,700-word Fable 5.1 memos that Fable 5 would have finished in 1,000 words, plus the re-runs when someone pastes the same tax outline uncached. according to the BLS Occupational Outlook Handbook, $81,680 was the median annual wage for accountants and auditors in May 2024, about $39 an hour on a 2,080-hour year. Ten extra review minutes on 40 memos a month is about 6.7 hours, or ~$260 of partner-adjacent time, before the API bill.

according to AICPA & CIMA, 400,000-plus CPAs sit in the U.S. professional body that still expects a reviewer on work that goes to a client or a file. A cheaper cache read does not replace that reviewer. It only makes the second pass cheaper if you actually cache the handbook.

according to the IRS Statistics of Income, 160 million-plus individual income-tax returns are in a recent Data Book cycle, which is the volume CAS and tax shops are approximating when they ask a model to narrate every exception. Verbosity at that scale is a budget line.

Hidden-cost line (one CAS pod, monthly)Fable 5 (manual paste)Fable 5.1 (uncapped paste)Fable 5.1 (cached + cap)
Model list in / out per 1M$10 / $50$10 / $50$10 / $50
Cache read per 1M$1.00$0.25$0.25
Estimated output tokens8.0M13.6M (~1.7×)9.0M
Output $ at $50 / 1M$400$680$450
Cache-read $ on 20M hits$20$5$5
Review hours at $39 / h8146
Review $$312$546$234
Illustrative total$732$1,231$689

Source: list and cache from Anthropic pricing, checked 2026-09-03. 1.7× output is the AA-direction estimate vs Fable 5, not your firm’s log. Replace the token counts with Console export before you change SKUs.

Manual review also hides a second bill: the re-prompt. A senior who hates a 1,800-word Fable 5.1 variance note will ask the model to “make it shorter,” which writes another 800 tokens at $50 per million and another five minutes of reading. Fable 5 did less of that because it started closer to the length the file needs. If you are not willing to set a 1,200-token cap and a reviewer checklist, stay on Fable 5 for routine CAS memos and reserve 5.1 for the research jobs where the extra pages are the point.

How we evaluated

We scored Claude Fable 5.1 against Claude Fable 5 as token-bill products for an accounting workflow, not as “which is smarter than Astra.” Astra stays off this vs page.

Evaluation criterionWeightProof testDisqualifier
Output-token $ at $50 / 1M30%2 weeks of output_tokensNo usage export
Cache-read $ ($1.00 vs $0.25)25%Hit rate ≥ 40%Re-pasting the handbook every call
Breaking API changes15%1 staging replayForced tool_choice still in code
Reviewer minutes per memo15%10 filesModel ships to client unedited
AWS retention / Covered Model10%1 BA checkBedrock aws_review by surprise
12-month price transparency5%1 invoiceFast mode or long-context mixed in

Fable 5.1 thinking is adaptive and always on, default effort high. That is a feature for hard research and a bill for “rewrite this three-line variance.” Effort and max_tokens are the controls. Fast mode is not the Fable lever; do not mix OpenAI’s 2× API-docs Fast mode into this comparison.

How the automation actually works

The automation is: freeze the handbook in cache, send only the client delta, cap output, hold for a reviewer, then write the workpaper system.

Worked example. A 22-person CAS firm posts payroll to Sage Intacct after Gusto runs. QuickBooks still holds AR. Official Intuit docs define Invoice.TotalAmt on the Invoice object in the QuickBooks Online Invoice API. When Invoice.TotalAmt on a $18,600 draft changes by $420, the firm today pastes the invoice, the Gusto register, and a 35-page revenue-recognition note into Fable 5.1. Three turns at 9,000 output_tokens each is 27,000 tokens, $1.35 of output at $50 per million, plus thinking. Cache the 35-page note (cache_read_input_tokens at $0.25 / 1M on Fable 5.1 vs $1.00 on Fable 5) and the next 40 invoices reuse it. US Tech Automations would subscribe to the invoice change, call Fable 5.1 with that cache prefix, and park a 400-word variance draft for the manager before the Intacct journal — the same event shape as Gusto to Sage Intacct payroll journals.

Invoicing cost work that is still a spreadsheet belongs in invoicing software cost for accounting firms, not in a longer Fable 5.1 essay.

according to Anthropic’s Claude API pricing, $0.25 per million tokens is the Fable 5.1 cache-hit price versus $1 per million on Fable 5, with 5-minute cache writes at $12.50 and 1-hour writes at $20. Batch is half off both models. That is the entire “upgrade” if your hit rate is real.

according to AWS What’s New on Fable 5.1, 5.1 is on Bedrock as a Covered Model: aws_review mode can retain traffic up to 30 days with AWS human review unless you are EFS-eligible for ZDR through 2026-12-31. Token bills are not the only line a CPA firm is buying.

Benchmarks: before vs after

Benchmark (counted 2026-09-03)Claude Fable 5Claude Fable 5.1
Input $ / 1M$10$10
Output $ / 1M$50$50
Cache read $ / 1M$1.00$0.25
Cache write 5 m / 1 h $ / 1M$12.50 / $20$12.50 / $20
Anthropic typical-bill estimatebaseline~25% lower; up to ~45% on agent loops
AA-direction output vs Fable 51.0×~1.7×
Forced tool_choice any/toolsupported400 error
Thinkingprior Fable 5 behavioralways on, default high
Max output128K128K
Context1M1M

Source: Anthropic pricing and Fable 5.1 “what’s new” docs; AA-direction output multiplier from the 2026-09-03 research count, not a substitute for your usage export.

The before/after that matters inside a firm is not the index. It is last month’s Console bill versus this month after you (a) switch the model id to claude-fable-5-1, (b) add cache_control, (c) set a 1,200-token cap on variance memos, and (d) stop sending the handbook as fresh input. Skip (b)–(d) and Fable 5.1 will look “more expensive than Fable 5” while doing the job you asked: write more.

A practical week-one test is 20 identical variance jobs: ten on claude-fable-5, ten on claude-fable-5-1, same cache prefix, same max_tokens. Export output_tokens, cache_read_input_tokens, and wall-clock review minutes. If 5.1 wins on review minutes and cache $ enough to cover extra output $, switch the default. If 5.1 only wins on “it felt smarter,” keep Fable 5 on the high-volume memo and put 5.1 on PBC research. That split is cheaper than a firm-wide model-id flip with no cap.

Build vs buy vs orchestrate

Build: keep Fable 5, add your own summarizer, and accept $1 cache reads. Buy: switch the model id, take $0.25 reads, and change tool_choice. Orchestrate: fire the model from payroll and AR events with a reviewer.

AP tools do not replace this decision. If bill-pay is the actual bottleneck, read Bill.com vs Ramp vs Brex for AP instead of stretching Fable 5.1 into a payment-approval bot.

PathWho owns retriesToken controlWhen it wins
Stay on Fable 5 in chatThe personNoneOne partner, short memos
Switch to Fable 5.1 API + cacheYour scriptmax_tokens, effort, cacheSame handbook, many clients
Bedrock Fable 5.1AWS + youSame, plus Covered Model rulesAlready on AWS, ZDR checked
Orchestrated event + holdWorkflow + reviewerPer-event capPayroll, AR, PBC every week

Only two models are the vs: Fable 5.1 and Fable 5. Orchestration is a runtime, not a third Anthropic SKU.

Pros and cons

Pros

Claude Fable 5.1

  • Cache reads $0.25 vs $1.00, the 75% cut.

  • Same $10 / $50 list as Fable 5.

  • Anthropic estimate ~25% cheaper typical bills when cache and loops behave.

  • Stronger long-running coding and document work than Fable 5 on Anthropic’s own pitch.

Claude Fable 5

  • Same list prices, less verbose on the AA-direction output comparison.

  • Forced tool_choice still works for stacks that send any / tool.

  • $1 cache reads are still a discount to raw input if you already cache.

  • Fewer breaking thinking-block rules than 5.1.

Cons

Claude Fable 5.1

  • More output tokens in independent evals (~1.7× vs Fable 5), so uncapped bills rise.

  • Forced tool_choice any/tool returns 400.

  • Thinking always on; earlier models cannot read 5.1 thinking blocks.

  • Bedrock Covered Model retention unless ZDR/EFS applies.

Claude Fable 5

  • Cache reads stay $1.00, four times Fable 5.1.

  • You miss the 5.1 document and agent improvements.

  • Teams delay the break-fix (tool_choice, thinking) until a rushed cutover.

  • Not the model Anthropic is pushing for long-horizon agents after 1 September 2026.

FAQs

Did Anthropic cut Fable 5.1 input and output prices?

No. List remains $10 input / $50 output per million tokens, same as Fable 5. The cut is cache reads, $1.00 down to $0.25. Treat any “Fable 5.1 is cheaper” claim as a cache-and-loop claim until your output_tokens prove it.

Why might our Fable 5.1 bill go up after the upgrade?

Because the model writes more. Independent evals saw on the order of 1.7× output versus Fable 5, and thinking is always on. If you paste full binders uncached, you buy verbosity at $50 / 1M. Cap max_tokens, cache the handbook, and lower effort on routine variance notes.

Is Fable 5.1 cheaper than GPT-6 Astra?

Not on Artificial Analysis intelligence cost per task: Astra $1.67, Fable 5.1 $3.69. List prices tie at $10 / $50. Fable 5.1 can be cheaper than Fable 5 on cache-heavy loops. Do not mix those three sentences.

What breaks in our integration if we switch model ids?

Forced tool_choice set to any or a named tool returns 400. Earlier models cannot read Fable 5.1 thinking blocks. Editing earlier turns can invalidate thinking. Replay staging with claude-fable-5-1 before you cut over payroll.

How do we keep output bills predictable for CAS work?

Export output_tokens and cache_read_input_tokens weekly. Set a 1,200-token cap on client-facing variance memos. Cache the 30-page policy pack. Hold every GL-bound draft for a person. Fable 5.1 cache hits cost $0.25 per 1M tokens.

Should tax and CAS use Bedrock or the first-party API?

Use the API if you want Anthropic’s cache price and you are not already standardized on AWS. Use Bedrock if AWS is the control plane, and read the Covered Model retention (up to 30 days, aws_review) unless ZDR/EFS applies through 2026-12-31. Token price is not the confidentiality decision.

Key Takeaways

  • Fable 5.1 vs Fable 5 is a verbosity-plus-cache decision; list I/O is a tie at $10 / $50.

  • Both models list at $10 / $50 per 1M tokens.

  • Cache reads fall from $1.00 to $0.25 on Fable 5.1, a 75% cut.

  • Uncapped Fable 5.1 can raise output $ even as cache $ falls.

  • Breaking changes (tool_choice 400, thinking blocks) belong in the cutover checklist.

  • Meter events, not chat windows, when the draft can hit a journal.

For the event-shaped version of that meter, start at US Tech Automations.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.