Skip to content
AI & Automation

Cache Pricing Census: Cut 75% Token Cost in 2026?

Sep 4, 2026

Frontier model cache pricing is the published rate a lab or cloud charges when a model rereads tokens it already saw, instead of billing those tokens at the full input list price. A close team that sends the same chart of accounts, the same policy pack, and the same entity tree on every run is buying that meter whether they named it or not. This census is a price map, not a model beauty contest.

The accounting category decision is which cache meter you will put on the close budget, not which frontier name is currently loudest. Accounting work is repetitive on purpose: the chart does not change between Tuesday and Wednesday, bank rules do not change because a new model shipped, and a 1099 mapping does not need to be re-explained at full input rates. Cache is how that repetition shows up on an invoice.

TL;DR: Claude Fable 5.1 list price matches GPT-6 Astra at $10 input and $50 output per 1M tokens, but Fable 5.1 cache reads sit at $0.25 per 1M (a 75% cut versus Fable 5) while Azure Foundry’s GPT-6 Astra short-context cache read sits at $1 per 1M and cache writes at $12.50 per 1M. Read the meters before you pick a default model for close. US Tech Automations shapes only when those token invoices must land next to a GL close with a human hold. no accounting vendor paid for inclusion.

What frontier cache pricing actually meters

A cache read is a bill for tokens the provider already stored from an earlier prefix. A cache write is a bill for storing that prefix. List input is the bill for tokens that miss the cache. Output is the bill for tokens the model generates. Four numbers, not one “price per 1M.”

Average month-end close cycle: 8-10 days according to Journal of Accountancy (2025 close-cycle benchmark), 8-10 business days for mid-market firms. Do not extend that band to a Fortune-500 3-5 day close. An 8-10 day close that re-sends the same policy pack on each nightly run is why cache meters matter to controllers: the calendar is already tight, and full input rates on unchanged context are a self-inflicted line item.

Firms adopting cloud-based workflow tools sit at 62% in the aggregate AICPA reading according to AICPA 2025 PCPS CPA Firm Top Issues Survey (2025), 62%. That figure is not a model ranking and not a claim about any named GL. It is why this census is written for accounting operators, not only for ML engineers: cloud workflow is already the default conversation, and token invoices now sit beside it.

Close work that already has a non-model playbook still belongs in the same budget conversation. Payroll, information returns, advisory packaging, and bank rec are the jobs that reuse context: payroll processing automation, 1099 processing automation, CPA advisory upsell automation, and bank reconciliation in 10 minutes. Cache pricing is what happens when those jobs start calling a frontier model on a schedule.

Key Takeaways

  • List price is tied: Fable 5.1 and GPT-6 Astra both publish $10 input / $50 output per 1M tokens on the cited 1 Sep and 3 Sep pages.

  • Cache is not tied: Fable 5.1 cache reads are $0.25 per 1M; Azure Foundry Astra short-context cache reads are $1 per 1M; Foundry cache writes are $12.50 per 1M.

  • Fable 5.1 cache reads are a 75% cut versus Fable 5 on the cited Anthropic and Verge notes.

  • Mid-market close still sits at 8-10 business days in the cited Journal of Accountancy band. Token cost is a close-pack line, not a research hobby.

  • This page is a census. It does not pick a winner.

Close-cycle pressure meets token invoices

If you close in 8-10 business days and you run a model every night of that window, you are buying nine or ten prefixes. If 80% of each prefix is the same chart and policy, you are deciding whether that 80% bills at $10, $1, $0.25, or a write rate of $12.50. The model card will not make that decision for you. The close calendar will.

A census is useful only if the units stay still. Every dollar figure in this article is per 1M tokens unless a sentence says otherwise. “Blended v5228” in the head query is a mix label, not a hidden SKU: blend list input, cache write, cache read, and output using your actual hit rate. Do not average the four published numbers and call it a discount.

Close files are prefix-heavy. The chart of accounts, entity tree, posting policy, and “do not invent accounts” instructions are the same on night two as on night one. Output is usually a short accounting exception list, not a novel. That shape is why cache reads dominate the bill when hit rate is high, and why cache writes dominate when someone rebuilds the prefix every run. A 10-K summarizer is a different shape. Do not copy a chatbot blend onto a close blend.

Hit rate is measured, not hoped. Export usage. Count cached tokens against prefix tokens for 8-10 nights. If the cached share is low, you are paying list input or writes. The fix is usually prefix stability (stop rewriting the chart dump) rather than a model swap. Swapping Fable 5.1 for Astra because a slide said “frontier” will not save a prefix you invalidate every run.

Write versus read is a once-versus-many decision. Foundry’s $12.50 write is higher than Fable 5.1’s $0.25 read and higher than Astra’s $1 read. That is expected: storage is not reread. A chart that changes twice a year should be written rarely and read nightly. A 1099 mapping that changes as returns arrive may be write-heavy and should be scoped to the changed slice, not the whole policy pack.

Output at $50 per 1M is the same on the tied list. Keep accounting exception lists short. A model that restates the entire policy in every answer is an output-rate problem you created. Ask for ids, pass/fail, and a reason code. The close pack already has the policy.

Prior-generation defaults linger in contracts. Fable 5 and GPT-5.6 Sol are on this landscape because finance teams do not rip models out on launch day. If you still run Fable 5, the 75% cache-read cut on 5.1 is the number to take to the contract owner. If you still run Sol, this census does not have a complete cache card — contact vendor and do not blend Sol invoices with Astra invoices as if the meters matched.

Provider consoles and cloud meters can disagree in naming. OpenAI’s list, Azure Foundry’s short-context table, and Anthropic’s launch note are three documents. Your invoice will use one of them. Put the document name on the close budget line so next year’s reviewer knows which $10 you meant.

Fable 5.1 cache $0.25 vs Astra cache $1

Fable 5.1 cache reads: $0.25 per 1M according to Anthropic Claude Fable and Mythos 5.1 (1 Sep 2026 launch), $0.25 per 1M tokens, a 75% cut versus Fable 5. That is the Fable side of the cache census. It is a read price, not a write price, and it is not an output price.

Fable 5.1 cache reads are a 75% cut versus Fable 5 according to The Verge (1 Sep 2026), 75%. Same cut, second publisher. Use it as a check on the launch note, not as a third rate.

Azure Foundry Standard Global short context for GPT-6 Astra publishes $10 input, $1 cached input, $12.50 cache writes, and $50 output per 1M tokens according to Azure Microsoft Foundry GPT-6 Astra (3 Sep 2026), $10 / $1 / $12.50 / $50. The cache-read comparison that operators actually need is $0.25 (Fable 5.1) versus $1 (Astra on that Foundry meter). Five times is not a rounding error. It is the difference between a nightly close prefix and a line item someone will ask about.

GPT-6 Astra’s own list is $10 input / $50 output per 1M tokens according to OpenAI GPT-6 Astra (2026), $10 / $50. That list matches Fable 5.1 on input and output. It does not match Fable 5.1 on cache reads. If you only screenshot the $10 / $50 row, you will think the models are tied. They are tied on list. They are not tied on cache.

Foundry cache writes at $12.50

Cache writes are the neglected meter. Foundry’s Astra short-context write price is $12.50 per 1M tokens on the same 3 Sep Azure post, $12.50. A write is not a read. If your orchestrator rebuilds the prefix every run, you are buying writes (or full input) instead of reads. If your orchestrator stores a stable chart-and-policy prefix, you buy the write once and the read many times.

That is the operational fork for an accounting team. A 1099 mapping that changes every night is a write-heavy job. A chart of accounts that changes twice a year is a read-heavy job. Putting both jobs on the same default model without measuring hit rate is how a “tied” $10 / $50 list becomes an untied invoice.

Artificial Analysis publishes independent model leaderboards according to Artificial Analysis. This census does not copy a quality index from that board. Quality is a separate column from cache dollars. Mix them only in your own scorecard, with your own close files, not with a scraped rank.

Neutral model landscape

Published rates for the two current-generation rows, restated as a census you can paste into a budget sheet. First column is the meter name; every data cell is a number or a stated contact-vendor gap.

Meter (per 1M tokens)Fable 5.1Astra Foundry shortFable 5 vs 5.1Astra list (OpenAI)
List input $1010contact vendor10
List output $5050contact vendor50
Cache read $0.25175% cut on 5.1contact vendor
Cache write $contact vendor12.50contact vendorcontact vendor
Close nights in cited mid-market band8-108-108-108-10
Publisher velocity pages (2026-06-14)3200320032003200

The table is a landscape, not a verdict. Each row is a genuine strength and a best-fit scenario. There is no winner column.

ModelList in / out per 1MCache read per 1MCache write per 1MGenuine strengthBest-fit close scenario
Claude Fable 5.1$10 / $50$0.25Contact vendorCited $0.25 cache-read on the 1 Sep launch noteStable chart + policy prefix reread nightly
GPT-6 Astra (Foundry Standard Global short)$10 / $50$1$12.50Foundry meter publishes write + readAzure-centered close stack that must see write cost
Claude Fable 5Contact vendor (prior gen)Higher than $0.25 (75% above 5.1 reads)Contact vendorPrior-gen Fable line still in some contractsExisting Fable 5 contract you have not moved
GPT-5.6 SolContact vendorContact vendorContact vendorPrior-gen Sol line in some stacksExisting Sol contract; confirm cache meters before mixing with Astra
USTA accounting two-week publish velocity (pages, 2026-06-14)320032003200First-party operating number, not a model scoreUse as a publisher artifact only

List price tie: $10 in / $50 out is the Fable 5.1 and Astra row above. Fable 5 and GPT-5.6 Sol stay on the map because many close stacks still run last-generation defaults. Contact vendor where this census does not have a published cell. The 3,200 figure is this accounting publisher artifact-backed June velocity ceiling (~3,200 accounting pages in two weeks for 2026 state of frontier). It does not mean Fable 5.1 indexes a 10-Q faster than Astra.

Token-bill recipe for a close file

An illustrative mid-market close runs 8 business days, 12 entities, and a 2,000,000-token nightly prefix of which 1,600,000 tokens are a stable chart-and-policy block. When the Anthropic usage object returns cache_read_input_tokens on that block, the Fable 5.1 read bills at $0.25 per 1M on the cited launch note, so 1.6 × $0.25 = $0.40 for that cached slice, while the same 1.6M tokens at Astra’s $1 Foundry cache read would be $1.60, and at $10 list input would be $16. Prerequisites: provider usage export, a stored prefix you actually reuse, and a reviewer who checks that the chart in the prefix still matches the GL. Outputs: a line-item comparison, not a promised close-day reduction. Nothing here is a live customer result.

A configurable US Tech Automations workflow can read that usage export, join it to a close-calendar date, and open a finance task when cache-hit rate on the stable block falls below a threshold you set, so the controller sees writes and misses before the invoice lands. Prerequisites: API usage credentials, a uniqueness key on run-id, and a accounting human hold before anyone changes the default model. Outputs: a G11143 pass/fail reason plus a accounting exception list. The workflow pricing page is the matching product route if that hold is in scope. Native provider consoles still show the raw meters.

MeterRate usedTokens in the exampleDollar mathOwner
Fable 5.1 cache read$0.25 / 1M1,600,000$0.40controller
Astra Foundry cache read$1 / 1M1,600,000$1.60controller
List input (either, on a miss)$10 / 1M1,600,000$16.00controller
Foundry cache write$12.50 / 1M1,600,000$20.00controller
Output (either list)$50 / 1M200,000$10.00controller

Those five rows are arithmetic on published rates plus an illustrative token mix. They are not a benchmark of model quality and not a discount you are owed.

A second numeric view is the same 1,600,000-token stable block across four billing choices, using only rates already cited. No quality rank is implied.

Billing choiceRate $ / 1MTokensLine $Nights in 8-day closeClose-window $
Fable 5.1 cache read0.2516000000.4083.20
Astra Foundry cache read1.0016000001.60812.80
List input miss10.00160000016.008128.00
Foundry cache write once12.50160000020.00120.00

Read the close-window column as arithmetic on an illustrative prefix, not as a promise. Real files add output, misses, and writes on nights the prefix changes. The point of the census is that $3.20, $12.80, $128, and $20 are different lines, and a $10 / $50 list screenshot hides that.

Prefix hygiene is the control that makes cache real. Freeze the chart dump. Version the policy pack. Do not interpolate last night’s exceptions back into the prefix. If a human edits the prefix, treat it as a write. If a job silently rebuilds it, treat it as a write you did not budget.

Entity count changes the prefix more than people expect. Twelve entities with twelve charts is not one chart. Either send a per-entity prefix (more writes, smaller reads) or a consolidated chart with entity ids (one write, larger reads). Pick one and measure. Mixing both in a “helpful” prompt is how hit rate collapses.

Glossary of cache meters

  • List input — tokens billed as new. $10 per 1M on both Fable 5.1 and Astra in this census.

  • List output — tokens the model writes. $50 per 1M on both in this census.

  • Cache write — tokens stored as a reusable prefix. $12.50 per 1M on the cited Foundry Astra short-context meter.

  • Cache read — tokens reread from that prefix. $0.25 per 1M on Fable 5.1; $1 per 1M on the cited Foundry Astra meter.

  • Hit rate — share of prefix tokens that actually read as cache. Your number, not the vendor’s.

  • Prefix — the stable front of the prompt (chart, policy, entity tree).

  • Miss — tokens that fall back to list input.

  • Blended rate — (write + read + miss + output) / total tokens for your file, not an average of list prices.

Who this accounting page is for

This census is for a controller, close lead, or firm technologist who already runs or is about to run a frontier model on repeating accounting files, and who needs the cache meters in one place. It assumes you already have a GL.

Red flags: skip a cache project when you do not reuse a prefix, when you have no usage export, or when nobody will own hit rate. Do not treat a model leaderboard as a close-calendar. Do not treat this publisher’s 3,200-page velocity figure as a model score.

A close team that runs a model once a quarter on a one-off memo is not this reader. Cache meters reward repetition. If the prefix is new every time, list input is the honest rate and this census is background reading.

A team that already exports usage, already stores a versioned chart dump, and already has a named owner for hit rate can stay inside the provider console. The census is still the price map. The extra workflow is optional.

Frontier cache FAQ

Is Fable 5.1 cheaper than GPT-6 Astra?

On list input and output they tie at $10 / $50 per 1M tokens in the cited notes; on cache reads Fable 5.1 is $0.25 versus Astra’s $1 on the cited Foundry short-context meter.

What is the Foundry cache write price?

Azure Foundry Standard Global short context for GPT-6 Astra lists cache writes at $12.50 per 1M tokens on the 3 Sep 2026 post.

Did Fable 5.1 cut cache reads versus Fable 5?

Yes. The cited Anthropic launch and Verge note put Fable 5.1 cache reads at $0.25 per 1M, a 75% cut versus Fable 5.

Should we mix Fable 5 and GPT-5.6 Sol into the same close default?

Only if the contract still requires them; confirm each generation’s cache meters before blending invoices, because this census does not publish a complete Sol or Fable 5 rate card.

Does a cache discount close the books faster?

No. The cited close cycle is still 8-10 business days for mid-market firms. Cache changes the token line, not the calendar by itself.

How should we pilot cache on a close file?

Run one prefix for 8-10 nights, export cache_read_input_tokens, and compare $0.25 / $1 / $10 / $12.50 math on the same token counts before you change the default model.

Price the cache, then price the close

Put Fable 5.1, GPT-6 Astra, Fable 5, and GPT-5.6 Sol on one sheet with four meters, not one. Then prove hit rate on a real chart-and-policy prefix.

The team at US Tech Automations can map a configurable usage-export-to-close-calendar trail with a human hold. Review workflow pricing after you have named the 2026 state of model meters, the GL, and the reviewer.

Industry context according to USTA model-wave 2026-09-03 knowledge pack (checked September 4, 2026).

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.

See how our Finance & Accounting AI agents work

US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.

Explore Finance & Accounting agents