DeepSeek peak-hour pricing [What It Changes]
TL;DR
DeepSeek peak-hour pricing is the vendor's weekday UTC clock that bills V4 API tokens at a published peak rate in two windows and at half that rate in every other hour.
As of August 16, 2026, DeepSeek's pricing table lists deepseek-v4-flash cache-miss input at $0.44 peak and $0.22 off-peak per million tokens, with output at $1.32 peak and $0.66 off-peak.
Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday; weekends and all other weekday hours are off-peak.
A 2-truck HVAC shop, a 10-person agency, or a solo clinic that runs overnight summaries, form-to-CRM copy, or chart reads on DeepSeek can cut the token line in half by moving non-urgent jobs out of those UTC windows.
Key Takeaways
The clock is the product: the same model, prompt, and token count cost twice as much if the request starts in a peak window.
Off-peak is not a cut on the old flat list. Codersera's August 13, 2026 V4 guide listed Flash at $0.14 / $0.28; both new tiers sit above that floor.
Cache hits stay cheap ($0.014 peak / $0.007 off-peak on Flash input). Thinking is default-on and bills as output, so max effort in a peak window is the expensive pair.
West Coast evenings and Hawaii afternoons overlap the first peak window; most East Coast 9-to-5 work does not. Rival APIs still sell flat lists or a 50% batch cut, not a weekday UTC clock.
What DeepSeek peak-hour pricing is
DeepSeek peak-hour pricing is a two-tier token clock on the DeepSeek V4 API: during two weekday UTC windows the published peak rate applies, and during every other hour the published off-peak rate applies at half of peak.
That is the whole mechanism. There is no sliding scale, no city-level tariff, and no separate "business hours" SKU. The Models & Pricing page states that prices are per 1 million tokens, that off-peak rates are half of the peak rates, and that peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, with all other hours off-peak.
If you run a 2-truck HVAC shop, this is not an abstract GPU story. The after-hours voicemail dump that you send to a model for a next-day dispatch brief is a token bill. If that job fires while UTC is inside a peak window, you pay the peak line. If it waits until off-peak, you pay half. A 10-person marketing agency that drafts client emails from form fills has the same choice: run the queue while the team is in Slack, or park it for a cheaper UTC hour. A solo clinic that turns visit notes into patient-facing summaries has the same clock, plus a duty to keep the work off any path that pretends the model is a clinician.
US Tech Automations is the site behind this explainer. Teams that already route documents through US Tech Automations workflows can add a UTC-window check before the model call rather than rebuild the rest of the stack.
What shipped, and when
DeepSeek did not hide the change. The API change log dated 2026-08-13 says the V4 family would adopt peak/off-peak pricing, with off-peak set at half of peak, "to allocate resources more reasonably." According to the V4-Pro GA note, the switch is dated 16:00 UTC, August 16, 2026.
That same note put V4-Pro into general availability on app, web, and API, added low / high / max thinking-effort controls, and added a native OpenAI Responses API path aimed at Codex. The first-call docs list deepseek-v4-flash (checkpoint DeepSeek-V4-Flash-0731), deepseek-v4-pro (DeepSeek-V4-Pro-0813), and deepseek-v4-flash-vision-exp. Consumer chat stays on chat.deepseek.com; API keys come from the DeepSeek Platform; the homepage is deepseek.com.
Weights are MIT-licensed. The V4-Pro card lists 1.6T total parameters with 49B activated and a 1M-token context. The V4-Flash card lists 284B total with 13B activated. The org index is huggingface.co/deepseek-ai; kernels live under github.com/deepseek-ai.
The published rates
According to DeepSeek's Models & Pricing page, cache-miss input for deepseek-v4-flash is $0.44 per 1M tokens at peak and $0.22 off-peak.
Peak Flash output is $1.32 per million tokens. That figure is the output row on the same pricing table.
Off-peak rates are half of peak rates. DeepSeek writes that rule in footnote (1) on the pricing page.
| Model | Cache-hit in, off-peak | Cache-hit in, peak | Cache-miss in, off-peak | Cache-miss in, peak | Output, off-peak | Output, peak |
|---|---|---|---|---|---|---|
| deepseek-v4-flash | $0.007 | $0.014 | $0.22 | $0.44 | $0.66 | $1.32 |
| deepseek-v4-pro | $0.022 | $0.044 | $0.66 | $1.32 | $1.98 | $3.96 |
| deepseek-v4-flash-vision-exp | $0.007 | $0.014 | $0.22 | $0.44 | $0.66 | $1.32 |
Sources: DeepSeek Models & Pricing.
Vision is not a separate SKU: images on deepseek-v4-flash-vision-exp convert to input tokens and bill with text, with a 384-token-per-image cap after resize, per the Vision guide. Context is 1M tokens and max output is 384K on the pricing page. Deduction is tokens × price, granted balance first, then topped-up balance.
What the old flat list was
The live DeepSeek table no longer shows the pre-August-16 flat rates. According to Codersera's V4 complete guide, last updated August 13, 2026, V4-Flash had been billed at $0.14 input and $0.28 output per million tokens, with a $0.0028 cache-hit input, until 16:00 UTC on August 16, 2026. That guide listed V4-Pro at $0.435 / $0.87 and warned that both new tiers would sit above the then-current flat list. Treat it as a snapshot of the old list. The live tariff is DeepSeek's own page.
The clock in US hours
DeepSeek bills in UTC. UTC does not observe daylight saving. Time and Date's UTC page, retrieved September 2, 2026, shows UTC with no DST and lists UTC as 7 hours ahead of San Francisco that day, which is Pacific Daylight Time (UTC−7).
| Peak window (UTC, Mon–Fri) | Pacific (UTC−7) | Eastern (UTC−4, US daylight) | Hawaii (UTC−10) |
|---|---|---|---|
| 01:00–04:00 | 18:00–21:00 | 21:00–00:00 | 15:00–18:00 |
| 06:00–10:00 | 23:00–03:00 | 02:00–06:00 | 20:00–00:00 |
| All other hours + Sat/Sun UTC | Off-peak | Off-peak | Off-peak |
Sources: DeepSeek peak-hour footnote; Time and Date UTC (Pacific offset as of 2026-09-02).
Two operational facts fall out of that grid. First, a Monday 01:00 UTC peak is still Sunday evening in the United States. Second, a Honolulu clinic that runs chart-summary jobs at 4 p.m. local time is inside the first peak window, while a Boston agency that runs the same job at 4 p.m. Eastern is not.
This is the same idea utilities have used for years. According to the U.S. Energy Information Administration, the 2025 U.S. annual average retail electricity price was 13.63¢ per kWh, with residential at 17.30¢, commercial at 13.41¢, and industrial at 8.62¢, and some utilities offer time-of-day pricing because wholesale costs rise in peak hours. The EIA Electric Power Monthly for June 2026 (released August 26, 2026) is the current monthly series. DeepSeek applied that pattern to tokens.
How a request actually bills
A request is not one number. It is cache-hit input + cache-miss input + output, and thinking mode changes the output pile.
Context caching is on by default. Overlapping prefixes can hit a disk cache. The response reports prompt_cache_hit_tokens and prompt_cache_miss_tokens. Hits bill at the cache-hit row. Misses bill at the cache-miss row. Cache construction takes seconds; unused cache is cleared in hours to days; hits are best-effort, not a guarantee.
Thinking mode is enabled by default at high effort. The model emits reasoning_content before content. Those chain-of-thought tokens are output tokens. Effort maps low → low, medium/high/xhigh → high, max → max. Tool-calling turns must pass reasoning_content back or the API returns HTTP 400.
Flash concurrency is capped at 2,500 in-flight requests. Rate Limit & Isolation lists 2,500 for Flash and vision-exp, 500 for Pro, account-wide, with HTTP 429 over the cap. Capacity expansion has no extra fee. A hung request that never starts inference is closed after 10 minutes.
If you already push form fills into a CRM with a tool from our form-to-CRM roundup, the DeepSeek step is just another node: same payload, different clock.
USTA analysis: one sample job, four bills
This is a USTA analysis. Inputs are the live DeepSeek Flash rows and the pre-change Flash list from Codersera. The sample is 20 million cache-miss input tokens plus 5 million output tokens — a long overnight batch of summaries, not a chat turn.
| Bill | Arithmetic | USD |
|---|---|---|
| Old flat Flash (Codersera, pre-2026-08-16) | (20 × $0.14) + (5 × $0.28) | $4.20 |
| New off-peak Flash | (20 × $0.22) + (5 × $0.66) | $7.70 |
| New peak Flash | (20 × $0.44) + (5 × $1.32) | $15.40 |
| Peak minus off-peak | $15.40 − $7.70 | $7.70 |
| Off-peak vs old flat | $7.70 / $4.20 | 1.83× |
| Peak vs old flat | $15.40 / $4.20 | 3.67× |
Sources: DeepSeek pricing; Codersera V4 guide. USTA analysis = arithmetic on those inputs only.
If those 20 million input tokens are all cache hits, output still dominates: peak becomes (20 × $0.014) + (5 × $1.32) = $6.88; off-peak becomes (20 × $0.007) + (5 × $0.66) = $3.44. Caching does not erase the clock. A 10-person agency that already uses US Tech Automations to push form-to-CRM jobs can queue non-urgent copy for an off-peak UTC hour and keep live client chat on the model already in the account.
How this sits next to other lists
DeepSeek is not the only paid token meter. It is the one that split the meter by hour of day.
According to OpenAI's API pricing, gpt-5.6-sol short-context input is $4.00 per 1M tokens, with $0.40 cached input and $20.00 output. gpt-5.6-luna is $0.20 / $0.02 / $1.20. Regional processing adds a 10% uplift on eligible models released on or after March 5, 2026. Fast mode is a speed tier, not a clock.
Anthropic's model table lists Claude Opus 4.7 at $5 input and $25 output per million tokens, with a Batch API at a 50% cut. Gemini Developer API pricing lists Gemini 3.8 Flash paid input at $0.75 per 1M tokens through December 31, 2026 and output at $3.75, plus a Batch API at a 50% cost reduction. Azure OpenAI likewise sells a Batch API at a 50% discount on Global Standard; GPT-Chat Latest Global is $5 / $30. Amazon Bedrock pricing is an AWS-billed catalog, not a UTC peak table. Claude consumer plans (Pro $17/month on annual billing) are seats, not API rows.
The clock did not make DeepSeek expensive versus those catalogs. It made DeepSeek's own yesterday look cheap.
| Vendor list (per 1M tokens) | Input | Output | Time-of-day split on that page? |
|---|---|---|---|
| DeepSeek V4-Flash, peak | $0.44 (cache miss) | $1.32 | Yes, UTC weekday windows |
| DeepSeek V4-Flash, off-peak | $0.22 (cache miss) | $0.66 | Yes |
| OpenAI gpt-5.6-sol, short context | $4.00 | $20.00 | No |
| Anthropic Claude Opus 4.7 | $5.00 | $25.00 | No |
| Gemini 3.8 Flash, paid, through 2026-12-31 | $0.75 | $3.75 | No |
Sources: DeepSeek; OpenAI; Anthropic; Gemini.
What a small shop should actually change
Live chat and anything a person is waiting on should keep using the model that already answers in seconds. Do not park a ringing phone on a UTC clock. Batch work is the lever: nightly HVAC recaps, agency drafts, clinic note clean-up, and "PDF to CRM note" jobs can wait for an off-peak UTC hour. Log created_at in UTC. DeepSeek's public table does not document a mid-request split, so treat a long thinking call as one window.
Thinking effort is the second lever. Thinking mode docs map low to simple tasks, high to daily agent work, and max to complex work. Max writes more output. For a form-to-CRM rewrite, disable thinking or set low. Cache is the third lever: keep a stable system prompt so later turns hit disk cache, as the cache guide requires a full prefix match. Self-hosting MIT weights from Hugging Face is a GPU bill, not DeepSeek peak-hour pricing.
HVAC dispatch summaries in US Tech Automations can wait for off-peak UTC hours when the work is a next-morning packet rather than a live callback. The same overnight-queue pattern shows up in executive-assistant automations and in the small-business automation survey. Agencies comparing Plutio alternatives, agency automation tools, or dispatch software should put a UTC column next to the model name.
Honest limits
DeepSeek can change the table again. The pricing page says prices may vary. The public docs do not publish a utilization chart for the two windows; "allocate resources more reasonably" is the stated reason on the change log. Cache hits are best-effort. Thinking tokens are not a separate SKU. Vision is experimental: the August 21, 2026 change-log entry added deepseek-v4-flash-vision-exp. Concurrency 429s are a throughput cap, not a price.
NIST's AI Risk Management Framework (AI RMF 1.0, January 2023, NIST.AI.100-1) is voluntary and does not set token prices. It does tell a clinic to map, measure, and manage the system that writes patient-facing text, including the vendor meter behind it.
Signal vs Speculation
Signal (sourced). DeepSeek publishes peak and off-peak token prices, names the UTC windows, and dates the cutover to 16:00 UTC on August 16, 2026. Off-peak is defined as half of peak. Flash, Pro, and vision-exp share that clock. Cache-hit input, cache-miss input, and output are separate rows. Concurrency caps are 500 / 2,500 / 2,500. Context is 1M. Max output is 384K. Thinking defaults on. OpenAI, Anthropic, Gemini, and Azure still publish flat (or batch-discounted) lists without a weekday UTC split on the pages cited above. EIA documents time-of-day electricity pricing as a demand tool.
Our read (12–36 months, small and mid-size businesses). If DeepSeek's two windows stay busy, other low-list hosts will copy a clock before they copy a 50% batch SKU, because a clock needs no new API. If the windows stay empty, DeepSeek will widen off-peak or cut peak, because a clock nobody hits does not allocate GPUs. Shops that already batch overnight win with UTC-aware queues, not a new vendor. Live US-daytime chat is mostly already off-peak; Hawaii afternoons and late Pacific evenings are the US slices on the peak rail. We do not claim other vendors will ship the same windows. We claim the shops that log UTC will notice first.
Frequently asked questions
Does DeepSeek peak-hour pricing apply on weekends?
No. Peak hours are Monday through Friday UTC only. Saturday and Sunday UTC are off-peak at the half-of-peak rate, per the pricing footnote.
Is off-peak cheaper than the old Flash list?
No. Using Codersera's pre-cutover Flash list of $0.14 / $0.28 and DeepSeek's live off-peak of $0.22 / $0.66, off-peak is higher than the old flat list on both input and output. Peak is higher still. The clock rebased the list up, then split it.
Do thinking tokens pay the peak rate?
They pay whatever output rate is in force for that request. Thinking is default-on, and reasoning_content is generated as output, per the thinking-mode guide. A max-effort call in a peak window is the high-cost pair.
Which US hours should an HVAC shop avoid?
If the shop is on Pacific time as of September 2, 2026, the first peak window is 18:00–21:00 local and the second is 23:00–03:00, Monday–Friday in UTC terms. Hawaii afternoons (15:00–18:00 HST) sit inside the first window. East Coast 9-to-5 does not. Convert with UTC, not with a wall clock labeled "business hours."
Does a cache hit dodge the clock?
No. Cache-hit input still has peak and off-peak rows ($0.014 / $0.007 on Flash). Output still doubles. Caching shrinks the input share; it does not delete the clock. See the cache guide and the pricing table.
Can I keep using the OpenAI SDK?
Yes. The first-call docs keep https://api.deepseek.com as the OpenAI-format base URL and https://api.deepseek.com/anthropic as the Anthropic-format base URL. The clock is on the bill, not on the SDK.
Glossary
DeepSeek peak-hour pricing: The weekday UTC two-window tariff on DeepSeek V4 API tokens, with off-peak defined as half of peak.
Peak hours: 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, as published by DeepSeek.
Off-peak: Every hour that is not a peak hour, including all of Saturday and Sunday UTC.
Cache-hit input: Input tokens served from DeepSeek's disk prefix cache, billed at the cache-hit row.
Cache-miss input: Input tokens that did not hit that cache, billed at the cache-miss row.
Thinking / reasoning_content: Default-on chain-of-thought tokens emitted before the visible answer; they bill as output.
Concurrency limit: Account-level cap on in-flight requests (500 Pro, 2,500 Flash and vision-exp) before HTTP 429.
What to do next
Read the live DeepSeek pricing table, log one week of jobs in UTC, and mark which rows sat in 01:00–04:00 or 06:00–10:00 UTC on weekdays. Move the rows that can wait. Leave the rows a person is waiting on.
If those jobs already sit in a workflow graph, plug the clock in as a gate, not as a rebuild. Open the agentic workflow builder and add a UTC-window check in front of the DeepSeek node. That is the whole operational change DeepSeek peak-hour pricing asks of a small shop.
About the Author

Helping businesses leverage automation for operational efficiency.
Related Articles
See how AI agents fit your team
US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.
View pricing & plans