DeepSeek raised its API prices on August 16, 2026 at 16:00 UTC, replacing a single flat rate per model with a two-tier peak and off-peak grid. The change was announced in the DeepSeek changelog on August 13 and the effective date never moved. It is presented as a scheduling incentive: off-peak rates are exactly half of peak rates, so shifting work out of the busy window halves the bill. What that framing hides is the base it is measured from. The off-peak rate — the cheaper of the two — already runs between 1.52 and 6.07 times the flat price it replaced, depending on the billing item. Peak runs between 3.03 and 12.14 times. Across all six billing items, on both deepseek-v4-pro and deepseek-v4-flash, there is no longer a single hour in the twenty-four at which DeepSeek costs what it cost on August 15.
Key takeaways
- Off-peak is a price increase, not a discount. The cheapest rate now available on
deepseek-v4-profor cache-miss input is $0.66 per million tokens, against $0.435 flat before. The smallest increase anywhere in the grid is 1.52 times. - Peak is 35 hours out of 168, off-peak is 133. The windows are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday; weekends are off-peak in full since August 23, 2026. That is the Chinese working week, which means European weekday mornings sit inside peak while North American working hours sit entirely outside it.
- The cache-hit line took the heaviest hit. On V4-Pro, a cache hit used to cost one one-hundred-twentieth of a cache miss. It now costs one thirtieth. The discount is still large, but it is exactly a quarter as deep as it was.
- DeepSeek changed its own description of the policy twice. On August 1 the pricing page said peak would be twice the regular price. By August 12 that footnote had been replaced with a warning of a significant overall increase. The delivered grid matches the second description, not the first.
- V4-Pro now costs exactly three times V4-Flash on cache-miss input and on output, against 3.11 times before. On cache hits the ratio moved much further, from 1.29 times to 3.14 times.
Update, August 28, 2026. DeepSeek narrowed its peak window a second time, six days after the increase this article covers. Effective 00:00 Beijing time on Sunday, August 23, 2026, weekends bill at off-peak rates in full, so peak now runs 01:00 to 04:00 and 06:00 to 10:00 UTC Monday through Friday only — 35 hours out of the 168 in a week rather than 49. The change was published on the pricing page and never appeared in the changelog. Every hour-weighted figure below has been recomputed against 35 hours; the price multipliers themselves are unaffected, because they do not depend on the clock. The finding does not soften the conclusion, it sharpens it: the off-peak tier is now reachable for more of the week, and it is still 1.52 to 6.07 times the flat rate it replaced.
What exactly changed on August 16, 2026?
Until 16:00 UTC on August 16, 2026, each DeepSeek model carried one price per billing item, charged identically at every hour. After that moment, each billing item carries two prices, selected by the clock. DeepSeek states the rule plainly on its pricing page: off-peak rates are half of peak rates, and peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, with all other hours off-peak. The three billing items are unchanged — input on a cache hit, input on a cache miss, and output — and so are the model strings, the 1M token context window and the 384K token maximum output.
Here is the full grid. The middle column is the flat rate that applied through August 15, read from an archived copy of the DeepSeek pricing page captured on August 14, 2026. The two right-hand columns are the live rates, first read on August 22, 2026 and re-read on August 28, 2026, unchanged between the two readings.
| Model and billing item | Flat, through Aug 15 | Off-peak, from Aug 16 | Peak, from Aug 16 |
|---|---|---|---|
| deepseek-v4-flash — input, cache hit | $0.0028 | $0.007 | $0.014 |
| deepseek-v4-flash — input, cache miss | $0.14 | $0.22 | $0.44 |
| deepseek-v4-flash — output | $0.28 | $0.66 | $1.32 |
| deepseek-v4-pro — input, cache hit | $0.003625 | $0.022 | $0.044 |
| deepseek-v4-pro — input, cache miss | $0.435 | $0.66 | $1.32 |
| deepseek-v4-pro — output | $0.87 | $1.98 | $3.96 |
All figures are list prices per million tokens, in US dollars, published by DeepSeek. If you are unsure why a cache hit and a cache miss are billed as separate items in the first place, our explainer on input, output and cached token pricing covers the mechanics that make the next three sections readable.
Why the off-peak rate is not a discount
A two-tier price is only a discount if the lower tier matches what came before. DeepSeek's does not. Dividing each new off-peak rate by the flat rate it replaced gives a multiplier above 1 on every one of the six billing items: 1.52 times on V4-Pro cache-miss input, 1.57 on V4-Flash cache-miss input, 2.28 on V4-Pro output, 2.36 on V4-Flash output, 2.50 on V4-Flash cache-hit input, and 6.07 on V4-Pro cache-hit input. The cheapest hour of the new grid is more expensive than every hour of the old one.
The peak column simply doubles those numbers, because DeepSeek sets off-peak at exactly half of peak. That relationship holds to the cent across all six items, and it is the source of the confusion: the doubling in this policy describes the gap between the two new tiers, not the gap between the new grid and the old one. A reader who takes "off-peak is half of peak" as a promise that off-peak equals the previous rate will underestimate the bill by somewhere between 52 percent and 507 percent.
| Billing item | Off-peak, as a multiple of the old flat rate | Peak, as a multiple of the old flat rate |
|---|---|---|
| deepseek-v4-pro — input, cache miss | 1.52x | 3.03x |
| deepseek-v4-flash — input, cache miss | 1.57x | 3.14x |
| deepseek-v4-pro — output | 2.28x | 4.55x |
| deepseek-v4-flash — output | 2.36x | 4.71x |
| deepseek-v4-flash — input, cache hit | 2.50x | 5.00x |
| deepseek-v4-pro — input, cache hit | 6.07x | 12.14x |
Our calculation, from DeepSeek's published rates. Rows are sorted by the off-peak multiplier, which is the number that matters for anyone who can schedule freely: it is the floor of what the service can now cost, and that floor has risen on every line.
How many hours are actually off-peak, and whose hours are they?
Peak covers 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, which is three hours plus four hours on each of five weekdays, or 35 hours out of the 168 in a week. The remaining 133 hours are off-peak, weekends included in full since August 23, 2026. A team that could shift every token into the off-peak window would still be paying the multipliers in the table above, but the window is genuinely wide — and, critically, it is not evenly inconvenient for everyone.
Converted out of UTC, the peak windows are 09:00 to 12:00 and 14:00 to 18:00 China Standard Time. This is not an inference: DeepSeek's own pricing footnote defined the windows in exactly those Beijing hours before the page switched to UTC notation. The policy is timed to the working day at the vendor's home market, and the consequences fall unevenly by geography.
- North America. The peak windows land at 21:00 to 00:00 and 02:00 to 06:00 US Eastern Daylight Time. A team working 09:00 to 18:00 Eastern never touches peak at all, without changing anything.
- Europe. The second window lands at 08:00 to 12:00 Central European Summer Time. A European team's entire morning is peak, at double the off-peak rate, and the fix is to push batch work past noon.
- Asia-Pacific. Teams working Chinese business hours are in the worst position, since both windows sit inside the working day. For them the off-peak rate is largely theoretical unless work is queued overnight or shifted to the weekend.
- Everyone, at the weekend. Since August 23, 2026 the geography stops mattering entirely from Saturday to Sunday Beijing time: every hour of both days bills off-peak, wherever the team sits. The asymmetry described above is a weekday asymmetry.
This asymmetry is the part of the story most worth acting on. The nominal increase is identical worldwide, but the achievable increase is not: a Western team that batches work already sits in the cheaper tier, while a Chinese team doing the same work interactively pays double on top of an increase that was already applied.
The cache-hit line is where the increase really lands
The single largest multiplier in the grid is V4-Pro cache-hit input, which went from $0.003625 to $0.022 per million tokens off-peak, a factor of 6.07, and to $0.044 at peak, a factor of 12.14. Context caching is the mechanism that made long-context agent loops affordable on DeepSeek, because the same system prompt and the same repository context are re-sent on every turn and billed at the cache rate rather than the full input rate.
The caching discount survives, but it is measurably shallower. Before the change, a V4-Pro cache hit cost one one-hundred-twentieth of a cache miss — the ratio is exact, since $0.435 divided by $0.003625 is 120. Afterwards it costs one thirtieth, since $0.66 divided by $0.022 is exactly 30. The discount is therefore precisely a quarter as deep as it was, in both tiers, because the doubling applies to hits and misses alike. Expressed as a share, a cache hit rose from 0.83 percent of a cache miss to 3.33 percent.
V4-Flash moved less violently on the same line: a cache hit was one fiftieth of a cache miss and is now roughly one thirty-first, going from 2.00 percent to 3.18 percent of the miss rate. The practical reading is that the workloads punished hardest by this change are exactly the ones DeepSeek's caching was best at — long-lived agent sessions on the large model, replaying a large stable context many times per task.
What a real workload costs now
Take a symmetric job that consumes one million cache-miss input tokens and emits one million output tokens. On V4-Flash that job cost $0.42 through August 15. It now costs $0.88 off-peak, 2.10 times more, and $1.76 at peak, 4.19 times more. On V4-Pro the same job cost $1.305, and now costs $2.64 off-peak, 2.02 times more, and $5.28 at peak, 4.05 times more. A balanced job therefore roughly doubles off-peak and roughly quadruples at peak, on both models.
For a workload that cannot schedule at all — traffic arriving uniformly around the clock — the blended rate is the 133-hour off-peak rate and the 35-hour peak rate weighted by hours across a full week, which works out to 1.21 times the off-peak rate. Applying that weighting to each billing item gives the increase an always-on service actually absorbs: 1.83 times on V4-Pro cache-miss input, 1.90 on V4-Flash cache-miss input, 2.75 on V4-Pro output, 2.85 on V4-Flash output, 3.02 on V4-Flash cache-hit input, and 7.33 on V4-Pro cache-hit input.
Two honest limits on those blended figures. They assume tokens are distributed evenly across the clock and across the week, which almost no real service does — consumer traffic is peaky and batch traffic is schedulable, and the two distort the number in opposite directions. And they are our arithmetic on DeepSeek's published rates, not a measurement of anyone's invoice. The unweighted multipliers in the earlier table are the ones that carry no assumptions at all.
What DeepSeek said it would do, and what it did
The policy was signposted for weeks before it landed, but the signposting changed shape twice. On August 1, 2026, the pricing page carried a footnote stating that the service would soon adopt peak and off-peak billing, that prices during peak hours would be twice the regular prices, and that the effective date would be subject to official announcement. Read literally, "twice the regular prices" implies peak at double the then-listed rate and off-peak at the then-listed rate. That is not what shipped.
By August 12 the footnote had been rewritten. It no longer described a doubling at peak; it stated that DeepSeek planned to raise overall pricing for its API services in the near future, with a significant increase expected, and asked users to plan accordingly. That second wording matches the delivered grid. The changelog entry of August 13 then fixed the effective date at 16:00 UTC on August 16, alongside the general availability release of V4-Pro, and the date held.
The gap between the first description and the outcome is quantifiable. If "peak equals twice the regular price" had been implemented literally, V4-Flash output would have peaked at $0.56 per million tokens. The delivered peak rate is $1.32, which is 2.36 times that. And on four of the six billing items — both cache-hit lines and both output lines — the delivered off-peak rate sits above what the originally described doubling would have produced at peak. On the two cache-miss input lines it sits below. DeepSeek corrected its own wording before the change took effect rather than after, which is the right order to do it in, but anyone who built a cost model against the August 1 footnote and did not re-read the page built the wrong model.
V4-Pro is now exactly three times V4-Flash
The increase was not applied uniformly across the two models, so the price ratio between them moved. Under the new grid, deepseek-v4-pro costs exactly 3.00 times deepseek-v4-flash on cache-miss input and on output, in both tiers. Under the old grid it was 3.11 times on those two items. The cache-hit line moved much further: V4-Pro used to cost 1.29 times V4-Flash on a cache hit, and now costs 3.14 times.
That last figure is the structural change hiding inside the price rise. The old grid gave V4-Pro a nearly free cache line, which meant a heavily cached workload on the large model was almost as cheap as the same workload on the small one. That is gone. The new grid prices the two models as consistent multiples of each other on every item, which is a cleaner design and a more expensive one, and it removes a specific arbitrage that made V4-Pro viable for context-heavy agents.
| Billing item | Pro as a multiple of Flash, through Aug 15 | Pro as a multiple of Flash, from Aug 16 |
|---|---|---|
| Input, cache hit | 1.29x | 3.14x |
| Input, cache miss | 3.11x | 3.00x |
| Output | 3.11x | 3.00x |
Our calculation from DeepSeek's published rates. The ratio is identical off-peak and at peak, since both tiers scale together.
This has a direct consequence for something we published three weeks ago. Our August 2 piece on DeepSeek's small model overtaking its own flagship rested on a price that has since moved and on a build of V4-Pro that has since been replaced. We have annotated that article rather than rewritten it, and the next section explains why its central comparison no longer holds.
What shipped alongside the increase
The August 13 changelog bundled the pricing announcement with the general availability release of DeepSeek-V4-Pro, published as DeepSeek-V4-Pro-0813. DeepSeek reports Terminal Bench 2.1 at 87.9, NL2Repo at 61.5, DeepSWE at 62.7 and Cybergym at 83.3 for the GA build, against 82.7, 54.2, 54.4 and 76.7 respectively for DeepSeek-V4-Flash-0731. The same entry adds native support for the OpenAI Responses API on both models, and three thinking effort levels — low, high and max — across V4-Pro and V4-Flash.
Those benchmark figures carry a caveat that matters more than usual. DeepSeek publishes them about its own models, and the harness configuration it documents elsewhere in the same changelog — DeepSeek Harness minimal mode, max effort level, top_p of 0.95 and temperature of 1.0 — appears as a footnote to the V4-Flash-0731 entry of July 31 and to the vision model entry of August 21, but is not restated for the V4-Pro GA numbers of August 13. Comparing the two columns assumes the same harness and the same effort level were used on both sides. That assumption is reasonable and DeepSeek has given no reason to doubt it, but it is an assumption rather than a documented fact, and self-reported benchmarks are not equivalent to independent measurement.
With that caveat stated, the direction is not in dispute: on every agentic benchmark DeepSeek publishes for both models, the V4-Pro GA build scores above V4-Flash-0731. The weights for DeepSeek-V4-Pro-0813 were published on Hugging Face the same day under the MIT license, which permits commercial use, modification and redistribution — the same licensing posture DeepSeek applied to V4-Flash, and the reason its releases keep mattering beyond its own API. Our DeepSeek V4 tool page carries the current specification, and our comparisons of Claude Opus 4.8 against DeepSeek V4 and GLM-5.2 against DeepSeek V4 cover the rivals DeepSeek benchmarks itself against.
The vision model arrives at the Flash rate
On August 21, 2026, DeepSeek added deepseek-v4-flash-vision-exp, its first multimodal model, described as experimental. It is priced identically to deepseek-v4-flash on all six cells of the grid — $0.007 and $0.014 for cache-hit input, $0.22 and $0.44 for cache-miss input, $0.66 and $1.32 for output — and carries the same concurrency limit of 2500. Images are converted to tokens by their dimensions and billed as input alongside text.
DeepSeek states that on pure-text capabilities the vision model is on par with the standard V4-Flash, and that on agent benchmarks requiring visual understanding it brings multimodal agent capability close to Claude Opus 4.8. Its published text-agent scores sit above V4-Flash-0731 on the rows both models report: Terminal Bench 2.1 at 83.9 against 82.7, NL2Repo at 57.7 against 54.2, DeepSWE at 59.3 against 54.4. All of these are vendor-reported, the model is labeled experimental, and we would treat any of them as provisional until an outside evaluator runs the same tasks.
The pricing decision is the interesting part. Adding a modality at zero premium, five days after raising every rate in the catalog, is a familiar shape: the increase pays for the capability expansion. Whether that trade reads as fair depends entirely on whether a given workload needs vision at all.
What would change this reading
Three developments would move the analysis above, and it is worth naming them now rather than defending the reading later.
- A further revision to the grid. DeepSeek's pricing page states explicitly that prices may vary and that it reserves the right to adjust them. It has now revised the policy description twice and the rates once inside a month. Every multiplier in this article is anchored to a flat rate that no longer exists, and would need re-anchoring if the grid moves again.
- A further change to the peak windows. The 35-hour peak block is what makes the off-peak tier reachable for Western teams, and DeepSeek has already moved it once, exempting weekends on August 23, 2026 without changing a single published rate. Widening it again, or shifting the hours, would change the geographic asymmetry described above the same way.
- Independent measurement of the V4-Pro GA build. The claim that V4-Pro now outscores V4-Flash rests entirely on DeepSeek's own benchmark table, published without a restated harness configuration for that entry. An outside evaluation would either confirm the ordering or restore the one our August 2 coverage documented.
The bottom line
DeepSeek's price rise is real, it is large, and the peak and off-peak structure makes it read smaller than it is. The honest one-line summary is that the vendor raised every rate in its catalog and then split the result into two tiers, one of which is labeled as the cheap one. Anyone comparing the off-peak column against a July invoice is comparing two different price lists.
What the change does not do is end DeepSeek's cost advantage in the abstract. V4-Flash at $0.22 per million cache-miss input tokens and $0.66 per million output tokens off-peak remains far below the Western frontier models it benchmarks itself against, and the weights remain MIT-licensed and self-hostable, which is a ceiling on how far any API price can rise before self-hosting becomes the rational answer. What the change does end is the specific arbitrage that made V4-Pro's near-free cache line attractive for context-heavy agents, and it ends the era in which DeepSeek's price list could be quoted from memory. From August 16, 2026, quoting a DeepSeek rate without saying which tier it belongs to is quoting half a number.
DeepSeek API price increase FAQ
When did the DeepSeek price increase take effect?
At 16:00 UTC on August 16, 2026. DeepSeek announced it in a changelog entry dated August 13, 2026, alongside the general availability release of DeepSeek-V4-Pro, and the effective date did not change between the announcement and the application.
Is DeepSeek's off-peak rate cheaper than what it charged before?
No. The off-peak rate is higher than the old flat rate on every one of the six billing items. The smallest increase is 1.52 times, on deepseek-v4-pro cache-miss input, which went from $0.435 to $0.66 per million tokens. The largest is 6.07 times, on deepseek-v4-pro cache-hit input, which went from $0.003625 to $0.022 per million tokens. Off-peak is the cheaper of the two new tiers, not a return to the old price.
What are DeepSeek's peak hours?
Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday. That is 7 hours a day on weekdays, or 35 hours out of the 168 in a week; the remaining 133 hours are off-peak, and weekends are off-peak in full. DeepSeek added the weekend exemption on August 23, 2026, one week after the increase itself. In local terms those windows are 09:00 to 12:00 and 14:00 to 18:00 China Standard Time, the hours DeepSeek used to define the policy on its pricing page before the page switched to UTC notation.
How much does DeepSeek V4-Flash cost now?
DeepSeek lists deepseek-v4-flash at $0.007 per million cache-hit input tokens off-peak and $0.014 at peak, $0.22 per million cache-miss input tokens off-peak and $0.44 at peak, and $0.66 per million output tokens off-peak and $1.32 at peak. Through August 15, 2026 the same three items were billed flat at $0.0028, $0.14 and $0.28.
How much does DeepSeek V4-Pro cost now?
DeepSeek lists deepseek-v4-pro at $0.022 per million cache-hit input tokens off-peak and $0.044 at peak, $0.66 per million cache-miss input tokens off-peak and $1.32 at peak, and $1.98 per million output tokens off-peak and $3.96 at peak. Through August 15, 2026 the same three items were billed flat at $0.003625, $0.435 and $0.87.
How much more expensive is DeepSeek at peak than it was before?
Between 3.03 and 12.14 times, depending on the billing item. The peak multipliers against the old flat rate are 3.03 on V4-Pro cache-miss input, 3.14 on V4-Flash cache-miss input, 4.55 on V4-Pro output, 4.71 on V4-Flash output, 5.00 on V4-Flash cache-hit input and 12.14 on V4-Pro cache-hit input. These are our calculations from DeepSeek's published list prices.
Did DeepSeek's context caching discount survive the price change?
Yes, but it is a quarter as deep on V4-Pro. A cache hit used to cost one one-hundred-twentieth of a cache miss on deepseek-v4-pro, since $0.435 divided by $0.003625 is exactly 120. It now costs one thirtieth, since $0.66 divided by $0.022 is exactly 30. As a share, a cache hit rose from 0.83 percent of a cache miss to 3.33 percent. On V4-Flash the same ratio moved from 50 to about 31.
Can I avoid the increase by scheduling my workload off-peak?
You can avoid the peak tier, but not the increase. Shifting every token into the off-peak window still leaves you paying between 1.52 and 6.07 times the flat rate that applied through August 15, 2026. Teams working North American hours are already entirely off-peak without changing anything, because both peak windows fall overnight in US time zones. European teams have their whole morning inside the second peak window.
What does a typical DeepSeek job cost after the change?
For a job consuming one million cache-miss input tokens and emitting one million output tokens, deepseek-v4-flash went from $0.42 flat to $0.88 off-peak and $1.76 at peak, and deepseek-v4-pro went from $1.305 flat to $2.64 off-peak and $5.28 at peak. That is roughly double off-peak and roughly quadruple at peak on both models. These are our calculations from DeepSeek's list prices, not measured invoices.
Did DeepSeek announce this increase accurately in advance?
It corrected its own description before the change took effect. On August 1, 2026 the pricing page said peak prices would be twice the regular prices, which read as though off-peak would stay at the existing rate. By August 12 that footnote had been replaced with a statement that DeepSeek planned to raise overall pricing significantly. The delivered grid matches the second wording. On four of the six billing items, the delivered off-peak rate is above what the originally described doubling would have produced at peak.
Is DeepSeek V4-Pro still three times the price of V4-Flash?
It is now exactly three times, on cache-miss input and on output, in both tiers. Before August 16 it was 3.11 times on those two items. The cache-hit line moved much further: V4-Pro cost 1.29 times V4-Flash on a cache hit under the old grid and costs 3.14 times under the new one, which removes most of the reason a context-heavy agent would have run on the larger model.
Does the new DeepSeek vision model cost extra?
No. DeepSeek added deepseek-v4-flash-vision-exp on August 21, 2026, its first multimodal model, and priced it identically to deepseek-v4-flash on all six cells of the grid, with the same concurrency limit of 2500. Images are converted into tokens based on their dimensions and billed as input tokens alongside text. DeepSeek labels the model experimental.
Sources
- DeepSeek API Docs - Models and Pricing (live peak and off-peak grid, peak-hour definition, concurrency limits, vision token billing; read August 22, 2026)
- DeepSeek API Docs - Change Log (entries of August 13, 2026 for the pricing adjustment and V4-Pro general availability, August 21, 2026 for the vision model, and July 31, 2026 for V4-Flash-0731 and its harness configuration footnote)
- DeepSeek API Docs - Models and Pricing, archived August 14, 2026 (the flat rates that applied through August 15, and the announced new grid)
- DeepSeek API Docs - Models and Pricing, archived August 1, 2026 (the original footnote describing peak prices as twice the regular prices, with Beijing-time windows)
- DeepSeek API Docs - Models and Pricing, archived August 12, 2026 (the replacement footnote warning of a significant overall increase)
- Hugging Face - deepseek-ai/DeepSeek-V4-Pro-0813 (weights published August 13, 2026 under the MIT license)
Every rate quoted in this article is a DeepSeek list price per million tokens, read from the vendor's own documentation or from a dated archive of it. Multipliers, hour counts, ratios and blended figures are our arithmetic on those rates and are labeled as such where they appear. DeepSeek states on the same page that prices may vary and that it reserves the right to adjust them.



