DeepSeek raises V4 developer prices by up to 1,100%AI generatedConfirmed

DeepSeek's live rate card and August 13 changelog confirm higher prices for V4-Flash and V4-Pro in every token category; Reuters and Chinese business outlets independently corroborate increases ranging from 50% to 1,100%.

AI Models

DeepSeek raises V4 developer prices by up to 1,100%

DeepSeek, the Hangzhou artificial-intelligence company, has raised what developers pay to use its V4 models, replacing flat ultra-cheap rates with higher peak and off-peak prices.

Published
NRB — News Republic Brigade

Other links

DeepSeek

Who's involved

  • DeepSeek

    Hangzhou-based Chinese AI developer behind the V4 models and the API whose prices have risen

    goal → manage heavy developer demand and computing capacity while turning low-cost adoption into a sustainable commercial service

  • DeepSeek API customers

    software developers and companies that pay to use DeepSeek models instead of running them themselves

    goal → keep inference bills predictable and decide whether to move work into cheaper hours or to other providers

  • OpenAI

    U.S. AI company whose GPT-5.6 Luna is among the industry's lowest-priced high-volume API models

    goal → win price-sensitive production workloads by competing on cost as well as capability

  • Google

    U.S. technology company selling Gemini models through its developer API

    goal → compete for high-volume agent and application workloads with its Flash models

  • Anthropic

    U.S. AI company behind the Claude model family

    goal → sell higher-priced models on capability and reliability while keeping Sonnet competitive on cost

  • Moonshot AI

    Beijing AI company behind the Kimi model family and a prominent Chinese DeepSeek rival

    goal → win developers for its Kimi K3 flagship despite higher token prices

In short

DeepSeek's low-price edge has narrowed sharply: every listed price for its V4 artificial-intelligence models is now higher than before. Customers who keep the same workloads will pay more, making cheaper hours or rival services more attractive. Higher bills are near-certain under the new rates; whether many customers will switch providers is uncertain. An API is the paid software connection developers use to send text to a model and receive generated text back, with charges based on tokens, or small pieces of text. DeepSeek's August 6 warning of a large increase was official. On August 13 it named DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813, introduced daily peak and off-peak rates, and set the change for 16:00 UTC on August 16, equivalent to midnight August 17 in Beijing.

The rise is much more than a peak-hour surcharge. V4-Flash moved from 0.02 yuan per million cache-hit input tokens, meaning reused stored input, 1 yuan for cache-miss or fresh input and 2 yuan for output to 0.05/1.5/4.5 yuan off-peak and 0.10/3/9 yuan at peak. That is 150%/50%/125% more off-peak and 400%/200%/350% more at peak. V4-Pro moved from 0.025/3/6 yuan to 0.15/4.5/13.5 off-peak and 0.30/9/27 peak, increases of 500%/50%/125% off-peak and 1,100%/200%/350% at peak. DeepSeek advertises a 50% off-peak discount, but that discount is from the new peak rate: every off-peak price is still above the old flat price.

DeepSeek remains cheap beside many premium rivals, but not in every comparison. V4-Flash now costs $0.007 cached/$0.22 fresh input/$0.66 output off-peak and $0.014/$0.44/$1.32 at peak. OpenAI, a U.S. AI company, charges $0.02/$0.20/$1.20 for standard short-context GPT-5.6 Luna: DeepSeek Flash has cheaper cached input, but fresh input is 10% dearer off-peak and 120% dearer at peak; peak output is 10% dearer and off-peak output 45% cheaper. Against Google, the U.S. technology company, Gemini 3.7 Flash costs $0.75 input/$3.75 output, so DeepSeek Flash remains substantially cheaper even at peak. DeepSeek Pro's peak $1.32/$3.96 input/output remains about 34%/60% below Anthropic Claude Sonnet 5's $2/$10 and 56%/74% below Moonshot AI's Kimi K3 at $3/$15. These are list-price comparisons, not direct measures of value: capability, token counting and workloads differ.

Previously in this story

How it unfolded

01

V4 starts DeepSeek's low-price push

2026-04-23 – 2026-04-23

DeepSeek set this course when it launched the V4 Preview family on April 24, promoting a "cost-effective" one-million-token context window and V4-Flash as the cheaper option. Reuters then recorded aggressive discounts. The first verified reversal found here came on June 29, when Odaily, citing Blue Whale Tech and multiple API users, reported a subscriber email announcing double peak-hour V4 prices; SCMP, a newspaper, independently obtained the email the next day. The broader increase began with DeepSeek's August 6 developer notice. IT Home, a technology news outlet, carried by Sina, had the earliest professional report found here at 09:49 Beijing time; that does not prove nobody published earlier.

1 source
02

DeepSeek cuts prices again

2026-04-26 – 2026-05-22

Three days after V4 appeared, DeepSeek pushed harder on price. Reuters reported a temporary 75% discount on V4-Pro and a tenfold cut in cache-hit input prices across its API lineup. On May 23, DeepSeek made the Pro promotion permanent, putting V4-Pro API rates at 0.025 to 6 yuan per million tokens depending on token type, versus 0.1 to 24 yuan before. DeepSeek said Pro had initially cost more because high-end computing capacity was constrained and that prices should fall as more Huawei Ascend 950 capacity arrived. At that point, the later broad increase looked less likely.

2 sources
03

Peak pricing appears

2026-06-28 – 2026-07-30

The first clear retreat from ultra-low flat pricing came at the end of June. Multiple API subscribers were reported to have received a DeepSeek email saying normal rates would stay unchanged but peak rates would double between 09:00-12:00 and 14:00-18:00 Beijing time. SCMP independently saw the email and said DeepSeek described the change as a way to distribute resources better and keep service stable. Meanwhile rivals kept cutting: OpenAI reduced GPT-5.6 Luna to $0.20 input and $1.20 output per million tokens on July 30, and DeepSeek released the improved V4-Flash-0731 a day later. The June email made a move away from flat pricing more likely, although the public record reviewed here does not establish whether that exact June schedule was billed continuously until the final August change.

4 sources
04

DeepSeek warns prices will rise

2026-08-02 – 2026-08-05

DeepSeek's price shift became clear within three days. On August 3, Reuters reported research firm Artificial Analysis's finding that V4-Flash was by far the cheapest prominent model in its benchmark set: $0.14 per million input tokens, $0.28 output and roughly three cents per benchmark task, against about $0.86 for Kimi K3, $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5. On August 6, DeepSeek warned developers that overall API pricing would rise and that the increase was expected to be large. It gave no models, percentages or start date. IT Home carried the earliest professional timestamp found here at 09:49 Beijing; SCMP independently reported the notice at 13:01 HKT. The warning made a broad increase likely, but specific multipliers circulating before August 13 were still unconfirmed.

3 sources
05

DeepSeek sets the new bill

2026-08-09 – 2026-08-15

Some rivals were still making AI access cheaper. Anthropic made Sonnet 5's introductory $2/$10 rate permanent on August 10, cancelling a planned rise to $3/$15, while Google launched Gemini 3.7 Flash on August 13 at introductory $0.75/$3.75 pricing. DeepSeek went the other way. Its August 13 V4-Pro general-availability announcement said both V4 models would use peak and off-peak billing, with off-peak exactly half the new peak rate. Reuters calculated increases of 50% to 1,100% from DeepSeek's earlier rates, depending on model, token type and time. WSJ, a news outlet, separately checked the scale through V4-Pro output: $3.96 per million at peak and $1.98 off-peak versus $0.87 before. DeepSeek said the aim was more rational use of resources. Claims that the move was mainly about margins, fundraising or an IPO remain unconfirmed.

6 sources
06

Higher prices go live

2026-08-16 – 2026-08-17

Once the new billing began, the full change was clear. V4-Flash cache-hit/cache-miss/output prices rose from 0.02/1/2 yuan per million tokens to 0.05/1.5/4.5 off-peak and 0.10/3/9 peak. V4-Pro moved from 0.025/3/6 to 0.15/4.5/13.5 off-peak and 0.30/9/27 peak. Peak hours are 09:00-12:00 and 14:00-18:00 Beijing time, seven hours a day; the other 17 hours are off-peak. DeepSeek's live dollar card lists Flash at $0.007/$0.22/$0.66 off-peak and $0.014/$0.44/$1.32 peak, and Pro at $0.022/$0.66/$1.98 off-peak and $0.044/$1.32/$3.96 peak. DeepSeek has not abandoned low pricing, but the earlier flat ultra-cheap offer has become a higher, capacity-sensitive one.

3 sources

Where things stand

The increase is live. DeepSeek's hosted V4 lineup now costs $0.007 cached/$0.22 fresh input/$0.66 output off-peak and $0.014/$0.44/$1.32 peak for V4-Flash-0731; V4-Pro-0813 costs $0.022/$0.66/$1.98 off-peak and $0.044/$1.32/$3.96 peak. Every category is more expensive than under the old rate card, even off-peak. The biggest increase is V4-Pro cache-hit input at peak, up 1,100%; the smallest are cache-miss inputs off-peak for both models, up 50%. DeepSeek says the time bands are meant to allocate resources more rationally and encourage customers to schedule workloads. Claims that fundraising, an IPO plan or a deliberate effort to maximize margins caused the increase remain unconfirmed.

DeepSeek is also no longer simply the cheapest in every raw-price comparison. Against OpenAI's standard short-context GPT-5.6 Luna rate of $0.20 input/$1.20 output, DeepSeek Flash costs more for fresh input even off-peak and more for both fresh input and output at peak, though cached input remains cheaper. OpenAI's listed long-context Luna rates rise to $0.40 input/$1.80 output. DeepSeek Flash remains cheaper than Google's introductory Gemini 3.7 Flash $0.75/$3.75 rate, while DeepSeek Pro remains well below Anthropic Sonnet 5's $2/$10 and Moonshot K3's $3/$15 in raw token prices. Those figures are not rankings of overall value: vendors split text into tokens differently, models produce different amounts of text, capability differs, and some providers charge more for longer context or different processing tiers. One point remains unresolved: the June peak-pricing email proves DeepSeek had already announced a move away from flat pricing, but the sources reviewed do not establish that the exact June schedule was billed continuously until August 17. The August 17 schedule, however, is now the official live rate card.

Sources

  • OpenAIGPT-5.6 Luna and Terra price cuts · 2026-07-30
  • IT Home via Sinabroad API hike warning, earliest professional timestamp found · 2026-08-06
  • Reuters50%-1,100% hike calculation · 2026-08-13
  • GoogleGemini 3.7 Flash rival pricing · 2026-08-13
  • Jiemian Newsimplementation and peak-hour cross-check · 2026-08-17
  • Anthropiccurrent Claude model pricing · 2026-08-18