DeepSeek to Quadruple V4 API Peak Prices in Shift to Tiered Billing
The AI provider is introducing a peak and off-peak pricing system starting August 16, 2026, to better manage resource allocation.
DeepSeek is implementing a significant price hike for its V4 AI models, transitioning to a tiered billing system based on usage hours. The move marks a departure from the company's previous strategy of aggressive, low-cost market penetration.
Starting August 16, 2026, DeepSeek will introduce a peak and off-peak billing structure. For the V4 Pro model, peak output tokens will rise from $0.87 to $3.96 per 1 million tokens. Similarly, the V4 Flash model will see peak output tokens increase from $0.28 to $1.32 per 1 million tokens. Under the new system, off-peak rates will be set at half of the peak rates. DeepSeek has defined peak hours as 01:00-04:00 and 06:00-10:00 UTC. The company stated that these changes are intended to "allocate resources more reasonably."
A Pivot in Pricing Strategy
This pricing shift follows a period of volatility regarding DeepSeek's cost structure. The current API pricing was originally a promotional offer scheduled to end on May 31. While DeepSeek previously claimed that these discounted rates would be made permanent, the company has since reversed that decision, coinciding with the rollout of the V4 Pro model.
DeepSeek has historically built its market position by offering frontier-level AI capabilities at a fraction of the cost of its Western competitors. By pivoting to a resource-managed revenue model, the company is signaling a transition from acquiring users through subsidies to establishing a more sustainable financial operation.
Market Implications
Despite the steep increase, DeepSeek's pricing remains competitive against several high-end industry benchmarks. For comparison, Moonshot's Kimi K3 costs $15 per 1 million output tokens, and OpenAI's GPT-5.6 Sol is priced at $30 per 1 million output tokens.
However, the hike narrows the gap between DeepSeek and other low-cost alternatives. For instance, GPT-5.6 Luna, priced at $1.20 per 1 million output tokens, is now cheaper than DeepSeek's V4 Flash during peak hours. This change may prompt developers and enterprises to re-evaluate their model selection based on the specific timing of their API calls to avoid the higher peak costs.
What to Watch
Industry observers will be monitoring whether other AI providers adopt similar time-of-use billing to manage GPU demand. It remains to be seen if the introduction of the peak/off-peak system will successfully stabilize DeepSeek's resource allocation or if the price increase will drive a migration of users toward more stable, low-cost competitors like the Luna series.