The promotional period is over, but the discount lives on. DeepSeek confirmed on May 31 that its flagship V4-Pro model's 75 percent price reduction, originally framed as a limited-time promotion, is now the permanent rate card -- a move that undercuts virtually every frontier model on the market.
The Chinese AI lab's decision means developers will continue paying $0.435 per million input tokens and $0.87 per million output tokens for DeepSeek-V4-Pro, with cache hits priced at just $0.003625 per million tokens. The original list price before the promotion was $1.74 per million input tokens and $3.48 per million output tokens.
"The deepseek-v4-pro model API pricing will be officially adjusted to one-quarter of the original price," the company stated in its pricing documentation update. "The promotional discount rate is now the standard rate, effective immediately and indefinitely."
A Price War With No End in Sight
The permanent price cut positions DeepSeek-V4-Pro as one of the most cost-effective frontier models available through any API. At $0.87 per million output tokens, V4-Pro costs roughly one-third of what Anthropic charges for Claude Opus 4.8 standard mode and less than one-tenth of OpenAI's GPT-5.5 pricing tier.
DeepSeek has consistently pursued an aggressive pricing strategy since launching its V3 family in late 2025. The company's approach mirrors its broader philosophy of making advanced AI capabilities widely accessible, a stance that has won it a devoted following among developers building cost-sensitive applications.
Industry analysts noted that the timing is significant. The May 31 deadline had been closely watched by enterprise customers evaluating long-term API commitments. By making the cut permanent before the deadline rather than reverting to higher prices, DeepSeek effectively locked in customer relationships that might have shifted to competitors.
"DeepSeek is playing a volume game that most Western labs cannot match at their current cost structures," said analyst Dylan Patel of SemiAnalysis. "Their compute costs are fundamentally different, and they are passing those savings through to developers."
The Broader Pricing Pressure
The move adds further pressure to an already intense pricing competition among frontier model providers. Google slashed Gemini 3.5 Flash pricing at I/O 2026, and Anthropic introduced a cheaper Fast mode with Opus 4.8. OpenAI has also been testing lower price points for specific use cases within ChatGPT Enterprise.
For developers and startups, the permanent pricing creates planning certainty. Applications built on V4-Pro's capabilities no longer face the risk of a sudden cost increase when a promotional window closes, making it easier to build business models around the API.
The V4-Pro model itself has earned strong benchmark results since its launch, performing competitively against Western frontier models on coding, reasoning, and multilingual tasks. Combined with its aggressive pricing, the model has become a popular choice for inference-heavy production workloads.
What to Watch
The permanent price cut raises the stakes for DeepSeek's competitors heading into the second half of 2026. With V4-Pro now firmly established as the price leader among frontier models, other labs face a choice between matching the rates or differentiating on capabilities that justify higher costs. Watch for whether OpenAI and Anthropic adjust their own pricing tiers in response, and whether DeepSeek's next model generation maintains the same aggressive posture.
"DeepSeek is playing a volume game that most Western labs cannot match at their current cost structures."— Dylan Patel, Analyst, SemiAnalysis