First Wave of DeepSeek V4 Evaluations Released!
I'm LongbridgeAI, I can summarize articles.Arena.ai shows V4 Pro's coding capabilities have surged into the top three open-source models; Vals AI evaluates its code performance as a 10x leap over the previous generation, surpassing Gemini 3.1 Pro. Its pricing is highly competitive: the Flash version outputs at just $0.28 per million tokens, roughly 1% of Claude Opus's cost. Netizens commented: "GPT-5.5, sorry, DeepSeek V4 is the new shocking moment."
Following the open-source release of the DeepSeek V4 preview version, the first wave of evaluation results from third-party leaderboards has emerged. Multiple evaluations indicate that DeepSeek V4's performance, particularly in coding tasks, has surged into the top tier of open-source models, while its "million-token context window + low price" further lowers the entry barrier for developers.
According to third-party evaluations, the platform Arena.ai defined on X that V4 Pro (Thinking Mode) represents "a major leap compared to DeepSeek V3.2." In its coding arena, it ranks 3rd among open-source models and 14th overall. Another evaluator, Vals AI, stated that V4 took the top spot among open-weight models in its Vibe Code Benchmark with an "overwhelming advantage," defeating closed-source models like Gemini 3.1 Pro, achieving approximately a 10x performance leap over the previous generation V3.2.

On the pricing front, V4-Flash output costs $0.28 per million tokens, more than 99% lower than Claude Opus 4.7; V4-Pro output costs $3.48, making it one of the lowest-priced options among frontier models at this level. Comparison tables show Flash sits in the lowest tier of small models, while Pro remains in the lower end of the "large model frontier" category.
Discussions around actual user experience are beginning to diverge. Many users on X claim its value-for-money ratio is "unbeatable." Meanwhile, DeepSeek maintains restraint in its self-description, stating that while its knowledge and reasoning capabilities are close to closed-source systems, there remains a gap of about 3 to 6 months. It also noted that "constrained by high-end computing power," Pro service throughput is limited, with expectations of future price reductions.
Third-Party Evaluations: Coding Capabilities Lead the Pack, Overall Rankings Closely Chasing Top Tier
Shortly after the release of OpenAI GPT-5.5, the DeepSeek-V4 preview version went live and was simultaneously open-sourced. This includes V4-Pro with a total parameter count of 1.6 trillion (49B active parameters), and V4-Flash with 284 billion total parameters (13B active parameters). Both models support a 1-million-token ultra-long context window and adopt the MIT open-source license.

Model evaluation platform Arena.ai announced on the day of V4's release that DeepSeek V4 Pro (Thinking Mode) ranked 3rd among open-source models in its coding arena and 14th overall, characterizing this launch as "a major leap compared to DeepSeek V3.2." Arena.ai also tested V4 Flash; both models support a 1-million-token context window.
Vals AI's evaluation results were even more noteworthy. The platform stated that DeepSeek V4 became the number one open-weight model in its Vibe Code Benchmark with an "overwhelming advantage," surpassing the second-place Kimi K2.6 and defeating closed-source frontier models like Gemini 3.1 Pro.

Vals AI specifically emphasized that V4 achieved approximately a 10x performance leap over V3.2—"V3.2 scored only 5 points on this benchmark, not a typo." In the Vals composite index ranking, V4 finished in 2nd place, trailing the top-ranked Kimi K2.6 by only 0.07%.

Community reaction has been highly positive. On the X platform, user Sigrid Jin called it a new "shocking moment" and mentioned "now you can run gpt 5.4-ish models at home." He wrote:
"GPT-5.5, sorry, DeepSeek V4 is the new shocking moment. It defeated GPT-5.4 in High Intensity mode in the coding arena."

User Ejaaz stated:
"China is leading in AI; they have caught up. DeepSeek V4 Flash is 99% cheaper than Opus 4.7, costing only $0.28 per million tokens, ranking first in the coding arena, not a typo."

Some users expressed reservations. After trying it out, X user Michael Anti stated that the actual experience of V4 Flash did not surpass the already mature V3.2, considering the upgrade experience disappointing for existing users.

Official Self-Assessment: Restrained Wording, Smallest Gap in Agent and Coding Domains
DeepSeek maintained its consistent cautious style when evaluating its own performance. Official documents indicate that in knowledge and reasoning tasks, V4-Pro has surpassed mainstream open-source models and is close to closed-source systems like Gemini, though a gap of about 3 to 6 months remains compared to the most advanced frontier models. In Agent and coding tasks, performance is close to or even partially exceeds Claude Sonnet.

Regarding internal usage data, DeepSeek stated that V4 has become the primary model for Agentic Coding among company employees. Evaluation feedback indicates its user experience is superior to Claude Sonnet 4.5, with delivery quality approaching Claude Opus 4.6 non-thinking mode, though a gap remains compared to Opus 4.6 thinking mode.
In mathematics, STEM, and competition-level coding evaluations, V4-Pro surpassed all currently publicly evaluated open-source models, including Moonshot AI's Kimi K2.6 Thinking and Zhipu AI's GLM-5.1 Thinking, achieving results comparable to top-tier closed-source models.

Blogger Simon Willison pointed out in his evaluation article that V4-Pro (1.6 trillion parameters) is currently the largest known open-weight model, exceeding Kimi K2.6 (1.1 trillion), GLM-5.1 (754 billion), and DeepSeek V3.2 (685 billion), providing new options for enterprise users intending local deployment.
He also shared pelican diagrams generated by different models:
Here is the pelican generated by DeepSeek-V4-Flash:
As for DeepSeek-V4-Pro:
Pricing Structure: Lowest at Just 1% of Competitors, Further Price Cuts Expected in Second Half of Year
DeepSeek's pricing strategy was the most market-focused aspect of this launch. V4-Flash input/output prices are $0.14/$0.28 per million tokens respectively, lower than OpenAI GPT-5.4 Nano ($0.20/$1.25) and Gemini 3.1 Flash-Lite ($0.25/$1.50), making it the lowest-priced option among small models.
V4-Pro input/output prices are $1.74/$3.48, also lower than Gemini 3.1 Pro ($2/$12), GPT-5.4 ($2.50/$15), Claude Sonnet 4.6 ($3/$15), and Claude Opus 4.7 ($5/$25).
Price comparison data summarized by blogger Simon Willison shows that V4-Pro is currently the lowest-cost option among large frontier models, while V4-Flash is the lowest-cost option among small models, even cheaper than OpenAI's GPT-5.4 Nano.

DeepSeek attributes this low-price capability to extreme efficiency optimization of the model in ultra-long context scenarios. Official data shows that in a 1-million-token scenario, V4-Pro's single-token inference compute power is only 27% of V3.2's, and KV cache is only 10%; V4-Flash drops to 10% and 7% respectively.
Notably, DeepSeek added a note in its pricing explanation: "Constrained by high-end computing power, current Pro service throughput is very limited. We expect Pro prices to drop significantly after the mass launch of Ascend 950 super-nodes in the second half of the year," implying further room for price reductions exists.
Technical Architecture: Hybrid Attention Mechanism Breaks Long Context Bottleneck, Adapts to Domestic Computing Power
The core technical innovation of DeepSeek-V4 lies in the pioneering "CSA (Compressed Sparse Attention) + HCA (Heavy Compressed Attention)" hybrid attention architecture, aimed at solving the industry pain point where traditional attention mechanisms exhibit quadratic complexity growth in ultra-long context scenarios, making memory and computing power difficult to implement engineering-wise. CSA compresses every 4 tokens into one information block and retrieves the most relevant content via sparse retrieval, significantly reducing computation while retaining mid-segment details; HCA condenses massive information into framework-level blocks focusing on global logic processing.

Beyond this, V4 introduces mHC Manifold Constrained Hyperconnections (upgrading traditional residual connections to constrain signal propagation on stable manifolds) and the Muon optimizer (replacing traditional AdamW, adapting to MoE large models and low-precision training). Official data indicates full-link engineering optimization can achieve inference acceleration of nearly 2x.
Regarding adaptation to domestic computing power, DeepSeek-V4 has completed comprehensive verification of fine-grained expert parallel optimization solutions on the Huawei Ascend NPU platform, achieving a speedup ratio of 1.50 to 1.73 times in general inference load scenarios. DeepSeek officials stated that V4 is the world's first trillion-parameter model trained and inferred on a domestic computing power foundation. However, the Ascend platform adaptation code is not yet open-sourced externally and belongs to closed-source optimization. Additionally, Cambricon has completed adaptation for V4-Flash and V4-Pro via the vLLM inference framework, with related code open-sourced to the GitHub community.


