DeepSeek's Updated V4 Pro Falls Short on General Benchmarks — But Leads Every Model Tested on Cybersecurity
I'm LongbridgeAI, I can summarize articles.DeepSeek released its updated V4-Pro model, which underperformed on general benchmarks compared to rivals like OpenAI and Kimi but led all tested models in cybersecurity vulnerability detection. Despite criticism over higher pricing (44 cents per million input tokens) and issues with precision, the model's strong security capabilities highlight a niche strength for Chinese AI in defensive research.
Chinese AI start-up DeepSeek has quietly released an updated version of its flagship model — DeepSeek-V4-Pro-0813 — to a mixed reception from the developer community. Benchmark results show the model struggling to match top-tier rivals on general capability measures, drawing disappointment over both performance and pricing. But in one specific domain, cybersecurity vulnerability detection, it has outperformed every other model tested — a finding that carries significant implications for the ongoing debate about Chinese AI's role in security research.
Key Points
- DeepSeek-V4-Pro-0813 scored 53 on the Artificial Analysis Intelligence Index — on par with Zhipu AI's GLM-5.2 but four points behind OpenAI's GPT-5.6 Terra and seven points behind Kimi K3
- On the Vals Index, the model ranked 12th — trailing OpenAI's previous-generation GPT-5.5 and lagging behind Kimi K3 and Anthropic's Claude Opus 5
- Belgian cybersecurity firm Aikido Security found DeepSeek-V4-Pro-0813 outperformed every other model tested at finding system vulnerabilities — despite suffering from poor precision
- The model detected significantly more vulnerabilities than Anthropic's Opus 5 and Alibaba's Qwen 3.8
- Pricing has drawn criticism — at approximately 44 US cents per million input tokens, it is described as "somewhat expensive" compared to the market median of 33 cents
- DeepSeek is expected to launch a harness product — a software framework similar to Claude Code — enabling the model to function as an autonomous AI agent
A Stealth Release With Mixed Results
The update arrived without fanfare. DeepSeek published a brief statement on its official website noting that the model offered significantly enhanced agent capabilities — then removed the statement by Thursday afternoon. No technical blog post accompanied the release, departing from the detailed documentation that typically accompanies major AI model launches.
Early benchmark results have been less than enthusiastic. On the Artificial Analysis Intelligence Index, DeepSeek-V4-Pro-0813 scored 53 — matching Zhipu AI's GLM-5.2, which was released in June, but falling four points behind the mid-tier Terra model in OpenAI's GPT-5.6 series and seven points behind Moonshot AI's Kimi K3. On the Vals Index, compiled by San Francisco-based Vals AI across multiple benchmarks, the model ranked 12th overall — trailing OpenAI's previous-generation GPT-5.5 and sitting well behind frontier systems including Kimi K3 and Anthropic's Claude Opus 5.
Two specific weaknesses stood out in Vals AI's evaluation: completing tasks within a sandboxed terminal environment and generating complex financial models in Excel spreadsheets. On social media, developers described the model as disappointing, with multiple users flagging issues including the model stopping prematurely during long coding tasks. DeepSeek did not respond to requests for comment.
The Cybersecurity Exception
The general benchmark story is not the only story. In cybersecurity vulnerability detection specifically, DeepSeek-V4-Pro-0813 produced results that have caught the attention of security researchers.
According to Belgian cybersecurity firm Aikido Security, the model outperformed every other system tested at finding system vulnerabilities. Aikido researcher Philippe Dourassov noted that while the model suffered from poor precision — flagging a higher proportion of false positives than competing systems — it detected significantly more vulnerabilities in total than Anthropic's Opus 5 and Alibaba's Qwen 3.8.
That trade-off between precision and recall is a known consideration in vulnerability scanning. A model that flags more vulnerabilities — including some false positives — may be more valuable in a first-pass security audit than one with higher precision but lower detection rates, because the cost of a missed critical vulnerability is typically far higher than the cost of investigating a false positive. The Aikido finding suggests DeepSeek-V4-Pro-0813 is optimised — whether by design or by the nature of its training — toward broad coverage rather than conservative precision.
The cybersecurity result lands in a charged context. The Bitcoin Red Team campaign, which found thousands of security vulnerabilities across Bitcoin open-source infrastructure, relied heavily on Chinese open-weight models including Kimi K3 after American frontier models repeatedly blocked legitimate defensive work. DeepSeek's demonstrated strength in vulnerability detection adds another data point to the emerging pattern of Chinese AI models outperforming restricted American systems on the specific task that security researchers most urgently need performed.
Pricing — A Complicated Picture
DeepSeek's reputation in the market has been built substantially on cost efficiency — the V4 Flash model released on July 31 was praised for its extreme cost-effectiveness and reportedly jolted Silicon Valley with its pricing. The V4 Pro update has not sustained that reputation.
At approximately 44 US cents per million input tokens, Artificial Analysis characterised the model as somewhat expensive compared to the market median of 33 US cents. Output tokens are priced at 87 US cents per million, described as moderately priced. At 6 US cents per task on the Artificial Analysis Intelligence Index, it is slightly more expensive to run than OpenAI's lightweight GPT-5.6 Luna at 5 US cents — though it remains approximately 93 percent cheaper than Moonshot AI's Kimi K3 at 84 US cents per task.
The pricing context is further complicated by DeepSeek's announcement last week that it is implementing a significant price increase across its models due to surging demand. The company advised users to plan their usage accordingly — suggesting that the current pricing, already described as above the market median, may not represent the floor.
What's Coming Next — DeepSeek Harness
Despite the mixed reception for V4 Pro, DeepSeek is preparing a product that could significantly extend the model's practical reach. The company is expected to launch DeepSeek Harness — a software framework similar to Anthropic's Claude Code — that enables large language models to function as autonomous AI agents capable of executing multi-step tasks with minimal human direction.
DeepSeek has already invited open-source developers to join beta testing for the coming Harness product, and registered a WeChat account for the team in July. The agentic framework is a logical next step for a company whose updated flagship explicitly emphasises enhanced agent capabilities — even if the general benchmark results for the underlying model have not matched expectations.
An autonomous agent framework built on a model with confirmed strong vulnerability detection performance — regardless of its ranking on general coding and reasoning benchmarks — could be a more consequential development for the cybersecurity landscape than the overall benchmark position suggests.
Sources
South China Morning Post reporting on DeepSeek-V4-Pro-0813 release, August 2026. Artificial Analysis Intelligence Index scores and pricing analysis, August 2026. Vals AI Vals Index benchmark results, August 2026. Aikido Security vulnerability detection evaluation, Philippe Dourassov, August 2026. DeepSeek official website statement on V4 Pro enhanced agent capabilities, August 2026. DeepSeek V4 Flash model release and Silicon Valley reaction, July 31, 2026. DeepSeek price increase announcement, August 2026. DeepSeek Harness beta testing invitation and WeChat account registration, July–August 2026.
