DeepSeek’s V4 Pro 0813 Launch Tightens Race With Kimi K3 on Cost and Agent Tasks

Chinese AI developer DeepSeek rolled out the official version of its V4 Pro model on August 12. Named DeepSeek-V4-Pro-0813, the release marks a step forward in the company’s two-track strategy that pairs a high-capability Pro variant with a lighter, faster Flash option. Early data show sharp gains on agent-focused benchmarks. Yet questions linger about how it stacks up against Moonshot AI’s Kimi K3 in real deployments.
The model arrives with a one-million-token context window. It supports a maximum output of 384,000 tokens. Those figures position it well for long codebases or extended agent workflows. Pricing sits at roughly $0.435 per million input tokens and $0.87 per million output tokens. That’s a fraction of what frontier models from OpenAI or Anthropic command.
Pandaily reported the launch details including structured JSON output, native tool calls, Responses API support and Anthropic API compatibility. Beta features cover conversation-prefix continuation and fill-in-the-middle completion in non-thinking mode. The system defaults to thinking mode for heavier reasoning. A separate endpoint offers lower latency without it.
DeepSeek claims major jumps from preview performance. Terminal-Bench 2.1 climbed from 72.1% to 87.9%. CyberGym rose from 52.7% to 83.3%. DeepSWE improved dramatically from 12.8% to 62.7%. Those numbers put the model near leaders such as Claude Opus 4.8 and Grok 4.6 on several agent tasks. Independent evaluators have yet to fully verify the gains.
Community reaction on X mixed excitement with caution. One developer noted using V4 Pro 0813 as a planner alongside Kimi K3 as reviewer. “Kimi passed DeepSeek’s plan on the first try two times in a row,” the user posted. Others highlighted the cost advantage. At one-thirtieth the output price of some competitors, the model could shift economics for high-volume agent work.
But raw capability tells only part of the story. Artificial Analysis data gives Kimi K3 a higher intelligence index of 60 compared with 53 for V4 Pro 0813. Kimi also supports image input while DeepSeek does not. Parameter counts differ too. V4 Pro uses 1.6 trillion total parameters with 49 billion active in its mixture-of-experts design. Kimi K3 scales larger at 2.8 trillion total and 104 billion active.
Speed favors DeepSeek. The comparison shows 83 output tokens per second versus 41 for Kimi. Time to first token measures 1.63 seconds against 2.97 seconds. Context windows sit close. DeepSeek offers 1 million tokens. Kimi edges it at 1.049 million. Both models carry open-source licenses for their weights, though practical self-hosting demands substantial hardware.
The Information first signaled the impending launch in a briefing that framed V4 Pro as a direct challenge to Kimi K3. It highlighted claimed outperformance on AIME math, LiveCodeBench coding and MMLU general knowledge. Those earlier assertions align with the latest agent benchmark gains. DeepSeek has not released full weights for the Pro tier yet. The Flash variant saw its official update in late July.
Founded in 2023 by Liang Wenfeng and backed by the High-Flyer hedge fund, DeepSeek built its reputation on aggressive open-source releases. Its V3 model earned praise for cost efficiency. V4 extends that approach. The company offers both OpenAI-compatible and Anthropic-compatible endpoints. Concurrency limits reach 500 for Pro and 2,500 for Flash. Cache-hit pricing drops even lower, rewarding repeated prompts.
Developers have begun testing the model in practical settings. Some report strong results on software engineering tasks that require tool use and iterative debugging. One X post described it handling entire repositories, identifying issues, calling tools, editing code and testing changes. Others found it less precise than Kimi K3 on pixel-perfect design-to-code conversions.
The release comes amid rapid movement across Chinese AI labs. Alibaba’s Qwen team recently opened weights for Qwen 3.8 Max. Moonshot continues to refine Kimi with multimodal strengths. DeepSeek’s bet centers on price-to-performance for agentic coding and long-context work. Its $0.18 per million tokens in some aggregated pricing data undercuts rivals by wide margins.
Yet benchmark fragmentation complicates direct comparisons. Companies report on different suites. Terminal-Bench, CyberGym and DeepSWE emphasize agent reliability over pure reasoning. Independent arenas such as Arena.ai’s Code Arena place early AutoEval scores for V4 Pro around eighth overall and second among open models. It trails Kimi K3 slightly in web development tasks but outperforms pricier closed models on cost-adjusted metrics.
Real-world reliability will decide the winner. One Reddit discussion questioned whether the large benchmark leap from 7.3% to 62.7% on DeepSWE translates to messy production repos with 30-step tasks. “Give it a messy repo, tools and a task that takes 30+ steps and see how often it actually finishes without going off the rails,” a commenter wrote.
DeepSeek has promised further updates. Responses API and Codex support for Pro were slated for early August and now appear live. A tool called DeepSeek Harness may follow to aid evaluation. The company also introduced peak and off-peak pricing in prior updates to manage demand.
For enterprises watching AI costs, the numbers matter. Output at $0.87 per million tokens makes million-token contexts practical. Previous models often became prohibitively expensive beyond 100,000 tokens. This shift could accelerate adoption of agent systems that maintain long histories or process massive codebases in one pass.
Analysts see the move as part of broader pressure on U.S. leaders. By offering near-frontier performance at commodity prices, Chinese developers force global players to reconsider pricing. Whether V4 Pro closes the gap fully with Kimi K3 on intelligence benchmarks remains open. Current data suggest a trade-off. Pay less and accept slightly lower raw scores. Or choose Kimi for multimodal work and top-end reasoning at higher cost.
The contest won’t end here. Both labs iterate quickly. DeepSeek’s latest release gives developers a powerful new option today. And the low barrier to entry means more teams will test it immediately. Results from those experiments will shape the next round of updates. For now, V4 Pro 0813 strengthens DeepSeek’s hand in the agent and coding arena while keeping expenses in check.