DeepSeek V4 Pro 0813
deepseek-v4-pro-0813- 🧠 1.6T-parameter Mixture-of-Experts flagship, roughly 49B parameters activated per token
- 📏 One-million-token context window, standard across DeepSeek's V4 services
- 🔧 Hybrid attention: Compressed Sparse Attention plus Heavily Compressed Attention
- ⚡ DeepSeek reports 27% of V3.2's per-token FLOPs at 1M context
- 🎯 Aimed at reasoning, coding, and long-horizon agentic workflows
- 💬 Thinking and non-thinking modes selectable per request
- 🔧 Function calling and web search enabled in this catalog deployment
- 🏢 From DeepSeek in Hangzhou; weights published on Hugging Face
DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…
Explore 6 more models by DeepSeek →DeepSeek V4 Pro 0813 is the dated August 2026 build of DeepSeek's flagship V4 Pro line: a Mixture-of-Experts model with roughly 1.6 trillion total parameters, about 49 billion activated per token, and a one-million-token context window, per DeepSeek's model card. It succeeds the April 2026 preview release [[sibling:deepseek-v4-pro|DeepSeek V4 Pro]], which introduced the same parameter scale and context length; 0813 is the later iteration of that endpoint.
Architecturally, the V4 generation departs from [[sibling:deepseek-v3.2|DeepSeek V3.2]] with a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention. On DeepSeek's own reported figures, at one-million-token context V4 Pro needs about 27% of the single-token inference FLOPs and 10% of the KV cache of V3.2 — the stated reason million-token context became the default across DeepSeek's services rather than a premium tier. Post-training follows a two-stage recipe: domain-specific experts trained separately with supervised fine-tuning and reinforcement learning, then consolidated into a single model through on-policy distillation.
Within the family, Pro is the heavyweight tier. The lighter siblings — [[sibling:deepseek-v4-flash|DeepSeek V4 Flash 0423]], [[sibling:deepseek-v4-flash-0731|DeepSeek V4 Flash 0731]], its low-latency [[sibling:deepseek-v4-flash-0731-fast|DeepSeek V4 Flash 0731 Fast]] variant, and the encrypted [[sibling:e2ee-deepseek-v4-flash|DeepSeek V4 Flash]] deployment — use a 284B-parameter, 13B-active configuration at the same context length. DeepSeek notes that Flash's smaller scale trails Pro on knowledge-heavy tasks and the most complex agentic workflows.
The model exposes thinking and non-thinking modes, and this deployment adds function calling and web search. It is text-only: suited to long-document analysis, software engineering, tool use, and multi-step agents rather than image understanding.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $1.4174 | $0.1417 | $4.2522 | chat | — |
| antseed-opal-badger-2580 0xc85d…2580 | 2 | gated | $0.825 | $0.825 | $2.475 | chat | openai-chat-completions |
| Cooper 0xacb4…00ae | 0 | gated | $0.50 | $0.075 | $0.95 | chat,coding,reasoning | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.