DeepSeek V4 Pro
deepseek-v4-proDeepSeek-V4-Pro- 🧠 Flagship 1.6T-parameter Mixture-of-Experts model with 49B active per token.
- 📏 1M-token context window, now standard across DeepSeek services.
- 🆕 Hybrid attention pairs Compressed Sparse Attention with Heavily Compressed Attention.
- ⚡ Uses 27% of V3.2's per-token FLOPs and 10% of its KV cache at 1M context.
- 💬 Offers non-thinking, thinking, and Think Max reasoning modes.
- 🔧 Function calling and tool use for long-horizon agentic workflows.
- 🔒 MIT-licensed, open-weights; released April 24, 2026.
- 📚 Post-trained via domain-expert SFT/RL plus on-policy distillation.
DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…
Explore 6 more models by DeepSeek →DeepSeek V4 Pro is the flagship of DeepSeek's two-tier V4 preview series, a Mixture-of-Experts language model with 1.6 trillion total parameters and 49 billion activated per token, supporting a one-million-token context window. It launched alongside its lighter sibling [[sibling:deepseek-v4-flash|DeepSeek V4 Flash]] (284B total / 13B active) under the MIT license on April 24, 2026, and is positioned for advanced reasoning, coding, and long-horizon agentic tasks.
The headline change over [[sibling:deepseek-v3.2|DeepSeek V3.2]] is architectural. V4 Pro introduces a hybrid attention design combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to make long-context processing far cheaper. Per DeepSeek's own model card, at a 1M-token context the model requires only 27% of the single-token inference FLOPs and 10% of the KV cache compared with V3.2. This sub-linear scaling is what makes million-token context economically practical, and 1M context is now the default across DeepSeek's official services.
Training adds a two-stage post-training pipeline: domain-specific experts are first cultivated independently through supervised fine-tuning and reinforcement learning with GRPO, then consolidated into one model via on-policy distillation. The model was pre-trained on tens of trillions of tokens and uses an FP4/FP8 mixed-precision scheme for its expert parameters.
V4 Pro exposes three reasoning effort levels, including a maximum-effort "Think Max" mode, and supports function calling and web search for tool-driven agents.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 99 | #4 | $0.865 | $0.165 | $1.898 | chat,reasoning,coding,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 83 | #1 | $0.54 | $0.027 | $1.08 | chat,open-source,math,coding,finance,reasoning,long-context,agents | openai-chat-completions |
| Open Ant 0xe4f6…5bc4 | 74 | #5 | $1.0725 | $0.2145 | $2.1457 | chat,coding,reasoning,tools,long-context | openai-chat-completions |
| surplusintelligence.ai 0x0e49…8927 | 68 | #3 | $0.6488 | $0.1238 | $1.4235 | agents,chat,cheap,code,coding,developer,frontier,function-calling,reasoning,research,tasks,tools,translate,web-search | openai-chat-completions |
| Open Bird 0xc0f1…8183 | 64 | #2 | $0.609 | $0.0508 | $1.218 | chat,coding,reasoning | openai-chat-completions |
| NovaRoute AI 0xc50d…ed7b | 60 | gated | $0.594 | $0.594 | $1.188 | chat,coding,code,reasoning,tasks,deepseek,value,surplus,openai-compatible,low-cost,verified,github,response-auth,base-usdc,monitored | openai-chat-completions |
| DeepArc 0xfc36…8842 | 53 | gated | $0.60 | $0.02 | $1.20 | chat,coding,reasoning | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 40 | gated | $0.825 | $0.165 | $1.6505 | chat,coding,reasoning,tools,long-context | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $1.485 | $0.297 | $2.9709 | base-usdc,chat,code,coding,deepseek,github,json,low-cost,math,moe,monitored,openai-compatible,reasoning,response-auth,surplus,tasks,tools,value,verified | — |
| D5V1N2 0xd5e7…7be0 | 26 | gated | $0.36 | $0.18 | $0.80 | chat,reasoning,coding,research,router,fallback,deepseek | openai-chat-completions |
| Meridian AI 0x8c8c…06f5 | 22 | gated | $0.0926 | $0.0926 | $0.1851 | chat,coding,reasoning | openai-chat-completions |
| uomi.ai 0x87df…48e3 | 19 | gated | $1.535 | $1.535 | $3.123 | chat,math,coding | openai-chat-completions |
| silent-validator 0xbb89…7301 | 19 | gated | $0.435 | $0.435 | $0.87 | chat,reasoning,tools,moe | openai-chat-completions |
| Ant Army 0xc8bd…f6c9 | 15 | gated | $0.80 | $0.80 | $3.00 | chat,coding,reasoning | openai-chat-completions |
| bartly.eth64.de 0x666e…4666 | 10 | gated | $0.50 | $0.10 | $1.20 | chat,coding,reasoning | openai-chat-completions |
| adfreellm Gateway 0xabd5…4e98 | 5 | gated | $0.1584 | $0.0158 | $0.3168 | chat,coding,reasoning,math | openai-chat-completions |
| Deep Node 0x668e…6cd0 | 4 | gated | $0.435 | $0.0036 | $0.87 | chat,coding,reasoning,research | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 3 | gated | $0.5899 | $0.33 | $1.2944 | chat,coding,math | openai-chat-completions |
| Jarvis 0x5beb…a1c9 | 2 | gated | $0.30 | $0.30 | $1.50 | chat,coding,research,reasoning | openai-chat-completions |
| ➤Bullet Ant 🐜 0xe924…8936 | 1 | gated | $3.00 | $3.00 | $5.00 | chat,coding,reasoning,anon | openai-chat-completions |
| antseed-opal-badger-2580 0xc85d…2580 | 1 | gated | $0.825 | $0.825 | $1.6505 | chat | openai-chat-completions |
| minion0x 0x215e…e2e3 | 1 | gated | $0.24 | $0.04 | $1.28 | chat,coding,math,finance,fast | openai-chat-completions |
| antseed-opal-marten-4af3 0x17f5…4af3 | 0 | gated | $10.00 | $10.00 | $10.00 | chat | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0 | gated | $1.60 | $1.60 | $3.50 | chat,reasoning,agent,function-calling | openai-chat-completions |
| Leftermute 0x388b…5389 | 0 | gated | $0.3495 | $0.3495 | $0.7668 | chat,coding,json,tools | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0 | gated | $0.1115 | $0.0056 | $0.223 | chat,open-source,math,coding,finance,reasoning,long-context,agents | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.