DeepSeek V4 Flash 0423
deepseek-v4-flashDeepSeek-V4-Flashdeepseek-v4-flash-0423deepseek/deepseek-v4-flash- 🆕 Efficiency-optimized member of DeepSeek's V4 preview series, released April 2026.
- 🧠 284B-parameter Mixture-of-Experts with only 13B active per token.
- 📏 One-million-token context window, now DeepSeek's default standard.
- 🔧 Hybrid attention pairs Compressed Sparse and Heavily Compressed Attention.
- ⚡ Tuned for fast, high-throughput, cost-efficient inference.
- 💬 Dual Thinking and Non-Thinking modes via one model.
- 🎯 Capable in reasoning, coding, function-calling, and agentic tool use.
- 🔒 Released under the permissive MIT license.
DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…
Explore 7 more models by DeepSeek →DeepSeek V4 Flash is the lightweight half of DeepSeek's V4 preview series, launched alongside DeepSeek V4 Pro on April 24, 2026. Where the Pro model carries 1.6 trillion total parameters with 49 billion active, Flash uses a much smaller 284-billion-parameter Mixture-of-Experts design activating just 13 billion parameters per token — positioning it as DeepSeek's economical, high-throughput option. Both models share a one-million-token context window, which the company states is now the default across its services.
The V4 family introduces a new hybrid attention mechanism combining Compressed Sparse Attention and Heavily Compressed Attention, plus DeepSeek Sparse Attention, to cut long-context compute and memory cost. DeepSeek reports that, at the 1M-token setting, the Pro variant needs only 27% of single-token inference FLOPs and 10% of the KV cache compared with the prior-generation DeepSeek V3.2, illustrating the architectural efficiency gains this generation targets.
Both V4 models support Thinking and Non-Thinking modes and an OpenAI- and Anthropic-compatible API. DeepSeek notes that Flash's maximum-effort mode can reach reasoning quality comparable to Pro when given a larger thinking budget, though its smaller scale leaves it slightly behind on pure-knowledge tasks and the most complex agentic workflows.
The model targets advanced reasoning, software engineering, tool use, and enterprise assistants, and ships under the MIT license.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | #4 | $0.085 | $0.014 | $0.175 | chat,reasoning,coding,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 91.01 | #2 | $0.05 | $0.01 | $0.10 | chat,fast,open-source,cheap,coding,privacy,reasoning,long-context,agents | openai-chat-completions |
| DeepArc 0xfc36…8842 | 84.89 | #7 | $0.20 | $0.01 | $0.45 | chat,coding,reasoning,fast | openai-chat-completions |
| Open Forge 0x1d90…b0aa | 75.54 | #1 | $0.00 | $0.00 | $0.00 | chat,reasoning,coding,fast,free | openai-chat-completions |
| Open Ant 0xe4f6…5bc4 | 74.05 | #5 | $0.138 | $0.028 | $0.275 | chat,coding,reasoning,tools,long-context,cheap | openai-chat-completions |
| NovaRoute AI 0xc50d…ed7b | 67.51 | #6 | $0.198 | $0.198 | $0.4455 | chat,coding,code,reasoning,tasks,deepseek,fast,value,surplus,openai-compatible,low-cost,verified,github,response-auth,base-usdc,monitored | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 66.94 | #3 | $0.069 | $0.014 | $0.1375 | chat,coding,reasoning,tools,long-context,cheap | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.35 | gated | $0.1151 | $0.0234 | $0.2293 | — | — |
| Phala 0x88c8…15d6 | 37.80 | gated | $0.20 | $0.07 | $0.40 | chat,confidential | openai-chat-completions |
| bartly.eth64.de 0x666e…4666 | 33.99 | gated | $0.09 | $0.015 | $0.20 | chat,coding,reasoning,fast | openai-chat-completions |
| D5V1N2 0xd5e7…7be0 | 33.17 | gated | $0.04 | $0.04 | $0.12 | chat,reasoning,coding,fast,cheap,router,fallback,deepseek | openai-chat-completions |
| Open Bird 0xc0f1…8183 | 18.32 | gated | $0.049 | $0.0098 | $0.098 | chat,coding | openai-chat-completions |
| Hana Gateway ✅ 0x4ae1…117b | 17.87 | gated | $0.02 | $0.005 | $0.04 | chat,coding,fast,tools | openai-chat-completions |
| Meridian AI 0x8c8c…06f5 | 5.64 | gated | $0.0191 | $0.0191 | $0.0383 | chat,coding | openai-chat-completions |
| uomi.ai 0x87df…48e3 | 4.80 | gated | $0.103 | $0.103 | $0.206 | chat,math,coding | openai-chat-completions |
| Jarvis 0x5beb…a1c9 | 1.76 | gated | $0.10 | $0.10 | $0.50 | chat,coding,research,fast | openai-chat-completions |
| Xerxes Inference 0xdc73…65ce | 0.52 | gated | $0.05 | $0.02 | $0.10 | chat,code,reasoning | openai-responses |
| Apex TEE Test 0xe672…7955 | 0.51 | gated | $0.00 | $0.00 | $0.00 | chat,fast,open-source,cheap,coding,privacy,reasoning,long-context,agents | openai-chat-completions |
| antseed-opal-marten-4af3 0x17f5…4af3 | 0.48 | gated | $10.00 | $10.00 | $10.00 | chat | openai-chat-completions |
| minion0x 0x215e…e2e3 | 0.34 | gated | $0.08 | $0.02 | $0.31 | chat,coding,math,fast | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0.02 | gated | $0.15 | $0.15 | $0.30 | chat,reasoning,agent,function-calling | openai-chat-completions |
| Colony Router 0x8aa0…5d3f | 0.00 | gated | $0.50 | $0.50 | $1.50 | — | openai-chat-completions |
| Deep Node 0x668e…6cd0 | 0.00 | gated | $0.14 | $0.0028 | $0.28 | chat,coding,fast,tasks | openai-chat-completions |
| AntFeed 0xddb6…1442 | 0.00 | gated | $0.1232 | $0.1232 | $0.2464 | chat,fast | openai-chat-completions |
| deepseek 0x47f9…ccd7 | 0.00 | gated | $0.07 | $0.003 | $0.28 | chat,fast | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 0.00 | gated | $0.058 | $0.028 | $0.1194 | chat,coding,math | openai-chat-completions |
| Leftermute 0x388b…5389 | 0.00 | gated | $0.0429 | $0.0429 | $0.0884 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.