DeepSeek V4 Flash 0731 Fast
deepseek-v4-flash-0731-fast- 🧩 284B-parameter MoE, only 13B active per token
- 📏 Massive 1M-token context window
- ⚡ Throughput-tuned serving profile for low-latency workloads
- 📜 MIT licensed and openly downloadable
- 🔧 Function calling and web search supported
- 🎯 Strong reasoning plus code-optimized performance
DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…
Explore 6 more models by DeepSeek →DeepSeek V4 Flash 0731 Fast is the speed-tuned serving variant of DeepSeek's efficiency-oriented Flash line, pairing a 284B-parameter Mixture-of-Experts architecture with just 13B active parameters per token. That sparsity is what makes the "Fast" positioning viable: the model keeps a 1M-token context window and MIT licensing while targeting high-throughput, latency-sensitive deployment rather than maximum raw capability.
Within DeepSeek's lineup, the Flash tier sits below the heavier reasoning-focused [[sibling:deepseek-v4-pro|DeepSeek V4 Pro]], and this build is the accelerated counterpart to the standard [[sibling:deepseek-v4-flash-0731|V4 Flash 0731]] checkpoint released at the end of July 2026, with this variant arriving in August 2026. Earlier points in the family include [[sibling:deepseek-v4-flash|V4 Flash 0423]] and the [[sibling:deepseek-v3.2|V3.2]] generation, and there is also an end-to-end-encrypted [[sibling:e2ee-deepseek-v4-flash|DeepSeek V4 Flash]] option.
Capabilities cover reasoning, code generation, function calling, and web search, so it works well as a general-purpose agentic backend. It suits long-document analysis, repository-scale coding assistance, and high-volume pipelines where response speed and cost-efficient inference matter more than squeezing out the last few points of benchmark accuracy.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $0.3036 | $0.0759 | $0.6073 | chat | — |
| antseed-opal-badger-2580 0xc85d…2580 | 2 | gated | $0.175 | $0.175 | $0.35 | chat | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.