DeepSeekDeepSeek·text

DeepSeek V4 Flash 0731

CodeReasoningWeb searchFunction calling
Advertised as deepseek-v4-flash-0731
Quick reference
DeepSeek V4 Flash 0731 — TLDR
  • 🧠 Sparse Mixture-of-Experts: 284B total parameters, about 13B active per token
  • 📏 One-million-token context window for long-document and repository-scale work
  • ⚡ Efficiency-tuned V4 variant aimed at fast, high-throughput inference
  • 🆕 July 2026 dated build of the V4 Flash line
  • 🔧 Supports function calling, code-optimized generation and web search
  • 🎯 Artificial Analysis measures 50 on its Intelligence Index at maximum effort
  • 🌐 Hybrid attention design cuts long-context compute and KV cache
  • 🏢 Built by DeepSeek, the Hangzhou lab backed by High-Flyer
💰 Best price on AntSeed
$0.016 / $0.03089%
per 1M · cheapest in / out
📏 Context
1M tokens
🐜 Sellers
11
advertising on AntSeed
Provider

DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…

Explore 6 more models by DeepSeek
About this model

DeepSeek V4 Flash 0731 is the efficiency-oriented member of DeepSeek's V4 generation: a sparse Mixture-of-Experts model with 284B total parameters and roughly 13B activated per token, paired with a one-million-token context window. It sits below the larger [[sibling:deepseek-v4-pro|DeepSeek V4 Pro]], which shares the same million-token context but is aimed at the heaviest knowledge and agentic workloads.

Architecturally, the V4 line departs from [[sibling:deepseek-v3.2|DeepSeek V3.2]] through a hybrid attention design that combines compressed sparse attention with heavily compressed attention. DeepSeek's own model card describes this as substantially reducing per-token inference FLOPs and KV cache footprint at million-token context relative to V3.2 — the efficiency recipe that lets the Flash tier serve very long inputs at high throughput. The card also outlines a two-stage post-training pipeline: domain-specific expert models trained with supervised fine-tuning and reinforcement learning, then consolidated into a single model via on-policy distillation.

Compared with the earlier April 2026 [[sibling:deepseek-v4-flash|DeepSeek V4 Flash]] release, the 0731 build is a dated refresh within the same family: parameter counts and context window are unchanged, so differences come from updated training rather than a new architecture. Independent evaluation from Artificial Analysis places the reasoning, maximum-effort configuration at 50 on its Intelligence Index.

Catalog capabilities include reasoning, code optimization, function calling and web search, suiting long-context agents, repository-scale coding and high-volume batch pipelines. An encrypted-serving variant is listed separately as [[sibling:e2ee-deepseek-v4-flash|DeepSeek V4 Flash]].

View source on GitHub ↗View model card on HuggingFace ↗
Sources
artificialanalysis.aiDeepSeek V4 Flash 0731 (max) - Intelligence, Performance & Price Analysis· artificialanalysis.aihuggingface.codeepseek-ai/DeepSeek-V4-Flash · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Usage on AntSeed
Tokens served
196.28M
input + output
Requests
20,764
settled calls
Buyers
19
distinct, on this model
Sellers used
9
of 11 advertising
Settled
$8.45
gross USDC, this model
Sellers serving DeepSeek V4 Flash 0731 (11)compare on the network explorer →
SellerReputationRoutingInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.