Qwen 3 Next 80b
qwen-3-next-80bqwen3-next-80b- 🧠 80B Mixture-of-Experts activating only ~3B parameters per token
- 🆕 First model in Alibaba's Qwen3-Next architecture series
- 🔧 Hybrid attention: Gated DeltaNet plus Gated Attention for efficiency
- 📏 Native 256K-token context for long-document work
- ⚡ High sparsity and multi-token prediction boost throughput
- 🎯 Function calling and web search supported
- 🔒 Apache 2.0 license, openly downloadable weights
- 🏢 Built by Alibaba's Qwen team, fp16 served here
Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research, developing the Qwen…
Explore 39 more models by Alibaba Group →Qwen 3 Next 80b is the first release in Alibaba's Qwen3-Next architecture line, an 80-billion-parameter Mixture-of-Experts model that activates roughly 3 billion parameters per token, drastically reducing FLOPs while preserving capacity. Its defining feature is a hybrid attention design combining Gated DeltaNet with Gated Attention, paired with a high-sparsity MoE and multi-token prediction for faster, cheaper inference on long inputs. Released under Apache 2.0, it ships with a native 256K context window and supports function calling and web search.
Compared with the broader Qwen3 generation, Alibaba reports meaningful efficiency gains. The team states the underlying base model reaches performance comparable to—or slightly better than—the dense Qwen3-32B while using less than 10% of its training GPU hours. On the long-context RULER benchmark, Alibaba reports that the Instruct version outperforms the earlier Qwen3 30B A3B across all tested lengths, and even surpasses the flagship Qwen 3 235B A22B Instruct 2507 within 256K context.
This positions Qwen 3 Next 80b as the efficiency-focused step in the family, trading the dense scaling of older Qwen3 models for a sparser, architecturally novel approach. Venice serves it at fp16, optimized for speed, with the weights also deployable through common engines such as vLLM and SGLang.
The model exists in two post-trained forms in Alibaba's release—an instruct variant for chat and agents and a separate thinking variant for complex reasoning—both sharing the same hybrid attention and MoE backbone. Within this catalog it is the newest entry in its architecture family, sitting alongside many other Qwen-derived text, image, and embedding siblings.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | gated | $0.175 | $0.175 | $0.95 | chat,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 91.02 | #1 | $0.16 | $0.16 | $0.86 | chat,open-source,tasks,long-context | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.37 | gated | $0.2235 | $0.2235 | $1.2131 | — | — |
| D5V1N2 0xd5e7…7be0 | 33.17 | gated | $0.174 | $0.174 | $1.092 | chat,reasoning,coding,research,router,fallback,qwen | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0.51 | gated | $0.24 | $0.24 | $1.3616 | chat,open-source,tasks,long-context,agents | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0.02 | gated | $0.15 | $0.15 | $1.50 | chat,function-calling | openai-chat-completions |
| AntFeed 0xddb6…1442 | 0.00 | gated | $0.35 | $0.35 | $2.3722 | chat,coding | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 0.00 | gated | $0.1194 | $0.1194 | $0.6479 | chat | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.