Qwen 3 235B A22B Thinking 2507
qwen-3-235b-a22b-thinking-2507qwen3-235b-a22b-thinking-2507qwen3-235b-thinking- 🧠 Reasoning-only model tuned for deep, multi-step problem solving
- 🏢 Built by Alibaba's Qwen team, Apache 2.0 licensed
- 🔧 Mixture-of-Experts: 235B total, 22B activated per token
- 📏 Model card cites native 262,144-token (256K) context
- 🆕 Split from dual-mode predecessor into dedicated thinking variant
- ⚡ Served here in FP8 for efficient deployment
- 🎯 Increased "thinking length" for highly complex tasks
- 🔧 Strong tool-calling and agentic use via Qwen-Agent
Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research, developing the Qwen…
Explore 39 more models by Alibaba Group →Qwen 3 235B A22B Thinking 2507 is a reasoning-specialized large language model from Alibaba's Qwen team, released under the Apache 2.0 license. Architecturally it is a Mixture-of-Experts transformer with 235 billion total parameters and roughly 22 billion activated per token, using 128 experts with 8 active per token across 94 layers with grouped-query attention. It always operates in "thinking" mode, emitting reasoning traces before its final answer, and is positioned for in-depth research, technical work, and long, complex documents.
The most direct family comparison is to the original Qwen3-235B-A22B, which uniquely combined thinking and non-thinking behavior in a single switchable model. With the 2507 refresh, Qwen split that design into two dedicated checkpoints: a non-thinking Qwen 3 235B A22B Instruct 2507 and this thinking-only release. Qwen describes the 2507 line as featuring significant enhancements over the previous version, including extended 256K long-context understanding, and notes this thinking variant has an increased thinking length recommended for the hardest reasoning tasks.
Per the model card, it natively handles a 262,144-token context, and an optional configuration extends inputs toward one million tokens with sparse attention. The catalog exposes a 128K-token window in an FP8 quantization for more efficient serving.
Within the broader Qwen lineup, it sits alongside vision-language siblings such as Qwen3 VL 235B and the efficiency-focused Qwen 3 Next 80b. It retains strong tool-calling and agentic integration through the Qwen-Agent framework.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | gated | $0.225 | $0.225 | $1.75 | chat,reasoning,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 91.02 | #1 | $0.225 | $0.23 | $1.00 | chat,open-source,math,coding,reasoning,long-context,agents | openai-chat-completions |
| Open Ant 0xe4f6…5bc4 | 74.05 | #2 | $0.45 | $0.1463 | $3.50 | chat,reasoning,tools,cheap | openai-chat-completions |
| Chutes 0xded6…657c | 44.01 | gated | $0.3288 | $0.0329 | $1.3153 | chat,reasoning,coding,math,tee | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.37 | gated | $0.33 | $0.33 | $2.5666 | — | — |
| D5V1N2 0xd5e7…7be0 | 33.17 | gated | $0.155 | $0.14 | $1.06 | chat,reasoning,research,thinking,router,fallback,qwen | openai-chat-completions |
| Meridian AI 0x8c8c…06f5 | 5.64 | gated | $0.0162 | $0.0162 | $0.0162 | chat,reasoning | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0.51 | gated | $0.108 | $0.108 | $0.108 | chat,open-source,math,coding,reasoning,long-context,agents | openai-chat-completions |
| AntFeed 0xddb6…1442 | 0.00 | gated | $0.45 | $0.45 | $2.1094 | reasoning | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 0.00 | gated | $0.1535 | $0.1535 | $1.1935 | chat,math | openai-chat-completions |
| Leftermute 0x388b…5389 | 0.00 | gated | $0.0409 | $0.0409 | $0.3182 | chat,coding,json,reasoning,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.