Qwen 3 235B A22B Thinking 2507

ReasoningWeb searchFunction calling
Advertised as qwen-3-235b-a22b-thinking-2507qwen3-235b-a22b-thinking-2507qwen3-235b-thinking
Quick reference
Qwen 3 235B A22B Thinking 2507 — TLDR
  • 🧠 Reasoning-only model tuned for deep, multi-step problem solving
  • 🏢 Built by Alibaba's Qwen team, Apache 2.0 licensed
  • 🔧 Mixture-of-Experts: 235B total, 22B activated per token
  • 📏 Model card cites native 262,144-token (256K) context
  • 🆕 Split from dual-mode predecessor into dedicated thinking variant
  • ⚡ Served here in FP8 for efficient deployment
  • 🎯 Increased "thinking length" for highly complex tasks
  • 🔧 Strong tool-calling and agentic use via Qwen-Agent
💰 Best price on AntSeed
$0.016 / $0.01693%
per 1M · cheapest in / out
📏 Context
128K tokens
🐜 Sellers
11
advertising on AntSeed
Provider

Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research, developing the Qwen…

Explore 39 more models by Alibaba Group
About this model

Qwen 3 235B A22B Thinking 2507 is a reasoning-specialized large language model from Alibaba's Qwen team, released under the Apache 2.0 license. Architecturally it is a Mixture-of-Experts transformer with 235 billion total parameters and roughly 22 billion activated per token, using 128 experts with 8 active per token across 94 layers with grouped-query attention. It always operates in "thinking" mode, emitting reasoning traces before its final answer, and is positioned for in-depth research, technical work, and long, complex documents.

The most direct family comparison is to the original Qwen3-235B-A22B, which uniquely combined thinking and non-thinking behavior in a single switchable model. With the 2507 refresh, Qwen split that design into two dedicated checkpoints: a non-thinking Qwen 3 235B A22B Instruct 2507 and this thinking-only release. Qwen describes the 2507 line as featuring significant enhancements over the previous version, including extended 256K long-context understanding, and notes this thinking variant has an increased thinking length recommended for the hardest reasoning tasks.

Per the model card, it natively handles a 262,144-token context, and an optional configuration extends inputs toward one million tokens with sparse attention. The catalog exposes a 128K-token window in an FP8 quantization for more efficient serving.

Within the broader Qwen lineup, it sits alongside vision-language siblings such as Qwen3 VL 235B and the efficiency-focused Qwen 3 Next 80b. It retains strong tool-calling and agentic integration through the Qwen-Agent framework.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
qwenlm.github.ioQwen3: Think Deeper, Act Faster | Qwen· qwenlm.github.iohuggingface.coQwen/Qwen3-235B-A22B-Thinking-2507 · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Usage on AntSeed
Tokens served
111.49k
input + output
Requests
612
settled calls
Buyers
5
distinct, on this model
Sellers used
6
of 11 advertising
Settled
$0.25
gross USDC, this model
Sellers serving Qwen 3 235B A22B Thinking 2507 (11)compare on the network explorer →
SellerReputationRoutingInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.