Qwen3 VL 235B

VisionWeb searchFunction calling
Advertised as qwen3-vl-235bqwen3-vl-235b-a22b
Quick reference
Qwen3 VL 235B — TLDR
  • - 🧠 Alibaba's most powerful vision-language model in the Qwen series.
  • - 🔧 Mixture-of-experts design, 235B total with ~22B active parameters.
  • - 👁️ Strong visual perception, multilingual OCR, and document understanding.
  • - 📏 256K-token context for long documents, images, and video.
  • - 🎯 Spatial reasoning and video dynamics comprehension upgraded this generation.
  • - 💬 Function-calling and web-search capable; agent-oriented interactions.
  • - 🔒 Apache-2.0 license, served here at FP8 quantization.
  • - 🆕 Instruct and reasoning-enhanced Thinking editions available upstream.
💰 Best price on AntSeed
$0.085 / $0.51159%
per 1M · cheapest in / out
📏 Context
128K tokens
🐜 Sellers
4
advertising on AntSeed
Provider

Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research, developing the Qwen…

Explore 39 more models by Alibaba Group
About this model

Qwen3-VL 235B is the flagship multimodal model in Alibaba's Qwen vision-language line, combining text generation with image and video understanding in a mixture-of-experts architecture that activates roughly 22B of its 235B parameters per token. Alibaba describes the Qwen3-VL generation as the most powerful vision-language series in the Qwen lineup to date, citing comprehensive upgrades across text understanding, visual perception and reasoning, extended context length, spatial and video comprehension, and agent interaction.

Relative to the Qwen3 text flagships Qwen 3 235B A22B Instruct 2507 and Qwen 3 235B A22B Thinking 2507, this model adds native visual input while, per Alibaba's model card, maintaining text-only performance comparable to the flagship Qwen3 language models. It also scales up the smaller VL sibling Qwen3 VL 30B A3B, offering a far larger expert pool for heavier perception and reasoning workloads.

The catalog deployment exposes a 256K-token context window and supports function-calling and web-search, suiting document AI, multilingual OCR, UI/software assistance, and vision-language agent workflows. Upstream, Qwen3-VL ships in both Instruct and reasoning-enhanced Thinking editions, the latter tuned for multimodal reasoning. The weights are released under the Apache-2.0 license, and this instance runs at FP8 precision, lowering memory requirements for the large MoE model.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
build.nvidia.comqwen3-235b-a22b Model by Qwen· build.nvidia.comhuggingface.coQwen/Qwen3-VL-235B-A22B-Instruct · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Usage on AntSeed
Tokens served
2.17k
input + output
Requests
5
settled calls
Buyers
2
distinct, on this model
Sellers used
2
of 4 advertising
Settled
$0.00
gross USDC, this model
Sellers serving Qwen3 VL 235B (4)compare on the network explorer →
SellerReputationRoutingInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.