GLM 5V Turbo
glm-5v-turboz-ai-glm-5v-turbo- 🆕 Z.ai's first native multimodal agent foundation model.
- 👁️ Natively handles image, video, and text inputs.
- 🔧 Built for vision-based coding and agent-driven tasks.
- 📏 Roughly 200K-token context window.
- 🌐 Includes multimodal tools like screenshots and webpage reading.
- 💬 Targeted at agent-driven engineering workflows.
- 🎯 Supports reasoning, function calling, and web search.
- 🏢 Released April 2026 by Chinese lab Z.ai (formerly Zhipu AI).
Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of large language…
Explore 14 more models by Z.ai →GLM-5V-Turbo is Z.ai's first native multimodal agent foundation model, built specifically for vision-based coding and agent-driven tasks. Unlike the text-only models in the GLM line, it natively processes image, video, and text inputs together rather than relying on intermediate text descriptions, and is aimed at agentic engineering workflows. According to the catalog, it carries roughly a 200K-token context window.
The key generational distinction is vision. Its same-family predecessor, [[sibling:z-ai-glm-5-turbo|GLM 5 Turbo]], is a text-only model tuned for agent execution such as tool calling and long action chains; GLM-5V-Turbo inherits that agentic positioning and adds native multimodal perception. This lets it work directly from visual inputs—such as screenshots, design mockups, and document layouts—within coding and agent loops.
On the tooling side, Z.ai's documentation describes an expanded multimodal toolchain that includes capabilities like taking screenshots and reading webpages, supporting a perceive-then-act style of operation. The model also supports reasoning, function calling, and web search per the catalog capabilities.
Within the broader family—which includes the text-focused [[sibling:zai-org-glm-5|GLM 5]] and the later [[sibling:zai-org-glm-5-2|GLM 5.2]]—GLM-5V-Turbo is positioned as the agent-first, vision-capable branch rather than a general-purpose text upgrade.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 99 | #2 | $0.75 | $0.15 | $2.50 | chat,reasoning,coding,vision,multimodal,web-search | openai-chat-completions |
| surplusintelligence.ai 0x0e49…8927 | 68 | #1 | $0.45 | $0.09 | $1.50 | agents,chat,cheap,code,coding,developer,fast,frontier,function-calling,multimodal,reasoning,research,tasks,tools,translate,vision,web-search | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $1.3207 | $0.27 | $4.4024 | chat,coding,json,math,multimodal,reasoning,tools | — |
| Meridian AI 0x8c8c…06f5 | 22 | gated | $0.3224 | $0.3224 | $1.0746 | chat,coding,reasoning,multimodal | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 3 | gated | $0.5115 | $0.30 | $1.705 | chat,coding,math | openai-chat-completions |
| ⚡ MetaSpark 0xf629…3ec7 | 2 | gated | $0.50 | $0.09 | $1.66 | chat,vision,coding | openai-chat-completions |
| Leftermute 0x388b…5389 | 0 | gated | $0.303 | $0.303 | $1.01 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.