GLM 4.6
glm-4.6zai-org-glm-4-6zai-org-glm-4.6- 🆕 Z.ai's open-weight upgrade over GLM-4.5 across coding and reasoning
- 📏 Context window expanded from 128K to 200K tokens
- ⚡ Over 30% more efficient token use than GLM-4.5
- 🧠 Stronger reasoning with tool use during inference
- 🔧 Improved tool-using and search-based agentic performance
- 🔒 Released openly under the permissive MIT License
- 💬 Toggleable deep-thinking ("thinking") mode per request
Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of large language…
Explore 17 more models by Z.ai →GLM 4.6 is a large language model from Z.ai (formerly Zhipu AI), whose GLM family of open-weight models is released under the MIT License. It positions itself as an agentic, reasoning, and coding foundation model, with the catalog listing reasoning, function-calling, and web-search capabilities and a context window near 200K tokens.
Compared to its same-family predecessor GLM-4.5, Z.ai reports several concrete improvements: the context window grew from 128K to 200K tokens for more complex agentic tasks, average token consumption dropped by over 30%, and reasoning now supports tool use during inference for stronger overall capability. The company also reports stronger performance in tool-using and search-based agents and better integration within agent frameworks versus GLM-4.5. Z.ai evaluated GLM-4.6 across eight public benchmarks and published its test questions and agent trajectories for reproduction.
GLM-4.6 supports a toggleable deep-thinking mode, enabled or disabled per request, building on the interleaved-thinking approach introduced with GLM-4.5.
Within the lineage, GLM-4.6 was succeeded by GLM 4.7. The family later expanded with GLM 5, GLM 5.1, and GLM 5.2, alongside lightweight variants like GLM 4.7 Flash.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | #3 | $0.425 | $0.15 | $1.375 | chat,reasoning,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 91.01 | #1 | $0.15 | $0.03 | $0.61 | chat,open-source,reasoning,long-context,agents | openai-chat-completions |
| DeepArc 0xfc36…8842 | 84.89 | #4 | $0.65 | $0.12 | $2.35 | chat,coding,reasoning,fast | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 66.90 | #2 | $0.215 | $0.04 | $0.875 | chat,reasoning,tools,cheap | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.35 | gated | $0.3821 | $0.0711 | $1.5552 | — | — |
| Open Bird 0xc0f1…8183 | 18.32 | gated | $0.195 | $0.195 | $0.95 | chat | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0.51 | gated | $0.1032 | $0.0216 | $0.4176 | chat,open-source,reasoning,long-context,agents | openai-chat-completions |
| AntFeed 0xddb6…1442 | 0.00 | gated | $0.473 | $0.473 | $1.8192 | chat | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 0.00 | gated | $0.2899 | $0.30 | $0.9378 | chat,math | openai-chat-completions |
| Leftermute 0x388b…5389 | 0.00 | gated | $0.0773 | $0.0773 | $0.25 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.