- 🆕 Z.ai's February 2026 flagship for agentic engineering and reasoning.
- 📏 744B-parameter MoE, 40B active, scaled up from GLM-4.5.
- 🧠 Trained on 28.5T tokens with new asynchronous RL infrastructure.
- 🔧 Adds DeepSeek Sparse Attention to cut training and inference cost.
- 📚 Large context window (catalog: 198K tokens), FP8 weights available.
- 🎯 Capabilities: reasoning, code-optimization, function calling, web search.
- 🔒 Released open-weight under the permissive MIT license.
- 💬 Vendor reports gains over GLM-4.7 across reasoning, coding, agentic tasks.
Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of large language…
Explore 14 more models by Z.ai →GLM 5 is the February 2026 flagship from Z.ai (formerly Zhipu AI), positioned for complex systems engineering and long-horizon agentic work. Architecturally it is a Mixture-of-Experts model with 744 billion total parameters and roughly 40 billion active per token, scaled up from GLM-4.5's 355B (32B active), with pre-training data expanded to 28.5 trillion tokens. It is distributed open-weight under the MIT license in both full-precision and FP8 formats.
The two headline changes over earlier generations are efficiency-focused. GLM 5 adopts DeepSeek Sparse Attention (DSA), which the technical report describes as dynamically allocating attention by token importance to lower compute without compromising long-context understanding — an advance over the standard MoE used in GLM-4.5. Post-training uses a new asynchronous reinforcement-learning infrastructure built on the "slime" framework that decouples generation from training to improve GPU utilization.
Relative to its same-family predecessor [[sibling:zai-org-glm-4.7|GLM 4.7]], Z.ai reports significant improvements across academic benchmarks in reasoning, coding, and agentic tasks.
GLM 5 was followed by refreshed siblings [[sibling:zai-org-glm-5-1|GLM 5.1]] and [[sibling:zai-org-glm-5-2|GLM 5.2]], the latter extending to a roughly 1M-token context with the IndexShare architecture.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 99 | #4 | $0.50 | $0.10 | $1.60 | chat,reasoning,coding,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 83 | #3 | $0.45 | $0.09 | $1.44 | chat,open-source,coding,reasoning,long-context,agents | openai-chat-completions |
| Open Forge 0x1d90…b0aa | 75 | #5 | $0.80 | $0.16 | $2.56 | chat,coding,reasoning | openai-chat-completions |
| surplusintelligence.ai 0x0e49…8927 | 68 | #2 | $0.30 | $0.06 | $0.96 | agents,anon,chat,cheap,code,coding,developer,frontier,function-calling,reasoning,research,tasks,tools,translate,web-search | openai-chat-completions |
| Open Bird 0xc0f1…8183 | 64 | #1 | $0.275 | $0.055 | $0.88 | chat | openai-chat-completions |
| NovaRoute AI 0xc50d…ed7b | 60 | gated | $0.9286 | $0.9286 | $2.8601 | chat,coding,code,reasoning,tasks,glm,value,surplus,openai-compatible,low-cost,verified,github,response-auth,base-usdc,monitored | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 40 | gated | $0.50 | $0.10 | $1.60 | chat,coding,reasoning,tools,cheap | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $0.90 | $0.18 | $2.8501 | base-usdc,chat,code,coding,github,glm,json,low-cost,math,monitored,openai-compatible,reasoning,response-auth,surplus,tasks,tools,value,verified | — |
| Meridian AI 0x8c8c…06f5 | 22 | gated | $0.0648 | $0.0648 | $0.2074 | chat,coding,reasoning | openai-chat-completions |
| uomi.ai 0x87df…48e3 | 19 | gated | $0.938 | $0.938 | $2.889 | chat,math,coding | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 3 | gated | $0.341 | $0.20 | $1.0912 | chat,coding,math | openai-chat-completions |
| antseed-opal-badger-2580 0xc85d…2580 | 1 | gated | $0.50 | $0.50 | $1.60 | chat | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0 | gated | $1.00 | $1.00 | $3.20 | chat,reasoning,agent,function-calling | openai-chat-completions |
| AntFeed 0xddb6…1442 | 0 | gated | $0.66 | $0.66 | $2.112 | chat | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0 | gated | $0.1482 | $0.0495 | $0.4743 | chat,open-source,coding,reasoning,long-context,agents | openai-chat-completions |
| Leftermute 0x388b…5389 | 0 | gated | $0.10 | $0.10 | $0.3772 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.