Z.aiZ.ai·text

GLM 5

CodeReasoningWeb searchFunction calling
Advertised as glm-5zai-org-glm-5
Quick reference
GLM 5 — TLDR
  • 🆕 Z.ai's February 2026 flagship for agentic engineering and reasoning.
  • 📏 744B-parameter MoE, 40B active, scaled up from GLM-4.5.
  • 🧠 Trained on 28.5T tokens with new asynchronous RL infrastructure.
  • 🔧 Adds DeepSeek Sparse Attention to cut training and inference cost.
  • 📚 Large context window (catalog: 198K tokens), FP8 weights available.
  • 🎯 Capabilities: reasoning, code-optimization, function calling, web search.
  • 🔒 Released open-weight under the permissive MIT license.
  • 💬 Vendor reports gains over GLM-4.7 across reasoning, coding, agentic tasks.
💰 Best price on AntSeed
$0.065 / $0.207
per 1M · cheapest in / out
📏 Context
198K tokens
🐜 Sellers
16
advertising on AntSeed
Provider

Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of large language…

Explore 14 more models by Z.ai
About this model

GLM 5 is the February 2026 flagship from Z.ai (formerly Zhipu AI), positioned for complex systems engineering and long-horizon agentic work. Architecturally it is a Mixture-of-Experts model with 744 billion total parameters and roughly 40 billion active per token, scaled up from GLM-4.5's 355B (32B active), with pre-training data expanded to 28.5 trillion tokens. It is distributed open-weight under the MIT license in both full-precision and FP8 formats.

The two headline changes over earlier generations are efficiency-focused. GLM 5 adopts DeepSeek Sparse Attention (DSA), which the technical report describes as dynamically allocating attention by token importance to lower compute without compromising long-context understanding — an advance over the standard MoE used in GLM-4.5. Post-training uses a new asynchronous reinforcement-learning infrastructure built on the "slime" framework that decouples generation from training to improve GPU utilization.

Relative to its same-family predecessor [[sibling:zai-org-glm-4.7|GLM 4.7]], Z.ai reports significant improvements across academic benchmarks in reasoning, coding, and agentic tasks.

GLM 5 was followed by refreshed siblings [[sibling:zai-org-glm-5-1|GLM 5.1]] and [[sibling:zai-org-glm-5-2|GLM 5.2]], the latter extending to a roughly 1M-token context with the IndexShare architecture.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
docs.z.aiGLM-5 - Overview - Z.AI DEVELOPER DOCUMENT· docs.z.aiarxiv.orgGLM-5: from Vibe Coding to Agentic Engineering· arxiv.orghuggingface.cozai-org/GLM-5 · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Usage on AntSeed
Tokens served
18.36M
input + output
Requests
1,175
settled calls
Buyers
14
distinct, on this model
Sellers used
11
of 16 advertising
Settled
$0.15
gross USDC, this model
Sellers serving GLM 5 (16)compare on the network explorer →
SellerReputationRoutingInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.