Z.aiZ.ai·text

GLM 4.7 Flash

ReasoningWeb searchFunction calling
Advertised as zai-org-glm-4-7-flashzai-org-glm-4.7-flash
Quick reference
GLM 4.7 Flash — TLDR
  • 🆕 Speed-optimized variant of GLM-4.7, released January 2026.
  • 🧠 30B-A3B mixture-of-experts with roughly 3B active parameters.
  • 📏 Roughly 128K-token context window for long inputs.
  • 🔧 Supports function calling and web search.
  • ⚡ Optimized for fast, low-latency inference.
  • 🔒 Open weights under the permissive MIT license.
  • 📚 Targets lightweight, cost-efficient deployment in the 30B class.
  • 🌐 Runs on common inference stacks like vLLM and SGLang.
💰 Best price on AntSeed
$0.030 / $0.171
per 1M · cheapest in / out
📏 Context
128K tokens
🐜 Sellers
3
advertising on AntSeed
Provider

Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of large language…

Explore 17 more models by Z.ai
About this model

GLM 4.7 Flash is Z.ai's speed-optimized, lightweight member of the GLM-4.7 generation, released in early 2026 alongside the full-size GLM 4.7. According to its Hugging Face model card, it is a 30B-A3B mixture-of-experts model—30 billion total parameters with roughly 3 billion active per token—positioned for strong balance of performance and efficiency in the 30B class. It ships as open weights under the MIT license and runs on common inference stacks such as vLLM and SGLang.

Compared with the full GLM 4.7, the Flash variant is built for faster, cheaper inference by activating fewer parameters per token, while retaining the generation's coding focus, function calling, and web-search support. The catalog lists a 128K-token context window and reasoning capability.

The wider GLM-4.7 line advanced over GLM 4.6; Z.ai's GLM-4.7 model card reports the full model reaching 42.8% on Humanity's Last Exam, a gain over GLM-4.6, with improved tool use.

This makes Flash well suited to high-volume, latency-sensitive tasks, with the full GLM 4.7 available for harder jobs.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
huggingface.cozai-org/GLM-4.7-Flash · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Usage on AntSeed
Tokens served
190.34k
input + output
Requests
91
settled calls
Buyers
3
distinct, on this model
Sellers used
3
of 3 advertising
Settled
$0.02
gross USDC, this model
Sellers serving GLM 4.7 Flash (3)compare on the network explorer →
SellerReputationRoutingInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.