GLM 4.7 Flash
glm-4-7-flashglm-4.7-flashglm-4.7-flash-p- 🧠 30B-A3B Mixture-of-Experts reasoning model, roughly 3B active parameters.
- 🔧 Optimized for agentic coding, tool use, and long-horizon planning.
- 📏 Context window near 200K tokens in this configuration.
- 🔒 Runs inside a Trusted Execution Environment with hardware attestation.
- 🆕 Uses interleaved/"preserved" thinking before each tool call.
- 🌐 Includes web search and reasoning capabilities; MIT-licensed.
- 🏢 Built by Z.ai (formerly Zhipu AI), released 2026.
Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of large language…
Explore 17 more models by Z.ai →GLM 4.7 Flash is the efficiency-tier member of Z.ai's GLM 4.7 generation, built as a 30-billion-parameter Mixture-of-Experts model that activates only around 3 billion parameters per token, making it suited to lightweight and local deployment. This particular listing runs the model inside a Trusted Execution Environment, adding hardware attestation evidence so users can independently verify that inference happens in a sealed enclave — a privacy-and-verifiability feature layered on top of the standard weights. It supports a context window approaching 200K tokens and is tuned specifically for coding, tool collaboration, and long-horizon agentic workflows.
Compared with the larger flagship GLM 4.7, this Flash variant trades raw capacity for speed and deployability while staying within the same release family. Against earlier GLM generations such as GLM 4.6, Z.ai positions the 4.7 line around stronger agentic coding, repository-level understanding, and "interleaved thinking," where the model reasons sequentially before each action rather than planning everything upfront.
Within Z.ai's broader catalog, it sits alongside successors like GLM 5 and GLM 5.1, remaining the compact, agent-focused option for developers prioritizing efficiency and verifiable execution.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | gated | $0.0625 | $0.0625 | $0.25 | chat,reasoning,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 91.01 | #1 | $0.02 | $0.0033 | $0.14 | chat,fast,open-source,cheap,reasoning,long-context,agents | openai-chat-completions |
| Open Forge 0x1d90…b0aa | 75.54 | #3 | $0.06 | $0.01 | $0.40 | chat,fast | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 66.94 | #2 | $0.03 | $0.005 | $0.20 | chat,reasoning,tools,cheap | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.35 | gated | $0.0534 | $0.0089 | $0.3563 | — | — |
| Open Bird 0xc0f1…8183 | 18.32 | gated | $0.03 | $0.005 | $0.20 | chat | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0.51 | gated | $0.054 | $0.009 | $0.36 | chat,fast,open-source,cheap,reasoning,long-context,agents | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0.02 | gated | $0.10 | $0.10 | $0.50 | chat,reasoning,agent,function-calling | openai-chat-completions |
| Leftermute 0x388b…5389 | 0.00 | gated | $0.0118 | $0.0118 | $0.05 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.