GLM 5.3 Flash
glm-5-3-flashglm-5.3-flashz-ai-glm-5-3-flashz-ai/glm-5.3-flash- 🏢 Z.ai's efficiency-focused Flash model, released August 2026, weights on Hugging Face
- 🧠 Mixture-of-experts: 320B total parameters, roughly 18B active per token
- 🆕 First GLM to combine sparse attention with linear attention
- 📏 Context window of 1,048,576 tokens
- 👁️ Handles visual context alongside text, returning text output
- 🔧 Function calling, web search and configurable reasoning effort
- 🎯 Built for long-horizon software engineering and sustained agentic loops
- 📚 Artificial Analysis measures 57 on its Intelligence Index
Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of large language…
Explore 17 more models by Z.ai →GLM 5.3 Flash is Z.ai's efficiency-oriented entry in the GLM-5 line, arriving days after the flagship GLM 5.3 and positioned for coding agents, complex reasoning and production workloads that combine text with visual context. It succeeds earlier compact releases such as GLM 4.7 Flash in the Flash tier.
Architecturally it departs from its predecessors. Z.ai describes a 320-billion-parameter mixture-of-experts network that activates roughly 18 billion parameters per token, and the first model in the GLM series to combine sparse attention with linear attention — a hybrid the company says lowers long-context serving cost while preserving precise long-context behaviour. The catalog context window is 1,048,576 tokens, and reasoning effort is configurable in the API. Function calling and web search are supported.
On generational gains, Z.ai's own model card and launch post report that GLM 5.3 Flash improves on GLM 5.2 across its benchmark suite and real-world workloads at substantially lower serving cost. Its published tables list, as vendor-run results, 63.4 on DeepSWE v1.1 against 46.2 for GLM 5.2; harnesses and context limits differ per test, so these figures are self-reported.
Independently, Artificial Analysis measures the model at 57 on its Intelligence Index. Weights are published on Hugging Face, continuing Z.ai's open-release practice for the GLM family.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| ▲ Apex Ant 0x73b4…e736 | 91.01 | #2 | $0.05 | $0.01 | $0.18 | chat,fast,open-source,cheap,coding,long-context,multimodal | openai-chat-completions |
| Open Forge 0x1d90…b0aa | 75.54 | #1 | $0.00 | $0.00 | $0.00 | chat,reasoning,coding,multimodal,fast,free | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 66.94 | #3 | $0.075 | $0.015 | $0.25 | chat,coding,reasoning,vision,multimodal,tools,long-context,cheap | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.35 | gated | $0.1053 | $0.0211 | $0.3509 | — | — |
| Phala 0x88c8…15d6 | 37.80 | gated | $0.15 | $0.03 | $0.50 | chat,confidential,reasoning,multimodal | openai-chat-completions |
| bartly.eth64.de 0x666e…4666 | 33.99 | gated | $0.09 | $0.02 | $0.28 | chat,coding,reasoning,fast,long-context,vision,multimodal,function-calling | openai-chat-completions |
| D5V1N2 0xd5e7…7be0 | 33.17 | gated | $0.094 | $0.072 | $0.315 | chat,coding,reasoning,fast,cheap,vision,video,multimodal,large-context,web-search,glm,router,fallback | openai-chat-completions |
| ZLKPro-Api 0x0b0b…f446 | 29.97 | gated | $0.02 | $0.01 | $0.07 | agent,chat,text,reasoning,research,smart,long-context,multimodal,visual | openai-chat-completions |
| Hana Gateway ✅ 0x4ae1…117b | 17.87 | gated | $0.03 | $0.03 | $0.09 | chat,coding,fast,tools | openai-chat-completions |
| zro 0x8398…2c06 | 16.73 | gated | $0.15 | $0.003 | $0.50 | chat,fast | openai-chat-completions |
| Katant 0xd2ae…d698 | 13.17 | gated | $0.015 | $0.003 | $0.05 | chat,fast | openai-chat-completions |
| Plaid Labs 0x06f4…9088 | 5.01 | gated | $0.03 | $0.007 | $0.12 | chat,coding,reasoning,cheap,fast,long-context,tools,agents,vision,multimodal,video,glm | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.