Google Gemma 4 26B A4B Instruct
gemma-4-26b-a4bgemma-4-26b-a4b-itgoogle-gemma-4-26b-a4b-instructgoogle-gemma-4-26b-a4b-it- 🧠 Mixture-of-Experts: 26B total parameters, only ~4B active per token.
- ⚡ Google reports it runs almost as fast as a 4B dense model.
- 📏 256K-token context window across the larger Gemma 4 models.
- 👁️ Accepts text, image, and video input.
- 🔧 Native function calling among its listed capabilities.
- 🧠 Configurable thinking modes for step-by-step reasoning.
- 🌐 Multilingual support across 140+ languages.
- 🔒 Open-weight, instruction-tuned release under Apache 2.0.
Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its DeepMind and Google…
Explore 17 more models by Google →Gemma 4 26B A4B Instruct is an open-weight, instruction-tuned model from Google DeepMind, released in April 2026 as part of the Gemma 4 family. Unlike its dense siblings, it uses a Mixture-of-Experts design: of its roughly 26 billion total parameters, only about 4 billion activate per token, so all weights load into memory while inference stays fast. Google describes it as running almost as quickly as a 4B dense model.
Compared to same-family predecessors, this model advances on several fronts. Where the earlier [[sibling:google-gemma-3-27b-it|Google Gemma 3 27B Instruct]] used a dense architecture, Gemma 4 introduces MoE variants alongside dense ones, plus built-in configurable reasoning modes and video input. Its context window reaches 256K tokens, and multilingual coverage spans over 140 languages.
Within Gemma 4 itself, the 26B A4B is positioned as the throughput-optimized counterpart to the dense [[sibling:google-gemma-4-31b-it|Google Gemma 4 31B Instruct]]. Sparse activation reduces compute per token relative to the dense 31B while aiming for comparable quality.
Both target consumer GPUs and workstations. It supports text, image, and video input natively, with audio featured on the smaller family members rather than this size. Its listed capabilities include reasoning, vision, function calling, and web search, offered as an open-weight, multilingual option.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 99 | gated | $0.0813 | $0.0813 | $0.25 | chat,reasoning,vision,video,multimodal,web-search | openai-chat-completions |
| surplusintelligence.ai 0x0e49…8927 | 68 | #1 | $0.039 | $0.015 | $0.12 | agents,anon,chat,cheap,function-calling,multimodal,reasoning,research,tasks,tools,translate,video,vision,web-search | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $0.1128 | $0.0436 | $0.36 | chat,coding,free,json,math,role-play,tools | — |
| uomi.ai 0x87df…48e3 | 19 | gated | $0.10 | $0.10 | $0.389 | chat,math,coding | openai-chat-completions |
| ZLKPro-Api 0x0b0b…f446 | 18 | gated | $0.001 | $0.0001 | $0.01 | free,chat,role-play | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 3 | gated | $0.0554 | $0.0554 | $0.1705 | chat,math | openai-chat-completions |
| antseed-opal-badger-2580 0xc85d…2580 | 1 | gated | $0.065 | $0.065 | $0.20 | chat | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0 | gated | $0.15 | $0.15 | $0.40 | chat,reasoning,vision,function-calling | openai-chat-completions |
| Leftermute 0x388b…5389 | 0 | gated | $0.0348 | $0.0348 | $0.0929 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.