Google Gemma 4 31B Instruct
gemma-4-31bgemma-4-31b-itgemma-4-31B-itgemma4-31bgoogle-gemma-4-31b-instructgoogle-gemma-4-31b-itgoogle/gemma-4-31b-it- 🧠 Dense 30.7B open model from Google DeepMind for reasoning
- 🆕 Configurable thinking modes toggled via a reasoning token
- 📏 256K-token context window for long documents and code
- 👁️ Handles text and image input; video processed as frames
- 🔧 Native function calling for agentic, tool-using workflows
- 🏢 Quantized checkpoints target consumer GPUs and workstations
- 🔒 Apache 2.0 license; open pre-trained and instruction-tuned weights
- 📚 Hybrid local/global attention with Proportional RoPE for long context
Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its DeepMind and Google…
Explore 21 more models by Google →Gemma 4 31B Instruct is the dense flagship of Google DeepMind's Gemma 4 family, a 30.7B-parameter multimodal model that accepts text and image input (and can process video as sequences of frames) while generating text output. It offers a 256K-token context window, native function calling, and configurable thinking modes, aimed at running reasoning, coding, and multimodal tasks under an Apache 2.0 license.
Architecturally it is a dense transformer paired with a vision encoder, using a hybrid attention scheme that interleaves local sliding-window layers with full global attention and Proportional RoPE (p-RoPE) for efficient long-context handling; quantization-aware and w4a16 checkpoints are published for smaller-footprint deployment.
Relative to the sibling Gemma 4 26B A4B Instruct, a Mixture-of-Experts variant with fewer active parameters, this 31B is dense—trading that inference efficiency for the family's highest-quality tier. Against the previous generation Gemma 3 27B, Google DeepMind highlights Gemma 4's built-in reasoning with configurable thinking, native system-prompt and function-calling support, and coding improvements.
Google DeepMind publishes instruction-tuned results in the official Gemma 4 31B model card, spanning reasoning, coding, vision, long-context, and safety tasks, and states the models undergo the same safety evaluations as its proprietary Gemini models.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | gated | $0.0875 | $0.0875 | $0.25 | chat,reasoning,vision,video,multimodal,web-search | openai-chat-completions |
| NovaRoute AI 0xc50d…ed7b | 67.51 | #1 | $0.009 | $0.009 | $0.024 | chat,coding,code,reasoning,tasks,gemma,value,surplus,openai-compatible,low-cost,verified,github,response-auth,base-usdc,monitored | openai-chat-completions |
| Chutes 0xded6…657c | 44.01 | gated | $0.132 | $0.0132 | $0.407 | chat,reasoning,vision,tee | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.35 | gated | $0.0901 | $0.0676 | $0.2704 | — | — |
| Phala 0x88c8…15d6 | 37.80 | gated | $0.15 | $0.075 | $0.46 | chat,confidential,reasoning,multimodal | openai-chat-completions |
| ZLKPro-Api 0x0b0b…f446 | 29.97 | gated | $0.00 | $0.00 | $0.00 | free,chat,role-play | openai-chat-completions |
| uomi.ai 0x87df…48e3 | 4.80 | gated | $0.096 | $0.096 | $0.296 | chat,math,coding | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0.51 | gated | $0.1089 | $0.1089 | $0.3226 | chat,open-source,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
| minion0x 0x215e…e2e3 | 0.34 | gated | $0.01 | $0.01 | $0.02 | chat,coding,math,fast | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0.02 | gated | $0.15 | $0.15 | $0.40 | chat,reasoning,vision,function-calling | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 0.00 | gated | $0.0409 | $0.09 | $0.1228 | chat,math | openai-chat-completions |
| Leftermute 0x388b…5389 | 0.00 | gated | $0.0348 | $0.0348 | $0.0929 | chat,coding,json,tools | openai-chat-completions |
| Inference Ready 0x6eb5…ad9a | 0.00 | gated | $0.02 | $0.02 | $0.06 | chat,coding,reasoning,multimodal | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.