GoogleGoogle·text

Google Gemma 4 31B Instruct

VisionReasoningWeb searchFunction calling
Quick reference
Google Gemma 4 31B Instruct — TLDR
  • 🧠 Dense 30.7B open model from Google DeepMind for reasoning
  • 🆕 Configurable thinking modes toggled via a reasoning token
  • 📏 256K-token context window for long documents and code
  • 👁️ Handles text and image input; video processed as frames
  • 🔧 Native function calling for agentic, tool-using workflows
  • 🏢 Quantized checkpoints target consumer GPUs and workstations
  • 🔒 Apache 2.0 license; open pre-trained and instruction-tuned weights
  • 📚 Hybrid local/global attention with Proportional RoPE for long context
💰 Best price on AntSeed
$0.0090 / $0.02493%
per 1M · cheapest in / out
📏 Context
256K tokens
🐜 Sellers
9
advertising on AntSeed
Provider

Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its DeepMind and Google…

Explore 13 more models by Google
About this model

Gemma 4 31B Instruct is the dense flagship of Google DeepMind's Gemma 4 family, a 30.7B-parameter multimodal model that accepts text and image input (and can process video as sequences of frames) while generating text output. It offers a 256K-token context window, native function calling, and configurable thinking modes, aimed at running reasoning, coding, and multimodal tasks under an Apache 2.0 license.

Architecturally it is a dense transformer paired with a vision encoder, using a hybrid attention scheme that interleaves local sliding-window layers with full global attention and Proportional RoPE (p-RoPE) for efficient long-context handling; quantization-aware and w4a16 checkpoints are published for smaller-footprint deployment.

Relative to the sibling [[sibling:google-gemma-4-26b-a4b-it|Gemma 4 26B A4B Instruct]], a Mixture-of-Experts variant with fewer active parameters, this 31B is dense—trading that inference efficiency for the family's highest-quality tier. Against the previous generation [[sibling:google-gemma-3-27b-it|Gemma 3 27B]], Google DeepMind highlights Gemma 4's built-in reasoning with configurable thinking, native system-prompt and function-calling support, and coding improvements.

Google DeepMind publishes instruction-tuned results in the official Gemma 4 31B model card, spanning reasoning, coding, vision, long-context, and safety tasks, and states the models undergo the same safety evaluations as its proprietary Gemini models.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
docs.api.nvidia.comgoogle / gemma-4-31b-it· docs.api.nvidia.combuild.nvidia.comgemma-4-31b-it Model by Google· build.nvidia.comhuggingface.cogoogle/gemma-4-31B · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Sellers serving Google Gemma 4 31B Instruct (9)
SellerReputationInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens, not settled amounts). Reputation = on-chain trust (0-100). Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase.