Gemma 4 31B
Google · Open weights · Released Apr 2, 2026 · google/gemma-4-31b-it
262K contextVisionToolsReasoningStructured outputPrompt cachingOpen weights
Available from 18 providersUSD / MTOK · VERIFIED 12 HOURS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | SMOKE | UPTIME |
|---|---|---|---|---|---|---|---|---|---|
| OPOpenInference | $0.08 | $0.35 | $0.01 | 262K | bf16 | 2.3s | 21 tok/s | 4/5 | 98.0% |
| $0.09 | $0.34 | $0.05 | 262K | fp4 | 0.84s | 56 tok/s | 5/5 | 99.9% | |
| CWCoreWeave | $0.10 | $0.34 | $0.10 | 262K | bf16 | 1.6s | 38 tok/s | 5/5 | 98.3% |
| $0.12 | $0.36 | $0.09 | 256K | bf16 | 1.5s | 81 tok/s | 5/5 | 99.7% | |
| CHChutes | $0.12 | $0.37 | $0.012 | 131K | fp4 | 9s | 30 tok/s | 5/5 | 96.5% |
| $0.13 | $0.38 | — | 262K | fp8 | — | — | — | 99.3% | |
| $0.13 | $0.40 | — | 262K | fp8 | 2.7s | 82 tok/s | 5/5 | 83.4% | |
| $0.14 | $0.40 | $0.14 | 262K | — | 1.3s | 62 tok/s | — | 99.4% | |
| $0.14 | $0.40 | — | 262K | — | 2.3s | 95 tok/s | — | 99.8% | |
| $0.14 | $0.40 | — | 262K | bf16 | 1.2s | 37 tok/s | — | 98.5% | |
| $0.15 | $0.40 | $0.06 | 262K | fp8 | 1.4s | 91 tok/s | — | 99.6% | |
| PHPhala | $0.15 | $0.46 | $0.075 | 262K | — | 1.2s | 104 tok/s | — | 99.4% |
| $0.27 | $0.76 | — | 131K | fp8 | — | — | — | 99.4% | |
| $0.28 | $0.86 | — | 262K | — | 2.2s | 25 tok/s | — | 95.0% | |
| $0.38 | $1.15 | — | 131K | — | 2.2s | 200 tok/s | — | 100.0% | |
| $0.39 | $0.97 | — | 262K | — | — | — | — | 98.9% | |
| MOModelRun | $0.75 | $1.00 | $0.75 | 262K | fp4 | 0.31s | 205 tok/s | — | 100.0% |
| $0.99 | $1.49 | $0.99 | 131K | fp16 | 0.43s | 665 tok/s | — | 99.9% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · SMOKE: of 5 programmatic capability probes passed (JSON schema, instruction following, tool call, long-context, benign compliance; retry-once) · screened via OpenRouter routing · v0.1 · methodology v1.1 · measured 43 hours ago
Pricing variants · USD / MTok in / out
STANDARD
$0.08 / $0.35
per MTok in / out
FREE
Free / Free
per MTok in / out
CACHE READ $0.01
About
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Specifications
Context window262,144
Max output262,144
Modalities inimage, text, video
Modalities outtext
LicenseOpen weights
ReleasedApr 2, 2026
TokenizerGemma
WeightsHugging Face ↗
Related
Embed badge
[](https://modelindex.ai/models/google/gemma-4-31b-it)Sources
OpenRouter API12 hours ago
AGGREGATORHugging Face12 hours ago
COMMUNITYModelIndex benchmarks43 hours ago
MEASUREDVerified 12 hours ago · list prices, not negotiated rates.