Gemma 4 31B
Google · Open weights · Released Apr 2, 2026 · google/gemma-4-31b-it
Best $0.09 in / $0.34 out via DeepInfra · fp4
262K contextVisionToolsReasoningStructured outputPrompt cachingOpen weights
Available from 15 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 24 HOURS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | SMOKE | UPTIME |
|---|---|---|---|---|---|---|---|---|---|
| $0.09 | $0.27 | — | — | — | — | — | — | — | |
| $0.09 | $0.34 | $0.05 | 262K | fp4 | ✕ | ✕ | 5/5 | 100.0% | |
| $0.10 | $0.30 | — | — | — | — | — | — | — | |
| $0.10 | $0.34 | $0.10 | 262K | fp4 | 3s | 46 tok/s | 5/5 | 100.0% | |
| $0.12 | $0.37 | $0.012 | 131K | fp4 | 3.2s | 29 tok/s | 5/5 | 59.3% | |
| $0.12 | $0.36 | $0.09 | 256K | fp4 | ✕ | ✕ | 5/5 | 99.8% | |
| $0.14 | $0.40 | $0.14 | 262K | bf16 | ✕ | ✕ | — | 98.6% | |
| $0.14 | $0.40 | — | 262K | — | ✕ | ✕ | — | 94.0% | |
| $0.14 | $0.40 | — | 262K | bf16 | ✕ | ✕ | — | 63.7% | |
| $0.15 | $0.40 | $0.06 | 262K | fp8 | 1.4s | 38 tok/s | — | 97.2% | |
| $0.20 | $0.40 | — | 262K | fp8 | — | — | — | 96.7% | |
| $0.361 | $1.09 | $0.18 | 262K | — | ✕ | ✕ | — | 96.8% | |
| $0.38 | $1.15 | — | 131K | — | 2s | 196 tok/s | — | 87.9% | |
| $0.75 | $1.00 | $0.20 | 262K | fp4 | ✕ | ✕ | — | 99.9% | |
| $0.75 | $1.00 | $0.25 | 262K | fp8 | 3.5s | 57 tok/s | 5/5 | 84.8% |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · SMOKE: of 5 programmatic capability probes passed (JSON schema, instruction following, tool call, long-context, benign compliance; retry-once) · screened via OpenRouter routing · v0.2 · methodology v1.1 · measured 24 hours ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Input price history · best listed price, observed
$0.09 · SEP 27, 2026$0.09 · TODAY
Pricing variants · USD / MTok in / out
STANDARD
$0.09 / $0.34
per MTok in / out
FREE
Free / Free
per MTok in / out
CACHE READ $0.05
Recent changesRSS ↗
Oct 5, 2026endpoint removed · dekallm16 → 15
Oct 1, 2026endpoint added · cheaperinference14 → 16
Oct 1, 2026endpoint added · dekallm14 → 16
Sep 30, 2026endpoint removed · reka15 → 14
Sep 27, 2026price decreased · input0.09 → 0.08
Sep 27, 2026price decreased · output0.34 → 0.3
Sep 26, 2026endpoint added · io-net13 → 15
Sep 26, 2026endpoint added · reka13 → 15
About
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Specifications
Context window262,144
Max output16,384
Modalities inimage, text, video
Modalities outtext
LicenseOpen weights
ReleasedApr 2, 2026
TokenizerGemma
WeightsHugging Face ↗
Related
Embed badge
[](https://modelindex.ai/models/google/gemma-4-31b-it)Sources
OpenRouter API23 minutes ago
AGGREGATORHugging Face23 minutes ago
COMMUNITYModelIndex benchmarks20 days ago
MEASUREDVerified 23 minutes ago · list prices, not negotiated rates.