Llama 4 Maverick
Meta · Open weights · Released Apr 5, 2025 · meta-llama/llama-4-maverick
1.05M contextVisionToolsStructured outputPrompt cachingOpen weights
Available from 5 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | SMOKE | UPTIME |
|---|---|---|---|---|---|---|---|---|---|
| DIDigitalOcean | $0.20 | $0.696 | — | 128K | — | 1s | 24 tok/s | 5/5 | 100.0% |
| $0.20 | $0.80 | — | 1.05M | fp8 | 0.37s | 95 tok/s | 4/5 | 99.9% | |
| $0.27 | $0.85 | — | 1.05M | fp8 | 0.37s | 110 tok/s | 4/5 | 99.9% | |
| $0.35 | $1.00 | $0.17 | 524K | fp8 | 0.87s | 106 tok/s | 5/5 | 100.0% | |
| $0.35 | $1.15 | — | 524K | — | 0.7s | 162 tok/s | 5/5 | 99.9% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · SMOKE: of 5 programmatic capability probes passed (JSON schema, instruction following, tool call, long-context, benign compliance; retry-once) · screened via OpenRouter routing · v0.2 · methodology v1.1 · measured 44 hours ago
About
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Specifications
Context window1,048,576
Max output16,384
Modalities intext, image
Modalities outtext
LicenseOpen weights
ReleasedApr 5, 2025
Knowledge cutoff2024-08-31
TokenizerLlama4
WeightsHugging Face ↗
Open weightsGATED
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/meta-llama/llama-4-maverick)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks44 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.