Llama 4 Scout
Meta · Open weights · Released Apr 5, 2025 · meta-llama/llama-4-scout
1.31M contextVisionToolsStructured outputPrompt cachingOpen weights
Available from 4 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.10 | $0.30 | — | 328K | fp8 | 0.41s | 50 tok/s | 99.9% | |
| $0.11 | $0.34 | $0.055 | 131K | — | 0.67s | 439 tok/s | 99.7% | |
| $0.18 | $0.59 | — | 131K | bf16 | 0.86s | 77 tok/s | 99.9% | |
| $0.25 | $0.70 | — | 1.31M | — | 0.72s | 208 tok/s | 99.9% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 44 hours ago
About
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Specifications
Context window1,310,720
Max output16,384
Modalities intext, image
Modalities outtext
LicenseOpen weights
ReleasedApr 5, 2025
Knowledge cutoff2024-08-31
TokenizerLlama4
WeightsHugging Face ↗
Open weightsGATED
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/meta-llama/llama-4-scout)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks44 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.