Llama 4 Scout
Meta · Open weights · Released Apr 5, 2025 · meta-llama/llama-4-scout
Best $0.10 in / $0.30 out via DeepInfra · fp8 · 328K ctx at this price (headline 1.31M)
1.31M contextVisionToolsStructured outputOpen weights
Available from 3 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 2 DAYS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.10 | $0.30 | — | 328K | fp8 | 0.43s | 62 tok/s | 100.0% | |
| $0.18 | $0.59 | — | 131K | bf16 | ✕ | ✕ | 97.1% | |
| $0.25 | $0.70 | — | 1.31M | — | ✕ | ✕ | — |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · methodology v1.1 · measured 2 days ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Recent changesRSS ↗
Sep 1, 2026endpoint removed · groq4 → 3
About
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Specifications
Context window1,310,720
Max output16,384
Modalities intext, image
Modalities outtext
LicenseOpen weights
ReleasedApr 5, 2025
Knowledge cutoff2024-08-31
TokenizerLlama4
WeightsHugging Face ↗
Open weightsGATED
safetensorspytorch
Related
Compare with
Embed badge
[](https://modelindex.ai/models/meta-llama/llama-4-scout)Sources
OpenRouter API23 minutes ago
AGGREGATORHugging Face23 minutes ago
COMMUNITYModelIndex benchmarks17 days ago
MEASUREDVerified 23 minutes ago · list prices, not negotiated rates.