Llama 3.1 8B Instruct
Meta · Open weights · Released Jul 23, 2024 · meta-llama/llama-3.1-8b-instruct
Best $0.02 in / $0.04 out via DeepInfra · fp8
131K contextToolsStructured outputPrompt cachingOpen weights
Available from 5 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 2 DAYS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | SMOKE | UPTIME |
|---|---|---|---|---|---|---|---|---|---|
| $0.02 | $0.04 | — | 131K | fp8 | ✕ | ✕ | 3/5 | 99.7% | |
| $0.02 | $0.05 | — | 16K | fp8 | ✕ | ✕ | 4/5 | 100.0% | |
| $0.05 | $0.08 | $0.025 | 131K | — | ✕ | ✕ | 4/5 | 100.0% | |
| $0.152 | $0.287 | — | 32K | fp8 | 2.4s | 14 tok/s | 3/5 | 100.0% | |
| $0.22 | $0.22 | $0.22 | 131K | bf16 | 0.4s | 147 tok/s | 4/5 | 100.0% |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · SMOKE: of 5 programmatic capability probes passed (JSON schema, instruction following, tool call, long-context, benign compliance; retry-once) · screened via OpenRouter routing · v0.2 · methodology v1.1 · measured 2 days ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
About
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Specifications
Context window131,072
Max output117,964
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJul 23, 2024
Knowledge cutoff2023-12-31
TokenizerLlama3
WeightsHugging Face ↗
Open weightsGATED
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/meta-llama/llama-3.1-8b-instruct)Sources
OpenRouter API23 minutes ago
AGGREGATORHugging Face23 minutes ago
COMMUNITYModelIndex benchmarks16 days ago
MEASUREDVerified 23 minutes ago · list prices, not negotiated rates.