Llama 3.1 8B Instruct
Meta · Open weights · Released Jul 23, 2024 · meta-llama/llama-3.1-8b-instruct
131K contextToolsStructured outputPrompt cachingOpen weights
Available from 5 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | SMOKE | UPTIME |
|---|---|---|---|---|---|---|---|---|---|
| $0.02 | $0.04 | — | 131K | fp8 | 1.1s | 37 tok/s | 3/5 | 99.8% | |
| $0.02 | $0.05 | — | 16K | fp8 | 0.61s | 170 tok/s | 4/5 | 99.7% | |
| $0.05 | $0.08 | $0.025 | 131K | — | 0.66s | 507 tok/s | 4/5 | 99.9% | |
| $0.152 | $0.287 | — | 32K | fp8 | 2.7s | 10 tok/s | 3/5 | 98.5% | |
| CWCoreWeave | $0.22 | $0.22 | $0.22 | 128K | bf16 | 0.42s | 147 tok/s | 4/5 | 100.0% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · SMOKE: of 5 programmatic capability probes passed (JSON schema, instruction following, tool call, long-context, benign compliance; retry-once) · screened via OpenRouter routing · v0.2 · methodology v1.1 · measured 44 hours ago
About
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Specifications
Context window131,072
Max output131,072
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJul 23, 2024
Knowledge cutoff2023-12-31
TokenizerLlama3
WeightsHugging Face ↗
Open weightsGATED
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/meta-llama/llama-3.1-8b-instruct)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks44 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.