Llama 3.3 70B Instruct
Meta · Open weights · Released Dec 6, 2024 · meta-llama/llama-3.3-70b-instruct
Best $0.10 in / $0.32 out via DeepInfra · fp8
131K contextToolsStructured outputPrompt cachingOpen weights
Available from 11 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 2 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | SMOKE | UPTIME |
|---|---|---|---|---|---|---|---|---|---|
| $0.10 | $0.32 | — | 131K | fp8 | 1.7s | 39 tok/s | 5/5 | 99.9% | |
| $0.135 | $0.40 | — | 12K | bf16 | ✕ | ✕ | 5/5 | 97.2% | |
| $0.20 | $0.52 | $0.10 | 131K | fp8 | 2.9s | 46 tok/s | 5/5 | 99.6% | |
| $0.22 | $0.50 | $0.11 | 131K | fp8 | ✕ | ✕ | 4/5 | 99.9% | |
| $0.293 | $2.25 | — | 24K | fp8 | 4.6s | 39 tok/s | 4/5 | 99.9% | |
| $0.45 | $0.90 | — | 131K | — | ✕ | ✕ | — | 99.9% | |
| $0.59 | $0.79 | $0.295 | 131K | — | 0.81s | 250 tok/s | — | 100.0% | |
| $0.71 | $0.71 | $0.71 | 128K | fp16 | 0.61s | 83 tok/s | — | 99.8% | |
| $0.72 | $0.72 | — | 128K | — | ✕ | ✕ | — | — | |
| $0.72 | $0.72 | — | 128K | — | — | — | — | — | |
| $1.04 | $1.04 | — | 131K | — | 1.1s | 33 tok/s | — | 98.0% |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · SMOKE: of 5 programmatic capability probes passed (JSON schema, instruction following, tool call, long-context, benign compliance; retry-once) · screened via OpenRouter routing · v0.2 · methodology v1.1 · measured 2 minutes ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Recent changesRSS ↗
Oct 5, 2026endpoint removed · primeintellect12 → 11
Oct 4, 2026endpoint added · primeintellect11 → 12
Sep 15, 2026endpoint removed · crusoe12 → 11
Sep 11, 2026endpoint removed · nebius13 → 12
About
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Specifications
Context window131,072
Max output16,384
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedDec 6, 2024
Knowledge cutoff2023-12-31
TokenizerLlama3
WeightsHugging Face ↗
Open weightsGATED
safetensorspytorch
Related
Embed badge
[](https://modelindex.ai/models/meta-llama/llama-3.3-70b-instruct)Sources
OpenRouter API23 minutes ago
AGGREGATORHugging Face23 minutes ago
COMMUNITYModelIndex benchmarks3 days ago
MEASUREDVerified 23 minutes ago · list prices, not negotiated rates.