Llama 3.3 70B Instruct vs Qwen3.8 Max Prime
API pricing, context, and independently measured performance, side by side. Llama 3.3 70B Instruct starts 98% cheaper on input than Qwen3.8 Max Prime at list prices.
Llama 3.3 70B Instruct
Meta · released Dec 6, 2024
Best input$0.10 / MTok
Best output$0.32 / MTok
Context131K
Providers11
Fastest measured0.57s° TTFT · 81 tok/s
Open weightsToolsStructured outputPrompt caching
Qwen3.8 Max Prime
Qwen · released Sep 23, 2026
Best input$4.00 / MTok
Best output$12.00 / MTok
Context1M
Providers1
Fastest measured—
VisionToolsReasoningStructured outputPrompt caching
Cheapest providers, measuredUSD / MTOK
| MODEL | PROVIDER | INPUT | OUTPUT | TTFT | TPS | SMOKE |
|---|---|---|---|---|---|---|
| Llama 3.3 70B Instruct | $0.10 | $0.32 | 1.7s° | 39° tok/s | 5/5 | |
| Llama 3.3 70B Instruct | $0.135 | $0.40 | 1.5s° | 48° tok/s | 5/5 | |
| Llama 3.3 70B Instruct | $0.20 | $0.52 | 2.9s° | 46° tok/s | 5/5 | |
| Qwen3.8 Max Prime | $4.00 | $12.00 | — | — | — |
TTFT/TPS: ModelIndex-measured medians (P50, trailing 21d) · chat10k profile (~10K in / 200 out) · ° = measured via OpenRouter routing (adds a hop) · SMOKE: of 5 programmatic capability probes passed · methodology v1.1 — full details on each model page