Models/Llama 3.3 70B Instruct vs Qwen3.8 2.4T A95B

Llama 3.3 70B Instruct vs Qwen3.8 2.4T A95B

API pricing, context, and independently measured performance, side by side. Llama 3.3 70B Instruct starts 95% cheaper on input than Qwen3.8 2.4T A95B at list prices.

Llama 3.3 70B Instruct
Meta · released Dec 6, 2024
Best input$0.10 / MTok
Best output$0.32 / MTok
Context131K
Providers13
Fastest measured0.59s° TTFT · 86 tok/s
Open weightsToolsStructured outputPrompt caching
Qwen3.8 2.4T A95B
Qwen · released Aug 12, 2026
Best input$2.00 / MTok
Best output$6.00 / MTok
Context1.05M
Providers6
Fastest measured1s° TTFT · 122 tok/s
Open weightsToolsReasoningStructured outputPrompt caching
Cheapest providers, measuredUSD / MTOK
MODELPROVIDERINPUTOUTPUTTTFTTPSSMOKE
Llama 3.3 70B InstructDeepInfra logoDeepInfra$0.10$0.321.9s°12° tok/s5/5
Llama 3.3 70B InstructNebius logoNebius$0.13$0.4015s°3° tok/s5/5
Llama 3.3 70B InstructNovita logoNovita$0.135$0.401.8s°25° tok/s5/5
Qwen3.8 2.4T A95BDeepInfra logoDeepInfra$2.00$6.001.3s°76° tok/s4/5
Qwen3.8 2.4T A95BSiliconFlow logoSiliconFlow$2.00$6.002.7s°49° tok/s5/5
Qwen3.8 2.4T A95BMOModal$2.00$6.001s°122° tok/s4/5
TTFT/TPS: ModelIndex-measured medians (P50, trailing 72h) · chat10k profile (~10K in / 200 out) · ° = measured via OpenRouter routing (adds a hop) · SMOKE: of 5 programmatic capability probes passed · methodology v1.1 — full details on each model page
All 13 providers for Llama 3.3 70B InstructAll 6 providers for Qwen3.8 2.4T A95BOpen in comparator