Models/Qwen/Qwen3 32B

Qwen3 32B

Qwen · Open weights · Released Apr 28, 2025 · qwen/qwen3-32b
Best $0.08 in / $0.28 out via DeepInfra · fp8 · 41K ctx at this price (headline 131K)
CompareEstimate cost
131K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 3 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 2 DAYS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
PROVIDERINPUTOUTPUTCACHECONTEXTQUANTTTFTTPSUPTIME
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · methodology v1.1 · measured 2 days ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Input price history · best listed price, observed
↑ 14% SINCE OCT 4, 2026
$0.07 · OCT 4, 2026$0.08 · TODAY
Recent changesRSS ↗
Oct 4, 2026price increased · input0.07 → 0.08
Oct 4, 2026price increased · output0.2 → 0.28
Oct 3, 2026endpoint added · tenstorrent2 → 3
Oct 3, 2026provider added · tenstorrent
Sep 4, 2026endpoint removed · nebius3 → 2
Sep 1, 2026endpoint removed · groq4 → 3
About

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

Specifications
Context window131,072
Max output16,384
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedApr 28, 2025
Knowledge cutoff2025-03-31
TokenizerQwen3
Open weights
Hugging FaceQwen/Qwen3-32B ↗
Downloads3.7M
Likes754
Licenseapache-2.0
safetensors
Embed badge
Qwen3 32B price and speed badge
[![Qwen3 32B](https://modelindex.ai/badge/qwen/qwen3-32b)](https://modelindex.ai/models/qwen/qwen3-32b)
Sources
OpenRouter API23 minutes ago
AGGREGATOR
Hugging Face23 minutes ago
COMMUNITY
ModelIndex benchmarks6 days ago
MEASURED
Verified 23 minutes ago · list prices, not negotiated rates.