Qwen3 32B
Qwen · Open weights · Released Apr 28, 2025 · qwen/qwen3-32b
131K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 4 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.08 | $0.28 | — | 41K | fp8 | 2.4s | 25 tok/s | 99.9% | |
| $0.10 | $0.30 | — | 41K | fp8 | 1.6s | 23 tok/s | 99.8% | |
| $0.14 | $0.57 | — | 131K | fp8 | 5.8s | 19 tok/s | 91.6% | |
| $0.29 | $0.59 | $0.145 | 131K | — | 0.92s | 351 tok/s | 100.0% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 43 hours ago
About
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Specifications
Context window131,072
Max output16,384
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedApr 28, 2025
Knowledge cutoff2025-03-31
TokenizerQwen3
WeightsHugging Face ↗
Related
Embed badge
[](https://modelindex.ai/models/qwen/qwen3-32b)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks43 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.