DeepSeek V4 Flash 0731 vs GLM 5.2
API pricing, context, and independently measured performance, side by side. DeepSeek V4 Flash 0731 starts 92% cheaper on input than GLM 5.2 at list prices.
DeepSeek V4 Flash 0731
DeepSeek · released Jul 31, 2026
Best input$0.04 / MTok
Best output$0.08 / MTok
Context1.31M
Providers30
Fastest measured0.48s° TTFT · 65 tok/s
Open weightsToolsReasoningStructured outputPrompt caching
GLM 5.2
Z.ai · released Jun 16, 2026
Best input$0.49 / MTok
Best output$1.54 / MTok
Context1.05M
Providers33
Fastest measured0.54s° TTFT · 131 tok/s
Open weightsToolsReasoningStructured outputPrompt caching
Cheapest providers, measuredUSD / MTOK
| MODEL | PROVIDER | INPUT | OUTPUT | TTFT | TPS | SMOKE |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | RERelace | $0.04 | $0.08 | 0.49s° | 65° tok/s | 5/5 |
| DeepSeek V4 Flash 0731 | OPOpenInference | $0.04 | $0.13 | 14.9s° | 13° tok/s | 3/5 |
| DeepSeek V4 Flash 0731 | DEDecart | $0.064 | $0.128 | 0.65s° | 51° tok/s | 4/5 |
| GLM 5.2 | BABaidu | $0.49 | $1.54 | 0.82s° | 72° tok/s | 3/5 |
| GLM 5.2 | SRSail Research | $0.50 | $3.15 | 1.5s° | 38° tok/s | 3/5 |
| GLM 5.2 | AMAmbient | $0.60 | $2.00 | 1.5s° | 167° tok/s | 3/5 |
TTFT/TPS: ModelIndex-measured medians (P50, trailing 72h) · chat10k profile (~10K in / 200 out) · ° = measured via OpenRouter routing (adds a hop) · SMOKE: of 5 programmatic capability probes passed · methodology v1.1 — full details on each model page