DeepSeek V4.1 Flash vs GLM 5.3 Flash
API pricing, context, and independently measured performance, side by side.
DeepSeek V4.1 Flash
DeepSeek · released Sep 10, 2026
Best input$0.02 / MTok
Best output$0.39 / MTok
Context1.05M
Providers33
Fastest measured0.29s° TTFT · 210 tok/s
Open weightsVisionToolsReasoningStructured outputPrompt caching
GLM 5.3 Flash
Z.ai · released Aug 26, 2026
Best input$0.02 / MTok
Best output$0.10 / MTok
Context1.31M
Providers37
Fastest measured0.49s° TTFT · 184 tok/s
Open weightsVisionToolsReasoningStructured outputPrompt caching
Cheapest providers, measuredUSD / MTOK
| MODEL | PROVIDER | INPUT | OUTPUT | TTFT | TPS | SMOKE |
|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.02 | $0.60 | 1.2s° | 23° tok/s | — | |
| DeepSeek V4.1 Flash | OPOpenInference | $0.03 | $0.50 | 5s° | 10° tok/s | — |
| DeepSeek V4.1 Flash | $0.06 | $0.39 | — | — | — | |
| GLM 5.3 Flash | OPOpenInference | $0.02 | $0.30 | 2.6s° | 47° tok/s | — |
| GLM 5.3 Flash | $0.025 | $0.50 | 0.93s° | 34° tok/s | — | |
| GLM 5.3 Flash | $0.03 | $0.10 | 2.6s | 70 tok/s | — |
TTFT/TPS: ModelIndex-measured medians (P50, trailing 21d) · chat10k profile (~10K in / 200 out) · ° = measured via OpenRouter routing (adds a hop) · SMOKE: of 5 programmatic capability probes passed · methodology v1.1 — full details on each model page