GLM 4.7 Flash
Z.ai · Open weights · Released Jan 19, 2026 · z-ai/glm-4.7-flash
203K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 4 providersUSD / MTOK · VERIFIED 4 MINUTES AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.06 | $0.40 | $0.01 | 203K | bf16 | 0.55s | 98 tok/s | 99.8% | |
| $0.06 | $0.40 | $0.01 | 128K | fp8 | 0.99s | 39 tok/s | 99.7% | |
| $0.06 | $0.40 | — | 131K | — | 0.85s | 52 tok/s | 95.0% | |
| $0.07 | $0.40 | $0.01 | 200K | bf16 | 1.9s | 102 tok/s | 75.5% |
TTFT/TPS: median (P50) over trailing 72h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 43 hours ago
Pricing variants · USD / MTok in / out
STANDARD
$0.06 / $0.40
per MTok in / out
CACHE READ $0.01
About
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Specifications
Context window202,752
Max output16,384
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJan 19, 2026
TokenizerOther
WeightsHugging Face ↗
Embed badge
[](https://modelindex.ai/models/z-ai/glm-4.7-flash)Sources
OpenRouter API4 minutes ago
AGGREGATORHugging Face4 minutes ago
COMMUNITYModelIndex benchmarks43 hours ago
MEASUREDVerified 4 minutes ago · list prices, not negotiated rates.