GLM 4.7 Flash
Z.ai · Open weights · Released Jan 19, 2026 · z-ai/glm-4.7-flash
Best $0.06 in / $0.40 out via Venice · fp8 · 128K ctx at this price (headline 200K)
200K contextToolsReasoningStructured outputPrompt cachingOpen weights
Available from 3 providersUSD / MTOK · PRICES 23 MINUTES AGO · SPEED 24 HOURS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.06 | $0.40 | $0.01 | 128K | fp8 | 1.5s | 34 tok/s | 97.1% | |
| $0.06 | $0.40 | — | 131K | — | ✕ | ✕ | 99.9% | |
| $0.07 | $0.40 | $0.01 | 200K | bf16 | ✕ | ✕ | 11.4% |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · ✕ = probed, endpoint returned no measurable stream · methodology v1.1 · measured 24 hours ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Pricing variants · USD / MTok in / out
STANDARD
$0.06 / $0.40
per MTok in / out
CACHE READ $0.01
Recent changesRSS ↗
Sep 10, 2026endpoint removed · deepinfra4 → 3
Sep 10, 2026context changed202752 → 200000
About
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Specifications
Context window200,000
Max output117,964
Modalities intext
Modalities outtext
LicenseOpen weights
ReleasedJan 19, 2026
TokenizerOther
WeightsHugging Face ↗
Related
Embed badge
[](https://modelindex.ai/models/z-ai/glm-4.7-flash)Sources
OpenRouter API23 minutes ago
AGGREGATORHugging Face23 minutes ago
COMMUNITYModelIndex benchmarks19 days ago
MEASUREDVerified 23 minutes ago · list prices, not negotiated rates.