← model directory

z-ai

Z.ai: GLM 4.7 Flash

z-ai/glm-4.7-flash

↓ runs free on your own hardware⚙ tool calling

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

specs & pricing
Type
text
Provider
z-ai
Model ID
z-ai/glm-4.7-flash
Capabilities
tools, reasoning
Context window
203K tokens
Self-hostable
Yes — runs on your own GPU
Input price
$0.066 / 1M tokens
Output price
$0.44 / 1M tokens

Cloud price is billed from prepaid credits when a request fails over to the cloud. Open-weight models run free on GPUs you own — the gateway routes to your nodes first.