API pricing per million tokens
- Input
- $0.4
- per 1M, DeepInfra (fp4)
- Output
- $1.75
- per 1M tokens
- Cached input
- $0.08
- no batch tier listed
- Context
- 205K
- 131K max output
Where to buy it (7)
DeepInfrafp4, fp4$0.4 / $1.75cached $0.08
Venicefp4, fp4$0.4 / $1.93cached $0.08
- AtlasCloudfp8, fp8$0.52 / $1.85cached $0.12
Novita AIfp8, fp8$0.54 / $1.98cached $0.099
Google Vertex AIdefault region$0.6 / $2.2$1 blended
Z.aifp4, fp4$0.6 / $2.2cached $0.11
- Mancer 2fp4, fp4$0.7 / $2.5$1.15 blended
Through a gateway
Your own numbers100M input and 20M output tokens a month: $75.00 direct.
Helicone$75.00+0%
Kong AI Gateway$75.00+0%
LiteLLM$75.00+0%
Vercel AI Gateway$75.00+0%
AWS Bedrock$75.00+0%
Azure AI Foundry$75.00+0%
Google Vertex AI$75.00+0%
Databricks Mosaic AI$75.00+0%
Hugging Face Inference$75.00+0%
Cloudflare AI Gateway$78.75+5%
Requesty$78.75+5%
Eden AI$79.13+5.5%
OpenRouter$79.13+5.5%
Snowflake Cortex AI$90.00+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
GLM 5.3 PrimeZ.ai · 1M$2.8 / $8.8in / out per 1M$2.8$8.8$0.561MSep 23, 2026
GLM 5.3 FlashXZ.ai · 1M$0.37 / $1.25in / out per 1M$0.37$1.25$0.091MSep 18, 2026
GLM 5.3 FlashZ.ai · via InferenceNet (fp4) · 1.31M$0.045 / $0.14in / out per 1M$0.045$0.14$0.011.31MAug 26, 2026
GLM 5.3Z.ai · via DeepInfra (fp4) · 1.31M$0.563 / $2.5in / out per 1M$0.563$2.5$0.1251.31MAug 18, 2026
Gemini 3.8 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MSep 2, 2026
Gemini 3.7 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MAug 13, 2026
Gemini 3.6 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MJul 21, 2026
Kimi K2.5Moonshot AI · via SiliconFlow (int4) · 262K$0.45 / $2.25in / out per 1M$0.45$2.25$0.07262KJan 27, 2026
About GLM 4.7
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution.
Takes text in and returns text. First listed Sep 23, 2026. Blended price $0.738 per 1M tokens at 3 input to 1 output.
Questions
How much does GLM 4.7 cost?
$0.4 per million input tokens and $1.75 per million output tokens, with cached input at $0.08.
What is the context window of GLM 4.7?
205K tokens, with up to 131K tokens of output.
Where is GLM 4.7 cheapest?
Of 7 routes tracked, DeepInfra is cheapest at $0.4 input and $1.75 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.