API pricing per million tokens
- Input
- $0.563
- per 1M, DeepInfra (fp4)
- Output
- $1.8
- per 1M tokens
- Cached input
- $0.105
- no batch tier listed
- Context
- 1M
- 131K max output
Where to buy it (27)
DeepInfrafp4, fp4$0.563 / $1.8cached $0.105
- StreamLakefp8, fp8$0.643 / $2.02cached $0.119
Novita AIfp8, fp8$0.65 / $2.04cached $0.121
DigitalOceandefault region$0.7 / $2.2cached $0.105
CoreWeavefp4, fp4$0.76 / $2.42cached $0.14
- AtlasCloudfp8, fp8$0.938 / $2.95cached $0.174
Qwen (Alibaba)fp8, fp8$0.966 / $3.04cached $0.193
- Inceptronfp4, fp4$1.01 / $3.24cached $0.194
- Phalafp8, fp8$1.26 / $3cached $0.22
SiliconFlowfp8, fp8$1.19 / $3.74cached $0.221
- Baidufp8, fp8$1.4 / $4.4cached $0.26
Basetenfp8, fp8$1.4 / $4.4cached $0.14
Cloudflare Workers AIdefault region$1.4 / $4.4cached $0.26
Fireworks AIdefault region$1.4 / $4.4cached $0.14
Friendlidefault region$1.4 / $4.4cached $0.26
- GMICloudfp8, fp8$1.4 / $4.4cached $0.26
Parasailfp4, fp4$1.4 / $4.4cached $0.26
Together AIdefault region$1.4 / $4.4cached $0.26
Venicefp8, fp8$1.4 / $4.4cached $0.26
- Waferdefault region$1.4 / $4.4cached $0.26
Z.aifp8, fp8$1.4 / $4.4cached $0.26
Basetenfast, fp8$2.1 / $6.6cached $0.21
Fireworks AIfast$2.1 / $6.6cached $0.21
Fireworks AIfast-us$2.1 / $6.6cached $0.21
Qwen (Alibaba)fast, fp8$2.31 / $7.26cached $0.462
- Baidufast, fp4$2.25 / $7.88cached $0.56
- Decartfast, fp4$2.25 / $8cached $0.48
Through a gateway
Your own numbers100M input and 20M output tokens a month: $92.25 direct.
Helicone$92.25+0%
Kong AI Gateway$92.25+0%
LiteLLM$92.25+0%
Vercel AI Gateway$92.25+0%
AWS Bedrock$92.25+0%
Azure AI Foundry$92.25+0%
Google Vertex AI$92.25+0%
Databricks Mosaic AI$92.25+0%
Hugging Face Inference$92.25+0%
Cloudflare AI Gateway$96.86+5%
Requesty$96.86+5%
Eden AI$97.32+5.5%
OpenRouter$97.32+5.5%
Snowflake Cortex AI$110.70+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
GLM 5.3 PrimeZ.ai · 1M$2.8 / $8.8in / out per 1M$2.8$8.8$0.561MSep 23, 2026
GLM 5.3 FlashXZ.ai · 1M$0.37 / $1.25in / out per 1M$0.37$1.25$0.091MSep 18, 2026
GLM 5.3 FlashZ.ai · via InferenceNet (fp4) · 1.31M$0.045 / $0.14in / out per 1M$0.045$0.14$0.011.31MAug 26, 2026
GLM 5.3Z.ai · via DeepInfra (fp4) · 1.31M$0.563 / $2.5in / out per 1M$0.563$2.5$0.1251.31MAug 18, 2026
Gemini 3.8 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MSep 2, 2026
Gemini 3.7 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MAug 13, 2026
Gemini 3.6 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MJul 21, 2026
Kimi K2.5Moonshot AI · via SiliconFlow (int4) · 262K$0.45 / $2.25in / out per 1M$0.45$2.25$0.07262KJan 27, 2026
About GLM 5.2
GLM 5.2 is a large-scale reasoning model from Z.ai.
Takes text in and returns text. First listed Sep 23, 2026. Blended price $0.872 per 1M tokens at 3 input to 1 output.
Questions
How much does GLM 5.2 cost?
$0.563 per million input tokens and $1.8 per million output tokens, with cached input at $0.105.
What is the context window of GLM 5.2?
1M tokens, with up to 131K tokens of output.
Where is GLM 5.2 cheapest?
Of 27 routes tracked, DeepInfra is cheapest at $0.563 input and $1.8 output per million tokens.
Has GLM 5.2's price changed?
Yes. The output price went from $2.27 to $7.88, first seen Sep 24, 2026.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.