API pricing per million tokens
- Input
- $0.563
- per 1M, DeepInfra (fp4)
- Output
- $2.5
- per 1M tokens
- Cached input
- $0.125
- no batch tier listed
- Context
- 1.31M
- 131K max output
Where to buy it (32)
DeepInfrafp4, fp4$0.563 / $2.5cached $0.125
- InferenceNetfp4, fp4$0.684 / $2.28cached $0.114
- Morphdefault region$0.774 / $2.43cached $0.127
Novita AIfp8, fp8$0.783 / $2.46cached $0.145
- Rekafp8, fp8$0.761 / $2.57cached $0.152
- Io Netfp8, fp8$0.767 / $2.6cached $0.153
- Phaladefault region$0.84 / $2.64cached $0.156
DigitalOceandefault region$0.91 / $2.86cached $0.169
- GMICloudfp8, fp8$0.98 / $3.08cached $0.182
- Inceptronfp4, fp4$0.901 / $3.53cached $0.23
- Sail Researchfp8, fp8$0.77 / $4cached $0.19
SiliconFlowfp8, fp8$1.12 / $3.52cached $0.208
- Decartfp4, fp4$1.19 / $3.74cached $0.196
Qwen (Alibaba)default region$1.19 / $3.74cached $0.238
- Sail Researchus, fp8$1.21 / $3.87cached $0.225
Friendlidefault region$1.26 / $3.96cached $0.234
- AkashMLfp8, fp8$1.3 / $4.4cached $0.26
- AtlasCloudfp8, fp8$1.4 / $4.4cached $0.26
- Baidufp8, fp8$1.4 / $4.4cached $0.26
Basetenfp4, fp4$1.4 / $4.4cached $0.14
Cloudflare Workers AIdefault region$1.4 / $4.4cached $0.26
Crusoefp4, fp4$1.4 / $4.4cached $0.26
Fireworks AIdefault region$1.4 / $4.4cached $0.26
Mistral AInvfp4, nvfp4$1.4 / $4.4cached $0.14
Modaldefault region$1.4 / $4.4cached $0.26
Parasailfp8, fp8$1.4 / $4.4cached $0.26
- PrimeIntellectdefault region$1.4 / $4.4cached $0.26
Together AIdefault region$1.4 / $4.4cached $0.26
Venicedefault region$1.4 / $4.4cached $0.26
- Waferdefault region$1.4 / $4.4cached $0.26
Z.aifp8, fp8$1.4 / $4.4cached $0.26
Basetenfast, fp8$2.1 / $6.6cached $0.21
Through a gateway
Your own numbers100M input and 20M output tokens a month: $106.25 direct.
Helicone$106.25+0%
Kong AI Gateway$106.25+0%
LiteLLM$106.25+0%
Vercel AI Gateway$106.25+0%
AWS Bedrock$106.25+0%
Azure AI Foundry$106.25+0%
Google Vertex AI$106.25+0%
Databricks Mosaic AI$106.25+0%
Hugging Face Inference$106.25+0%
Cloudflare AI Gateway$111.56+5%
Requesty$111.56+5%
Eden AI$112.09+5.5%
OpenRouter$112.09+5.5%
Snowflake Cortex AI$127.50+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
GLM 5.3 PrimeZ.ai ยท 1M$2.8 / $8.8in / out per 1M$2.8$8.8$0.561MSep 23, 2026
GLM 5.3 FlashXZ.ai ยท 1M$0.37 / $1.25in / out per 1M$0.37$1.25$0.091MSep 18, 2026
GLM 5.3 FlashZ.ai ยท via InferenceNet (fp4) ยท 1.31M$0.045 / $0.14in / out per 1M$0.045$0.14$0.011.31MAug 26, 2026
GLM 5.2Z.ai ยท via DeepInfra (fp4) ยท 1M$0.563 / $1.8in / out per 1M$0.563$1.8$0.1051MJun 16, 2026
Muse Spark 1.3Meta ยท 1M$1.25 / $4.25in / out per 1M$1.25$4.25$0.151MSep 2, 2026
Gemini 3.8 FlashGoogle Gemini API ยท 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MSep 2, 2026
Muse Spark 1.1Meta ยท 1M$1.25 / $4.25in / out per 1M$1.25$4.25$0.151MJul 16, 2026
Gemini 3.7 FlashGoogle Gemini API ยท 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MAug 13, 2026
About GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks.
Takes text in and returns text. First listed Sep 23, 2026. Blended price $1.05 per 1M tokens at 3 input to 1 output. Overall score 8 on leadersboard.fru.dev (rank 20).
Questions
How much does GLM 5.3 cost?
$0.563 per million input tokens and $2.5 per million output tokens, with cached input at $0.125.
What is the context window of GLM 5.3?
1.31M tokens, with up to 131K tokens of output.
Where is GLM 5.3 cheapest?
Of 32 routes tracked, DeepInfra is cheapest at $0.563 input and $2.5 output per million tokens.
Has GLM 5.3's price changed?
Yes. The output price went from $1.76 to $4.4, first seen Sep 24, 2026.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.