API pricing per million tokens
- Input
- $0.045
- per 1M, InferenceNet (fp4)
- Output
- $0.14
- per 1M tokens
- Cached input
- $0.01
- no batch tier listed
- Context
- 1.31M
- 944K max output
Where to buy it (30)
- InferenceNetfp4, fp4$0.045 / $0.14cached $0.01
DeepInfrafp4, fp4$0.075 / $0.25cached $0.015
- Relacedefault region$0.07 / $0.28cached $0.02
- Inceptronfp8, fp8$0.09 / $0.28cached $0.07
- GMICloudfp8, fp8$0.09 / $0.3cached $0.018
- Morphdefault region$0.088 / $0.308cached $0.018
- Waferdefault region$0.089 / $0.35cached $0.03
- Io Netfp8, fp8$0.125 / $0.42cached $0.025
- OpenInferencefp4, fp4$0.1 / $0.5cached $0.025
- Phalafp8, fp8$0.128 / $0.425cached $0.025
Novita AIfp8, fp8$0.132 / $0.44cached $0.026
- StreamLakefp8, fp8$0.141 / $0.47cached $0.028
- Sail Researchus, fp8$0.143 / $0.475cached $0.029
- AtlasCloudfp8, fp8$0.15 / $0.5cached $0.03
Basetenfp8, fp8$0.15 / $0.5cached $0.03
CoreWeavenvfp4, nvfp4$0.15 / $0.5cached $0.05
Crusoefp4, fp4$0.15 / $0.5cached $0.03
DigitalOceandefault region$0.15 / $0.5cached $0.03
Fireworks AIdefault region$0.15 / $0.5cached $0.03
Friendlidefault region$0.15 / $0.5cached $0.03
- Near AIfp8, fp8$0.15 / $0.5cached $0.035
Parasailfp8, fp8$0.15 / $0.5cached $0.03
- Rekafp8, fp8$0.15 / $0.5cached $0.03
SiliconFlowfp8, fp8$0.15 / $0.5cached $0.03
Together AIdefault region$0.15 / $0.5cached $0.03
Venicedefault region$0.15 / $0.5cached $0.03
Z.aifp8, fp8$0.15 / $0.5cached $0.03
- NextBitfp8, fp8$0.165 / $0.55cached $0.033
Cloudflare Workers AIdefault region$0.3 / $1cached $0.03
Modalfp8, fp8$0.45 / $1.5cached $0.09
Through a gateway
Your own numbers100M input and 20M output tokens a month: $7.30 direct.
Helicone$7.30+0%
Kong AI Gateway$7.30+0%
LiteLLM$7.30+0%
Vercel AI Gateway$7.30+0%
AWS Bedrock$7.30+0%
Azure AI Foundry$7.30+0%
Google Vertex AI$7.30+0%
Databricks Mosaic AI$7.30+0%
Hugging Face Inference$7.30+0%
Cloudflare AI Gateway$7.67+5%
Requesty$7.67+5%
Eden AI$7.70+5.5%
OpenRouter$7.70+5.5%
Snowflake Cortex AI$8.76+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
GLM 5.3 PrimeZ.ai · 1M$2.8 / $8.8in / out per 1M$2.8$8.8$0.561MSep 23, 2026
GLM 5.3 FlashXZ.ai · 1M$0.37 / $1.25in / out per 1M$0.37$1.25$0.091MSep 18, 2026
GLM 5.3Z.ai · via DeepInfra (fp4) · 1.31M$0.563 / $2.5in / out per 1M$0.563$2.5$0.1251.31MAug 18, 2026
GLM 5.2Z.ai · via DeepInfra (fp4) · 1M$0.563 / $1.8in / out per 1M$0.563$1.8$0.1051MJun 16, 2026
Muse Spark 1.3 ContributorMeta · 1M$0.1 / $0.2in / out per 1M$0.1$0.2$0.0021MSep 2, 2026
Muse Spark 1.2 ContributorMeta · 1M$0.1 / $0.2in / out per 1M$0.1$0.2$0.0021MAug 21, 2026
DeepSeek V4 Flash 0731DeepSeek · via StreamLake (fp8) · 1.31M$0.053 / $0.158in / out per 1M$0.053$0.158$0.00171.31MJul 31, 2026
Qwen3.7 FlashQwen (Alibaba) · 1M$0.03 / $0.13in / out per 1M$0.03$0.13$0.0061MJul 27, 2026
About GLM 5.3 Flash
GLM-5.3-Flash is a native multimodal model from Z.ai.
Takes text, image, video in and returns text. First listed Sep 23, 2026. Blended price $0.069 per 1M tokens at 3 input to 1 output.
Questions
How much does GLM 5.3 Flash cost?
$0.045 per million input tokens and $0.14 per million output tokens, with cached input at $0.01.
What is the context window of GLM 5.3 Flash?
1.31M tokens, with up to 944K tokens of output.
Where is GLM 5.3 Flash cheapest?
Of 30 routes tracked, InferenceNet is cheapest at $0.045 input and $0.14 output per million tokens.
Has GLM 5.3 Flash's price changed?
Yes. The output price went from $0.5 to $0.36, first seen Sep 24, 2026.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.