API pricing per million tokens
- Input
- $0.43
- per 1M, Venice (fp4)
- Output
- $1.75
- per 1M tokens
- Cached input
- $0.08
- no batch tier listed
- Context
- 205K
- 16K max output
Where to buy it (4)
Venicefp4, fp4$0.43 / $1.75cached $0.08
DeepInfrafp4, fp4$0.5 / $2cached $0.1
Novita AIbf16, bf16$0.55 / $2.2cached $0.11
Z.aifp4, fp4$0.6 / $2.2cached $0.11
Through a gateway
Your own numbers100M input and 20M output tokens a month: $78.00 direct.
Helicone$78.00+0%
Kong AI Gateway$78.00+0%
LiteLLM$78.00+0%
Vercel AI Gateway$78.00+0%
AWS Bedrock$78.00+0%
Azure AI Foundry$78.00+0%
Google Vertex AI$78.00+0%
Databricks Mosaic AI$78.00+0%
Hugging Face Inference$78.00+0%
Cloudflare AI Gateway$81.90+5%
Requesty$81.90+5%
Eden AI$82.29+5.5%
OpenRouter$82.29+5.5%
Snowflake Cortex AI$93.60+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
GLM 5.3 PrimeZ.ai · 1M$2.8 / $8.8in / out per 1M$2.8$8.8$0.561MSep 23, 2026
GLM 5.3 FlashXZ.ai · 1M$0.37 / $1.25in / out per 1M$0.37$1.25$0.091MSep 18, 2026
GLM 5.3 FlashZ.ai · via InferenceNet (fp4) · 1.31M$0.045 / $0.14in / out per 1M$0.045$0.14$0.011.31MAug 26, 2026
GLM 5.3Z.ai · via DeepInfra (fp4) · 1.31M$0.563 / $2.5in / out per 1M$0.563$2.5$0.1251.31MAug 18, 2026
Gemini 3.8 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MSep 2, 2026
Gemini 3.7 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MAug 13, 2026
Gemini 3.6 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MJul 21, 2026
Kimi K2.5Moonshot AI · via SiliconFlow (int4) · 262K$0.45 / $2.25in / out per 1M$0.45$2.25$0.07262KJan 27, 2026
About GLM 4.6
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
Takes text in and returns text. First listed Sep 23, 2026. Blended price $0.76 per 1M tokens at 3 input to 1 output. Overall score 5 on leadersboard.fru.dev (rank 30).
Questions
How much does GLM 4.6 cost?
$0.43 per million input tokens and $1.75 per million output tokens, with cached input at $0.08.
What is the context window of GLM 4.6?
205K tokens, with up to 16K tokens of output.
Where is GLM 4.6 cheapest?
Of 4 routes tracked, Venice is cheapest at $0.43 input and $1.75 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.