API pricing per million tokens
glm-4-7-flash Open weightsChecked 7h ago
- Input
- $0.06
- per 1M, Venice (fp8)
- Output
- $0.4
- per 1M tokens
- Cached input
- $0.01
- no batch tier listed
- Context
- 200K
- 118K max output
Where to buy it (3)
Venicefp8, fp8$0.06 / $0.4cached $0.01
Cloudflare Workers AIdefault region$0.06 / $0.4$0.145 blended
Novita AIbf16, bf16$0.07 / $0.4cached $0.01
Through a gateway
Your own numbers100M input and 20M output tokens a month: $14.00 direct.
Helicone$14.00+0%
Kong AI Gateway$14.00+0%
LiteLLM$14.00+0%
Vercel AI Gateway$14.00+0%
AWS Bedrock$14.00+0%
Azure AI Foundry$14.00+0%
Google Vertex AI$14.00+0%
Databricks Mosaic AI$14.00+0%
Hugging Face Inference$14.00+0%
Cloudflare AI Gateway$14.70+5%
Requesty$14.70+5%
Eden AI$14.77+5.5%
OpenRouter$14.77+5.5%
Snowflake Cortex AI$16.80+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
ModelInputOutputCachedContextReleased
GLM 5.3 PrimeZ.ai · 1M$2.8 / $8.8in / out per 1M$2.8$8.8$0.561MSep 23, 2026
GLM 5.3 FlashXZ.ai · 1M$0.37 / $1.25in / out per 1M$0.37$1.25$0.091MSep 18, 2026
GLM 5.3 FlashZ.ai · via InferenceNet (fp4) · 1.31M$0.045 / $0.14in / out per 1M$0.045$0.14$0.011.31MAug 26, 2026
GLM 5.3Z.ai · via DeepInfra (fp4) · 1.31M$0.563 / $2.5in / out per 1M$0.563$2.5$0.1251.31MAug 18, 2026
DeepSeek V4.1 FlashDeepSeek · 1M$0.15 / $0.6in / out per 1M$0.15$0.6$0.0031MSep 10, 2026
GPT-6 LunaOpenAI · 1.05M$0.1 / $0.5in / out per 1M$0.1$0.5$0.011.05MSep 22, 2026
GPT-6 Luna ProOpenAI · 1.05M$0.1 / $0.5in / out per 1M$0.1$0.5$0.011.05MSep 22, 2026
Qwen3.8 Omni FlashQwen (Alibaba) · 1M$0.15 / $0.47in / out per 1M$0.15$0.47$0.0161MSep 21, 2026
About GLM 4.7 Flash
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency.
Takes text in and returns text. First listed Sep 23, 2026. Blended price $0.145 per 1M tokens at 3 input to 1 output.
Across fru.devCompany profile
Questions
How much does GLM 4.7 Flash cost?
$0.06 per million input tokens and $0.4 per million output tokens, with cached input at $0.01.
What is the context window of GLM 4.7 Flash?
200K tokens, with up to 118K tokens of output.
Where is GLM 4.7 Flash cheapest?
Of 3 routes tracked, Venice is cheapest at $0.06 input and $0.4 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.