API pricing per million tokens
- Input
- $0.3
- per 1M, Novita AI (bf16)
- Output
- $0.9
- per 1M tokens
- Cached input
- $0.055
- no batch tier listed
- Context
- 131K
- 33K max output
Where to buy it (2)
Novita AIbf16, bf16$0.3 / $0.9cached $0.055
Z.aifp8, fp8$0.3 / $0.9cached $0.05
Through a gateway
Your own numbers100M input and 20M output tokens a month: $48.00 direct.
Helicone$48.00+0%
Kong AI Gateway$48.00+0%
LiteLLM$48.00+0%
Vercel AI Gateway$48.00+0%
AWS Bedrock$48.00+0%
Azure AI Foundry$48.00+0%
Google Vertex AI$48.00+0%
Databricks Mosaic AI$48.00+0%
Hugging Face Inference$48.00+0%
Cloudflare AI Gateway$50.40+5%
Requesty$50.40+5%
Eden AI$50.64+5.5%
OpenRouter$50.64+5.5%
Snowflake Cortex AI$57.60+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
GLM 5.3 PrimeZ.ai · 1M$2.8 / $8.8in / out per 1M$2.8$8.8$0.561MSep 23, 2026
GLM 5.3 FlashXZ.ai · 1M$0.37 / $1.25in / out per 1M$0.37$1.25$0.091MSep 18, 2026
GLM 5.3 FlashZ.ai · via InferenceNet (fp4) · 1.31M$0.045 / $0.14in / out per 1M$0.045$0.14$0.011.31MAug 26, 2026
GLM 5.3Z.ai · via DeepInfra (fp4) · 1.31M$0.563 / $2.5in / out per 1M$0.563$2.5$0.1251.31MAug 18, 2026
DeepSeek V4.1 FlashDeepSeek · 1M$0.15 / $0.6in / out per 1M$0.15$0.6$0.0031MSep 10, 2026
Kimi K2.5Moonshot AI · via SiliconFlow (int4) · 262K$0.45 / $2.25in / out per 1M$0.45$2.25$0.07262KJan 27, 2026
Gemini 3.1 Flash LiteGoogle Gemini API · 1M$0.25 / $1.5in / out per 1M$0.25$1.5$0.0251MMay 7, 2026
Gemini 3.1 Flash Lite PreviewGoogle Gemini API · 1M$0.25 / $1.5in / out per 1M$0.25$1.5$0.0251MMar 3, 2026
About GLM 4.6V
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media.
Takes image, text, video in and returns text. First listed Sep 23, 2026. Blended price $0.45 per 1M tokens at 3 input to 1 output.
Questions
How much does GLM 4.6V cost?
$0.3 per million input tokens and $0.9 per million output tokens, with cached input at $0.055.
What is the context window of GLM 4.6V?
131K tokens, with up to 33K tokens of output.
Where is GLM 4.6V cheapest?
Of 2 routes tracked, Novita AI is cheapest at $0.3 input and $0.9 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.