API pricing per million tokens
- Input
- $0.39
- per 1M, list
- Output
- $2.34
- per 1M tokens
- Cached input
- n/a
- no batch tier listed
- Context
- 262K
- 236K max output
Where to buy it (10)
Qwen (Alibaba) (list)the lab's own API$0.39 / $2.34$0.878 blended
DeepInfrafp8, fp8+24%$0.45 / $3cached $0.22
Parasailfp8, fp8+45%$0.5 / $3.6cached $0.3
- AtlasCloudfp8, fp8+47%$0.55 / $3.5cached $0.55
DigitalOceandefault region+47%$0.55 / $3.5cached $0.11
- Phaladefault region+47%$0.55 / $3.5cached $0.225
- GMICloudfp8, fp8+54%$0.6 / $3.6$1.35 blended
Novita AIdefault region+54%$0.6 / $3.6$1.35 blended
- StreamLakedefault region+54%$0.6 / $3.6cached $0.12
Venicedefault region+92%$0.75 / $4.5$1.69 blended
Through a gateway
Your own numbers100M input and 20M output tokens a month: $85.80 direct.
Helicone$85.80+0%
Kong AI Gateway$85.80+0%
LiteLLM$85.80+0%
Vercel AI Gateway$85.80+0%
AWS Bedrock$85.80+0%
Azure AI Foundry$85.80+0%
Google Vertex AI$85.80+0%
Databricks Mosaic AI$85.80+0%
Hugging Face Inference$85.80+0%
Cloudflare AI Gateway$90.09+5%
Requesty$90.09+5%
Eden AI$90.52+5.5%
OpenRouter$90.52+5.5%
Snowflake Cortex AI$102.96+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
Qwen3.8 Max PrimeQwen (Alibaba) · 1M$4 / $12in / out per 1M$4$12$0.51MSep 23, 2026
Qwen3.8 Omni FlashQwen (Alibaba) · 1M$0.15 / $0.47in / out per 1M$0.15$0.47$0.0161MSep 21, 2026
Qwen3.8 Max (0902)Qwen (Alibaba) · 1M$2 / $6in / out per 1M$2$6$0.251MSep 3, 2026
Qwen3.8 FlashQwen (Alibaba) · 1M$0.15 / $0.47in / out per 1M$0.15$0.47$0.0161MAug 26, 2026
Gemini 3.8 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MSep 2, 2026
GLM 5.3Z.ai · via DeepInfra (fp4) · 1.31M$0.563 / $2.5in / out per 1M$0.563$2.5$0.1251.31MAug 18, 2026
Gemini 3.7 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MAug 13, 2026
GLM 4.6Z.ai · via Venice (fp4) · 205K$0.43 / $1.75in / out per 1M$0.43$1.75$0.08205KSep 30, 2025
About Qwen3.5 397B A17B
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.
Takes text, image, video in and returns text. First listed Sep 23, 2026. Blended price $0.878 per 1M tokens at 3 input to 1 output.
Questions
How much does Qwen3.5 397B A17B cost?
$0.39 per million input tokens and $2.34 per million output tokens.
What is the context window of Qwen3.5 397B A17B?
262K tokens, with up to 236K tokens of output.
Where is Qwen3.5 397B A17B cheapest?
Of 10 routes tracked, Qwen (Alibaba) is cheapest at $0.39 input and $2.34 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.