API pricing per million tokens
- Input
- $2
- per 1M, list
- Output
- $6
- per 1M tokens
- Cached input
- $0.25
- no batch tier listed
- Context
- 1M
- 131K max output
Where to buy it (7)
DeepInfrafp4, fp4$2 / $6cached $0.2
Qwen (Alibaba) (list)the lab's own API$2 / $6cached $0.25
Modaldefault region$2 / $6cached $0.25
Novita AIdefault region$2 / $6cached $0.25
SiliconFlowfp8, fp8$2 / $6cached $0.25
Together AIdefault region$2 / $6cached $0.25
Venicedefault region$2 / $6cached $0.25
Through a gateway
Your own numbers100M input and 20M output tokens a month: $320.00 direct.
Helicone$320.00+0%
Kong AI Gateway$320.00+0%
LiteLLM$320.00+0%
Vercel AI Gateway$320.00+0%
AWS Bedrock$320.00+0%
Azure AI Foundry$320.00+0%
Google Vertex AI$320.00+0%
Databricks Mosaic AI$320.00+0%
Hugging Face Inference$320.00+0%
Cloudflare AI Gateway$336.00+5%
Requesty$336.00+5%
Eden AI$337.60+5.5%
OpenRouter$337.60+5.5%
Snowflake Cortex AI$384.00+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
Qwen3.8 Max PrimeQwen (Alibaba) · 1M$4 / $12in / out per 1M$4$12$0.51MSep 23, 2026
Qwen3.8 Omni FlashQwen (Alibaba) · 1M$0.15 / $0.47in / out per 1M$0.15$0.47$0.0161MSep 21, 2026
Qwen3.8 Max (0902)Qwen (Alibaba) · 1M$2 / $6in / out per 1M$2$6$0.251MSep 3, 2026
Qwen3.8 FlashQwen (Alibaba) · 1M$0.15 / $0.47in / out per 1M$0.15$0.47$0.0161MAug 26, 2026
GPT-5.6 SolOpenAI · 1.05M$2 / $10in / out per 1M$2$10$0.21.05MJul 9, 2026
Muse Spark 1.3Meta · 1M$1.25 / $4.25in / out per 1M$1.25$4.25$0.151MSep 2, 2026
Gemini 3.8 FlashGoogle Gemini API · 1M$0.75 / $3.75in / out per 1M$0.75$3.75$0.0751MSep 2, 2026
Gemini 3.1 Pro PreviewGoogle Gemini API · 1M$2 / $12in / out per 1M$2$12$0.21MFeb 19, 2026
About Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total.
Takes text in and returns text. First listed Sep 23, 2026. Blended price $3 per 1M tokens at 3 input to 1 output. Overall score 5 on leadersboard.fru.dev (rank 27).
Questions
How much does Qwen3.8 2.4T A95B cost?
$2 per million input tokens and $6 per million output tokens, with cached input at $0.25.
What is the context window of Qwen3.8 2.4T A95B?
1M tokens, with up to 131K tokens of output.
Where is Qwen3.8 2.4T A95B cheapest?
Of 7 routes tracked, DeepInfra is cheapest at $2 input and $6 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.