API pricing per million tokens
llama-3-3-70b-instruct Open weightsChecked 8h ago
- Input
- $0.1
- per 1M, DeepInfra (turbo)
- Output
- $0.32
- per 1M tokens
- Cached input
- n/a
- no batch tier listed
- Context
- 131K
- 16K max output
Where to buy it (14)
DeepInfraturbo, fp8$0.1 / $0.32$0.155 blended
Novita AIbf16, bf16$0.135 / $0.4$0.201 blended
- AkashMLfp8, fp8$0.2 / $0.52cached $0.1
Parasailfp8, fp8$0.22 / $0.5cached $0.11
SambaNovadefault region$0.45 / $0.9$0.563 blended
Groqdefault region$0.59 / $0.79cached $0.295
CoreWeavefp16, fp16$0.71 / $0.71cached $0.71
Google Vertex AIdefault region$0.72 / $0.72$0.72 blended
Google Vertex AIus-central1$0.72 / $0.72$0.72 blended
Databricks Mosaic AI Gateway + Foundation Model APIsdefault region, 7.143 DBU per 1M input tokens, 21.429 DBU per 1M output tokens$0.5 / $1.5$0.75 blended
IBM watsonx.aidefault region, USD 0.7526 per 1M tokens (single price listed)$0.753 / $0.753$0.753 blended
Cloudflare Workers AIfp8, fp8$0.293 / $2.25$0.783 blended
Snowflake Cortex AI (AI_COMPLETE)default region, 0.432 AI Credits per 1M input tokens, 0.432 per 1M output tokens$0.864 / $0.864$0.864 blended
Together AIdefault region$1.04 / $1.04$1.04 blended
Through a gateway
Your own numbers100M input and 20M output tokens a month: $16.40 direct.
Helicone$16.40+0%
Kong AI Gateway$16.40+0%
LiteLLM$16.40+0%
Vercel AI Gateway$16.40+0%
AWS Bedrock$16.40+0%
Azure AI Foundry$16.40+0%
Google Vertex AI$16.40+0%
Databricks Mosaic AI$16.40+0%
Hugging Face Inference$16.40+0%
Cloudflare AI Gateway$17.22+5%
Requesty$17.22+5%
Eden AI$17.30+5.5%
OpenRouter$17.30+5.5%
Snowflake Cortex AI$19.68+20%
Martian Gatewayn/a
Portkeyn/a
TrueFoundry AI Gatewayn/a
IBM watsonx.ain/a
OCI Generative AIn/a
BigQuery MLn/a
Cloudflare Workers AIn/a
NVIDIA NIMn/a
Related models
ModelInputOutputCachedContextReleased
Muse Spark 1.3Meta · 1M$1.25 / $4.25in / out per 1M$1.25$4.25$0.151MSep 2, 2026
Muse Spark 1.3 ContributorMeta · 1M$0.1 / $0.2in / out per 1M$0.1$0.2$0.0021MSep 2, 2026
Muse Spark 1.2 ContributorMeta · 1M$0.1 / $0.2in / out per 1M$0.1$0.2$0.0021MAug 21, 2026
Muse Glimmer 30BMeta · via Phala · 131K$0.3 / $1.1in / out per 1M$0.3$1.1$0.04131KAug 9, 2026
DeepSeek V4.1 FlashDeepSeek · 1M$0.15 / $0.6in / out per 1M$0.15$0.6$0.0031MSep 10, 2026
GPT-6 LunaOpenAI · 1.05M$0.1 / $0.5in / out per 1M$0.1$0.5$0.011.05MSep 22, 2026
GPT-6 Luna ProOpenAI · 1.05M$0.1 / $0.5in / out per 1M$0.1$0.5$0.011.05MSep 22, 2026
Qwen3.8 Omni FlashQwen (Alibaba) · 1M$0.15 / $0.47in / out per 1M$0.15$0.47$0.0161MSep 21, 2026
About Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).
Takes text in and returns text. First listed Sep 23, 2026. Blended price $0.155 per 1M tokens at 3 input to 1 output.
Questions
How much does Llama 3.3 70B Instruct cost?
$0.1 per million input tokens and $0.32 per million output tokens.
What is the context window of Llama 3.3 70B Instruct?
131K tokens, with up to 16K tokens of output.
Where is Llama 3.3 70B Instruct cheapest?
Of 14 routes tracked, DeepInfra is cheapest at $0.1 input and $0.32 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.