Skip to content

Ways to buy model access

Gateways, cloud and data platforms, and developer hubs: what each adds to the token price.Updated 5h ago

22 of 22 gateways and platforms

Gateway or platformFee on tokens
Independent gateways
Helicone
0% markup on credits; plans by requests and seats (Pro $79 a month) · BYOK 0% · 100+ modelsFallbacks, Caching, Guardrails, Observability, Rate limits
None
Kong AI Gateway
$100 a month per model proxied (Konnect Plus); no token fee · BYOK no · See page modelsFallbacks, Caching, Guardrails, Observability, Rate limits
None
LiteLLM
Open source, self-hosted with your own keys; enterprise license by quote · BYOK 0% · 100+ modelsFallbacks, Caching, Guardrails, Observability, Rate limits
None
Vercel AI Gateway
Provider list prices, no markup or platform fee on tokens · BYOK 0% · See page modelsFallbacks, Caching, Observability, Rate limits
None
Cloudflare AI Gateway
Core features free; 5% fee on Unified Billing credits · BYOK yes · See page modelsFallbacks, Caching, Guardrails, Observability, Rate limits
5%
Requesty
5% on top of model costs, pay as you go · BYOK yes · 600+ modelsFallbacks, Caching, Guardrails, Observability
5%
Eden AI
Provider prices passed through; 5.5% fee on credit purchases · BYOK yes · 500+ modelsFallbacks, Observability
5.5%
OpenRouter
Provider prices passed through; 5.5% fee on card credit purchases (5% crypto) · BYOK 5% · 500+ modelsFallbacks, Caching, Observability
5.5%
Martian GatewayUnverified
No public pricing · BYOK no · 200+ modelsObservability
Not published
Portkey
Per logged request: free 10k logs a month, Production from $49 a month · BYOK yes · 250+ modelsFallbacks, Caching, Guardrails, Observability, Rate limits
$49/mo
TrueFoundry AI GatewayUnverified
Per user and per request ($25 per user a month on Pro) · BYOK no · 1,600+ modelsFallbacks, Guardrails, Observability, Rate limits
$25/mo
Cloud model platforms
AWS Bedrock
Pass-through per 1M tokens at the lab's list price for Global cross-region inference; Geo and in-region inference cost 10% more; billed on your AWS account · BYOK no · 19 models
List price
Azure AI Foundry
Per 1M (or 1K) tokens on your Azure bill; Global deployment at OpenAI list price, Data Zone and Regional deployments cost 10% more; PTUs for provisioned throughput · BYOK no · 2 models
List price
Google Vertex AI
Per 1M tokens on your Google Cloud bill; partner models at the lab's list price on the Global endpoint, regional/multi-region endpoints about 10% more · BYOK no · 2 models
List price
IBM watsonx.ai
Per 1M tokens (billed as Resource Units, 1 RU = 1,000 tokens incl. input and output); Essentials plan from USD 0/month, Standard from USD 1110/month · BYOK no · See page models
Not published
OCI Generative AIUnverified
Mixed: per 1M tokens for Google, xAI and OpenAI gpt-oss models; per 10,000 transactions (1 transaction = 1 character of prompt plus response) for Meta and Cohere; dedicated AI clusters per unit hour · BYOK no · 3 models
Not published
Data platforms
Databricks Mosaic AIUnverified
DBUs per 1M tokens (Standard pay per token; Priority and Batch tiers separate) · BYOK no · See page models
List price
Snowflake Cortex AI
Snowflake AI Credits per 1M tokens (input and output priced separately) · BYOK no · 6 models
20%
BigQuery ML
BigQuery bytes processed (on-demand or editions) plus the underlying Vertex AI model charges billed directly by that service; no separate BigQuery per-token rate for LLM calls · BYOK no · See page models
Not published
Developer hubs
Hugging Face Inference
Pass-through at each provider's rate with no Hugging Face markup; pay as you go on your HF account, or bill directly to the provider with your own key · BYOK yes · 200+ models
List price
Cloudflare Workers AI
Neurons: $0.011 per 1,000 Neurons, with per-model token prices published as the Neuron equivalent · BYOK no · See page models
Not published
NVIDIA NIMUnverified
Hosted API endpoints free for NVIDIA Developer Program prototyping; production self-hosted NIM requires an NVIDIA AI Enterprise license from $4,500 per GPU per year (about $1 per GPU per hour in the cloud); no per-token price · BYOK no · n/a models
Not published

A tick means the vendor's own pricing page or docs say so; a dash means we did not find it stated. An amber dot marks a row with facts we could not confirm.

Claude Sonnet 5 through each

Calculator

100M input and 20M output tokens a month: $400.00 direct from the lab.

Questions

What is an AI gateway?

A proxy between your app and model providers that gives you one API for many models, with fallbacks when a provider fails, caching, logs, rate limits and spend controls.

OpenRouter vs Vercel AI Gateway: which is cheaper?

Vercel AI Gateway states no markup and no platform fee on tokens. OpenRouter passes provider prices through but charges a fee when you buy credits (5.5% by card). With your own keys, OpenRouter's first allowance is free and then 5%.

What is BYOK?

Bring your own key: you give the gateway your own provider API key, pay the provider directly, and the gateway charges nothing or a small fee for routing.

Is Claude or GPT more expensive on Bedrock, Azure or Vertex AI?

Usually not in the default region: AWS Bedrock, Azure AI Foundry and Google Vertex AI mostly charge the lab's own list price, and some regional endpoints add about 10%. Each platform page lists every model it sells with the premium against the direct API.

How do Snowflake Cortex and Databricks price LLM calls?

In their own units: Snowflake bills Cortex AI functions in credits per million tokens and Databricks bills Foundation Model APIs in DBUs per million tokens. Their pages here convert those to dollars and state the credit or DBU price assumed.

Can I self-host an AI gateway?

Yes. LiteLLM, Portkey (gateway core) and Helicone publish open-source gateways you can run yourself; you pay only for hosting.

Pricing as read from each vendor's own pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Weekly: LLM API list prices that changed, Thursday mornings.