Ways to buy model access
Gateways, cloud and data platforms, and developer hubs: what each adds to the token price.Updated 5h ago
22 of 22 gateways and platforms
| Gateway or platform | Fee on tokens | Pricing model | BYOK | Models | Fallback | Cache | Guardrails | Logs | Rate limits | Residency | Private net |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Independent gateways | |||||||||||
0% markup on credits; plans by requests and seats (Pro $79 a month) · BYOK 0% · 100+ modelsFallbacks, Caching, Guardrails, Observability, Rate limits | None | 0% markup on credits; plans by requests and seats (Pro $79 a month) | Yes, 0% | 100+ | |||||||
$100 a month per model proxied (Konnect Plus); no token fee · BYOK no · See page modelsFallbacks, Caching, Guardrails, Observability, Rate limits | None | $100 a month per model proxied (Konnect Plus); no token fee | No | See page | |||||||
Open source, self-hosted with your own keys; enterprise license by quote · BYOK 0% · 100+ modelsFallbacks, Caching, Guardrails, Observability, Rate limits | None | Open source, self-hosted with your own keys; enterprise license by quote | Yes, 0% | 100+ | |||||||
Provider list prices, no markup or platform fee on tokens · BYOK 0% · See page modelsFallbacks, Caching, Observability, Rate limits | None | Provider list prices, no markup or platform fee on tokens | Yes, 0% | See page | |||||||
Core features free; 5% fee on Unified Billing credits · BYOK yes · See page modelsFallbacks, Caching, Guardrails, Observability, Rate limits | 5% | Core features free; 5% fee on Unified Billing credits | Yes | See page | |||||||
5% on top of model costs, pay as you go · BYOK yes · 600+ modelsFallbacks, Caching, Guardrails, Observability | 5% | 5% on top of model costs, pay as you go | Yes | 600+ | |||||||
Provider prices passed through; 5.5% fee on credit purchases · BYOK yes · 500+ modelsFallbacks, Observability | 5.5% | Provider prices passed through; 5.5% fee on credit purchases | Yes | 500+ | |||||||
Provider prices passed through; 5.5% fee on card credit purchases (5% crypto) · BYOK 5% · 500+ modelsFallbacks, Caching, Observability | 5.5% | Provider prices passed through; 5.5% fee on card credit purchases (5% crypto) | Yes, 5% | 500+ | |||||||
No public pricing · BYOK no · 200+ modelsObservability | Not published | No public pricing | No | 200+ | |||||||
Per logged request: free 10k logs a month, Production from $49 a month · BYOK yes · 250+ modelsFallbacks, Caching, Guardrails, Observability, Rate limits | $49/mo | Per logged request: free 10k logs a month, Production from $49 a month | Yes | 250+ | |||||||
Per user and per request ($25 per user a month on Pro) · BYOK no · 1,600+ modelsFallbacks, Guardrails, Observability, Rate limits | $25/mo | Per user and per request ($25 per user a month on Pro) | No | 1,600+ | |||||||
| Cloud model platforms | |||||||||||
Pass-through per 1M tokens at the lab's list price for Global cross-region inference; Geo and in-region inference cost 10% more; billed on your AWS account · BYOK no · 19 models | List price | Pass-through per 1M tokens at the lab's list price for Global cross-region inference; Geo and in-region inference cost 10% more; billed on your AWS account | No | 19 | |||||||
Per 1M (or 1K) tokens on your Azure bill; Global deployment at OpenAI list price, Data Zone and Regional deployments cost 10% more; PTUs for provisioned throughput · BYOK no · 2 models | List price | Per 1M (or 1K) tokens on your Azure bill; Global deployment at OpenAI list price, Data Zone and Regional deployments cost 10% more; PTUs for provisioned throughput | No | 2 | |||||||
Per 1M tokens on your Google Cloud bill; partner models at the lab's list price on the Global endpoint, regional/multi-region endpoints about 10% more · BYOK no · 2 models | List price | Per 1M tokens on your Google Cloud bill; partner models at the lab's list price on the Global endpoint, regional/multi-region endpoints about 10% more | No | 2 | |||||||
Per 1M tokens (billed as Resource Units, 1 RU = 1,000 tokens incl. input and output); Essentials plan from USD 0/month, Standard from USD 1110/month · BYOK no · See page models | Not published | Per 1M tokens (billed as Resource Units, 1 RU = 1,000 tokens incl. input and output); Essentials plan from USD 0/month, Standard from USD 1110/month | No | See page | |||||||
Mixed: per 1M tokens for Google, xAI and OpenAI gpt-oss models; per 10,000 transactions (1 transaction = 1 character of prompt plus response) for Meta and Cohere; dedicated AI clusters per unit hour · BYOK no · 3 models | Not published | Mixed: per 1M tokens for Google, xAI and OpenAI gpt-oss models; per 10,000 transactions (1 transaction = 1 character of prompt plus response) for Meta and Cohere; dedicated AI clusters per unit hour | No | 3 | |||||||
| Data platforms | |||||||||||
DBUs per 1M tokens (Standard pay per token; Priority and Batch tiers separate) · BYOK no · See page models | List price | DBUs per 1M tokens (Standard pay per token; Priority and Batch tiers separate) | No | See page | |||||||
Snowflake AI Credits per 1M tokens (input and output priced separately) · BYOK no · 6 models | 20% | Snowflake AI Credits per 1M tokens (input and output priced separately) | No | 6 | |||||||
BigQuery bytes processed (on-demand or editions) plus the underlying Vertex AI model charges billed directly by that service; no separate BigQuery per-token rate for LLM calls · BYOK no · See page models | Not published | BigQuery bytes processed (on-demand or editions) plus the underlying Vertex AI model charges billed directly by that service; no separate BigQuery per-token rate for LLM calls | No | See page | |||||||
| Developer hubs | |||||||||||
Pass-through at each provider's rate with no Hugging Face markup; pay as you go on your HF account, or bill directly to the provider with your own key · BYOK yes · 200+ models | List price | Pass-through at each provider's rate with no Hugging Face markup; pay as you go on your HF account, or bill directly to the provider with your own key | Yes | 200+ | |||||||
Neurons: $0.011 per 1,000 Neurons, with per-model token prices published as the Neuron equivalent · BYOK no · See page models | Not published | Neurons: $0.011 per 1,000 Neurons, with per-model token prices published as the Neuron equivalent | No | See page | |||||||
Hosted API endpoints free for NVIDIA Developer Program prototyping; production self-hosted NIM requires an NVIDIA AI Enterprise license from $4,500 per GPU per year (about $1 per GPU per hour in the cloud); no per-token price · BYOK no · n/a models | Not published | Hosted API endpoints free for NVIDIA Developer Program prototyping; production self-hosted NIM requires an NVIDIA AI Enterprise license from $4,500 per GPU per year (about $1 per GPU per hour in the cloud); no per-token price | No | n/a | |||||||
A tick means the vendor's own pricing page or docs say so; a dash means we did not find it stated. An amber dot marks a row with facts we could not confirm.
Claude Sonnet 5 through each
Calculator100M input and 20M output tokens a month: $400.00 direct from the lab.
Helicone$400.00+0%
Kong AI Gateway$400.00+0%
LiteLLM$400.00+0%
Vercel AI Gateway$400.00+0%
AWS Bedrock$400.00+0%
Azure AI Foundry$400.00+0%
Google Vertex AI$400.00+0%
Databricks Mosaic AI$400.00+0%
Hugging Face Inference$400.00+0%
Cloudflare AI Gateway$420.00+5%
Requesty$420.00+5%
Eden AI$422.00+5.5%
OpenRouter$422.00+5.5%
Snowflake Cortex AI$480.00+20%
Questions
What is an AI gateway?
A proxy between your app and model providers that gives you one API for many models, with fallbacks when a provider fails, caching, logs, rate limits and spend controls.
OpenRouter vs Vercel AI Gateway: which is cheaper?
Vercel AI Gateway states no markup and no platform fee on tokens. OpenRouter passes provider prices through but charges a fee when you buy credits (5.5% by card). With your own keys, OpenRouter's first allowance is free and then 5%.
What is BYOK?
Bring your own key: you give the gateway your own provider API key, pay the provider directly, and the gateway charges nothing or a small fee for routing.
Is Claude or GPT more expensive on Bedrock, Azure or Vertex AI?
Usually not in the default region: AWS Bedrock, Azure AI Foundry and Google Vertex AI mostly charge the lab's own list price, and some regional endpoints add about 10%. Each platform page lists every model it sells with the premium against the direct API.
How do Snowflake Cortex and Databricks price LLM calls?
In their own units: Snowflake bills Cortex AI functions in credits per million tokens and Databricks bills Foundation Model APIs in DBUs per million tokens. Their pages here convert those to dollars and state the credit or DBU price assumed.
Can I self-host an AI gateway?
Yes. LiteLLM, Portkey (gateway core) and Helicone publish open-source gateways you can run yourself; you pay only for hosting.
Pricing as read from each vendor's own pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.