LLM API cost calculator
Your tokens per month, priced on every model, route and gateway.
Cached input is billed at the model's cache-read price where one is published; batch applies only where the lab sells a batch tier.
GPT-5.4
$2.5 in, $15 out per 1M
On a cloud or data platform
Databricks Mosaic AI Gateway + Foundation Model APIs$2.5 in, $15 out per 1M$675.00
Azure AI Foundry$2.5 in, $15 out per 1M$675.00
Azure AI Foundry (us)$2.75 in, $16.5 out per 1M$742.50
Azure AI Foundry (eu)$2.75 in, $16.5 out per 1M$742.50
AWS Bedrock (us-east-1)$2.75 in, $16.5 out per 1M$742.50
Snowflake Cortex AI (AI_COMPLETE)$3 in, $18 out per 1M$810.00
Platform prices without caching or batch discounts.
Through a gateway
HeliconeNo token markup, plan from $79/mo$567.00BYOK $567.00
Kong AI GatewayNo token markup, plan from $100/mo$567.00no BYOK
LiteLLMNo token markup, self-hosted$567.00BYOK $567.00
Vercel AI GatewayNo token markup$567.00BYOK $567.00
Cloudflare AI Gateway5% on credit purchases$595.35BYOK fee not published
Requesty5% on credit purchases$595.35BYOK fee not published
Eden AI5.5% on credit purchases$598.18BYOK fee not published
OpenRouter5.5% on credit purchases$598.18BYOK $567.00
Martian GatewayNo token markup publishedn/ano BYOK
PortkeyNo token markup published, plan from $49/mon/aBYOK fee not published
TrueFoundry AI GatewayNo token markup published, plan from $25/mon/ano BYOK
Monthly bill by model
The 24 cheapest of 253. Bars are on a log scale. Tap a model to price it through each gateway.
Questions
How do I estimate my monthly LLM API cost?
Multiply the millions of input tokens you send by the input price, add the millions of output tokens times the output price, and subtract the discount on any input served from cache. The calculator does this for every model at once.
How many tokens is a word?
For English text, about 1.3 tokens per word, or roughly 750 words per 1,000 tokens. Code and other languages use more tokens per word.
How much does prompt caching save?
Cached input is typically billed at 10% to 50% of the normal input price, depending on the provider. Long, repeated system prompts and documents benefit most.
What is the Batch API discount?
OpenAI, Anthropic, Google, Mistral and AWS Bedrock sell asynchronous batch processing at about half price, with results returned within 24 hours.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.