LLM API cost calculator
Your tokens per month, then a model, then every way to buy it.
Step 1: Your usage
Step 2: Model
10 popular models, cheapest first at your usage
| Model | In / out per 1M | Per month |
|---|---|---|
| $0.15 / $0.6 | $25.94 | |
| $0.25 / $1.5 | $56.70 | |
| $0.43 / $1.75 | $78.55 | |
| $0.563 / $2.5 | $109.00 | |
| $0.75 / $3.75 | $151.35 | |
| $0.75 / $3.75 | $151.35 | |
| $1.25 / $4.25 | $203.45 | |
| $1.25 / $4.25 | $203.45 | |
| $1.25 / $4.25 | $203.45 | |
| $1.6 / $4.8 | $254.40 |
Step 3: Result
CompareModel pageGemini 3.1 Flash Lite Preview at 120M input and 25M output tokens a month, 40% of input cached: from $56.70 a month.
| Way to buy | In / out per 1M | Per month | Vs cheapest |
|---|---|---|---|
The lab's own list price · $0.25 / $1.5 | $0.25 / $1.5 | $56.70Cheapest | Cheapest |
Gateway, no token markup, plan from $79/mo · $0.25 / $1.5 | $0.25 / $1.5 | $56.70BYOK at costSame | Same |
Gateway, no token markup, plan from $100/mo · $0.25 / $1.5 | $0.25 / $1.5 | $56.70Same | Same |
Gateway, no token markup, self-hosted · $0.25 / $1.5 | $0.25 / $1.5 | $56.70BYOK at costSame | Same |
Gateway, no token markup · $0.25 / $1.5 | $0.25 / $1.5 | $56.70BYOK at costSame | Same |
Gateway, 5% on credits · $0.263 / $1.58 | $0.263 / $1.58 | $59.54+$2.84 (+5.0%) | +$2.84 (+5.0%) |
Gateway, 5% on credits · $0.263 / $1.58 | $0.263 / $1.58 | $59.54+$2.84 (+5.0%) | +$2.84 (+5.0%) |
Gateway, 5.5% on credits · $0.264 / $1.58 | $0.264 / $1.58 | $59.82+$3.12 (+5.5%) | +$3.12 (+5.5%) |
Gateway, 5.5% on credits · $0.264 / $1.58 | $0.264 / $1.58 | $59.82BYOK at cost+$3.12 (+5.5%) | +$3.12 (+5.5%) |
Cached input is billed at the cache-read price where one is published; batch applies only where the lab sells a batch tier. Platform rows are list prices without caching or batch. Gateway rows add the gateway's fee on credits; BYOK is the cost with your own provider key. Not priced here (no published token fee): Martian Gateway, Portkey, TrueFoundry AI Gateway.
Questions
How do I estimate my monthly LLM API cost?
Multiply the millions of input tokens you send by the input price, add the millions of output tokens times the output price, and subtract the discount on any input served from cache. The calculator does this for every model at once.
How many tokens is a word?
For English text, about 1.3 tokens per word, or roughly 750 words per 1,000 tokens. Code and other languages use more tokens per word.
How much does prompt caching save?
Cached input is typically billed at 10% to 50% of the normal input price, depending on the provider. Long, repeated system prompts and documents benefit most.
What is the Batch API discount?
OpenAI, Anthropic, Google, Mistral and AWS Bedrock sell asynchronous batch processing at about half price, with results returned within 24 hours.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.