Ways to buy 13
| Way to buy | Input | Output | In / out per 1M | Cached | Vs cheapest |
|---|---|---|---|---|---|
turbo, fp4 · cached $0.05 | $0.09 | $0.34 | $0.09 / $0.34Cheapest | $0.05 | Cheapest |
fp4 · cached $0.1 | $0.1 | $0.34 | $0.1 / $0.34+5% vs cheapest | $0.1 | +5% vs cheapest |
fp4 · cached $0.09 | $0.12 | $0.36 | $0.12 / $0.36+18% vs cheapest | $0.09 | +18% vs cheapest |
fp4 · cached $0.012 | $0.12 | $0.37 | $0.12 / $0.37+20% vs cheapest | $0.012 | +20% vs cheapest |
fp8 | $0.13 | $0.38 | $0.13 / $0.38+26% vs cheapest | n/a | +26% vs cheapest |
bf16 · cached $0.14 | $0.14 | $0.4 | $0.14 / $0.4+34% vs cheapest | $0.14 | +34% vs cheapest |
Default region | $0.14 | $0.4 | $0.14 / $0.4+34% vs cheapest | n/a | +34% vs cheapest |
bf16 | $0.14 | $0.4 | $0.14 / $0.4+34% vs cheapest | n/a | +34% vs cheapest |
fp8 · cached $0.06 | $0.15 | $0.4 | $0.15 / $0.4+39% vs cheapest | $0.06 | +39% vs cheapest |
ultra, fp8 | $0.27 | $0.76 | $0.27 / $0.76+157% vs cheapest | n/a | +157% vs cheapest |
Default region | $0.38 | $1.15 | $0.38 / $1.15+275% vs cheapest | n/a | +275% vs cheapest |
| ModelRun fp4 · cached $0.75 | $0.75 | $1 | $0.75 / $1+433% vs cheapest | $0.75 | +433% vs cheapest |
fp8 · cached $0.25 | $0.75 | $1 | $0.75 / $1+433% vs cheapest | $0.25 | +433% vs cheapest |
USD per 1M tokens. Regions are separate routes.
Through a gateway
Your own numbers100M input and 20M output tokens a month: $15.80 direct.
Helicone$15.80+0%
Kong AI Gateway$15.80+0%
LiteLLM$15.80+0%
Vercel AI Gateway$15.80+0%
Cloudflare AI Gateway$16.59+5%
Requesty$16.59+5%
Eden AI$16.67+5.5%
OpenRouter$16.67+5.5%
No published token fee: Martian Gateway, Portkey, TrueFoundry AI Gateway.
Related models
| In / out per 1M | Cheapest route | ||||||
|---|---|---|---|---|---|---|---|
Google · 1MCheapest: $0.75 on Google Gemini API (flex) | $0.75 | $3.75 | $0.75 / $3.75 | $1.5 | 1M | $0.75 on Google Gemini API (flex) | · |
Google · 1MCheapest: $0.75 on Google Vertex AI (global-flex) | $0.75 | $3.75 | $0.75 / $3.75100% | $1.5 | 1M | $0.75 on Google Vertex AI (global-flex) | 100% |
Google · 1MCheapest: $0.425 on Google Vertex AI (global-flex) | $0.3 | $2.5 | $0.3 / $2.5 | $0.85 | 1M | $0.425 on Google Vertex AI (global-flex) | · |
Google · 1MCheapest: $0.75 on Google Vertex AI (global-flex) | $0.75 | $3.75 | $0.75 / $3.75 | $1.5 | 1M | $0.75 on Google Vertex AI (global-flex) | · |
DeepSeek · 1MCheapest: $0.1 on Relace | $0.15 | $0.6 | $0.15 / $0.6 | $0.262 | 1M | $0.1 on Relace | · |
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex) | $0.1 | $0.5 | $0.1 / $0.5 | $0.2 | 1.05M | $0.1 on OpenAI (flex) | · |
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex) | $0.1 | $0.5 | $0.1 / $0.5 | $0.2 | 1.05M | $0.1 on OpenAI (flex) | · |
Qwen (Alibaba) · 1M | $0.15 | $0.47 | $0.15 / $0.47 | $0.23 | 1M | Lab list price | · |
About Gemma 4 31B
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.
Takes image, text, video in and returns text. First listed Sep 23, 2026. Up to 16K output tokens. Blended price $0.153 per 1M tokens at 3 input to 1 output. Model id gemma-4-31b-it.
Price history
No list-price history: this model is sold by hosts rather than its lab.
Across fru.devCompany profile7 acquisitionsPay schedule7 conferences
Questions
How much does Gemma 4 31B cost?
$0.09 per million input tokens and $0.34 per million output tokens, with cached input at $0.05.
What is the context window of Gemma 4 31B?
262K tokens, with up to 16K tokens of output.
Where is Gemma 4 31B cheapest?
Of 13 routes tracked, DeepInfra is cheapest at $0.09 input and $0.34 output per million tokens.
List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.