Skip to content

Gemma 4 31B

Google · Open weights · 262K context · released Apr 2, 2026 · checked 10h ago

Ways to buy 13

Way to buyIn / out per 1M
DeepInfra
turbo, fp4 · cached $0.05
$0.09 / $0.34Cheapest
CoreWeave
fp4 · cached $0.1
$0.1 / $0.34+5% vs cheapest
Venice
fp4 · cached $0.09
$0.12 / $0.36+18% vs cheapest
Chutes
fp4 · cached $0.012
$0.12 / $0.37+20% vs cheapest
DeepInfra
fp8
$0.13 / $0.38+26% vs cheapest
Crusoe
bf16 · cached $0.14
$0.14 / $0.4+34% vs cheapest
Friendli
Default region
$0.14 / $0.4+34% vs cheapest
Novita AI
bf16
$0.14 / $0.4+34% vs cheapest
Parasail
fp8 · cached $0.06
$0.15 / $0.4+39% vs cheapest
DeepInfra
ultra, fp8
$0.27 / $0.76+157% vs cheapest
SambaNova
Default region
$0.38 / $1.15+275% vs cheapest
ModelRun
fp4 · cached $0.75
$0.75 / $1+433% vs cheapest
SiliconFlow
fp8 · cached $0.25
$0.75 / $1+433% vs cheapest

USD per 1M tokens. Regions are separate routes.

Through a gateway

Your own numbers

100M input and 20M output tokens a month: $15.80 direct.

No published token fee: Martian Gateway, Portkey, TrueFoundry AI Gateway.

Related models

In / out per 1M
Gemini 3.8 Flash
Google · 1MCheapest: $0.75 on Google Gemini API (flex)
$0.75 / $3.75
Gemini 3.7 Flash
Google · 1MCheapest: $0.75 on Google Vertex AI (global-flex)
$0.75 / $3.75100%
Gemini 3.5 Flash Lite
Google · 1MCheapest: $0.425 on Google Vertex AI (global-flex)
$0.3 / $2.5
Gemini 3.6 Flash
Google · 1MCheapest: $0.75 on Google Vertex AI (global-flex)
$0.75 / $3.75
DeepSeek V4.1 Flash
DeepSeek · 1MCheapest: $0.1 on Relace
$0.15 / $0.6
GPT-6 Luna
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex)
$0.1 / $0.5
GPT-6 Luna Pro
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex)
$0.1 / $0.5
Qwen3.8 Omni Flash
Qwen (Alibaba) · 1M
$0.15 / $0.47

About Gemma 4 31B

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.

Takes image, text, video in and returns text. First listed Sep 23, 2026. Up to 16K output tokens. Blended price $0.153 per 1M tokens at 3 input to 1 output. Model id gemma-4-31b-it.

Price history

No list-price history: this model is sold by hosts rather than its lab.

Questions

How much does Gemma 4 31B cost?

$0.09 per million input tokens and $0.34 per million output tokens, with cached input at $0.05.

What is the context window of Gemma 4 31B?

262K tokens, with up to 16K tokens of output.

Where is Gemma 4 31B cheapest?

Of 13 routes tracked, DeepInfra is cheapest at $0.09 input and $0.34 output per million tokens.

List prices from providers' pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Weekly: LLM API list prices that changed, Thursday mornings.