Skip to content
Data platform

BigQuery ML

Model prices vs the direct API

cloud.google.comChecked Sep 23, 2026
Pricing page Docs
Fee on tokens
Not published
BigQuery bytes processed (on-demand or editions) plus the underlying Vertex AI model charges billed directly by that service; no separate BigQuery per-token rate for LLM calls
BYOK fee
No BYOK
Entry plan
Usage based
Hosting
Hosted

Features

  • Fallbacks
  • Load balancing
  • Caching
  • Guardrails
  • Observability
  • Rate limits
  • Budgets
  • Prompts
  • OpenAI-compatible
  • Bring your own key
  • Data residency
  • Private networking

What it costs on real workloads

100M input and 20M output tokens a month, paying with gateway credits.

The terms, in their words

BigQuery ML remote models that call Vertex AI (Gemini Enterprise Agent Platform) models such as Gemini and Claude from SQL via ML.GENERATE_TEXT.

Markup
BigQuery bytes processed (on-demand or editions) plus the underlying Vertex AI model charges billed directly by that service; no separate BigQuery per-token rate for LLM calls
Providers
Google models on Agent Platform, Anthropic Claude models enabled on Agent Platform, and open models deployed to Agent Platform
  • The bytes processed by BigQuery are billed according to standard pricing such as on-demand or editions pricing.
  • For remote endpoint model pricing, you are billed directly by the above services.
  • Google models hosted on Gemini Enterprise Agent Platform ML.GENERATE_TEXT ML.GENERATE_EMBEDDING Generative AI on Gemini Enterprise Agent Platform batch pricing

Source: cloud.google.com/bigquery/pricing

BigQuery ML vs

BigQuery, the Gemini Enterprise Agent Platform (formerly Vertex AI), the Gemini API, Looker, AlloyDB and Spanner.

Pricing as read from the gateway's own pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Weekly: LLM API list prices that changed, Thursday mornings.