Skip to content

Hugging Face Inference

Developer hub · list price · BYOK · checked Sep 23, 2026

About Hugging Face Inference

Hugging Face's router that sends OpenAI-compatible requests to partner inference providers (Groq, Together, Fireworks, Cerebras, and others) under one HF token and bill.

Fee on tokens
List price
Pricing model
Pass-through at each provider's rate with no Hugging Face markup; pay as you go on your HF account, or bill directly to the provider with your own key
Bring your own key
Yes, fee not stated
Entry plan
Usage based
Hosting
Hosted

Features stated: none found. Not stated: Fallbacks, Load balancing, Caching, Guardrails, Observability, Rate limits, Budgets, Prompts, OpenAI-compatible, Bring your own key, Data residency, Private networking.

The terms, in their words

Markup
Pass-through at each provider's rate with no Hugging Face markup; pay as you go on your HF account, or bill directly to the provider with your own key
Free tier
Monthly credits: $0.10 for free users, $2.00 for PRO, $2.00 per seat for Team/Enterprise
Providers
200+ models from about 17 listed providers (Baseten, Cerebras, Cohere, DeepInfra, Fal AI, Featherless AI, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeedAI, Z.ai)
  • Access 200+ models from leading AI inference providers with centralized, transparent, pay-as-you-go pricing.
  • Hugging Face charges you the same rates as the provider, with no additional fees. We just pass through the provider costs directly.
  • Custom Provider Key : You can bring your own provider key to use with the Inference Providers.

Source: huggingface.co/docs/inference-providers/pricing

Hugging Face Inference vsCloudflare Workers AINVIDIA NIMHeliconeKong AI GatewayLiteLLMVercel AI Gateway

Pricing as read from the gateway's own pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Weekly: LLM API list prices that changed, Thursday mornings.