Hugging Face Inference
Developer hub · list price · BYOK · checked Sep 23, 2026
About Hugging Face Inference
Hugging Face's router that sends OpenAI-compatible requests to partner inference providers (Groq, Together, Fireworks, Cerebras, and others) under one HF token and bill.
- Fee on tokens
- List price
- Pricing model
- Pass-through at each provider's rate with no Hugging Face markup; pay as you go on your HF account, or bill directly to the provider with your own key
- Bring your own key
- Yes, fee not stated
- Entry plan
- Usage based
- Hosting
- Hosted
Features stated: none found. Not stated: Fallbacks, Load balancing, Caching, Guardrails, Observability, Rate limits, Budgets, Prompts, OpenAI-compatible, Bring your own key, Data residency, Private networking.
The terms, in their words
- Markup
- Pass-through at each provider's rate with no Hugging Face markup; pay as you go on your HF account, or bill directly to the provider with your own key
- Free tier
- Monthly credits: $0.10 for free users, $2.00 for PRO, $2.00 per seat for Team/Enterprise
- Providers
- 200+ models from about 17 listed providers (Baseten, Cerebras, Cohere, DeepInfra, Fal AI, Featherless AI, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeedAI, Z.ai)
- Access 200+ models from leading AI inference providers with centralized, transparent, pay-as-you-go pricing.
- Hugging Face charges you the same rates as the provider, with no additional fees. We just pass through the provider costs directly.
- Custom Provider Key : You can bring your own provider key to use with the Inference Providers.
Source: huggingface.co/docs/inference-providers/pricing
Hugging Face Inference vsCloudflare Workers AINVIDIA NIMHeliconeKong AI GatewayLiteLLMVercel AI Gateway
Pricing as read from the gateway's own pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.