Developer hub
NVIDIA NIM
Model prices vs the direct API
build.nvidia.comUnverified
- Fee on tokens
- Not published
- Hosted API endpoints free for NVIDIA Developer Program prototyping; production self-hosted NIM requires an NVIDIA AI Enterprise license from $4,500 per GPU per year (about $1 per GPU per hour in the cloud); no per-token price
- BYOK fee
- No BYOK
- Entry plan
- Usage based
- Hosting
- Hosted
Features
- Fallbacks
- Load balancing
- Caching
- Guardrails
- Observability
- Rate limits
- Budgets
- Prompts
- OpenAI-compatible
- Bring your own key
- Data residency
- Private networking
What it costs on real workloads
Claude Sonnet 5n/avs $400.00 direct
GPT-6 Soln/avs $400.00 direct
Gemini 3.8 Flashn/avs $150.00 direct
DeepSeek V4.1 Flashn/avs $27.00 direct
100M input and 20M output tokens a month, paying with gateway credits.
The terms, in their words
NVIDIA's API catalog of hosted model endpoints for prototyping plus downloadable NIM inference microservices for self-hosting on NVIDIA GPUs.
- Markup
- Hosted API endpoints free for NVIDIA Developer Program prototyping; production self-hosted NIM requires an NVIDIA AI Enterprise license from $4,500 per GPU per year (about $1 per GPU per hour in the cloud); no per-token price
- Free tier
- Free inference endpoints for NVIDIA Developer Program members (prototyping only); 90-day NVIDIA AI Enterprise trial
- These licenses start at $4500 per GPU per year or ~ $1 per GPU per hour in the cloud. Pricing is based on the number of GPUs, not the number of NIMs.
- Members of the NVIDIA Developer Program have free access to NIM API endpoints for prototyping
- Using NIM in production requires an NVIDIA AI Enterprise license.
NVIDIA NIM vs
Pricing as read from the gateway's own pages on the date shown; your contract may differ. Logos via logo.dev; trademarks belong to their owners.