Products

Four ways to run inference,
one API and console

Start on shared capacity and move to dedicated when traffic justifies it. The model IDs, keys, logs and billing stay the same the whole way.

01/Available now
Serverless InferencePay per token, no infrastructure
  • Per-token pricing, prepaid balance
  • Shared capacity with automatic failover
  • Rate limits by tier
  • Best for: getting started, variable traffic
02/Early access
Dedicated EndpointsIsolated capacity, predictable latency
  • Per-minute GPU pricing
  • Capacity reserved for you
  • No shared rate limits
  • Best for: steady production load
03/Early access
Coding PlanA fixed price per term for coding tools
  • Fixed price per term
  • Uses your existing API keys
  • Included usage in USD at list prices
  • Best for: coding agents and IDE tools
04/Talk to sales
GPU CloudBare metal and clusters, reserved by term
  • Bare metal and clusters, reserved by term
  • You run your own stack
  • Longer terms lower the rate
  • Best for: training and custom serving
Comparison

Which one fits

You can run more than one at a time. Most teams start on serverless and add a dedicated endpoint for the one model that carries steady load.

Serverless InferenceDedicated EndpointsCoding PlanGPU Cloud
BillingPer tokenPer GPU-minuteFixed price per termReserved by term, quoted per deal
CapacitySharedIsolatedSharedIsolated
Rate limitsBy tierNoneBy tierNone
API keysProject keysProject keysUses your existing API keysNot applicable
Cold startNoneOn scale from zeroNoneYou manage
Custom modelsNoYesNoYes
Setup effortChange three settingsA wizardSet base URL and model in the toolYou build it
StatusAvailable nowEarly accessEarly accessTalk to sales

Start on serverless today

Move to a dedicated endpoint later without changing your client code beyond the model name.