Products
Four ways to run inference,
one API and console
Start on shared capacity and move to dedicated when traffic justifies it. The model IDs, keys, logs and billing stay the same the whole way.
01/Available now
Serverless InferencePay per token, no infrastructure
- Per-token pricing, prepaid balance
- Shared capacity with automatic failover
- Rate limits by tier
- Best for: getting started, variable traffic
02/Early access
Dedicated EndpointsIsolated capacity, predictable latency
- Per-minute GPU pricing
- Capacity reserved for you
- No shared rate limits
- Best for: steady production load
03/Early access
Coding PlanA fixed price per term for coding tools
- Fixed price per term
- Uses your existing API keys
- Included usage in USD at list prices
- Best for: coding agents and IDE tools
04/Talk to sales
GPU CloudBare metal and clusters, reserved by term
- Bare metal and clusters, reserved by term
- You run your own stack
- Longer terms lower the rate
- Best for: training and custom serving
Comparison
Which one fits
You can run more than one at a time. Most teams start on serverless and add a dedicated endpoint for the one model that carries steady load.
| Serverless Inference | Dedicated Endpoints | Coding Plan | GPU Cloud | |
|---|---|---|---|---|
| Billing | Per token | Per GPU-minute | Fixed price per term | Reserved by term, quoted per deal |
| Capacity | Shared | Isolated | Shared | Isolated |
| Rate limits | By tier | None | By tier | None |
| API keys | Project keys | Project keys | Uses your existing API keys | Not applicable |
| Cold start | None | On scale from zero | None | You manage |
| Custom models | No | Yes | No | Yes |
| Setup effort | Change three settings | A wizard | Set base URL and model in the tool | You build it |
| Status | Available now | Early access | Early access | Talk to sales |
Start on serverless today
Move to a dedicated endpoint later without changing your client code beyond the model name.