Capacity that is only yours
One model, on GPUs reserved for your project. No shared rate limits, latency you can predict, and the same API, keys and logs you already use.
When serverless stops being the right fit
Shared capacity is the right default. These are the signals that it is time to reserve your own.
Tier limits apply per key, project and organization. A dedicated endpoint has none — throughput is bounded by the replicas you run, not by a shared quota.
On shared capacity, first-token latency moves with everyone else's traffic. On reserved GPUs the p95 stays flat because nothing else is queued ahead of you.
Fine-tuned checkpoints, LoRA adapters or a model not in the catalog. Import from a model repository or object storage, then deploy it to an endpoint.
What you choose when you create one
- 01/Model and version
A catalog model pinned to a version, or one of your own custom model versions after it passes validation.
- 02/GPU and region
GPU type and count per replica, plus the region. The wizard shows the memory the model needs and the hourly price before you confirm.
- 03/Replicas
Minimum and maximum replicas with a target concurrency. Scale to zero is optional; the first request after that pays a cold start.
- 04/Access
Call it as model "ep/your-slug" with any project key, or restrict it to keys scoped to that endpoint.
Billed only while replicas are ready
Provisioning, image pull and model load are not billed. A replica starts billing the moment it passes the readiness check and stops the moment it is drained.
| GPU | Memory | Typical model size | Price |
|---|---|---|---|
| NVIDIA H100 80GB | 80 GB | up to 70B at FP8 | Quoted per deployment |
| NVIDIA H200 141GB | 141 GB | up to 140B at FP8 | Quoted per deployment |
| NVIDIA L40S 48GB | 48 GB | up to 32B at FP8 | Quoted per deployment |
What the console gives you
Concurrency, latency, throughput, GPU utilization and replica count, at 1h, 24h and 7d.
Startup, load and out-of-memory errors, tailed live. Request logs stay in the same place as serverless.
Config changes create a new revision and roll it out one replica at a time. A failed update keeps the old revision serving.
Every revision is kept. Rolling back is redeploying an earlier config, no rebuild needed.
Common questions
Only the model string. The base URL, the key, the response shape and the request logs are identical to serverless. Many teams run both and route a percentage of traffic to the endpoint to compare.
Reserve capacity for your busiest model
Tell us the model and the traffic shape. We will size the endpoint with you and enable self-service creation on your account.