Dedicated EndpointsEarly accessRolling out to enterprise organizations

Capacity that is only yours

One model, on GPUs reserved for your project. No shared rate limits, latency you can predict, and the same API, keys and logs you already use.

Dedicated endpoints are in early access. Self-service creation is being rolled out organization by organization — request access and we will enable it on your account, usually within two business days.
Why move

When serverless stops being the right fit

Shared capacity is the right default. These are the signals that it is time to reserve your own.

Serverless Inference
KeyProjectOrganizationTier
Dedicated Endpoints
Replicas
You are hitting rate limits

Tier limits apply per key, project and organization. A dedicated endpoint has none — throughput is bounded by the replicas you run, not by a shared quota.

First-token latencyIllustration
Shared capacityReserved GPUs
Latency varies too much

On shared capacity, first-token latency moves with everyone else's traffic. On reserved GPUs the p95 stays flat because nothing else is queued ahead of you.

Model repositoryObject storage
ep/your-slug
You need your own weights

Fine-tuned checkpoints, LoRA adapters or a model not in the catalog. Import from a model repository or object storage, then deploy it to an endpoint.

Configuration

What you choose when you create one

  1. 01/Model and version

    A catalog model pinned to a version, or one of your own custom model versions after it passes validation.

  2. 02/GPU and region

    GPU type and count per replica, plus the region. The wizard shows the memory the model needs and the hourly price before you confirm.

  3. 03/Replicas

    Minimum and maximum replicas with a target concurrency. Scale to zero is optional; the first request after that pays a cold start.

  4. 04/Access

    Call it as model "ep/your-slug" with any project key, or restrict it to keys scoped to that endpoint.

Billing

Billed only while replicas are ready

Provisioning, image pull and model load are not billed. A replica starts billing the moment it passes the readiness check and stops the moment it is drained.

Ask for a quote
GPUMemoryTypical model sizePrice
NVIDIA H100 80GB80 GBup to 70B at FP8Quoted per deployment
NVIDIA H200 141GB141 GBup to 140B at FP8Quoted per deployment
NVIDIA L40S 48GB48 GBup to 32B at FP8Quoted per deployment
BilledReady replicas, by the minute
Not billedQueueing, image pull, model load, scaled-to-zero time
PrepaidA short coverage window (15 minutes by default) is paid from your balance ahead and renewed while the endpoint runs
Operations

What the console gives you

Monitoring

Concurrency, latency, throughput, GPU utilization and replica count, at 1h, 24h and 7d.

Engine logs

Startup, load and out-of-memory errors, tailed live. Request logs stay in the same place as serverless.

Rolling updates

Config changes create a new revision and roll it out one replica at a time. A failed update keeps the old revision serving.

Rollback

Every revision is kept. Rolling back is redeploying an earlier config, no rebuild needed.

FAQ

Common questions

Only the model string. The base URL, the key, the response shape and the request logs are identical to serverless. Many teams run both and route a percentage of traffic to the endpoint to compare.

Reserve capacity for your busiest model

Tell us the model and the traffic shape. We will size the endpoint with you and enable self-service creation on your account.