GPU CloudTalk to salesContract capacity, quoted per deal

The machines under the inference platform

Bare metal and managed clusters with InfiniBand, high-speed storage and the same region footprint our own inference runs on. Reserved by the month or the year.

What comes with it

Beyond the GPUs

Storage

Shared parallel filesystem for datasets and checkpoints, plus object storage in the same region so training reads do not cross the internet.

Networking

Private VPC, no egress charge between your nodes, and an optional private link to your own cloud account.

Isolation

Single-tenant hardware. Nothing else is scheduled on your nodes, including our own serverless traffic.

Access

SSH to bare metal, or a managed Kubernetes control plane with the GPU operator and drivers already in place.

Use cases

What teams run on it

Training and fine-tuning

Multi-node runs that need collective communication to be fast and stable for days at a time. This is what the InfiniBand fabric is for.

Batch and offline jobs

Embedding a corpus, scoring a dataset, generating synthetic data. Cheaper per unit than the API when the volume is large and latency does not matter.

Your own stack

Run vLLM, SGLang, TensorRT-LLM or a proprietary engine with the flags you want, on hardware you control end to end.

Fleet

GPU configurations

Every node ships with NVMe local scratch and non-blocking InfiniBand inside a rack. Availability and lead time vary by region — ask for current stock.

NodeMemoryInterconnectPrice
8 × NVIDIA H100 SXM640 GB3.2 Tb/s IBQuoted
8 × NVIDIA H200 SXM1,128 GB3.2 Tb/s IBQuoted
8 × NVIDIA B200 SXM1,440 GB3.2 Tb/s IBQuoted
8 × NVIDIA L40S PCIe384 GB200 Gb/s RoCEQuoted
CommitMonthly, 6-month or annual. Longer terms lower the hourly rate.
Lead timeDays for single nodes, weeks for multi-rack clusters.
LocationWhere capacity is available; confirmed with each quote.

Pricing depends on term, location and cluster size, and is quoted per deal.

Choosing

Do you actually need machines?

Most inference workloads are cheaper and less work on the API. Rent hardware when one of these is true.

If you need toUse
Call catalog models and pay only for what you useServerless Inference
Serve one model with isolated capacity and no rate limitDedicated Endpoints
Train, fine-tune or run a custom engine yourselfGPU Cloud
Keep everything inside your own datacenterTalk to us about on-premise
Talk to sales

Tell us the shape of the job

How many GPUs, for how long, and in which region. We come back with current availability, a lead time and a quote — usually within one business day.

  • Single-tenant hardware, no shared scheduling
  • Terms from one month, discounts for longer commits
  • Same regions as the inference platform
Fields marked * are required.We use these details only to answer you. See the privacy policy.
FAQ

Common questions

Yes. Single nodes have the shortest lead time and are the usual way to benchmark before committing to a cluster. Multi-node reservations need the InfiniBand fabric planned as a unit, which is what adds weeks.

Get availability and a quote

Send the GPU count, the term and the region. You will hear back with real numbers, not a brochure.