The machines under the inference platform
Bare metal and managed clusters with InfiniBand, high-speed storage and the same region footprint our own inference runs on. Reserved by the month or the year.
Beyond the GPUs
Shared parallel filesystem for datasets and checkpoints, plus object storage in the same region so training reads do not cross the internet.
Private VPC, no egress charge between your nodes, and an optional private link to your own cloud account.
Single-tenant hardware. Nothing else is scheduled on your nodes, including our own serverless traffic.
SSH to bare metal, or a managed Kubernetes control plane with the GPU operator and drivers already in place.
What teams run on it
Multi-node runs that need collective communication to be fast and stable for days at a time. This is what the InfiniBand fabric is for.
Embedding a corpus, scoring a dataset, generating synthetic data. Cheaper per unit than the API when the volume is large and latency does not matter.
Run vLLM, SGLang, TensorRT-LLM or a proprietary engine with the flags you want, on hardware you control end to end.
GPU configurations
Every node ships with NVMe local scratch and non-blocking InfiniBand inside a rack. Availability and lead time vary by region — ask for current stock.
| Node | Memory | Interconnect | Price |
|---|---|---|---|
| 8 × NVIDIA H100 SXM | 640 GB | 3.2 Tb/s IB | Quoted |
| 8 × NVIDIA H200 SXM | 1,128 GB | 3.2 Tb/s IB | Quoted |
| 8 × NVIDIA B200 SXM | 1,440 GB | 3.2 Tb/s IB | Quoted |
| 8 × NVIDIA L40S PCIe | 384 GB | 200 Gb/s RoCE | Quoted |
Pricing depends on term, location and cluster size, and is quoted per deal.
Do you actually need machines?
Most inference workloads are cheaper and less work on the API. Rent hardware when one of these is true.
| If you need to | Use |
|---|---|
| Call catalog models and pay only for what you use | Serverless Inference |
| Serve one model with isolated capacity and no rate limit | Dedicated Endpoints |
| Train, fine-tune or run a custom engine yourself | GPU Cloud |
| Keep everything inside your own datacenter | Talk to us about on-premise |
Tell us the shape of the job
How many GPUs, for how long, and in which region. We come back with current availability, a lead time and a quote — usually within one business day.
- Single-tenant hardware, no shared scheduling
- Terms from one month, discounts for longer commits
- Same regions as the inference platform
Common questions
Yes. Single nodes have the shortest lead time and are the usual way to benchmark before committing to a cluster. Multi-node reservations need the InfiniBand fabric planned as a unit, which is what adds weeks.
Get availability and a quote
Send the GPU count, the term and the region. You will hear back with real numbers, not a brochure.