Serverless InferenceAvailable nowOpenAI and Anthropic compatible

Call open models like you call OpenAI

Prepaid balance, per-token billing, no infrastructure to run. Change three settings in your existing OpenAI code and every request becomes traceable and explainable.

01Gateway overhead target<50ms

Design target for the time the gateway adds on top of the upstream first token. Measured latency per model is on each model page.

02SLA uptime99.9%

Monthly, backed by service credits.

03Price increase notice30d

Published prices rise only after 30 days.

04Deprecation notice30d

By email and response header.

What you get

Built for running in production

pythonThree lines change.
base_url="https://gridport.ai/v1",api_key=os.environ["GRIDPORT_API_KEY"],model="deepseek-ai/DeepSeek-V4-Flash",
OpenAI-compatible

Parameters the platform does not know are ignored and named in the x-tf-ignored-params response header instead of failing the call. A parameter that needs a capability the model lacks, such as tools on a model without tool calling, is refused with 400 unsupported_parameter.

Route
  1. 1route-aTimed out at 800 ms · not billed
  2. 2route-bServed · switched before the first byte
Automatic failover

Where a model has more than one route, a failure before the first byte switches route transparently and writes the switch to the response headers. Each model page shows how many routes it has.

req_01JQ8N4W7X2K502
upstream_timeout

This was not a problem with your request.

Replay in PlaygroundRouting trail
Request-level logs

Every response carries a Request ID. The log shows the error reason, the routing trail across attempts, tokens and cost for 90 days.

production-backendsk-tf-…a41c
2 modelsIP allowlist$50 / month
Budget this month$38.20 / $50
Keys and budgets

Model and IP allowlists, per-key budgets, project budgets and three-stage balance alerts. Every block returns a specific error code.

Repeated prefix
New tokens
Cached input rateStandard input rate
Prompt caching

Repeated prefixes are billed at the model's cached input rate, listed on each model page. Cache state is reported per request.

Playground

Try any model with real parameters, see first-token latency and cost per reply, then export the exact request as code.

How it works

What happens to a request

Every call goes through the same pipeline. Each step either passes the request on or returns a specific error code you can act on.

  1. 01/Authenticate

    The key is hashed and checked against project, organization and model allowlists. A disabled key stops working within five seconds.

  2. 02/Limit and check balance

    Rate limits apply per key, project, organization and tier — the strictest wins. Available balance is checked before any upstream call.

  3. 03/Route

    The model resolves to a version, then to a route. If the first route fails before the first byte, the gateway switches and records the attempt.

  4. 04/Meter

    Input, cached input and output are metered separately at the price in effect when the request started, then settled into the ledger.

Pricing

Per million tokens

Prepaid balance. Cached input is billed at a lower rate. Batch jobs use the batch price: half the standard rate unless the model lists its own.

Full price list and rate limits
FAQ

Common questions

Three settings change: base_url, api_key and model. Nothing else in your code changes. The migration assistant checks a pasted request parameter by parameter and suggests a model mapping.

Your first call in ten minutes

Create a project key and compare actual usage before committing to a plan.