Rate limits

Limits apply per organization, project, API key and model, and a request has to fit every one of them. Organization limits are shared across projects; the table shows the configured tier baseline, and key limits, project caps and dedicated capacity may differ.

Tiers

Where limits apply

Each request counts against every layer that has a limit, and the first one without room refuses it with 429. error.dimension names that layer:

LayerSet byerror.dimension
OrganizationThe tier’s requests and tokens per minute, shared by every project. Tiers rise with paid usage, or GridPort sets one by contract.org_rpm · org_tpm
ProjectA cap under Settings → Project rate limits, set by owners, admins and project admins, so one team cannot use up the organization’s limit.project_rpm · project_tpm
API keyThe key’s own RPM and TPM.key_rpm · key_tpm
ModelA limit for one model, from the tier or set by GridPort for the organization or a project. It applies to that model only; error.model names it.org_model_rpm · project_model_rpm (· _tpm)

Headers and 429

http
x-ratelimit-limit-requests: 600
x-ratelimit-remaining-requests: 597
x-ratelimit-limit-tokens: 2000000
x-ratelimit-remaining-tokens: 1998112

HTTP/1.1 429 Too Many Requests
retry-after: 12
{"error":{"code":"rate_limit_exceeded","type":"rate_limit_error","message":"…"}}

The x-ratelimit headers describe the limit that binds first: when a key, project or model limit has less room left than the organization’s, the headers show that limit and what is left of it.

Token limits are checked against an estimate before the call — the input plus max_tokens — and settled with the real usage afterwards. The official SDKs retry 429 with backoff automatically; keep that on and respect retry-after in your own clients.

A request whose estimate is larger than a whole minute of the limit can never be admitted, so it is not answered with 429: it gets 400 with param max_tokens and the largest max_tokens that would fit. A max_tokens above what the model can write is lowered to the model's ceiling and reported in x-tf-max-output-tokens.

Raising limits

When an automatic ladder is configured, Settings shows the next tier and outstanding requirements; a tier GridPort set by contract stays until GridPort changes it. Key and project ceilings remain independent. Inspect error.dimension: key_rpm/key_tpm points to API Keys; project_rpm/project_tpm to Settings → Project rate limits; org_rpm/org_tpm to Organization Settings; a per-model limit is GridPort’s to raise. Concurrency and upstream failures require reducing in-flight traffic or waiting for capacity, not a payment. Errors