Rate limits
Limits apply per organization, project, API key and model, and a request has to fit every one of them. Organization limits are shared across projects; the table shows the configured tier baseline, and key limits, project caps and dedicated capacity may differ.
Tiers
Where limits apply
Each request counts against every layer that has a limit, and the first one without room refuses it with 429. error.dimension names that layer:
Headers and 429
x-ratelimit-limit-requests: 600 x-ratelimit-remaining-requests: 597 x-ratelimit-limit-tokens: 2000000 x-ratelimit-remaining-tokens: 1998112 HTTP/1.1 429 Too Many Requests retry-after: 12 {"error":{"code":"rate_limit_exceeded","type":"rate_limit_error","message":"…"}}
The x-ratelimit headers describe the limit that binds first: when a key, project or model limit has less room left than the organization’s, the headers show that limit and what is left of it.
Token limits are checked against an estimate before the call — the input plus max_tokens — and settled with the real usage afterwards. The official SDKs retry 429 with backoff automatically; keep that on and respect retry-after in your own clients.
A request whose estimate is larger than a whole minute of the limit can never be admitted, so it is not answered with 429: it gets 400 with param max_tokens and the largest max_tokens that would fit. A max_tokens above what the model can write is lowered to the model's ceiling and reported in x-tf-max-output-tokens.
Raising limits
When an automatic ladder is configured, Settings shows the next tier and outstanding requirements; a tier GridPort set by contract stays until GridPort changes it. Key and project ceilings remain independent. Inspect error.dimension: key_rpm/key_tpm points to API Keys; project_rpm/project_tpm to Settings → Project rate limits; org_rpm/org_tpm to Organization Settings; a per-model limit is GridPort’s to raise. Concurrency and upstream failures require reducing in-flight traffic or waiting for capacity, not a payment. Errors