Response cache
Exact-match caching for non-streaming chat: a byte-identical request within the TTL is answered from cache at zero cost and logged as a hit.
Enable per request
curl
curl https://gridport.ai/v1/chat/completions -H "Authorization: Bearer $GRIDPORT_API_KEY" \ -H "Content-Type: application/json" \ -H "x-tf-cache: true" -H "x-tf-cache-ttl: 600" \ -d '{"model":"deepseek-ai/DeepSeek-V4-Flash","messages":[{"role":"user","content":"Define idempotency."}],"temperature":0}'
The key covers the model, messages and every sampling parameter; user and stream fields are ignored. The response carries x-tf-cache: miss the first time and hit afterwards, with usage.cost 0 and tf.cache "hit". TTL is 60 s to 24 h (default one hour). Streaming requests are never cached.
When it helps
Deterministic prompts (temperature 0) that repeat across users — classification, extraction, canned answers. It does not help for creative generation or when the prompt embeds a timestamp or user id. Semantic caching is deliberately not offered: exact matching is the only kind whose correctness can be explained.