Models/deepseek-ai/DeepSeek-V4-Pro

DeepSeek V4 Pro

ActiveUnknown
deepseek-ai/DeepSeek-V4-Pro·deepseek-ai · MIT · updated 2026-04-24
RecommendedReasoningMathCoding
Overview

1.6T MoE flagship, 49B active. Extended reasoning traces before the answer.

Use cases: Reasoning · Math · Coding
Context1,048,576
Max output65,536
QuantizationFP8
Capabilities
ToolsParallel toolsStructured outputStreamReasoningCacheBatchVision
GridPort billing priceUSD / 1M tokens
UnitStandardBatch
Input tokens$0.66—
Cached input$0.022—
Output tokens$1.98—
Live metricslast hour

Last hour of customer requests; each metric needs at least 20 valid samples.Samples this hour: 0 for TTFT, 0 for throughput · as of 2026-10-07 00:35 EDT

TTFT p50—
Throughput p50—
Success rate—
Routes2 · failover on
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://gridport.ai/v1",
    api_key=os.environ["GRIDPORT_API_KEY"],
)
stream = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4-Pro",  # pin a version: "deepseek-ai/DeepSeek-V4-Pro@2026-04-24"
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()
Known limits
01Prompt cache belongs to a route: a failover to the backup route starts with a cold cache.02Rate limits apply per organization, project and key, and per model where one is set; the strictest applies.03Streaming responses carry usage in the final chunk only when stream_options.include_usage is set.
Versions
VersionStatusReleasedRetires
2026-04-24defaultActive2026-04-24—