GLM-5.3
ActiveUnknownzai-org/GLM-5.3·zai-org · GLM-5.3 License · updated 2026-08-14
RecommendedCodingAgentChinese
Overview
743B MoE aimed at agentic software engineering.
Use cases: Coding · Agent · ChineseContext1,048,576
Max output32,768
QuantizationFP8
Capabilities
ToolsParallel toolsStructured outputStreamReasoningCacheBatchVision
GridPort billing priceUSD / 1M tokens
UnitStandardBatch
Input tokens$1.40—
Cached input$0.26—
Output tokens$4.40—
Live metricslast hour
Last hour of customer requests; each metric needs at least 20 valid samples.Samples this hour: 0 for TTFT, 0 for throughput · as of 2026-10-07 00:33 EDT
TTFT p50—
Throughput p50—
Success rate—
Routes2 · failover on
import os from openai import OpenAI client = OpenAI( base_url="https://gridport.ai/v1", api_key=os.environ["GRIDPORT_API_KEY"], ) stream = client.chat.completions.create( model="zai-org/GLM-5.3", # pin a version: "zai-org/GLM-5.3@2026-08-14" messages=[{"role": "user", "content": "Hello"}], stream=True, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) print()
Known limits
01Prompt cache belongs to a route: a failover to the backup route starts with a cold cache.02Rate limits apply per organization, project and key, and per model where one is set; the strictest applies.03Streaming responses carry usage in the final chunk only when stream_options.include_usage is set.
Versions
VersionStatusReleasedRetires
2026-08-14defaultActiveReleased 2026-08-14Retires —