Models/Qwen/Qwen3-Embedding-8B

Qwen3 Embedding 8B

ActiveUnknown
Qwen/Qwen3-Embedding-8B·Qwen · Apache 2.0 · updated 2026-01-22
RecommendedEmbeddingMultilingual
Overview

Top open-weight model on MTEB v2 retrieval, 4096 dimensions.

Use cases: Embedding · Multilingual
Context32,768
Max output—
Capabilities
ToolsParallel toolsStructured outputStreamReasoningCacheBatchVision
GridPort billing priceUSD / 1M tokens
UnitStandardBatch
Input tokens$0.010—
Live metricslast hour

Last hour of customer requests; each metric needs at least 20 valid samples.Samples this hour: 0 for TTFT, 0 for throughput · as of 2026-10-07 00:35 EDT

TTFT p50—
Throughput p50—
Success rate—
Routes2 · failover on
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://gridport.ai/v1",
    api_key=os.environ["GRIDPORT_API_KEY"],
)
resp = client.embeddings.create(
    model="Qwen/Qwen3-Embedding-8B",  # pin a version: "Qwen/Qwen3-Embedding-8B@2026-01-22"
    input=["The quick brown fox","jumps over the lazy dog"],
)
print(len(resp.data), "vectors of", len(resp.data[0].embedding), "dimensions")
Known limits
01Prompt cache belongs to a route: a failover to the backup route starts with a cold cache.02Rate limits apply per organization, project and key, and per model where one is set; the strictest applies.03Streaming responses carry usage in the final chunk only when stream_options.include_usage is set.
Versions
VersionStatusReleasedRetires
2026-01-22defaultActive2026-01-22—