xpu liveBETA
Models›Qwen3 Embedding 8B

Qwen3 Embedding 8B

Model reference ↗

Provider prices 1

Lowest input price per billing unit. Conditions in Details.

ProviderPriceProvider model IDStatusContext / max outputSource
Fireworks AINot collectedAwaiting verified pricingNot yet capturedCatalog-listed
Conditions

Catalog-listed; account access and regional availability may apply.

Context tokens 40960Official source ↗2026-09-28

Qwen3 Embedding 8B accepts text and produces embedding. Deployment limits and API options depend on the selected provider.

CONTEXT TOKENS32,768Developer model card
MAX OUTPUT TOKENSNot suppliedDeveloper model card
PROVIDERS1Linked in this catalog
SOURCE CHECKED2026-09-29Sources reviewed; unknown fields labeled

Specifications

Catalog model ID
qwen/qwen3-embedding-8b
Developer model ID / revision
Not verified in cited sources
Input modalities
textDeveloper model card ↗
Output modalities
embeddingDeveloper model card ↗
Context / input token limit
32,768Developer model card ↗
Maximum output tokens
Not verified in cited sources
Tool calling
Not supportedDeveloper model card ↗
Structured output
Not supportedDeveloper model card ↗
API parameters (source-specific)
Not verified in cited sources
Weight access
Open weightsDeveloper model card ↗
Weight license / access terms
Apache 2.0Developer model card ↗
Total parameters (billions)
8Developer model card ↗
Active parameters (billions)
Not verified in cited sources
Embedding dimensions
32–4,096Developer model card ↗
Download size (GB)
Depends on checkpoint and precision; no verified file size
Release date
Not verified in cited sources
Notes
Media limits on this page describe the linked provider endpoint; other deployments may differ. The model card specifies 32K input context and configurable embedding dimensions. A vector dimension is not an output-token limit.

Catalog entries link to their specification and availability sources. This is a curated selection, not every model each provider offers. Best prices are the lowest collected USD tariffs, not a total-cost estimate. Token offers are selected by input price, with output from the same tariff; workload mix may change the cheapest provider. Listed variants may require batch processing, off-peak hours, or specific media settings. Price coverage varies by provider; collected tariffs include their source and observation time. Benchmark scores are a dated snapshot, not a live feed or a universal measure of quality. They apply to the shown evaluation configuration; listed prices may use different settings. Models without verified scores are not assumed to be weaker.