DeepSeek V4.1 Flash
| Provider | Price | Provider model ID | Status | Context / max output | Source |
|---|---|---|---|---|---|
Off-peak Input $0.15USD / 1M tokensOutput $0.6USD / 1M tokens DeepSeek · Off-peakCache Read: $0.003 Input: $0.15 Output: $0.6 Time-of-day tariff; consult official source for UTC peak/off-peak windows and alias behavior. Checked 2026-10-03T14:45:20.473Z Price source ↗Peak Tariff not listedInput $0.3USD / 1M tokensOutput $1.2USD / 1M tokens DeepSeek · PeakCache Read: $0.006 Input: $0.3 Output: $1.2 Time-of-day tariff; consult official source for UTC peak/off-peak windows and alias behavior. Checked 2026-10-03T14:45:20.473Z Price source ↗ | deepseek-flash | Catalog-listedConditionsCatalog-listed; account access and regional availability may apply. | Context tokens 1000000Output tokens 384000 | Official source ↗2026-09-28 | |
Standard list Tariff not listedInput $0.2USD / 1M tokensOutput $0.6USD / 1M tokens DeepInfra · Standard listInput: $0.2 Output: $0.6 Published cents-per-token list rate. Account discounts and cache/service-tier modifiers are not applied. Checked 2026-10-03T14:45:20.473Z Price source ↗ | deepseek-ai/DeepSeek-V4.1-Flash | Catalog-listedConditionsCatalog-listed; account access and regional availability may apply. | Context tokens 1048576 | Official source ↗2026-09-28 | |
Standard Input $0.3USD / 1M tokensOutput $1.2USD / 1M tokens Fireworks AI · StandardInput: $0.3 Cache Read: $0.006 Output: $1.2 Public serverless pricing. Model display names are retained until endpoint mapping is confirmed. Checked 2026-10-03T14:45:20.473Z Price source ↗Priority Tariff not listedInput $0.375USD / 1M tokensOutput $1.5USD / 1M tokens Fireworks AI · PriorityInput: $0.375 Cache Read: $0.0075 Output: $1.5 Public serverless pricing. Model display names are retained until endpoint mapping is confirmed. Checked 2026-10-03T14:45:20.473Z Price source ↗ | Not yet captured | Catalog-listedConditionsServerless catalog listing verified; provider-specific API ID not yet captured. | Not supplied | Official source ↗2026-09-28 | |
Router catalog Tariff not listedInput $0.3USD / 1M tokensOutput $1.2USD / 1M tokens OpenRouter · Router catalogInput: $0.3 Output: $1.2 Cache Read: $0.006 OpenRouter route catalog price, not a direct-provider price. Routing, modality and endpoint conditions can change the actual bill. Checked 2026-10-03T14:45:20.473Z Price source ↗ | deepseek/deepseek-v4.1-flash | Catalog-listedConditionsCatalog-listed; account access and regional availability may apply. | Context tokens 1048576Output tokens 384000 | Official source ↗2026-09-28 | |
Standard Tariff not listedInput $0.3USD / 1M tokensOutput $1.2USD / 1M tokens Together AI · StandardInput: $0.3 Cache Read: $0.006 Output: $1.2 Serverless token list rate; verify exact deployment and quantization. Checked 2026-10-03T14:45:20.473Z Price source ↗ | deepseek-ai/DeepSeek-V4.1-Flash | Catalog-listedConditionsCatalog-listed; account access and regional availability may apply. | Context tokens 1000000 | Official source ↗2026-09-28 |
DeepSeek V4.1 Flash accepts text, image and produces text. Tool calling and structured outputs are available on supported routes.
Specifications
- Catalog model ID
- deepseek/deepseek-v4.1-flash
- Developer model ID / revision
- Not verified in cited sources
- Input modalities
- text, imageOpenRouter catalog ↗
- Output modalities
- textOpenRouter catalog ↗
- Context / input token limit
- 1,048,576OpenRouter catalog ↗
- Maximum output tokens
- 943,718OpenRouter catalog ↗
- Tool calling
- SupportedOpenRouter catalog ↗
- Structured output
- SupportedOpenRouter catalog ↗
- API parameters (source-specific)
- frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_pOpenRouter catalog ↗
- API parameter scope
- OpenRouter API; support varies by route
- Weight access
- Open weightsDeveloper model card ↗
- Weight license / access terms
- MITDeveloper model card ↗
- Total parameters (billions)
- Not verified in cited sources
- Active parameters (billions)
- Not verified in cited sources
- Parameter-count details
- 552B backbone parameters plus 196B Engram conditional memory; 8B active during prefill and 16B during decoding.Developer model card ↗
- Download size (GB)
- Depends on checkpoint and precision; no verified file size
- Release date
- Not verified in cited sources
- Added to OpenRouter
- 2026-09-10OpenRouter catalog ↗
- Notes
- API parameter names are reported by OpenRouter and do not describe the developer’s native API. Token limits labeled OpenRouter are route catalog limits, not universal model limits. Media limits on this page describe the linked provider endpoint; other deployments may differ.
Catalog entries link to their specification and availability sources. This is a curated selection, not every model each provider offers. Best prices are the lowest collected USD tariffs, not a total-cost estimate. Token offers are selected by input price, with output from the same tariff; workload mix may change the cheapest provider. Listed variants may require batch processing, off-peak hours, or specific media settings. Price coverage varies by provider; collected tariffs include their source and observation time. Benchmark scores are a dated snapshot, not a live feed or a universal measure of quality. They apply to the shown evaluation configuration; listed prices may use different settings. Models without verified scores are not assumed to be weaker.