← model directory

deepseek

DeepSeek: DeepSeek V4 Flash 0731

deepseek/deepseek-v4-flash-0731

⚙ tool calling

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

specs & pricing
Type
text
Provider
deepseek
Model ID
deepseek/deepseek-v4-flash-0731
Capabilities
tools, reasoning
Context window
1.3M tokens
Input price
$0.066 / 1M tokens
Output price
$0.13 / 1M tokens

Cloud price is billed from prepaid credits when a request fails over to the cloud. Open-weight models run free on GPUs you own — the gateway routes to your nodes first.

What it costs to run

cloud · 1M in + 1M out

$0.20

$0.066 for a million input tokens plus $0.13 for a million output tokens, billed from credits only when a request fails over to the cloud.

your hardware

cloud only

DeepSeek: DeepSeek V4 Flash 0731 is a hosted model, so it is always served from the cloud. Route easy traffic to an open model on your own hardware and keep this one for the requests that need it.

/// initialize

Call DeepSeek: DeepSeek V4 Flash 0731 through one OpenAI-compatible endpoint.

Mix it with open models on your own hardware — the gateway routes each request to the cheapest place that can serve it.

no credit card · 2 nodes free · openai-compatible

Frequently asked questions

What is the context window of DeepSeek: DeepSeek V4 Flash 0731?
DeepSeek: DeepSeek V4 Flash 0731 accepts up to 1,310,720 tokens (1.3M) of context per request.
How much does DeepSeek: DeepSeek V4 Flash 0731 cost per million tokens?
Through Wide Area Intelligence, DeepSeek: DeepSeek V4 Flash 0731 costs $0.066 per 1M input tokens and $0.13 per 1M output tokens when served from the cloud, billed from prepaid credits. A workload of 1M input plus 1M output tokens costs $0.20.

More from deepseek

← browse all models