DeepSeek V4 Pro 🧠 Model Open source
DeepSeek's MIT-licensed V4 models with 1M-token context and very low API prices
- Maker
- DeepSeek
- Latest release
- 2026-08
- Weights
- Open
- Licence
- MIT
- Apps offering it
- 2
Where you can use DeepSeek V4 Pro
Apps in this directory that let you generate with DeepSeek V4 Pro, from their own documentation.
Good for
About DeepSeek V4 Pro
What it is
DeepSeek is a Chinese AI lab that publishes open-weight language models. Its current generation, DeepSeek V4, launched in April 2026 in two sizes: V4 Pro (1.6T parameters, 49B active) and V4 Flash (284B, 13B active). DeepSeek V4 Pro left preview and became generally available on the API and in the chat app on 13 August 2026. Both models have a 1M-token context window.
Key features
- 1M-token context and up to 384K output tokens per response
- Reasoning effort setting (low, high, max) so simple rewrites stay cheap
- OpenAI-compatible API, including the Responses API
- Flash supports image input, JSON output and tool calls; V4 Pro is text-only
- Weights on Hugging Face under MIT, so you can fine-tune on your own style
Where you can use it
In DeepSeek's free chat app (web and mobile, with Expert Mode on V4 Pro) and the DeepSeek API. The weights run on vLLM or SGLang and through many hosting providers. For content automation, n8n has a DeepSeek Chat Model node.
Pricing and rights
Since 16 August 2026 the API uses off-peak and peak rates; peak (weekday mornings UTC) costs double. Off-peak list prices per million tokens: Flash input from $0.003 on cache hits and $0.60 output; V4 Pro input from $0.022 on cache hits and $1.98 output. The MIT licence allows commercial use of the weights and outputs. The chat app is free, so there is no paid consumer plan.
Who it is for
Budget-conscious publishers running high-volume drafting, rewriting or translation, and teams who want open weights they can host themselves.
Verdict
DeepSeek V4 offers frontier-class context at a fraction of Western API prices, with a truly permissive licence. The flagship is text-only, peak pricing complicates budgeting, and some clients will object to sending data to a China-based service.
Pros
- Open weights under the MIT licence, commercial use allowed
- 1M-token context and up to 384K output tokens
- Off-peak API output prices of about $0.60 (Flash) and $1.98 (V4 Pro) per million tokens
- Free chat app with an Expert Mode on V4 Pro
Cons
- V4 Pro has no vision input; only the Flash model reads images
- Peak-hour API prices are double the off-peak rates
- Self-hosting V4 Pro (1.6T parameters) needs data-centre hardware
- Data processed on the hosted service is subject to Chinese jurisdiction, a concern for some clients
Similar AI models
All language models →Claude Fable 5.1 🧠 ModelFreemium
Anthropic's Claude models (Fable 5.1, Opus 5.5, Sonnet 5, Haiku 4.5) with 1M-token context
GPT-6 Astra 🧠 ModelFreemium
OpenAI's GPT-6 family (Astra, Sol, Luna) behind ChatGPT and the OpenAI API
Gemini 3.1 Pro 🧠 ModelFreemium
Google's Gemini models (3.1 Pro, 3.8 Flash) in the Gemini app, NotebookLM and the Gemini API
Qwen3.8 🧠 ModelOpen source
Alibaba's Qwen3.8 family, from Apache-2.0 open weights (27B to 2.4T) to the hosted Qwen3.8-Max
Grok 4.7 🧠 ModelFreemium
xAI's Grok models, built into Grok and X, with a 500K-token context on Grok 4.7
Mistral Medium 3.5 🧠 ModelOpen source
Mistral AI's European models: open-weight Medium 3.5, Large 3 and Small 4, used in Mistral Vibe