Models
DeepSeek V4 preview ships open weights with a 1M-token context and Huawei Ascend support
DeepSeek released V4-Pro (1.6T parameters, 49B active) and V4-Flash (284B, 13B active) under the MIT license, both with 1M-token context. Huawei says its chips run them.
HackHoster Team · · 9 min read

At a glance
- DeepSeek published open weights for V4-Pro (1.6T parameters, 49B active) and V4-Flash (284B, 13B active) on April 24, 2026 under the MIT license.
- Both models take a 1-million-token context, and DeepSeek made 1M the default across its own services, up from 128K for V3.
- The model card says V4-Pro needs 27% of the per-token inference compute and 10% of the KV cache of DeepSeek-V3.2 at 1M tokens.
- Huawei says its whole Ascend supernode line supports V4, and Reuters reports Huawei chips trained part of V4-Flash.
- Old API names deepseek-chat and deepseek-reasoner now route to V4-Flash and stop working after July 24, 2026.
DeepSeek released a preview of its V4 model family on April 24, with open weights on Hugging Face under the MIT license. There are two mixture-of-experts models:
- DeepSeek-V4-Pro: 1.6 trillion total parameters, 49 billion active per token
- DeepSeek-V4-Flash: 284 billion total parameters, 13 billion active per token
Both support a 1-million-token context window, and DeepSeek says 1M is now the default context length across its official services, up from the 128,000 tokens of the V3 generation. Base checkpoints are published next to the post-trained models, which gives you a clean starting point for your own fine-tuning.
On the same day, Huawei said its entire Ascend supernode product line supports the V4 series. According to Reuters, Huawei's chips were also used for part of V4-Flash's training. DeepSeek has not fully specified what hardware trained V4, and its earlier V3 and R1 models were trained on Nvidia chips.
How DeepSeek got here
DeepSeek is unusual among frontier labs in that it grew out of a quantitative hedge fund. High-Flyer, co-founded by Liang Wenfeng in 2015, announced an AI research lab in April 2023 and spun it off as DeepSeek that July, remaining its principal backer. Liang had begun buying Nvidia GPUs years before, reportedly accumulating around 10,000 A100 cards before US export controls tightened, a stockpile few Chinese companies matched.
The releases that followed set the pattern. DeepSeek-V2, in May 2024, triggered a price war in the Chinese model market. V3 arrived in December 2024. Then R1, the reasoning model, launched on mobile in January 2025 and reached the top of the US App Store's free chart; on January 27, 2025, DeepSeek's rise was blamed for an 18% drop in Nvidia's share price. Since January 2025 the company has shipped its models under permissive open-source licenses, usually MIT, which is why V4's weights are downloadable at all.

That history shapes how V4 lands. Tech Xplore, reporting the launch, quoted Omdia chief analyst Lian Jye Su saying the benchmark results make V4 look competitive against US rivals, and Marina Zhang of the University of Technology Sydney calling the rollout a milestone for China's AI industry. Morningstar's Ivan Su was more measured, describing V4 as a competent follow-up but not the kind of shock R1 was. The same report notes that Anthropic and OpenAI have accused DeepSeek of distilling their models, an accusation DeepSeek has not conceded.
What a million tokens actually costs
The headline number is the context window, and the interesting engineering is in making it affordable rather than merely possible.
Why the KV cache matters. When a transformer generates text, it stores the key and value vectors for every token it has already seen so it doesn't have to recompute them. That store, the KV cache, grows with the length of the conversation and must sit in accelerator memory for each active request. On long contexts it, not the weights, usually decides how many sessions fit on a machine.

The main architectural change in V4 is a hybrid attention design that combines what DeepSeek calls Compressed Sparse Attention and Heavily Compressed Attention. The model card says that in the 1M-token setting, V4-Pro needs 27% of the per-token inference compute and 10% of the KV cache of DeepSeek-V3.2. A tenfold cut in cache changes serving math directly: the same memory holds ten times as many long-context sessions.
The mixture-of-experts structure handles the other half of the cost. The idea goes back to the 1990s as an ensemble method, dividing a problem space among specialized learners with a gating function to weight them. The modern version, the sparsely-gated layer introduced by Shazeer and colleagues in 2017, keeps only the top few experts for each input and zeroes the rest, so the model's total parameter count can be very large while the compute per token stays small. Training such a layer needs an auxiliary loss to stop the gate from favoring a handful of experts and leaving the others untrained. That is why V4-Pro lists 1.6 trillion total parameters but only 49 billion active, and Flash 284 billion against 13 billion active. The trade-off is memory: every expert has to be loaded even though most sit idle for any given token.

Other details from the model card:
- Precision: the weights ship in mixed precision, with MoE expert parameters in FP4 and most other parameters in FP8. The base checkpoints use FP8 mixed precision.
- Training: more than 32 trillion tokens of pretraining, using the Muon optimizer and a technique DeepSeek calls manifold-constrained hyper-connections, which strengthens the model's residual connections while keeping expressivity.
- Post-training: DeepSeek first trains separate domain experts with supervised fine-tuning and reinforcement learning using GRPO, then merges their skills into one model through on-policy distillation.


The base models tell their own story
Because DeepSeek publishes base checkpoints, the model card also compares them directly with V3.2-Base, before any instruction tuning. Two numbers stand out. On SimpleQA Verified, a factual-recall test, V3.2-Base scores 28.3 and V4-Pro-Base 55.2. On FACTS Parametric, which probes what the model knows without retrieval, the jump is 27.1 to 62.6. MultiLoKo, a multilingual set, goes from 38.7 to 51.1, and LongBench-V2 from 40.2 to 51.5.
| Base model benchmark | V3.2-Base | V4-Flash-Base | V4-Pro-Base |
|---|---|---|---|
| MMLU-Pro | 65.5 | 68.3 | 73.5 |
| SimpleQA Verified | 28.3 | 30.1 | 55.2 |
| FACTS Parametric | 27.1 | 33.9 | 62.6 |
| LongBench-V2 | 40.2 | 44.7 | 51.5 |
| HumanEval | 62.8 | 69.5 | 76.8 |
| BigCodeBench | 63.9 | 56.8 | 59.2 |
Those gains concentrate in Pro rather than Flash, which is what you would expect if they come from the larger parameter count soaking up more of 32 trillion tokens. Note also that both V4 base models score below V3.2-Base on BigCodeBench, a reminder that a new generation is not uniformly better at everything.
Three reasoning modes, and what they buy
V4 has three reasoning modes: non-thinking, Think High and Think Max. For Think Max, DeepSeek recommends setting the context window to at least 384K tokens, a sign of how long its reasoning can run.
The model card's own mode comparison is the clearest evidence of what thinking is worth. On Humanity's Last Exam, V4-Pro goes from 7.7 without thinking to 34.5 at Think High and 37.7 at Think Max. On LiveCodeBench, 56.8 to 89.8 to 93.5. On the Apex Shortlist, 9.2 to 85.5 to 90.2. Agentic coding moves far less: SWE-bench Verified goes 73.6, 79.4, 80.6. In other words, thinking transforms hard reasoning tasks and barely shifts tasks that are mostly about navigating a codebase.
| Benchmark | V4-Pro non-thinking | V4-Pro Think High | V4-Pro Think Max |
|---|---|---|---|
| Humanity's Last Exam | 7.7 | 34.5 | 37.7 |
| LiveCodeBench | 56.8 | 89.8 | 93.5 |
| GPQA Diamond | 72.9 | 89.1 | 90.1 |
| SWE-bench Verified | 73.6 | 79.4 | 80.6 |
| Terminal-Bench 2.0 | 59.1 | 63.3 | 67.9 |
| MRCR at 1M tokens | 44.7 | 83.3 | 83.5 |
How it compares
The model card compares V4-Pro in Think Max against closed models at high reasoning settings. These are DeepSeek's own figures, run by DeepSeek.
| Benchmark | V4-Pro Max | Claude Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro |
|---|---|---|---|---|
| SWE-bench Verified | 80.6 | 80.8 | — | 80.6 |
| LiveCodeBench | 93.5 | 88.8 | — | 91.7 |
| Codeforces (rating) | 3206 | — | 3168 | 3052 |
| Terminal-Bench 2.0 | 67.9 | 65.4 | 75.1 | 68.5 |
| GPQA Diamond | 90.1 | 91.3 | 93.0 | 94.3 |
| Humanity's Last Exam | 37.7 | 40.0 | 39.8 | 44.4 |
| SimpleQA Verified | 57.9 | 46.2 | 45.3 | 75.6 |
| MRCR at 1M tokens | 83.5 | 92.9 | — | 76.3 |
The picture is mixed in a specific way. V4-Pro leads on competitive programming, holds level on agentic coding, and trails on general knowledge and the hardest reasoning sets. Reuters, citing DeepSeek's paper, reports that V4-Pro beats all open models in maximum reasoning mode but still trails frontier closed systems such as Gemini 3.1 Pro and GPT-5.4 in some areas. Tech Xplore reports DeepSeek's own framing: ahead of GPT-5.2 and Gemini 3.0 Pro, marginally behind GPT-5.4 and Gemini 3.1 Pro.

Flash is the version most teams will actually run. Its non-thinking mode scores 73.7 on SWE-bench Verified, within a point of Pro's 73.6, and its Think Max mode reaches 79.0 against Pro's 80.6. The gap widens on knowledge-heavy tests: SimpleQA Verified is 34.1 for Flash at Think Max against 57.9 for Pro, which is what you would expect from a model with roughly a fifth of the parameters.
Migrating, if you call the hosted API
If you call DeepSeek's hosted API, there is a migration to plan. The new model IDs are deepseek-v4-pro and deepseek-v4-flash, both with thinking and non-thinking modes. DeepSeek says you can keep your base URL and just change the model name, and that the API accepts both OpenAI-style chat completions and Anthropic-style requests, with the announcement naming agent tools such as Claude Code, OpenClaw and OpenCode as compatible clients.
The deadline is the part to diary. DeepSeek says the old deepseek-chat and deepseek-reasoner names, which now route to V4-Flash in non-thinking and thinking modes respectively, will stop working after July 24, 2026 at 15:59 UTC. Any project with those strings hardcoded needs an update before then.
Reuters reports that Pro costs up to 12 times as much as Flash because of tight high-end compute capacity, and that prices are expected to fall sharply once Huawei's Ascend 950 supernodes are deployed at scale in the second half of the year.
Running the open weights
DeepSeek recommends Transformers, vLLM or SGLang for local deployment, with temperature and top_p both set to 1.0. One detail will trip up existing tooling: the release does not include a Jinja chat template. Instead DeepSeek ships an encoding folder with Python scripts and test cases showing how to turn OpenAI-format messages into the input string the model expects, including the thinking modes. Frameworks that rely on the tokenizer's built-in template will need adapting.

Be realistic about hardware. Even with FP4 experts, Flash's 284 billion parameters need multiple data-center GPUs, and Pro at 1.6 trillion is a cluster job. For most individual developers, Flash through an API or a hosted provider is the practical option, and its 13 billion active parameters help keep it cheaper to serve. The MIT license allows commercial use and fine-tuning without the custom restrictions that come with some other open-weight models, and because the base checkpoints ship alongside the post-trained ones, you can start fine-tuning from a model that hasn't already been shaped by DeepSeek's own preference data.
Practical tip. Before planning around the 1M window, measure it on your own data. The model card's own long-context scores, 83.5 on MRCR at 1M and 62.0 on CorpusQA at 1M, are well below its short-context results, so retrieval over a long prompt is not free accuracy.
If you are evaluating V4 for a project, a sensible order is: run Flash in non-thinking mode first to set a latency and cost baseline; turn on thinking only for the tasks where the mode table shows a real gain; and test long-context behavior with documents from your own domain rather than trusting the benchmark.
A few other things worth checking before committing:
- The template gap is real work. If your stack assumes
tokenizer.apply_chat_template, budget time to port DeepSeek's encoding scripts, and write tests against the cases they ship. - Non-thinking mode is a different model in practice. The 7.7-to-37.7 spread on Humanity's Last Exam means a prompt tuned in one mode tells you little about the other.
- Cheap context is not free context. The attention and cache savings cut what a long prompt costs to serve; they do not make the model better at using it.
- Keep a path off the hosted API. Since the weights are MIT-licensed and the base checkpoints are published, a project that starts on DeepSeek's servers can move to self-hosting or another provider without a rewrite.
The caveats
This is a preview, and DeepSeek may change the models before a final release. Every benchmark number above is DeepSeek's own and had no independent reproduction at launch, which matters more than usual given the spread between modes and the number of competitor scores quoted without a stated harness.
The Huawei claim is narrower than the headlines suggest. Reuters reports that Huawei chips were used for part of V4-Flash's training and that the Ascend supernode line supports the V4 series for inference. That is not the same as saying the family was trained end to end on domestic hardware, and DeepSeek has not said what else was involved.
There are policy considerations too. If you use DeepSeek's hosted API, your prompts go to servers run by a Chinese company, which matters for some data policies; self-hosting the weights avoids that. And the distillation accusations from Anthropic and OpenAI, reported by Tech Xplore, remain unresolved, which is worth knowing if your use of the model carries licensing scrutiny.
What to watch
As of April 25, the open questions are when the preview becomes a final release, whether independent evaluations reproduce the benchmark numbers, and whether Pro's price falls as Reuters' sources expect once Ascend 950 supernodes ship in the second half of the year. The July 24 deprecation of the old API names is the one hard date on the calendar.
Sources
- DeepSeek-V4-Pro model card (Hugging Face)
- DeepSeek V4 preview release (DeepSeek API docs)
- Factbox, DeepSeek-V4, the Chinese AI model adapted for Huawei chips (Reuters via Investing.com)
- DeepSeek rolls out V4 update with 1 million-token context and stronger reasoning (Tech Xplore)
- DeepSeek (Wikipedia)
- Mixture of experts (Wikipedia)
More from the blog

Models ·
OpenAI begins a staged rollout of GPT-6 Astra, starting with its security program
GPT-6 Astra went to OpenAI's Daybreak security customers first. It is OpenAI's first model rated Critical for cyber capability, and its reasoning is harder to monitor.
12 min read

Models ·
Z.ai reveals stealth model Ox Alpha as GLM-5.3-Flash and releases MIT-licensed weights
The 320B-parameter mixture-of-experts model activates 18B per token and mixes linear with sparse attention. How it works, what it takes to run, and what to check first.
11 min read

Models ·
NVIDIA releases Nemotron 3 Ultra, a 550B open-weights model built for long-running agents
NVIDIA published weights, data and recipes for a 550-billion-parameter hybrid Mamba-attention model. It computes like a 55B model but needs memory for all 550B.
11 min read