DeepSeek V4 Pro (Non-reasoning)
DeepSeek-V4-Pro (0424) is the initial preview version of DeepSeek's flagship fourth-generation large language model, released on April 24, 2026. Built on a Mixture-of-Experts (MoE) architecture, the model contains a total of 1.6 trillion parameters, with 49 billion parameters activated per token during inference. It is designed to compete with frontier closed-source models in complex reasoning, software engineering, and multi-step agentic tasks while maintaining the efficiency and accessibility of open-weight systems.
Architecture and Innovation
The model introduces several structural advancements intended to handle ultra-long contexts and improve training stability. A core innovation is the Hybrid Attention Architecture, which integrates Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). This mechanism significantly reduces the compute and memory requirements for its 1-million-token context window; at maximum context, the V4-Pro requires only 10% of the KV cache and 27% of the inference FLOPs compared to the previous V3 generation. The architecture also utilizes Manifold-Constrained Hyper-Connections (mHC) to stabilize signal propagation across deep layers and the Muon optimizer for faster convergence during its pre-training on 32 trillion tokens.
Capabilities and Performance
DeepSeek-V4-Pro is optimized for "agentic" workflows, showing high performance in coding benchmarks such as LiveCodeBench and SWE-bench. It features a native dual-mode system allowing for both "Thinking" and "Non-Thinking" responses, with three distinct effort levels (low, high, and max) to adapt reasoning depth to task complexity. This version led open-weights models in world knowledge and STEM reasoning at the time of its release, trailing only the most advanced closed-source systems like Gemini-3.1-Pro and GPT-5.4.
Post-training for the 0424 build involved a specialized two-stage process. The model first undergoes independent cultivation of domain-specific experts using Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO), followed by a consolidation phase using on-policy distillation to unify these proficiencies into a single cohesive model.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow DeepSeek V4 Pro (Non-reasoning) ranks
DeepSeek V4 Pro (Non-reasoning) is highlighted in the table below. Switch the metric to see how the ordering changes.