DeepSeek has officially unveiled its latest flagship model, DeepSeek V4, marking a significant leap in AI efficiency and capability. The model features a massive 1.6 trillion total parameters with a Mixture-of-Experts (MoE) architecture, activating only 49 billion parameters per token, ensuring high performance while maintaining computational efficiency.
One of the standout features is the support for a 1 million token context window, achieved through DeepSeek's proprietary DSA (Dynamic Sparse Attention) combined with Hybrid Attention mechanisms. This allows the model to process extremely long documents, codebases, or conversations without losing coherence.
DeepSeek V4 also introduces the Muon optimizer, a novel optimization algorithm designed to improve training stability and convergence speed. Additionally, the model comes with native agent integrations, including compatibility with Claude Code and OpenCode, enabling developers to build autonomous coding agents directly on top of the model.
Alongside the full V4 model, DeepSeek released V4-Flash, a distilled variant with 284 billion total parameters and 13 billion active parameters, offering a more lightweight option for deployment. The company has also launched a new API for both models, providing developers with easy access to the latest capabilities.
DeepSeek V4 is now available for preview, with the company promising further details and benchmarks in the coming weeks.