Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper DeepSeek

DeepSeek's DSpark Boosts LLM Inference Speed by Up to 85% with Speculative Decoding

AI By Crimson AI DeepSeek Blog 26 July 2026 · 15:32 22 views
Share: X Telegram

DeepSeek introduces DSpark, a speculative decoding system that accelerates LLM inference by 57–85% on V4, Qwen, and Gemma models without retraining.

Key points

DeepSeek has unveiled DSpark, a speculative decoding framework that significantly speeds up large language model (LLM) inference without requiring any retraining. The system achieves byte-identical outputs while delivering per-user speed improvements of 57–85% on models such as V4, Qwen, and Gemma.

Speculative decoding works by using a small, fast draft model to generate candidate tokens, which are then verified by the larger target model. This approach reduces the number of sequential calls to the large model, cutting latency while preserving output quality. DSpark builds on this concept with optimizations tailored for real-world deployment.

DeepSeek's benchmarks show consistent gains across different model architectures, making DSpark a practical solution for reducing inference costs and improving user experience. The company emphasizes that the method requires no changes to the original model weights or training pipeline.

ModelSpeed Improvement
V457–85%
Qwen57–85%
Gemma57–85%
Source
DeepSeek · DeepSeek Blog
Related news
DeepSeek
DeepSeek 10 Aug 2026

DeepSeek V4 Launches with 1.6T Parameters, Claims 10-50x Cheaper Pricing

DeepSeek unveiled its V4 model family on April 24, 2026, featuring two text-only variants with up to 1.6 trillion parameters and a...

20
DeepSeek
DeepSeek 4 Aug 2026

DeepSeek Retires Legacy API Aliases: Price Changes and Migration Guide

DeepSeek retired the legacy API aliases 'deepseek-chat' and 'deepseek-reasoner' on July 24, 2026, at 15:59 UTC. Developers must up...

33
DeepSeek
DeepSeek 2 Aug 2026

DeepSeek-V4-Flash Goes Official: Agent Benchmarks Surpass V4-Pro-Preview

DeepSeek has launched the official deepseek-v4-flash API in public beta, featuring improved agent capabilities that outperform V4-...

34