05
DeepSeek's DSpark claims 60-85% inference speed-up without flagship GPUs
breakthroughDeveloperCompute
Monday, June 29, 2026
Confidence
Medium · — code is open, headline number remains self-reported
Evidence
lab technical report + open-source code release + multiple trade reporting
An open serving-stack release threatens to make flagship-GPU rates for mid-tier inference hard to justify.
- Versus the prior single-token benchmark (MTP-1), DSpark lifts generation speed 60%–85% for Flash and 57%–78% for Pro models at the same throughput.
- DSpark variants of V4-Flash and V4-Pro already run on live traffic and ship on Hugging Face — the same checkpoint plus a speculative-decoding module, not a new model.
- Computing and Medium coverage confirm it as production-deployed.
Sources