Advertisement
Faster AI, lower costs: DSpark eases inference bottlenecks and chip strain, says DeepSeek
Start-up unveils speculative decoding framework that speeds up inference by up to 85 per cent amid China’s push to overcome US AI curbs
2-MIN READ2-MIN
Listen

Ben Jiangin Beijing
Chinese artificial intelligence start-up DeepSeek has rolled out a major upgrade to its flagship V4 model aimed at sharply accelerating AI response generation, as competition among Chinese developers increasingly shifts to reducing serving costs and enhancing user experience.
DeepSeek, by adopting what it called a speculative decoding framework, DSpark, said it increased per-user response speeds by up to 85 per cent, an efficiency gain that could reduce AI systems’ reliance on larger, more powerful chip infrastructure.
AI models’ conventional token-by-token output often slowed when responses were lengthy, leading to low utilisation of graphics processing units (GPU) and high user-perceived waiting time, which was a “primary bottleneck in serving AI”, the company said in research published on Saturday.
Select Voice
Select Speed
1x
AI-generated voice