DeepSeek's New Model Sets a Template for Powerful LLMs that Run Lean
The model charges a fraction of a cent per million tokens and will replace V4-Pro as DeepSeek intensifies pressure on rivals.
- On Thursday, Sep 10, Chinese startup DeepSeek unveiled its V4.1 Flash model, a cost-optimized platform charging as little as a fraction of a cent per million tokens while claiming to outperform mainstays from Anthropic and Z.AI.
- DeepSeek redesigned its architecture to reduce key-value cache consumption to between 13 percent and 25 percent of prior requirements, while introducing N-gram parameters that function as implicit memory to boost intelligence without proportional memory increases.
- At 763 billion parameters, the model outperforms the V4-Pro in coding tasks, though it trails flagship models from Anthropic and OpenAI. Starting Sep 14, DeepSeek will retire the V4-Pro by automatically rerouting all inference tasks to the V4.1 Flash at cheaper rates.
- Industry competition is intensifying as Alibaba recently revealed its Qwen 3.8-Flash-Next model, which similarly employs N-gram techniques to challenge ChatGPT-developer OpenAI and other US rivals in a broadening price battle.
- By offloading N-gram weights to cheaper system RAM, the model reduces minimum GPU memory requirements from 763 GB to around 567 GB, enabling efficient deployment across high-throughput production environments without sacrificing performance.
17 Articles
17 Articles
China's DeepSeak has unveiled AI Model V4.1 Flash, which reduces costs and memory usage. By improving KV Cache and SSD efficiency, it has lowered the consumption of computational resources; this news caused global semiconductor and AI-related stocks to fall together. Experts predict that intensifying price competition in the AI industry due to technological standardization will threaten corporate survival.
DeepSeek has introduced its new V4.1 Flash model with 763 billion parameters. Architectural innovations and the use of N-grams are setting new standards in LLM efficiency. This article, "DeepSeek Sets Powerful and Efficient LLM Standards with Its New Model," first appeared on TechInside.
DeepSeek V4.1 Flash replaces V4-Pro for the time being and offers more performance at lower costs, but with unusually high token consumption.
DeepSeek returns with DeepSeek V4.1-Flash, a new multimodal model designed for IA agents, software development and very long contexts. Despite its suffix "Flash", it is not a small model in the classical sense: its architecture totals 552 billion parameters, but only activates a fraction at each stage in order to significantly reduce the [...] Read more DeepSeek V4.1-Flash: The new open source model wants to beat V4-Pro in code while at the same…
Coverage Details
Bias Distribution
- 67% of the sources lean Right
Factuality
To view factuality data please Upgrade to Premium
















