Skip to main content
See every side of every news story
Published loading...Updated

DeepSeek's New Model Sets a Template for Powerful LLMs that Run Lean

The model charges a fraction of a cent per million tokens and will replace V4-Pro as DeepSeek intensifies pressure on rivals.

  • On Thursday, Sep 10, Chinese startup DeepSeek unveiled its V4.1 Flash model, a cost-optimized platform charging as little as a fraction of a cent per million tokens while claiming to outperform mainstays from Anthropic and Z.AI.
  • DeepSeek redesigned its architecture to reduce key-value cache consumption to between 13 percent and 25 percent of prior requirements, while introducing N-gram parameters that function as implicit memory to boost intelligence without proportional memory increases.
  • At 763 billion parameters, the model outperforms the V4-Pro in coding tasks, though it trails flagship models from Anthropic and OpenAI. Starting Sep 14, DeepSeek will retire the V4-Pro by automatically rerouting all inference tasks to the V4.1 Flash at cheaper rates.
  • Industry competition is intensifying as Alibaba recently revealed its Qwen 3.8-Flash-Next model, which similarly employs N-gram techniques to challenge ChatGPT-developer OpenAI and other US rivals in a broadening price battle.
  • By offloading N-gram weights to cheaper system RAM, the model reduces minimum GPU memory requirements from 763 GB to around 567 GB, enabling efficient deployment across high-throughput production environments without sacrificing performance.
Insights by Ground AI

17 Articles

Lean Right

China's DeepSeak has unveiled AI Model V4.1 Flash, which reduces costs and memory usage. By improving KV Cache and SSD efficiency, it has lowered the consumption of computational resources; this news caused global semiconductor and AI-related stocks to fall together. Experts predict that intensifying price competition in the AI industry due to technological standardization will threaten corporate survival.

DeepSeek has introduced its new V4.1 Flash model with 763 billion parameters. Architectural innovations and the use of N-grams are setting new standards in LLM efficiency. This article, "DeepSeek Sets Powerful and Efficient LLM Standards with Its New Model," first appeared on TechInside.

DeepSeek V4.1 Flash replaces V4-Pro for the time being and offers more performance at lower costs, but with unusually high token consumption.

·Germany
Read Full Article

DeepSeek returns with DeepSeek V4.1-Flash, a new multimodal model designed for IA agents, software development and very long contexts. Despite its suffix "Flash", it is not a small model in the classical sense: its architecture totals 552 billion parameters, but only activates a fraction at each stage in order to significantly reduce the [...] Read more DeepSeek V4.1-Flash: The new open source model wants to beat V4-Pro in code while at the same…

Think freely.Subscribe and get full access to Ground NewsSubscriptions start at $9.99/yearSubscribe

Bias Distribution

  • 67% of the sources lean Right
67% Right

Factuality Info Icon

To view factuality data please Upgrade to Premium

Ownership

Info Icon

To view ownership data please Upgrade to Vantage

Pandaily broke the news on Thursday, September 10, 2026.
Too Big Arrow Icon
Sources are mostly out of (0)

Similar News Topics

News
Feed Dots Icon
For You
Search Icon
Search
Blindspot LogoBlindspotLocal