Skip to main content
See every side of every news story
Published loading...Updated

DeepSeek Brings His Attention Cache to 890 Bytes per Token, the Cost Post Became Critical for Agents

Summary by Actu IA
DeepSeek published on September 17, 2026 the technical report of V4.1-Flash, multimodal model with a mix of 552 billion parameters served since September 10, with a context of 1 million chips. Its architecture with causative encoder-decoder activates 16 billion parameters at decoding but 8 billion at pre-filling, dominant phase of agent loads. The reuse of the key-value cache between layers and a storage in FP4 bring the overall footprint of thi…
DisclaimerThis story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.

1 Articles

DeepSeek published on September 17, 2026 the technical report of V4.1-Flash, multimodal model with a mix of 552 billion parameters served since September 10, with a context of 1 million chips. Its architecture with causative encoder-decoder activates 16 billion parameters at decoding but 8 billion at pre-filling, dominant phase of agent loads. The reuse of the key-value cache between layers and a storage in FP4 bring the overall footprint of thi…

Think freely.Subscribe and get full access to Ground NewsSubscriptions start at $9.99/yearSubscribe

Bias Distribution

  • There is no tracked Bias information for the sources covering this story.

Factuality Info Icon

To view factuality data please Upgrade to Premium

Ownership

Info Icon

To view ownership data please Upgrade to Vantage

Actu IA broke the news on Friday, September 18, 2026.
Too Big Arrow Icon
Sources are mostly out of (0)
News
Feed Dots Icon
For You
Search Icon
Search
Blindspot LogoBlindspotLocal