• bitcoinBitcoin(BTC)$77,999.00-1.49%
  • ethereumEthereum(ETH)$2,466.97-1.15%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$716.28-5.04%
  • rippleXRP(XRP)$1.38-3.37%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$100.93-2.80%
  • tronTRON(TRX)$0.3401460.30%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,219.79-1.54%
  • HyperliquidHyperliquid(HYPE)$83.07-3.56%
  • dogecoinDogecoin(DOGE)$0.085006-6.05%
  • RainRain(RAIN)$0.0162361.08%
  • USDSUSDS(USDS)$1.00-0.02%
  • moneroMonero(XMR)$510.382.80%
  • whitebitWhiteBIT Coin(WBT)$80.60-1.51%
  • chainlinkChainlink(LINK)$11.79-4.94%
  • leo-tokenLEO Token(LEO)$9.190.09%
  • cardanoCardano(ADA)$0.212636-2.95%
  • stellarStellar(XLM)$0.179386-4.75%
  • bitcoin-cashBitcoin Cash(BCH)$247.32-4.33%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$52.30-3.64%
  • CantonCanton(CC)$0.102256-4.66%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-2.49%
  • uniswapUniswap(UNI)$6.00-10.44%
  • avalanche-2Avalanche(AVAX)$7.74-2.86%
  • hedera-hashgraphHedera(HBAR)$0.076101-3.90%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$2.40-1.57%
  • suiSui(SUI)$0.76-6.41%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.88%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056825-5.18%
  • MemeCoreMemeCore(M)$1.211.57%
  • tether-goldTether Gold(XAUT)$4,387.98-0.15%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • BittensorBittensor(TAO)$252.17-2.75%
  • okbOKB(OKB)$112.45-1.75%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.02%
  • mantleMantle(MNT)$0.59-7.71%
  • AsterAster(ASTER)$0.71-5.70%
  • aaveAave(AAVE)$123.92-3.87%
  • pax-goldPAX Gold(PAXG)$4,391.58-0.17%
  • polkadotPolkadot(DOT)$1.10-7.13%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0559430.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

September 10, 2026
in AI & Technology
Reading Time: 15 mins read
A A
DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
ShareShareShareShareShare

Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. DeepSeek AI built its newest release around that exact bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters, 196B additional Engram parameters, and a 1M-token context window. It activates 8B parameters per token during prefill and 16B during decode. The main number is a global KV cache footprint of 890 bytes per token, about 1/4 of DeepSeek-V4-Flash and roughly 437x smaller than DeepSeek-V1.

Is it deployable? Yes. Open weights ship under an MIT license with vLLM, SGLang, and Transformers paths on Hugging Face, and the research team describes a public API with low, high, and max reasoning tiers.

YOU MAY ALSO LIKE

2028 Volvo XC40 First Look: Hello new tech, goodbye EV

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

Causal Encoder-Decoder: Half the Prefill

The 40-layer backbone is split into a 20-layer causal encoder and a 20-layer decoder. Inspired by YOCO, the decoder does not compute its own global KV. Instead, per-layer projection weights derive it from the final encoder hidden state. Prompt tokens therefore stop at the encoder, which nearly halves prefill compute. Sliding-window attention (SWA) with a 128-token window still runs in every layer, so decoder SWA states are rebuilt by replaying only the last 128 prompt tokens. The research team calls this Decoder SWA Bounded Replay.

Compressed Sparse Attention 2 (CSA2)

DeepSeek-V4 mixed CSA with Heavily Compressed Attention. V4.1-Flash uses pure CSA2 and attacks cache size along the layer axis. Each CSA2 layer is statically assigned one of 3 modes:

  • Full: computes its own main KV, projects indexer K from it, and selects fresh Top-512 indices.
  • Reindex: reuses main KV and indexer K from the last Full layer but rescores them with its own indexer Q.
  • Reuse: reuses both the main KV and the latest Top-K indices, skipping the indexer entirely.

Every layer keeps its own main Q and SWA KV. The 18 CSA2 encoder layers use a compression ratio of 2 in 3 groups of 6 (1 Full, 5 Reuse). The 20 decoder layers use ratio 1 in 5 groups of 4: the first is Full plus 3 Reuse, the rest Reindex plus 3 Reuse. A Hierarchical Sparse Indexer in the decoder lets the Full layer build a candidate pool of up to 16,384 positions (2,048 blocks of 8), so later Reindex layers score a bounded set instead of the entire context.

FP4 KV, Bounded Replay, and Other Extensions

The main KV cache is quantized to E2M1 with one E4M3 scale per 16 channels, following NVFP4 without its global scale. This is introduced through quantization-aware training in post-training and nearly halves storage against V4’s FP8 cache.

At the deployment level, SWA KV is no longer persisted to SSD. It lives in a distributed pool carved from 10% of host DRAM with a TTL of minutes, while global KV keeps a guaranteed 72-hour lifetime. On a miss, Encoder SWA Bounded Replay recomputes only 128 tokens instead of layers times window.

Other changes include Single-Pass mHC, which shifts input-mixing coefficients by one block so a fused Mega-mHC kernel can halve activation memory traffic, the Engram conditional memory module at layers 1 and 14, DSpark speculative decoding trained after pre-training with the backbone frozen, and head-wise Muon. Single-token decode FLOPs rise by only 1/4 when context grows from 4K to 1M.

Training and Results

Pre-training covers 45T multimodal tokens at a 7:1 text-to-multimodal ratio. Sparse attention is trained from scratch at 64K sequence length with no dense warmup, and context is extended to 1M at 34T tokens. The base model matches DeepSeek-V4-Pro-Base on world knowledge and coding while using 1/3 of the total and 1/4 of the activated parameters.

Post-training introduces no new algorithms. Gains come from large-scale synthesis of verifiable agent tasks, RL across heterogeneous scaffolds (Claude Code, Codex, OpenCode, Pi, mini-SWE, DeepSeek Harness), and on-policy distillation from over 40 teachers. Selected max-effort results:

Benchmark DS-V4.1-Flash DS-V4-Flash Opus-5 GPT-5.6 Sol
Terminal-Bench 2.1 90.6 82.7 89.1 88.8
DeepSWE v1.1 74.2 54.4 74.0 73.0
Terminal-Bench 4.0 31.2 7.0 51.8 39.9
Automation-Bench 54.8 37.7 50.3 45.8
GPQA Diamond 90.9 89.9 93.4 94.1
Codeforces (rating) 3471 3289 n/a n/a

Interactive Explainer

Key Takeaways

  • Global KV cache falls to 890 bytes per token, about 1/4 of V4-Flash and 437x below V1.
  • CED runs only 20 encoder layers in prefill, activating 8B parameters against 16B in decode.
  • CSA2 shares main KV, indexer K, and Top-K indices across layers in Full, Reindex, and Reuse modes.
  • FP4 main KV plus SWA Bounded Replay cut persistent cache to about 1/8 of V4-Flash.
  • Beats Opus-5 and GPT-5.6 Sol on Terminal-Bench 2.1 and DeepSWE v1.1 with MIT weights.

Check out the Model on Hugging Face and the Technical Report. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

2028 Volvo XC40 First Look: Hello new tech, goodbye EV
AI & Technology

2028 Volvo XC40 First Look: Hello new tech, goodbye EV

September 10, 2026
Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI
AI & Technology

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

September 10, 2026
LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Next Post
The Worst Isn't Over For Camping World Holdings

The Worst Isn't Over For Camping World Holdings

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
A dog missing for over 5 years reunited with his family

A dog missing for over 5 years reunited with his family

September 4, 2026
2028 Volvo XC40 First Look: Hello new tech, goodbye EV

2028 Volvo XC40 First Look: Hello new tech, goodbye EV

September 10, 2026
China-linked hackers backdoored executives’ laptops via USB, exploiting a fix companies had but weren’t using

China-linked hackers backdoored executives’ laptops via USB, exploiting a fix companies had but weren’t using

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!