• bitcoinBitcoin(BTC)$76,954.00-1.27%
  • ethereumEthereum(ETH)$2,475.15-1.62%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$718.12-0.71%
  • rippleXRP(XRP)$1.400.35%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$100.82-0.98%
  • tronTRON(TRX)$0.338681-0.48%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,142.180.38%
  • HyperliquidHyperliquid(HYPE)$79.35-0.59%
  • dogecoinDogecoin(DOGE)$0.082660-1.96%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$516.970.01%
  • RainRain(RAIN)$0.013253-12.37%
  • whitebitWhiteBIT Coin(WBT)$79.60-1.39%
  • chainlinkChainlink(LINK)$11.38-0.14%
  • leo-tokenLEO Token(LEO)$8.990.32%
  • cardanoCardano(ADA)$0.205090-2.49%
  • stellarStellar(XLM)$0.1946534.34%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$222.27-0.59%
  • USD1USD1(USD1)$1.00-0.01%
  • uniswapUniswap(UNI)$6.655.57%
  • litecoinLitecoin(LTC)$52.55-2.36%
  • CantonCanton(CC)$0.095328-0.43%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.34-0.89%
  • hedera-hashgraphHedera(HBAR)$0.0773591.27%
  • avalanche-2Avalanche(AVAX)$7.531.89%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • nearNEAR Protocol(NEAR)$2.39-1.06%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.76%
  • suiSui(SUI)$0.71-1.85%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057267-3.20%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,276.66-0.48%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$226.39-4.26%
  • MemeCoreMemeCore(M)$1.11-0.89%
  • okbOKB(OKB)$113.04-0.90%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.18%
  • aaveAave(AAVE)$127.240.14%
  • BitwayBitway(BTW)$0.7111.12%
  • AsterAster(ASTER)$0.69-1.45%
  • pax-goldPAX Gold(PAXG)$4,277.70-0.53%
  • mantleMantle(MNT)$0.56-2.06%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057419-0.33%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Moonshot AI Releases π‘¨π’•π’•π’†π’π’•π’Šπ’π’ π‘Ήπ’†π’”π’Šπ’…π’–π’‚π’π’” to Replace Fixed Residual Mixing with Depth-Wise Attention for Better Scaling in Transformers

March 16, 2026
in AI & Technology
Reading Time: 6 mins read
A A
Moonshot AI Releases π‘¨π’•π’•π’†π’π’•π’Šπ’π’ π‘Ήπ’†π’”π’Šπ’…π’–π’‚π’π’” to Replace Fixed Residual Mixing with Depth-Wise Attention for Better Scaling in Transformers
ShareShareShareShareShare

Residual connections are one of the least questioned parts of modern Transformer design. In PreNorm architectures, each layer adds its output back into a running hidden state, which keeps optimization stable and allows deep models to train. Moonshot AI researchers argue that this standard mechanism also introduces a structural problem: all prior layer outputs are accumulated with fixed unit weights, which causes hidden-state magnitude to grow with depth and progressively weakens the contribution of any single layer.

The research team proposes Attention Residuals (AttnRes) as a drop-in replacement for standard residual accumulation. Instead of forcing every layer to consume the same uniformly mixed residual stream, AttnRes lets each layer aggregate earlier representations using softmax attention over depth. The input to layer (l) is a weighted sum of the token embedding and previous layer outputs, where the weights are computed over prior depth positions rather than over sequence positions. The core idea is simple: if attention improved sequence modeling by replacing fixed recurrence over time, a similar idea can be applied to the depth dimension of a network.

YOU MAY ALSO LIKE

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus

Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI

https://github.com/MoonshotAI/Attention-Residuals/tree/master?tab=readme-ov-file

Why Standard Residuals Become a Bottleneck

The research team identified three issues with standard residual accumulation. First, there is no selective access: all layers receive the same aggregated state even though attention layers and feed-forward or MoE layers may benefit from different mixtures of earlier information. Second, there is irreversible loss: once information is blended into a single residual stream, later layers cannot selectively recover specific earlier representations. Third, there is output growth: deeper layers tend to produce larger outputs to remain influential inside an ever-growing accumulated state, which can destabilize training.

This is the research team’s main framing: standard residuals behave like a compressed recurrence over layers. AttnRes replaces that fixed recurrence with explicit attention over previous layer outputs.

Full AttnRes: Attention Over All Previous Layers

In Full AttnRes, each layer computes attention weights over all preceding depth sources. The default design does not use an input-conditioned query. Instead, each layer has a learned layer-specific pseudo-query vector wl ∈ Rd, while keys and values come from the token embedding and previous layer outputs after RMSNorm. The RMSNorm step is important because it prevents large-magnitude layer outputs from dominating the depth-wise attention weights.

Full AttnRes is straightforward, but it increases cost. Per token, it requires O(L2 d) arithmetic and (O(Ld)) memory to store layer outputs. In standard training this memory largely overlaps with activations already needed for backpropagation, but under activation re-computation and pipeline parallelism the overhead becomes more significant because those earlier outputs must remain available and may need to be transmitted across stages.

Block AttnRes: A Practical Variant for Large Models

To make the method usable at scale, Moonshot AI research team introduces Block AttnRes. Instead of attending over every earlier layer output, the model partitions layers into N blocks. Within each block, outputs are accumulated into a single block representation, and attention is applied only over those block-level representations plus the token embedding. This reduces memory and communication overhead from O(Ld) to O(Nd).

The research team describes cache-based pipeline communication and a two-phase computation strategy that make Block AttnRes practical in distributed training and inference. This results in less than 4% training overhead under pipeline parallelism, while the repository reports less than 2% inference latency overhead on typical workloads.

Scaling Results

The research team evaluates five model sizes and compares three variants at each size: a PreNorm baseline, Full AttnRes, and Block AttnRes with about eight blocks. All variants within each size group share the same hyperparameters chosen under the baseline, which the research team note makes the comparison conservative. The fitted scaling laws are reported as:

Baseline: L = 1.891 x C-0.057
Block AttnRes: L = 1.870 x C-0.058
Full AttnRes: L = 1.865 x C-0.057

The practical implication is that AttnRes achieves lower validation loss across the tested compute range, and the Block AttnRes matches the loss of a baseline trained with about 1.25Γ— more compute.

Integration into Kimi Linear

Moonshot AI also integrates AttnRes into Kimi Linear, its MoE architecture with 48B total parameters and 3B activated parameters, and pre-trains it on 1.4T tokens. According to the research paper, AttnRes mitigates PreNorm dilution by keeping output magnitudes more bounded across depth and distributing gradients more uniformly across layers. Another implementation detail is that all pseudo-query vectors are initialized to zero so the initial attention weights are uniform across source layers, effectively reducing AttnRes to equal-weight averaging at the start of training and avoiding early instability.

On downstream evaluation, the reported gains are consistent across all listed tasks. It reports improvements from 73.5 to 74.6 on MMLU, 36.9 to 44.4 on GPQA-Diamond, 76.3 to 78.0 on BBH, 53.5 to 57.1 on Math, 59.1 to 62.2 on HumanEval, 72.0 to 73.9 on MBPP, 82.0 to 82.9 on CMMLU, and 79.6 to 82.5 on C-Eval.

Key Takeaways

  • Attention Residuals replaces fixed residual accumulation with softmax attention over previous layers.
  • The default AttnRes design uses a learned layer-specific pseudo-query, not an input-conditioned query.
  • Block AttnRes makes the method practical by reducing depth-wise memory and communication from O(Ld) to O(Nd).
  • Moonshot research teamreports lower scaling loss than the PreNorm baseline, with Block AttnRes matching about 1.25Γ— more baseline compute.
  • In Kimi Linear, AttnRes improves results across reasoning, coding, and evaluation benchmarks with limited overhead.

Check outΒ Paper andΒ Repo.Β Also,Β feel free to follow us onΒ TwitterΒ and don’t forget to join ourΒ 120k+ ML SubRedditΒ and Subscribe toΒ our Newsletter. Wait! are you on telegram?Β now you can join us on telegram as well.

The post Moonshot AI Releases π‘¨π’•π’•π’†π’π’•π’Šπ’π’ π‘Ήπ’†π’”π’Šπ’…π’–π’‚π’π’” to Replace Fixed Residual Mixing with Depth-Wise Attention for Better Scaling in Transformers appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus
AI & Technology

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus

September 15, 2026
Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI
AI & Technology

Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI

September 15, 2026
Double The Range And Smarter Safety, Too
AI & Technology

Double The Range And Smarter Safety, Too

September 15, 2026
Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second
AI & Technology

Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second

September 15, 2026
Next Post
Suspected terror attack on power substation

Suspected terror attack on power substation

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump says war with Iran will be over after midterm elections

Trump says war with Iran will be over after midterm elections

September 14, 2026
Clancy’s attorney says he’s not open to a deal requiring jail time

Clancy’s attorney says he’s not open to a deal requiring jail time

September 15, 2026
Sonoco Products Company (SON) Presents at UBS Global Materials Conference 2026 – Slideshow

Sonoco Products Company (SON) Presents at UBS Global Materials Conference 2026 – Slideshow

September 12, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

Β©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. β€œTradepoint.ioβ€œ, β€œInstant Investing” and β€œMy Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. Β© 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

Β© 2023 - TradePoint.io - All Rights Reserved!