• bitcoinBitcoin(BTC)$76,760.00-3.02%
  • ethereumEthereum(ETH)$2,428.10-4.39%
  • tetherTether(USDT)$1.00-0.04%
  • binancecoinBNB(BNB)$721.55-0.79%
  • rippleXRP(XRP)$1.39-4.52%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.67-3.70%
  • tronTRON(TRX)$0.332654-2.33%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.04-0.03%
  • zcashZcash(ZEC)$1,128.59-4.77%
  • HyperliquidHyperliquid(HYPE)$77.20-5.56%
  • dogecoinDogecoin(DOGE)$0.081841-3.84%
  • RainRain(RAIN)$0.0144610.68%
  • USDSUSDS(USDS)$1.00-0.03%
  • moneroMonero(XMR)$505.23-1.15%
  • whitebitWhiteBIT Coin(WBT)$79.04-3.45%
  • chainlinkChainlink(LINK)$11.26-3.48%
  • leo-tokenLEO Token(LEO)$8.89-1.16%
  • cardanoCardano(ADA)$0.201937-5.65%
  • stellarStellar(XLM)$0.189803-2.30%
  • Ethena USDeEthena USDe(USDE)$1.00-0.06%
  • daiDai(DAI)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$221.94-2.42%
  • USD1USD1(USD1)$1.00-0.05%
  • litecoinLitecoin(LTC)$51.81-4.49%
  • uniswapUniswap(UNI)$6.44-1.78%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.33-2.24%
  • CantonCanton(CC)$0.093478-3.87%
  • hedera-hashgraphHedera(HBAR)$0.077596-1.00%
  • avalanche-2Avalanche(AVAX)$7.48-1.56%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.38-6.77%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.55%
  • suiSui(SUI)$0.71-4.50%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • crypto-com-chainCronos(CRO)$0.056956-4.29%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,300.37-0.21%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$224.37-5.62%
  • MemeCoreMemeCore(M)$1.111.27%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$111.26-2.74%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.19%
  • aaveAave(AAVE)$125.89-2.35%
  • BitwayBitway(BTW)$0.701.68%
  • pax-goldPAX Gold(PAXG)$4,303.47-0.26%
  • AsterAster(ASTER)$0.69-1.71%
  • mantleMantle(MNT)$0.55-4.18%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056969-2.27%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks

September 14, 2026
in AI & Technology
Reading Time: 16 mins read
A A
Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks
ShareShareShareShareShare

Backpropagation is a global algorithm: a forward pass, then a backward pass, then a weight update, each locked behind the previous one. Brains have no known mechanism for that kind of network-wide phase locking, which is why local-learning alternatives such as predictive coding (PC) keep drawing research interest. Sakana AI researchers propose Augmented Lagrangian Predictive Coding (PC-ALM), a variant of PC that keeps every update layer-local yet recovers backprop-aligned credit signals. The research team reports training residual MLPs up to 1000 layers within about 2 percentage points of backprop on MNIST.

Is it deployable? Yes, as research code: an MIT-licensed JAX reference implementation runs on CPU and reproduces the paper’s width-depth grid. It is a training method, not a model, and has only been tested on small image benchmarks.

Why standard PC stalls in deep, narrow networks

PC treats every hidden activation as an optimization variable and penalizes the squared mismatch between each layer’s activation and the prediction arriving from the layer below. Inference is gradient descent on that energy; learning is a Hebbian-like weight step. The catch is that supervision enters at the output and must diffuse through a chain of local compromises. In deep, narrow networks the credit signal fades long before it reaches the input. Innocenti et al. characterized this PC-BP gap as a function of width and depth, and it is worst when width is smaller than depth.

What PC-ALM changes

PC-ALM starts from the constrained view of training: minimize the supervised loss subject to hi=σ(Wihi−1)h_i = \sigma(W_i h_{i-1}) at every layer. PC is the quadratic-penalty relaxation of that problem. PC-ALM uses the augmented Lagrangian instead, attaching a Lagrange multiplier λi∈ℝdisuch thatdim(λi)=dim(hi)\lambda_i \in \mathbb{R}^{d_i} \quad \text{such that} \quad \text{dim}(\lambda_i) = \text{dim}(h_i) to each layer constraint while keeping PC’s penalty. Setting λ = 0 recovers PC exactly.

Inference alternates 2 local steps: a primal gradient step on the activations, and a dual step λi←λi+αri\lambda_i \leftarrow \lambda_i + \alpha r_i that accumulates the layer’s prediction error. Completing the square shows each primal step is a standard PC step with the prediction target shifted by −λi/ρ-\lambda_i/\rho. After T steps the weight update acts on the composite signal λi+ρri\lambda_i + \rho r_i. The research team read this as a PI controller per layer: the prediction error is the proportional term and the multiplier is the integral term. α = 0 gives PC; α = ρ with the inner problem solved exactly gives the classical method of multipliers.

Exact backprop gradients in the linear case

LeCun observed in 1988 that the Lagrange multipliers of a constrained network equal the backprop adjoints at a KKT point. The team proves that in linear PC networks, under a spectral-radius stability condition, PC-ALM converges to that KKT point: activations return to their forward-pass values while each λi\lambda_i integrates to the exact BP adjoint. The per-mode stability bound is ηhσi2(2ρ+α)4\eta_h \sigma_i^2 (2\rho + \alpha) , which reduces to PC’s condition at α = 0. Unlike PC’s monotone gradient flow, PC-ALM’s iteration matrix has complex eigenvalues that produce damped oscillations; α sets their frequency but not their decay rate.

YOU MAY ALSO LIKE

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models – Unite.AI

Are Older MacBooks Still Worth Buying In 2026?

Results

The research team sweeps residual MLPs with width and depth from 8 to 128 on Fashion-MNIST and MNIST under the mean-field parameterization of Innocenti et al., training for 1 epoch. With an inference budget of T = 2L, PC-ALM matches backprop across every width, depth, and activation (identity, tanh, ReLU), while PC drops sharply in deep, narrow cells. The repo’s reference cell (width 32, depth 32, ReLU, Fashion-MNIST) reports 78.66% test accuracy for BP, 68.13% for PC, and 77.75% for PC-ALM, with gradient cosine to BP rising from 0.604 to 0.909.

The research extends the picture: 1000-layer residual MLPs on MNIST (width 32, ReLU, 5 epochs) stay within roughly 2 points of BP, and PC-ALM improves over PC on every benchmark tried, including ResNet-18 on CIFAR-10 and Tiny ImageNet.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models – Unite.AI
AI & Technology

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models – Unite.AI

September 15, 2026
Are Older MacBooks Still Worth Buying In 2026?
AI & Technology

Are Older MacBooks Still Worth Buying In 2026?

September 15, 2026
This Is A Great Place To Store Your Old Hard Drives And Keep Them Safe
AI & Technology

This Is A Great Place To Store Your Old Hard Drives And Keep Them Safe

September 15, 2026
Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron – Unite.AI
AI & Technology

Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron – Unite.AI

September 15, 2026
Next Post
Ex-Apple Researchers’ Nuance Labs Lands M Series A Backed by Nvidia – Unite.AI

Ex-Apple Researchers’ Nuance Labs Lands $50M Series A Backed by Nvidia – Unite.AI

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump administration proposes changes to census that could exclude millions – The Washington Post

Trump administration proposes changes to census that could exclude millions – The Washington Post

September 9, 2026
You Can Now Plan IRL Events On Snapchat

You Can Now Plan IRL Events On Snapchat

September 10, 2026
Day’s first moment of silence remembers the victims of 9/11

Day’s first moment of silence remembers the victims of 9/11

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!