• bitcoinBitcoin(BTC)$83,928.00-2.89%
  • ethereumEthereum(ETH)$2,657.46-3.22%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$763.75-3.25%
  • rippleXRP(XRP)$1.49-4.18%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$114.14-2.77%
  • tronTRON(TRX)$0.338639-0.62%
  • zcashZcash(ZEC)$1,530.42-0.94%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.43%
  • HyperliquidHyperliquid(HYPE)$92.77-2.99%
  • dogecoinDogecoin(DOGE)$0.092191-7.27%
  • moneroMonero(XMR)$549.39-4.18%
  • whitebitWhiteBIT Coin(WBT)$84.26-2.81%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.23-5.89%
  • cardanoCardano(ADA)$0.238075-4.43%
  • RainRain(RAIN)$0.012446-7.08%
  • leo-tokenLEO Token(LEO)$8.97-0.01%
  • stellarStellar(XLM)$0.202990-4.95%
  • bitcoin-cashBitcoin Cash(BCH)$351.307.59%
  • uniswapUniswap(UNI)$9.12-0.87%
  • nearNEAR Protocol(NEAR)$4.25-3.91%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$60.25-2.94%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.25-7.44%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.108053-4.36%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-2.28%
  • hedera-hashgraphHedera(HBAR)$0.090120-6.03%
  • suiSui(SUI)$0.96-4.26%
  • BittensorBittensor(TAO)$292.56-6.31%
  • shiba-inuShiba Inu(SHIB)$0.000006-6.36%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.061721-7.41%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.21-7.03%
  • tether-goldTether Gold(XAUT)$4,285.06-1.20%
  • BitwayBitway(BTW)$0.9815.58%
  • okbOKB(OKB)$118.34-3.15%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.23%
  • aaveAave(AAVE)$139.57-2.46%
  • mantleMantle(MNT)$0.65-1.20%
  • EthenaEthena(ENA)$0.2088801.27%
  • OndoOndo(ONDO)$0.415667-3.16%
  • Pump.funPump.fun(PUMP)$0.004025-10.87%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Generalization of Gradient Descent in Over-Parameterized ReLU Networks: Insights from Minima Stability and Large Learning Rates

June 16, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Generalization of Gradient Descent in Over-Parameterized ReLU Networks: Insights from Minima Stability and Large Learning Rates
ShareShareShareShareShare

Gradient descent-trained neural networks operate effectively even in overparameterized settings with random weight initialization, often finding global optimum solutions despite the non-convex nature of the problem. These solutions, achieving zero training error, surprisingly do not overfit in many cases, a phenomenon known as “benign overfitting.” However, for ReLU networks, interpolating solutions can lead to overfitting. Moreover, the best solutions usually don’t interpolate the data in noisy data scenarios. Practical training often stops before reaching full interpolation to avoid entering unstable regions or spiky, non-robust solutions.

Researchers from UC Santa Barbara, Technion, and UC San Diego explore the generalization of two-layer ReLU neural networks in 1D nonparametric regression with noisy labels. They present a new theory showing that gradient descent with a fixed learning rate converges to local minima representing smooth, sparsely linear functions. These solutions, which do not interpolate, avoid overfitting and achieve near-optimal mean squared error (MSE) rates. Their analysis highlights that large learning rates induce implicit sparsity and that ReLU networks can generalize well even without explicit regularization or early stopping. This theory moves beyond traditional kernel and interpolation frameworks.

In overparameterized neural networks, most research focuses on generalization within the interpolation regime and benign overfitting. These typically require explicit regularization or early stopping to handle noisy labels. However, recent findings indicate that gradient descent with a large learning rate can achieve sparse, smooth functions that generalize well, even without explicit regularization. This method diverges from traditional theories, which rely on interpolation, demonstrating that gradient descent induces an implicit bias resembling L1-regularization. The study also connects to the hypothesis that “flat local minima generalize better” and provides insights into achieving optimal rates in nonparametric regression without weight decay.

The study addresses the setup and notation for studying generalization in two-layer ReLU neural networks. The model is trained using gradient descent on a dataset with noisy labels, focusing on regression problems. Key concepts include stable local minima, which are twice differentiable and lie within a specific distance from the global minimum. The study also explores the “Edge of Stability” regime, where the Hessian’s largest eigenvalue reaches a critical value related to the learning rate. For nonparametric regression, the target function is from a bounded variation class. The analysis demonstrates that gradient descent cannot find stable interpolating solutions in noisy settings, leading to smoother, non-interpolating functions.

The study’s main results explore stable solutions for gradient descent (GD) on ReLU neural networks across three aspects. First, it examines the implicit bias of stable solutions in the function space under large learning rates, revealing that they are inherently smoother and simpler. Second, it derives generalization bounds for these solutions in distribution-free and non-parametric regression settings, showing they avoid overfitting. Lastly, the analysis demonstrates that GD achieves optimal rates for estimating bounded variation functions within specific intervals, confirming the effective generalization performance of large learning rate GD solutions even in noisy environments.

In conclusion, the study explores how gradient descent-trained two-layer ReLU neural networks generalize through the lens of minima stability and the Edge-of-Stability phenomena. It focuses on univariate inputs with noisy labels and shows that gradient descent with a typical learning rate cannot interpolate data. The study demonstrates that local smoothness of the training loss implies a first-order total variation constraint on the neural network’s function, leading to a vanishing generalization gap within the data support’s strict interior. Additionally, these stable solutions achieve near-optimal rates for estimating first-order bounded variation functions under a mild assumption. The simulations validate the findings, showing that large learning rate training induces sparse linear spline fits.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit

Excited to share our latest work that shows “Large Step Size Training Cannot Overfit” in univariate ReLU nets 🔥🔥🔥 https://t.co/eFHXlJE9yT
For the first time, we understand how *flatness*, *edge-of-stability* and *large stepsize* imply (near-optimal) generalization. 🧵1/

— Yu-Xiang Wang (@yuxiangw_cs) June 13, 2024


YOU MAY ALSO LIKE

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

Never Use ChatGPT For These Five Tasks

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age
AI & Technology

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

September 23, 2026
Never Use ChatGPT For These Five Tasks
AI & Technology

Never Use ChatGPT For These Five Tasks

September 23, 2026
Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes
AI & Technology

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

September 23, 2026
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
AI & Technology

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

September 23, 2026
Next Post
Big pharma companies combating the tampering of life-saving drugs

Big pharma companies combating the tampering of life-saving drugs

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Drone attacks escalate in war with Ukraine

Drone attacks escalate in war with Ukraine

September 17, 2026
U.S. economy adds 162,000 jobs in August, unemployment rate at 4.1%

U.S. economy adds 162,000 jobs in August, unemployment rate at 4.1%

September 17, 2026
NASDAQ Is EXPLODING: What You Need To Know!

NASDAQ Is EXPLODING: What You Need To Know!

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!