• bitcoinBitcoin(BTC)$85,569.004.92%
  • ethereumEthereum(ETH)$2,738.152.70%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$788.502.01%
  • rippleXRP(XRP)$1.526.02%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.114.77%
  • tronTRON(TRX)$0.3487471.61%
  • zcashZcash(ZEC)$1,498.98-1.50%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$94.410.51%
  • dogecoinDogecoin(DOGE)$0.10016613.20%
  • moneroMonero(XMR)$575.23-2.44%
  • whitebitWhiteBIT Coin(WBT)$86.093.37%
  • RainRain(RAIN)$0.013695-3.36%
  • chainlinkChainlink(LINK)$12.973.23%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2451665.68%
  • leo-tokenLEO Token(LEO)$8.950.41%
  • stellarStellar(XLM)$0.2129927.38%
  • nearNEAR Protocol(NEAR)$4.421.96%
  • uniswapUniswap(UNI)$9.044.46%
  • bitcoin-cashBitcoin Cash(BCH)$266.004.72%
  • Ethena USDeEthena USDe(USDE)$1.00-0.04%
  • avalanche-2Avalanche(AVAX)$10.74-5.44%
  • litecoinLitecoin(LTC)$61.104.75%
  • CantonCanton(CC)$0.1185215.77%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.03%
  • suiSui(SUI)$1.028.63%
  • hedera-hashgraphHedera(HBAR)$0.0929987.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.94%
  • BittensorBittensor(TAO)$323.6619.95%
  • shiba-inuShiba Inu(SHIB)$0.0000069.28%
  • crypto-com-chainCronos(CRO)$0.0661447.33%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.37-10.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,332.47-0.54%
  • okbOKB(OKB)$121.631.84%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.04%
  • aaveAave(AAVE)$144.064.43%
  • EthenaEthena(ENA)$0.2180901.88%
  • pepePepe(PEPE)$0.00000527.78%
  • BitwayBitway(BTW)$0.793.07%
  • Pump.funPump.fun(PUMP)$0.0045753.89%
  • OndoOndo(ONDO)$0.4363341.70%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Introduces WARP: A Novel Reinforcement Learning from Human Feedback RLHF Method to Align LLMs and Optimize the KL-Reward Pareto Front of Solutions

June 29, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Google DeepMind Introduces WARP: A Novel Reinforcement Learning from Human Feedback RLHF Method to Align LLMs and Optimize the KL-Reward Pareto Front of Solutions
ShareShareShareShareShare

Reinforcement learning from human feedback (RLHF) encourages generations to have high rewards, using a reward model trained on human preferences to align large language models (LLMs). However, RLHF has several unresolved issues. First, the fine-tuning process is often limited to small datasets, causing the model to become too specialized and miss the wide range of knowledge it learned during pre-training. This can lower the LLM’s reasoning abilities and performance on NLP benchmarks. Second, trying to maximize an imperfect reward model (RM) can lead to problems, as the LLM might find ways to exploit flaws in the RM. Lastly, RLHF can reduce the variety of outputs, causing the model to collapse to produce similar responses.

This paper discusses two related topics. The first topic is how to merge models. Recently, the idea of merging deep models in the weight space, rather than in the prediction space as traditionally done in ensembling, has gained great attention. This method is called weight averaging (WA), and the most common form of WA is LERP. This form was initially used to average checkpoints from a single run, uniformly or with an exponential moving average (EMA). The second topic is the benefits of model merging, where WA improves generalization by reducing variance, memorization, and flattening the loss landscape. Moreover, merging weights combines their strengths, which is useful in multi-task setups.

YOU MAY ALSO LIKE

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

Why It’s Important To Unplug Your PC During A Power Outage

A team from Google DeepMind has proposed Weight Averaged Rewarded Policies (WARP), a method to align LLMs and optimize the Kullback-Leibler(KL)-reward Pareto front of solutions. WARP uses three types of WA at three stages of the alignment process for distinct reasons.  First, it uses the exponential moving average of the policy in the KL regularization as a flexible reference point. Second, it merges fine-tuned policies into an improved policy through spherical interpolation. Third, it linearly interpolates between the merged model and the initialization, to get back features from pre-training. This process is repeated, where each final model serves as a starting point for the next iteration, and enhances the KL-reward Pareto front, obtaining better rewards at fixed KL. 

In the experiment carried out by the team, Gemma “7B” LLM is considered and fine-tuned with RLHF into a better conversational agent. Moreover, the REINFORCE policy gradient is also utilized to optimize the KL-regularized reward. After that, on-policy samples are generated using the dataset which includes conversation prompts, with a temperature of 0.9, batch size of 128, Adam optimizer with learning rate 10−6, warmup of 100 steps, and SLERP is applied to the 28 layers separately. It’s important to note that this experiment relies on the high-capacity reward model, the largest available, which prevents the use of an oracle control RM.

Side-by-side comparisons were made for the trained policies against Mistral and Mixtral LLMs. Each policy generated a candidate answer from a set of prompts as described in the Gemma tech report. Similar to Gemini 1.5, side-by-side preference rates were calculated with “much better”, “better” and “slightly better” receiving scores of ±1.5, ±1, and ±0.5 respectively, and ties receiving a score of 0. A positive score means better policies. The results validate that WARP is efficient, as the proposed policies were preferred over the Mistral variants and outperformed the previous Gemma “7B” releases.

In conclusion, a team from Google DeepMind has introduced (WARP), a novel RLHF method to align LLMs and optimize the KL-reward Pareto front of solutions. It uses three distinct stages of model merging, (a) exponential moving average as a dynamic anchor during RL, (b) spherical interpolation to combine multiple policies rewarded independently, and (c) interpolation towards the shared initialization. This iterative application of WARP improves the KL-reward Pareto front, aligning the LLMs while protecting the knowledge from pre-training, and compares favorably against state-of-the-art baselines. In the future, WARP could help create safe and powerful AI systems by improving alignment and encouraging further study of model merging techniques. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 45k+ ML SubReddit


🚀 Create, edit, and augment tabular data with the first compound AI system, Gretel Navigator, now generally available! [Advertisement]


Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028
AI & Technology

These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028

September 21, 2026
Next Post
Biden says Gaza hospitals ‘must be protected’

Biden says Gaza hospitals 'must be protected'

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
1984 track champ preps South LA bakery for 2028 Olympics

1984 track champ preps South LA bakery for 2028 Olympics

September 15, 2026
Thai police crackdown on sex workers ahead of USS Lincoln’s arrival

Thai police crackdown on sex workers ahead of USS Lincoln’s arrival

September 20, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!