• bitcoinBitcoin(BTC)$79,082.002.41%
  • ethereumEthereum(ETH)$2,543.801.59%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$726.010.73%
  • rippleXRP(XRP)$1.467.78%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.182.33%
  • tronTRON(TRX)$0.340412-0.22%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,197.3810.54%
  • HyperliquidHyperliquid(HYPE)$81.224.32%
  • dogecoinDogecoin(DOGE)$0.0847910.81%
  • RainRain(RAIN)$0.014329-6.38%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$514.47-2.87%
  • whitebitWhiteBIT Coin(WBT)$81.812.11%
  • chainlinkChainlink(LINK)$11.742.91%
  • leo-tokenLEO Token(LEO)$9.00-0.71%
  • cardanoCardano(ADA)$0.2124992.11%
  • stellarStellar(XLM)$0.1938548.19%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$226.771.35%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$53.85-1.78%
  • uniswapUniswap(UNI)$6.697.09%
  • CantonCanton(CC)$0.0975402.10%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-0.09%
  • hedera-hashgraphHedera(HBAR)$0.0779702.62%
  • avalanche-2Avalanche(AVAX)$7.622.83%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • nearNEAR Protocol(NEAR)$2.538.86%
  • shiba-inuShiba Inu(SHIB)$0.0000051.46%
  • suiSui(SUI)$0.742.38%
  • crypto-com-chainCronos(CRO)$0.0595672.41%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$237.581.33%
  • tether-goldTether Gold(XAUT)$4,288.73-1.32%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.09-4.35%
  • okbOKB(OKB)$114.380.78%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.08%
  • aaveAave(AAVE)$131.013.82%
  • AsterAster(ASTER)$0.700.86%
  • mantleMantle(MNT)$0.571.35%
  • pax-goldPAX Gold(PAXG)$4,293.32-1.28%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0579471.79%
  • BitwayBitway(BTW)$0.66-6.64%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This Paper Reveals Insights from Reproducing OpenAI’s RLHF (Reinforcement Learning from Human Feedback) Work: Implementation and Scaling Explored

March 29, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This Paper Reveals Insights from Reproducing OpenAI’s RLHF (Reinforcement Learning from Human Feedback) Work: Implementation and Scaling Explored
ShareShareShareShareShare

In recent years, there has been an enormous development in pre-trained large language models (LLMs). These LLMs are trained to predict the next token given the previous tokens and provide a suitable prompt. They can solve various natural language processing (NLP) tasks. However, the next-token prediction objective deviates from the fundamental aim of “outputting contents that humans prefer.” 

To address this gap, Reinforcement Learning from Human Feedback (RLHF) is introduced as a pipeline to collect pair-wise human preferences, train a reward model (RM) to model these preferences, and use Reinforcement Learning (RL) to create a model that outputs contents that humans prefer. It has proven challenging to reproduce OpenAI’s RLHF pipeline in the open-source community for several reasons:

  1. RL and RLHF have many subtle implementation details that can significantly impact training stability.
  2. The models are challenging to evaluate for the following tasks: e.g., assessing the quality of 800 lines of generated code snippets for a coding task.
  3. They take a long time to train and iterate.

Hugging Face, Mila and Fuxi AI lab researchers have undertaken a unique approach, presenting a high-precision reproduction of the Reinforcement Learning from Human Feedback (RLHF) scaling behaviors reported in OpenAI’s seminal TL;DR summarization work. They meticulously created an RLHF pipeline, focusing on over 20 key implementation details. They adopted a unified learning rate for SFT, RM, and PPO training to enhance reproducibility. 

They used the transformers library’s implementation of the Pythia models in conjunction with deepspeed’s ZeRO Stage 2 to help fit the models into the GPU memory; for 6.9B PPO training, they also transferred the reference policy and reward model to the CPU. The dropout layers were turned off during training. This is important for PPO training, especially because with dropout activated, the log probabilities of tokens will not be reproducible, making calculating the KL penalty unreliable while also causing the ratios of the PPO to be not 1s during the first epoch, causing PPO optimization problems. For consistency, they also turn off dropout for SFT and RM training. 

The PPO implementation optimizes the RLHF objective, leading to a significant increase in the score total. Their best 6.9B model is preferred by GPT nearly 80% of the time, demonstrating its practical superiority. For their 1B-sized model, the average preference consistency in multiple random experiments is close to 0.4, indicating that the 1B model has captured a different set of preferences, a finding with important implications. It is shown that PPO models outperform SFT models across all summary lengths, further reinforcing the practical relevance of the research.

In conclusion, Mila and Fuxi AI lab researchers have successfully reproduced the RLHF scaling behaviors reported in OpenAI’s seminal TL;DR summarization work with high precision. Their RLHF-trained Pythia models have demonstrated significant gains in response quality that scale with model size. Notably, their 2.8B and 6.9B models have outperformed OpenAI’s released 1.3B checkpoint, underscoring the importance of model size in achieving superior results.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 39k+ ML SubReddit


YOU MAY ALSO LIKE

NVIDIA Adds RTX PRO 5500 Blackwell GPU with 84 GB GDDR7 Memory – Unite.AI

You Can Use Gemini To Help You Organize Your Files On Google Drive

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

NVIDIA Adds RTX PRO 5500 Blackwell GPU with 84 GB GDDR7 Memory – Unite.AI
AI & Technology

NVIDIA Adds RTX PRO 5500 Blackwell GPU with 84 GB GDDR7 Memory – Unite.AI

September 14, 2026
You Can Use Gemini To Help You Organize Your Files On Google Drive
AI & Technology

You Can Use Gemini To Help You Organize Your Files On Google Drive

September 14, 2026
Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI
AI & Technology

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

September 14, 2026
How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error
AI & Technology

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

September 14, 2026
Next Post
Top GOP donor paid for Clarence Thomas’s grandnephew’s school tuition: Report

Top GOP donor paid for Clarence Thomas’s grandnephew’s school tuition: Report

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Microsoft’s Data Center Plans Face Big Costs

Microsoft’s Data Center Plans Face Big Costs

September 12, 2026
Canoodling lawyers’ firm rocked by another major defection as top attorney bolts

Canoodling lawyers’ firm rocked by another major defection as top attorney bolts

September 10, 2026
LIVE: Vance, Trump deliver remarks at the Republican midterm convention | NBC News

LIVE: Vance, Trump deliver remarks at the Republican midterm convention | NBC News

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!