• bitcoinBitcoin(BTC)$79,211.00-0.79%
  • ethereumEthereum(ETH)$2,491.690.01%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$739.96-1.27%
  • rippleXRP(XRP)$1.40-1.04%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$104.12-1.46%
  • tronTRON(TRX)$0.334350-0.26%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,155.17-4.83%
  • HyperliquidHyperliquid(HYPE)$85.08-2.57%
  • dogecoinDogecoin(DOGE)$0.0900970.71%
  • RainRain(RAIN)$0.016310-2.83%
  • USDSUSDS(USDS)$1.00-0.02%
  • moneroMonero(XMR)$520.85-1.35%
  • chainlinkChainlink(LINK)$12.763.40%
  • whitebitWhiteBIT Coin(WBT)$76.594.14%
  • leo-tokenLEO Token(LEO)$9.25-0.84%
  • cardanoCardano(ADA)$0.2196550.31%
  • stellarStellar(XLM)$0.1909573.91%
  • bitcoin-cashBitcoin Cash(BCH)$261.011.93%
  • daiDai(DAI)$1.00-0.02%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • uniswapUniswap(UNI)$6.89-2.98%
  • litecoinLitecoin(LTC)$55.081.64%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.105384-3.87%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.40-1.43%
  • hedera-hashgraphHedera(HBAR)$0.0821682.20%
  • avalanche-2Avalanche(AVAX)$8.146.50%
  • suiSui(SUI)$0.833.73%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000050.29%
  • nearNEAR Protocol(NEAR)$2.31-4.50%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057034-0.62%
  • tether-goldTether Gold(XAUT)$4,407.54-0.22%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.152.16%
  • BittensorBittensor(TAO)$259.15-2.22%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$114.971.46%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.02%
  • AsterAster(ASTER)$0.770.28%
  • mantleMantle(MNT)$0.623.72%
  • aaveAave(AAVE)$132.18-0.51%
  • pax-goldPAX Gold(PAXG)$4,409.82-0.26%
  • OndoOndo(ONDO)$0.3843741.78%
  • polkadotPolkadot(DOT)$1.0913.59%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Do You Really Need Reinforcement Learning (RL) in RLHF? A New Stanford Research Proposes DPO (Direct Preference Optimization): A Simple Training Paradigm For Training Language Models From Preferences Without RL

June 3, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Do You Really Need Reinforcement Learning (RL) in RLHF? A New Stanford Research Proposes DPO (Direct Preference Optimization): A Simple Training Paradigm For Training Language Models From Preferences Without RL
ShareShareShareShareShare

When trained on massive datasets, huge unsupervised LMs acquire powers that surprise even their creators. These models, however, are trained on information produced by people with a diverse range of motivations, objectives, and abilities. Not all of these ambitions and abilities may be emulated. It is important to carefully select the model’s desired responses and behavior from its vast store of information and skills to create reliable, effective, and manageable systems.  

Without using explicit reward modeling or reinforcement learning, Stanford University and CZ researchers demonstrate how to optimize a language model to conform to human tastes. Their work shows that the RL-based objective employed by present approaches can be optimized exactly with a simple binary cross-entropy objective, considerably streamlining the preference learning process and demonstrating how this can be done in practice. 

They propose Direct Preference Optimization (DPO). This new algorithm implicitly achieves the same objective as existing RLHF algorithms (reward maximization with a KL-divergence constraint) but is easier to construct and train. While the DPO update intuitively boosts the log ratio of preferred to dispreferred replies, it also includes a dynamic, per-example significance weight that stops the model from degrading.

🚀 JOIN the fastest ML Subreddit Community

Like other algorithms, DPO evaluates the consistency of a reward function with empirical preference data using a theoretical preference model. While conventional approaches define a preference loss using the preference model to train a reward model, DPO instead trains a policy that maximizes the learned reward model using a variable switch. Therefore, DPO may optimize a policy with a simple binary cross-entropy goal given a dataset of human preferences over model responses without explicitly learning a reward function or sampling from the policy during training. 

The work’s findings demonstrate that DPO is as effective as state-of-the-art approaches, such as PPO-based RLHF, for preference-based learning on various tasks, including sentiment modulation, summarization, and dialogue, with language models containing up to 6B parameters. 58% of people prefer DPO summaries to PPO summaries (human evaluations), and 61% prefer DPO summaries to human evaluations in the test set. On Anthropic HH, 60% of the time, single-turn responses from DPOs are preferred over selective completions. 

The team states that DPO has many potential uses beyond only training language models based on human preferences. For example, it can train generative models in various modalities.

The proposed model evaluations go as high as 6B parameters, but the team believes that further work should explore scaling DPO to state-of-the-art models with orders of magnitude more data. The researchers also discovered that the prompt affects GPT -4’s computed win rates. In the future, they plan to investigate the most effective means of eliciting expert opinions from machines. 


Check Out The Paper. Don’t forget to join our 22k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

When Is It No Longer Worth Repairing Your Phone And Buying A New One Instead

Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI

Tanushree Shenwai is a consulting intern at MarktechPost. She is currently pursuing her B.Tech from the Indian Institute of Technology(IIT), Bhubaneswar. She is a Data Science enthusiast and has a keen interest in the scope of application of artificial intelligence in various fields. She is passionate about exploring the new advancements in technologies and their real-life application.


➡️ Ultimate Guide to Data Labeling in Machine Learning

Credit: Source link

ShareTweetSendSharePin

Related Posts

When Is It No Longer Worth Repairing Your Phone And Buying A New One Instead
AI & Technology

When Is It No Longer Worth Repairing Your Phone And Buying A New One Instead

September 7, 2026
Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI
AI & Technology

Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI

September 7, 2026
Grupo Financiero Inbursa Adopts Harvey Across Its Legal Organization – Unite.AI
AI & Technology

Grupo Financiero Inbursa Adopts Harvey Across Its Legal Organization – Unite.AI

September 7, 2026
How To Find And Hide An App On Android Auto
AI & Technology

How To Find And Hide An App On Android Auto

September 7, 2026
Next Post
Uber Deal with China’s Didi Said to be Worth  Billion

Uber Deal with China's Didi Said to be Worth $35 Billion

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
US unemployment claims rise to 206,000 — but remain at historically low levels

US unemployment claims rise to 206,000 — but remain at historically low levels

September 3, 2026
Nolan Wells’ state autopsy completed 

Nolan Wells’ state autopsy completed 

September 1, 2026
The True Cost of a Car Is Far More Than the Sticker Price

The True Cost of a Car Is Far More Than the Sticker Price

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!