• bitcoinBitcoin(BTC)$78,197.00-0.39%
  • ethereumEthereum(ETH)$2,466.30-0.58%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$730.43-2.85%
  • rippleXRP(XRP)$1.40-1.41%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.28-0.65%
  • tronTRON(TRX)$0.3394970.46%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-1.11%
  • zcashZcash(ZEC)$1,241.596.31%
  • HyperliquidHyperliquid(HYPE)$84.940.76%
  • dogecoinDogecoin(DOGE)$0.086638-3.35%
  • RainRain(RAIN)$0.015922-2.73%
  • USDSUSDS(USDS)$1.00-0.02%
  • moneroMonero(XMR)$509.352.46%
  • whitebitWhiteBIT Coin(WBT)$80.70-0.73%
  • chainlinkChainlink(LINK)$11.77-5.77%
  • leo-tokenLEO Token(LEO)$9.18-0.23%
  • cardanoCardano(ADA)$0.213298-3.00%
  • stellarStellar(XLM)$0.182098-3.17%
  • bitcoin-cashBitcoin Cash(BCH)$254.19-0.96%
  • daiDai(DAI)$1.00-0.03%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$53.52-1.13%
  • CantonCanton(CC)$0.104482-1.93%
  • uniswapUniswap(UNI)$6.33-5.61%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-1.70%
  • avalanche-2Avalanche(AVAX)$7.85-1.44%
  • hedera-hashgraphHedera(HBAR)$0.077048-2.59%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$2.477.53%
  • suiSui(SUI)$0.78-3.58%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.12%
  • crypto-com-chainCronos(CRO)$0.058787-0.27%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • MemeCoreMemeCore(M)$1.20-1.69%
  • tether-goldTether Gold(XAUT)$4,399.710.87%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$255.14-0.72%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$112.55-0.87%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.05%
  • mantleMantle(MNT)$0.61-4.03%
  • AsterAster(ASTER)$0.73-2.43%
  • aaveAave(AAVE)$125.33-2.47%
  • pax-goldPAX Gold(PAXG)$4,403.410.94%
  • polkadotPolkadot(DOT)$1.11-10.15%
  • Pump.funPump.fun(PUMP)$0.0043221.13%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Stanford and UT Austin Researchers Propose Contrastive Preference Learning (CPL): A Simple Reinforcement Learning RL-Free Method for RLHF that Works with Arbitrary MDPs and off-Policy Data

October 31, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Stanford and UT Austin Researchers Propose Contrastive Preference Learning (CPL): A Simple Reinforcement Learning RL-Free Method for RLHF that Works with Arbitrary MDPs and off-Policy Data
ShareShareShareShareShare

The challenge of matching human preferences to big pretrained models has gained prominence in the study as these models have grown in performance. This alignment becomes particularly challenging when there are unavoidably poor behaviours in bigger datasets. For this issue, reinforcement learning from human input, or RLHF has become popular. RLHF approaches use human preferences to distinguish between acceptable and bad behaviours to improve a known policy. This approach has demonstrated encouraging outcomes when used to adjust robot rules, enhance image generation models, and fine-tune large language models (LLMs) using less-than-ideal data. There are two stages to this procedure for the majority of RLHF algorithms. 

First, user preference data is gathered to train a reward model. An off-the-shelf reinforcement learning (RL) algorithm optimizes that reward model. Regretfully, there needs to be a correction in the foundation of this two-phase paradigm. Human preferences must be allocated by the discounted total of rewards or partial return of each behaviour segment for algorithms to develop reward models from preference data. Recent research, however, challenges this theory, suggesting that human preferences should be based on the regret of each action under the ideal policy of the expert’s reward function. Human evaluation is probably intuitively focused on optimality rather than whether situations and behaviours provide greater rewards. 

Therefore, the optimal advantage function, or the negated regret, may be the ideal number to learn from feedback rather than the reward. Two-phase RLHF algorithms use RL in their second phase to optimize the reward function known in the first phase. In real-world applications, temporal credit assignment presents a variety of optimization difficulties for RL algorithms, including the instability of approximation dynamic programming and the high variance of policy gradients. As a result, earlier works restrict their reach to avoid these problems. For example, contextual bandit formulation is assumed by RLHF approaches for LLMs, where the policy is given a single reward value in response to a user question. 

The single-step bandit assumption is broken because user interactions with LLMs are multi-step and sequential, even while this lessens the requirement for long-horizon credit assignment and, as a result, the high variation of policy gradients. Another example is the application of RLHF to low-dimensional state-based robotics issues, which works well for approximation dynamic programming. However, it has yet to be scaled to higher-dimensional continuous control domains with picture inputs, which are more realistic. In general, RLHF approaches require reducing the optimisation constraints of RL by making restricted assumptions about the sequential nature of problems or dimensionality. They generally mistakenly believe that the reward function alone determines human preferences.

In contrast to the widely used partial return model, which considers the total rewards, researchers from Stanford University, UMass Amherst and UT Austin provide a novel family of RLHF algorithms in this study that employs a regret-based model of preferences. In contrast to the partial return model, the regret-based approach gives precise information on the best course of action. Fortunately, this removes the necessity for RL, enabling us to tackle RLHF issues with high-dimensional state and action spaces in the generic MDP framework. Their fundamental finding is to create a bijection between advantage functions and policies by combining the regret-based preference framework with the Maximum Entropy (MaxEnt) principle. 

They can establish a purely supervised learning objective whose optimum is the best policy under the expert’s reward by trading optimization over advantages for optimization over policies. Because their method resembles widely recognized contrastive learning objectives, they call it Contrastive Preference Learning—three main benefits of CPL over earlier efforts. First, because CPL matches the optimal advantage exclusively using supervised goals—rather than using dynamic programming or policy gradients—it can scale as well as supervised learning. Second, CPL is completely off-policy, making using any offline, less-than-ideal data source possible. Lastly, CPL enables preference searches over sequential data for learning on arbitrary Markov Decision Processes (MDPs). 

As far as they know, previous techniques for RLHF have yet to satisfy all three of these requirements simultaneously. They illustrate CPL’s performance on sequential decision-making issues using sub-optimal and high-dimensional off-policy inputs to prove that it adheres to the abovementioned three tenets. Interestingly, they demonstrate that CPL may learn temporally extended manipulation rules in the MetaWorld Benchmark by efficiently utilising the same RLHF fine-tuning process as dialogue models. To be more precise, they use supervised learning from high-dimensional picture observations to pre-train policies, which they then fine-tune using preferences. CPL can match the performance of earlier RL-based techniques without the need for dynamic programming or policy gradients. It is also four times more parameter efficient and 1.6 times quicker simultaneously. On five tasks out of six, CPL outperforms RL baselines when utilizing denser preference data. Researchers can avoid the necessity for reinforcement learning (RL) by employing the concept of maximum entropy to create Contrastive Preference Learning (CPL), an algorithm for learning optimum policies from preferences without learning reward functions.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 32k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

Blizzard Employees Have Ratified Their First Union Contracts

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

Blizzard Employees Have Ratified Their First Union Contracts
AI & Technology

Blizzard Employees Have Ratified Their First Union Contracts

September 9, 2026
OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI
AI & Technology

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI

September 9, 2026
Google and NASA JPL Unveil AI Model Mapping Global Methane Plumes – Unite.AI
AI & Technology

Google and NASA JPL Unveil AI Model Mapping Global Methane Plumes – Unite.AI

September 9, 2026
Lightfield Raises M Series A Led by a16z to Accelerate Growth – Unite.AI
AI & Technology

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

September 9, 2026
Next Post
#WestVirginia Woman Wakes Up From Coma, Accuses Brother For Attack

#WestVirginia Woman Wakes Up From Coma, Accuses Brother For Attack

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
U.S. military intentionally allowing some Iranian projectiles through defenses

U.S. military intentionally allowing some Iranian projectiles through defenses

September 5, 2026
Tropical Storm Bertha remains relatively weak as Gulf Coast deals with rough surf

Tropical Storm Bertha remains relatively weak as Gulf Coast deals with rough surf

September 7, 2026
Iran attacks Kuwait with missiles and drones after US strikes – Fox News

Iran attacks Kuwait with missiles and drones after US strikes – Fox News

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!