• bitcoinBitcoin(BTC)$86,352.001.11%
  • ethereumEthereum(ETH)$2,747.790.60%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$789.440.36%
  • rippleXRP(XRP)$1.626.38%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$118.691.72%
  • tronTRON(TRX)$0.343778-1.30%
  • zcashZcash(ZEC)$1,621.127.80%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.77%
  • HyperliquidHyperliquid(HYPE)$97.092.79%
  • dogecoinDogecoin(DOGE)$0.1018012.16%
  • moneroMonero(XMR)$571.27-0.55%
  • whitebitWhiteBIT Coin(WBT)$86.761.03%
  • chainlinkChainlink(LINK)$13.010.77%
  • cardanoCardano(ADA)$0.2575534.96%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013077-4.31%
  • leo-tokenLEO Token(LEO)$8.980.26%
  • stellarStellar(XLM)$0.2201413.61%
  • bitcoin-cashBitcoin Cash(BCH)$353.4533.03%
  • uniswapUniswap(UNI)$10.3716.46%
  • nearNEAR Protocol(NEAR)$4.504.21%
  • litecoinLitecoin(LTC)$63.955.19%
  • avalanche-2Avalanche(AVAX)$11.144.03%
  • Ethena USDeEthena USDe(USDE)$1.000.08%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.114558-3.93%
  • USD1USD1(USD1)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0991325.95%
  • suiSui(SUI)$1.020.49%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.461.96%
  • shiba-inuShiba Inu(SHIB)$0.0000062.24%
  • BittensorBittensor(TAO)$313.26-1.65%
  • crypto-com-chainCronos(CRO)$0.0676663.46%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.29-4.02%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,326.740.11%
  • okbOKB(OKB)$125.182.87%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BitwayBitway(BTW)$0.9315.58%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • aaveAave(AAVE)$151.136.42%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • mantleMantle(MNT)$0.697.32%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.23%
  • EthenaEthena(ENA)$0.2162782.15%
  • OndoOndo(ONDO)$0.4376480.99%
  • pepePepe(PEPE)$0.000005-4.94%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

HyPO: A Hybrid Reinforcement Learning Algorithm that Uses Offline Data for Contrastive-based Preference Optimization and Online Unlabeled Data for KL Regularization

July 29, 2024
in AI & Technology
Reading Time: 6 mins read
A A
HyPO: A Hybrid Reinforcement Learning Algorithm that Uses Offline Data for Contrastive-based Preference Optimization and Online Unlabeled Data for KL Regularization
ShareShareShareShareShare

A critical aspect of AI research involves fine-tuning large language models (LLMs) to align their outputs with human preferences. This fine-tuning ensures that AI systems generate useful, relevant, and aligned responses with user expectations. The current paradigm in AI emphasizes learning from human preference data to refine these models, addressing the complexity of manually specifying reward functions for various tasks. The two predominant techniques in this area are online reinforcement learning (RL) and offline contrastive methods, each offering unique advantages and challenges.

A central challenge in fine-tuning LLMs to reflect human preferences is the limited coverage of static datasets. These datasets may need to adequately represent the diverse and dynamic range of human preferences in real-world applications. The issue of dataset coverage becomes particularly pronounced when models are trained exclusively on pre-collected data, potentially leading to suboptimal performance. This problem underscores the need for methods to effectively leverage static datasets and real-time data to enhance model alignment with human preferences.

YOU MAY ALSO LIKE

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

The Pros And Cons Of Using A Password Manager Over An Authenticator App

Existing techniques for preference fine-tuning in LLMs include online RL methods, such as Proximal Policy Optimization (PPO), and offline contrastive methods, like Direct Preference Optimization (DPO). Online RL methods involve a two-stage procedure where a reward model is trained on a fixed offline preference dataset, followed by RL training using on-policy data. This approach benefits from real-time feedback but is computationally intensive. In contrast, offline contrastive methods optimize policies based solely on pre-collected data, avoiding the need for real-time sampling but potentially suffering from overfitting and limited generalization capabilities.

Researchers from Carnegie Mellon University, Aurora Innovation, and Cornell University introduced a novel method called Hybrid Preference Optimization (HyPO). This hybrid approach combines the power of both online and offline techniques, aiming to improve model performance while maintaining computational efficiency. HyPO integrates offline data for initial preference optimization. It uses online unlabeled data for Kullback-Leibler (KL) regularization, ensuring the model remains close to a reference policy and better generalizes beyond the training data.

HyPO utilizes a sophisticated algorithmic framework that leverages offline data for the DPO objective and online samples to control the reverse KL divergence. The algorithm iteratively updates the model’s parameters by optimizing the DPO loss while incorporating a KL regularization term derived from online samples. This hybrid approach effectively addresses the deficiencies of purely offline methods, such as overfitting and insufficient dataset coverage, by incorporating the strengths of online RL methods without their computational complexity.

The performance of HyPO was evaluated on several benchmarks, including the TL;DR summarization task and general chat benchmarks like AlpacaEval 2.0 and MT-Bench. The results were impressive, with HyPO achieving a win rate of 46.44% on the TL;DR task using the Pythia 1.4B model, compared to 42.17% for the DPO method. For the Pythia 2.8B model, HyPO achieved a win rate of 50.50%, significantly outperforming DPO’s 44.39%. Additionally, HyPO demonstrated superior control over reverse KL divergence, with values of 0.37 and 2.51 for the Pythia 1.4B and 2.8B models, respectively, compared to 0.16 and 2.43 for DPO.

In general chat benchmarks, HyPO also showed notable improvements. For instance, in the MT-Bench evaluation, HyPO fine-tuned models achieved scores of 8.43 and 8.09 in the first and second turn averages, respectively, surpassing the DPO-fine-tuned models’ scores of 8.31 and 7.89. Similarly, in the AlpacaEval 2.0, HyPO achieved 30.7% and 32.2% win rates for the 1st and 2nd turns, compared to DPO’s 28.4% and 30.9%.

The empirical results highlight HyPO’s ability to mitigate overfitting issues commonly observed in offline contrastive methods. For example, when trained on the TL;DR dataset, HyPO maintained a mean validation KL score significantly lower than that of DPO, indicating better alignment with the reference policy and reduced overfitting. This ability to leverage online data for regularization helps HyPO achieve more robust performance across various tasks.

In conclusion, the introduction of hybrid preference optimization (HyPO), which effectively combines offline and online data, addresses the limitations of existing methods and enhances the alignment of large language models with human preferences. The performance improvements demonstrated in empirical evaluations underscore the potential of HyPO to deliver more accurate and reliable AI systems.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 47k+ ML SubReddit

Find Upcoming AI Webinars here


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks
AI & Technology

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

September 23, 2026
The Pros And Cons Of Using A Password Manager Over An Authenticator App
AI & Technology

The Pros And Cons Of Using A Password Manager Over An Authenticator App

September 23, 2026
How To Hide Or Replace The Audio Button In iMessages
AI & Technology

How To Hide Or Replace The Audio Button In iMessages

September 22, 2026
Improve Your Apple CarPlay Experience By Doing These Simple Things
AI & Technology

Improve Your Apple CarPlay Experience By Doing These Simple Things

September 22, 2026
Next Post
U.S. withholds thousands of bombs from Israel

U.S. withholds thousands of bombs from Israel

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
A Toaster With A Vision

A Toaster With A Vision

September 16, 2026
Anthropic Launches Claude Code Projects in Beta: Parallel Cloud Sessions That Keep Running After You Close Your Laptop

Anthropic Launches Claude Code Projects in Beta: Parallel Cloud Sessions That Keep Running After You Close Your Laptop

September 17, 2026
SpaceX May Buy Data From Failed Startups for AI Models

SpaceX May Buy Data From Failed Startups for AI Models

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!