• bitcoinBitcoin(BTC)$76,695.00-0.80%
  • ethereumEthereum(ETH)$2,475.77-2.34%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$715.14-2.87%
  • rippleXRP(XRP)$1.34-2.38%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.56-2.59%
  • tronTRON(TRX)$0.340691-0.09%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,088.90-6.20%
  • HyperliquidHyperliquid(HYPE)$77.41-3.60%
  • dogecoinDogecoin(DOGE)$0.083300-2.08%
  • RainRain(RAIN)$0.0152610.83%
  • moneroMonero(XMR)$530.510.38%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$79.53-1.05%
  • chainlinkChainlink(LINK)$11.25-2.65%
  • leo-tokenLEO Token(LEO)$9.05-0.64%
  • cardanoCardano(ADA)$0.204309-2.10%
  • stellarStellar(XLM)$0.178157-1.96%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$223.48-3.34%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.49-0.94%
  • uniswapUniswap(UNI)$6.26-2.96%
  • CantonCanton(CC)$0.095115-3.40%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-2.37%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0751820.64%
  • avalanche-2Avalanche(AVAX)$7.30-1.97%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.60%
  • nearNEAR Protocol(NEAR)$2.31-2.53%
  • suiSui(SUI)$0.71-2.70%
  • crypto-com-chainCronos(CRO)$0.0585840.45%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.32-0.02%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.14-2.65%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.72-1.24%
  • BittensorBittensor(TAO)$233.78-1.13%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • aaveAave(AAVE)$125.36-1.35%
  • pax-goldPAX Gold(PAXG)$4,350.50-0.04%
  • AsterAster(ASTER)$0.691.00%
  • mantleMantle(MNT)$0.56-1.67%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0575070.97%
  • polkadotPolkadot(DOT)$1.01-3.51%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from NVIDIA and the University of Maryland Propose ODIN: A Reward Disentangling Technique that Mitigates Hacking in Reinforcement Learning from Human Feedback (RLHF)

February 25, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Researchers from NVIDIA and the University of Maryland Propose ODIN: A Reward Disentangling Technique that Mitigates Hacking in Reinforcement Learning from Human Feedback (RLHF)
ShareShareShareShareShare

The well-known Artificial Intelligence (AI)-based chatbot, i.e., ChatGPT, which has been built on top of GPT’s transformer architecture, uses the technique of Reinforcement Learning from Human Feedback (RLHF). RLHF is an increasingly important method for utilizing the potential of pre-trained Large Language Models (LLMs) to generate more helpful, truthful responses that are in line with human preferences.

In RLHF, a language model is trained to produce responses that maximize the learned reward through reinforcement learning, after which a reward model is trained based on human preferences for particular prompts. Since gathering human ratings is typically less complicated than gathering demos for supervised fine-tuning, this approach streamlines the process of collecting data. 

However, reward hacking is a subtle problem with RLHF, where the policy gets a large reward without meeting the real objectives. This happens as a result of the reward model’s limited Out-Of-Distribution (OOD) generalization and potential imperfections in representing human preferences. Being a strong LLM, the language model can provide OOD examples to take advantage of flaws in the reward model. 

The scenario is further complicated by human preference data, which is frequently skewed and inconsistent due to task complexity and subjectivity, defects in rating standards, and the low caliber of raters. Verbosity is a popular example of reward hacking, in which models produce more tokens to appear more thorough or better formatted in responses, but there is no real improvement in quality.

In order to address these issues, recent research from NVIDIA and the University of Maryland has aimed to mitigate reward hacking by examining how RL algorithms and incentive models affect verbosity and performance. The team has presented an evaluation technique to compare various training setups and account for biases in model-based evaluations. The technique has provided a comprehensive knowledge of various response durations by evaluating performance on the Pareto front of evaluation score vs. length. 

This process is intended to analyze the trade-off between the LLM’s assessment score and response duration, allowing for a systematic comparison of different training settings. By varying the training hyperparameters, it can be evaluated how these modifications affect the ratio of verbosity to answer quality.

The study looks at RL hyperparameters and techniques, such as reward clipping and length penalty, to lessen reward hacking on length. The primary goal is to remove the spurious length signal from the reward, even though various tuning procedures can yield better outcomes. To accomplish this, the team has suggested a two-head reward model that separates representations for length from true preferences. The length head is deleted during RL. 

The suggested reward disentangling technique, ODIN, has been used with the help of which, even with a more costly tuning budget, the policy was able to attain a larger Pareto front than prior results. Proximal Policy Optimisation (PPO) and ReMax both benefit from ODIN’s effectiveness, indicating that it can be used to enhance other RL-tuning methods and lessen length hacking.

In conclusion, this method’s experimental results have shown a noteworthy decrease in the reward model’s association with response duration. The derived strategy performs significantly better when the quality of the information is prioritized over verbosity. This method successfully reduces the problem of response length-related reward hacking, improving the dependability and utility of LLMs trained using the RLHF paradigm.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 37k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🚀 LLMWare Launches SLIMs: Small Specialized Function-Calling Models for Multi-Step Automation [Check out all the models]


Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Why Do Routers Have So Many Antennas?
AI & Technology

Why Do Routers Have So Many Antennas?

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
Next Post
Dozens killed in rebel attack on school in Uganda

Dozens killed in rebel attack on school in Uganda

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Fevertree Drinks PLC 2026 Q2 – Results – Earnings Call Presentation (OTCMKTS:FQVTY) 2026-09-10

Fevertree Drinks PLC 2026 Q2 – Results – Earnings Call Presentation (OTCMKTS:FQVTY) 2026-09-10

September 10, 2026
Patriots WR A.J. Brown ruled out vs. Seahawks with ankle injury – The New York Times

Patriots WR A.J. Brown ruled out vs. Seahawks with ankle injury – The New York Times

September 10, 2026
Ford under fire from GOP, Trump administration over China ties

Ford under fire from GOP, Trump administration over China ties

September 9, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!