• bitcoinBitcoin(BTC)$80,297.00-0.90%
  • ethereumEthereum(ETH)$2,573.20-1.93%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$748.76-1.50%
  • rippleXRP(XRP)$1.38-2.19%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$108.56-2.81%
  • tronTRON(TRX)$0.3406620.91%
  • zcashZcash(ZEC)$1,449.72-7.31%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.31%
  • HyperliquidHyperliquid(HYPE)$91.35-1.77%
  • dogecoinDogecoin(DOGE)$0.085073-1.98%
  • moneroMonero(XMR)$520.04-8.55%
  • whitebitWhiteBIT Coin(WBT)$81.69-1.73%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013402-3.68%
  • chainlinkChainlink(LINK)$11.99-2.52%
  • cardanoCardano(ADA)$0.219614-0.97%
  • leo-tokenLEO Token(LEO)$8.900.16%
  • stellarStellar(XLM)$0.190127-0.69%
  • uniswapUniswap(UNI)$8.72-4.35%
  • bitcoin-cashBitcoin Cash(BCH)$246.160.23%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.00%
  • nearNEAR Protocol(NEAR)$3.47-5.34%
  • litecoinLitecoin(LTC)$56.74-0.22%
  • USD1USD1(USD1)$1.000.00%
  • avalanche-2Avalanche(AVAX)$9.6212.56%
  • CantonCanton(CC)$0.104048-4.99%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.65%
  • MemeCoreMemeCore(M)$1.5823.17%
  • hedera-hashgraphHedera(HBAR)$0.0818194.33%
  • suiSui(SUI)$0.821.16%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.22%
  • crypto-com-chainCronos(CRO)$0.058400-1.17%
  • BittensorBittensor(TAO)$252.770.10%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,370.22-0.04%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$115.51-0.66%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.34%
  • aaveAave(AAVE)$136.82-3.94%
  • OndoOndo(ONDO)$0.4092943.33%
  • AsterAster(ASTER)$0.74-2.82%
  • EthenaEthena(ENA)$0.1965138.20%
  • mantleMantle(MNT)$0.59-1.87%
  • pax-goldPAX Gold(PAXG)$4,360.87-0.06%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Generalizable Reward Model (GRM): An Efficient AI Approach to Improve the Generalizability and Robustness of Reward Learning for LLMs

July 12, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Generalizable Reward Model (GRM): An Efficient AI Approach to Improve the Generalizability and Robustness of Reward Learning for LLMs
ShareShareShareShareShare

Pretrained large models have shown impressive abilities in many different fields. Recent research focuses on ensuring these models align with human values and avoid harmful behaviors. To achieve this, alignment methods are crucial, where two primary methods are supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). RLHF is useful in generalizing the reward model to new prompt-response pairs. However, it faces the challenge of training a reward model that works well with unseen data. One common problem is “overoptimization” or “reward hacking”. Increasing the size of the reward model and the amount of training data can help solve this issue, but it is not practical in real-world situations.

This paper discusses two approaches in the related work. The first approach is Reward Modeling, where reward models are trained on human preference data to guide the RLHF process or prompt optimization. Recent research focuses on developing better reward models to improve the performance of large language models (LLMs) in RLHF. This includes enhancing reward modeling by improving the quality or quantity of preference data. The second approach is Mitigating Overoptimization in RLHF, where reward models often overfit and have trouble generalizing beyond the training data, leading to the issue of overoptimization. One can penalize overly confident model outputs using label smoothing or SFT regularization to reduce this problem.

YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

How To Record Audio On Your iPhone

Researchers from HKUST, Georgia Institute of Technology, and the University of Illinois Urbana-Champaign have introduced the Generalizable Reward Model (GRM), which uses text-generation regularization on hidden states to improve the performance of reward models. Their study shows that all three types of text-generation regularization work well with GRM, with SFT regularization being the most effective and reliable solution. The results demonstrate that GRM greatly enhances the accuracy of reward models in various out-of-distribution (OOD) tasks. Moreover,  it consistently boosts the performance of RLHF and helps in reducing the problem of overoptimization.

The Unified-Feedback dataset is used for training reward models, and it is one of the largest collections of pairwise feedback datasets. All reward models are trained on a subset of 400K and 40K instances from the Unified-Feedback dataset and evaluated on an 8K-instance hold-out eval set. Moreover, while evaluating model performance on OOD preference data, datasets like HHH-Alignment, MT-Bench Human Judgements, and RewardBench are used. The HHH-Alignment dataset evaluates language models on helpfulness, honesty, and harmlessness, while the MT-Bench dataset contains human preferences for model responses to MT-bench questions. 

Here are the results after evaluating GRM:

  • GRM greatly improves the generalization ability of reward models, leading to better performance on both (in-distribution) ID and OOD evaluation sets.
  • All three types of text-generation regularization losses can enhance generalization, with SFT regularization being the most effective and consistent.
  • It shows strong performance even with limited datasets, outperforming baselines with a huge margin.
  • GRM efficiently reduces the overoptimization problem in BoN and PPO and is robust against label noise in the preference data.

In conclusion, researchers have proposed the Generalizable Reward Model (GRM), an efficient method, that aims to improve the generalizability and robustness of reward learning for LLMs. GRM uses regularization techniques on the hidden states of reward models, which significantly improves the generalization performance of reward models for unseen data. Moreover, the proposed approach effectively reduces the problem of overoptimization in RLHF. These results will support future research in creating stronger reward models, helping to align large models more efficiently and solutions with cost-effectiveness.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
How To Record Audio On Your iPhone
AI & Technology

How To Record Audio On Your iPhone

September 20, 2026
What Is The Difference Between Apple CarPlay And CarPlay Ultra?
AI & Technology

What Is The Difference Between Apple CarPlay And CarPlay Ultra?

September 19, 2026
The Pros And Cons Of Using Wired Vs. Wireless Xbox Controllers
AI & Technology

The Pros And Cons Of Using Wired Vs. Wireless Xbox Controllers

September 19, 2026
Next Post
CPI DATA INFLATION REPORT LIVESTREAM

CPI DATA INFLATION REPORT LIVESTREAM

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Reddington on the balance between empathy and the law

Reddington on the balance between empathy and the law

September 15, 2026
TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

September 19, 2026
Tibet side of border shows flood’s devastating impact

Tibet side of border shows flood’s devastating impact

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!