• bitcoinBitcoin(BTC)$78,559.00-0.77%
  • ethereumEthereum(ETH)$2,489.48-0.20%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$752.471.55%
  • rippleXRP(XRP)$1.432.43%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.84-0.32%
  • tronTRON(TRX)$0.3391971.53%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.050.00%
  • zcashZcash(ZEC)$1,178.890.84%
  • HyperliquidHyperliquid(HYPE)$84.00-1.48%
  • dogecoinDogecoin(DOGE)$0.0902790.08%
  • RainRain(RAIN)$0.0166431.94%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$81.4511.54%
  • chainlinkChainlink(LINK)$12.68-2.17%
  • moneroMonero(XMR)$500.98-2.81%
  • leo-tokenLEO Token(LEO)$9.200.10%
  • cardanoCardano(ADA)$0.2245241.46%
  • stellarStellar(XLM)$0.190530-0.62%
  • bitcoin-cashBitcoin Cash(BCH)$257.43-1.21%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • uniswapUniswap(UNI)$6.87-0.42%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.1070400.86%
  • litecoinLitecoin(LTC)$54.32-1.84%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.410.47%
  • hedera-hashgraphHedera(HBAR)$0.079751-3.42%
  • avalanche-2Avalanche(AVAX)$8.01-1.86%
  • suiSui(SUI)$0.820.25%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.82%
  • nearNEAR Protocol(NEAR)$2.351.16%
  • crypto-com-chainCronos(CRO)$0.0599604.80%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.196.33%
  • tether-goldTether Gold(XAUT)$4,386.38-0.49%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$261.600.41%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.14-0.83%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • mantleMantle(MNT)$0.642.52%
  • polkadotPolkadot(DOT)$1.2211.84%
  • AsterAster(ASTER)$0.76-1.93%
  • aaveAave(AAVE)$129.26-2.41%
  • pax-goldPAX Gold(PAXG)$4,388.58-0.50%
  • OndoOndo(ONDO)$0.377411-2.75%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Can (Very) Simple Math Informs RLHF For Large Language Models LLMs? This AI Paper Says Yes!

June 7, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Can (Very) Simple Math Informs RLHF For Large Language Models LLMs? This AI Paper Says Yes!
ShareShareShareShareShare

Incorporating human input is a key component of the recent impressive improvements in large language model (LLM) capacities, such as ChatGPT and GPT-4. To use human feedback effectively, a reward model that incorporates human preferences, values, and ethical issues must first be trained. The LLMs are then adjusted using reinforcement learning under the direction of the reward model. This procedure, also known as reinforcement learning from human feedback (RLHF), successfully coordinates LLMs with human purpose, significantly enhancing the caliber of interpersonal communication. 

It isn’t easy to create a reward system that is functional and based on human preferences. It becomes very challenging when a human labeler fails to provide a numerical grade to a response or completion for a particular prompt. Instead, pairwise comparisons of completions in terms of quality are far simpler for people to make, and this approach was used in the creation of InstructGPT. In particular, a human labeler sorts the completions from highest to lowest perceived quality after being shown many completions produced by the LLMs for the same prompt.

The replies are then rewarded according to a reward model developed after training a neural network to match the ranks of human preferences nearly as feasible. Despite certain advantages, such as removing calibration problems, rankings do not adequately reflect the various reward distributions of multiple prompts. This is so that it is clear how much better one completion is than another when ranked higher. Since some RLHF prompts are open-ended or, to put it another way, reliant on the user’s history, the reward distribution might range over a wide range; thus, this worry is particularly relevant. 

🚀 JOIN the fastest ML Subreddit Community

In contrast, some prompts are closed-ended, producing responses that should receive a high or low score, resulting in an approximately two-point mass distribution for the reward distribution. Examples of the first kind of prompts include “Prove the Pythagorean theorem” and “Is chicken a dinosaur.” Examples of the second kind include “prove the Pythagorean theorem” and “write a short story about how AI will look like in 100 years.” The incentive model may only be able to assist LLMs in appropriately measuring uncertainty if they consider the subtleties of various cues.

Researchers from Stanford University, Princeton University, and the University of Pennsylvania make documentation of an unexpected phenomenon that shows how training a reward model on preference rankings can provide the same reward distribution independent of the prompts. This event, which takes place during the last stage of training, is known as reward collapse. It’s interesting to note that before this event was proved empirically, their theoretical analysis anticipated it. They demonstrate that a straightforward optimization program or even more simply, a closed-form expression may be used to infer the collapse reward distribution numerically. Their prediction of reward collapse is in very good accord with the empirical findings. 

Their second major contribution is introducing a principled strategy to prevent reward collapse using data from the same optimization program that helped forecast its occurrence. Reward collapse is undesirable because it ignores the minute distinctions between different prompts and might result in the miscalibration of human choice when LLMs are trained using reinforcement learning and the reward model. Early termination of the reward model’s training is a simple solution to this problem, but it is rather arbitrary and can be difficult to decide when to end. 

In essence, they suggest training the reward model with different utility functions based on the prompts, such that the resultant reward distribution may be either broadly scattered or tightly concentrated, depending on whether the prompt is open-ended or closed-ended. This prompt-aware technique has the obvious benefit of analytical analysis, allowing for complete customization of the reward distribution’s structure as needed. Their findings demonstrate that reward collapse may be significantly reduced by utilizing this prompt-aware technique.


Check Out The Paper and Github link. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

What Is The Purpose Of LiDAR On Your iPhone And How Do You Use It?

New Accenture Gemini Enterprise Business Group Targets Agentic AI Scaling – Unite.AI

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


Check out https://aitoolsclub.com to find 100’s of Cool AI Tools

Credit: Source link

ShareTweetSendSharePin

Related Posts

What Is The Purpose Of LiDAR On Your iPhone And How Do You Use It?
AI & Technology

What Is The Purpose Of LiDAR On Your iPhone And How Do You Use It?

September 8, 2026
New Accenture Gemini Enterprise Business Group Targets Agentic AI Scaling – Unite.AI
AI & Technology

New Accenture Gemini Enterprise Business Group Targets Agentic AI Scaling – Unite.AI

September 8, 2026
Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI
AI & Technology

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI

September 8, 2026
What Is Considered Good Speed For Home Internet And How Can You Test It?
AI & Technology

What Is Considered Good Speed For Home Internet And How Can You Test It?

September 8, 2026
Next Post
More employers moving to relax educational requirements

More employers moving to relax educational requirements

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
New social media scam uses AI sob stories to target sympathetic buyers

New social media scam uses AI sob stories to target sympathetic buyers

September 4, 2026
Trump-backed Steve Hilton on his bid to sway anti-Trump voters in California gubernatorial race

Trump-backed Steve Hilton on his bid to sway anti-Trump voters in California gubernatorial race

September 6, 2026
West Virginia hit by heavy flooding

West Virginia hit by heavy flooding

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!