• bitcoinBitcoin(BTC)$84,240.00-0.07%
  • ethereumEthereum(ETH)$2,686.420.55%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$777.351.66%
  • rippleXRP(XRP)$1.532.79%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$116.772.28%
  • tronTRON(TRX)$0.3404850.05%
  • zcashZcash(ZEC)$1,536.121.30%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.99%
  • HyperliquidHyperliquid(HYPE)$93.961.13%
  • dogecoinDogecoin(DOGE)$0.0957643.90%
  • moneroMonero(XMR)$568.603.21%
  • whitebitWhiteBIT Coin(WBT)$84.36-0.24%
  • chainlinkChainlink(LINK)$13.127.14%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2472554.14%
  • RainRain(RAIN)$0.012079-1.42%
  • leo-tokenLEO Token(LEO)$8.91-1.10%
  • stellarStellar(XLM)$0.2117465.21%
  • bitcoin-cashBitcoin Cash(BCH)$335.96-1.39%
  • nearNEAR Protocol(NEAR)$4.616.83%
  • uniswapUniswap(UNI)$9.260.93%
  • litecoinLitecoin(LTC)$71.3717.08%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.400.95%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.1143085.39%
  • USD1USD1(USD1)$1.00-0.03%
  • suiSui(SUI)$1.026.42%
  • hedera-hashgraphHedera(HBAR)$0.0926913.06%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.420.96%
  • shiba-inuShiba Inu(SHIB)$0.0000063.76%
  • BittensorBittensor(TAO)$295.382.98%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.0628893.00%
  • BitwayBitway(BTW)$1.053.87%
  • MemeCoreMemeCore(M)$1.221.15%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,262.07-0.56%
  • okbOKB(OKB)$119.541.47%
  • OndoOndo(ONDO)$0.5123.76%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.02%
  • mantleMantle(MNT)$0.684.32%
  • aaveAave(AAVE)$145.354.94%
  • EthenaEthena(ENA)$0.2208887.37%
  • polkadotPolkadot(DOT)$1.165.65%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Researchers Propose GenRM: Training Verifiers with Next-Token Prediction to Leverage the Text Generation Capabilities of LLMs

September 2, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Google DeepMind Researchers Propose GenRM: Training Verifiers with Next-Token Prediction to Leverage the Text Generation Capabilities of LLMs
ShareShareShareShareShare

Generative AI, an area of artificial intelligence, focuses on creating systems capable of producing human-like text and solving complex reasoning tasks. These models are essential in various applications, including natural language processing. Their primary function is to predict subsequent words in a sequence, generate coherent text, and even solve logical and mathematical problems. However, despite their impressive capabilities, generative AI models often need help with the accuracy and reliability of their outputs, which is particularly problematic in reasoning tasks where a single error can invalidate an entire solution.

One significant issue within this field is the tendency of generative AI models to produce outputs that, while confident and convincing, may need to be corrected. This challenge is critical in areas where precision is paramount, such as education, finance, and healthcare. The core of the problem lies in the models’ inability to consistently generate correct answers, which undermines their potential in high-stakes applications. Improving the accuracy and reliability of these AI systems is thus a priority for researchers who aim to enhance the trustworthiness of AI-generated solutions.

YOU MAY ALSO LIKE

Congressman Calls for National Data Center Strategy

New York Times Cooking Is Coming To Meta’s AI And Display Glasses

Existing methods to address these issues involve discriminative reward models (RMs), which classify potential answers as correct or incorrect based on their assigned scores. These models, however, need to fully leverage the generative abilities of large language models (LLMs). Another common approach is the LLM-as-a-Judge method, where pre-trained language models evaluate the correctness of solutions. While this method taps into the generative capabilities of LLMs, it often fails to match the performance of more specialized verifiers, particularly in reasoning tasks requiring nuanced judgment.

Researchers from Google DeepMind, University of Toronto, MILA and UCLA have introduced a novel approach called Generative Reward Modeling (GenRM). This method redefines the verification process by framing it as a next-token prediction task, a fundamental capability of LLMs. Unlike traditional discriminative RMs, GenRM integrates the text-generation strengths of LLMs into the verification process, allowing the model to generate and evaluate potential solutions simultaneously. This approach also supports Chain-of-Thought (CoT) reasoning, where the model generates intermediate reasoning steps before arriving at a final decision. The GenRM method, therefore, not only assesses the correctness of solutions but also enhances the overall reasoning process by enabling more detailed and structured evaluations.

The GenRM methodology employs a unified training approach combining solution generation and verification. This is achieved by training the model to predict the correctness of a solution through next-token prediction, a technique that leverages the inherent generative abilities of LLMs. In practice, the model generates intermediate reasoning steps—CoT rationales—which are then used to verify the final solution. This process integrates seamlessly with existing AI training techniques, allowing for the simultaneous improvement of generation and verification capabilities. Furthermore, the GenRM model benefits from additional inference-time computation, such as majority voting aggregating multiple reasoning paths to arrive at the most accurate solution.

The performance of the GenRM model, particularly when paired with CoT reasoning, significantly surpasses traditional verification methods. In a series of rigorous tests, including tasks related to grade-school math and algorithmic problem-solving, the GenRM model demonstrated a remarkable improvement in accuracy. Specifically, the researchers reported a 16% to 64% increase in the percentage of correctly solved problems compared to discriminative RMs and LLM-as-a-Judge methods. For example, when verifying outputs from the Gemini 1.0 Pro model, the GenRM approach improved the problem-solving success rate from 73% to 92.8%. This substantial performance boost highlights the model’s ability to mitigate errors that standard verifiers often overlook, particularly in complex reasoning scenarios. Furthermore, the researchers observed that the GenRM model scales effectively with increased dataset size and model capacity, further enhancing its applicability across various reasoning tasks.

In conclusion, the introduction of the GenRM method by researchers at Google DeepMind marks a significant advancement in generative AI, particularly in addressing the verification challenges associated with reasoning tasks. The GenRM model offers a more reliable and accurate approach to solving complex problems by unifying solution generation and verification into a single process. This method improves the accuracy of AI-generated solutions and enhances the overall reasoning process, making it a valuable tool for future AI applications across multiple domains. As generative AI continues to evolve, the GenRM approach provides a solid foundation for further research and development, particularly in areas where precision and reliability are crucial.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 50k+ ML SubReddit

Here is a highly recommended webinar from our sponsor: ‘Building Performant AI Applications with NVIDIA NIMs and Haystack’


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

▶• ılıılıılıılıılı Upcoming Live Session: ‘Building Performant AI Applications with NVIDIA NIMs and Haystack’.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Congressman Calls for National Data Center Strategy
AI & Technology

Congressman Calls for National Data Center Strategy

September 24, 2026
New York Times Cooking Is Coming To Meta’s AI And Display Glasses
AI & Technology

New York Times Cooking Is Coming To Meta’s AI And Display Glasses

September 24, 2026
Trump-Xi Summit Puts Global AI Race in Focus
AI & Technology

Trump-Xi Summit Puts Global AI Race in Focus

September 24, 2026
BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost
AI & Technology

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost

September 24, 2026
Next Post
Who is Usha Vance, wife of vice presidential nominee JD Vance?

Who is Usha Vance, wife of vice presidential nominee JD Vance?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Initial hope is that Caleb Williams has a Grade 1 hamstring strain – NBC Sports

Initial hope is that Caleb Williams has a Grade 1 hamstring strain – NBC Sports

September 21, 2026
Rapid‑Fire: Best Vs. Worst Charts In The Market

Rapid‑Fire: Best Vs. Worst Charts In The Market

September 23, 2026
Hundreds evacuated as Texas wildfire explodes

Hundreds evacuated as Texas wildfire explodes

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!