• bitcoinBitcoin(BTC)$75,643.00-1.53%
  • ethereumEthereum(ETH)$2,392.97-3.08%
  • tetherTether(USDT)$1.00-0.04%
  • binancecoinBNB(BNB)$708.62-0.89%
  • rippleXRP(XRP)$1.28-7.75%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$96.73-3.57%
  • tronTRON(TRX)$0.334887-0.78%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-2.43%
  • zcashZcash(ZEC)$1,170.303.10%
  • HyperliquidHyperliquid(HYPE)$77.44-1.56%
  • dogecoinDogecoin(DOGE)$0.079441-3.64%
  • RainRain(RAIN)$0.0138103.05%
  • USDSUSDS(USDS)$1.00-0.04%
  • moneroMonero(XMR)$508.13-1.23%
  • whitebitWhiteBIT Coin(WBT)$77.73-2.20%
  • leo-tokenLEO Token(LEO)$8.89-0.78%
  • chainlinkChainlink(LINK)$10.73-5.27%
  • cardanoCardano(ADA)$0.193559-4.88%
  • stellarStellar(XLM)$0.174969-8.28%
  • Ethena USDeEthena USDe(USDE)$1.00-0.06%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$218.69-0.78%
  • USD1USD1(USD1)$1.00-0.05%
  • litecoinLitecoin(LTC)$50.94-3.04%
  • uniswapUniswap(UNI)$6.25-4.98%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-1.73%
  • CantonCanton(CC)$0.090906-4.32%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.074230-2.87%
  • avalanche-2Avalanche(AVAX)$7.23-3.04%
  • nearNEAR Protocol(NEAR)$2.32-1.59%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.71%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • suiSui(SUI)$0.68-2.89%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,323.581.29%
  • crypto-com-chainCronos(CRO)$0.054932-3.85%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-2.39%
  • BittensorBittensor(TAO)$215.45-4.68%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$110.21-1.95%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.05%
  • BitwayBitway(BTW)$0.777.99%
  • pax-goldPAX Gold(PAXG)$4,328.641.35%
  • aaveAave(AAVE)$119.20-5.70%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057073-0.04%
  • AsterAster(ASTER)$0.67-2.26%
  • mantleMantle(MNT)$0.54-2.77%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Introduces ReasonEval: A New Machine Learning Method to Evaluate Mathematical Reasoning Beyond Accuracy

April 11, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper Introduces ReasonEval: A New Machine Learning Method to Evaluate Mathematical Reasoning Beyond Accuracy
ShareShareShareShareShare

Mathematical reasoning is vital for problem-solving and decision-making, particularly in large language models (LLMs). Evaluating LLMs’ mathematical reasoning usually focuses on the final result rather than the reasoning process intricacies. Current methodologies, like the OpenLLM leaderboard, primarily use overall accuracy, potentially overlooking logical errors or inefficient steps. Enhanced evaluation approaches are necessary to uncover underlying issues and improve LLMs’ reasoning.

Existing approaches typically evaluate mathematical reasoning in LLMs by comparing final answers with ground truth and computing overall accuracy. However, some methods assess reasoning quality by comparing generated solution steps with reference ones. Despite datasets providing ground truth, diverse reasoning paths to the same answer challenge reliance on any single reference. Prompting-based methods directly ask LLMs, often GPT-4, to judge generated solutions, but their high computational cost and transparency issues hinder the practicality of iterative model development.

Researchers from Shanghai Jiao Tong University, Shanghai Artificial Intelligence Laboratory, Yale University, Carnegie Mellon University, and Generative AI Research Lab (GAIR) introduced REASONEVAL, a new approach to evaluating reasoning quality beyond final-answer accuracy. It utilizes validity and redundancy metrics to characterize reasoning steps’ quality, which is automatically assessed by accompanying LLMs. REASONEVAL relies on base models with robust mathematical knowledge, trained on high-quality labeled data, to instantiate its evaluation framework.

REASONEVAL focuses on multi-step reasoning tasks, assessing the quality of reasoning beyond final-answer accuracy. It evaluates each reasoning step for validity and redundancy, categorizing them into positive, neutral, or negative labels. Step-level scores are computed based on validity and redundancy and then aggregated to generate solution-level scores. The method utilizes various LLMs with different base models, sizes, and training strategies. Training data is sourced from PRM800K, a dataset of labeled step-by-step solutions collected by human annotators.

REASONEVAL achieves state-of-the-art performance on human-labeled datasets and can accurately detect different errors generated by perturbation. It reveals that enhanced final-answer accuracy doesn’t consistently improve the quality of reasoning steps for complex mathematical problems. The method’s assessment also aids in data selection. Observations highlight significant decreases in validity scores for logical and calculation errors, while redundancy scores remain stable. REASONEVAL distinguishes between errors affecting validity and those introducing redundancy.

In conclusion, the research introduces REASONEVAL, an effective metric for assessing reasoning step quality based on correctness and efficiency. Experimentation confirms its ability to identify diverse errors and competitive performance compared to existing methods. REASONEVAL exposes inconsistencies between final-answer accuracy and reasoning step quality while also proving effective in data selection for training.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


YOU MAY ALSO LIKE

Cohere CEO Warns Against an AI Safety ‘Cartel’

David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Cohere CEO Warns Against an AI Safety ‘Cartel’
AI & Technology

Cohere CEO Warns Against an AI Safety ‘Cartel’

September 16, 2026
David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own
AI & Technology

David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own

September 16, 2026
Carney Calls for Global Tech Body to Boost AI Guardrails
AI & Technology

Carney Calls for Global Tech Body to Boost AI Guardrails

September 16, 2026
When AI Goes Rogue, Who’s Legally Responsible?
AI & Technology

When AI Goes Rogue, Who’s Legally Responsible?

September 16, 2026
Next Post
Stay Tuned NOW with Gadi Schwartz – March 19 | NBC News NOW

Stay Tuned NOW with Gadi Schwartz - March 19 | NBC News NOW

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NASA And IBM Made An AI Model For Exploring The Moon

NASA And IBM Made An AI Model For Exploring The Moon

September 10, 2026
1984 track champ preps South LA bakery for 2028 Olympics

1984 track champ preps South LA bakery for 2028 Olympics

September 15, 2026
Mike Tomlin says he’s been working on a city in Minecraft

Mike Tomlin says he’s been working on a city in Minecraft

September 15, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!