• bitcoinBitcoin(BTC)$78,742.00-0.08%
  • ethereumEthereum(ETH)$2,495.370.95%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$755.452.62%
  • rippleXRP(XRP)$1.422.95%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.850.57%
  • tronTRON(TRX)$0.3392121.60%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.020.00%
  • zcashZcash(ZEC)$1,184.232.53%
  • HyperliquidHyperliquid(HYPE)$83.92-1.57%
  • dogecoinDogecoin(DOGE)$0.0904401.66%
  • RainRain(RAIN)$0.0167152.40%
  • USDSUSDS(USDS)$1.000.02%
  • chainlinkChainlink(LINK)$12.67-1.82%
  • whitebitWhiteBIT Coin(WBT)$80.1510.35%
  • moneroMonero(XMR)$496.20-6.35%
  • cardanoCardano(ADA)$0.2292575.38%
  • leo-tokenLEO Token(LEO)$9.180.32%
  • stellarStellar(XLM)$0.1921111.89%
  • bitcoin-cashBitcoin Cash(BCH)$257.19-0.88%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • uniswapUniswap(UNI)$6.901.24%
  • litecoinLitecoin(LTC)$55.02-0.72%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.1059700.04%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.411.21%
  • hedera-hashgraphHedera(HBAR)$0.080602-0.21%
  • avalanche-2Avalanche(AVAX)$8.05-0.04%
  • suiSui(SUI)$0.832.62%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000061.56%
  • nearNEAR Protocol(NEAR)$2.414.33%
  • crypto-com-chainCronos(CRO)$0.0607666.46%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,395.34-0.34%
  • MemeCoreMemeCore(M)$1.186.13%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$261.822.14%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.740.43%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.36%
  • AsterAster(ASTER)$0.76-0.72%
  • mantleMantle(MNT)$0.62-0.25%
  • polkadotPolkadot(DOT)$1.1911.49%
  • aaveAave(AAVE)$130.74-0.16%
  • pax-goldPAX Gold(PAXG)$4,398.14-0.32%
  • OndoOndo(ONDO)$0.3846941.11%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Yale and Google DeepMind Unlock Math Problem-Solving Success with Advanced Fine-Tuning Techniques on Large Language Models

October 26, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Researchers from Yale and Google DeepMind Unlock Math Problem-Solving Success with Advanced Fine-Tuning Techniques on Large Language Models
ShareShareShareShareShare

Even the most advanced large language models (LLMs), such as GPT-4 and PaLM 2, find it difficult to solve mathematical issues since they call for imagination, mathematical reasoning, and computation. The chance of LLMs being able to discover a proper answer is considerably higher when they are permitted to tackle the problem many times. Therefore, LLMs already demonstrate the potential to improve on this arithmetic problem-solving challenge. For instance, the pre-trained PaLM 2- L can reach about 33.4% accuracy with greedy decoding. However, 79.4% of the time, there is at least one accurate answer (pass@64) when sampling 64 solutions using temperature sampling (Table 1). 

Table 1: Results of the fine-tuning of supervised solutions. The MATH dataset and the PRM800K dataset, which are two different sources of training data, are contrasted.

This significant performance disparity shows that LLMs may be able to generate accurate answers but have difficulty differentiating between proper and erroneous solutions. Therefore, to narrow the performance as mentioned above difference, they investigate task-specific fine-tuning techniques that might enhance the LLM’s capacity for solution development and assessment. 

They examine three fine-tuning techniques: 

(1) SSFT, supervised step-by-step solution fine-tuning. They study if the pre-trained LLMs may profit from a supervised fine-tuning step as a starting point technique. 

They adjust the LLMs to provide the whole solution and answer. 

(2) Solution-cluster Reranking (SCR). They keep perfecting the generator as a solution evaluator for candidate solution reranking to improve the LLM’s capability to evaluate solutions. While earlier research has looked at such a solution sample-rank or reranking, they offer a novel method combining the advantages of majority voting with reranking while lowering ranking costs. To be more precise, as a preliminary stage in majority voting, they first sort the candidate replies into several groups based on their mathematical equivalency. Then, to enhance the outcomes of the majority vote even more, they apply the solution evaluator to the solutions in the most frequent clusters. 

(3) Sequential multi-tasking fine-tuning. In addition to the solution assessment task, they are also interested in enhancing the LLM’s performance on the solution-generating task and determining if the solution evaluation task’s training objective may help the model generate solutions. 

To achieve this, they provide a sequential multi-task learning environment where the solution assessment task is framed as a natural language generation problem, such that its training goal may offer a valuable supervision signal to the solution generation model. In further detail, they adjust the model in three stages: (1) as a generator (SSFT), (2) as a solution evaluator (SCR), and (3) again as a generator (SSFT). 

They do extensive research using PaLM 2-S* and PaLM 2-L, the small and big forms of PaLM 2, on the difficult MATH dataset, which results in the following conclusions: 

• Since SSFT benefits more from fine-grained, well-formatted answers, the caliber and style of the step-by-step solutions can significantly influence the refined model. 

• Reranking only the most common solution clusters can result in better performance than reranking all of the solutions, and it can also improve computational efficiency, which is why they think it would be a better standard practice for future work. 

• They demonstrate the benefit of training the model for both solution generation and evaluation tasks and present a successful attempt at leveraging the learning signal of a binary evaluation task for a generation model. Their proposed multi-task sequential fine-tuning can more effectively improve the performance of the solution generation model compared with supervised solution fine-tuning only.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 32k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on WhatsApp. Join our AI Channel on Whatsapp..


YOU MAY ALSO LIKE

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI

What Is Considered Good Speed For Home Internet And How Can You Test It?

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI
AI & Technology

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI

September 8, 2026
What Is Considered Good Speed For Home Internet And How Can You Test It?
AI & Technology

What Is Considered Good Speed For Home Internet And How Can You Test It?

September 8, 2026
What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?
AI & Technology

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

September 8, 2026
What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI
AI & Technology

What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI

September 8, 2026
Next Post
At Least 4 Dead After Alaska Teen Shoots 3 Siblings Then Themself

At Least 4 Dead After Alaska Teen Shoots 3 Siblings Then Themself

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Morning News NOW Full Episode – July 28

Morning News NOW Full Episode – July 28

September 4, 2026
Global Forex Shifts: Yen Carry Trade Unwinds Amid Policy Divergence

Global Forex Shifts: Yen Carry Trade Unwinds Amid Policy Divergence

September 7, 2026
Common Problems With Apple Wallet And How To Fix Them

Common Problems With Apple Wallet And How To Fix Them

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!