• bitcoinBitcoin(BTC)$78,372.00-0.06%
  • ethereumEthereum(ETH)$2,473.81-0.37%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$738.64-1.36%
  • rippleXRP(XRP)$1.41-1.03%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.93-0.23%
  • tronTRON(TRX)$0.3396090.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-1.61%
  • zcashZcash(ZEC)$1,252.128.56%
  • HyperliquidHyperliquid(HYPE)$85.371.34%
  • dogecoinDogecoin(DOGE)$0.088014-1.92%
  • RainRain(RAIN)$0.016007-3.45%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.91-0.45%
  • moneroMonero(XMR)$505.871.09%
  • chainlinkChainlink(LINK)$11.94-5.11%
  • leo-tokenLEO Token(LEO)$9.17-0.34%
  • cardanoCardano(ADA)$0.215917-2.23%
  • stellarStellar(XLM)$0.184221-2.32%
  • bitcoin-cashBitcoin Cash(BCH)$255.68-0.44%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$53.78-0.90%
  • CantonCanton(CC)$0.104212-3.11%
  • uniswapUniswap(UNI)$6.57-2.92%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-1.90%
  • avalanche-2Avalanche(AVAX)$7.89-1.28%
  • hedera-hashgraphHedera(HBAR)$0.077550-2.27%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$2.527.17%
  • suiSui(SUI)$0.79-2.67%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.43%
  • crypto-com-chainCronos(CRO)$0.059062-1.18%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.20-2.26%
  • tether-goldTether Gold(XAUT)$4,395.270.61%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$259.250.90%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.75-1.07%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.01%
  • mantleMantle(MNT)$0.62-2.53%
  • AsterAster(ASTER)$0.74-1.12%
  • aaveAave(AAVE)$127.37-1.17%
  • Pump.funPump.fun(PUMP)$0.0046685.93%
  • polkadotPolkadot(DOT)$1.12-9.62%
  • pax-goldPAX Gold(PAXG)$4,398.290.67%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Exploring New Frontiers in AI: Google DeepMind’s Research on Advancing Machine Learning with ReSTEM Self-Training Beyond Human-Generated Data

December 13, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Exploring New Frontiers in AI: Google DeepMind’s Research on Advancing Machine Learning with ReSTEM Self-Training Beyond Human-Generated Data
ShareShareShareShareShare

Large Language Models (LLMs) are transforming deep learning by demonstrating astounding powers to produce text of human caliber and perform a wide range of language tasks. Getting high-quality human data is a major barrier, even while supervised fine-tuning (SFT) using human-collected data further improves their performance on tasks of interest. This is especially taxing on intricate problem-solving assignments requiring substantial resources and specialized knowledge. To overcome this obstacle, model-generated synthetic data shows promise as a scalable and affordable solution if its quality can be guaranteed. 

Researchers from Google Deepmind and Mila in this study investigate a more straightforward scenario in which an external scalar feedback signal functions as a quality indicator for each generated sample, even if LLMs can self-evaluate created data. The research team proposes a straightforward yet effective self-training technique for language models, which involves only two skills: 1) creating samples from the model and 2) assessing these samples using a scoring mechanism. This approach allows us to study training on data created by the model. The research team utilizes the nomenclature of Reinforced Self-Training and refers to this technique as ReST𝐃𝑀 to achieve uniformity and clarity. The research team demonstrates how ReST𝐃𝑀 may be thought of as using expectation maximization for reinforcement learning. 

In particular, ReST𝐃𝑀 switches between the phases for expectation and maximization in the following way: 1. Generate (E-step): For every input context, the language model produces several output samples. After that, the research team gathers the training dataset by filtering these samples using a binary reward. 2. Improve (M-step): The original language model is supervised and fine-tuned using the training dataset from the preceding Generate phase. The next Generate phase then makes use of the adjusted model. ReST𝐃𝑀 and its variants have demonstrated efficacy in enhancing language models in many fields, such as machine translation, semantic parsing, and preference alignment.

ReST𝐃𝑀 was mostly employed in earlier studies on very small language models (up to 7B parameters), with limited scalability for bigger models. Their work intends to complement these efforts by comparing the scalability and effectiveness of synthetic data created by models to human-provided data in two challenging but understudied domains: code generation (APPS) and competition-level mathematical problem-solving (MATH). Their findings demonstrate that applying ReST𝐃𝑀 to PaLM 2 models at various sizes significantly improves mathematical reasoning and code generation skills.

Surprisingly, models refined on artificial data produced by the model outperform those trained on data supplied by humans by a large margin. Furthermore, the improvement diminishes after a few cycles of ReST𝐃𝑀, indicating the possibility of overfitting on a limited number of training cases. Moreover, models optimized using ReST𝐃𝑀 enhance pass@k and majority voting capabilities. Lastly, these refined models demonstrate enhanced performance on similar but distinct benchmarks, including Big-Bench Hard tasks, coding (HumanEval), and arithmetic problems (GSM8K and Hungarian HS finals). Lastly, ablation studies are carried out to investigate the effects of training problems, iterations, and the amount of model-generated solutions on ReST𝐸𝑀 fine-tuning.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🐝 [Free Webinar] Alexa, Upgrade my App: Integrating Voice AI into Your Strategy (Dec 15 2023)

Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI
AI & Technology

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI

September 9, 2026
Lightfield Raises M Series A Led by a16z to Accelerate Growth – Unite.AI
AI & Technology

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

September 9, 2026
Everything Announced During Nintendo Direct
AI & Technology

Everything Announced During Nintendo Direct

September 9, 2026
Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI
AI & Technology

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

September 9, 2026
Next Post
This Morning’s Top Headlines – Sept. 27

This Morning’s Top Headlines – Sept. 27

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Fetterman: ‘I am a Democrat’ but no longer believe the party is pro-Israel

Fetterman: ‘I am a Democrat’ but no longer believe the party is pro-Israel

September 2, 2026
Your Credit Card’s Purchase Protection Covers What Your Wallet Doesn’t

Your Credit Card’s Purchase Protection Covers What Your Wallet Doesn’t

September 3, 2026
Microchip Technology: This Chip Stock Is Cheaper Than It Seems (NASDAQ:MCHP)

Microchip Technology: This Chip Stock Is Cheaper Than It Seems (NASDAQ:MCHP)

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!