• bitcoinBitcoin(BTC)$78,292.001.94%
  • ethereumEthereum(ETH)$2,520.811.71%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$721.420.71%
  • rippleXRP(XRP)$1.425.90%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$102.763.06%
  • tronTRON(TRX)$0.338276-0.18%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,166.039.29%
  • HyperliquidHyperliquid(HYPE)$80.243.42%
  • dogecoinDogecoin(DOGE)$0.0837921.39%
  • RainRain(RAIN)$0.014302-5.77%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$512.91-1.25%
  • whitebitWhiteBIT Coin(WBT)$81.001.72%
  • chainlinkChainlink(LINK)$11.532.74%
  • leo-tokenLEO Token(LEO)$9.00-0.46%
  • cardanoCardano(ADA)$0.2091922.43%
  • stellarStellar(XLM)$0.1913987.80%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$223.821.11%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.24-1.05%
  • uniswapUniswap(UNI)$6.576.41%
  • CantonCanton(CC)$0.0971362.28%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.350.56%
  • hedera-hashgraphHedera(HBAR)$0.0777193.40%
  • avalanche-2Avalanche(AVAX)$7.583.70%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • nearNEAR Protocol(NEAR)$2.487.33%
  • shiba-inuShiba Inu(SHIB)$0.0000051.73%
  • suiSui(SUI)$0.722.92%
  • crypto-com-chainCronos(CRO)$0.0594553.73%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,290.33-1.11%
  • BittensorBittensor(TAO)$232.980.09%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-3.44%
  • okbOKB(OKB)$113.481.27%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.06%
  • aaveAave(AAVE)$128.472.77%
  • AsterAster(ASTER)$0.702.50%
  • mantleMantle(MNT)$0.572.70%
  • pax-goldPAX Gold(PAXG)$4,295.21-1.08%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0578802.30%
  • Pump.funPump.fun(PUMP)$0.0036964.55%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from CMU and Apple Unveils WRAP: A Game-Changer for Pre-training Language Models with Synthetic Data

February 5, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper from CMU and Apple Unveils WRAP: A Game-Changer for Pre-training Language Models with Synthetic Data
ShareShareShareShareShare

Large Language Models (LLMs) have gathered a massive amount of attention and popularity among the Artificial Intelligence (AI) community in recent months. These models have demonstrated great capabilities in tasks including text summarization, question answering, code completion, content generation, etc. 

LLMs are frequently trained on inadequate web-scraped data. Most of the time, this data is loud, unstructured, and not necessarily expressed clearly. Following the existing scaling principles, which indicate that as the size of the model increases, computational power and data quantity should also increase proportionately, comes as a challenge.

There are two main limitations. Firstly, there is the significant computational cost and time involved in pre-training. Secondly, there is the impending problem of the scarcity of high-quality data available on the Internet. In recent research, a team of researchers from Apple and Carnegie Mellon University has addressed these issues by introducing the idea of Web Rephrase Augmented Pre-training (WRAP). 

WRAP is an innovative method that makes use of an already-existing, instruction-tuned LLM. This LLM is used to paraphrase online pages into particular styles, including mimicking the tone of Wikipedia or converting text into an answer-question format. The main goal of WRAP is to improve LLMs’ pre-training by adding both genuine and artificially rephrased data. 

The primary features of WRAP are as follows:

  1. Pre-training Efficiency: Applying WRAP to the noisy C4 dataset considerably speeds up pre-training, around three times faster. This effectiveness is critical in reducing the high expenses and time commitment usually related to LLM training.
  1. Enhancement of Model Performance: WRAP makes the model perform better when run within the same computational budget. Using different subsets of the Pile, a large-scale dataset used for training and assessing LLMs reduces ambiguity by more than 10%. It improves zero-shot question-answer accuracy by over 2% for 13 different activities.
  1. Rephrasing Web Documents: WRAP uses a medium-sized LLM to paraphrase documents from the web into several styles. This method is different from creating new data because it improves already-existing content while preserving the original information’s quality and diversity.

There are two main benefits to the synthetic data produced by WRAP. Firstly, it includes a range of styles that reflect the diversity of languages used in applications farther down the line. With this diversity, the LLM is better prepared for a wider variety of real-world events. Secondly, the synthetic data rephrased is of a higher quality than the raw web-scraped data. This quality enhancement results from language that is more ordered and cohesive, as this promotes more efficient model learning.

In conclusion, WRAP is a big advancement in the field of LLM pre-training. Through the use of superior-quality, different-style synthetic data, WRAP not only expedites the training process but also improves the overall performance of LLMs. Given the abundance of low-quality web data and the resource-intensive nature of classic LLM training approaches, this approach presents a possible way forward. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

The EPA Wants To Stop Regulating Power Plant Emissions

Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🎯 [FREE AI WEBINAR] ‘Using ANN for Vector Search at Speed & Scale (Demo on AWS)’ (Feb 5, 2024)


Credit: Source link

ShareTweetSendSharePin

Related Posts

The EPA Wants To Stop Regulating Power Plant Emissions
AI & Technology

The EPA Wants To Stop Regulating Power Plant Emissions

September 14, 2026
Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data
AI & Technology

Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data

September 14, 2026
How To Force Quit On Your Windows PC
AI & Technology

How To Force Quit On Your Windows PC

September 14, 2026
NVIDIA Adds RTX PRO 5500 Blackwell GPU with 84 GB GDDR7 Memory – Unite.AI
AI & Technology

NVIDIA Adds RTX PRO 5500 Blackwell GPU with 84 GB GDDR7 Memory – Unite.AI

September 14, 2026
Next Post
Wildfire burns on Spanish island amid European heatwave

Wildfire burns on Spanish island amid European heatwave

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Paycheck-to-Paycheck And Can’t Tell Our Kids No

Paycheck-to-Paycheck And Can’t Tell Our Kids No

September 12, 2026
Anthropic CEO Says It’s Time to Slow AI Model Advances – Bloomberg.com

Anthropic CEO Says It’s Time to Slow AI Model Advances – Bloomberg.com

September 12, 2026
Your Employer’s Life Insurance Coverage Is Probably Not Enough

Your Employer’s Life Insurance Coverage Is Probably Not Enough

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!