• bitcoinBitcoin(BTC)$76,546.001.11%
  • ethereumEthereum(ETH)$2,459.452.89%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$726.162.01%
  • rippleXRP(XRP)$1.302.72%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.913.91%
  • tronTRON(TRX)$0.333954-0.57%
  • zcashZcash(ZEC)$1,467.8915.55%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.022.04%
  • HyperliquidHyperliquid(HYPE)$81.723.75%
  • dogecoinDogecoin(DOGE)$0.0816303.15%
  • moneroMonero(XMR)$509.283.71%
  • USDSUSDS(USDS)$1.000.04%
  • whitebitWhiteBIT Coin(WBT)$78.911.45%
  • RainRain(RAIN)$0.012898-1.94%
  • chainlinkChainlink(LINK)$11.335.75%
  • leo-tokenLEO Token(LEO)$8.920.60%
  • cardanoCardano(ADA)$0.2013404.99%
  • stellarStellar(XLM)$0.1859116.87%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$231.937.46%
  • daiDai(DAI)$1.00-0.01%
  • uniswapUniswap(UNI)$7.3418.68%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$53.465.55%
  • CantonCanton(CC)$0.10143411.84%
  • nearNEAR Protocol(NEAR)$2.9318.07%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.65%
  • avalanche-2Avalanche(AVAX)$7.594.89%
  • hedera-hashgraphHedera(HBAR)$0.0759504.35%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000058.40%
  • suiSui(SUI)$0.736.61%
  • crypto-com-chainCronos(CRO)$0.0579935.19%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,353.450.25%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$227.956.08%
  • MemeCoreMemeCore(M)$1.140.90%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$112.032.35%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • AsterAster(ASTER)$0.748.30%
  • aaveAave(AAVE)$125.839.07%
  • BitwayBitway(BTW)$0.70-7.87%
  • pax-goldPAX Gold(PAXG)$4,354.910.18%
  • mantleMantle(MNT)$0.574.42%
  • Pump.funPump.fun(PUMP)$0.0039378.14%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Large Language Model (LLM) Training Data Is Running Out. How Close Are We To The Limit?

May 14, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Large Language Model (LLM) Training Data Is Running Out. How Close Are We To The Limit?
ShareShareShareShareShare

In the quickly developing fields of Artificial Intelligence and Data Science, the volume and accessibility of training data are critical factors in determining the capabilities and potential of Large Language Models (LLMs). Large volumes of textual data are used by these models to train and improve their language understanding skills.

A recent tweet from Mark Cummins discusses how near we are to exhausting the global reservoir of text data required for training these models, given the exponential expansion in data consumption and the demanding specifications of next-generation LLMs. To explore this question, we share some textual sources currently available in different media and compare them to the increasing needs of sophisticated AI models.

  1. Web Data: Just the English text portion of the FineWeb dataset, which is a subset of the Common Crawl web data, has an astounding 15 trillion tokens. The corpus can double in size when top-notch non-English web content is added. 
  1. Code Repositories: Approximately 0.78 trillion tokens are contributed by publicly available code, such as that which is compiled in the Stack v2 dataset. While this may appear insignificant in comparison to other sources, the total amount of code worldwide is projected to be significant, amounting to tens of trillions of tokens. 
  1. Academic Publications and Patents: The total volume of academic publications and patents is approximately 1 trillion tokens, which is a sizable but unique subset of textual data.
  1. Books: With over 21 trillion tokens, digital book collections from sites like Google Books and Anna’s Archive make up a massive body of textual content. When every distinct book in the world is taken into account, the total token count rises to 400 trillion tokens. 
  1. Social Media Archives: User-generated material is hosted on platforms such as Weibo and Twitter, which together account for a token count of roughly 49 trillion. With 140 trillion tokens, Facebook stands out in particular. This is a significant but mostly unreachable resource because of privacy and ethical issues.
  1. Transcribing Audio: The training corpus gains around 12 trillion tokens from publicly accessible audio sources such as YouTube and TikTok.
  1. Private Communications: Emails and stored instant conversations add up to a massive amount of text data, roughly 1,800 trillion tokens when added together. Access to this data is limited, which raises privacy and ethical questions.

There are ethical and logistical obstacles to future growth as the current LLM training datasets get close to the 15 trillion token level, which represents the amount of high-quality English text that is available. Reaching out to other resources like books, audio transcriptions, and different language corpora could result in small improvements, possibly increasing the maximum amount of readable, high-quality text to 60 trillion tokens. 

However, token counts in private data warehouses run by Google and Facebook go into the quadrillions outside the purview of ethical business ventures. Because of the limitations imposed by limited and morally acceptable text sources, the future course of LLM development depends on the creation of synthetic data. Since access to private data reservoirs is prohibited, data synthesis appears to be a key future direction for AI research. 

In conclusion, there is an urgent need for unique ways of LLM teaching, given the combination of growing data needs and limited text resources. In order to overcome the approaching limits of LLM training data, synthetic data becomes increasingly important as existing datasets get closer to saturation. This paradigm shift draws attention to how the field of AI research is changing and forces a deliberate turn towards synthetic data synthesis in order to maintain ongoing advancement and ethical compliance.


YOU MAY ALSO LIKE

Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches

OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


[Recommended Read] Rightsify’s GCX: Your Go-To Source for High-Quality, Ethically Sourced, Copyright-Cleared AI Music Training Datasets with Rich Metadata


Credit: Source link

ShareTweetSendSharePin

Related Posts

Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches
AI & Technology

Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches

September 17, 2026
OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI
AI & Technology

OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI

September 17, 2026
Europe’s EU Kids Act Would Ban Social Media Access For Children Under 13
AI & Technology

Europe’s EU Kids Act Would Ban Social Media Access For Children Under 13

September 17, 2026
NVIDIA And Google’s New Coalition Wants To Speed Up AI Data Center Power Grid Connections
AI & Technology

NVIDIA And Google’s New Coalition Wants To Speed Up AI Data Center Power Grid Connections

September 17, 2026
Next Post
‘The cries have been heard’: Father of school shooting victim praises Crumbley conviction

'The cries have been heard': Father of school shooting victim praises Crumbley conviction

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
S&P 500 ends higher as strong inflation data cements rate-hike bets – Reuters

S&P 500 ends higher as strong inflation data cements rate-hike bets – Reuters

September 11, 2026
Paramount’s CEO has ‘tricks up his sleeve’ to get his Warner acquisition cleared in California

Paramount’s CEO has ‘tricks up his sleeve’ to get his Warner acquisition cleared in California

September 11, 2026
Philippine coast guard rescues passengers after ferry fire

Philippine coast guard rescues passengers after ferry fire

September 14, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!