• bitcoinBitcoin(BTC)$75,643.00-1.53%
  • ethereumEthereum(ETH)$2,392.97-3.08%
  • tetherTether(USDT)$1.00-0.04%
  • binancecoinBNB(BNB)$708.62-0.89%
  • rippleXRP(XRP)$1.28-7.75%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$96.73-3.57%
  • tronTRON(TRX)$0.334887-0.78%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-2.43%
  • zcashZcash(ZEC)$1,170.303.10%
  • HyperliquidHyperliquid(HYPE)$77.44-1.56%
  • dogecoinDogecoin(DOGE)$0.079441-3.64%
  • RainRain(RAIN)$0.0138103.05%
  • USDSUSDS(USDS)$1.00-0.04%
  • moneroMonero(XMR)$508.13-1.23%
  • whitebitWhiteBIT Coin(WBT)$77.73-2.20%
  • leo-tokenLEO Token(LEO)$8.89-0.78%
  • chainlinkChainlink(LINK)$10.73-5.27%
  • cardanoCardano(ADA)$0.193559-4.88%
  • stellarStellar(XLM)$0.174969-8.28%
  • Ethena USDeEthena USDe(USDE)$1.00-0.06%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$218.69-0.78%
  • USD1USD1(USD1)$1.00-0.05%
  • litecoinLitecoin(LTC)$50.94-3.04%
  • uniswapUniswap(UNI)$6.25-4.98%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-1.73%
  • CantonCanton(CC)$0.090906-4.32%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.074230-2.87%
  • avalanche-2Avalanche(AVAX)$7.23-3.04%
  • nearNEAR Protocol(NEAR)$2.32-1.59%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.71%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • suiSui(SUI)$0.68-2.89%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,323.581.29%
  • crypto-com-chainCronos(CRO)$0.054932-3.85%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-2.39%
  • BittensorBittensor(TAO)$215.45-4.68%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$110.21-1.95%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.05%
  • BitwayBitway(BTW)$0.777.99%
  • pax-goldPAX Gold(PAXG)$4,328.641.35%
  • aaveAave(AAVE)$119.20-5.70%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057073-0.04%
  • AsterAster(ASTER)$0.67-2.26%
  • mantleMantle(MNT)$0.54-2.77%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Redefining Efficiency: Beyond Compute-Optimal Training to Predict Language Model Performance on Downstream Tasks

March 18, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Redefining Efficiency: Beyond Compute-Optimal Training to Predict Language Model Performance on Downstream Tasks
ShareShareShareShareShare

In artificial intelligence, scaling laws serve as useful guides for developing Large Language Models (LLMs). Like skilled directors, these laws coordinate models’ growth, revealing development patterns that go beyond mere computation. With each step forward, these models become more sophisticated, unlocking the intricacies of human expression with careful accuracy. Besides, scaling laws provide limitless potential for language, poised at the edge of comprehension and creation. It is usually studied in the compute-optimal training regime and predicts loss on next-token prediction.

However, there are gaps between current scaling studies and how language models are ultimately trained and evaluated. Training LLMs are expensive, and often over-trained to reduce inference costs and compare them based on downstream task performance. Training high-quality models requires a complex recipe of algorithmic techniques and training data. Researchers often use reliable extrapolation for the final training run, making it commonplace for training state-of-the-art language models such as Chinchilla 70B, PaLM 540B, and GPT-4.

Researchers from different universities experimented by creating a testbed of 104 models with 0.011B to 6.9B parameters trained with various numbers of tokens on three different data datasets: RedPajama, C4, and Refined Web to determine when scaling is predictable in the over-trained regime. This has helped predict the validation loss of a 1.4B parameter, 900B token run, and a 6.9B parameter, 138B token run. It relates the perplexity of a language model to its downstream task performance via a power law, which is used to predict top-1 error averages over downstream tasks for the two models above that take less computing time.

It has been observed that scaling laws when applied to smaller models trained closer to the compute-optimal, can effectively forecast the performance of larger models subject to more extensive over-training. However, predicting errors on individual tasks proves challenging. Hence, aggregate performance is reliably forecasted based on a model’s perplexity relative to models trained on the same dataset. During the research, it was found that, for a set of model configurations with a constant ratio of training tokens to parameters, the models’ reducible loss L′ follows consistent power laws (L′=λ·C−αc) in the amount of training computed C. So, if the ratio of tokens to parameters increases, the scaling exponent αC remains the same while the scalar λ changes.

To gauge the extent of over-training, token multipliers are used for well-known models. For instance, Chinchilla 70B is trained with a token multiplier of 20, while LLaMA-2 7B uses a token multiplier 290. Token multipliers from 5 to 640 are considered to ensure coverage of popular models and relevance for future models that may be trained on even more tokens. Analysis of data points trained on three datasets shows that exponential decay of average top-1 error as C4 eval loss on the x-axis decreases, as shown in the figure:

For the average error over 46 evaluations and the average error on a subset of 17 assessments, performance can be 10 points above random chance for at least one 0.154B scale model. These observations suggest that average top-1 error should be predictable with reliable loss estimates.

In conclusion, this research efficiently handles both the topics: scaling in the over-trained regime and downstream performance prediction. It shows that the loss scaling behavior of models trained past compute-optimal in the overtrained regime is predictable. Also, using the proposed scaling law, one can predict the downstream average task performance of more expensive runs using smaller-scale proxies. However, future development in scaling laws could focus on incorporating hyperparameters and developing an analytical theory to explain instances where scaling fails.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 38k+ ML SubReddit


YOU MAY ALSO LIKE

Cohere CEO Warns Against an AI Safety ‘Cartel’

David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Cohere CEO Warns Against an AI Safety ‘Cartel’
AI & Technology

Cohere CEO Warns Against an AI Safety ‘Cartel’

September 16, 2026
David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own
AI & Technology

David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own

September 16, 2026
Carney Calls for Global Tech Body to Boost AI Guardrails
AI & Technology

Carney Calls for Global Tech Body to Boost AI Guardrails

September 16, 2026
When AI Goes Rogue, Who’s Legally Responsible?
AI & Technology

When AI Goes Rogue, Who’s Legally Responsible?

September 16, 2026
Next Post
Adobe Q1: AI Is Not An Enemy, But An Ally – Buy The Weakness (Rating Upgrade) (ADBE)

Adobe Q1: AI Is Not An Enemy, But An Ally - Buy The Weakness (Rating Upgrade) (ADBE)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How the 9/11 Generation Is Passing Down Their Shared History | Sept. 11

How the 9/11 Generation Is Passing Down Their Shared History | Sept. 11

September 13, 2026
AI researcher says there is ‘substantial probability’ AI could kill all humans in next decade

AI researcher says there is ‘substantial probability’ AI could kill all humans in next decade

September 14, 2026
My Wife Keeps Spending Behind My Back

My Wife Keeps Spending Behind My Back

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!