• bitcoinBitcoin(BTC)$77,131.000.36%
  • ethereumEthereum(ETH)$2,513.542.77%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$723.931.67%
  • rippleXRP(XRP)$1.350.44%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.812.48%
  • tronTRON(TRX)$0.338224-0.61%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.34%
  • zcashZcash(ZEC)$1,156.674.84%
  • HyperliquidHyperliquid(HYPE)$78.98-0.58%
  • dogecoinDogecoin(DOGE)$0.0838000.24%
  • RainRain(RAIN)$0.015443-2.04%
  • USDSUSDS(USDS)$1.000.02%
  • moneroMonero(XMR)$514.641.03%
  • whitebitWhiteBIT Coin(WBT)$80.100.73%
  • chainlinkChainlink(LINK)$11.49-0.33%
  • leo-tokenLEO Token(LEO)$9.15-0.41%
  • cardanoCardano(ADA)$0.204818-0.94%
  • stellarStellar(XLM)$0.1773320.81%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$226.440.76%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$52.951.06%
  • CantonCanton(CC)$0.096742-1.76%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.35%
  • uniswapUniswap(UNI)$5.99-0.24%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.41-1.45%
  • hedera-hashgraphHedera(HBAR)$0.074120-1.60%
  • nearNEAR Protocol(NEAR)$2.41-3.21%
  • shiba-inuShiba Inu(SHIB)$0.0000050.89%
  • suiSui(SUI)$0.72-1.55%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0561130.06%
  • MemeCoreMemeCore(M)$1.182.09%
  • tether-goldTether Gold(XAUT)$4,346.480.58%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$112.751.83%
  • BittensorBittensor(TAO)$233.64-2.14%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.05%
  • aaveAave(AAVE)$123.961.63%
  • mantleMantle(MNT)$0.581.42%
  • pax-goldPAX Gold(PAXG)$4,352.060.67%
  • AsterAster(ASTER)$0.68-3.46%
  • polkadotPolkadot(DOT)$1.03-7.18%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054585-3.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MosaicML Proposes Modifying Chinchilla Scaling Laws to Account for Inference Costs when Determining Optimal LLM Size

January 5, 2024
in AI & Technology
Reading Time: 4 mins read
A A
MosaicML Proposes Modifying Chinchilla Scaling Laws to Account for Inference Costs when Determining Optimal LLM Size
ShareShareShareShareShare

LLMs represent a significant leap in understanding and generating human language. These models are instrumental in various AI applications, from automated translation to conversational agents. Their development involves a delicate balance between enhancing capabilities and managing computational costs, a challenge that continues to evolve with the technology.

A central issue in LLM advancement is optimizing the model’s scale in terms of its size and training data. The goal is to improve performance without incurring prohibitive computational expenses. Increasing the model size traditionally leads to better performance but at the cost of higher training and inference expenses. Finding an efficient way to scale these models, balancing quality against computational expenditure, is a pressing concern in the field.

The prevailing approach to scaling LLMs has been guided by established scaling laws, notably the Chinchilla scaling laws developed by DeepMind. These laws provide a framework for increasing model parameters and training data to enhance quality. However, they predominantly focus on the computational costs during the training phase, overlooking the substantial expenses incurred during the model’s inference stage.

Researchers from MosaicML introduce an approach to scaling LLMs that incorporates training and inference costs. The modified Chinchilla scaling laws presented in the research aim to determine the optimal balance between model parameters, pre-training data size, and the quality of the model, factoring in the costs associated with both training and inference phases. This method significantly shifts from traditional scaling practices, prioritizing a more holistic view of computational expenses.

The methodology adopted in this study involves a comprehensive analysis of the trade-off between training and inference costs. The researchers developed a new formula to calculate the optimal size of LLMs, especially under significant inference demand. This formula suggests training models with fewer parameters for a longer duration than Chinchilla’s scaling laws previously recommended. The study aims to achieve a balance that reduces the overall computational burden without compromising the model’s performance.

The study demonstrates that smaller and more efficiently trained models become more cost-effective as inference demands increase. For example, a model with the quality of a Chinchilla-7B, under high inference demand, can be optimally trained with fewer parameters and more data. This strategic adjustment substantially reduces total computational costs, making the deployment of LLMs more efficient and economically viable.

In conclusion, this research presents several key highlights:

  • A modification of the Chinchilla scaling laws, integrating inference costs into the model scaling equation.
  • A strategic recommendation is to train smaller models for longer periods, optimizing for high inference demands.
  • Demonstrated cost-efficiency with smaller models under high inference loads, reducing overall computational expenses.
  • A pivotal step towards more resource-efficient AI, enhancing the sustainability of large language model development.

Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, LinkedIn Group, Twitter, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🐝 Get stunning professional headshots effortlessly with Aragon- TRY IT NOW!.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
AI & Technology

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

September 11, 2026
Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak
AI & Technology

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

September 11, 2026
New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
AI & Technology

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
Next Post
Pandas set to leave National Zoo and return to China after more than five decades

Pandas set to leave National Zoo and return to China after more than five decades

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Federal investigators probe Amazon cargo jet’s fiery runway crash that killed 5 in Miami – AP News

Federal investigators probe Amazon cargo jet’s fiery runway crash that killed 5 in Miami – AP News

September 7, 2026
Star Group: Overlooked, Outperforming, And Cheap, What's Not To Like?

Star Group: Overlooked, Outperforming, And Cheap, What's Not To Like?

September 6, 2026
Smithsonian chief to retire as Trump fights to reshape the institution – The Washington Post

Smithsonian chief to retire as Trump fights to reshape the institution – The Washington Post

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!