• bitcoinBitcoin(BTC)$76,415.000.65%
  • ethereumEthereum(ETH)$2,448.051.94%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$733.631.88%
  • rippleXRP(XRP)$1.290.49%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.173.03%
  • tronTRON(TRX)$0.334808-0.38%
  • zcashZcash(ZEC)$1,468.6311.17%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.19%
  • HyperliquidHyperliquid(HYPE)$83.156.23%
  • dogecoinDogecoin(DOGE)$0.0814881.42%
  • moneroMonero(XMR)$513.023.95%
  • USDSUSDS(USDS)$1.000.03%
  • whitebitWhiteBIT Coin(WBT)$78.801.15%
  • RainRain(RAIN)$0.012801-2.94%
  • chainlinkChainlink(LINK)$11.343.95%
  • leo-tokenLEO Token(LEO)$8.930.33%
  • cardanoCardano(ADA)$0.2009693.27%
  • stellarStellar(XLM)$0.1841902.15%
  • uniswapUniswap(UNI)$7.7420.39%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$233.056.91%
  • daiDai(DAI)$1.00-0.04%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.835.05%
  • CantonCanton(CC)$0.1001154.77%
  • nearNEAR Protocol(NEAR)$3.0317.75%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.78%
  • avalanche-2Avalanche(AVAX)$7.603.07%
  • hedera-hashgraphHedera(HBAR)$0.0750092.75%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • shiba-inuShiba Inu(SHIB)$0.0000054.13%
  • suiSui(SUI)$0.734.08%
  • crypto-com-chainCronos(CRO)$0.0573383.18%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,343.321.63%
  • MemeCoreMemeCore(M)$1.185.94%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$229.324.59%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$112.441.92%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.33%
  • AsterAster(ASTER)$0.757.71%
  • aaveAave(AAVE)$128.429.98%
  • pax-goldPAX Gold(PAXG)$4,343.061.62%
  • BitwayBitway(BTW)$0.70-5.85%
  • mantleMantle(MNT)$0.573.23%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0583041.78%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Cerebras & Neural Magic Introduce Sparse Llama: The First Production LLM based on Llama at 70% Sparsity

May 18, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Researchers from Cerebras & Neural Magic Introduce Sparse Llama: The First Production LLM based on Llama at 70% Sparsity
ShareShareShareShareShare

Natural Language Processing (NLP) is a cutting-edge field that enables machines to understand, interpret, & generate human language. It has applications in various domains, such as language translation, text summarization, sentiment analysis, and the development of conversational agents. Large language models (LLMs) have significantly advanced these applications by leveraging vast data to perform tasks with high accuracy, almost matching human performance.

Today’s primary challenge in NLP is the enormous computational and energy demands required to train and deploy these LLMs. Their sheer size often limits these models, making them expensive and less accessible to a broader audience. The high computational cost and significant energy impact restrict the usability of these models, emphasizing the need to reduce the computational footprint without compromising accuracy. Addressing this challenge is crucial for making these powerful tools more widely available and sustainable.

Various methods have been employed to mitigate these challenges and reduce LLMs’ size and computational requirements. Quantization is one technique that reduces the number of bits required to represent each model parameter, while pruning involves removing less important weights to streamline the model. However, both methods face significant hurdles in maintaining high accuracy, especially for complex tasks. Current techniques often struggle to achieve meaningful compression ratios without damaging model performance, particularly at high sparsity levels.

Researchers from Neural Magic, Cerebras Systems, and IST Austria have introduced a novel approach to create sparse foundational versions of large language models. They specifically targeted the LLaMA-2 7B model, aiming to combine the SparseGPT pruning method with sparse pretraining techniques. This innovative method seeks to achieve high sparsity levels while preserving or enhancing the model’s accuracy. The researchers’ approach involves initially pruning the model to 50% sparsity, followed by further iterative training and pruning steps to reach 70% sparsity. 

The method begins with sparse pretraining on subsets of high-quality datasets such as SlimPajama and The Stack. The sparse pretraining process includes fine-tuning with per-layer distillation, ensuring the model retains high accuracy across various complex tasks, including chat, code generation, and instruction following. This detailed process involves training the 50% sparse model until convergence and then pruning it further to achieve the 70% target. The weights are pruned and frozen, and sparsity masks are enforced during training to maintain the desired sparsity levels. This iterative process is crucial for maintaining high recovery levels after fine-tuning.

The sparse models demonstrated the ability to achieve up to 70% sparsity while fully recovering accuracy for fine-tuning tasks. Training acceleration on Cerebras CS-3 chips closely matched theoretical scaling, showcasing the efficiency of the approach. Inference speeds increased significantly, with improvements of up to 3x on CPUs using Neural Magic’s DeepSparse engine and 1.7x on GPUs using the nm-vllm engine. Additionally, the combination of sparsity and quantization resulted in total speedups on CPUs reaching up to 8.6x, highlighting the method’s efficiency and effectiveness.

The study’s results underscore the potential of combining sparsity with quantization to achieve dramatic speedups and performance gains. The sparse pretraining methodology proved particularly effective, demonstrating high recovery at up to 70% sparsity levels. The integration of Cerebras’s CS-3 AI accelerator for sparse pretraining further highlighted the advantages of this approach, enabling near-ideal speedups and significantly reducing computational requirements.

In conclusion, this research successfully addresses the challenge of reducing the computational demands of LLMs while maintaining their performance. The innovative sparse pretraining and deployment techniques introduced by the Neural Magic, Cerebras Systems, and IST Austria researchers offer a promising solution to the problem. This approach not only enhances the efficiency and accessibility of NLP models but also sets the stage for future advancements in the field.


Check out the Paper and Model. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI

Candy Crush Developers Are Planning A Strike For Next Week

Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI
AI & Technology

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI

September 17, 2026
Candy Crush Developers Are Planning A Strike For Next Week
AI & Technology

Candy Crush Developers Are Planning A Strike For Next Week

September 17, 2026
Anthropic Launches Life Sciences Verification Program in Beta – Unite.AI
AI & Technology

Anthropic Launches Life Sciences Verification Program in Beta – Unite.AI

September 17, 2026
Lofi Girl Returns With A New House Music Station And Vinyl Compilation
AI & Technology

Lofi Girl Returns With A New House Music Station And Vinyl Compilation

September 17, 2026
Next Post
Hwang Lied to Banks ‘Over and Over,’ Prosecutors Say

Hwang Lied to Banks 'Over and Over,' Prosecutors Say

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

September 12, 2026
Amazon packages spill onto a highway after a crash

Amazon packages spill onto a highway after a crash

September 15, 2026
Saudi Arabian oil pipeline system hit by projectiles triggering fires – CNN

Saudi Arabian oil pipeline system hit by projectiles triggering fires – CNN

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!