• bitcoinBitcoin(BTC)$76,344.000.57%
  • ethereumEthereum(ETH)$2,433.791.46%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$723.871.90%
  • rippleXRP(XRP)$1.290.68%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.802.77%
  • tronTRON(TRX)$0.3346940.04%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.64%
  • zcashZcash(ZEC)$1,334.2010.91%
  • HyperliquidHyperliquid(HYPE)$79.832.69%
  • dogecoinDogecoin(DOGE)$0.0808451.45%
  • USDSUSDS(USDS)$1.000.01%
  • RainRain(RAIN)$0.013372-2.64%
  • moneroMonero(XMR)$494.37-1.41%
  • whitebitWhiteBIT Coin(WBT)$78.540.72%
  • chainlinkChainlink(LINK)$11.143.20%
  • leo-tokenLEO Token(LEO)$8.930.55%
  • cardanoCardano(ADA)$0.1981102.11%
  • stellarStellar(XLM)$0.1812053.63%
  • Ethena USDeEthena USDe(USDE)$1.000.05%
  • daiDai(DAI)$1.00-0.02%
  • bitcoin-cashBitcoin Cash(BCH)$224.292.66%
  • USD1USD1(USD1)$1.00-0.02%
  • uniswapUniswap(UNI)$6.827.66%
  • litecoinLitecoin(LTC)$52.673.69%
  • CantonCanton(CC)$0.10168811.58%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.320.34%
  • nearNEAR Protocol(NEAR)$2.7815.51%
  • avalanche-2Avalanche(AVAX)$7.513.54%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0744790.34%
  • shiba-inuShiba Inu(SHIB)$0.0000053.07%
  • suiSui(SUI)$0.724.62%
  • crypto-com-chainCronos(CRO)$0.0576403.73%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,307.96-0.72%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$226.534.81%
  • MemeCoreMemeCore(M)$1.132.39%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$111.561.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • AsterAster(ASTER)$0.749.79%
  • BitwayBitway(BTW)$0.70-9.29%
  • aaveAave(AAVE)$122.122.19%
  • pax-goldPAX Gold(PAXG)$4,310.00-0.81%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0589203.63%
  • mantleMantle(MNT)$0.562.85%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from China Introduces ShortGPT: A Novel Artificial Intelligence Approach to Pruning Large Language Models (LLMs) based on Layer Redundancy

March 11, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper from China Introduces ShortGPT: A Novel Artificial Intelligence Approach to Pruning Large Language Models (LLMs) based on Layer Redundancy
ShareShareShareShareShare

Recent advancements in Large Language Models (LLMs) have led to models containing billions or even trillions of parameters, achieving remarkable performance across domains. However, their massive size poses challenges in practical deployment due to stringent hardware requirements. Research has focused on scaling models to enhance performance, guided by established scaling laws. This escalation underscores the need to address hardware limitations to facilitate the widespread utilization of these powerful LLMs.

Prior works address the challenge of deploying massive trained models by focusing on model compression techniques. These techniques, including quantization and pruning, aim to reduce inference costs. While quantization lowers precision, pruning removes redundant parameters without retraining. Recent advancements in pruning techniques have shown promise in simplifying model compression for large language models, highlighting the importance of exploring efficient pruning approaches tailored for such models.

The researchers from Baichuan Inc. and the Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences, present a unique approach, ShortGPT, to analyze layer-wise redundancy in LLMs using Block Influence (BI), measuring hidden state transformations. Their method significantly outperforms previous complex pruning techniques by identifying and removing redundant layers based on BI scores. They demonstrate that LLMs exhibit substantial layer redundancy, offering a straightforward yet effective pruning strategy. This method, orthogonal to quantization, reduces parameters and computation while maintaining high performance, paving the way for more efficient LLM training.

Their proposed LLM layer deletion approach begins by quantifying layer redundancy, particularly in Transformer-based architectures. BI metric assesses each layer’s impact on hidden state transformations during inference. Layers with low BI scores, indicating minimal impact, are removed to reduce inference costs without compromising model performance. The method involves constructing a calibration set, collecting hidden states, calculating BI scores, and iteratively deleting less important layers based on BI rankings.

The proposed method’s comparative experiments against benchmarks (including MMLU, CMMLU, and CMNLI) and baseline techniques (including LLMPru, SliceGPT, and LaCo) are commonly used in LLM evaluation. Results show that the model pruned using the proposed approach consistently outperforms baseline methods across multiple natural language benchmarks. Also, reducing the number of layers proves more effective than reducing embedding dimensions, indicating deeper redundancy within the models.

In conclusion, the researchers from Baichuan Inc. and the Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences present ShortGPT, a unique LLM pruning approach based on layer redundancy and attention entropy. Results show significant layer-wise redundancy in LLMs, enabling the removal of minimally contributing layers without compromising performance. The proposed strategy maintains up to 95% of model performance while reducing parameter count and computational requirements by around 25%, surpassing previous pruning methods. This approach, simple yet effective, suggests depth-based redundancy in LLMs and offers compatibility with other compression techniques for versatile model size reduction.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 38k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel

You may also like our FREE AI Courses….


YOU MAY ALSO LIKE

Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🚀 [FREE AI WEBINAR] ‘Building with Google’s New Open Gemma Models’ (March 11, 2024) [Promoted]


Credit: Source link

ShareTweetSendSharePin

Related Posts

Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI
AI & Technology

Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI

September 17, 2026
OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training
AI & Technology

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training

September 17, 2026
An iOS 27 Bug Can Temporarily Freeze Your iPhone
AI & Technology

An iOS 27 Bug Can Temporarily Freeze Your iPhone

September 17, 2026
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
AI & Technology

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

September 17, 2026
Next Post
You Now Have a New Rule…

You Now Have a New Rule...

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Universal Music Group Is Collaborating With ElevenLabs On A New AI-Powered Creation Platform

Universal Music Group Is Collaborating With ElevenLabs On A New AI-Powered Creation Platform

September 10, 2026
Amazon, Qualcomm Deal Broadens the AI Chip Race | Bloomberg Tech 9/08/2026

Amazon, Qualcomm Deal Broadens the AI Chip Race | Bloomberg Tech 9/08/2026

September 12, 2026
Mail-in ballot fight heads to Supreme Court for third time

Mail-in ballot fight heads to Supreme Court for third time

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!