• bitcoinBitcoin(BTC)$76,430.000.77%
  • ethereumEthereum(ETH)$2,439.261.78%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$724.351.61%
  • rippleXRP(XRP)$1.300.40%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.612.73%
  • tronTRON(TRX)$0.3354200.23%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.65%
  • zcashZcash(ZEC)$1,355.2815.61%
  • HyperliquidHyperliquid(HYPE)$79.142.49%
  • dogecoinDogecoin(DOGE)$0.0809241.19%
  • USDSUSDS(USDS)$1.000.02%
  • moneroMonero(XMR)$498.61-1.46%
  • whitebitWhiteBIT Coin(WBT)$78.630.94%
  • RainRain(RAIN)$0.012862-8.15%
  • chainlinkChainlink(LINK)$11.153.40%
  • leo-tokenLEO Token(LEO)$8.940.53%
  • cardanoCardano(ADA)$0.1970811.24%
  • stellarStellar(XLM)$0.1822723.42%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$220.950.52%
  • USD1USD1(USD1)$1.00-0.02%
  • uniswapUniswap(UNI)$6.706.35%
  • litecoinLitecoin(LTC)$52.182.31%
  • CantonCanton(CC)$0.0989608.94%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.320.23%
  • nearNEAR Protocol(NEAR)$2.7015.46%
  • avalanche-2Avalanche(AVAX)$7.522.93%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.073741-0.92%
  • suiSui(SUI)$0.724.85%
  • shiba-inuShiba Inu(SHIB)$0.0000051.96%
  • crypto-com-chainCronos(CRO)$0.0584205.46%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,308.80-0.30%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$224.893.87%
  • MemeCoreMemeCore(M)$1.121.73%
  • okbOKB(OKB)$111.640.65%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.27%
  • AsterAster(ASTER)$0.737.64%
  • BitwayBitway(BTW)$0.72-7.56%
  • aaveAave(AAVE)$122.291.66%
  • pax-goldPAX Gold(PAXG)$4,310.33-0.38%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0587903.16%
  • mantleMantle(MNT)$0.562.65%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from ETH Zurich, EPFL, and Microsoft Introduce QuaRot: A Machine Learning Method that Enables 4-bit Inference of LLMs by Removing the Outlier Features

April 5, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Researchers from ETH Zurich, EPFL, and Microsoft Introduce QuaRot: A Machine Learning Method that Enables 4-bit Inference of LLMs by Removing the Outlier Features
ShareShareShareShareShare

Large language models (LLMs) have revolutionized various applications across industries by providing advanced natural language processing capabilities. These models’ ability to generate, understand, and interpret human language has opened new avenues for technological advancements. However, their significant computational, memory, and energy demands hinder LLMs’ deployment and operational efficiency, especially during the inference phase. The challenge stems from the extensive number of parameters within these models, necessitating considerable data storage and manipulation resources.

Researchers have turned to quantization to tackle these issues. This process reduces the precision of the model’s parameters to achieve lower memory consumption and faster computation times. However, a persistent challenge in this process is the presence of outliers within the data. These outliers can drastically affect the model’s accuracy when significantly reduced precision. 

QuaRot is a breakthrough approach by researchers from ETH Zurich, EPFL, Microsoft Research, IST Austria, and NeuralMagic. It offers a promising solution by applying a novel quantization scheme based on rotations to mitigate the effects of outliers. It’s an innovative technique that employs randomized Hadamard transformations and leverages computational invariance, a principle ensuring that these transformations do not alter the final output of the model. This method allows for a comprehensive 4-bit quantization encompassing all model components, including weights, activations, and the key-value (KV) cache. By doing so, QuaRot significantly diminishes the model’s computational and memory requirements.

The efficacy of QuaRot is underscored by its performance on the LLAMA 2-70B model. The method achieved remarkable outcomes, demonstrating that a quantized model could retain up to 99% of its zero-shot performance capabilities post-quantization. The approach enabled up to 2.16 times speedup during the prefill phase of inference, a stage traditionally known for being compute-bound. It also facilitated a substantial reduction in memory usage, achieving up to 3.39 times savings during the decoding stage, a phase typically memory-bound. These improvements are pivotal, as they reduce operational costs and energy consumption associated with running such advanced models.

By enabling end-to-end 4-bit inference without significant performance loss, the method allows for the broader adoption and deployment of LLMs across various devices, including those with limited computational resources. This access to advanced language models holds the potential to drive innovation and expand the applicability of LLMs in sectors where computational resources are a limiting factor.

In conclusion, QuaRot marks a significant leap forward in optimizing large language models. QuaRot successfully addresses the longstanding challenge of efficiently quantizing LLMs while maintaining high accuracy through its innovative use of randomized Hadamard transformations and computational invariance. The method’s ability to significantly reduce memory usage and computational demands is evidenced by its LLAMA 2-70B model performance.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 39k+ ML SubReddit


YOU MAY ALSO LIKE

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
AI & Technology

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

September 17, 2026
House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI
AI & Technology

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

September 16, 2026
Snap Introduces A Standalone AI Assistant, Specs Intelligence
AI & Technology

Snap Introduces A Standalone AI Assistant, Specs Intelligence

September 16, 2026
Standalone AR Glasses Are Here
AI & Technology

Standalone AR Glasses Are Here

September 16, 2026
Next Post
One year since WSJ journalist Evan Gershkovich was detained in Russia

One year since WSJ journalist Evan Gershkovich was detained in Russia

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
OpenAI Launches Misalignment Reporting Framework With Six Incident Reports – Unite.AI

OpenAI Launches Misalignment Reporting Framework With Six Incident Reports – Unite.AI

September 16, 2026
US Air Force colonel shot down over Iran recounts 'modern-day miracle' – Fox News

US Air Force colonel shot down over Iran recounts 'modern-day miracle' – Fox News

September 14, 2026
Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!