• bitcoinBitcoin(BTC)$77,258.00-1.91%
  • ethereumEthereum(ETH)$2,412.41-2.42%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$683.09-1.32%
  • rippleXRP(XRP)$1.35-2.84%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.81-3.38%
  • tronTRON(TRX)$0.322142-3.03%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.054.38%
  • HyperliquidHyperliquid(HYPE)$82.88-1.66%
  • zcashZcash(ZEC)$830.04-2.97%
  • dogecoinDogecoin(DOGE)$0.081393-1.93%
  • RainRain(RAIN)$0.0171502.50%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$503.70-3.72%
  • leo-tokenLEO Token(LEO)$9.39-1.96%
  • whitebitWhiteBIT Coin(WBT)$71.07-2.03%
  • chainlinkChainlink(LINK)$11.16-1.56%
  • cardanoCardano(ADA)$0.196740-1.18%
  • stellarStellar(XLM)$0.174960-1.79%
  • bitcoin-cashBitcoin Cash(BCH)$245.89-0.77%
  • daiDai(DAI)$1.00-0.02%
  • CantonCanton(CC)$0.113362-6.26%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$49.451.54%
  • uniswapUniswap(UNI)$5.8810.29%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-5.51%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.073617-0.29%
  • avalanche-2Avalanche(AVAX)$7.20-0.86%
  • shiba-inuShiba Inu(SHIB)$0.0000051.04%
  • suiSui(SUI)$0.72-1.59%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.054858-3.63%
  • tether-goldTether Gold(XAUT)$4,326.11-2.73%
  • nearNEAR Protocol(NEAR)$1.87-3.00%
  • MemeCoreMemeCore(M)$1.06-3.67%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$110.20-1.61%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.28%
  • BittensorBittensor(TAO)$218.72-5.70%
  • aaveAave(AAVE)$127.301.78%
  • AsterAster(ASTER)$0.70-0.90%
  • pax-goldPAX Gold(PAXG)$4,334.22-2.77%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057023-0.77%
  • MorphoMorpho(MORPHO)$2.613.75%
  • mantleMantle(MNT)$0.53-2.83%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet SpQR (Sparse-Quantized Representation): A Compressed Format And Quantization Technique That Enables Near-Lossless Large Language Model Weight Compression

June 9, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meet SpQR (Sparse-Quantized Representation): A Compressed Format And Quantization Technique That Enables Near-Lossless Large Language Model Weight Compression
ShareShareShareShareShare

Large Language Models (LLMs) have demonstrated incredible capabilities in recent times. Learning from massive amounts of data, these models have been performing tasks with amazing applications, including human-like textual content generation, question-answering, code completion, text summarization, creation of highly-skilled virtual assistants, and so on. Though LLMs have been performing greatly, now there has been a shift toward developing smaller models trained on even more data. Smaller models require less computational resources as compared to the larger ones; for example, the LLaMA model having 7 billion parameters and trained on 1 trillion tokens, produces results that are 25 times better than those of the much bigger GPT-3 model despite being 25 times smaller.

Compressing the LLMs so that they fit into memory-limited devices, laptops, and mobile phones accompanies challenges such as difficulty in maintaining generative quality, accuracy degradation in 3 to 4-bit quantization techniques in models with 1 to 10 Billion parameters, etc. The limitations are due to the sequential nature of LLM generation, where little errors can add up to produce outputs that are seriously damaged, to avoid which it is important to design low-bit-width quantization methods that do not reduce predictive performance compared to the original 16-bit model.

To overcome the accuracy limitations, a team of researchers has introduced Sparse-Quantized Representation (SpQR), a compressed format and quantization technique. This hybrid sparse-quantized format enables nearly lossless compression of precise pretrained LLMs down to 3–4 bits per parameter. It is the first weight quantization technique to achieve such compression ratios with an end-to-end accuracy error of less than 1% in comparison to the dense baseline, as evaluated by perplexity.

🚀 JOIN the fastest ML Subreddit Community

SpQR makes use of two ways. Firstly, it begins by locating outlier weights that, when quantized, give excessively high errors, and these weights are stored in high precision, while the remaining weights are stored in a much lower format, typically 3 bits. Secondly, SpQR employs a variant of grouped quantization with very small group size, such as 16 contiguous elements, and even the quantization scales themselves can be represented in a 3-bit format.

For converting a pretrained LLM into the SpQR format, the team has adopted an extended version of the post-training quantization (PTQ) approach, which, inspired by GPTQ, passes calibration data through the uncompressed model. SpQR allows for running 33 billion parameter LLMs on a single 24 GB consumer GPU without any performance degradation while providing a 15% speedup at 4.75 bits. This makes powerful LLMs accessible to consumers without suffering from any performance penalties.

SpQR offers effective methods for encoding and decoding weights into their format at runtime. These algorithms are made to maximize the SpQR memory compression advantages. A powerful GPU inference algorithm has also been created for SpQR, enabling faster inference than 16-bit baselines while maintaining comparable levels of accuracy. Because of this, SpQR provides memory compression benefits of more than 4x, making it very effective for use on devices with limited memory. In conclusion, SpQR seems like a promising technique as it efficiently addresses the challenge of accuracy loss associated with low-bit quantization in LLMs.


Check Out The Paper and Github. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


Check out https://aitoolsclub.com to find 100’s of Cool AI Tools

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
AI & Technology

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

September 1, 2026
Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer
AI & Technology

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

September 1, 2026
The New Street Fighter Movie Trailer Looks Fun In All The Right Ways
AI & Technology

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways

September 1, 2026
Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI
AI & Technology

Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI

September 1, 2026
Next Post
TheStreet: Avoid Macy’s, Dept. Stores are in Decline Says Cramer

TheStreet: Avoid Macy's, Dept. Stores are in Decline Says Cramer

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Meet the Press NOW — August 6

Meet the Press NOW — August 6

August 28, 2026
Full Episode: TODAY Show – Aug. 10

Full Episode: TODAY Show – Aug. 10

August 26, 2026
Inside the evacuation zone as Washington state wildfires rage

Inside the evacuation zone as Washington state wildfires rage

August 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!