• bitcoinBitcoin(BTC)$78,831.000.28%
  • ethereumEthereum(ETH)$2,495.290.13%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$741.46-1.51%
  • rippleXRP(XRP)$1.42-0.43%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$103.59-0.28%
  • tronTRON(TRX)$0.3401660.35%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-1.85%
  • zcashZcash(ZEC)$1,279.348.28%
  • HyperliquidHyperliquid(HYPE)$86.593.29%
  • dogecoinDogecoin(DOGE)$0.089147-1.03%
  • RainRain(RAIN)$0.016248-2.33%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$511.342.02%
  • whitebitWhiteBIT Coin(WBT)$81.47-0.08%
  • chainlinkChainlink(LINK)$12.02-5.28%
  • leo-tokenLEO Token(LEO)$9.18-0.23%
  • cardanoCardano(ADA)$0.217838-3.14%
  • stellarStellar(XLM)$0.185563-2.63%
  • bitcoin-cashBitcoin Cash(BCH)$258.690.28%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$54.460.14%
  • uniswapUniswap(UNI)$6.63-3.56%
  • CantonCanton(CC)$0.103643-3.93%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-1.33%
  • nearNEAR Protocol(NEAR)$2.6411.69%
  • avalanche-2Avalanche(AVAX)$7.94-0.76%
  • hedera-hashgraphHedera(HBAR)$0.078215-1.91%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.80-2.87%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.76%
  • crypto-com-chainCronos(CRO)$0.059839-0.37%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,410.770.40%
  • MemeCoreMemeCore(M)$1.18-0.70%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$258.66-1.76%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.37-0.64%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.03%
  • mantleMantle(MNT)$0.63-2.83%
  • AsterAster(ASTER)$0.75-1.98%
  • aaveAave(AAVE)$129.49-0.06%
  • Pump.funPump.fun(PUMP)$0.0047848.68%
  • polkadotPolkadot(DOT)$1.14-7.08%
  • pax-goldPAX Gold(PAXG)$4,415.120.44%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Research Introduces Atom: A Low-Bit Quantization Technique for Efficient and Accurate Large Language Model (LLM) Serving

November 8, 2023
in AI & Technology
Reading Time: 4 mins read
A A
This AI Research Introduces Atom: A Low-Bit Quantization Technique for Efficient and Accurate Large Language Model (LLM) Serving
ShareShareShareShareShare

Large Language Models are the most recent introduction in the Artificial Intelligence community, which has taken the world by storm. These models, due to their incredible capabilities, are being used by everyone, be it researchers, scientists or even students. With their human-imitating potential to answer questions, generate content, summarise text, complete codes and so on, these models have come a long way. 

LLMs are needed in a number of domains, including sentiment analysis, intelligent chatbots, and content creation. These models utilise a lot of computational power, because of which GPU resources are effectively used to increase throughput. This is done by batching several user requests, and to further improve memory efficiency and computing capacity, LLM quantisation techniques are used. However, existing quantisation approaches, like 8-bit weight-activation quantisation, don’t really take advantage of what newer GPUs can accomplish. Since the integer operators on these GPUs are 4-bit, the current quantisation techniques are not designed for maximum efficiency. 

To address this issue, a team of researchers has introduced Atom, a new method that maximises the serving throughput of LLMs. Atom is a low-bit quantisation technique created to increase throughput significantly without sacrificing precision. It uses low-bit operators and low-bit quantisation to reduce memory usage in order to achieve this. It uses a special combination of fine-grained and mixed-precision quantisation to retain excellent accuracy.

The team has shared that Atom has been evaluated in terms of 4-bit weight-activation quantisation configurations when serving. The results demonstrated that Atom can maintain latency within the same goal range while improving end-to-end throughput by up to 7.73 times when compared to the typical 16-bit floating-point (FP16) approach and 2.53 times when compared to 8-bit integer (INT8) quantisation. This makes Atom a viable solution for catering to the increasing demand for their services because it maintains the desired level of response time and greatly increases the speed at which LLMs can process requests.

The researchers have summarised the primary contributions as follows.

  1. LLM serving has been thoroughly analysed as the first step in the study’s performance analysis. The important performance benefits that come from using low-bit weight-activation quantisation approaches have been identified.
  1. A unique and precise low-bit weight-activation quantisation technique called Atom has been presented. 
  1. The team has shared that Atom employs a variety of strategies to guarantee peak performance. It uses mixed precision, which uses reduced precision for the remaining key activations and weights while maintaining accuracy for the former. Fine-grained group quantisation has been used to reduce mistakes during the quantisation process.
  1. Atom employs dynamic activation quantisation, which reduces quantisation mistakes by adjusting to the unique distribution of each input. To further improve overall performance, the method additionally takes care of the KV-cache’s quantisation. 
  1. The research has also proposed an integrated framework for long-term management (LLM) servicing. The team has codesigned an effective inference system, constructing low-bit GPU kernels and showing off Atom’s useful end-to-end throughput and latency in an actual setting.
  1. Atom’s performance has been thoroughly assessed, which shows that Atom greatly increases LLM serving throughput, with throughput gains of up to 7.7x possible at the expense of a minuscule loss of accuracy.

Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 32k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

Everything Announced During Nintendo Direct

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

Lightfield Raises M Series A Led by a16z to Accelerate Growth – Unite.AI
AI & Technology

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

September 9, 2026
Everything Announced During Nintendo Direct
AI & Technology

Everything Announced During Nintendo Direct

September 9, 2026
Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI
AI & Technology

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

September 9, 2026
Lyft Is Now Offering Waymo Rides In Nashville
AI & Technology

Lyft Is Now Offering Waymo Rides In Nashville

September 9, 2026
Next Post
Florida, Kentucky Abortion Restriction Laws Blocked By Courts

Florida, Kentucky Abortion Restriction Laws Blocked By Courts

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
First American Pope prepares for “Concert for Peace”

First American Pope prepares for “Concert for Peace”

September 3, 2026
Watches, warnings discontinued as Hurricane Lowell pulls away from Hawaii – Hawaii News Now

Watches, warnings discontinued as Hurricane Lowell pulls away from Hawaii – Hawaii News Now

September 8, 2026
Live updates: Lindsay Clancy jury deliberates for 6th day – cnn.com

Live updates: Lindsay Clancy jury deliberates for 6th day – cnn.com

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!