• bitcoinBitcoin(BTC)$78,558.00-0.84%
  • ethereumEthereum(ETH)$2,489.130.01%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$750.561.48%
  • rippleXRP(XRP)$1.421.96%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.60-0.36%
  • tronTRON(TRX)$0.3382371.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.050.00%
  • zcashZcash(ZEC)$1,150.92-0.35%
  • HyperliquidHyperliquid(HYPE)$84.41-1.22%
  • dogecoinDogecoin(DOGE)$0.090114-0.09%
  • RainRain(RAIN)$0.0165941.80%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$81.446.26%
  • chainlinkChainlink(LINK)$12.59-1.12%
  • moneroMonero(XMR)$500.49-3.86%
  • leo-tokenLEO Token(LEO)$9.22-0.02%
  • cardanoCardano(ADA)$0.2221961.21%
  • stellarStellar(XLM)$0.189285-1.75%
  • bitcoin-cashBitcoin Cash(BCH)$257.67-0.94%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • CantonCanton(CC)$0.1082673.08%
  • USD1USD1(USD1)$1.00-0.03%
  • uniswapUniswap(UNI)$6.79-1.98%
  • litecoinLitecoin(LTC)$54.35-1.16%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.400.17%
  • hedera-hashgraphHedera(HBAR)$0.079776-2.85%
  • avalanche-2Avalanche(AVAX)$8.02-0.44%
  • suiSui(SUI)$0.82-1.11%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.60%
  • nearNEAR Protocol(NEAR)$2.351.64%
  • crypto-com-chainCronos(CRO)$0.0595174.45%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • MemeCoreMemeCore(M)$1.215.41%
  • tether-goldTether Gold(XAUT)$4,361.75-1.02%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$260.070.18%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$114.05-0.48%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.18%
  • polkadotPolkadot(DOT)$1.2821.53%
  • mantleMantle(MNT)$0.642.00%
  • AsterAster(ASTER)$0.75-1.94%
  • aaveAave(AAVE)$129.47-1.99%
  • pax-goldPAX Gold(PAXG)$4,361.21-1.12%
  • OndoOndo(ONDO)$0.379257-1.32%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Intel Researchers Propose a New Artificial Intelligence Approach to Deploy LLMs on CPUs More Efficiently

November 10, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Intel Researchers Propose a New Artificial Intelligence Approach to Deploy LLMs on CPUs More Efficiently
ShareShareShareShareShare

Large Language Models (LLMs) have taken the world by storm because of their remarkable performances and potential across a diverse range of tasks. They are best known for their capabilities in text generation, language understanding, text summarization and many more. The downside to their widespread adoption is the astronomical size of their model parameters, which requires significant memory capacity and specialized hardware for inference. As a result, deploying these models has been quite challenging.

One way the computational power required for inference could be reduced is by using quantization methods, i.e. reducing the precision of weights and activation functions of an artificial neural network. INT8 and weight-only quantization are a couple of ways the inference cost could be improved. These methods, however, are generally optimized for CUDA and may not necessarily work on CPUs.

The authors of this research paper from Intel have proposed an effective way of efficiently deploying LLMs on CPUs. Their approach supports automatic INT-4 weight-only quantization (low precision is applied to model weights only while that of activation functions is kept high) flow. They have also designed a specific LLM runtime that has highly optimized kernels that accelerate the inference process on CPUs.

The quantization flow is developed on the basis of an Intel Neural Compressor and allows for tuning on different quantization recipes, granularities, and group sizes to generate an INT4 model that meets the accuracy target. The model is then passed to the LLM runtime, a specialized environment designed to evaluate the performance of the quantized model. The runtime has been designed to provide an efficient inference of LLMs on CPUs.

For their experiments, the researchers selected some of the popular LLMs having a diverse range of parameter sizes (from 7B to 20B). They evaluated the performance of FP32 and INT4 models using open-source datasets. They observed that the accuracy of the quantized model on the selected datasets was nearly at par with that of the FP32 model. Additionally, they did a comparative analysis of the latency of the next token generation and found that the LLM runtime outperforms the ggml-based solution by up to 1.6 times.

In conclusion, this research paper presents a solution to one of the biggest challenges associated with LLMs, i.e., inference on CPUs. Traditionally, these models require specialized hardware like GPUs, which render them inaccessible for many organizations. This paper presents an INT4 model quantization along with a specialized LLM runtime to provide an efficient inference of LLMs on CPUs. When evaluated on a set of popular LLMs, the method demonstrated an advantage over ggml-based solutions and gave an accuracy on par with that of FP32 models. There is, however, scope for further improvement, and the researchers plan on empowering generative AI on PCs to meet the growing demands of AI-generated content.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 32k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

NVIDIA’s DLSS 5 Adds Subtle Details To NBA 2K27, But Demands A Lot More Power

Cognition Raises Over $2B Series E at $48B Valuation to Scale Devin Agents – Unite.AI

I am a Civil Engineering Graduate (2022) from Jamia Millia Islamia, New Delhi, and I have a keen interest in Data Science, especially Neural Networks and their application in various areas.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

NVIDIA’s DLSS 5 Adds Subtle Details To NBA 2K27, But Demands A Lot More Power
AI & Technology

NVIDIA’s DLSS 5 Adds Subtle Details To NBA 2K27, But Demands A Lot More Power

September 8, 2026
Cognition Raises Over B Series E at B Valuation to Scale Devin Agents – Unite.AI
AI & Technology

Cognition Raises Over $2B Series E at $48B Valuation to Scale Devin Agents – Unite.AI

September 8, 2026
What Is The Purpose Of LiDAR On Your iPhone And How Do You Use It?
AI & Technology

What Is The Purpose Of LiDAR On Your iPhone And How Do You Use It?

September 8, 2026
New Accenture Gemini Enterprise Business Group Targets Agentic AI Scaling – Unite.AI
AI & Technology

New Accenture Gemini Enterprise Business Group Targets Agentic AI Scaling – Unite.AI

September 8, 2026
Next Post
Former Senate Sergeant-At-Arms Michael Stenger Found Dead In His Home

Former Senate Sergeant-At-Arms Michael Stenger Found Dead In His Home

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NBC Nightly News with Tom Llamas Full Episode – July 22

NBC Nightly News with Tom Llamas Full Episode – July 22

September 6, 2026
How Microsoft Is Helping Napster Reinvent Itself

How Microsoft Is Helping Napster Reinvent Itself

September 3, 2026
Jobs Growth Update: Modest Improvement In August

Jobs Growth Update: Modest Improvement In August

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!