• bitcoinBitcoin(BTC)$84,288.00-2.56%
  • ethereumEthereum(ETH)$2,664.25-3.09%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$763.38-3.30%
  • rippleXRP(XRP)$1.51-4.19%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$114.43-2.71%
  • tronTRON(TRX)$0.339534-0.60%
  • zcashZcash(ZEC)$1,570.041.30%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.35%
  • HyperliquidHyperliquid(HYPE)$93.37-2.01%
  • dogecoinDogecoin(DOGE)$0.093446-6.79%
  • moneroMonero(XMR)$550.27-4.01%
  • whitebitWhiteBIT Coin(WBT)$84.55-2.65%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.25-6.30%
  • cardanoCardano(ADA)$0.238945-5.65%
  • RainRain(RAIN)$0.012492-7.09%
  • leo-tokenLEO Token(LEO)$8.980.01%
  • stellarStellar(XLM)$0.203902-4.99%
  • bitcoin-cashBitcoin Cash(BCH)$340.204.38%
  • nearNEAR Protocol(NEAR)$4.461.94%
  • uniswapUniswap(UNI)$9.260.14%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • litecoinLitecoin(LTC)$60.01-2.62%
  • avalanche-2Avalanche(AVAX)$10.36-6.96%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.107813-7.03%
  • hedera-hashgraphHedera(HBAR)$0.090456-7.62%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-2.55%
  • suiSui(SUI)$0.96-5.05%
  • shiba-inuShiba Inu(SHIB)$0.000006-6.45%
  • BittensorBittensor(TAO)$293.11-6.64%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.061478-7.91%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.23-6.23%
  • tether-goldTether Gold(XAUT)$4,285.88-0.99%
  • BitwayBitway(BTW)$0.9610.30%
  • okbOKB(OKB)$118.23-3.27%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.13%
  • aaveAave(AAVE)$140.26-2.57%
  • mantleMantle(MNT)$0.65-0.85%
  • EthenaEthena(ENA)$0.203578-1.94%
  • OndoOndo(ONDO)$0.414072-4.10%
  • Pump.funPump.fun(PUMP)$0.004070-8.51%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Neural Magic Releases Fully Quantized FP8 Version of Meta’s Llama 3.1 405B Model: FP8 Dynamic Quantization and FP8 Static Quantization

July 30, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Neural Magic Releases Fully Quantized FP8 Version of Meta’s Llama 3.1 405B Model: FP8 Dynamic Quantization and FP8 Static Quantization
ShareShareShareShareShare

Neural Magic has recently announced a significant breakthrough in AI model compression, introducing a fully quantized FP8 version of Meta’s Llama 3.1 405B model. This achievement marks a milestone in AI, allowing the massive 405 billion parameter model to fit seamlessly on any 8xH100 or 8xA100 system without the common out-of-memory (OOM) errors typically encountered with the original FP8 and FP16 versions. The new model solves memory constraints and enhances inference speeds by over 2X, leveraging faster memory and computing capabilities and eliminating the need for CPU offloading or distribution across multiple nodes.

Neural Magic provides two key versions of the model:

YOU MAY ALSO LIKE

Never Use ChatGPT For These Five Tasks

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

The fully quantized FP8 version, Meta-Llama-3.1-405B-Instruct-FP8-dynamic, maintains the architecture of Meta-Llama-3.1, designed for an assistant-like chat in multiple languages. However, it is restricted to usage in English and for lawful applications only. Released under version 1.0, this model was developed by Neural Magic and operates under the llama3.1 license.

Quantization and Optimization

The model achieves remarkable efficiency through weight and activation quantization to the FP8 data type. This process reduces the number of bits per parameter from 16 to 8, halving the disk size and GPU memory requirements. Consequently, the model can be loaded and evaluated on a single node of 8xH100 GPUs instead of requiring multiple nodes.

The quantization process involves symmetric per-channel quantization, where a linear scaling per output dimension maps the FP8 representations of the quantized weights and activations. Activations are quantized dynamically on a per-token basis. This was accomplished using LLM Compressor with 512 sequences from UltraChat, ensuring optimal performance.

Deployment and Evaluation

Neural Magic’s quantized model can be deployed efficiently using the vLLM backend. The deployment process involves using the `vllm` and `transformers` libraries in Python, as demonstrated in the provided code snippets. The example highlights the integration of the model with vLLM, showcasing the ease of generating text using the optimized model.

The model was evaluated on several benchmarks, including MMLU, ARC-Challenge, GSM-8K, Hellaswag, Winogrande, and TruthfulQA. The evaluation utilized Neural Magic’s fork of the ‘lm-evaluation-harness’ and the vLLM engine. The quantized model, Meta-Llama-3.1-405B-Instruct-FP8-dynamic, achieved an average score of 86.55 on the OpenLLM benchmark, closely mirroring the unquantized model’s score of 86.63, demonstrating a near-perfect recovery of 99.91%.

Reproduction and Accuracy

Neural Magic provides detailed commands for reproducing the evaluation results across various benchmarks. These commands illustrate the robustness of the quantized model, maintaining high accuracy across different tasks and few-shot settings. For instance, the model achieved a 99.91% recovery rate on MMLU (5-shot) and 100.2% on Winogrande (5-shot), underscoring its reliability and precision.

Conclusion

In conclusion, the release of the fully quantized FP8 version of Meta’s Llama 3.1 405B model by Neural Magic by effectively reducing memory requirements and enhancing inference speeds, this model opens new avenues for efficient and scalable AI applications. The success of this quantization effort, with minimal loss in accuracy, highlights the potential for further innovations in the field, making powerful AI models more accessible & practical for various users.


Check out the FP8 Dynamic Quantization and FP8 Static Quantization. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 47k+ ML SubReddit

Find Upcoming AI Webinars here


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Never Use ChatGPT For These Five Tasks
AI & Technology

Never Use ChatGPT For These Five Tasks

September 23, 2026
Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes
AI & Technology

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

September 23, 2026
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
AI & Technology

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

September 23, 2026
Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning
AI & Technology

Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning

September 23, 2026
Next Post
Columbia cancels main commencement amid pro-Palestinian protests

Columbia cancels main commencement amid pro-Palestinian protests

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How To Use Xbox Mode On Your Windows PC

How To Use Xbox Mode On Your Windows PC

September 20, 2026
Kroger yanks Red Bull nationwide as energy drink’s premium price comes under fire

Kroger yanks Red Bull nationwide as energy drink’s premium price comes under fire

September 18, 2026
This Robot Looks DEMONIC… On Purpose?

This Robot Looks DEMONIC… On Purpose?

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!