• bitcoinBitcoin(BTC)$79,107.002.26%
  • ethereumEthereum(ETH)$2,540.411.27%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$727.130.82%
  • rippleXRP(XRP)$1.467.66%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.642.38%
  • tronTRON(TRX)$0.340613-0.25%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,185.447.01%
  • HyperliquidHyperliquid(HYPE)$81.844.44%
  • dogecoinDogecoin(DOGE)$0.0851740.80%
  • RainRain(RAIN)$0.014356-6.60%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$512.96-4.02%
  • whitebitWhiteBIT Coin(WBT)$81.831.88%
  • chainlinkChainlink(LINK)$11.682.24%
  • leo-tokenLEO Token(LEO)$9.00-0.58%
  • cardanoCardano(ADA)$0.2141942.30%
  • stellarStellar(XLM)$0.1940217.70%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$227.241.33%
  • USD1USD1(USD1)$1.000.03%
  • litecoinLitecoin(LTC)$54.35-1.34%
  • uniswapUniswap(UNI)$6.553.11%
  • CantonCanton(CC)$0.0972281.44%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.09%
  • hedera-hashgraphHedera(HBAR)$0.0784022.39%
  • avalanche-2Avalanche(AVAX)$7.612.38%
  • nearNEAR Protocol(NEAR)$2.568.42%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000051.94%
  • suiSui(SUI)$0.742.44%
  • crypto-com-chainCronos(CRO)$0.0594741.81%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$237.750.32%
  • tether-goldTether Gold(XAUT)$4,312.76-0.87%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.09-4.84%
  • okbOKB(OKB)$114.480.91%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.09%
  • aaveAave(AAVE)$129.311.64%
  • AsterAster(ASTER)$0.700.48%
  • mantleMantle(MNT)$0.571.24%
  • pax-goldPAX Gold(PAXG)$4,316.99-0.87%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0581352.10%
  • BitwayBitway(BTW)$0.68-2.98%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Nvidia triples and Intel doubles generative AI inference performance on new MLPerf benchmark

March 27, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Nvidia triples and Intel doubles generative AI inference performance on new MLPerf benchmark
ShareShareShareShareShare

Join us in Atlanta on April 10th and explore the landscape of security workforce. We will explore the vision, benefits, and use cases of AI for security teams. Request an invite here.


MLCommons is out today with its MLPerf 4.0 benchmarks for inference, once again showing the relentless pace of software and hardware improvements.

YOU MAY ALSO LIKE

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

As generative AI continues to develop and gain adoption, there is a clear need for a vendor-neutral set of performance benchmarks, which is what MLCommons provides with the MLPerf set of benchmarks. There are multiple MLPerf benchmarks with training and inference being among the most useful. The new MLPerf 4.0 Inference results are the first update on inference benchmarks since the MLPerf 3.1 results were released in September 2023. 

Needless to say, a lot has happened in the AI world over the last six months, and the big hardware vendors including Nvidia and Intel have been busy improving both hardware and software to further optimize inference.  The MLPerf 4.0 inference results show marked improvements for both Nvidia and Intel’s technologies.

The MLPerf inference benchmark has also changed. With the MLPerf 3.1 benchmark large language models (LLMs) were included with the GPT-J 6B (billion) parameter model to perform text summarization. With the new MLPerf 4.0 benchmark the popular Llama 2 70 billion parameter open model is being benchmarked for question and answer (Q&A). MLPerf 4 also for the first time includes a benchmark for gen AI image generation with Stable Diffusion.

VB Event

The AI Impact Tour – Atlanta

Continuing our tour, we’re headed to Atlanta for the AI Impact Tour stop on April 10th. This exclusive, invite-only event, in partnership with Microsoft, will feature discussions on how generative AI is transforming the security workforce. Space is limited, so request an invite today.

Request an invite

“MLPerf is really sort of the industry standard benchmark for helping to improve speed efficiency and accuracy for AI,” MLCommons Founder and Executive Director David Kanter said in a press briefing.

Why AI benchmarks matter

There are more than 8,500 performance results in the MLCommons’ latest benchmark, testing all manner of combinations and permutations of hardware, software and AI inference use cases. Kanter emphasized that there is a real purpose to the MLPerf benchmarking process.

“To remind people of the principle behind benchmarks. really the goal is to set up good metrics for the performance of AI,” he said. “The whole point is that once we can measure these things, we can start improving them.”

With MLCommons another goal is to help align the whole industry together. The benchmark results are all conducted on tests with similar datasets and configuration parameters across different hardware and software. The results are seen by all the submitters to a given test, such that if there are any questions from a different submitter, they can be addressed. 

Ultimately the standardized approach to measuring AI performance is about enabling enterprises to make informed decisions.

“This is helping to inform buyers, helping them make decisions and understand how systems, whether they’re on premises systems, cloud systems or embedded systems, perform on relevant workloads,” Kanter said. “If you’re looking to buy a system to run large language model inference, you can use benchmarks to help guide you, for what those systems should look like.”

Nvidia triples AI inference performance, with the same hardware

Once again, Nvidia dominates the MLPerf benchmarks with a series of impressive results.

While it’s to be expected that new hardware would yield better performance, Nvidia is also able to get better performance out of its existing hardware. Using Nvidia’s TensorRT-LLM open-source inference technology, Nvidia was able to nearly triple the inference performance for text summarization with the GPT-J LLM on its  H100 Hopper GPU.

In a briefing with press and analysts, Dave Salvator, director of accelerated computing products at Nvidia emphasized that the performance boost has occurred in only six months.

“We’ve gone in and been able to triple the amount of performance that we’re seeing and we’re very, very pleased with this result,” Salvator said. “Our engineering team just continues to do great work to find ways to extract more performance from the Hopper architecture.”

Nvidia just announced its newest generation Blackwell GPU last week at GTC, which is the successor to the Hopper architecture. In response to a question from VentureBeat, Salvator said he wasn’t sure exactly when Blackwell-based GPUs would be benchmarked for MLPerf, but he hoped it would be as soon as possible.

Even before Blackwell is benchmarked, the MLPerf 4.0 results mark the debut of H200 GPU results which further improve on the H100’s inference capabilities The H200 results are up to 45% faster than the H100 when evaluated using Llama 2 for inference.

Intel reminds industry that CPUs still matter for inference too

Intel is also a very active participant in the MLPerf 4.0 benchmarks with both its Habana AI accelerator and Xeon CPU technologies.

With Gaudi, Intel’s actual performance results trail the Nvidia H100 though the company claims it offers better price per performance. What is perhaps more interesting are the impressive gains coming from the 5th Gen Intel Xeon processor for inference.

In a briefing with press and analysts, Ronak Shah, AI product director for Xeon at Intel commented that the 5th Gen Intel Xeon was 1.42 times faster for inference than the previous 4th Gen Intel Xeon across a range of MLPerf categories. Looking specifically at just the GPT-J LLM text summarization use case, the 5th Gen Xeon was up to 1.9 times faster.

“We recognize that for many enterprise customers that are deploying their AI solutions, they’re going to be doing it in a mixed general purpose and AI environment,” Shah said. “So we designed CPUs that mesh together, strong general purpose capabilities with strong AI capabilities with our AMX engine.”

VB Daily

Stay in the know! Get the latest news in your inbox daily

By subscribing, you agree to VentureBeat’s Terms of Service.

Thanks for subscribing. Check out more VB newsletters here.

An error occured.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI
AI & Technology

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

September 14, 2026
How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error
AI & Technology

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

September 14, 2026
Temporal Raises 0M Series E at .55B Valuation to Expand Operations – Unite.AI
AI & Technology

Temporal Raises $550M Series E at $12.55B Valuation to Expand Operations – Unite.AI

September 14, 2026
What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?
AI & Technology

What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?

September 14, 2026
Next Post
Gun safety bill advances in Texas legislature

Gun safety bill advances in Texas legislature

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Watches, warnings discontinued as Hurricane Lowell pulls away from Hawaii – Hawaii News Now

Watches, warnings discontinued as Hurricane Lowell pulls away from Hawaii – Hawaii News Now

September 8, 2026
A ‘Much Larger Pullback’ Is Justified: How to Prepare

A ‘Much Larger Pullback’ Is Justified: How to Prepare

September 11, 2026
Stay Tuned NOW Streaming Behind The Scenes! – Sept 09

Stay Tuned NOW Streaming Behind The Scenes! – Sept 09

September 14, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!