• bitcoinBitcoin(BTC)$79,801.000.16%
  • ethereumEthereum(ETH)$2,492.491.42%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$751.33-1.76%
  • rippleXRP(XRP)$1.420.15%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$106.303.29%
  • tronTRON(TRX)$0.3351040.56%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.061.52%
  • zcashZcash(ZEC)$1,187.6116.78%
  • HyperliquidHyperliquid(HYPE)$88.834.01%
  • dogecoinDogecoin(DOGE)$0.0895772.27%
  • RainRain(RAIN)$0.0169543.30%
  • moneroMonero(XMR)$530.51-2.21%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.293.55%
  • whitebitWhiteBIT Coin(WBT)$73.520.45%
  • leo-tokenLEO Token(LEO)$9.330.74%
  • cardanoCardano(ADA)$0.2189051.33%
  • stellarStellar(XLM)$0.1858181.12%
  • bitcoin-cashBitcoin Cash(BCH)$258.313.37%
  • daiDai(DAI)$1.000.03%
  • uniswapUniswap(UNI)$7.0112.98%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • CantonCanton(CC)$0.1096110.04%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.240.85%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-0.27%
  • hedera-hashgraphHedera(HBAR)$0.080845-0.21%
  • avalanche-2Avalanche(AVAX)$7.661.77%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • suiSui(SUI)$0.800.14%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.15%
  • nearNEAR Protocol(NEAR)$2.429.31%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0571321.28%
  • tether-goldTether Gold(XAUT)$4,420.19-0.13%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.131.31%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.33-1.14%
  • BittensorBittensor(TAO)$244.593.53%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • AsterAster(ASTER)$0.78-3.24%
  • aaveAave(AAVE)$135.273.98%
  • mantleMantle(MNT)$0.603.91%
  • pax-goldPAX Gold(PAXG)$4,426.03-0.17%
  • OndoOndo(ONDO)$0.3791192.42%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056570-0.34%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MLPerf 3.1 adds large language model benchmarks for inference

September 11, 2023
in AI & Technology
Reading Time: 5 mins read
A A
MLPerf 3.1 adds large language model benchmarks for inference
ShareShareShareShareShare

Head over to our on-demand library to view sessions from VB Transform 2023. Register Here


MLCommons is growing its suite of MLPerf AI benchmarks with the addition of testing for large language models (LLMs) for inference and a new benchmark that measures performance of storage systems for machine learning (ML) workloads.

YOU MAY ALSO LIKE

Don’t Get Rid Of Your Old Phone, Turn It Into A Security Camera

What Is Model Routing? How AI Systems Choose the Right Model for Every Request – Unite.AI

MLCommons is a vendor neutral, multi-stakeholder organization that aims to provide a level playing field for vendors to report on different aspects of AI performance with the MLPerf set of benchmarks. The new MLPerf Inference 3.1 benchmarks released today are the second major update of the results this year, following the 3.0 results that came out in April. The MLPerf 3.1 benchmarks include a large set of data with more than 13,500 performance results.

Submitters include: ASUSTeK, Azure, cTuning, Connect Tech, Dell, Fujitsu, Giga Computing, Google, H3C, HPE, IEI, Intel, Intel-Habana-Labs, Krai, Lenovo, Moffett, Neural Magic, Nvidia, Nutanix, Oracle, Qualcomm, Quanta Cloud Technology, SiMA, Supermicro, TTA and xFusion. 

Continued performance improvement

A common theme across MLPerf benchmarks with each update is the continued improvement in performance for vendors — and the MLPerf 3.1 Inference results follow that pattern. While there are multiple types of testing and configurations for the inference benchmarks, MLCommons founder and executive director David Kanter said in a press briefing that many submitters improved their performance by 20% or more over the 3.0 benchmark.

Event

VB Transform 2023 On-Demand

Did you miss a session from VB Transform 2023? Register to access the on-demand library for all of our featured sessions.

 

Register Now

Beyond continued performance gains, MLPerf is continuing to expand with the 3.1 inference benchmarks.

“We’re evolving the benchmark suite to reflect what’s going on,” he said. “Our LLM benchmark is brand new this quarter and really reflects the explosion of generative AI large language models.”

What the new MLPerf Inference 3.1 LLM benchmarks are all about

This isn’t the first time MLCommons has attempted to benchmark LLM performance.

Back in June, the MLPerf 3.0 Training benchmarks added LLMs for the first time. Training LLMs, however, is a very different task than running inference operations.

“One of the critical differences is that for inference, the LLM is fundamentally performing a generative task as it’s writing multiple sentences,” Kanter said.

The MLPerf Training benchmark for LLM makes use of the GPT-J 6B (billion) parameter model  to perform text summarization on the CNN/Daily Mail dataset. Kanter emphasized that while the MLPerf training benchmark focuses on very large foundation models, the actual task MLPerf is performing with the inference benchmark is representative of a wider set of use cases that more organizations can deploy. 

“Many folks simply don’t have the compute or the data to support a really large model,” said Kanter. “The actual task we’re performing with our inference benchmark is text summarization.”

Inference isn’t just about GPUs — at least according to Intel

While high-end GPU accelerators are often at the top of the MLPerf listing for training and inference, the big numbers are not what all organizations are looking for — at least according to Intel.

Intel silicon is well represented on the MLPerf Inference 3.1 with results submitted for Habana Gaudi accelerators, 4th Gen Intel Xeon Scalable processors and Intel Xeon CPU Max Series processors. According to Intel, the 4th Gen Intel Xeon Scalable performed well on the GPT-J news summarization task, summarizing one paragraph per second in real-time server mode.

In response to a question from VentureBeat during the Q&A portion of the MLCommons press briefing, Intel’s senior director of AI products Jordan Plawner commented that there is diversity in what organizations need for inference.

“At the end of the day, enterprises, businesses and organizations need to deploy AI in production and that clearly needs to be done in all kinds of compute,” said Plawner. “To have so many representatives of both software and hardware showing that it [inference] can be run in all kinds of compute is really a leading indicator of where the market goes next, which is now scaling out AI models, not just building them.”

Nvidia claims Grace Hopper MLPef Inference gains, with more to come

Courtesy Nvidia

While Intel is keen to show how CPUs are valuable for inference, GPUs from Nvidia are well represented in the MLPerf Inference 3.1 benchmarks.

The MLPerf Inference 3.1 benchmarks are the first time Nvidia’s GH200 Grace Hopper Superchip was included. The Grace Hopper superchip pairs an Nvidia CPU, along with a GPU to optimize AI workloads.

“Grace Hopper made a very strong first showing delivering up to 17% more performance versus our H100 GPU submissions, which we’re already delivering across the board leadership,” Dave Salvator, director of AI at Nvidia, said during a press briefing.

The Grace Hopper is intended for the largest and most demanding workloads, but that’s not all that Nvidia is going after. The Nvidia L4 GPUs were also highlighted by Salvator for their MLPerf Inference 3.1 results.

“L4  also had a very strong showing up to 6x more performance versus the best x86 CPUs submitted this round,” he said.

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Don’t Get Rid Of Your Old Phone, Turn It Into A Security Camera
AI & Technology

Don’t Get Rid Of Your Old Phone, Turn It Into A Security Camera

September 6, 2026
What Is Model Routing? How AI Systems Choose the Right Model for Every Request – Unite.AI
AI & Technology

What Is Model Routing? How AI Systems Choose the Right Model for Every Request – Unite.AI

September 6, 2026
UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents
AI & Technology

UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents

September 6, 2026
Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed
AI & Technology

Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed

September 6, 2026
Next Post
How data and automation are unlocking the future of subscription businesses

How data and automation are unlocking the future of subscription businesses

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Apple iTunes Still Exists, But Not The Way It Used To

Apple iTunes Still Exists, But Not The Way It Used To

September 5, 2026
Dating app burglary suspect arrested again

Dating app burglary suspect arrested again

September 1, 2026
AI is redefining the workforce — and most planning models aren’t ready

AI is redefining the workforce — and most planning models aren’t ready

September 1, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!