• bitcoinBitcoin(BTC)$81,626.000.76%
  • ethereumEthereum(ETH)$2,641.101.92%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$768.971.24%
  • rippleXRP(XRP)$1.433.13%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$111.730.94%
  • tronTRON(TRX)$0.337915-0.38%
  • zcashZcash(ZEC)$1,520.661.80%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.19%
  • HyperliquidHyperliquid(HYPE)$93.181.24%
  • dogecoinDogecoin(DOGE)$0.0887610.97%
  • moneroMonero(XMR)$572.553.99%
  • whitebitWhiteBIT Coin(WBT)$83.380.29%
  • RainRain(RAIN)$0.0139298.34%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.593.06%
  • cardanoCardano(ADA)$0.2267713.15%
  • leo-tokenLEO Token(LEO)$8.940.69%
  • stellarStellar(XLM)$0.1984733.33%
  • uniswapUniswap(UNI)$8.902.44%
  • bitcoin-cashBitcoin Cash(BCH)$253.850.32%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • nearNEAR Protocol(NEAR)$3.56-4.72%
  • daiDai(DAI)$1.000.01%
  • litecoinLitecoin(LTC)$57.953.29%
  • CantonCanton(CC)$0.1109872.32%
  • USD1USD1(USD1)$1.000.02%
  • avalanche-2Avalanche(AVAX)$9.4516.11%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.391.56%
  • hedera-hashgraphHedera(HBAR)$0.0808351.94%
  • suiSui(SUI)$0.844.81%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.12%
  • BittensorBittensor(TAO)$272.567.98%
  • crypto-com-chainCronos(CRO)$0.0599310.53%
  • MemeCoreMemeCore(M)$1.28-3.74%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,373.730.43%
  • okbOKB(OKB)$120.353.56%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.18%
  • aaveAave(AAVE)$142.863.42%
  • OndoOndo(ONDO)$0.43723110.49%
  • mantleMantle(MNT)$0.635.53%
  • AsterAster(ASTER)$0.771.46%
  • EthenaEthena(ENA)$0.19784219.95%
  • Pump.funPump.fun(PUMP)$0.004117-5.05%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MMLU-Pro: An Enhanced Benchmark Designed to Evaluate Language Understanding Models Across Broader and More Challenging Tasks

June 6, 2024
in AI & Technology
Reading Time: 5 mins read
A A
MMLU-Pro: An Enhanced Benchmark Designed to Evaluate Language Understanding Models Across Broader and More Challenging Tasks
ShareShareShareShareShare

Recent advancements in large language models (LLMs) have significantly transformed the field of natural language processing (NLP), but their performance on existing benchmarks has begun to plateau. This stagnation makes it difficult to discern differences in model capabilities, hindering progress in AI research. Benchmarks like the Massive Multitask Language Understanding (MMLU) have played a crucial role in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, the need for more challenging and discriminative benchmarks has become apparent as models improve. The performance saturation on these benchmarks limits the ability to effectively evaluate newer, more advanced models. Additionally, the existing benchmarks often feature questions that are predominantly knowledge-driven with limited reasoning requirements, leading to inflated performance metrics and reduced robustness due to sensitivity to prompt variations.

Current methods, such as the original MMLU and other benchmarks like GLUE, SuperGLUE, and BigBench, have played pivotal roles in advancing language understanding tasks. However, these benchmarks primarily focus on knowledge-driven questions with limited reasoning requirements, leading to performance saturation among top-tier models like GPT-4, Gemini, and Claude. These methods also exhibit non-robustness to minor prompt variations, resulting in significant fluctuations in model scores and overestimating LLMs’ true performance. The typical multiple-choice question format, often limited to four options, fails to differentiate closely performing systems and does not adequately challenge the models’ reasoning capabilities​​.

Researchers from the University of Waterloo, the University of Toronto, and Carnegie Mellon University propose a new benchmark/leaderboard, MMLU-Pro, which addresses these limitations by incorporating more challenging, reasoning-intensive tasks and increasing the number of distractor options from three to nine. This benchmark spans 14 diverse domains, encompassing over 12,000 questions, thus providing a broader and more discriminative evaluation. MMLU-Pro also involves a two-round expert review process to reduce dataset noise and enhance question quality. This novel approach significantly raises the benchmark’s difficulty level and robustness, making it better suited for assessing the advanced reasoning capabilities of state-of-the-art LLMs​​.

MMLU-Pro’s dataset construction involves integrating questions from various high-quality sources, including the original MMLU, STEM websites, TheoremQA, and SciBench, ensuring a diverse and challenging question set. The dataset is filtered and refined through a rigorous process, removing overly simple or erroneous questions and augmenting the question options to ten, which necessitates more discerning reasoning for correct selection. The benchmark also evaluates models’ performance across 24 different prompt styles to assess robustness and minimize prompt variability impacts. Notable technical aspects include leveraging the capabilities of advanced LLMs like GPT-4-Turbo for option augmentation and ensuring consistency and accuracy through expert verification​​.

MMLU-Pro presents significant challenges even for the leading models. For instance, GPT-4o, the strongest model tested, achieved an overall accuracy of 72.6%, while other top-tier models like GPT-4-Turbo reached 63.7%. These results indicate substantial room for improvement, highlighting the benchmark’s effectiveness in differentiating models’ reasoning capabilities. The benchmark’s robustness is evidenced by reduced variability in model scores under different prompts, with a maximum impact of 3.74% compared to up to 10.98% on the original MMLU. This stability enhances the reliability of evaluations and the benchmark’s utility in advancing AI language understanding capabilities.

In Conclusion, MMLU-Pro represents a significant advancement in benchmarking LLMs by addressing the limitations of existing methods and enhancing the assessment of multi-task language understanding and reasoning capabilities. The benchmark introduces more complex, reasoning-intensive tasks and increases the number of distractor options, significantly improving its robustness and discriminative power. The evaluations show that even top-performing models face substantial challenges on MMLU-Pro, indicating its effectiveness in pushing the boundaries of AI capabilities. This benchmark is poised to play a crucial role in the future development and evaluation of LLMs, driving advancements in AI research by overcoming critical challenges in model evaluation​​.


Check out the Paper and Leaderboard. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 43k+ ML SubReddit | Also, check out our AI Events Platform


YOU MAY ALSO LIKE

How To Block And Unblock A Number On Your Android Phone

Google Gemini Also Escaped Its Testing Environment And Hacked Three Companies

Aswin AK is a consulting intern at MarkTechPost. He is pursuing his Dual Degree at the Indian Institute of Technology, Kharagpur. He is passionate about data science and machine learning, bringing a strong academic background and hands-on experience in solving real-life cross-domain challenges.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Block And Unblock A Number On Your Android Phone
AI & Technology

How To Block And Unblock A Number On Your Android Phone

September 19, 2026
Google Gemini Also Escaped Its Testing Environment And Hacked Three Companies
AI & Technology

Google Gemini Also Escaped Its Testing Environment And Hacked Three Companies

September 19, 2026
What Is AI Agent Memory? Short-Term, Long-Term, Episodic, and Semantic Memory Explained – Unite.AI
AI & Technology

What Is AI Agent Memory? Short-Term, Long-Term, Episodic, and Semantic Memory Explained – Unite.AI

September 19, 2026
Linkup Research Releases SPARSEUP: A 149M-Parameter Open-Source Sparse Embedding Model
AI & Technology

Linkup Research Releases SPARSEUP: A 149M-Parameter Open-Source Sparse Embedding Model

September 19, 2026
Next Post
More than 50,000 Americans died by suicide in 2023 — more than any year on record

More than 50,000 Americans died by suicide in 2023 — more than any year on record

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI

OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI

September 17, 2026
She’s Draining Her Life Savings To Fund Her Kid’s College

She’s Draining Her Life Savings To Fund Her Kid’s College

September 19, 2026
Sinkhole threatens mansion once linked to Nicolas Cage

Sinkhole threatens mansion once linked to Nicolas Cage

September 15, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!