• bitcoinBitcoin(BTC)$75,971.00-3.01%
  • ethereumEthereum(ETH)$2,421.33-3.06%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$716.45-0.55%
  • rippleXRP(XRP)$1.39-0.14%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.46-1.89%
  • tronTRON(TRX)$0.337013-1.09%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-1.04%
  • zcashZcash(ZEC)$1,119.81-1.86%
  • HyperliquidHyperliquid(HYPE)$77.41-2.68%
  • dogecoinDogecoin(DOGE)$0.081932-2.10%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$516.681.49%
  • whitebitWhiteBIT Coin(WBT)$78.49-2.93%
  • RainRain(RAIN)$0.012732-14.68%
  • chainlinkChainlink(LINK)$11.25-0.97%
  • leo-tokenLEO Token(LEO)$8.78-2.33%
  • cardanoCardano(ADA)$0.202621-2.68%
  • stellarStellar(XLM)$0.1932311.69%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$221.63-0.41%
  • USD1USD1(USD1)$1.00-0.04%
  • litecoinLitecoin(LTC)$52.04-3.13%
  • uniswapUniswap(UNI)$6.31-0.34%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.33-1.57%
  • CantonCanton(CC)$0.093499-2.08%
  • hedera-hashgraphHedera(HBAR)$0.0774851.07%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$7.44-0.05%
  • nearNEAR Protocol(NEAR)$2.35-2.09%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.35%
  • suiSui(SUI)$0.70-2.18%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • crypto-com-chainCronos(CRO)$0.056838-4.06%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,276.830.14%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$223.53-3.50%
  • MemeCoreMemeCore(M)$1.111.43%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$111.62-2.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.11%
  • aaveAave(AAVE)$125.970.15%
  • BitwayBitway(BTW)$0.700.99%
  • pax-goldPAX Gold(PAXG)$4,281.000.11%
  • AsterAster(ASTER)$0.68-1.17%
  • mantleMantle(MNT)$0.55-2.95%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056951-0.30%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Research Introduces ‘MEGAVERSE’ for Benchmarking Large Language Models Across Languages, Modalities, Models, and Tasks

April 14, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Microsoft Research Introduces ‘MEGAVERSE’ for Benchmarking Large Language Models Across Languages, Modalities, Models, and Tasks
ShareShareShareShareShare

On many tasks and benchmarks, Large Language Models (LLMs) have outperformed earlier generations of language models, and on occasion, they have even come close to matching or surpassing human performance. While some models may seem to have impressive skills, it is not always easy to tell if that is due to enhanced model capabilities or something else entirely, such as contamination in test datasets or a lack of datasets that accurately assess their abilities. Because of this, research into determining LLMs has grown in stature.

Most studies that have attempted to assess LLMs, whether through human review, qualitative tests for specific competencies, or benchmarking, have primarily focused on the English language. This research has uncovered a significant disparity in the proficiency of LLMs in English compared to other languages. However, evaluating LLMs in languages other than English poses numerous challenges, including the scarcity of multilingual benchmarks for reasoning, conversation, and dialogue across various language families.

The findings from earlier studies on MEGA provide valuable insights into the multilingual capabilities of LLMs. Compared to state-of-the-art (SOTA) tuned language models like TULRv6, GPT-4 demonstrates commendable performance. However, it’s important to note that GPT models exhibit lower performance, particularly those designed for low-resource languages and languages written in scripts other than Latin.

Researchers from Microsoft Corporation expanded coverage to 22 datasets and 83 languages, including many low-resource African languages, by building on the MEGA benchmark and adding 6 new datasets.

This work provides valuable insights for developers and researchers. In particular, the team found that bigger commercial models like GPT-4 and Gemini-pro perform better than smaller ones like Gemma, Llama, and Mistral on low-resource languages. This pattern holds across most of the datasets examined, indicating that smaller models have difficulty with multilingual performance. This suggests that approaches like fine-tuning, language family-based models, and language-specific models should be investigated further to improve multilingual performance. 

Regarding the multimodal datasets, GPT-4-Vision performed better than LLaVA and Gemini-Pro-Vision. The efficiency of the Language Model is related to the fertility of tokenizers. The work also depicts the fertility analysis of each tokenizer, suggesting that tokenizer fertility was lower for Latin script languages like English and Spanish than for morphologically complicated languages like Telugu, Malay, and Malayalam.

Due to computational and time constraints, the researchers highlight that they could only conduct the contamination research on 7B variations of their open-source models and not all datasets. Further, dataset contamination is a major problem with benchmarking studies conducted in languages other than English. According to their contamination analysis on commercial and free source models, almost all models use MEGAVERSE datasets. It is important to avoid including newly created multilingual evaluation datasets in LLM training data because of the difficulty in doing so owing to financial and resource limitations. To accomplish this, the team aims to improve its capacity to detect contamination and implement safeguards to prevent it from happening again in the future.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


Want to get in front of 1.5 Million AI Audience? Work with us here


YOU MAY ALSO LIKE

Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron – Unite.AI

2 Ways Android Users Can Take Advantage Of Apple’s MagSafe Accessories

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron – Unite.AI
AI & Technology

Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron – Unite.AI

September 15, 2026
2 Ways Android Users Can Take Advantage Of Apple’s MagSafe Accessories
AI & Technology

2 Ways Android Users Can Take Advantage Of Apple’s MagSafe Accessories

September 15, 2026
Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus
AI & Technology

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus

September 15, 2026
Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI
AI & Technology

Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI

September 15, 2026
Next Post
NYPD officers save man who fell onto train tracks

NYPD officers save man who fell onto train tracks

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
White House may impose ‘limited safeguards’ on AI to prevent apocalypse, mass extinction: source

White House may impose ‘limited safeguards’ on AI to prevent apocalypse, mass extinction: source

September 14, 2026
Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data

Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data

September 14, 2026
Bunching Deductions Into One Year Can Beat the Standard Deduction

Bunching Deductions Into One Year Can Beat the Standard Deduction

September 12, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!