• bitcoinBitcoin(BTC)$78,458.00-0.70%
  • ethereumEthereum(ETH)$2,484.70-0.01%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$752.311.85%
  • rippleXRP(XRP)$1.421.62%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.37-0.33%
  • tronTRON(TRX)$0.3389171.40%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,180.654.04%
  • HyperliquidHyperliquid(HYPE)$85.02-0.08%
  • dogecoinDogecoin(DOGE)$0.089995-0.61%
  • RainRain(RAIN)$0.016202-0.52%
  • USDSUSDS(USDS)$1.000.03%
  • whitebitWhiteBIT Coin(WBT)$81.266.29%
  • moneroMonero(XMR)$505.58-2.44%
  • chainlinkChainlink(LINK)$12.53-1.59%
  • leo-tokenLEO Token(LEO)$9.210.03%
  • cardanoCardano(ADA)$0.2199650.01%
  • stellarStellar(XLM)$0.187928-2.94%
  • bitcoin-cashBitcoin Cash(BCH)$258.820.51%
  • daiDai(DAI)$1.000.02%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.1073322.32%
  • litecoinLitecoin(LTC)$54.38-1.34%
  • uniswapUniswap(UNI)$6.74-1.85%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.400.87%
  • hedera-hashgraphHedera(HBAR)$0.079305-3.06%
  • avalanche-2Avalanche(AVAX)$8.01-0.84%
  • suiSui(SUI)$0.81-0.64%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.18%
  • nearNEAR Protocol(NEAR)$2.310.40%
  • crypto-com-chainCronos(CRO)$0.0588163.94%
  • paypal-usdPayPal USD(PYUSD)$1.000.03%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.236.80%
  • tether-goldTether Gold(XAUT)$4,359.41-1.29%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$260.310.75%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$113.86-1.48%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • polkadotPolkadot(DOT)$1.2517.47%
  • mantleMantle(MNT)$0.631.88%
  • AsterAster(ASTER)$0.75-2.59%
  • aaveAave(AAVE)$128.85-2.30%
  • pax-goldPAX Gold(PAXG)$4,361.54-1.28%
  • OndoOndo(ONDO)$0.375020-1.99%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet XTREME-UP: A Benchmark for Evaluating Multilingual Models with Scarce Data Evaluation, Focusing on Under-Represented Languages

May 24, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meet XTREME-UP: A Benchmark for Evaluating Multilingual Models with Scarce Data Evaluation, Focusing on Under-Represented Languages
ShareShareShareShareShare
➡️ Annotate all types of unstructured data rapidly and accurately with customizable annotation tasks with Kili Technology:

The fields of Artificial Intelligence and Machine Learning are solely dependent upon data. Everyone is deluged with data from different sources like social media, healthcare, finance, etc., and this data is of great use to applications involving Natural Language Processing. But even with so much data, readily usable data is scarce for training an NLP model for a particular task. Finding high-quality data with usefulness and good-quality filters is a difficult task. Specifically talking about developing NLP models for different languages, the lack of data for most languages comes as a limitation that hinders progress in NLP for under-represented languages (ULs). 

The emerging tasks like news summarization, sentiment analysis, question answering, or the development of a virtual assistant all heavily rely on data availability in high-resource languages. These tasks are dependent upon technologies like language identification, automatic speech recognition (ASR), or optical character recognition (OCR), which are mostly unavailable for under-represented languages, to overcome which it is important to build datasets and evaluate models on tasks that would be beneficial for UL speakers. 

Recently, a team of researchers from GoogleAI has proposed a benchmark called XTREME-UP (Under-Represented and User-Centric with Paucal Data) that evaluates multilingual models on user-centric tasks in a few-shot learning setting. It primarily focuses on activities that technology users often perform in their day-to-day lives, such as information access and input/output activities that enable other technologies. The three main features that distinguish XTREME-UP are – its use of scarce data, its user-centric design, and its focus on under-represented languages.

🚀 JOIN the fastest ML Subreddit Community

With XTREME-UP, the researchers have introduced a standardized multilingual in-language fine-tuning setting in place of the conventional cross-lingual zero-shot option. This method considers the amount of data that can be generated or annotated in an 8-hour period for a particular language, thus aiming to give the ULs a more useful evaluation setup. 

XTREME-UP assesses the performance of language models across 88 under-represented languages in 9 significant user-centric technologies, some of which include Automatic Speech Recognition (ASR), Optical Character Recognition (OCR), Machine Translation (MT), and information access tasks that have general utility. The researchers have developed new datasets specifically for operations like OCR, autocomplete, semantic parsing, and transliteration in order to evaluate the capabilities of the language models. They have also improved and polished the currently existing datasets for other tasks in the same benchmark.

XTREME-UP has one of its key abilities to assess various modeling situations, including both text-only and multi-modal scenarios with visual, audio, and text inputs. It also offers methods for supervised parameter adjustment and in-context learning, allowing for a thorough assessment of various modeling approaches. The tasks in XTREME-UP involve enabling access to language technology, enabling information access as part of a larger system such as question answering, information extraction, and virtual assistants, followed by making information accessible in the speaker’s language.

Consequently, XTREME-UP is a great benchmark that addresses the data scarcity challenge in highly multilingual NLP systems. It is a standardized evaluation framework for under-represented language and seems really useful for future NLP research and developments.


Check out the Paper and Github. Don’t forget to join our 21k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


➡️ Ultimate Guide to Data Labeling in Machine Learning

Credit: Source link

ShareTweetSendSharePin

Related Posts

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction
AI & Technology

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

September 8, 2026
SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas
AI & Technology

SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas

September 8, 2026
What Is Roku’s Secret Menu And How Do You Unlock It?
AI & Technology

What Is Roku’s Secret Menu And How Do You Unlock It?

September 8, 2026
NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels
AI & Technology

NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels

September 8, 2026
Next Post
Salesforce Is Still Considering an Acquisition of Social Media Giant Twitter, Reports Say

Salesforce Is Still Considering an Acquisition of Social Media Giant Twitter, Reports Say

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Nvidia buying AI startup Hugging Face for whopping B

Nvidia buying AI startup Hugging Face for whopping $13B

September 3, 2026
German chancellor vows to stay in office despite AfD triumph in state election – The Guardian

German chancellor vows to stay in office despite AfD triumph in state election – The Guardian

September 8, 2026
How To See What’s Taking Up Space On Your Windows PC

How To See What’s Taking Up Space On Your Windows PC

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!