• bitcoinBitcoin(BTC)$84,897.001.70%
  • ethereumEthereum(ETH)$2,715.110.98%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$771.870.41%
  • rippleXRP(XRP)$1.500.54%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$119.551.27%
  • tronTRON(TRX)$0.334431-0.99%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-0.71%
  • zcashZcash(ZEC)$1,335.02-5.81%
  • HyperliquidHyperliquid(HYPE)$88.30-1.31%
  • dogecoinDogecoin(DOGE)$0.093911-0.82%
  • chainlinkChainlink(LINK)$14.29-0.58%
  • moneroMonero(XMR)$547.160.21%
  • whitebitWhiteBIT Coin(WBT)$84.821.63%
  • USDSUSDS(USDS)$1.000.04%
  • cardanoCardano(ADA)$0.246447-0.36%
  • RainRain(RAIN)$0.012038-1.90%
  • leo-tokenLEO Token(LEO)$8.92-1.24%
  • stellarStellar(XLM)$0.219194-4.24%
  • nearNEAR Protocol(NEAR)$4.89-6.44%
  • bitcoin-cashBitcoin Cash(BCH)$307.300.45%
  • uniswapUniswap(UNI)$8.981.55%
  • litecoinLitecoin(LTC)$68.762.57%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • suiSui(SUI)$1.180.85%
  • avalanche-2Avalanche(AVAX)$10.920.55%
  • CantonCanton(CC)$0.120645-4.37%
  • Blockchain USDBlockchain USD(USDB)$0.871,000.00%
  • daiDai(DAI)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.102709-1.19%
  • USD1USD1(USD1)$1.000.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.574.84%
  • BitwayBitway(BTW)$1.416.22%
  • quant-networkQuant(QNT)$254.30-14.00%
  • BittensorBittensor(TAO)$304.521.27%
  • shiba-inuShiba Inu(SHIB)$0.0000060.66%
  • crypto-com-chainCronos(CRO)$0.0683361.37%
  • tether-goldTether Gold(XAUT)$4,153.16-0.10%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.05%
  • aaveAave(AAVE)$176.219.30%
  • Pump.funPump.fun(PUMP)$0.005753-0.03%
  • okbOKB(OKB)$121.10-0.36%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • EthenaEthena(ENA)$0.241288-8.63%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • OndoOndo(ONDO)$0.493614-2.20%
  • MemeCoreMemeCore(M)$1.050.03%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.01%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet XTREME-UP: A Benchmark for Evaluating Multilingual Models with Scarce Data Evaluation, Focusing on Under-Represented Languages

May 24, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meet XTREME-UP: A Benchmark for Evaluating Multilingual Models with Scarce Data Evaluation, Focusing on Under-Represented Languages
ShareShareShareShareShare
➡️ Annotate all types of unstructured data rapidly and accurately with customizable annotation tasks with Kili Technology:

The fields of Artificial Intelligence and Machine Learning are solely dependent upon data. Everyone is deluged with data from different sources like social media, healthcare, finance, etc., and this data is of great use to applications involving Natural Language Processing. But even with so much data, readily usable data is scarce for training an NLP model for a particular task. Finding high-quality data with usefulness and good-quality filters is a difficult task. Specifically talking about developing NLP models for different languages, the lack of data for most languages comes as a limitation that hinders progress in NLP for under-represented languages (ULs). 

The emerging tasks like news summarization, sentiment analysis, question answering, or the development of a virtual assistant all heavily rely on data availability in high-resource languages. These tasks are dependent upon technologies like language identification, automatic speech recognition (ASR), or optical character recognition (OCR), which are mostly unavailable for under-represented languages, to overcome which it is important to build datasets and evaluate models on tasks that would be beneficial for UL speakers. 

Recently, a team of researchers from GoogleAI has proposed a benchmark called XTREME-UP (Under-Represented and User-Centric with Paucal Data) that evaluates multilingual models on user-centric tasks in a few-shot learning setting. It primarily focuses on activities that technology users often perform in their day-to-day lives, such as information access and input/output activities that enable other technologies. The three main features that distinguish XTREME-UP are – its use of scarce data, its user-centric design, and its focus on under-represented languages.

🚀 JOIN the fastest ML Subreddit Community

With XTREME-UP, the researchers have introduced a standardized multilingual in-language fine-tuning setting in place of the conventional cross-lingual zero-shot option. This method considers the amount of data that can be generated or annotated in an 8-hour period for a particular language, thus aiming to give the ULs a more useful evaluation setup. 

XTREME-UP assesses the performance of language models across 88 under-represented languages in 9 significant user-centric technologies, some of which include Automatic Speech Recognition (ASR), Optical Character Recognition (OCR), Machine Translation (MT), and information access tasks that have general utility. The researchers have developed new datasets specifically for operations like OCR, autocomplete, semantic parsing, and transliteration in order to evaluate the capabilities of the language models. They have also improved and polished the currently existing datasets for other tasks in the same benchmark.

XTREME-UP has one of its key abilities to assess various modeling situations, including both text-only and multi-modal scenarios with visual, audio, and text inputs. It also offers methods for supervised parameter adjustment and in-context learning, allowing for a thorough assessment of various modeling approaches. The tasks in XTREME-UP involve enabling access to language technology, enabling information access as part of a larger system such as question answering, information extraction, and virtual assistants, followed by making information accessible in the speaker’s language.

Consequently, XTREME-UP is a great benchmark that addresses the data scarcity challenge in highly multilingual NLP systems. It is a standardized evaluation framework for under-represented language and seems really useful for future NLP research and developments.


Check out the Paper and Github. Don’t forget to join our 21k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End

Florida County Finds 11 Unpermitted Flock Cameras With Unidentified Owners

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


➡️ Ultimate Guide to Data Labeling in Machine Learning

Credit: Source link

ShareTweetSendSharePin

Related Posts

A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End
AI & Technology

A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End

October 2, 2026
Florida County Finds 11 Unpermitted Flock Cameras With Unidentified Owners
AI & Technology

Florida County Finds 11 Unpermitted Flock Cameras With Unidentified Owners

October 1, 2026
OpenAI Fires Three Employees Who Allegedly Shared Info With An External AI Safety Group
AI & Technology

OpenAI Fires Three Employees Who Allegedly Shared Info With An External AI Safety Group

October 1, 2026
Microsoft Launches MAI-Transcribe-2-Streaming and Two MAI-Voice Models – Unite.AI
AI & Technology

Microsoft Launches MAI-Transcribe-2-Streaming and Two MAI-Voice Models – Unite.AI

October 1, 2026
Next Post
Salesforce Is Still Considering an Acquisition of Social Media Giant Twitter, Reports Say

Salesforce Is Still Considering an Acquisition of Social Media Giant Twitter, Reports Say

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NBC Nightly News with Tom Llamas Full Episode – Aug. 19

NBC Nightly News with Tom Llamas Full Episode – Aug. 19

September 27, 2026
TOP 5 STOCKS TO WATCH THIS WEEK | SEPTEMBER 2026

TOP 5 STOCKS TO WATCH THIS WEEK | SEPTEMBER 2026

September 28, 2026
Still Serving: Inside the Most Bombed McDonald’s in The World

Still Serving: Inside the Most Bombed McDonald’s in The World

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!