• bitcoinBitcoin(BTC)$76,954.00-1.27%
  • ethereumEthereum(ETH)$2,475.15-1.62%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$718.12-0.71%
  • rippleXRP(XRP)$1.400.35%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$100.82-0.98%
  • tronTRON(TRX)$0.338681-0.48%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,142.180.38%
  • HyperliquidHyperliquid(HYPE)$79.35-0.59%
  • dogecoinDogecoin(DOGE)$0.082660-1.96%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$516.970.01%
  • RainRain(RAIN)$0.013253-12.37%
  • whitebitWhiteBIT Coin(WBT)$79.60-1.39%
  • chainlinkChainlink(LINK)$11.38-0.14%
  • leo-tokenLEO Token(LEO)$8.990.32%
  • cardanoCardano(ADA)$0.205090-2.49%
  • stellarStellar(XLM)$0.1946534.34%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$222.27-0.59%
  • USD1USD1(USD1)$1.00-0.01%
  • uniswapUniswap(UNI)$6.655.57%
  • litecoinLitecoin(LTC)$52.55-2.36%
  • CantonCanton(CC)$0.095328-0.43%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.34-0.89%
  • hedera-hashgraphHedera(HBAR)$0.0773591.27%
  • avalanche-2Avalanche(AVAX)$7.531.89%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • nearNEAR Protocol(NEAR)$2.39-1.06%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.76%
  • suiSui(SUI)$0.71-1.85%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057267-3.20%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,276.66-0.48%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$226.39-4.26%
  • MemeCoreMemeCore(M)$1.11-0.89%
  • okbOKB(OKB)$113.04-0.90%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.18%
  • aaveAave(AAVE)$127.240.14%
  • BitwayBitway(BTW)$0.7111.12%
  • AsterAster(ASTER)$0.69-1.45%
  • pax-goldPAX Gold(PAXG)$4,277.70-0.53%
  • mantleMantle(MNT)$0.56-2.06%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057419-0.33%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Releases WAXAL: A Multilingual African Speech Dataset for Training Automatic Speech Recognition and Text-to-Speech Models

March 17, 2026
in AI & Technology
Reading Time: 4 mins read
A A
Google AI Releases WAXAL: A Multilingual African Speech Dataset for Training Automatic Speech Recognition and Text-to-Speech Models
ShareShareShareShareShare

Speech technology still has a data distribution problem. Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems have improved rapidly for high-resource languages, but many African languages remain poorly represented in open corpora. A team of researchers from Google and other collaborators introduce WAXAL, an open multilingual speech dataset for African languages covering 24 languages, with an ASR component built from transcribed natural speech and a TTS component built from studio-quality single-speaker recordings.

WAXAL is structured as two separate resources because ASR and TTS have different data requirements. The ASR side is designed around diverse speakers, natural environments, and spontaneous language production. The TTS side is designed around controlled recording conditions, phonetically balanced scripts, and cleaner single-speaker audio suited for synthesis. That separation is technically important: a dataset that is useful for robust recognition in noisy real-world settings is usually not the same dataset that produces strong single-speaker TTS models.

YOU MAY ALSO LIKE

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus

Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI

https://arxiv.org/pdf/2602.02734

How the ASR data was collected

The ASR portion of WAXAL was collected using image-prompted speech. Speakers were shown images and asked to describe what they saw in their native language, which is a more natural setup than simple prompted reading. Recordings were captured in speakers’ natural environments, each with a minimum duration of 15 seconds. The collection process also tracked metadata such as speaker age, gender, language, and recording environment. Only a subset of the full collected audio was transcribed: the research team states that the current ASR release includes transcriptions for about 10% of the total recorded audio. Those transcriptions were produced by paid local linguistic experts, using local scripts where available and English-alphabet transliteration otherwise.

This is important for anyone building multilingual ASR systems. Image-prompted speech tends to capture more natural lexical and syntactic variation than tightly scripted reading, but it also makes transcription harder and increases variation across speakers, domains, and acoustic conditions. WAXAL leans into that tradeoff rather than avoiding it. The result is not a perfectly clean benchmark dataset; it is closer to a field-collected multilingual ASR data with real variability baked in.

How the TTS data was collected

The TTS side of WAXAL was built very differently. The TTS dataset was designed for high-quality, single-speaker synthetic voices. For each target language, the research team created a phonetically balanced script of approximately 108,500 words. They contracted 72 community participants, evenly split between male and female voice actors, and recorded them in professional studio-like environments to reduce background noise and preserve audio fidelity. The target was approximately 16 hours of clean edited audio per voice actor.

This is the right design choice for synthesis. TTS models care much more about consistency in pronunciation, recording conditions, microphone quality, and speaker identity than ASR systems do. WAXAL therefore avoids the common mistake of treating ‘speech data’ as a single category, when in practice ASR and TTS pipelines want very different supervision signals.

Key Takeaways

  • WAXAL is an open multilingual speech corpus built for low-resource African language ASR and TTS.
  • The ASR data uses image-prompted, natural speech collected in real-world environments.
  • The TTS data uses studio-quality, single-speaker recordings with phonetically balanced scripts.

Check out Paper and Dataset here. Also, feel free to follow us on Twitter and don’t forget to join our 120k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Google AI Releases WAXAL: A Multilingual African Speech Dataset for Training Automatic Speech Recognition and Text-to-Speech Models appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus
AI & Technology

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus

September 15, 2026
Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI
AI & Technology

Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI

September 15, 2026
Double The Range And Smarter Safety, Too
AI & Technology

Double The Range And Smarter Safety, Too

September 15, 2026
Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second
AI & Technology

Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second

September 15, 2026
Next Post
Gulf Smelter Cuts Tighten Aluminum Outlook

Gulf Smelter Cuts Tighten Aluminum Outlook

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Your Employer’s Life Insurance Coverage Is Probably Not Enough

Your Employer’s Life Insurance Coverage Is Probably Not Enough

September 11, 2026
Husband Only Contributes 17% To Our Household

Husband Only Contributes 17% To Our Household

September 12, 2026
New Hampshire Senate Primary Election 2026 Live Results: Chris Pappas, Karishma Manzur, John Sununu and More – NBC News

New Hampshire Senate Primary Election 2026 Live Results: Chris Pappas, Karishma Manzur, John Sununu and More – NBC News

September 9, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!