• bitcoinBitcoin(BTC)$77,119.00-1.60%
  • ethereumEthereum(ETH)$2,459.31-0.74%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$715.06-1.06%
  • rippleXRP(XRP)$1.35-3.09%
  • usd-coinUSDC(USDC)$1.000.03%
  • solanaSolana(SOL)$99.82-2.22%
  • tronTRON(TRX)$0.339832-0.17%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.97%
  • zcashZcash(ZEC)$1,087.56-12.82%
  • HyperliquidHyperliquid(HYPE)$79.35-5.66%
  • dogecoinDogecoin(DOGE)$0.084099-2.18%
  • RainRain(RAIN)$0.015727-4.13%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$517.600.56%
  • whitebitWhiteBIT Coin(WBT)$79.83-1.31%
  • chainlinkChainlink(LINK)$11.53-2.45%
  • leo-tokenLEO Token(LEO)$9.10-1.04%
  • cardanoCardano(ADA)$0.209584-1.90%
  • stellarStellar(XLM)$0.176810-1.83%
  • bitcoin-cashBitcoin Cash(BCH)$228.41-9.20%
  • daiDai(DAI)$1.00-0.02%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$53.170.49%
  • CantonCanton(CC)$0.098702-5.02%
  • uniswapUniswap(UNI)$6.09-0.04%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-1.62%
  • hedera-hashgraphHedera(HBAR)$0.075578-1.64%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.52-4.08%
  • nearNEAR Protocol(NEAR)$2.46-2.51%
  • suiSui(SUI)$0.74-3.58%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.56%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056488-2.44%
  • tether-goldTether Gold(XAUT)$4,327.61-1.82%
  • MemeCoreMemeCore(M)$1.16-5.34%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$109.21-3.76%
  • BittensorBittensor(TAO)$236.80-8.12%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.07%
  • polkadotPolkadot(DOT)$1.132.06%
  • AsterAster(ASTER)$0.71-2.20%
  • mantleMantle(MNT)$0.58-3.98%
  • aaveAave(AAVE)$122.78-2.02%
  • pax-goldPAX Gold(PAXG)$4,333.62-1.77%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.055903-1.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meta AI Launches Massively Multilingual Speech (MMS) Project: Introducing Speech-To-Text, Text-To-Speech, And More For 1,000+ Languages

May 31, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Meta AI Launches Massively Multilingual Speech (MMS) Project: Introducing Speech-To-Text, Text-To-Speech, And More For 1,000+ Languages
ShareShareShareShareShare

Significant advancements in speech technology have been made over the past decade, allowing it to be incorporated into various consumer items. It takes a lot of labeled data, in this case, many thousands of hours of audio with transcriptions, to train a good machine learning model for such jobs. This information only exists in some languages. For instance, out of the 7,000+ languages in use today, only about 100 are supported by current voice recognition algorithms. 

Recently, the amount of labeled data needed to construct speech systems have been drastically reduced because of self-supervised speech representations. Despite progress, major current efforts still only cover around 100 languages. 

Facebook’s Massively Multilingual Speech (MMS) project combines wav2vec 2.0 with a new dataset that contains labeled data for over 1,100 languages and unlabeled data for almost 4,000 languages to address some of these obstacles. Based on their findings, the Massively Multilingual Speech models are superior to the state-of-the-art methods and support ten times as many languages. 

🚀 JOIN the fastest ML Subreddit Community

Since the greatest available speech datasets only include up to 100 languages, their initial goal was to collect audio data for hundreds of languages. As a result, they looked to religious writings like the Bible, which have been translated into many languages and whose translations have been extensively examined for text-based language translation research. People have recorded themselves reading these translations and made the audio files available online. This research compiled a collection of New Testament readings in over 1,100 languages, yielding an average of 32 hours of data per language.

Their investigation reveals that the proposed models perform similarly well for male and female voices, even though this data is from a specific domain and is typically read by male speakers. Even though the recordings are religious, the research indicates that this does not unduly bias the model toward producing more religious language. According to the researchers, this is because they employ a Connectionist Temporal Classification strategy, which is more limited than large language models (LLMs) or sequence-to-sequence models for voice recognition.

The team preprocessed tha data by combining a highly efficient forced alignment approach that can handle recordings that are 20 minutes or longer with an alignment model that was trained using data from over 100 different languages. To eliminate possibly skewed information, they used numerous iterations of this procedure plus a cross-validation filtering step based on model accuracy. They integrated the alignment technique into PyTorch and made the alignment model publicly available so that other academics may use it to generate fresh speech datasets.

There is insufficient information to train traditional supervised speech recognition models with only 32 hours of data per language. The team relied on wav2vec 2.0 to train effective systems, drastically decreasing the quantity of previously required labeled data. Specifically, they used over 1,400 unique languages to train self-supervised models on over 500,000 hours of voice data, approximately five times more languages than any previous effort. 

The researchers employed pre-existing benchmark datasets like FLEURS to assess the performance of models trained on the Massively Multilingual Speech data. Using a 1B parameter wav2vec 2.0 model, they trained a multilingual speech recognition system on over 1,100 languages. The performance degrades slightly as the number of languages grows: The character mistake rate only goes up by roughly 0.4% from 61 to 1,107 languages, while the language coverage goes up by nearly 18 times.

Comparing the Massively Multilingual Speech data to OpenAI’s Whisper, the researchers discovered that models trained on the former achieve half the word error rate. At the same time, the latter covers 11 times as many languages. This illustrates that the model can compete favorably with the state-of-the-art in voice recognition.

The team also used their datasets and publicly available datasets like FLEURS and CommonVoice to train a language identification (LID) model for more than 4,000 languages. Then it tested it on the FLEURS LID challenge. The findings show that performance is still excellent even when 40 times as many languages are supported. They also developed speech synthesis systems for more than 1,100 languages. The majority of existing text-to-speech algorithms are trained on single-speaker voice datasets. 

The team foresees a world where one model can handle many speech tasks across all languages. While they did train individual models for each task—recognition, synthesis, and identification of language—they believe that in the future, a single model will be able to handle all of these functions and more, improving performance in every area.


Check out the Paper, Blog, and Github Link. Don’t forget to join our 22k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

How These XL Phones Compete

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

Tanushree Shenwai is a consulting intern at MarktechPost. She is currently pursuing her B.Tech from the Indian Institute of Technology(IIT), Bhubaneswar. She is a Data Science enthusiast and has a keen interest in the scope of application of artificial intelligence in various fields. She is passionate about exploring the new advancements in technologies and their real-life application.


➡️ Ultimate Guide to Data Labeling in Machine Learning

Credit: Source link

ShareTweetSendSharePin

Related Posts

How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots
AI & Technology

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
AI & Technology

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100
AI & Technology

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

September 10, 2026
Next Post
Abortion lawsuit in Texas could restrict medical pill nationwide

Abortion lawsuit in Texas could restrict medical pill nationwide

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Western Digital Corporation (WDC) Presents at Citi’s 2026 Global TMT Conference Transcript

Western Digital Corporation (WDC) Presents at Citi’s 2026 Global TMT Conference Transcript

September 9, 2026
Where Baked by Melissa finds her inspiration

Where Baked by Melissa finds her inspiration

September 5, 2026
UFC Paris live results: Dan Hooker vs. Salahdine Parnasse updates, round-by-round scoring for today's fight – Yahoo Sports

UFC Paris live results: Dan Hooker vs. Salahdine Parnasse updates, round-by-round scoring for today's fight – Yahoo Sports

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!