• bitcoinBitcoin(BTC)$77,021.00-2.39%
  • ethereumEthereum(ETH)$2,431.57-2.57%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$708.69-4.57%
  • rippleXRP(XRP)$1.36-4.22%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.53-3.76%
  • tronTRON(TRX)$0.3386620.08%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.80%
  • zcashZcash(ZEC)$1,168.32-8.59%
  • HyperliquidHyperliquid(HYPE)$80.71-6.14%
  • dogecoinDogecoin(DOGE)$0.083733-6.73%
  • RainRain(RAIN)$0.015924-2.22%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$503.950.28%
  • whitebitWhiteBIT Coin(WBT)$79.53-2.50%
  • chainlinkChainlink(LINK)$11.65-3.57%
  • leo-tokenLEO Token(LEO)$9.190.15%
  • cardanoCardano(ADA)$0.210019-3.35%
  • stellarStellar(XLM)$0.177442-5.28%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$227.75-11.56%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$52.36-2.76%
  • CantonCanton(CC)$0.101040-3.72%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.34-3.60%
  • uniswapUniswap(UNI)$5.90-10.64%
  • hedera-hashgraphHedera(HBAR)$0.075691-3.27%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.60-4.08%
  • nearNEAR Protocol(NEAR)$2.44-4.81%
  • suiSui(SUI)$0.75-7.37%
  • shiba-inuShiba Inu(SHIB)$0.000005-6.14%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.056358-5.86%
  • tether-goldTether Gold(XAUT)$4,362.47-1.19%
  • MemeCoreMemeCore(M)$1.16-1.34%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$111.02-2.42%
  • BittensorBittensor(TAO)$242.66-7.44%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.16%
  • mantleMantle(MNT)$0.58-9.12%
  • AsterAster(ASTER)$0.71-5.20%
  • pax-goldPAX Gold(PAXG)$4,363.88-1.24%
  • aaveAave(AAVE)$122.01-5.69%
  • polkadotPolkadot(DOT)$1.09-4.55%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0561330.94%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CMU Researchers Introduce the Open Whisper-Style Speech Model: Advancing Open-Source Solutions for Efficient and Transparent Speech Recognition Training

October 3, 2023
in AI & Technology
Reading Time: 4 mins read
A A
CMU Researchers Introduce the Open Whisper-Style Speech Model: Advancing Open-Source Solutions for Efficient and Transparent Speech Recognition Training
ShareShareShareShareShare

Natural language processing (NLP) has paid much attention to large-scale Transformers. These models, trained on large datasets, have demonstrated amazing emergent abilities in various downstream applications. Notably, comparable pre-training methods have been successfully used in voice processing. A promising path to creating universal speech models that can handle many speech tasks inside a single model is large-scale supervised learning. A collection of multilingual, multitask models called OpenAI Whisper [15] was developed using 680k hours of labeled voice data that was carefully selected from various online sources.

The complete process for model building (from data preparation to training) is still unavailable to the general public despite the publication of pre-trained Whisper models and inference code, which has been a usual situation for large language models (LLMs). This restriction raises several issues:

  1. Because users are unaware of the actual training data, using pre-trained models on new benchmarks carries the danger of data leakage.
  2. Users lack access to the training dynamics. So, researchers have difficulty understanding the underlying mechanisms and illuminating strategies for improving the model’s performance. 
  3. Dealing with problems relating to robustness, fairness, bias, and toxicity, all of which typically arise due to the data and training process, is made far more difficult by the lack of access to the entire model development pipeline.

By pushing for the publication of comprehensive training pipelines, there has recently been a determined movement to promote open science in the field of LLM research. This inspired the research team from Carnegie Mellon University, Shanghai Jiao Tong University, and Honda Research Institute to create the Open Whisper-style Speech Model (OWSM)2, which uses an open-source toolbox and publicly available data to replicate whisper-style training. To handle crucial tasks, including language identification (LID), multilingual automated speech recognition (ASR), and utterance-level segmentation, OWSM adopts the Whisper framework. 

Notably, OWSM also displays several technical innovations. Instead of just any-to-English translation, it handles any-to-any speech translation. OWSM also uses a variety of tactics to improve efficiency. The entire pipeline, including data preparation, training, inference, and scoring, will be covered by reproducible recipes. The team also plans to make pre-trained models and training logs available, letting researchers dig into the mechanics of the training procedure and obtain important knowledge for their research. 

While OWSM performs similarly to Whisper or even better on some metrics, its goal is not to engage in a protracted arms race with Whisper. The team’s largest dataset only makes up about 25% of the training set used by Whisper, and they cannot execute numerous trial runs because of resource constraints. 

In the future, the team plans to explore the following directions:

  1. The current OWSM still lags behind Whisper in many benchmarks. The researchers believe that using more sophisticated encoder or decoder architectures, gathering more varied ASR and ST data from open sources, and incorporating self-supervised speech representations similar to Google USM can help.
  2. They also intend to add other speech-processing tasks to the multitask framework, such as spoken language comprehension and speech synthesis based on discrete representations, to create “universal speech models.”

Check out the Paper and Code. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 31k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

IBM and NASA Open-Source Lunar Foundation Model With SomBench Dataset – Unite.AI

NASA And IBM Made An AI Model For Exploring The Moon

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

IBM and NASA Open-Source Lunar Foundation Model With SomBench Dataset – Unite.AI
AI & Technology

IBM and NASA Open-Source Lunar Foundation Model With SomBench Dataset – Unite.AI

September 10, 2026
NASA And IBM Made An AI Model For Exploring The Moon
AI & Technology

NASA And IBM Made An AI Model For Exploring The Moon

September 10, 2026
Fujitsu Signs New Palantir AIP Agreement, Becomes Global FDE Partner – Unite.AI
AI & Technology

Fujitsu Signs New Palantir AIP Agreement, Becomes Global FDE Partner – Unite.AI

September 10, 2026
AppleCare One Now Has A  Tier Per Month For Families
AI & Technology

AppleCare One Now Has A $50 Tier Per Month For Families

September 10, 2026
Next Post
Good Samaritans Offer Free Mobile Charging on Wall Street

Good Samaritans Offer Free Mobile Charging on Wall Street

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Lawyers tackles challenges in Trump administration’s interpretation of immigration law

Lawyers tackles challenges in Trump administration’s interpretation of immigration law

September 6, 2026
Canada Imposes New Tariffs on U.S. Goods, as Trade War Intensifies – The New York Times

Canada Imposes New Tariffs on U.S. Goods, as Trade War Intensifies – The New York Times

September 8, 2026
Live updates: Rescues underway as two workers pulled alive more than a week after Nepal-China floods – CNN

Live updates: Rescues underway as two workers pulled alive more than a week after Nepal-China floods – CNN

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!