• bitcoinBitcoin(BTC)$77,260.00-1.45%
  • ethereumEthereum(ETH)$2,466.28-0.53%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$714.93-0.82%
  • rippleXRP(XRP)$1.35-2.56%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$99.78-2.06%
  • tronTRON(TRX)$0.338684-0.32%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.79%
  • zcashZcash(ZEC)$1,102.95-10.24%
  • HyperliquidHyperliquid(HYPE)$79.95-4.61%
  • dogecoinDogecoin(DOGE)$0.083959-2.09%
  • RainRain(RAIN)$0.015733-3.59%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$510.82-0.52%
  • whitebitWhiteBIT Coin(WBT)$79.97-1.20%
  • chainlinkChainlink(LINK)$11.52-2.87%
  • leo-tokenLEO Token(LEO)$9.09-1.13%
  • cardanoCardano(ADA)$0.208474-2.92%
  • stellarStellar(XLM)$0.176484-2.49%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$226.76-9.32%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$52.950.27%
  • CantonCanton(CC)$0.098918-5.62%
  • uniswapUniswap(UNI)$6.152.20%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-2.04%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.075021-2.50%
  • avalanche-2Avalanche(AVAX)$7.52-4.12%
  • nearNEAR Protocol(NEAR)$2.48-0.37%
  • suiSui(SUI)$0.74-4.47%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.62%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056401-2.76%
  • tether-goldTether Gold(XAUT)$4,350.73-1.52%
  • MemeCoreMemeCore(M)$1.17-4.19%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$111.19-1.74%
  • BittensorBittensor(TAO)$235.83-7.64%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.14%
  • polkadotPolkadot(DOT)$1.153.62%
  • AsterAster(ASTER)$0.71-2.51%
  • mantleMantle(MNT)$0.58-3.34%
  • aaveAave(AAVE)$123.20-1.29%
  • pax-goldPAX Gold(PAXG)$4,355.47-1.52%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.055617-1.51%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Researchers Introduce SpeechX: A Versatile Speech Generation Model Capable of Zero-Shot TTS and Various Speech Transformation Tasks

August 19, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Microsoft Researchers Introduce SpeechX: A Versatile Speech Generation Model Capable of Zero-Shot TTS and Various Speech Transformation Tasks
ShareShareShareShareShare

Multiple machine learning applications, including text, vision, and audio, have seen rapid and significant developments in the technology of generative models. The industry and society have felt significant effects of these developments. Notably, generative models with multi-modal input have become a truly innovative development. Zero-shot text-to-speech (TTS) is a well-known speech generation problem in the speech domain that uses audio-text input. Using just a small audio clip of the intended talker, zero-shot TTS includes turning a text source into speech with that talker’s voice qualities and speaking manner. Fixed dimensional speaker embeddings were used in early research of zero-shot TTS. This method did not effectively support speaker cloning capabilities and restricted its use to TTS alone. 

Recent strategies, however, have included broader concepts such as masked speech prediction and neural codec language modelling. These cutting-edge methods use the audio from the target speaker without compressing it into a one-dimensional representation. As a result, these models have displayed new features, such as voice conversion and speech editing, in addition to their exceptional zero-shot TTS performance. This increased adaptability can greatly expand the potential of speech-generating models. Despite their amazing accomplishments, these current generative models nevertheless have several limits, particularly when handling diverse audio-text-based speech-generating tasks that include converting input speech. 

For example, current voice editing algorithms are limited to processing only clean signals and cannot change spoken content while maintaining background noise. Additionally, the approach discussed places major limitations on its practical applicability by requiring the noisy signal to be surrounded by clean speech segments to complete denoising. Target speaker extraction is a job that is particularly helpful in the context of changing unclean speech. Target speaker extraction is the process of removing a target speaker’s voice from a speech mixture that contains several talkers. You can specify the speaker you want by playing a little speech clip of them. As mentioned, the current generation of generative speech models cannot handle this job despite its potential importance. 

Regression models have historically been used for reliable signal recovery in classical methods for speech enhancement tasks like denoising and target speaker extraction. However, these earlier techniques sometimes need different expert models for every job, which is not optimal given the variety of acoustic disruptions that may occur. Apart from small studies concentrating primarily on certain speech improvement tasks, much research has yet to be done on complete audio text-based speech enhancement models that use reference transcriptions to produce understandable speech. The development of audio-text-based generative speech models integrating generation and transformation capacities takes critical research relevance in light of the factors above and the successful precedents in other disciplines. 

Fig. 1: SpeechX’s general layout. SpeechX uses a neural codec language model that has been trained on the text and acoustic token stream to perform a variety of audio-text-based speech generation tasks, such as noise suppression, speech removal, target speaker extraction, zero-shot TTS, clean speech editing, and noisy speech editing. For certain jobs, text input is not required.

These models have the broad capacity to handle various voice-generating jobs. They suggest that such models should include the following crucial characteristics: 

• Versatility: The unified audio-text-based generative speech models must be able to perform various tasks requiring voice generation from audio and text inputs, similar to unified or foundation models produced in other machine learning domains. Not just zero-shot TTS but also many types of speech alteration, including, for example, speech augmentation and speech editing, should be included in these activities.

• Tolerance: Since unified models are likely to be used in acoustically difficult contexts, they must demonstrate tolerance to diverse acoustic distortions. These models can be useful in real-world situations where background noise is common since they provide dependable performance. 

• Extensibility: Flexible architectures must be used by the unified models to enable smooth task support expansions. One way to do this is to provide room for new components, such as extra modules or input tokens. The models will be better able to adapt to new speech-generating jobs because of this flexibility efficiently. Researchers from Microsoft Corporation in this paper introduce a flexible speech generation model to achieve this goal. It is capable of performing multiple tasks, such as zero-shot TTS, noise suppression using an optional transcript input, speech removal, target speaker extraction using an optional transcript input, and speech editing for both quiet and noisy acoustic environments (Fig. 1). They designate SpeechX1 as their recommended model. 

As with VALL-E, SpeechX adopts a language modeling approach that generates codes of a neural codec model, or acoustic tokens, based on textual and acoustic inputs. To enable the handling of diverse tasks, they incorporate additional tokens in a multi-task learning setup, where the tokens collectively specify the task to be executed. Experimental results, using 60K hours of speech data from LibriLight as a training set, demonstrate the efficacy of SpeechX, showcasing comparable or superior performance compared to expert models in all the tasks above. Notably, SpeechX exhibits novel or expanded capabilities, such as preserving background sounds during speech editing and leveraging reference transcriptions for noise suppression and target speaker extraction. Audio samples showcasing the capabilities of their proposed SpeechX model are available at https://aka.ms/speechx.


Check out the Paper and Project Page. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 28k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

How These XL Phones Compete

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Use SQL to predict the future (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots
AI & Technology

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
AI & Technology

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100
AI & Technology

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

September 10, 2026
Next Post
Big Returns Will Bloom at Flowers Foods, Edwards Lifesciences

Big Returns Will Bloom at Flowers Foods, Edwards Lifesciences

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
LA’s Lucas Museum sued by contractor over -a-year state land deal

LA’s Lucas Museum sued by contractor over $30-a-year state land deal

September 9, 2026
Trump takes questions on Iran and cyclosporiasis outbreak during White House event

Trump takes questions on Iran and cyclosporiasis outbreak during White House event

September 5, 2026
OpenAI’s GPT-Live-1 Arrives in the API at alt=

OpenAI’s GPT-Live-1 Arrives in the API at $0.05 Per Minute – Unite.AI

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!