• bitcoinBitcoin(BTC)$79,344.001.30%
  • ethereumEthereum(ETH)$2,501.711.24%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$750.46-0.12%
  • rippleXRP(XRP)$1.432.54%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$104.261.52%
  • tronTRON(TRX)$0.3390550.18%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,266.599.85%
  • HyperliquidHyperliquid(HYPE)$86.314.07%
  • dogecoinDogecoin(DOGE)$0.0908641.68%
  • RainRain(RAIN)$0.016183-3.93%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$81.994.70%
  • moneroMonero(XMR)$494.40-2.59%
  • chainlinkChainlink(LINK)$12.11-2.90%
  • leo-tokenLEO Token(LEO)$9.180.02%
  • cardanoCardano(ADA)$0.2202611.44%
  • stellarStellar(XLM)$0.1889490.34%
  • bitcoin-cashBitcoin Cash(BCH)$258.471.32%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.31-1.83%
  • CantonCanton(CC)$0.1060521.58%
  • uniswapUniswap(UNI)$6.69-4.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.401.06%
  • hedera-hashgraphHedera(HBAR)$0.078679-1.52%
  • avalanche-2Avalanche(AVAX)$7.97-0.81%
  • nearNEAR Protocol(NEAR)$2.6013.87%
  • suiSui(SUI)$0.810.58%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • shiba-inuShiba Inu(SHIB)$0.0000050.40%
  • crypto-com-chainCronos(CRO)$0.060159-5.83%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,405.870.10%
  • MemeCoreMemeCore(M)$1.19-0.12%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$265.504.91%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$114.56-0.51%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.42%
  • mantleMantle(MNT)$0.642.71%
  • AsterAster(ASTER)$0.75-0.49%
  • polkadotPolkadot(DOT)$1.189.48%
  • aaveAave(AAVE)$129.45-0.19%
  • Pump.funPump.fun(PUMP)$0.0046126.58%
  • pax-goldPAX Gold(PAXG)$4,409.300.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Researchers Unveil PromptTTS 2: Revolutionizing Text-to-Speech with Enhanced Voice Variability and Cost-Effective Prompt Generation

September 12, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Microsoft Researchers Unveil PromptTTS 2: Revolutionizing Text-to-Speech with Enhanced Voice Variability and Cost-Effective Prompt Generation
ShareShareShareShareShare

The intelligibility and naturalness of synthesized speech have improved due to recent developments in text-to-speech systems. Large-scale TTS systems have been created for multi-speaker settings, and some TTS systems have reached a quality equivalent to single-speaker recordings. Despite these advancements, modeling voice variability is still difficult since different ways of saying the same phrase can communicate additional information, such as emotion and tone. Traditional TTS techniques frequently rely on speaker information or speech prompts to simulate the variability in voice. Still, these techniques are not user-friendly because the speaker ID is pre-defined, and the appropriate speech prompt is difficult to discover or doesn’t exist. 

A more promising approach for modeling voice variability is to utilize text prompts that specify voice features since natural language is a handy interface for users to convey their intent on voice production. This strategy makes it simple to create voices using text prompts. TTS systems based on text prompts are typically trained using a dataset of speech and the text prompt that corresponds to it. The text prompt describing the variability or style of the voice is used to condition how the model generates the voice. 

Text prompt TTS systems continue to face two main difficulties: 

• One-to-Many Challenge: Because voice quality varies from person to person, it is hard for written instructions to represent all speech aspects accurately. Different voice samples may ineluctably correlate to the same prompt. The one-to-many phenomena make TTS model training more challenging and can result in over-fitting or mode collapse. As far as they know, no procedures have been created expressly to address the one-to-many problem in TTS systems based on text prompts.

• Data-Scale Challenge: Since text prompts are uncommon on the internet, compiling a dataset of text prompts defining the voice isn’t easy. 

As a result, vendors are hired to create prompts, which is both expensive and time-consuming. The prompt datasets are typically tiny or private, making it difficult to do further research on prompt-based TTS systems. In their work, they provide PromptTTS 2, which makes a variation network proposal to model the voice variability information of speech not captured by the prompts. It uses the big language model to produce high-quality prompts to overcome the challenges above. They suggest a variation network to anticipate the missing information about voice variability from the text prompt for the one-to-many challenge. The reference speech, thought to include all information on voice variability, is used to train the variation network. 

A text prompt encoder for text prompts, a reference speech encoder for reference speech, and a TTS module to synthesize speech based on the representations retrieved by the text prompt encoder and reference speech encoder make up the TTS model in PromptTTS 2. Based on the immediate representation from text prompt encoder 3, a variation network is trained to predict the reference representation from the reference voice encoder. They may modify the qualities of synthesized speech by using the diffusion model in the variation network to select diverse information about voice variability from Gaussian noise conditioned on text prompts, giving users more freedom when producing voices.

Researchers from Microsoft suggest a pipeline to automatically create text prompts for speech using a speech understanding model to recognize voice characteristics from speech and a big language model to construct text prompts depending on recognition results to address the data-scale difficulty. In particular, they use a speech understanding model to identify the attribute values for each speech sample inside a speech dataset to describe the voice from various features. The text prompt is then created by putting these phrases together, with each attribute’s description given in its sentence. In contrast to earlier studies, which relied on vendors to construct and combine phrases, PromptTTS 2 uses massive language models that have proven capable of performing a range of tasks at a level comparable to that of a person. 

They give LLM instructions to write excellent prompts that include the qualities and connect the phrases into a thorough prompt. Thanks to this completely automated workflow, there is no longer any need for human intervention in prompt authoring. The following is a summary of this paper’s contributions: 

• To solve the one-to-many problem in text prompt-based TTS systems, they build a diffusion model-based variation network to describe the voice variability not covered by the text prompt. The voice variability may be managed by selecting samples from various Gaussian noises conditioned on the text prompt during inference. 

• They build and publish a text prompt dataset produced by a pipeline for text prompt creation and a big language model. The pipeline lessens dependency on providers by producing prompts of high quality. 

• Using 44K hours of speech data, they test PromptTTS 2 on a sizable speech dataset. According to experimental findings, PromptTTS 2 surpasses earlier studies in producing voices that more closely match the text prompt while supporting limiting vocal variability by sampling from Gaussian noise.


Check out the Paper and Samples. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?
AI & Technology

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

September 9, 2026
Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds
AI & Technology

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

September 9, 2026
Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
AI & Technology

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

September 9, 2026
OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI
AI & Technology

OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI

September 9, 2026
Next Post
What Key Analyst Cut to iPhone X Estimates Means for Apple

What Key Analyst Cut to iPhone X Estimates Means for Apple

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Semi-truck hit by train after getting stuck on the tracks

Semi-truck hit by train after getting stuck on the tracks

September 5, 2026
Wildberries says warehouses struck in drone attack

Wildberries says warehouses struck in drone attack

September 5, 2026
Home Equity Is an Asset – But Tapping It Carelessly Is Dangerous

Home Equity Is an Asset – But Tapping It Carelessly Is Dangerous

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!