• bitcoinBitcoin(BTC)$78,458.00-0.70%
  • ethereumEthereum(ETH)$2,483.67-0.03%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$752.321.87%
  • rippleXRP(XRP)$1.421.68%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.35-0.30%
  • tronTRON(TRX)$0.3389031.35%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,182.224.08%
  • HyperliquidHyperliquid(HYPE)$84.83-0.30%
  • dogecoinDogecoin(DOGE)$0.089949-0.56%
  • RainRain(RAIN)$0.016223-0.40%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$81.256.28%
  • moneroMonero(XMR)$504.76-2.05%
  • chainlinkChainlink(LINK)$12.50-1.73%
  • leo-tokenLEO Token(LEO)$9.20-0.12%
  • cardanoCardano(ADA)$0.2198020.21%
  • stellarStellar(XLM)$0.187945-2.51%
  • bitcoin-cashBitcoin Cash(BCH)$258.05-0.05%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.1075822.54%
  • litecoinLitecoin(LTC)$54.35-1.61%
  • uniswapUniswap(UNI)$6.75-1.52%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.400.86%
  • hedera-hashgraphHedera(HBAR)$0.079245-3.14%
  • avalanche-2Avalanche(AVAX)$7.99-1.08%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.81-0.55%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.27%
  • nearNEAR Protocol(NEAR)$2.310.17%
  • crypto-com-chainCronos(CRO)$0.0588774.12%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.236.89%
  • tether-goldTether Gold(XAUT)$4,354.04-1.21%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$260.580.98%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.80-1.78%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.11%
  • polkadotPolkadot(DOT)$1.2417.80%
  • mantleMantle(MNT)$0.631.96%
  • AsterAster(ASTER)$0.75-2.74%
  • aaveAave(AAVE)$128.90-2.08%
  • pax-goldPAX Gold(PAXG)$4,356.92-1.22%
  • OndoOndo(ONDO)$0.374420-2.03%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Innovative technology from Typecast allows generative AI to transfer human emotion

November 14, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Innovative technology from Typecast allows generative AI to transfer human emotion
ShareShareShareShareShare

VentureBeat presents: AI Unleashed – An exclusive executive event for enterprise data leaders. Hear from top industry leaders on Nov 15. Reserve your free pass


Language is fundamental to human interaction — but so, too, is the emotion behind it. 

YOU MAY ALSO LIKE

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas

Expressing happiness, sadness, anger, frustration or other feelings helps convey our messages and connect us. 

While generative AI has excelled in many other areas, it has struggled to pick up on these nuances and process the intricacies of human emotion. 

Typecast, a startup using AI to create synthetic voices and videos, says it is revolutionizing in this area with its new Cross-Speaker Emotion Transfer.

VB Event

AI Unleashed

Don’t miss out on AI Unleashed on November 15! This virtual event will showcase exclusive insights and best practices from data leaders including Albertsons, Intuit, and more.

 

Register for free here

The technology allows users to apply emotions recorded from another’s voice to their own while maintaining their unique style, thus enabling faster, more efficient content creation. It is available today through Typecast’s My Voice Maker feature. 

“AI actors have yet to fully capture the emotional range of humans, which is their biggest limiting factor,” said Taesu Kim, CEO and cofounder of the Seoul, South Korea-based Neosapience and Typecast.

With the new Typecast Cross-Speaker Emotion Transfer, “anyone can use AI actors with real emotional depth based on only a small sample of their voice.” 

Decoding emotion

Although emotions usually fall within seven categories — happiness, sadness, anger, fear, surprise and disgust, based on universal facial movements — this is not enough to express the wide variety of emotions in generated speech, Kim noted. 

Speaking is not just a one-to-one mapping between given text and output speech, he pointed out.

“Humans can speak the same sentence in thousands of different ways,” he told VentureBeat in an exclusive interview. We can also show various different emotions in the same sentence (or even the same word). 

For example, recording the sentence “How can you do this to me?” with the emotion prompt “In a sad voice, as if disappointed” would be completely different from the emotion prompt “Angry, like scolding.” 

Similarly, an emotion described in the prompt, “So sad because her father passed away but showing a smile on her face” is complicated and not easily defined in one given category. 

“Humans can speak with different emotions and this leads to rich and diverse conversations,” Kim and other researchers write in a paper on their new technology.

Emotional text-to-speech limitations

Text-to-speech technology has seen significant gains in just a short period of time, led by models ChatGPT, LaMDA, LLama, Bard, Claude and other incumbents and new entrants. 

Emotional text-to-speech has shown considerable progress, too, but it requires a large amount of labeled data that is not easily accessible, Kim explained. Capturing the subtleties of different emotions through voice recordings has been time-consuming and arduous.

Furthermore, “it is extremely hard to record multiple sentences for a long time while consistently preserving emotion,” Kim and his colleagues write. 

In traditional emotional speech synthesis, all training data must have an emotion label, he explained. These methods often require additional emotion encoding or reference audio. 

But this poses a fundamental challenge, as there must be available data for every emotion and every speaker. Furthermore, existing approaches are exposed to mislabeling problems as they have difficulty extracting intensity. 

Cross-speaker emotion transfer becomes ever more difficult when an unseen emotion is assigned to a speaker. The technology has so far performed poorly, as it is unnatural for emotional speech to be produced by a neutral speaker instead of the original speaker. Additionally, emotion intensity control is often not possible. 

“Even if it is possible to acquire an emotional speech dataset,” Kim and his fellow researchers write, “there is still a limitation in controlling emotion intensity.”

Leveraging deep neural networks, unsupervised learning

To address this problem, the researchers first input emotion labels into a generative deep neural network — what Kim called a world first. While successful, this method was not enough to express sophisticated emotions and speaking styles. 

The researchers then built an unsupervised learning algorithm that discerned speaking styles and emotions from a large database. During training, the entire model was trained without any emotion label, Kim said. 

This provided representative numbers from given speeches. While not interpretable to humans, these representations can be used in text-to-speech algorithms to express emotions existing in a database. 

The researchers further trained a perception neural network to translate natural language emotion descriptions into representations. 

“With this technology, the user doesn’t need to record hundreds or thousands of different speaking styles/emotions because it learns from a large database of various emotional voices,” said Kim. 

Adapting to voice characteristics from just snippets

The researchers achieved “transferable and controllable emotion speech synthesis” by leveraging latent representation, they write. Domain adversarial training and cycle-consistency loss disentangle the speaker from style. 

The technology learns from vast quantities of recorded human voices — via audiobooks, videos and other mediums — to analyze and understand emotional patterns, tones and inflections. 

The method successfully transfers emotion to a neutral reading-style speaker with just a handful of labeled samples, Kim explained, and emotion intensity can be controlled by an easy and intuitive scalar value.

This helps to achieve emotion transfer in a natural way without changing identity, he said. Users can record a basic snippet of their voice and apply a range of emotions and intensity, and the AI can adapt to specific voice characteristics. 

Users can select different types of emotional speech recorded by someone else and apply that style to their voice while still preserving their own unique voice identity. By recording just five minutes of their voice, they can express happiness, sadness, anger or other emotions even if they spoke in a normal tone.

Typecast’s technology has been used by Samsung Securities in South Korea (a Samsung Group subsidiary), LG Electronics in Korea and others, and the company has raised $26.8 billion since its founding in 2017. The startup is now working to apply its core technologies in speech synthesis to facial expressions, Kim said.

Controllability critical to generative AI

The media environment is a rapidly-changing one, Kim pointed out. 

In the past, text-based blogs were the most popular corporate media format. But now, short-form videos reign supreme, and companies and individuals must produce much more audio and video content, more frequently. 

“To deliver a corporate message, high-quality expressive voice is essential,” Kim said. 

Fast, affordable production is of utmost importance, he added — manual work by human actors is simply inefficient. 

“Controllability in generative AI is crucial to content creation,” said Kim. “We believe these technologies help ordinary people and companies to unleash their creative potential and improve their productivity.”

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction
AI & Technology

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

September 8, 2026
SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas
AI & Technology

SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas

September 8, 2026
What Is Roku’s Secret Menu And How Do You Unlock It?
AI & Technology

What Is Roku’s Secret Menu And How Do You Unlock It?

September 8, 2026
NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels
AI & Technology

NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels

September 8, 2026
Next Post
Opal's Tadpole proves webcams don't need to be big or boring

Opal's Tadpole proves webcams don't need to be big or boring

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
South Carolina to hold first Democratic 2028 primary, indicating DNC’s push to court minority voters

South Carolina to hold first Democratic 2028 primary, indicating DNC’s push to court minority voters

September 5, 2026
‘It’s our turn to hit them’: Trump vows retaliation against Iran after thwarted surprise attack

‘It’s our turn to hit them’: Trump vows retaliation against Iran after thwarted surprise attack

September 2, 2026
‘Spider-Man’ helps someone cross the street

‘Spider-Man’ helps someone cross the street

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!