• bitcoinBitcoin(BTC)$81,256.000.12%
  • ethereumEthereum(ETH)$2,639.220.13%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$762.22-0.35%
  • rippleXRP(XRP)$1.431.68%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$111.25-2.18%
  • tronTRON(TRX)$0.3393110.23%
  • zcashZcash(ZEC)$1,479.72-0.98%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-2.87%
  • HyperliquidHyperliquid(HYPE)$91.960.24%
  • dogecoinDogecoin(DOGE)$0.0890860.82%
  • moneroMonero(XMR)$542.68-3.71%
  • whitebitWhiteBIT Coin(WBT)$82.95-0.50%
  • RainRain(RAIN)$0.0138762.52%
  • USDSUSDS(USDS)$1.00-0.03%
  • chainlinkChainlink(LINK)$12.510.96%
  • cardanoCardano(ADA)$0.2296152.70%
  • leo-tokenLEO Token(LEO)$8.90-0.04%
  • stellarStellar(XLM)$0.1990202.50%
  • uniswapUniswap(UNI)$8.66-4.55%
  • bitcoin-cashBitcoin Cash(BCH)$254.42-0.11%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.54-3.86%
  • daiDai(DAI)$1.000.02%
  • litecoinLitecoin(LTC)$57.770.87%
  • CantonCanton(CC)$0.1113461.18%
  • USD1USD1(USD1)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$9.6817.24%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.391.02%
  • hedera-hashgraphHedera(HBAR)$0.0815902.39%
  • suiSui(SUI)$0.876.89%
  • MemeCoreMemeCore(M)$1.4712.17%
  • shiba-inuShiba Inu(SHIB)$0.0000061.34%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BittensorBittensor(TAO)$264.395.48%
  • crypto-com-chainCronos(CRO)$0.059539-0.76%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • tether-goldTether Gold(XAUT)$4,372.61-0.11%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$118.551.21%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.14%
  • aaveAave(AAVE)$142.111.34%
  • AsterAster(ASTER)$0.771.65%
  • mantleMantle(MNT)$0.630.08%
  • EthenaEthena(ENA)$0.20485921.03%
  • OndoOndo(ONDO)$0.4209035.39%
  • Pump.funPump.fun(PUMP)$0.004207-4.22%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Introduces Video-to-Audio V2A Technology: Synchronizing Audiovisual Generation

June 23, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Google DeepMind Introduces Video-to-Audio V2A Technology: Synchronizing Audiovisual Generation
ShareShareShareShareShare

Sound is indispensable for enriching human experiences, enhancing communication, and adding emotional depth to media. While AI has made significant progress in various domains, incorporating sound in video-generating models with the same sophistication and nuance as human-created content remains challenging. Producing scores for these silent videos is a significant next step in making generated films.

Google DeepMind introduces video-to-audio (V2A) technology that enables synchronized audiovisual creation. Using a combination of video pixels and text instructions in natural language, V2A creates immersive audio for the on-screen action. The team tried autoregressive and diffusion methods to find the best scalable AI architecture; the results for generating audio using the diffusion method were the most convincing and realistic regarding the synchronization of audio and visuals.

YOU MAY ALSO LIKE

SpaceX Targets September 28 For Starship’s First Orbital Flight

TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

The first step of their video-to-audio technology is compressing the input video. The audio is repeatedly cleaned up from background noise using the diffusion model. Visual input and natural language prompts are used to steer this process, which generates realistic, synced audio that closely follows the instructions. Decoding, waveform generation, and merging the audio and visual data constitute the final step in the audio output process.

Before iteratively running the video and audio prompt input through the diffusion model, V2A encodes them. The next step is to create compressed audio decoded into a waveform. The researchers supplemented the training process with additional information, such as transcripts of spoken dialogue and AI-generated annotations with extensive descriptions of sound, to improve the model’s ability to produce high-quality audio and to train it to make specific sounds.

The presented technology learns to respond to the information in the transcripts or annotations by associating distinct audio occurrences with different visual sceneries by training on video, audio, and the added annotations. To make shots with a dramatic score, realistic sound effects, or dialogue that complements the characters and tone of a video, V2A technology can be paired with video generation models like Veo.

With its ability to create scores for a wide range of classic videos, such as silent films and archival footage, V2A technology opens up a world of creative possibilities. The most exciting aspect is that it can generate as many soundtracks as users desire for any video input. Users can define a “positive prompt” to guide the output towards desired sounds or a “negative prompt” to steer it away from unwanted noises. This flexibility gives users unprecedented control over V2A’s audio output, fostering a spirit of experimentation and enabling them to quickly find the perfect match for their creative vision.

The team is dedicated to ongoing research and development to address a range of issues. They are aware that the quality of the audio output is dependent on the video input, and distortions or artifacts in the video that are outside the training distribution of the model can lead to noticeable audio degradation. They are working on improving lip-syncing for videos with voiceovers. By analyzing the input transcripts, V2A aims to create speech that is perfectly synchronized with the mouth movements of the characters. The team is also aware of the incongruity that can occur when the video model doesn’t correspond to the transcript, leading to eerie lip-syncing. They are actively working to resolve these issues, demonstrating their commitment to maintaining high standards and continuously improving the technology.

The team is actively seeking input from prominent creators and filmmakers, recognizing their invaluable insights and contributions to the development of V2A technology. This collaborative approach ensures that V2A technology can positively influence the creative community, meeting their needs and enhancing their work. To further protect AI-generated content from any abuse, they have integrated the SynthID toolbox into the V2A study and watermarked it all, demonstrating their commitment to the ethical use of the technology.


Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.

[Announcing Gretel Navigator] Create, edit, and augment tabular data with the first compound AI system trusted by EY, Databricks, Google, and Microsoft

Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceX Targets September 28 For Starship’s First Orbital Flight
AI & Technology

SpaceX Targets September 28 For Starship’s First Orbital Flight

September 19, 2026
TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text
AI & Technology

TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

September 19, 2026
Why Is Your iPad Not Charging (And How To Fix It)
AI & Technology

Why Is Your iPad Not Charging (And How To Fix It)

September 19, 2026
How To Block And Unblock A Number On Your Android Phone
AI & Technology

How To Block And Unblock A Number On Your Android Phone

September 19, 2026
Next Post
GOP congressman calls for more transparency from Biden on Hamas

GOP congressman calls for more transparency from Biden on Hamas

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump says war with Iran will end after Election Day

Trump says war with Iran will end after Election Day

September 14, 2026
Virginia mom convicted after 5-year-old son walks alone

Virginia mom convicted after 5-year-old son walks alone

September 15, 2026
Canoodling Central Park lawyer out at Wachtell law firm following embarrassing scandal

Canoodling Central Park lawyer out at Wachtell law firm following embarrassing scandal

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!