• bitcoinBitcoin(BTC)$80,832.004.51%
  • ethereumEthereum(ETH)$2,504.714.82%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$721.254.50%
  • rippleXRP(XRP)$1.445.78%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.523.33%
  • tronTRON(TRX)$0.3300431.57%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.91%
  • HyperliquidHyperliquid(HYPE)$86.876.39%
  • zcashZcash(ZEC)$939.8914.93%
  • dogecoinDogecoin(DOGE)$0.0870105.64%
  • RainRain(RAIN)$0.0171242.35%
  • moneroMonero(XMR)$517.041.38%
  • USDSUSDS(USDS)$1.000.02%
  • chainlinkChainlink(LINK)$11.836.26%
  • whitebitWhiteBIT Coin(WBT)$73.804.20%
  • leo-tokenLEO Token(LEO)$9.330.91%
  • cardanoCardano(ADA)$0.2209367.30%
  • stellarStellar(XLM)$0.1830632.61%
  • bitcoin-cashBitcoin Cash(BCH)$255.073.84%
  • daiDai(DAI)$1.000.02%
  • CantonCanton(CC)$0.1120582.65%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • USD1USD1(USD1)$1.000.03%
  • uniswapUniswap(UNI)$6.3810.04%
  • litecoinLitecoin(LTC)$51.062.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.372.34%
  • hedera-hashgraphHedera(HBAR)$0.0786064.38%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.504.01%
  • suiSui(SUI)$0.782.52%
  • shiba-inuShiba Inu(SHIB)$0.0000052.33%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0574545.81%
  • tether-goldTether Gold(XAUT)$4,463.951.59%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • nearNEAR Protocol(NEAR)$1.943.28%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.04-2.93%
  • okbOKB(OKB)$108.214.03%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.02%
  • BittensorBittensor(TAO)$225.813.95%
  • aaveAave(AAVE)$133.314.81%
  • pax-goldPAX Gold(PAXG)$4,473.121.51%
  • AsterAster(ASTER)$0.71-2.11%
  • mantleMantle(MNT)$0.571.08%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0575422.61%
  • OndoOndo(ONDO)$0.3624223.18%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Exploring AVFormer: Google AI’s Innovative Approach to Augment Audio-Only Models with Visual Information & Streamlined Domain Adaptation

June 7, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Exploring AVFormer: Google AI’s Innovative Approach to Augment Audio-Only Models with Visual Information & Streamlined Domain Adaptation
ShareShareShareShareShare

One of the biggest obstacles facing automated speech recognition (ASR) systems is their inability to adapt to novel, unbounded domains. Audiovisual ASR (AV-ASR) is a technique for enhancing the accuracy of ASR systems in multimodal video, especially when the audio is loud. This feature is invaluable for movies shot “in the wild” when the speaker’s mouth might not be in view. Models for this task are often large and comprise both visual and audio encoders and datasets for this task tend to be small.

As other AVASR works, it is only taught and tested using instructional videos. As trials by Google’s research team demonstrate, it performs badly when applied to novel domains using only a single training data set. However, several newly released massive audio-only models have been greatly optimized using self-supervised pretraining and tremendous supervised training on audio-only data from audiobooks like LibriLight and LibriSpeech. Models with billions of parameters, widespread availability, and impressive cross-domain generalization are all features of this class of models. The idea is to recycle the massive investment in such models’ training by reusing their weights. Inspiring them are recent efforts that modify frozen foundation models for use in a variety of domains.

While these models retain the advantages of audio-only pretraining for zero-shot generalization, they now integrate visual inputs in a lightweight manner to enable AV-ASR. The AVFormer framework uses light projection layers and trainable adaptors to infuse visual input into a static ASR model. 

🚀 JOIN the fastest ML Subreddit Community

Researchers demonstrate that these can be taught with minimal extra training time and parameters on a modest amount of poorly labeled video data. This reduces the potential for domain shift and catastrophic forgetting associated with end-to-end finetuning. They also incorporate a basic curricular plan during training to guarantee consistency in the finetuning of these adapters, which they demonstrate is essential for the model to interpret auditory and visual data in tandem correctly. Finally, they show that the model beats state-of-the-art zero-shot approaches on three AV-ASR benchmarks from various domains while maintaining respectable performance on baselines that rely just on audio.

Zero-shot generalization across all AV domains is the target without sacrificing quality on audio-only benchmarks. A state-of-the-art ASR model is used as a starting point and then modified for use in unrestricted AV-ASR. The following two elements are used to include visual features derived from a robust pretrained visual model into the model:

  • They use a linear projection of visual elements to incorporate audio tokens.
  • To facilitate domain adaptation, they introduce minimally invasive adapters into the ASR model’s encoder before it is frozen.

Here are some of the architecture’s most crucial parts:

  • Encoder and decoder for frozen conformers
  • Layers of the optical encoder and projection are used for projecting and extracting features from images.
  • Adaptation layers were added to the core infrastructure, specifically for the audio spectrum. 

To facilitate domain adaptation across multiple modalities, the architecture features a frozen Conformer encoder-decoder model and a frozen CLIP encoder (frozen layers shown in grey with a lock symbol), as well as two lightweight trainable modules, a visual projection layer (shown in orange) and bottleneck adapters (shown in blue). Researchers recommend a two-stage approach to curriculum learning, with the first phase focusing on training the adapters (blue) without any visual tokens and the second phase tuning the visual projection layer (orange) while keeping the rest of the model static.

Researchers evaluate AVFormer’s zero-shot performance on the How2, VisSpeech, and Ego4D AV-ASR benchmarks compared to BEST-RQ, the audio version of the model, and AVATAR, the state-of-the-art AV-ASR. When both AVATAR and BEST-RQ are trained on LibriSpeech and the complete HowTo100M dataset, AVFormer still surpasses them. Notably, this requires training 600M parameters for BEST-RQ but only 4M parameters for AVFormer; therefore, it only needs a small subset of the training dataset (5% of HowTo100M). In addition, they compare AVFormer to an audio-only baseline called LibriSpeech and find that it outperforms both.

The state-of-the-art in zero-shot performance on many AV-ASR datasets is compared. LibriSpeech, an audio-only platform, also features performances. Lower WER percentages indicate higher performance. While the entirety of AVATAR and BEST-RQ are finetuned on HowTo100M, AVFormer’s small collection of finetuned parameters allows it to function effectively with as little as 5% of the dataset.

Researchers unveil AVFormer, an efficient tool for converting static examples of state-of-the-art ASR models into those suitable for AVASR. This method is realistic and effective, as seen by its zero-shot efficiency. Tuning the full parameter set of pre-trained models becomes problematic as ASR models grow in size and complexity across domains. The method is parameter efficient, allowing for simultaneous domain transfer and visual input blending.


Check Out The Paper and Blog Article. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Mobile Games Designed to Be Addictive Get More Kid-Friendly

Nvidia Makes $3.5 Billion Bet on MediaTek

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


Check out https://aitoolsclub.com to find 100’s of Cool AI Tools

Credit: Source link

ShareTweetSendSharePin

Related Posts

Mobile Games Designed to Be Addictive Get More Kid-Friendly
AI & Technology

Mobile Games Designed to Be Addictive Get More Kid-Friendly

September 4, 2026
Nvidia Makes .5 Billion Bet on MediaTek
AI & Technology

Nvidia Makes $3.5 Billion Bet on MediaTek

September 4, 2026
Nvidia Makes MediaTek Partnership Even Bigger, Huang Says
AI & Technology

Nvidia Makes MediaTek Partnership Even Bigger, Huang Says

September 4, 2026
We Accelerate Every AI Model in the World, Huang Says
AI & Technology

We Accelerate Every AI Model in the World, Huang Says

September 4, 2026
Next Post
Whole Foods Gets Upgrade on Price and Cost Cuts

Whole Foods Gets Upgrade on Price and Cost Cuts

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Video shows police rescue two people from water off pier

Video shows police rescue two people from water off pier

September 3, 2026
CCTV-Affiliated Account Attacks Anthropic, Sets Terms for US-China AI Talks – Unite.AI

CCTV-Affiliated Account Attacks Anthropic, Sets Terms for US-China AI Talks – Unite.AI

August 31, 2026
Odd Lots: Why Many Communities Are Skeptical of the AI Boom

Odd Lots: Why Many Communities Are Skeptical of the AI Boom

August 29, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!