• bitcoinBitcoin(BTC)$78,852.002.04%
  • ethereumEthereum(ETH)$2,529.941.02%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$724.750.52%
  • rippleXRP(XRP)$1.424.96%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$102.881.87%
  • tronTRON(TRX)$0.340800-0.06%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.00%
  • zcashZcash(ZEC)$1,144.102.89%
  • HyperliquidHyperliquid(HYPE)$80.813.18%
  • dogecoinDogecoin(DOGE)$0.0847040.16%
  • RainRain(RAIN)$0.014364-6.47%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$513.95-3.97%
  • whitebitWhiteBIT Coin(WBT)$81.561.79%
  • chainlinkChainlink(LINK)$11.581.57%
  • leo-tokenLEO Token(LEO)$8.99-0.69%
  • cardanoCardano(ADA)$0.2110851.14%
  • stellarStellar(XLM)$0.1946098.30%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$225.950.48%
  • USD1USD1(USD1)$1.000.03%
  • litecoinLitecoin(LTC)$54.03-1.18%
  • uniswapUniswap(UNI)$6.441.27%
  • CantonCanton(CC)$0.0968531.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-0.09%
  • hedera-hashgraphHedera(HBAR)$0.0773861.14%
  • avalanche-2Avalanche(AVAX)$7.592.24%
  • nearNEAR Protocol(NEAR)$2.537.58%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000050.66%
  • suiSui(SUI)$0.731.70%
  • crypto-com-chainCronos(CRO)$0.0591641.25%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,312.28-0.86%
  • BittensorBittensor(TAO)$235.65-0.55%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-3.70%
  • okbOKB(OKB)$114.100.84%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.11%
  • aaveAave(AAVE)$127.330.21%
  • BitwayBitway(BTW)$0.713.70%
  • AsterAster(ASTER)$0.700.51%
  • mantleMantle(MNT)$0.570.83%
  • pax-goldPAX Gold(PAXG)$4,317.33-0.85%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0574780.81%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MIT Researchers Propose A New Multimodal Technique That Blends Machine Learning Methods To Learn More Similarly To Humans

June 9, 2023
in AI & Technology
Reading Time: 5 mins read
A A
MIT Researchers Propose A New Multimodal Technique That Blends Machine Learning Methods To Learn More Similarly To Humans
ShareShareShareShareShare

Artificial intelligence is revolutionary in all the major use cases and applications we encounter daily. One such area revolves around a lot of audio and visual media. Think about all the AI-powered apps that can generate funny videos, and artistically astounding images, copy a celebrity’s voice, or note down the entire lecture for you with just one click. All of these models require a huge corpus of data to train. And most of the successful systems rely on annotated datasets to teach themselves. 

The biggest challenge is to store and annotate this data and transform it into usable data points which models can ingest. Easier said than done; companies need help gathering and creating gold-standard data points every year. 

Now, researchers from MIT, the MIT-IBM Watson AI Lab, IBM Research, and other institutions have developed a groundbreaking technique that can efficiently solve these issues by analyzing unlabeled audio and visual data. This model has a lot of promise and potential to improve how current models train. This method resonates with many models, such as speech recognition models, transcribing and audio creation engines, and object detection. It combines two self-supervised learning architectures, contrastive learning, and masked data modeling. This approach follows one basic idea: replicate how humans perceive and understand the world and then replicate the same behavior. 

🚀 JOIN the fastest ML Subreddit Community

As explained by Yuan Gong, an MIT Postdoc, self-supervised learning is essential because if you look at how humans gather and learn from the data, a big portion is without direct supervision. The goal is to enable the same procedure in machines, allowing them to learn as many features as possible from unlabelled data. This training becomes a strong foundation that can be utilized and improved with the help of supervised learning or reinforcement learning, depending on the use cases. 

The technique used here is contrastive audio-visual masked autoencoder (CAV-MAE), which uses a neural network to extract and map meaningful latent representations from audio and visual data. The models can be trained on large datasets of 10-second YouTube clips, utilizing audio and video components. The researchers claimed that CAV-MAE is much better than any other previous approaches because it explicitly emphasizes the association between audio and visual data, which other methods don’t incorporate. 

The CAV-MAE method incorporates two approaches: masked data modeling and contrastive learning. Masked data modeling involves:

  • Taking a video and its matched audio waveform.
  • Converting the audio to a spectrogram.
  • Masking 75% of the audio and video data.

The model then recovers the missing data through a joint encoder/decoder. The reconstruction loss, which measures the difference between the reconstructed prediction and the original audio-visual combination, is used to train the model. The main aim of this approach is to map similar representations close to one another. It does so by associating the relevant parts of audio and video data, such as connecting the mouth movements of spoken words. 

The testing of CAV-MAE-based models with other models proved to be very insightful. The tests were conducted on audio-video retrieval and audio-visual classification tasks. The results demonstrated that contrastive learning and masked data modeling are complementary methods. CAV-MAE outperformed previous techniques in event classification and remained competitive with models trained using industry-level computational resources. In addition, multi-modal data significantly improved fine-tuning of single-modality representation and performance on audio-only event classification tasks.

The researchers at MIT believe that CAV-MAE represents a breakthrough in progress in self-supervised audio-visual learning. They envision that its use cases can range from action recognition, including sports, education, entertainment, motor vehicles, and public safety, to cross-linguistic automatic speech recognition and audio-video generations. While the current method focuses on audio-visual data, the researchers aim to extend it to other modalities, recognizing that human perception involves multiple senses beyond audio and visual cues. 

It will be interesting to see how this approach performs over time and how many existing models try to incorporate such techniques. 

The researchers hope that as machine learning advances, techniques like CAV-MAE will become increasingly valuable, enabling models to understand better and interpret the world.


Check Out The Paper and MIT Blog. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

Anant is a Computer science engineer currently working as a data scientist with experience in Finance and AI products as a service. He is keen to build AI-powered solutions that create better data points and solve daily life problems in an impactful and efficient way.


Check out https://aitoolsclub.com to find 100’s of Cool AI Tools

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI
AI & Technology

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

September 14, 2026
How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error
AI & Technology

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

September 14, 2026
Temporal Raises 0M Series E at .55B Valuation to Expand Operations – Unite.AI
AI & Technology

Temporal Raises $550M Series E at $12.55B Valuation to Expand Operations – Unite.AI

September 14, 2026
What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?
AI & Technology

What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?

September 14, 2026
Next Post
I Have 0,000 in Non-Mortgage Debt!

I Have $300,000 in Non-Mortgage Debt!

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump blasts Anthropic CEO Dario Amodei over AI warning: ‘SICK conspiracy’

Trump blasts Anthropic CEO Dario Amodei over AI warning: ‘SICK conspiracy’

September 14, 2026
Plane carrying Zelensky threatened by drone in Moldova, Norway’s premier says – The Washington Post

Plane carrying Zelensky threatened by drone in Moldova, Norway’s premier says – The Washington Post

September 10, 2026
Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants

Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!