• bitcoinBitcoin(BTC)$78,842.00-2.11%
  • ethereumEthereum(ETH)$2,455.46-1.89%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$695.70-2.54%
  • rippleXRP(XRP)$1.43-4.83%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$96.56-4.57%
  • tronTRON(TRX)$0.338129-1.92%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.87%
  • HyperliquidHyperliquid(HYPE)$81.861.33%
  • dogecoinDogecoin(DOGE)$0.086386-6.26%
  • zcashZcash(ZEC)$783.59-7.21%
  • RainRain(RAIN)$0.01767420.36%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$72.78-2.11%
  • leo-tokenLEO Token(LEO)$9.31-0.51%
  • chainlinkChainlink(LINK)$11.34-3.75%
  • moneroMonero(XMR)$447.41-0.32%
  • cardanoCardano(ADA)$0.210091-6.69%
  • stellarStellar(XLM)$0.183051-7.11%
  • bitcoin-cashBitcoin Cash(BCH)$264.98-3.54%
  • CantonCanton(CC)$0.118517-5.30%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.43-4.33%
  • litecoinLitecoin(LTC)$50.03-4.47%
  • hedera-hashgraphHedera(HBAR)$0.078084-4.63%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.37-3.52%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.55%
  • suiSui(SUI)$0.76-6.86%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.058948-4.01%
  • tether-goldTether Gold(XAUT)$4,614.95-0.38%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • MemeCoreMemeCore(M)$1.183.03%
  • uniswapUniswap(UNI)$4.25-2.83%
  • nearNEAR Protocol(NEAR)$1.86-5.42%
  • okbOKB(OKB)$113.97-2.92%
  • BittensorBittensor(TAO)$233.04-4.40%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • pax-goldPAX Gold(PAXG)$4,623.08-0.40%
  • aaveAave(AAVE)$127.78-2.80%
  • AsterAster(ASTER)$0.700.32%
  • Pump.funPump.fun(PUMP)$0.004837-0.76%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0587601.87%
  • OndoOndo(ONDO)$0.367087-5.72%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Introduces MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks

May 11, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Google AI Introduces MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks
ShareShareShareShareShare

The idea on which vision-language fundamental models are constructed is that a single pre-training can be used to adapt to a wide variety of downstream activities. There are two widely used but distinct training scenarios: 

  • Contrastive learning in the style of CLIP. It trains the model to predict if image-text pairs correctly match, effectively building visual and text representations for the corresponding image and text inputs. It enables image-text and text-image retrieval tasks like selecting the image that best matches a specific description.
  • Next-token prediction: It learns to generate text by predicting the most probable next token in a sequence. It supports text-generative tasks like Image Captioning and Visual Question Answering (VQA) while contrastive learning.

While both methods have shown promising results, pre-trained models not transferable to other tasks tend to perform poorly on text-generation tasks and vice versa. It’s also common for complex or inefficient approaches to be used while adapting to new tasks.

To train jointly for these competing aims and to provide the groundwork for numerous vision-language tasks either directly or by easy adaptation, a recent Google study presents MaMMUT, a simple architecture for joint learning for multimodal tasks. MaMMUT is a condensed multimodal model with only 2B parameters, and it may be trained to achieve contrastive, text-generating, and localization-aware goals. Its simple design—just one image encoder and one text decoder—makes it easy to recycle the two independently. 

🚀 JOIN the fastest ML Subreddit Community

The proposed model comprises a single visual encoder and a single text-decoder linked via cross-attention and trains concurrently on contrastive and text-generative types of losses. Previous work either doesn’t address image-text retrieval tasks or just applies some losses to select aspects of the model. Jointly training contrastive losses and text-generative captioning-like losses is necessary to enable multimodal tasks and fully use the decoder-only model.

There is a considerable performance gain with a smaller model size (nearly half the parameters) for decoder-only models in language learning. One of the biggest obstacles to using them in multimodal situations is reconciling contrastive learning (which relies on unconditional sequence-level representation) and captioning (which optimizes the likelihood of a token based on the tokens that came before it). The researchers offer a two-pass technique to learn these incompatible text representations within the decoder jointly.

Their initial run at learning the caption generation challenge uses cross-attention and causal masking so that the text features can pay attention to the image features and make sequential token predictions. They turn off cross-attention and causal masking to learn the contrastive task on the second pass. While the picture features will remain hidden from the text features, the text features will be able to attend in both directions on all text tokens simultaneously. Both tasks, which were previously difficult to reconcile, may now be handled by the same decoder thanks to the two-pass technique. Even though this model architecture is quite simple, it can serve as a basis for various multimodal tasks.

Since the architecture is trained for several separate tasks, it may be easily integrated into many applications, including image-text and text-image retrieval, visual quality assessment, and captioning. The researchers use sparse video tubes to directly access spatiotemporal information from video for lightweight adaptation. Training to detect bounding boxes via an object-detection head is also required to transfer the model to Open-Vocabulary Detection.

Despite its compact design, MaMMUT provides superior or competitive results in various areas, including image-text and text-image retrieval, video question answering (VideoQA), video captioning, open-vocabulary identification, and VQA. The team highlights that their model achieves better results than much larger models like Flamingo, which is tailored to image+video pre-training and already pre-trained on image-text and video-text data.


Check out the Paper and Google blog. Don’t forget to join our 21k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together

Adam Mosseri Says It’s ‘News To Me’ That Instagram Employees Limited His Exposure To Teen Safety Data

Tanushree Shenwai is a consulting intern at MarktechPost. She is currently pursuing her B.Tech from the Indian Institute of Technology(IIT), Bhubaneswar. She is a Data Science enthusiast and has a keen interest in the scope of application of artificial intelligence in various fields. She is passionate about exploring the new advancements in technologies and their real-life application.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together
AI & Technology

Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together

August 25, 2026
Adam Mosseri Says It’s ‘News To Me’ That Instagram Employees Limited His Exposure To Teen Safety Data
AI & Technology

Adam Mosseri Says It’s ‘News To Me’ That Instagram Employees Limited His Exposure To Teen Safety Data

August 25, 2026
Here Are Three Sick Gibson Guitar Controllers Made For Stage Tour’s Dec. 10 Release
AI & Technology

Here Are Three Sick Gibson Guitar Controllers Made For Stage Tour’s Dec. 10 Release

August 25, 2026
Lego Skylines Is A Cozy City Builder From The Studio That’s Repairing Cities: Skylines II
AI & Technology

Lego Skylines Is A Cozy City Builder From The Studio That’s Repairing Cities: Skylines II

August 25, 2026
Next Post
‘Bloomberg Technology’ Full Show (11/11/2022)

'Bloomberg Technology' Full Show (11/11/2022)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Morning News NOW Full Episode – Aug. 13

Morning News NOW Full Episode – Aug. 13

August 24, 2026
How To Limit Instagram From Using Your Data For AI And Ads

How To Limit Instagram From Using Your Data For AI And Ads

August 22, 2026
Good News: Family’s sweet birthday tradition goes viral

Good News: Family’s sweet birthday tradition goes viral

August 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!