• bitcoinBitcoin(BTC)$77,253.000.53%
  • ethereumEthereum(ETH)$2,512.542.73%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$735.043.29%
  • rippleXRP(XRP)$1.361.57%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.632.26%
  • tronTRON(TRX)$0.3395630.01%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.34%
  • zcashZcash(ZEC)$1,130.155.73%
  • HyperliquidHyperliquid(HYPE)$78.67-0.06%
  • dogecoinDogecoin(DOGE)$0.0844420.96%
  • RainRain(RAIN)$0.015328-2.29%
  • moneroMonero(XMR)$525.383.86%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$80.170.82%
  • chainlinkChainlink(LINK)$11.48-0.02%
  • leo-tokenLEO Token(LEO)$9.140.45%
  • cardanoCardano(ADA)$0.2087680.74%
  • stellarStellar(XLM)$0.1808642.96%
  • bitcoin-cashBitcoin Cash(BCH)$230.571.75%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$53.881.83%
  • CantonCanton(CC)$0.0988770.98%
  • uniswapUniswap(UNI)$6.122.24%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.65%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.45-0.25%
  • hedera-hashgraphHedera(HBAR)$0.074490-0.85%
  • nearNEAR Protocol(NEAR)$2.38-1.39%
  • shiba-inuShiba Inu(SHIB)$0.0000052.64%
  • suiSui(SUI)$0.73-1.62%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0573811.76%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.192.81%
  • tether-goldTether Gold(XAUT)$4,349.340.91%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.405.01%
  • BittensorBittensor(TAO)$234.14-0.43%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.10%
  • aaveAave(AAVE)$125.513.07%
  • mantleMantle(MNT)$0.581.19%
  • pax-goldPAX Gold(PAXG)$4,354.780.88%
  • AsterAster(ASTER)$0.68-2.76%
  • polkadotPolkadot(DOT)$1.05-6.70%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054651-2.30%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Researchers Propose a Novel AI Method Called Sparse Fine-grained Contrastive Alignment (SPARC) for Fine-Grained Vision-Language Pretraining

January 24, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Google DeepMind Researchers Propose a Novel AI Method Called Sparse Fine-grained Contrastive Alignment (SPARC) for Fine-Grained Vision-Language Pretraining
ShareShareShareShareShare

Contrastive pre-training using large, noisy image-text datasets has become popular for building general vision representations. These models align global image and text features in a shared space through similar and dissimilar pairs, excelling in tasks like image classification and retrieval. However, they need help with fine-grained tasks such as localization and spatial relationships. Recent efforts incorporate losses between image patches and text tokens to capture finer details, improving performance in fine-grained retrieval, image classification, object detection, and segmentation. Despite these advancements, challenges like computational expense and reliance on pretrained models persist.

Researchers from Google DeepMind have developed SPARse Fine-grained Contrastive Alignment (SPARC), a method for pretraining fine-grained multimodal representations from image-text pairs. SPARC focuses on learning groups of image patches corresponding to individual words in captions. It utilizes a sparse similarity metric to compute language-grouped vision embeddings for each token, allowing detailed information capture in a computationally efficient manner. SPARC combines fine-grained sequence-wise loss with a contrastive loss, enhancing performance in coarse-grained tasks like classification and fine-grained tasks like retrieval, object detection, and segmentation. The method also improves model faithfulness and captioning in foundational vision-language models.

Contrastive image-text pre-training methods like CLIP and ALIGN have popularized learning general visual representations by leveraging textual supervision from large-scale data scraped from the internet.FILIP proposes a cross-modal late interaction mechanism to optimize the token-wise maximum similarity between image and text tokens, addressing the problem of coarse visual representation in global matching. PACL starts from CLIP-pre-trained vision and text encoders and trains an adapter through a contrastive objective to improve fine-grained understanding. GLoRIA builds localized visual representations by contrasting attention-weighted patch embeddings with text tokens, but it becomes computationally intensive for large batch sizes. 

SPARC is a method for pretraining fine-grained multimodal representations from image-text pairs. It uses a sparse similarity metric between image patches and language tokens to learn a grouping of image patches for each token in the caption. The token and language-grouped vision embeddings are then contrasted through a fine-grained sequence-wise loss that only depends on individual samples, enabling detailed information to be learned computationally inexpensively. SPARC combines this fine-grained loss with a contrastive loss between global image and text embeddings to encode global and local information simultaneously. 

The SPARC study assesses its performance across image-level tasks like classification and region-level tasks such as retrieval, object detection, and segmentation. It outperforms other methods in both task types and enhances model faithfulness and captioning in foundational vision-language models. In the evaluation, zero-shot segmentation is conducted by computing patch embeddings and determining class matches through cosine similarity with text embeddings of ground-truth classes. Intersection over Union (IoU) is then calculated to measure the accuracy of predicted and ground-truth segmentations for each class.

SPARC improves performance over competing approaches in image-level tasks (classification) and region-level tasks (retrieval, object detection, and segmentation). SPARC achieves improved model faithfulness and captioning in foundational vision-language models. The evaluation of SPARC includes zero-shot segmentation, where patch embeddings of an image are compared to text embeddings of ground-truth classes. The matching class for each patch is assigned based on maximum cosine similarity, and IoU is calculated for each class. The study mentions using Flamingo’s Perceiver Resampler in training SPARC, which suggests incorporating this method in the experimental setup.

In conclusion, SPARC is a method that helps pretrain fine-grained multimodal representations from image-text pairs. To achieve this, it uses fine-grained contrastive alignment and a contrastive loss between global image and text embeddings. SPARC outperforms competing approaches in image-level tasks such as classification and region-level tasks such as retrieval, object detection, and segmentation. SPARC improves model faithfulness and captioning in foundational vision-language models. To evaluate SPARC, zero-shot segmentation is used where patch embeddings of an image are compared to text embeddings of ground-truth classes. The study suggests using Flamingo’s Perceiver Resampler in training SPARC and recommends incorporating it in the experimental setup.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

d-Matrix Plugs Into Nvidia’s AI Ecosystem

Everything You Need to Know About Apple’s iPhone Duo

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.



Credit: Source link

ShareTweetSendSharePin

Related Posts

d-Matrix Plugs Into Nvidia’s AI Ecosystem
AI & Technology

d-Matrix Plugs Into Nvidia’s AI Ecosystem

September 12, 2026
Everything You Need to Know About Apple’s iPhone Duo
AI & Technology

Everything You Need to Know About Apple’s iPhone Duo

September 12, 2026
BofA: Apple’s Ternus Era Starts With Innovation, AI
AI & Technology

BofA: Apple’s Ternus Era Starts With Innovation, AI

September 12, 2026
Musk’s Boring Co. Gets  Billion Valuation
AI & Technology

Musk’s Boring Co. Gets $23 Billion Valuation

September 12, 2026
Next Post
Nightly News Full Broadcast – Aug. 1

Nightly News Full Broadcast - Aug. 1

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Columbia Select Mid Cap Growth Fund Q2 2026 Portfolio Review

Columbia Select Mid Cap Growth Fund Q2 2026 Portfolio Review

September 8, 2026
OpenAI says autonomous agent hacked a startup

OpenAI says autonomous agent hacked a startup

September 6, 2026
Secret Service agent from Vice President Vance detail removed

Secret Service agent from Vice President Vance detail removed

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!