• bitcoinBitcoin(BTC)$85,330.004.54%
  • ethereumEthereum(ETH)$2,728.962.43%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$786.211.73%
  • rippleXRP(XRP)$1.525.57%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$116.523.77%
  • tronTRON(TRX)$0.3486331.68%
  • zcashZcash(ZEC)$1,492.65-0.92%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.29%
  • HyperliquidHyperliquid(HYPE)$94.220.15%
  • dogecoinDogecoin(DOGE)$0.09967811.65%
  • moneroMonero(XMR)$577.26-6.53%
  • whitebitWhiteBIT Coin(WBT)$85.812.99%
  • RainRain(RAIN)$0.013673-2.41%
  • chainlinkChainlink(LINK)$12.892.61%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2447854.97%
  • leo-tokenLEO Token(LEO)$8.940.29%
  • stellarStellar(XLM)$0.2117036.05%
  • nearNEAR Protocol(NEAR)$4.330.45%
  • uniswapUniswap(UNI)$8.872.65%
  • bitcoin-cashBitcoin Cash(BCH)$264.983.36%
  • Ethena USDeEthena USDe(USDE)$1.00-0.04%
  • avalanche-2Avalanche(AVAX)$10.68-4.51%
  • CantonCanton(CC)$0.1189644.93%
  • litecoinLitecoin(LTC)$60.693.52%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.04%
  • suiSui(SUI)$1.025.08%
  • hedera-hashgraphHedera(HBAR)$0.0931727.34%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.432.11%
  • BittensorBittensor(TAO)$316.9517.37%
  • shiba-inuShiba Inu(SHIB)$0.0000068.04%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0653174.92%
  • MemeCoreMemeCore(M)$1.36-11.13%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,321.27-0.78%
  • okbOKB(OKB)$121.481.77%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.04%
  • aaveAave(AAVE)$141.671.70%
  • BitwayBitway(BTW)$0.809.82%
  • pepePepe(PEPE)$0.00000527.18%
  • EthenaEthena(ENA)$0.210410-2.92%
  • mantleMantle(MNT)$0.645.58%
  • OndoOndo(ONDO)$0.431980-0.47%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Unveils PaliGemma: A Versatile 3B Vision-Language Model VLM with Large-Scale Ambitions

July 12, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Google DeepMind Unveils PaliGemma: A Versatile 3B Vision-Language Model VLM with Large-Scale Ambitions
ShareShareShareShareShare

Vision-language models have evolved significantly over the past few years, with two distinct generations emerging. The first generation, exemplified by CLIP and ALIGN, expanded on large-scale classification pretraining by utilizing web-scale data without requiring extensive human labeling. These models used caption embeddings obtained from language encoders to broaden the vocabulary for classification and retrieval tasks. The second generation, akin to T5 in language modeling, unified captioning and question-answering tasks through generative encoder-decoder modeling. Models like Flamingo, BLIP-2, and PaLI further scaled up these approaches. Recent developments have introduced an additional “instruction tuning” step to enhance user-friendliness. Alongside these advancements, systematic studies have aimed to identify the critical factors in vision-language models. 

Building on this progress, DeepMind researchers present PaliGemma, an open vision-language model combining the strengths of the PaLI vision-language model series with the Gemma family of language models. This innovative approach builds upon the success of previous PaLI iterations, which demonstrated impressive scaling capabilities and performance improvements. PaliGemma integrates a 400M SigLIP vision model with a 2B Gemma language model, resulting in a sub-3B vision-language model that rivals the performance of much larger predecessors like PaLI-X, PaLM-E, and PaLI-3. The Gemma component, derived from the same technology powering the Gemini models, contributes its auto-regressive decoder-only architecture to enhance PaliGemma’s capabilities—this fusion of advanced vision and language processing techniques positions PaliGemma as a significant advancement in multimodal AI.

YOU MAY ALSO LIKE

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

PaliGemma’s architecture comprises three key components: a SigLIP ViTSo400m image encoder, a Gemma-2B v1.0 decoder-only language model, and a linear projection layer. The image encoder transforms input images into a sequence of tokens, while the language model processes text using its SentencePiece tokenizer. The linear projection layer aligns the dimensions of image and text tokens, allowing them to be concatenated. This simple yet effective design enables PaliGemma to handle various tasks, including image classification, captioning, and visual question-answering, through a flexible image+text in, text out API.

The model’s input sequence structure is carefully designed for optimal performance. Image tokens are placed at the beginning, followed by a BOS token, prefix tokens (task description), a SEP token, suffix tokens (prediction), an EOS token, and PAD tokens. This arrangement allows for full attention across the entire input, enabling image tokens to consider the task context when updating their representations. The suffix, which forms the output, is covered by an auto-regressive mask to maintain the generation process’s integrity.

PaliGemma’s training process involves multiple stages to ensure comprehensive visual-language understanding. It begins with unimodal pretraining of individual components, followed by multimodal pretraining on a diverse mixture of tasks. Notably, the image encoder is not frozen during this stage, allowing for improved spatial and relational understanding. The training continues with a resolution increase stage, enhancing the model’s ability to handle high-resolution images and complex tasks. Finally, a transfer stage adapts the base model to specific tasks or use cases, demonstrating PaliGemma’s versatility and effectiveness across various applications.

The results demonstrate PaliGemma’s impressive performance across a wide range of visual-language tasks. The model excels in image captioning, achieving high scores on benchmarks like COCO-Captions and TextCaps. In visual question answering, PaliGemma shows strong performance on various datasets, including VQAv2, GQA, and ScienceQA. The model also performs well on more specialized tasks such as chart understanding (ChartQA) and OCR-related tasks (TextVQA, DocVQA). Notably, PaliGemma exhibits significant improvements when increasing image resolution from 224px to 448px and 896px, especially for tasks involving fine-grained details or text recognition. The model’s versatility is further demonstrated by its ability to handle video input tasks and image segmentation challenges.

Researchers also present the noteworthy findings from the PaliGemma research:

  • Simple square resizing (224×224) performs as well as complex aspect-ratio preserving techniques for segmentation tasks.
  • Researchers introduced CountBenchQA, a new dataset addressing limitations in TallyQA for assessing VLMs’ counting abilities.
  • Discrepancies were found in previously published WidgetCaps numbers, invalidating some comparisons.
  • Image annotations (e.g., red boxes) are as effective as text prompts for indicating widgets to be captioned.
  • RoPE interpolation for image tokens during resolution upscaling (Stage 2) showed no significant benefits.
  • PaliGemma demonstrates unexpected zero-shot generalization to 3D renders from Objaverse without specific training.
  • The model achieves state-of-the-art performance on MMVP, significantly outperforming larger models like GPT4-V and Gemini.

This research introduces PaliGemma, a robust, compact open-base VLM that excels in transfer learning across diverse tasks. This research demonstrates that smaller VLMs can achieve state-of-the-art performance on a wide spectrum of benchmarks, challenging the notion that larger models are always superior. By releasing the base model without instruction tuning, the researchers aim to provide a valuable foundation for further studies in instruction tuning and specific applications. This approach encourages a clearer distinction between base models and fine-tuned versions in VLM research, potentially opening new avenues for more efficient and versatile AI systems in the field of visual-language understanding.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
Next Post
Tesla shareholders to vote on massive  billion package for Elon Musk

Tesla shareholders to vote on massive $56 billion package for Elon Musk

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own

David Sacks: Anthropic, OpenAI Can Slow Down AI on Their Own

September 16, 2026
Nebraska Democrat looks to oust GOP congressman as Iran war crosses 6 months

Nebraska Democrat looks to oust GOP congressman as Iran war crosses 6 months

September 21, 2026
Lindsay Clancy jury deadlocked again as judge orders jurors to keep deliberating

Lindsay Clancy jury deadlocked again as judge orders jurors to keep deliberating

September 19, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!