• bitcoinBitcoin(BTC)$75,894.00-2.42%
  • ethereumEthereum(ETH)$2,403.00-4.05%
  • tetherTether(USDT)$1.00-0.04%
  • binancecoinBNB(BNB)$712.89-1.09%
  • rippleXRP(XRP)$1.29-9.04%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$97.20-4.37%
  • tronTRON(TRX)$0.332845-1.47%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-2.43%
  • zcashZcash(ZEC)$1,122.81-2.35%
  • HyperliquidHyperliquid(HYPE)$77.50-2.65%
  • dogecoinDogecoin(DOGE)$0.080038-4.29%
  • RainRain(RAIN)$0.014078-1.11%
  • USDSUSDS(USDS)$1.00-0.04%
  • moneroMonero(XMR)$506.25-1.59%
  • whitebitWhiteBIT Coin(WBT)$78.02-3.08%
  • chainlinkChainlink(LINK)$10.87-5.80%
  • leo-tokenLEO Token(LEO)$8.83-1.42%
  • cardanoCardano(ADA)$0.194407-5.97%
  • stellarStellar(XLM)$0.176264-9.89%
  • Ethena USDeEthena USDe(USDE)$1.00-0.06%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$218.85-1.52%
  • USD1USD1(USD1)$1.00-0.03%
  • litecoinLitecoin(LTC)$51.03-3.53%
  • uniswapUniswap(UNI)$6.28-4.49%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-2.57%
  • CantonCanton(CC)$0.091369-4.84%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.074399-4.63%
  • avalanche-2Avalanche(AVAX)$7.27-3.59%
  • nearNEAR Protocol(NEAR)$2.34-3.77%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.45%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • suiSui(SUI)$0.69-4.28%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.055494-5.65%
  • tether-goldTether Gold(XAUT)$4,323.840.30%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.133.18%
  • BittensorBittensor(TAO)$218.12-6.17%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$111.31-1.69%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.04%
  • pax-goldPAX Gold(PAXG)$4,328.680.35%
  • aaveAave(AAVE)$121.18-4.93%
  • BitwayBitway(BTW)$0.69-3.51%
  • AsterAster(ASTER)$0.68-2.34%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572270.30%
  • mantleMantle(MNT)$0.55-5.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers at Intel Labs Introduce LLaVA-Gemma: A Compact Vision-Language Model Leveraging the Gemma Large Language Model in Two Variants (Gemma-2B and Gemma-7B)

April 7, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Researchers at Intel Labs Introduce LLaVA-Gemma: A Compact Vision-Language Model Leveraging the Gemma Large Language Model in Two Variants (Gemma-2B and Gemma-7B)
ShareShareShareShareShare

Recent advancements in large language models (LLMs) and Multimodal Foundation Models (MMFMs) have spurred interest in large multimodal models (LMMs). Models like GPT-4, LLaVA, and their derivatives have shown remarkable performance in vision-language tasks such as Visual Question Answering and image captioning. However, their high computational demands have prompted exploration into smaller-scale LMMs.

Researchers from Cognitive AI, Intel Labs, introduce LLaVA-Gemma, a suite of vision-language assistants trained from Gemma LLM variants, Gemma-2B and Gemma-7B and inspired by progress in small yet capable visual language models (VLMs) like LLaVA-Phi. LLaVA-Gemma allows researchers to investigate the trade-offs between computational efficiency and the richness of visual and linguistic understanding by possessing two variants with different parameter sizes. Also, the researchers examine how a massively increased token set affects multi-modal performance.

LLaVA-Gemma follows the LLaVA framework with modifications, combining a pretrained vision encoder (like CLIP) and a pretrained language model (such as Gemma) via an MLP connector. It undergoes a two-stage training process: pretraining the MLP connector on a custom dataset, then jointly finetuning the language model and connector on multimodal instruction tuning examples. Deviations include using Gemma models for language backbone, employing the larger DINOv2 image encoder for vision, and exploring skipping the initial pretraining stage for improved performance. Both pretraining and finetuning stages are conducted with and without initial pretraining.

For the 2B backbone, DinoV2 variants outperform CLIP variants on all benchmarks except POPE-F1 and MMVP. Comparing the training and eval speed for the two model sizes, The training time for the Gemma-2B model on 8 Intel Gaudi 2® AI accelerators was 4 hours, while the larger Gemma-7B model required 16 hours to train under the same conditions. This indicates that the Gemma-7B model, with its increased parameter count, takes approximately four times longer to train than the Gemma-2B model. The relative speed of the Gemma7B model is thus 0.25x compared to the Gemma-2B model. These results highlight the trade-off between model size and training efficiency, with larger models requiring significantly more computational resources and time.

Contributions to this research are as follows:

1. Researchers introduce LLaVA-Gemma, an MMFM leveraging compact, powerful Gemma language models for efficient multimodal interactions. 

2. They extensively evaluate Gemma-2B and Gemma-7B model variants, providing valuable insights into the tradeoffs between computational efficiency and the richness of visual and linguistic understanding in LLMs.

3. They present a deep exploration into alternate design choices and visualize attention with relevancy maps to enhance their understanding of the model’s performance and attention.

In conclusion, The research introduces LLaVA-Gemma, a compact vision-language model utilizing Gemma LLM in two variants, Gemma-2B and Gemma-7B. This research provides a unique opportunity for researchers to explore the trade-offs between computational efficiency and multimodal understanding in small-scale models. Evaluations demonstrate the versatility and effectiveness of LLaVA-Gemma across a range of datasets, highlighting its potential as a benchmark for future research in small-scale vision-language models.


Check out the Paper and HF Page. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 39k+ ML SubReddit


YOU MAY ALSO LIKE

Canon’s R8 II Camera Borrowed Its Styling From A Classic SLR Film Camera

How To Get Spotify’s Best Audio Quality

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Canon’s R8 II Camera Borrowed Its Styling From A Classic SLR Film Camera
AI & Technology

Canon’s R8 II Camera Borrowed Its Styling From A Classic SLR Film Camera

September 16, 2026
How To Get Spotify’s Best Audio Quality
AI & Technology

How To Get Spotify’s Best Audio Quality

September 15, 2026
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
AI & Technology

Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend

September 15, 2026
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
AI & Technology

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

September 15, 2026
Next Post
Park Aerospace: Growth Still On The Horizon (NYSE:PKE)

Park Aerospace: Growth Still On The Horizon (NYSE:PKE)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Bitwise XRP ETF: The Institutional Demand Is Real. The Price Signal Isn't

Bitwise XRP ETF: The Institutional Demand Is Real. The Price Signal Isn't

September 14, 2026
0,000 In Debt And Getting A Divorce

$130,000 In Debt And Getting A Divorce

September 9, 2026
Chinese company releases video of humanoid robot walking off assembly line

Chinese company releases video of humanoid robot walking off assembly line

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!