• bitcoinBitcoin(BTC)$80,034.000.58%
  • ethereumEthereum(ETH)$2,480.171.28%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$776.408.15%
  • rippleXRP(XRP)$1.421.98%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$104.102.68%
  • tronTRON(TRX)$0.3341350.89%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.64%
  • HyperliquidHyperliquid(HYPE)$85.560.78%
  • zcashZcash(ZEC)$1,027.251.19%
  • dogecoinDogecoin(DOGE)$0.09320710.92%
  • RainRain(RAIN)$0.0171353.44%
  • moneroMonero(XMR)$545.784.64%
  • USDSUSDS(USDS)$1.00-0.04%
  • chainlinkChainlink(LINK)$12.093.99%
  • whitebitWhiteBIT Coin(WBT)$73.580.67%
  • leo-tokenLEO Token(LEO)$9.25-0.02%
  • cardanoCardano(ADA)$0.2203843.90%
  • stellarStellar(XLM)$0.1851383.83%
  • bitcoin-cashBitcoin Cash(BCH)$257.602.17%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.1104193.28%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.828.87%
  • USD1USD1(USD1)$1.00-0.02%
  • uniswapUniswap(UNI)$6.809.12%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.433.78%
  • hedera-hashgraphHedera(HBAR)$0.0811875.09%
  • avalanche-2Avalanche(AVAX)$7.633.81%
  • suiSui(SUI)$0.806.53%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000066.56%
  • nearNEAR Protocol(NEAR)$2.2612.84%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0570321.83%
  • tether-goldTether Gold(XAUT)$4,427.270.40%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.121.51%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$114.125.58%
  • BittensorBittensor(TAO)$235.235.74%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.36%
  • AsterAster(ASTER)$0.786.25%
  • aaveAave(AAVE)$133.101.87%
  • mantleMantle(MNT)$0.580.91%
  • pax-goldPAX Gold(PAXG)$4,434.580.44%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0574190.55%
  • OndoOndo(ONDO)$0.3688874.37%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

ChatGPT with Eyes and Ears: BuboGPT is an AI Approach That Enables Visual Grounding in Multi-Modal LLMs

August 14, 2023
in AI & Technology
Reading Time: 4 mins read
A A
ChatGPT with Eyes and Ears: BuboGPT is an AI Approach That Enables Visual Grounding in Multi-Modal LLMs
ShareShareShareShareShare

Large Language Models (LLMs) have emerged as game changers in the natural language processing domain. They are becoming a key part of our daily lives. The most famous example of an LLM is ChatGPT, and it is safe to assume almost everybody knows about it at this point, and most of us use it on a daily basis.

LLMs are characterized by their huge size and capacity to learn from vast amounts of text data. This enables them to generate coherent and contextually relevant human-like text. These models are built based on deep learning architectures, such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers), which uses attention mechanisms to capture long-range dependencies in a language.

By leveraging pre-training on large-scale datasets and fine-tuning on specific tasks, LLMs have shown remarkable performance in various language-related tasks, including text generation, sentiment analysis, machine translation, and question-answering. As LLMs continue to improve, they hold immense potential to revolutionize natural language understanding and generation, bridging the gap between machines and human-like language processing.

On the other hand, some people thought LLMs were not using their full potential as they are limited to text input only. They have been working on extending the potential of LLMs beyond language. Some of the studies have successfully integrated LLMs with various input signals, such as images, videos, speech, and audio, to build powerful multi-modal chatbots. 

Though, there is still a long way to go here as most of these models lack the understanding of the relationships between visual objects and other modalities. While visually-enhanced LLMs can generate high-quality descriptions, they do so in a black-box manner without explicitly relating to the visual context. 

Establishing an explicit and informative correspondence between text and other modalities in multi-modal LLMs can enhance user experience and enable a new set of applications for these models. Let us meet with BuboGPT, which tackles this limitation.

BuboGPT is the first attempt to incorporate visual grounding into LLMs by connecting visual objects with other modalities. BuboGPT enables joint multi-modal understanding and chatting for text, vision, and audio by learning a shared representation space that aligns well with pre-trained LLMs.

Visual grounding is not an easy task to achieve, so that plays a crucial part in BuboGPT’s pipeline. To achieve this, BuboGPT builds a pipeline based on a self-attention mechanism. This mechanism establishes fine-grained relations between visual objects and modalities.

The pipeline includes three modules: a tagging module, a grounding module, and an entity-matching module. The tagging module generates relevant text tags/labels for the input image, the grounding module localizes semantic masks or boxes for each tag, and the entity-matching module uses LLM reasoning to retrieve matched entities from the tags and image descriptions. By connecting visual objects and other modalities through language, BuboGPT enhances the understanding of multi-modal inputs.

To enable a multi-modal understanding of arbitrary combinations of inputs, BuboGPT employs a two-stage training scheme similar to Mini-GPT4. In the first stage, it uses ImageBind as the audio encoder, BLIP-2 as the vision encoder, and Vicuna as the LLM to learn a Q-former that aligns vision or audio features with language. In the second stage, it performs multi-modal instruct tuning on a high-quality instruction-following dataset. 

The construction of this dataset is crucial for the LLM to recognize provided modalities and whether the inputs are well-matched. Therefore, BuboGPT builds a novel high-quality dataset with subsets for vision instruction, audio instruction, sound localization with positive image-audio pairs, and image-audio captioning with negative pairs for semantic reasoning. By introducing negative image-audio pairs, BuboGPT learns better multi-modal alignment and exhibits stronger joint understanding capabilities.


Check out the Paper, Github, and Project. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 28k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

New Twitter Rebrands To Tweet.app After Court’s Double-Edged Ruling

How To Check Your MacBook’s Hard Drive Health

Ekrem Çetinkaya received his B.Sc. in 2018, and M.Sc. in 2019 from Ozyegin University, Istanbul, Türkiye. He wrote his M.Sc. thesis about image denoising using deep convolutional networks. He received his Ph.D. degree in 2023 from the University of Klagenfurt, Austria, with his dissertation titled “Video Coding Enhancements for HTTP Adaptive Streaming Using Machine Learning.” His research interests include deep learning, computer vision, video encoding, and multimedia networking.


🔥 Use SQL to predict the future (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

New Twitter Rebrands To Tweet.app After Court’s Double-Edged Ruling
AI & Technology

New Twitter Rebrands To Tweet.app After Court’s Double-Edged Ruling

September 5, 2026
How To Check Your MacBook’s Hard Drive Health
AI & Technology

How To Check Your MacBook’s Hard Drive Health

September 5, 2026
Remote Work As A Worm, Colorful Platformers And Other New Indie Games Worth Checking Out
AI & Technology

Remote Work As A Worm, Colorful Platformers And Other New Indie Games Worth Checking Out

September 5, 2026
OpenAI Plans Misalignment Incident Reporting Framework After Wiki Incident – Unite.AI
AI & Technology

OpenAI Plans Misalignment Incident Reporting Framework After Wiki Incident – Unite.AI

September 5, 2026
Next Post
Jim Cramer on P&G, Boeing and Dow Chemical Earnings

Jim Cramer on P&G, Boeing and Dow Chemical Earnings

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Senate committee postpones Blanche attorney general vote

Senate committee postpones Blanche attorney general vote

September 2, 2026
Cracker Barrel CEO replaced one year after rebrand triggered customer backlash

Cracker Barrel CEO replaced one year after rebrand triggered customer backlash

September 4, 2026
The Pros And Cons Of Using A TV As Your Computer Monitor

The Pros And Cons Of Using A TV As Your Computer Monitor

September 1, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!