• bitcoinBitcoin(BTC)$77,769.000.94%
  • ethereumEthereum(ETH)$2,570.435.25%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$729.663.17%
  • rippleXRP(XRP)$1.371.72%
  • usd-coinUSDC(USDC)$1.00-0.03%
  • solanaSolana(SOL)$101.892.40%
  • tronTRON(TRX)$0.335865-0.81%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.42%
  • zcashZcash(ZEC)$1,172.443.39%
  • HyperliquidHyperliquid(HYPE)$81.462.31%
  • dogecoinDogecoin(DOGE)$0.0854572.55%
  • RainRain(RAIN)$0.015783-0.65%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$516.542.35%
  • whitebitWhiteBIT Coin(WBT)$81.031.81%
  • chainlinkChainlink(LINK)$11.751.61%
  • leo-tokenLEO Token(LEO)$9.12-0.76%
  • cardanoCardano(ADA)$0.2090721.00%
  • stellarStellar(XLM)$0.1812232.63%
  • bitcoin-cashBitcoin Cash(BCH)$233.013.24%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.05%
  • litecoinLitecoin(LTC)$53.863.51%
  • CantonCanton(CC)$0.098185-1.19%
  • uniswapUniswap(UNI)$6.143.29%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.371.97%
  • nearNEAR Protocol(NEAR)$2.585.39%
  • avalanche-2Avalanche(AVAX)$7.570.03%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • hedera-hashgraphHedera(HBAR)$0.0753770.61%
  • shiba-inuShiba Inu(SHIB)$0.0000054.91%
  • suiSui(SUI)$0.740.23%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0570921.25%
  • MemeCoreMemeCore(M)$1.182.01%
  • tether-goldTether Gold(XAUT)$4,362.060.08%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.782.82%
  • BittensorBittensor(TAO)$236.96-1.08%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.06%
  • mantleMantle(MNT)$0.604.24%
  • aaveAave(AAVE)$125.813.67%
  • pax-goldPAX Gold(PAXG)$4,365.890.13%
  • AsterAster(ASTER)$0.69-0.74%
  • polkadotPolkadot(DOT)$1.07-1.52%
  • OndoOndo(ONDO)$0.3578803.35%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from China Introduce ImageBind-LLM: A Multi-Modality Instruction Tuning Method of Large Language Models (LLMs) via ImageBind

September 17, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Researchers from China Introduce ImageBind-LLM: A Multi-Modality Instruction Tuning Method of Large Language Models (LLMs) via ImageBind
ShareShareShareShareShare

Researchers have recently seen significant improvements in large language models’ (LLMs) instruction tuning. ChatGPT and GPT-4 are general-purpose talking systems that obey human commands in language and visuals. However, they are still unreplicable because of the closed-source constraint. Alpaca, LLaMAAdapter, and related efforts offer to modify the publicly accessible LLaMA into language instruction models using self-generated data in response to this. LLaVA, LLaMA-Adapter, and others integrate visual understanding capabilities into LLMs for image-conditioned generation to accomplish picture instruction tailoring. 

Despite the success of current instruction tuning techniques, more is needed to create an LLM for broad multimodality instructions, such as text, picture, audio, 3D point clouds, and video. The authors of this study from Shanghai Artificial Intelligence Laboratory, CUHK MMLab and vivo AI Lab introduce the ImageBind-LLM multimodality instruction-following model, which effectively fine-tunes LLaMA under the direction of the joint embedding space in the pre-trained ImageBind. As shown in Figure 1, their ImageBind-LLM (b) can respond to input instructions of numerous modalities in addition to pictures, distinct from earlier visual instruction models (a), demonstrating promising extensibility and generalization capacity.

They specifically propose solely using the vision-language data for tweaking multimodality instruction due to ImageBind’s image-aligned multimodality embedding space. For a picture-caption pair, they first extract the global image feature using ImageBind’s frozen image encoder before embedding transformation using a learnable bind network. The converted picture feature is subsequently applied to all transformer layer word tokens in LLaMA, creating the visual context for generating the appropriate textual caption. In contrast to the zero-initialized attention in the LLaMA-Adapter series, their visual injection mechanism is simple and weighted by a trainable zero-initialized gating factor. 

In this effective way, as the training progresses, the instruction cues of ImageBind’s multimodality embeddings may be gradually introduced into LLaMA without interfering with the original language understanding. Using ImageBind for modality-specific encodings, such as text, picture, audio, and video, their ImageBind-LLM acquires the competence to obey instructions of diverse modalities after the basic vision-language training. They use the pre-trained 3D encoder in Point-Bind to encode the input 3D point clouds for instructions in 3D domains. They also provide a training-free visual cache approach for embedding augmentation during inference to address the modality gap between image training and text, audio, 3D, or video-conditioned production. 

Figure 1 compares our multi-modality vs. visual instruction models ImageBind-LLM. ImageBind-LLM performs a universal multi-modality instruction tuning for image, text, audio, video, and 3D, in contrast to earlier efforts [1-3] that are exclusively conditioned on image modality.

The cache model comprises millions of picture features in the training datasets retrieved by ImageBind, which enhances text/audio/3D/video embeddings by obtaining comparable visual characteristics (Tip-Adapter). As a result, verbal replies to multimodal instructions are of greater quality. They test ImageBind-LLM’s multimodality instruction-following capabilities in various circumstances and consistently find it to perform better. 

Overall, their ImageBind-LLM demonstrates the four qualities listed below.

• Instructions with many modes. ImageBind-LLM is optimized to respond to general multimodality inputs, such as image, text, audio, 3D point clouds, and video, and their embedding-space arithmetic represented by ImageBind and Point-Bind. This is different from earlier language and image instruction models. 

• Efficiency Tuning. During training, they freeze ImageBind’s image encoder and adjust partial weights in LLaMA using parameter-efficient approaches like LoRA and bias-norm tuning. They also train the zero-initialized gating factors and the extra bind network. 

• Zero-initialized Injection without Attention. They employ a learnable gating method for progressive knowledge injection, which is more straightforward and efficient, and incorporate the multimodality requirements with all word tokens of LLaMA directly instead of introducing additional instruction signals through attention layers. 

• Retrieval from a cross-modal cache. They offer a visual cache model from image features extracted by ImageBind, which performs cross-modality retrieval for embedding augmentation to address the modality disparity between training (single picture) and inference (many modalities).


Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI

Upgraded In All The Right Places

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI
AI & Technology

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI

September 11, 2026
Upgraded In All The Right Places
AI & Technology

Upgraded In All The Right Places

September 11, 2026
Apple’s iPhone Handoff Feature Will Cost You  A Month On T-Mobile
AI & Technology

Apple’s iPhone Handoff Feature Will Cost You $5 A Month On T-Mobile

September 11, 2026
Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
AI & Technology

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

September 11, 2026
Next Post
Full Show: Bloomberg Technology (10/30)

Full Show: Bloomberg Technology (10/30)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
G-III Apparel Group, Ltd. 2027 Q2 – Results – Earnings Call Presentation (NASDAQ:GIII) 2026-09-07

G-III Apparel Group, Ltd. 2027 Q2 – Results – Earnings Call Presentation (NASDAQ:GIII) 2026-09-07

September 7, 2026
Cardiff Oncology, Inc. (CRDF) Presents at 8th Annual RAS-Targeted Drug Development Summit – Slideshow (NASDAQ:CRDF) 2026-09-10

Cardiff Oncology, Inc. (CRDF) Presents at 8th Annual RAS-Targeted Drug Development Summit – Slideshow (NASDAQ:CRDF) 2026-09-10

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!