• bitcoinBitcoin(BTC)$78,763.00-0.52%
  • ethereumEthereum(ETH)$2,495.060.08%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$751.701.62%
  • rippleXRP(XRP)$1.421.35%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.56-0.28%
  • tronTRON(TRX)$0.3387101.08%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,181.613.82%
  • HyperliquidHyperliquid(HYPE)$85.670.95%
  • dogecoinDogecoin(DOGE)$0.089996-0.79%
  • RainRain(RAIN)$0.016084-1.18%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$81.606.49%
  • moneroMonero(XMR)$498.45-3.24%
  • chainlinkChainlink(LINK)$12.46-1.92%
  • leo-tokenLEO Token(LEO)$9.230.33%
  • cardanoCardano(ADA)$0.217526-1.41%
  • stellarStellar(XLM)$0.187871-1.73%
  • bitcoin-cashBitcoin Cash(BCH)$258.10-0.92%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • CantonCanton(CC)$0.1091132.35%
  • USD1USD1(USD1)$1.000.00%
  • uniswapUniswap(UNI)$6.77-3.62%
  • litecoinLitecoin(LTC)$54.12-3.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.390.32%
  • hedera-hashgraphHedera(HBAR)$0.078724-4.55%
  • avalanche-2Avalanche(AVAX)$7.97-1.00%
  • suiSui(SUI)$0.81-1.97%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.78%
  • nearNEAR Protocol(NEAR)$2.28-1.89%
  • crypto-com-chainCronos(CRO)$0.0606105.88%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.214.04%
  • tether-goldTether Gold(XAUT)$4,368.24-1.16%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$255.14-1.95%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.36-2.20%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.42%
  • mantleMantle(MNT)$0.630.25%
  • polkadotPolkadot(DOT)$1.2012.77%
  • AsterAster(ASTER)$0.75-2.58%
  • aaveAave(AAVE)$129.06-2.01%
  • pax-goldPAX Gold(PAXG)$4,372.06-1.17%
  • Pump.funPump.fun(PUMP)$0.004425-1.33%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Introduces LLaVA-Plus: A General-Purpose Multimodal Assistant that Expands the Capabilities of Large Multimodal Models

November 17, 2023
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper Introduces LLaVA-Plus: A General-Purpose Multimodal Assistant that Expands the Capabilities of Large Multimodal Models
ShareShareShareShareShare

Creating general-purpose assistants that can efficiently carry out various real-world activities by following users’ (multimodal) instructions has long been a goal in artificial intelligence. The area has recently seen increased interest in creating foundation models with emerging multimodal understanding and generating skills in open-world challenges. How to create multimodal, general-purpose assistants for computer vision and vision-language activities still needs to be discovered, despite the effectiveness of employing large language models (LLMs) like ChatGPT to produce general-purpose assistants for natural language tasks. 

The current endeavors aimed at creating multimodal agents may be generally divided into two groups: 

(i) End-to-end training using LLMs, in which a succession of Large Multimodal Models (LMMs) are created by continuously training LLMs to learn how to interpret visual information using image-text data and multimodal instruction-following data. Both open-sourced models like LLaVA and MiniGPT-4 and private models like Flamingo and multimodal GPT-4 have shown impressive visual understanding and reasoning skills. While these end-to-end training approaches work well for assisting LMMs in acquiring emergent skills (like in-context learning), creating a cohesive architecture that can smoothly integrate a broad range of abilities—like image segmentation and generation—that are essential for multimodal applications in the real world is still a difficult task. 

(ii) Tool chaining with LLMs, in which the prompts are carefully designed to allow LLMs to call upon various tools (such as vision models that have already been trained) to do desired (sub-)tasks, all without requiring further model training. VisProg, ViperGPT, Visual ChatGPT, X-GPT, and MM-REACT are well-known works. The strength of these approaches is their ability to handle a wide range of visual tasks using (new) tools that can be developed cheaply and integrated into an AI agent. Prompting, however, needs to be more flexible and reliable to enable multimodal agents to reliably choose and activate the right tools (from a broad and varied toolset) and compose their outcomes to provide final solutions for multimodal tasks in the actual world on the go. 

Figure 1: A graphic representation of the possibilities of LLaVA-Plus made possible via skill acquisition.

Researchers from Tsinghua University, Microsoft Research, University of Wisconsin-Madison, HKUST, and IDEA Research in this paper introduce LLaVA-Plus (Large Language and Vision Assistants that Plug and Learn to Use Skills), a multimodal assistant with a broad range of applications that acquires tool usage skills through an end-to-end training methodology that methodically enhances LMMs’ capabilities through visual instruction tweaking. To their knowledge, this is the first documented attempt to combine the advantages of the previously described tool chaining and end-to-end training techniques. The skill repository that comes with LLaVA-Plus has a large selection of vision and vision-language tools. The design is an example of the “Society of Mind” theory, in which individual tools are created for certain tasks and have limited use on their own; nevertheless, when these tools are combined, they provide emergent skills that demonstrate greater intelligence. 

For instance, given users’ multimodal inputs, LLaVA-Plus may create a new workflow instantly, choose and activate pertinent tools from the skill library, and assemble the outcomes of their execution to complete various real-world tasks that are not visible during model training. Through instruction tweaking, LLaVA-Plus may be enhanced over time by adding additional capabilities or instruments. Consider a brand-new multimodal tool created for a certain use case or ability. To build instruction-following data for tuning, they gather relevant user instructions that require this tool along with their execution outcomes or the results that follow. Following instruction tweaking, LLaVA-Plus gains more capabilities as it learns to use this new tool to accomplish jobs previously impossible. 

Additionally, LLaVA-Plus deviates from previous studies on tool usage training for LLMs by utilizing visual cues exclusively in conjunction with multimodal tools. On the other hand, LLaVA-Plus enhances LMM’s capacity for planning and reasoning by using unprocessed visual signals for all the human-AI contact sessions. To summarize, the contributions of their paper are as follows: 

• Use data for a new multimodal instruction-following tool. Using ChatGPT and GPT-4 as labeling tools, they describe a new pipeline for selecting vision-language instruction-following data that is intended for use as a tool in human-AI interaction sessions. 

• A new, large multimodal helper. They have created LLaVA-Plus, a multimodal assistant with a broad range of uses that expands on LLaVA by integrating an extensive and varied collection of external tools that can be quickly chosen, assembled, and engaged to complete tasks. Figure 1 illustrates how LLaVA-Plus greatly expands the possibilities of LMM. Their empirical investigation verifies the efficacy of LLaVA-Plus by showing consistently better results on several benchmarks, especially the new SoTA on VisiT-Bench with a wide range of real-world activities. 

• Source-free. The materials they will make publicly available are the produced multimodal instruction data, the codebase, the LLaVA-Plus checkpoints, and a visual chat demo.


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

How To Change And Customize Your Apple CarPlay Display

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Join The AI Startup Newsletter To Learn About Latest AI Startups

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Change And Customize Your Apple CarPlay Display
AI & Technology

How To Change And Customize Your Apple CarPlay Display

September 8, 2026
Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction
AI & Technology

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

September 8, 2026
SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas
AI & Technology

SpaceX’s Recovered Starship 40 Will Take Months To Get Back To Texas

September 8, 2026
What Is Roku’s Secret Menu And How Do You Unlock It?
AI & Technology

What Is Roku’s Secret Menu And How Do You Unlock It?

September 8, 2026
Next Post
Baltimore Symphony Orchestra welcomes first Black music director

Baltimore Symphony Orchestra welcomes first Black music director

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Nvidia buying AI startup Hugging Face for whopping B

Nvidia buying AI startup Hugging Face for whopping $13B

September 3, 2026
Qwen Developers Open-Sources zg (zvec-grep): A Local-First Search Layer Unifying ripgrep, BM25, and Vector Search

Qwen Developers Open-Sources zg (zvec-grep): A Local-First Search Layer Unifying ripgrep, BM25, and Vector Search

September 2, 2026
Beyerdynamic’s New Aventho Y Headphones Last Up To 90 Hours Per Charge

Beyerdynamic’s New Aventho Y Headphones Last Up To 90 Hours Per Charge

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!