• bitcoinBitcoin(BTC)$79,834.00-0.08%
  • ethereumEthereum(ETH)$2,502.980.62%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$750.38-2.68%
  • rippleXRP(XRP)$1.42-0.01%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$105.681.84%
  • tronTRON(TRX)$0.3359950.41%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.064.99%
  • zcashZcash(ZEC)$1,231.4221.26%
  • HyperliquidHyperliquid(HYPE)$87.312.09%
  • dogecoinDogecoin(DOGE)$0.090303-0.64%
  • RainRain(RAIN)$0.016742-1.86%
  • moneroMonero(XMR)$541.60-1.40%
  • USDSUSDS(USDS)$1.000.03%
  • chainlinkChainlink(LINK)$12.876.37%
  • whitebitWhiteBIT Coin(WBT)$73.650.11%
  • leo-tokenLEO Token(LEO)$9.340.54%
  • cardanoCardano(ADA)$0.220615-0.04%
  • stellarStellar(XLM)$0.1857580.25%
  • bitcoin-cashBitcoin Cash(BCH)$258.710.29%
  • daiDai(DAI)$1.000.01%
  • uniswapUniswap(UNI)$7.170.56%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • CantonCanton(CC)$0.1100531.00%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$54.58-0.66%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.43-0.75%
  • hedera-hashgraphHedera(HBAR)$0.0814780.84%
  • avalanche-2Avalanche(AVAX)$7.761.79%
  • suiSui(SUI)$0.810.90%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$2.4813.44%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.21%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0576861.67%
  • tether-goldTether Gold(XAUT)$4,413.01-0.36%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.130.28%
  • BittensorBittensor(TAO)$265.1111.91%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.300.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.46%
  • AsterAster(ASTER)$0.780.05%
  • aaveAave(AAVE)$134.59-0.09%
  • mantleMantle(MNT)$0.612.85%
  • pax-goldPAX Gold(PAXG)$4,418.55-0.37%
  • OndoOndo(ONDO)$0.3809192.02%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056634-1.26%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Transforming AI Interaction: LLaVAR Outperforms in Visual and Text-Based Comprehension, Marking a New Era in Multimodal Instruction-Following Models

July 2, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Transforming AI Interaction: LLaVAR Outperforms in Visual and Text-Based Comprehension, Marking a New Era in Multimodal Instruction-Following Models
ShareShareShareShareShare

By combining several activities into one instruction, instruction tuning enhances generalization to new tasks. Such capacity to respond to open-ended questions has contributed to the recent chatbot explosion since ChatGPT 2. Visual encoders like CLIP-ViT have recently been added to conversation agents as part of visual instruction-tuned models, allowing for human-agent interaction based on pictures. However, they need help comprehending text inside images, maybe due to the training data’s predominance of natural imagery (e.g., Conceptual Captions and COCO). However, reading comprehension is essential for daily visual perception in humans. Fortunately, OCR techniques make it possible to recognize words from photos. 

The computation (larger context lengths) is increased (naively) by adding recognized texts to the input of visual instruction-tuned models without completely using the encoding capacity of visual encoders. To do this, they suggest gathering instruction-following data that necessitates comprehension of words inside pictures to improve the visual instruction-tuned model end-to-end. By combining manually given directions (such as “Identify any text visible in the image provided.”) with the OCR results, they specifically first gather 422K noisy instruction-following data using text-rich3 images. 

These massive noisy-aligned data significantly enhance the feature alignment between the language decoder and the visual features. Additionally, they ask text-only GPT-4 to produce 16K conversations using OCR results and image captions as high-quality examples of how to follow instructions. Each conversation may contain many turns of question-and-answer pairs. To produce sophisticated instructions depending on the input, this approach necessitates that GPT-4 denoise the OCR data and create unique questions (Figure 1). They supplement the pretraining and finetuning stages of LLaVA correspondingly using noisy and high-quality examples to assess the efficacy of the data that has been obtained. 

🔥 Join The Fastest Growing ML Subreddit
Figure 1 shows how accurate statistics on instruction following are gathered. | https://arxiv.org/pdf/2306.17107.pdf

Researchers from Georgia Tech, Adobe Research, and Stanford University develop LLaVAR, which stands for Large Language and Vision Assistant that Can Read. To better encode minute textual features, they experiment with scaling the input resolution from 2242 to 3362 compared to the original LLaVA. According to the assessment technique, empirically, they give the findings on four text-based VQA datasets together with the ScienceQA finetuning outcomes. Additionally, they use 50 text-rich pictures from LAION and 30 natural images from COCO in the GPT-4-based instruction-following assessment. Additionally, they offer qualitative analysis to measure more sophisticated instruction-following abilities (e.g., on posters, website screenshots, and tweets). 

In conclusion, their contributions include the following: 

• They gather 16K high-quality and 422K noisy instruction-following data. Both have been demonstrated to improve visual instruction tuning. The improved capacity allows their model, LLaVAR, to deliver end-to-end interactions based on diverse online material, including text and images, while only modestly enhancing the model’s performance on natural photos. 

• The training and assessment data, as well as the model milestones, are made publicly available.


YOU MAY ALSO LIKE

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

How To Send High-Quality Images And Videos From Android To iPhone

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
AI & Technology

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

September 6, 2026
How To Send High-Quality Images And Videos From Android To iPhone
AI & Technology

How To Send High-Quality Images And Videos From Android To iPhone

September 6, 2026
What Is Vibe Coding And Why Does It Get So Much Hate?
AI & Technology

What Is Vibe Coding And Why Does It Get So Much Hate?

September 6, 2026
My Content Tracker Idea Became a Real App – Unite.AI
AI & Technology

My Content Tracker Idea Became a Real App – Unite.AI

September 6, 2026
Next Post
When Do We Need An Estate Planner For Our Will? | Mama Bear Legal Forms

When Do We Need An Estate Planner For Our Will? | Mama Bear Legal Forms

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Meta prices Muse Voice Transcribe at alt=

Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?

September 2, 2026
Amb. Mike Waltz says military has ‘everything that it needs’ amid supply concerns: Full interview

Amb. Mike Waltz says military has ‘everything that it needs’ amid supply concerns: Full interview

September 5, 2026
Don’t Get Rid Of Your Old Phone, Turn It Into A Security Camera

Don’t Get Rid Of Your Old Phone, Turn It Into A Security Camera

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!