• bitcoinBitcoin(BTC)$78,348.00-1.00%
  • ethereumEthereum(ETH)$2,478.37-0.84%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$720.38-4.42%
  • rippleXRP(XRP)$1.39-3.24%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$101.91-2.53%
  • tronTRON(TRX)$0.3398340.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.84%
  • zcashZcash(ZEC)$1,232.100.24%
  • HyperliquidHyperliquid(HYPE)$83.91-3.15%
  • dogecoinDogecoin(DOGE)$0.085767-5.16%
  • RainRain(RAIN)$0.0163481.82%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$515.032.46%
  • whitebitWhiteBIT Coin(WBT)$80.87-1.30%
  • chainlinkChainlink(LINK)$11.84-5.81%
  • leo-tokenLEO Token(LEO)$9.190.09%
  • cardanoCardano(ADA)$0.213797-3.16%
  • stellarStellar(XLM)$0.180826-4.71%
  • bitcoin-cashBitcoin Cash(BCH)$250.22-3.39%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.104368-3.42%
  • litecoinLitecoin(LTC)$52.71-3.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-1.78%
  • uniswapUniswap(UNI)$6.03-13.19%
  • avalanche-2Avalanche(AVAX)$7.83-2.41%
  • hedera-hashgraphHedera(HBAR)$0.076997-2.97%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$2.484.03%
  • suiSui(SUI)$0.77-6.22%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.14%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.057923-2.86%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.223.09%
  • tether-goldTether Gold(XAUT)$4,417.720.44%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$254.60-1.78%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.13-1.17%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.08%
  • mantleMantle(MNT)$0.60-5.99%
  • AsterAster(ASTER)$0.72-5.09%
  • aaveAave(AAVE)$124.73-3.86%
  • pax-goldPAX Gold(PAXG)$4,421.940.46%
  • polkadotPolkadot(DOT)$1.11-5.56%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0564140.67%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Microsoft and Georgia Tech Introduce VCoder: Versatile Vision Encoders for Multimodal Large Language Models

December 27, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Researchers from Microsoft and Georgia Tech Introduce VCoder: Versatile Vision Encoders for Multimodal Large Language Models
ShareShareShareShareShare

In the evolving landscape of artificial intelligence and machine learning, the integration of visual perception with language processing has become a frontier of innovation. This integration is epitomized in the development of Multimodal Large Language Models (MLLMs), which have shown remarkable prowess in a range of vision-language tasks. However, these models often falter in basic object perception tasks, such as accurately identifying and counting objects within a visual scene. This discrepancy points to a critical need for improvement in the perceptual capabilities of MLLMs, particularly in accurately recognizing both salient and background entities.

The main challenge this research confronts is enhancing the MLLMs’ ability to perceive objects in a visual scene accurately. Current MLLMs, while adept at complex reasoning tasks, often overlook finer details and background elements, leading to inaccuracies in object perception. This issue is further compounded when models are required to count objects or identify less prominent entities in an image. The goal is to refine these models to achieve a more holistic and accurate understanding of visual scenes without compromising their reasoning abilities.

The Versatile vision enCoders (VCoder) method introduced by researchers from Georgia Tech, Microsoft Research, and Picsart AI Research represents an innovative solution to this challenge. VCoder improves MLLMs by incorporating additional perception modalities, such as segmentation or depth maps, into the models. This approach aims to enhance the model’s understanding of the visual world, thereby improving their perception and reasoning capabilities. VCoder operates by using additional vision encoders that project information from perception modalities into the LLM’s space. This involves identifying and reducing higher-order components in weight matrices, focusing on specific layers within the Transformer model. The method is designed to sharpen the models’ object-level perception skills, including counting, without the need for additional training or parameters.

VCoder’s performance was rigorously evaluated against various benchmarks to assess its effectiveness in enhancing object perception tasks. It demonstrated notable improvements in accuracy, particularly in scenarios involving less frequently represented information in training data. This advancement in the models’ robustness and factuality is a significant step forward in the development of MLLMs that are equally adept at perception and reasoning.

The study illustrates that while MLLMs have made significant strides in complex visual reasoning tasks, they often display subpar performance in simpler tasks like counting objects. VCoder, by feeding extra perception modalities as control inputs through additional vision encoders, provides a novel solution to this problem. The researchers used images from the COCO dataset and outputs from off-the-shelf vision perception models to create a COCO Segmentation Text dataset for training and evaluating MLLMs on object perception tasks. They introduced metrics like count score, hallucination score, and depth score to assess object perception abilities in MLLMs.

Extensive experimental evidence proved VCoder’s improved object-level perception skills over existing Multimodal LLMs, including GPT-4V. VCoder was effective in enhancing model performance on less frequently represented information in the training data, indicating an increase in the model’s robustness and factuality. The method allowed MLLMs to handle nuanced and less common data better, thus broadening their applicability and effectiveness.

In conclusion, the VCoder technique marks a significant advance in the optimization of MLLMs. Adopting a selective approach to reducing components in weight matrices successfully enhances these models’ efficiency without imposing additional computational burdens. This approach not only elevates the performance of MLLMs in familiar tasks but also expands their capabilities in processing and understanding complex visual scenes. The research opens new avenues for developing more refined and efficient language models that are proficient in both perception and reasoning.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🚀 Boost your LinkedIn presence with Taplio: AI-driven content creation, easy scheduling, in-depth analytics, and networking with top creators – Try it free now!.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI
AI & Technology

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

September 10, 2026
LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Next Post
Sam Altman, Jony Ive poach Apple iPhone design boss to work on AI devices: report

Sam Altman, Jony Ive poach Apple iPhone design boss to work on AI devices: report

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
What Trump and Iran are signaling about war as American death toll rises

What Trump and Iran are signaling about war as American death toll rises

September 7, 2026
Beats Can Beat Viruses?! What’s Going On?

Beats Can Beat Viruses?! What’s Going On?

September 7, 2026
Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!