• bitcoinBitcoin(BTC)$76,980.00-1.20%
  • ethereumEthereum(ETH)$2,462.90-0.15%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$712.60-0.67%
  • rippleXRP(XRP)$1.34-2.94%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.26-1.88%
  • tronTRON(TRX)$0.338097-0.70%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.80%
  • zcashZcash(ZEC)$1,100.14-10.16%
  • HyperliquidHyperliquid(HYPE)$79.22-4.38%
  • dogecoinDogecoin(DOGE)$0.083469-2.03%
  • RainRain(RAIN)$0.015644-3.24%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$508.620.19%
  • whitebitWhiteBIT Coin(WBT)$79.73-0.98%
  • chainlinkChainlink(LINK)$11.38-3.50%
  • leo-tokenLEO Token(LEO)$9.04-1.68%
  • cardanoCardano(ADA)$0.202931-4.68%
  • stellarStellar(XLM)$0.174325-2.80%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$224.87-8.55%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$52.25-0.20%
  • CantonCanton(CC)$0.096234-4.93%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.45%
  • uniswapUniswap(UNI)$5.98-0.47%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.37-4.68%
  • hedera-hashgraphHedera(HBAR)$0.073720-3.11%
  • nearNEAR Protocol(NEAR)$2.461.76%
  • suiSui(SUI)$0.73-4.39%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.89%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056301-0.43%
  • MemeCoreMemeCore(M)$1.18-1.53%
  • tether-goldTether Gold(XAUT)$4,341.39-0.85%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$114.232.12%
  • BittensorBittensor(TAO)$231.82-7.80%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.20%
  • mantleMantle(MNT)$0.58-2.45%
  • aaveAave(AAVE)$122.15-0.81%
  • pax-goldPAX Gold(PAXG)$4,346.43-0.79%
  • AsterAster(ASTER)$0.69-3.46%
  • polkadotPolkadot(DOT)$1.08-1.11%
  • OndoOndo(ONDO)$0.346527-1.83%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Microsoft and Georgia Tech Introduce VCoder: Versatile Vision Encoders for Multimodal Large Language Models

December 27, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Researchers from Microsoft and Georgia Tech Introduce VCoder: Versatile Vision Encoders for Multimodal Large Language Models
ShareShareShareShareShare

In the evolving landscape of artificial intelligence and machine learning, the integration of visual perception with language processing has become a frontier of innovation. This integration is epitomized in the development of Multimodal Large Language Models (MLLMs), which have shown remarkable prowess in a range of vision-language tasks. However, these models often falter in basic object perception tasks, such as accurately identifying and counting objects within a visual scene. This discrepancy points to a critical need for improvement in the perceptual capabilities of MLLMs, particularly in accurately recognizing both salient and background entities.

The main challenge this research confronts is enhancing the MLLMs’ ability to perceive objects in a visual scene accurately. Current MLLMs, while adept at complex reasoning tasks, often overlook finer details and background elements, leading to inaccuracies in object perception. This issue is further compounded when models are required to count objects or identify less prominent entities in an image. The goal is to refine these models to achieve a more holistic and accurate understanding of visual scenes without compromising their reasoning abilities.

The Versatile vision enCoders (VCoder) method introduced by researchers from Georgia Tech, Microsoft Research, and Picsart AI Research represents an innovative solution to this challenge. VCoder improves MLLMs by incorporating additional perception modalities, such as segmentation or depth maps, into the models. This approach aims to enhance the model’s understanding of the visual world, thereby improving their perception and reasoning capabilities. VCoder operates by using additional vision encoders that project information from perception modalities into the LLM’s space. This involves identifying and reducing higher-order components in weight matrices, focusing on specific layers within the Transformer model. The method is designed to sharpen the models’ object-level perception skills, including counting, without the need for additional training or parameters.

VCoder’s performance was rigorously evaluated against various benchmarks to assess its effectiveness in enhancing object perception tasks. It demonstrated notable improvements in accuracy, particularly in scenarios involving less frequently represented information in training data. This advancement in the models’ robustness and factuality is a significant step forward in the development of MLLMs that are equally adept at perception and reasoning.

The study illustrates that while MLLMs have made significant strides in complex visual reasoning tasks, they often display subpar performance in simpler tasks like counting objects. VCoder, by feeding extra perception modalities as control inputs through additional vision encoders, provides a novel solution to this problem. The researchers used images from the COCO dataset and outputs from off-the-shelf vision perception models to create a COCO Segmentation Text dataset for training and evaluating MLLMs on object perception tasks. They introduced metrics like count score, hallucination score, and depth score to assess object perception abilities in MLLMs.

Extensive experimental evidence proved VCoder’s improved object-level perception skills over existing Multimodal LLMs, including GPT-4V. VCoder was effective in enhancing model performance on less frequently represented information in the training data, indicating an increase in the model’s robustness and factuality. The method allowed MLLMs to handle nuanced and less common data better, thus broadening their applicability and effectiveness.

In conclusion, the VCoder technique marks a significant advance in the optimization of MLLMs. Adopting a selective approach to reducing components in weight matrices successfully enhances these models’ efficiency without imposing additional computational burdens. This approach not only elevates the performance of MLLMs in familiar tasks but also expands their capabilities in processing and understanding complex visual scenes. The research opens new avenues for developing more refined and efficient language models that are proficient in both perception and reasoning.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

How These XL Phones Compete

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🚀 Boost your LinkedIn presence with Taplio: AI-driven content creation, easy scheduling, in-depth analytics, and networking with top creators – Try it free now!.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
AI & Technology

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

September 11, 2026
How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots
AI & Technology

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
AI & Technology

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
Next Post
Sam Altman, Jony Ive poach Apple iPhone design boss to work on AI devices: report

Sam Altman, Jony Ive poach Apple iPhone design boss to work on AI devices: report

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Russia launches deadly strikes on Kyiv as pause during US envoy visits ends – Al Jazeera

Russia launches deadly strikes on Kyiv as pause during US envoy visits ends – Al Jazeera

September 8, 2026
For years, they warned AI could kill all humans. Now people are listening. – The Washington Post

For years, they warned AI could kill all humans. Now people are listening. – The Washington Post

September 10, 2026
District attorney defends Nolan Wells investigation

District attorney defends Nolan Wells investigation

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!