• bitcoinBitcoin(BTC)$76,750.00-0.65%
  • ethereumEthereum(ETH)$2,482.99-1.66%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$719.08-1.16%
  • rippleXRP(XRP)$1.35-1.64%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.62-2.05%
  • tronTRON(TRX)$0.338581-0.43%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • zcashZcash(ZEC)$1,058.23-5.94%
  • HyperliquidHyperliquid(HYPE)$77.82-1.66%
  • dogecoinDogecoin(DOGE)$0.082665-2.54%
  • RainRain(RAIN)$0.015162-3.57%
  • moneroMonero(XMR)$515.72-4.27%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$79.56-0.87%
  • chainlinkChainlink(LINK)$11.26-2.13%
  • leo-tokenLEO Token(LEO)$9.04-1.07%
  • cardanoCardano(ADA)$0.204349-1.75%
  • stellarStellar(XLM)$0.178112-1.07%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$222.05-1.88%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.140.75%
  • uniswapUniswap(UNI)$6.19-3.66%
  • CantonCanton(CC)$0.095240-2.50%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-2.54%
  • hedera-hashgraphHedera(HBAR)$0.0755190.95%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.32-0.97%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.99%
  • nearNEAR Protocol(NEAR)$2.31-2.70%
  • suiSui(SUI)$0.70-3.43%
  • crypto-com-chainCronos(CRO)$0.057444-5.29%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,339.18-0.25%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.13-3.90%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.04-0.53%
  • BittensorBittensor(TAO)$232.980.30%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.04%
  • BitwayBitway(BTW)$0.7231.38%
  • aaveAave(AAVE)$124.97-0.74%
  • pax-goldPAX Gold(PAXG)$4,342.32-0.28%
  • AsterAster(ASTER)$0.69-0.49%
  • mantleMantle(MNT)$0.560.81%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056930-0.18%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Enhancing Vision-Language Models with Chain of Manipulations: A Leap Towards Faithful Visual Reasoning and Error Traceability

February 16, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Enhancing Vision-Language Models with Chain of Manipulations: A Leap Towards Faithful Visual Reasoning and Error Traceability
ShareShareShareShareShare

Big Vision Language Models (VLMs) trained to comprehend vision have shown viability in broad scenarios like visual question answering, visual grounding, and optical character recognition, capitalizing on the strength of Large Language Models (LLMs) in general knowledge of the world.

Humans mark or process the provided photos for convenience and rigor to address the intricate visual challenges; this process is known as manipulation. In the initial training round, most VLMs learned a plethora of intrinsic multimodal abilities, such as grounding boxes and word recognition. Models can execute evidential visual reasoning for issue-solving by mimicking basic human-like behaviors (e.g., cropping, zooming in). However, this approach for model training is not used due to two significant obstacles. 

  1. The first and foremost requirement is producing copious amounts of training data using the evidential visual reasoning paths from preexisting language instruction-answer pairs.
  2. Training VLMs of dedicated architectures while maintaining their preset capabilities is challenging because building a general mechanism with varied manipulations is difficult.

A new study by Tsinghua University and Zhipu AI explores Chain of Manipulations (CoM), a generic mechanism that allows VLMs to execute evidential visual reasoning. VLMs acquire various visual contents (e.g., boxes, messages, images) by applying a sequence of manipulations to the visual input. They initially established an automated data creation platform based on the preexisting image-question-answer corpus. A linguistic annotator with access to a set of manipulations is asked to supply reasoning steps for a specific query, and basic visual tools are used to get the corresponding returns that the manipulations have asked for. Next, the researchers find all the possible manipulation returns and do a traverse on the resulting tree to find all the possible paths that, when combined, lead to the correct answer. 

To build general and reasoning multimodal skills, they offer CogCoM, a 17B VLM trained with a memory-based compatible architecture and a fusion of four categories of data based on the produced data. To arrive at its conclusion, the model uses reasoning to actively adopt various modifications to gain visual contents (such as the new picture img1) and referential regions bbx1 and bbx2. They also present a testbed with detailed visual issues involving reasoning processes and a key points-aware measure to investigate the accuracy of both the final result and the solving process since evaluation resources are scarce. 

The team carries out comprehensive trials on eight benchmarks spanning three classes of abilities: visual grounding (RefCOCO, RefCOCO+, and RefCOCOg), hallucination validation (POPE), and a suggested reasoning examination benchmark (AutoCoM-test). The outcomes demonstrate that methodology consistently provides competitive or better performance. According to the inquiry on the proposed testbed, by combining the reasoning chains produced, CogCoM quickly reaches competitive performance with only a few training steps.

The team discovered that the language solution processes lack variety and that visual tools aren’t always accurate, leading to many unfavorable paths (although making good use of them would be useful). They recommend highlighting these restrictions with dedicated reminders and enhanced visual aids. Additionally, their present model may have performance drops because it re-inputs the altered photos using strict instructions. Incorporating the physical manipulations into the vector space calculations is anticipated to enhance this.

The researchers believe that the suggested visual reasoning process may accelerate VLM development in the area of complicated visual problem-solving. Furthermore, the data generation system that has been introduced has the potential to be used in various training scenarios, which could help advance data-driven machine learning. 


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Which Is Better For Charging Your MacBook?

Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🚀 LLMWare Launches SLIMs: Small Specialized Function-Calling Models for Multi-Step Automation [Check out all the models]


Credit: Source link

ShareTweetSendSharePin

Related Posts

Which Is Better For Charging Your MacBook?
AI & Technology

Which Is Better For Charging Your MacBook?

September 14, 2026
Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI
AI & Technology

Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI

September 13, 2026
How To Fix iMessage “Not Delivered” Error On iPhones
AI & Technology

How To Fix iMessage “Not Delivered” Error On iPhones

September 13, 2026
How To Adjust The Liquid Glass Effect On Your iPhone With iOS 27
AI & Technology

How To Adjust The Liquid Glass Effect On Your iPhone With iOS 27

September 13, 2026
Next Post
Takeaways from Fani Willis’ stunning testimony in Georgia

Takeaways from Fani Willis’ stunning testimony in Georgia

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Nearly 50,000 Yemeni civilians flee as Houthis advance along Red Sea coast – The Guardian

Nearly 50,000 Yemeni civilians flee as Houthis advance along Red Sea coast – The Guardian

September 12, 2026
Carclo plc (CCEGF) Shareholder/Analyst Call Transcript

Carclo plc (CCEGF) Shareholder/Analyst Call Transcript

September 9, 2026
Salesforce Debuts Job-Ready Agentforce Agents and Long-Horizon Runtime – Unite.AI

Salesforce Debuts Job-Ready Agentforce Agents and Long-Horizon Runtime – Unite.AI

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!