• bitcoinBitcoin(BTC)$81,243.005.40%
  • ethereumEthereum(ETH)$2,507.585.19%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$727.335.89%
  • rippleXRP(XRP)$1.457.78%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.794.11%
  • tronTRON(TRX)$0.3303351.74%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.91%
  • HyperliquidHyperliquid(HYPE)$86.956.56%
  • zcashZcash(ZEC)$945.9016.67%
  • dogecoinDogecoin(DOGE)$0.0875907.60%
  • RainRain(RAIN)$0.0171352.51%
  • moneroMonero(XMR)$520.583.63%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$11.867.30%
  • whitebitWhiteBIT Coin(WBT)$74.054.86%
  • leo-tokenLEO Token(LEO)$9.351.19%
  • cardanoCardano(ADA)$0.2210239.97%
  • stellarStellar(XLM)$0.1846265.09%
  • bitcoin-cashBitcoin Cash(BCH)$256.795.52%
  • daiDai(DAI)$1.000.01%
  • CantonCanton(CC)$0.1125663.54%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • USD1USD1(USD1)$1.000.03%
  • uniswapUniswap(UNI)$6.399.82%
  • litecoinLitecoin(LTC)$51.283.16%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.373.35%
  • hedera-hashgraphHedera(HBAR)$0.0791685.90%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.504.66%
  • suiSui(SUI)$0.785.38%
  • shiba-inuShiba Inu(SHIB)$0.0000054.03%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0574926.14%
  • tether-goldTether Gold(XAUT)$4,474.162.01%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • nearNEAR Protocol(NEAR)$1.954.53%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • MemeCoreMemeCore(M)$1.04-3.13%
  • okbOKB(OKB)$109.755.41%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.02%
  • BittensorBittensor(TAO)$226.434.82%
  • aaveAave(AAVE)$133.455.38%
  • AsterAster(ASTER)$0.72-0.26%
  • pax-goldPAX Gold(PAXG)$4,482.911.98%
  • mantleMantle(MNT)$0.571.41%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0575483.13%
  • OndoOndo(ONDO)$0.3634924.76%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Allen Institute for AI Introduce VISPROG: A Neuro-Symbolic Approach to Solving Complex and Compositional Visual Tasks Given Natural Language Instructions

June 26, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Researchers from Allen Institute for AI Introduce VISPROG: A Neuro-Symbolic Approach to Solving Complex and Compositional Visual Tasks Given Natural Language Instructions
ShareShareShareShareShare

The search for general-purpose AI systems has facilitated the development of capable end-to-end trainable models, many of which aim to provide a simple natural language interface for a user to engage with the model. Massive-scale unsupervised pretraining followed by supervised multitask training has been the most common method for developing these systems. They eventually want these systems to execute to scale to the indefinitely long tail of difficult jobs. However, this strategy needs a carefully selected dataset for each task. By breaking down difficult activities stated in natural language into simpler phases that can be handled by specialized end-to-end trained models or other programs, they study the usage of big language models to handle the long tail of complex tasks in this work. 

Tell a computer vision program to “Tag the seven main characters from the TV show Big Bang Theory in this image.” The system must first comprehend the purpose of the instruction before carrying out the following steps: detecting faces, retrieving the list of Big Bang Theory’s main characters from a knowledge base, classifying faces using the list of characters, and tagging the image with the names and faces of the characters that were recognized. While several vision and language systems can carry out each task, natural language task execution is outside the purview of end-to-end trained systems. 

Figure 1: A modular and interpretable neuro-symbolic system for compositional visual reasoning – VISPROG. VISPROG develops a program for every new instruction using in-context learning in GPT-3, given a few instances of natural language instructions and the necessary high-level programs, and then runs the program on the input image(s) to get the prediction. Additionally, VISPROG condenses the intermediate outputs into an understandable visual justification. We use VISPROG to do jobs that call for assembling a variety of modules for knowledge retrieval, arithmetic, and logical operations, as well as for analyzing and manipulating images

Researchers from Allen Institute for AI propose VISPROG, a program that takes as input visual information (a single picture or a collection of images) and a natural language command, creates a series of instructions, or a visual program, as they can be called, and then executes these instructions to produce the required result. Each line of a visual program calls one of the many modules the system now supports. Modules can be pre-built language models, OpenCV image processing subroutines, or arithmetic and logical operators. They can also be pre-built computer vision models. The inputs created by running earlier lines of code are consumed by modules, producing intermediate outputs that can be used later.

🔥 Unleash the power of Live Proxies: Private, undetectable residential and mobile IPs.

In the example mentioned earlier, a face detector, GPT-3 as a knowledge retrieval system, and CLIP as an open-vocabulary image classifier are all used by the visual program created by VISPROG to provide the necessary output (see Fig. 1). The generation and execution of programs for vision applications are both enhanced by VISPROG. Neural Module Networks (NMN) combine specialized, differentiable neural modules to create a question-specific, end-to-end trainable network for the visual question answering (VQA) problem. These methods either train a layout generator using REINFORCE’s weak answer supervision or brittle, pre-built semantic parsers to generate the layout of modules deterministically. 

In contrast, VISPROG allows users to build complicated programs without prior training using a potent language model (GPT-3) and limited in-context examples. Invoking trained state-of-the-art models, non-neural Python subroutines, and greater levels of abstraction than NMNs, VISPROG programs are likewise more abstract than NMNs. Due to these benefits, VISPROG is a quick, effective, and versatile neuro-symbolic system. Additionally, VISPROG is very interpreted. First, VISPROG creates simple-to-understand programs whose logical accuracy may be checked by the user. Second, by breaking the prediction down into manageable parts, VISPROG enables the user to examine the results of intermediate phases to spot flaws and, if necessary, make corrections to the logic. 

A completed program with intermediate step outputs (such as text, bounding boxes, segmentation masks, produced pictures, etc.) connected to show the flow of information serves as a visual justification for the prediction. They employ VISPROG for four distinct activities to show off its versatility. These tasks involve common skills (such as picture parsing) but also require specialized thinking and visual manipulation skills. These tasks include:

  1. Answering compositional visual questions.
  2. Zero-shot NLVR on picture pairings.
  3. Factual knowledge object labeling from NL instructions.
  4. Language-guided image manipulation. 

They stress that none of the modules or the language model have been altered in any manner. It takes a few in-context examples with natural language commands and the appropriate programs to adapt VISPROG to any task. VISPROG is simple to use and has substantial gains over a base VQA model on the compositional VQA test of 2.7 points, good zero-shot accuracy on NLVR of 62.4%, and pleasing qualitative and quantitative results on knowledge tagging and picture editing tasks.


Check Out The Paper, Github, and Project Page. Don’t forget to join our 25k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Nvidia Makes MediaTek Partnership Even Bigger, Huang Says

We Accelerate Every AI Model in the World, Huang Says

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Nvidia Makes MediaTek Partnership Even Bigger, Huang Says
AI & Technology

Nvidia Makes MediaTek Partnership Even Bigger, Huang Says

September 4, 2026
We Accelerate Every AI Model in the World, Huang Says
AI & Technology

We Accelerate Every AI Model in the World, Huang Says

September 4, 2026
Meta Agrees to Settle Lawsuit Over Social Media Claims
AI & Technology

Meta Agrees to Settle Lawsuit Over Social Media Claims

September 4, 2026
How Apple’s New CEO Will Shape the iPhone Maker
AI & Technology

How Apple’s New CEO Will Shape the iPhone Maker

September 3, 2026
Next Post
Logistics of Vaccine Distribution Are Formidable, Says DeWitt

Logistics of Vaccine Distribution Are Formidable, Says DeWitt

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Rare addax calf born at Buffalo Zoo

Rare addax calf born at Buffalo Zoo

August 29, 2026
WWE’s Bella Twins prepare for SummerSlam match with Minnesota quiz

WWE’s Bella Twins prepare for SummerSlam match with Minnesota quiz

August 31, 2026
First deaths from lettuce-linked outbreak

First deaths from lettuce-linked outbreak

August 30, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!