• bitcoinBitcoin(BTC)$78,923.000.45%
  • ethereumEthereum(ETH)$2,495.970.95%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$751.210.09%
  • rippleXRP(XRP)$1.432.82%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$104.011.03%
  • tronTRON(TRX)$0.3385810.45%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,214.207.81%
  • HyperliquidHyperliquid(HYPE)$85.792.04%
  • dogecoinDogecoin(DOGE)$0.0900830.52%
  • RainRain(RAIN)$0.016045-1.19%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$81.747.49%
  • moneroMonero(XMR)$505.42-2.25%
  • chainlinkChainlink(LINK)$12.48-1.44%
  • leo-tokenLEO Token(LEO)$9.19-0.46%
  • cardanoCardano(ADA)$0.2187990.65%
  • stellarStellar(XLM)$0.189018-0.74%
  • bitcoin-cashBitcoin Cash(BCH)$258.45-0.02%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • CantonCanton(CC)$0.1088002.69%
  • USD1USD1(USD1)$1.000.00%
  • uniswapUniswap(UNI)$6.85-2.56%
  • litecoinLitecoin(LTC)$54.02-1.83%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.390.25%
  • hedera-hashgraphHedera(HBAR)$0.079048-2.52%
  • avalanche-2Avalanche(AVAX)$7.98-1.02%
  • suiSui(SUI)$0.81-0.83%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000050.11%
  • nearNEAR Protocol(NEAR)$2.321.63%
  • crypto-com-chainCronos(CRO)$0.0597563.17%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,397.68-0.41%
  • MemeCoreMemeCore(M)$1.180.63%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$256.460.36%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.24-0.84%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.01%
  • mantleMantle(MNT)$0.631.12%
  • AsterAster(ASTER)$0.76-0.93%
  • polkadotPolkadot(DOT)$1.1911.21%
  • aaveAave(AAVE)$129.00-1.94%
  • pax-goldPAX Gold(PAXG)$4,401.73-0.43%
  • Pump.funPump.fun(PUMP)$0.0044604.20%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from UCLA and Google Propose AVIS: A Groundbreaking AI Framework for Autonomous Information Seeking in Visual Question Answering

September 7, 2023
in AI & Technology
Reading Time: 6 mins read
A A
Researchers from UCLA and Google Propose AVIS: A Groundbreaking AI Framework for Autonomous Information Seeking in Visual Question Answering
ShareShareShareShareShare

GPT3, LaMDA, PALM, BLOOM, and LLaMA are just a few examples of large language models (LLMs) that have demonstrated their ability to store and apply vast amounts of information. New skills such as in-context learning, code creation, and common sense reasoning are displayed. A recent push has been to train LLMs to simultaneously process visual and linguistic data. GPT4, Flamingo, and PALI are three illustrious examples of VLMs. They established new benchmarks for numerous tasks, including picture captioning, visual question answering, and open vocabulary recognition. While state-of-the-art LLMs do far better than humans can on tasks involving textual information retrieval, state-of-the-art VLMs struggle with visual information-seeking datasets like Infoseek, Oven, and OK-VQA.

For many reasons, it is difficult for today’s most advanced vision-language models (VLMs) to respond satisfactorily to such inquiries. Youngsters need to be taught to recognize fine-grained categories and specifics in images. Second, their reasoning must be more robust because they employ a smaller language model than state-of-the-art Large Language Models (LLMs). Finally, unlike image search engines, they do not examine the query image against a large corpus of images tagged with different metadata. In this study, researchers from the University of California, Los Angeles (UCLA) and Google provide a novel approach to overcoming these obstacles by merging LLMs with three different types of tools, resulting in state-of-the-art performance on visual information-seeking tasks.

  • Computer programs that help with visual information extraction include object detectors, optical character recognition software, picture captioning models, and visual quality assessment software.
  • An online resource for discovering data and information about the outside world
  • A method of finding relevant results in an image search by mining the metadata of visually related images.

The method employs a planner driven by an LLM to decide which tool to employ and what query to send to it on the fly. In addition, researchers use a reasoner powered by LLM to examine the results of the tools and pull out the relevant data.

To begin, the LLM simplifies a query into a strategy, a program, or a set of instructions. After this, the appropriate APIs are activated to gather data. While promising in simple visual-language challenges, this approach often needs to be revised in more complex real-world scenarios. An all-encompassing strategy cannot be determined from such an initial query. Instead, it calls for continuous iteration in response to ongoing data. The capacity to make decisions on the fly is the key innovation of the proposed strategy. Planning for questions that require visual information is a multi-step process due to the complexity of the assignment. The planner must decide which API to use and what query to submit at each stage. It can only anticipate the utility of answers from sophisticated APIs like image search or predict their output after calling them. Therefore, researchers choose a dynamic strategy rather than the traditional methods, which include upfront planning of process stages and API calls.

Researchers perform a user study to understand better how people make choices while interacting with APIs to find visual information. For the Large Language Model (LLM) to make educated choices about API selection and query formulation, they compile this information into a systematic framework. There are two major ways in which the system benefits from the user data collected. They begin by building a transition graph by deducing the order of user actions. This graph defines the boundaries between states and the steps that can be taken in each. Second, they provide the planner and the reasoner with useful examples of user decision-making.

Key Contributions

  • The team proposes an innovative visual question-answering framework, which uses a large language model (LLM) to strategize the use of external tools dynamically and the investigation of their outputs, therefore learning the knowledge required to deliver answers to the questions posed.
  • The team uses the findings from the user study on how people make decisions to create a systematic plan. This framework instructs the Large Language Model (LLM) to mimic human decision-making when selecting APIs and building queries.
  • The strategy outperforms state-of-the-art solutions on Infoseek and OK-VQA, two benchmarks for knowledge-based visual question answering. In particular, compared to PALI’s 16.0% accuracy on the Infoseek (unseen entity split) dataset, our results are substantially higher at 50.7%.

APIs and other Tools

AVIS (Autonomous Visual Information Seeking with Large Language Models) needs a robust set of resources to respond to visual inquiries requiring proper in-depth information retrieval.

  • Image Captioning Model
  • Visual Question Answering Model
  • Object Detection
  • Image Search
  • OCR
  • Web Search
  • LLM Short QA

Limitations

Currently, AVIS’s primary function is to provide visual responses to questions. The researchers plan to broaden the scope of the LLM-driven dynamic decision-making system to incorporate additional reasoning applications. The current framework also requires the PALM model, a computationally complex LLM. They want to determine if smaller, less computationally intensive language models can make the same decisions.

To sum it up, UCLA and Google researchers have proposed a new method that gives Large Language Models (LLM) access to a wide range of resources for processing visually oriented knowledge queries. The methodology is founded on user study data on human decision-making. It uses a structured framework in which an LLM-powered planner chooses which tools to utilize and how to build queries on the fly. The selected tool’s output will be processed, and a reasoner powered by 9 LLM will extract key information. A visual question is broken down into smaller pieces, and the planner and reasoner work together to solve each one using a variety of tools until they have accumulated enough data to answer the issue.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🚀 Check out Hostinger AI Website Builder (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
AI & Technology

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

September 9, 2026
NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI
AI & Technology

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI

September 9, 2026
How To Change And Customize Your Apple CarPlay Display
AI & Technology

How To Change And Customize Your Apple CarPlay Display

September 8, 2026
Is There Any Benefit To Restarting Your PC Regularly?
AI & Technology

Is There Any Benefit To Restarting Your PC Regularly?

September 8, 2026
Next Post
Politics Push Stock Plunger

Politics Push Stock Plunger

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Comedian Leslie Jones says she created her own success by taking the long path

Comedian Leslie Jones says she created her own success by taking the long path

September 4, 2026
With plans to renovate a D.C. golf course, Trump tees up another fight with residents

With plans to renovate a D.C. golf course, Trump tees up another fight with residents

September 5, 2026
Why PayPal Trades Below The Offer It Turned Down (NASDAQ:PYPL)

Why PayPal Trades Below The Offer It Turned Down (NASDAQ:PYPL)

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!