• bitcoinBitcoin(BTC)$76,660.00-1.63%
  • ethereumEthereum(ETH)$2,374.54-3.25%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$683.44-0.45%
  • rippleXRP(XRP)$1.32-3.67%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$98.46-3.71%
  • tronTRON(TRX)$0.322867-1.75%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • HyperliquidHyperliquid(HYPE)$81.36-2.03%
  • zcashZcash(ZEC)$805.78-4.35%
  • dogecoinDogecoin(DOGE)$0.080797-2.26%
  • RainRain(RAIN)$0.0167540.42%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$519.69-1.56%
  • leo-tokenLEO Token(LEO)$9.26-1.20%
  • whitebitWhiteBIT Coin(WBT)$70.37-2.12%
  • chainlinkChainlink(LINK)$11.03-3.02%
  • cardanoCardano(ADA)$0.193687-2.37%
  • stellarStellar(XLM)$0.172582-2.30%
  • bitcoin-cashBitcoin Cash(BCH)$244.10-0.76%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.112792-5.41%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$48.79-0.03%
  • uniswapUniswap(UNI)$6.005.64%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.32-2.48%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.073332-1.51%
  • avalanche-2Avalanche(AVAX)$7.11-1.76%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.52%
  • suiSui(SUI)$0.71-1.43%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • tether-goldTether Gold(XAUT)$4,310.33-1.59%
  • crypto-com-chainCronos(CRO)$0.054378-3.01%
  • nearNEAR Protocol(NEAR)$1.85-4.08%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.05-3.08%
  • okbOKB(OKB)$105.98-4.59%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.20%
  • BittensorBittensor(TAO)$216.97-4.23%
  • aaveAave(AAVE)$126.740.99%
  • AsterAster(ASTER)$0.700.19%
  • pax-goldPAX Gold(PAXG)$4,319.04-1.55%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056795-0.56%
  • mantleMantle(MNT)$0.550.86%
  • MorphoMorpho(MORPHO)$2.53-0.98%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers From UC Berkeley and Google Introduce an AI Framework that Formulates Visual Question Answering as Modular Code Generation

June 16, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Researchers From UC Berkeley and Google Introduce an AI Framework that Formulates Visual Question Answering as Modular Code Generation
ShareShareShareShareShare

The domain of Artificial Intelligence (AI) is evolving and advancing with the release of every new model and solution. Large Language Models (LLMs), which have recently got very popular due to their incredible abilities, are the main reason for the rise in AI. The subdomains of AI, be it Natural Language Processing, Natural Language Understanding, or Computer Vision, all of these are progressing, and for all good reasons. One research area that has recently garnered a lot of interest from AI and deep learning communities is Visual Question Answering (VQA). VQA is the task of answering open-ended text-based questions about an image. 

Systems adopting Visual Question Answering attempt to appropriately answer questions in natural language regarding an input in the form of an image, and these systems are designed in a way that they understand the contents of an image similar to how humans do and thus effectively communicate the findings. Recently, a team of researchers from UC Berkeley and Google Research has proposed an approach called CodeVQA that addresses visual question answering using modular code generation. CodeVQA formulates VQA as a program synthesis problem and utilizes code-writing language models which take questions as input and generate code as output.

This framework’s main goal is to create Python programs that can call pre-trained visual models and combine their outputs to provide answers. The produced programs manipulate the visual model outputs and derive a solution using arithmetic and conditional logic. In contrast to previous approaches, this framework uses pre-trained language models, pre-trained visual models based on image-caption pairings, a small number of VQA samples, and pre-trained visual models to support in-context learning. 

🚀 JOIN the fastest ML Subreddit Community

To extract specific visual information from the image, such as captions, pixel locations of things, or image-text similarity scores, CodeVQA uses primitive visual APIs wrapped around Visual Language Models. The created code coordinates various APIs to gather the necessary data, then uses the full expressiveness of Python code to analyze the data and reason about it using math, logical structures, feedback loops, and other programming constructs to arrive at a solution.

For evaluation, the team has compared the performance of this new technique to a few-shot baseline that does not use code generation to gauge its effectiveness. COVR and GQA were the two benchmark datasets used in the evaluation, among which the GQA dataset includes multihop questions created from scene graphs of individual Visual Genome photos that humans have manually annotated, and the COVR dataset contains multihop questions about sets of images in the Visual Genome and imSitu datasets. The outcomes showed that CodeVQA performed better on both datasets than the baseline. In particular, it showed an improvement in the accuracy by at least 3% on the COVR dataset and by about 2% on the GQA dataset.

The team has mentioned that CodeVQA is simple to deploy and utilize because it doesn’t require any additional training. It makes use of pre-trained models and a limited number of VQA samples for in-context learning, which aids in tailoring the created programs to particular question-answer patterns. To sum up, this framework is powerful and makes use of the strength of pre-trained LMs and visual models, providing a modular and code-based approach to VQA.


Check Out The Paper and GitHub link. Don’t forget to join our 24k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

GTA VI Extended Look Got 31 Million Views On Netflix Despite Six-Hour Exclusivity Window

Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


➡️ Try: Ake: A Superb Residential Proxy Network (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

GTA VI Extended Look Got 31 Million Views On Netflix Despite Six-Hour Exclusivity Window
AI & Technology

GTA VI Extended Look Got 31 Million Views On Netflix Despite Six-Hour Exclusivity Window

September 2, 2026
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
AI & Technology

Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

September 2, 2026
This Is The Best Setting And Placement For Your Dolby Atmos Soundbar
AI & Technology

This Is The Best Setting And Placement For Your Dolby Atmos Soundbar

September 2, 2026
Aramco Digital and Avathon Partner on Autonomous Operations AI – Unite.AI
AI & Technology

Aramco Digital and Avathon Partner on Autonomous Operations AI – Unite.AI

September 1, 2026
Next Post
The 13 most anticipated games unveiled this summer | The DeanBeat

The 13 most anticipated games unveiled this summer | The DeanBeat

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Stock Market Today: Nvidia Stock Rises as Bullish Forecast Lifts Market Spirits – WSJ

Stock Market Today: Nvidia Stock Rises as Bullish Forecast Lifts Market Spirits – WSJ

August 27, 2026
Devolver Digital Squares Up To Rockstar With A Mascot Platformer Out The Same Day As GTA 6

Devolver Digital Squares Up To Rockstar With A Mascot Platformer Out The Same Day As GTA 6

August 26, 2026
Panic Is Refunding Tariff Fees Paid By Playdate Owners

Panic Is Refunding Tariff Fees Paid By Playdate Owners

August 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!