• bitcoinBitcoin(BTC)$76,776.00-0.65%
  • ethereumEthereum(ETH)$2,477.74-2.37%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$716.85-2.54%
  • rippleXRP(XRP)$1.34-1.92%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.94-1.82%
  • tronTRON(TRX)$0.3406600.07%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,090.70-5.10%
  • HyperliquidHyperliquid(HYPE)$77.60-3.12%
  • dogecoinDogecoin(DOGE)$0.083328-1.82%
  • RainRain(RAIN)$0.0152621.27%
  • moneroMonero(XMR)$530.76-0.02%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$79.66-0.87%
  • chainlinkChainlink(LINK)$11.28-2.28%
  • leo-tokenLEO Token(LEO)$9.06-0.66%
  • cardanoCardano(ADA)$0.205628-1.21%
  • stellarStellar(XLM)$0.179081-1.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$224.47-2.52%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.53-0.52%
  • uniswapUniswap(UNI)$6.25-2.01%
  • CantonCanton(CC)$0.095159-2.83%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.78%
  • hedera-hashgraphHedera(HBAR)$0.0758231.84%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.35-1.11%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.91%
  • nearNEAR Protocol(NEAR)$2.30-2.47%
  • suiSui(SUI)$0.71-1.79%
  • crypto-com-chainCronos(CRO)$0.0584140.25%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.31-0.05%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.14-2.93%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$112.61-1.13%
  • BittensorBittensor(TAO)$235.050.34%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.20%
  • aaveAave(AAVE)$126.680.53%
  • pax-goldPAX Gold(PAXG)$4,350.12-0.09%
  • AsterAster(ASTER)$0.701.68%
  • mantleMantle(MNT)$0.57-1.27%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572970.87%
  • BitwayBitway(BTW)$0.6620.51%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CMU Researchers Introduce VisualWebArena: An AI Benchmark Designed to Evaluate the Performance of Multimodal Web Agents on Realistic and Visually Stimulating Challenges

February 10, 2024
in AI & Technology
Reading Time: 4 mins read
A A
CMU Researchers Introduce VisualWebArena: An AI Benchmark Designed to Evaluate the Performance of Multimodal Web Agents on Realistic and Visually Stimulating Challenges
ShareShareShareShareShare

The field of Artificial Intelligence (AI) has always had a long-standing goal of automating everyday computer operations using autonomous agents. Basically, the web-based autonomous agents with the ability to reason, plan, and act are a potential way to automate a variety of computer operations. However, the main obstacle to accomplishing this goal is creating agents that can operate computers with ease, process textual and visual inputs, understand complex natural language commands, and carry out activities to accomplish predetermined goals. The majority of currently existing benchmarks in this area have predominantly concentrated on text-based agents.

In order to address these challenges, a team of researchers from Carnegie Mellon University has introduced VisualWebArena, a benchmark designed and developed to evaluate the performance of multimodal web agents on realistic and visually stimulating challenges. This benchmark includes a wide range of complex web-based challenges that assess several aspects of autonomous multimodal agents’ abilities.

In VisualWebArena, agents are required to read image-text inputs accurately, decipher natural language instructions, and perform activities on websites in order to accomplish user-defined goals. A comprehensive assessment has been carried out on the most advanced Large Language Model (LLM)–based autonomous agents, which include many multimodal models. Text-only LLM agents have been found to have certain limitations through both quantitative and qualitative analysis. The gaps in the capabilities of the most advanced multimodal language agents have also been disclosed, thus offering insightful information.

The team has shared that VisualWebArena consists of 910 realistic activities in three different online environments, i.e., Reddit, Shopping, and Classifieds. While the Shopping and Reddit environments are carried over from WebArena, the Classifieds environment is a new addition to real-world data. Unlike WebArena, which does not have this visual need, all challenges offered in VisualWebArena are notable for being visually anchored and requiring a thorough grasp of the content for effective resolution. Since images are used as input, about 25.2% of the tasks require understanding interleaving.

The study has thoroughly compared the current state-of-the-art Large Language Models and Vision-Language Models (VLMs) in terms of their autonomy. The results have demonstrated that powerful VLMs outperform text-based LLMs on VisualWebArena tasks. The highest-achieving VLM agents have shown to attain a success rate of 16.4%, which is significantly lower than the human performance of 88.7%.

An important discrepancy between open-sourced and API-based VLM agents has also been found, highlighting the necessity of thorough assessment metrics. A unique VLM agent has also been suggested, which draws inspiration from the Set-of-Marks prompting strategy. This new approach has shown significant performance benefits, especially on graphically complex web pages, by streamlining the action space. By addressing the shortcomings of LLM agents, this VLM agent has offered a possible way to improve the capabilities of autonomous agents in visually complex web contexts.

In conclusion, VisualWebArena is an amazing solution for providing a framework for assessing multimodal autonomous language agents as well as offering knowledge that may be applied to the creation of more powerful autonomous agents for online tasks.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

How To Get Your Cut Of PlayStation’s $7.85 Million Settlement

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🎯 [FREE AI WEBINAR] ‘Actions in GPTs: Developer Tips, Tricks & Techniques’ (Feb 12, 2024)


Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Get Your Cut Of PlayStation’s .85 Million Settlement
AI & Technology

How To Get Your Cut Of PlayStation’s $7.85 Million Settlement

September 13, 2026
AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks
AI & Technology

Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Next Post
He sounds stretched thin

He sounds stretched thin

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Student Loan Autopay Now Cuts Your Rate a Full Point

Student Loan Autopay Now Cuts Your Rate a Full Point

September 11, 2026
Why that Elizabeth Holmes clip is so excruciating, according to expert

Why that Elizabeth Holmes clip is so excruciating, according to expert

September 9, 2026
2028 Volvo XC40 First Look: Hello new tech, goodbye EV

2028 Volvo XC40 First Look: Hello new tech, goodbye EV

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!