• bitcoinBitcoin(BTC)$76,430.000.77%
  • ethereumEthereum(ETH)$2,439.261.78%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$724.351.61%
  • rippleXRP(XRP)$1.300.40%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.612.73%
  • tronTRON(TRX)$0.3354200.23%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.65%
  • zcashZcash(ZEC)$1,355.2815.61%
  • HyperliquidHyperliquid(HYPE)$79.142.49%
  • dogecoinDogecoin(DOGE)$0.0809241.19%
  • USDSUSDS(USDS)$1.000.02%
  • moneroMonero(XMR)$498.61-1.46%
  • whitebitWhiteBIT Coin(WBT)$78.630.94%
  • RainRain(RAIN)$0.012862-8.15%
  • chainlinkChainlink(LINK)$11.153.40%
  • leo-tokenLEO Token(LEO)$8.940.53%
  • cardanoCardano(ADA)$0.1970811.24%
  • stellarStellar(XLM)$0.1822723.42%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$220.950.52%
  • USD1USD1(USD1)$1.00-0.02%
  • uniswapUniswap(UNI)$6.706.35%
  • litecoinLitecoin(LTC)$52.182.31%
  • CantonCanton(CC)$0.0989608.94%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.320.23%
  • nearNEAR Protocol(NEAR)$2.7015.46%
  • avalanche-2Avalanche(AVAX)$7.522.93%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.073741-0.92%
  • suiSui(SUI)$0.724.85%
  • shiba-inuShiba Inu(SHIB)$0.0000051.96%
  • crypto-com-chainCronos(CRO)$0.0584205.46%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,308.80-0.30%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$224.893.87%
  • MemeCoreMemeCore(M)$1.121.73%
  • okbOKB(OKB)$111.640.65%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.27%
  • AsterAster(ASTER)$0.737.64%
  • BitwayBitway(BTW)$0.72-7.56%
  • aaveAave(AAVE)$122.291.66%
  • pax-goldPAX Gold(PAXG)$4,310.33-0.38%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0587903.16%
  • mantleMantle(MNT)$0.562.65%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Blink: A New Multimodal LLM Benchmark that Evaluates Core Visual Perception Abilities not Found in Existing Evaluations

April 23, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Blink: A New Multimodal LLM Benchmark that Evaluates Core Visual Perception Abilities not Found in Existing Evaluations
ShareShareShareShareShare

Earlier, with the adoption of computer vision, its studies weren’t content to only scan 2D arrays of flat “patterns.” Rather, they sought to understand images as projections of 3D scenes. Initially, researchers created several intermediate tasks to help with this pursuit. These included learning about optical properties like reflectance, three-dimensional primitives using multi-view reasoning, geometric reasoning using depth estimation, visual correspondence, recognition, keypoint grounding for affordance, and intrinsic images for forensics. Studies have led to constructing new tasks, largely articulated in natural language, in the present era of large language models (LLMs), emphasizing the vision-language relationship learned by multimodal LLMs and less on such perceptual tasks. This could be because of the intrinsic imprecision of language, which makes it difficult to utilize it to mediate many conventional computer vision tasks (e.g., pinpointing a spatial key point through language is tough).

A collaborative effort by researchers from the University of Pennsylvania, the University of Washington, the Allen Institute for AI, the University of California, and Columbia University, this study delves into crucial yet overlooked aspects of visual perception in evaluating multimodal LLMs. Despite their widespread use as evaluation metrics for seminal models like GPT-4V and Gemini-Pro, many of these standards conflate perception with linguistic understanding and reasoning. This work reveals that a ‘blind’ GPT-4 performs well on these ‘multimodal tasks’ when a task-agnostic dense caption is used in place of the picture. 

The study introduces Blink, a novel benchmark for multimodal language models (LLMs) that uniquely focuses on core visual perception abilities not addressed in other evaluations. From basic pattern matching to intermediate reasoning and advanced visual understanding (like visual similarity), Blink’s fourteen classic computer vision challenges encompass a comprehensive range. The image assignments are deliberately challenging, designed to require a genuine understanding of the image’s content rather than relying on superficial labeling.

The researchers revamped every old task by making it a question-and-answer session with picture or textual answers. Blink has 3,800 questions and 7,300 photos, with each question potentially containing many images selected from various datasets. These photographs depict sights inside and outside homes, cities, and nature. Either human beings or datasets are used to generate the questions and options. A human can usually answer every question (except the IQ test) in the Blink of an eye. 

On Blink, the team thoroughly assesses seventeen multimodal LLMs ranging in size from seven to thirty-four bits. Contrary to popular belief, these issues are quite easy for humans to solve (95.70% average accuracy). However, current equipment finds them incredibly challenging, with the GPT-4V model only managing an average accuracy of 51.26%. This is 44.44% poorer than humans and 13.17% better than random guessing. In addition, Blink compared multimodal LLMs to expert vision models and discovered that the latter performs substantially better. On visual correspondence estimation, for instance, the expert beats GPT-4V by 62.8%, relative depth estimation by 38.7%, and multi-view reasoning by 34.6% in terms of absolute accuracy. 

The research findings challenge previous estimates of multimodal LLMs’ perceptual capacities, suggesting they may have been overstated. Moreover, these models could potentially benefit from incorporating insights from specialist models that excel in specific domains. The team envisions Blink as a valuable platform for exploring how multimodal LLMs can integrate more conventional ideas of perception with their state-of-the-art generating capabilities, paving the way for future advancements in the field. 


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


YOU MAY ALSO LIKE

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
AI & Technology

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

September 17, 2026
House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI
AI & Technology

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

September 16, 2026
Snap Introduces A Standalone AI Assistant, Specs Intelligence
AI & Technology

Snap Introduces A Standalone AI Assistant, Specs Intelligence

September 16, 2026
Standalone AR Glasses Are Here
AI & Technology

Standalone AR Glasses Are Here

September 16, 2026
Next Post
ReFantazio, a fantasy RPG from the Persona 5 team, comes out in October

ReFantazio, a fantasy RPG from the Persona 5 team, comes out in October

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
John Hancock Multimanager 2045 Lifetime Portfolio Q2 2026 Commentary (Mutual Fund:JLJAX)

John Hancock Multimanager 2045 Lifetime Portfolio Q2 2026 Commentary (Mutual Fund:JLJAX)

September 16, 2026
ASE Technology Stock Can Miss August’s Run Rate And Still Beat Q3 Guidance (NYSE:ASX)

ASE Technology Stock Can Miss August’s Run Rate And Still Beat Q3 Guidance (NYSE:ASX)

September 15, 2026
Hurricane Lowell barrels towards Hawaii

Hurricane Lowell barrels towards Hawaii

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!