• bitcoinBitcoin(BTC)$84,459.001.51%
  • ethereumEthereum(ETH)$2,701.842.53%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$775.131.23%
  • rippleXRP(XRP)$1.556.06%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$118.134.49%
  • tronTRON(TRX)$0.337185-0.42%
  • zcashZcash(ZEC)$1,579.847.71%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • HyperliquidHyperliquid(HYPE)$93.283.46%
  • dogecoinDogecoin(DOGE)$0.0962314.57%
  • moneroMonero(XMR)$564.883.36%
  • chainlinkChainlink(LINK)$13.9614.91%
  • whitebitWhiteBIT Coin(WBT)$84.281.28%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2516677.44%
  • RainRain(RAIN)$0.011884-1.08%
  • leo-tokenLEO Token(LEO)$8.84-0.88%
  • stellarStellar(XLM)$0.21944110.96%
  • bitcoin-cashBitcoin Cash(BCH)$338.203.14%
  • nearNEAR Protocol(NEAR)$4.9819.07%
  • uniswapUniswap(UNI)$9.314.45%
  • litecoinLitecoin(LTC)$70.625.56%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • CantonCanton(CC)$0.11821010.63%
  • daiDai(DAI)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$10.363.00%
  • USD1USD1(USD1)$1.000.02%
  • suiSui(SUI)$1.0410.36%
  • hedera-hashgraphHedera(HBAR)$0.0928394.19%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.421.20%
  • shiba-inuShiba Inu(SHIB)$0.0000064.47%
  • BittensorBittensor(TAO)$302.448.27%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0651707.55%
  • BitwayBitway(BTW)$1.098.06%
  • MemeCoreMemeCore(M)$1.20-3.13%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,300.121.14%
  • OndoOndo(ONDO)$0.5531.21%
  • okbOKB(OKB)$119.571.33%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.15%
  • EthenaEthena(ENA)$0.22436411.65%
  • aaveAave(AAVE)$145.577.02%
  • mantleMantle(MNT)$0.683.17%
  • MorphoMorpho(MORPHO)$2.929.75%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Explores If Human Visual Perception can Help Computer Vision Models Outperform in Generalized Tasks

October 20, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper Explores If Human Visual Perception can Help Computer Vision Models Outperform in Generalized Tasks
ShareShareShareShareShare

Human beings possess innate extraordinary perceptual judgments, and when computer vision models are aligned with them, model’s performance can be improved manifold. Various attributes such as scene layout, subject location, camera pose, color, perspective, and semantics help us have a clear picture of the world and objects within. The alignment of vision models with visual perception makes them sensitive to these attributes and more human-like. While it has been established that molding vision models along the lines of human perception helps attain specific goals in certain contexts, such as image generation, their impact in general-purpose roles is yet to be ascertained. Inferences drawn from research until now are nuanced with naive incorporation of human perception abilities, badly harming models and distorting representations. It is also argued whether the model actually matters or whether the results depend upon objective function and training data. Furthermore, labels’ sensitivity and implications make the puzzle more complicated. All these factors further complicate understanding human perceptual abilities regarding vision tasks.

Researchers from MIT and UC Berkeley analyze this question in depth. Their paper “When Does Perceptual Alignment Benefit Vision Representations?” investigates how a human vision perceptual aligned model performs on various downstream visual tasks. The authors finetuned state-of-the-art models ViTs on human similarity judgments for image triplets and evaluated them across standard vision benchmarks. They introduce the idea of a second pretraining stage, which aligns the feature representations from large vision models with human judgments before applying them to downstream tasks. 

YOU MAY ALSO LIKE

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

To understand this further, we first discuss the image triplets mentioned above. The authors used the renowned synthetic NIGHTS dataset with image triplets annotated with forced choice human similarity judgments where humans chose two images with the highest similarity to the first image. They formulate a patch alignment objective function to catch spatial representations present in patch tokens and translate visual attributes from global annotations; instead of computing the loss just between global CLS tokens of Vision Transformer, they focused CLS and pooled patch embeddings of ViT for this purpose to optimize local patch features jointly with the global image label.After this, various state-of-the-art Vision Transformer models, such as DINO, CLIP, etc, were finetuned on the above data using Low-Rank Adaptation (LoRA). The authors also incorporated synthetic images in triplets with SynCLR to compute the performance delta.

These models performed better in vision tasks than the base Vision Transformers. For Dense prediction tasks, human-aligned models outperformed base models in over  75 % of the cases in case of both semantic segmentation and depth estimation. Moving on in the realm of generative vision and LLMs, task of Retrieval-Augmented Generation were checked by humanizing a vision language model. Results again favored prompts retrieved by human-aligned models as they boosted classification accuracy across domains. Further, in the task of object counting, these modified models outperformed the base in more than 95 % of the cases. A similar trend persists in instance retrieval. These models failed on classification tasks due to their high level of semantic understanding.

The authors also addressed whether training data had a more significant role than the training method. For this purpose, more datasets with image triplets were considered. The results were astonishing, with the NIGHTS dataset offering the most considerable impact and the rest barely affected. The perceptual cues captured in NIGHTS play a crucial role in this with its features like style, pose, color, and object count. Others failed due to the inability to capture required mid-level perceptual features.

Overall, human-aligned vision models performed well in most cases. However, these models are prone to overfitting and bias propagation. Thus, if the quality and diversity of human annotation are ensured, visual intelligence could be taken a notch above.


Check out the Paper, GitHub, and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 50k+ ML SubReddit.

[Upcoming Live Webinar- Oct 29, 2024] The Best Platform for Serving Fine-Tuned Models: Predibase Inference Engine (Promoted)


Adeeba Alam Ansari is currently pursuing her Dual Degree at the Indian Institute of Technology (IIT) Kharagpur, earning a B.Tech in Industrial Engineering and an M.Tech in Financial Engineering. With a keen interest in machine learning and artificial intelligence, she is an avid reader and an inquisitive individual. Adeeba firmly believes in the power of technology to empower society and promote welfare through innovative solutions driven by empathy and a deep understanding of real-world challenges.

Listen to our latest AI podcasts and AI research videos here ➡️


Credit: Source link

ShareTweetSendSharePin

Related Posts

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU
AI & Technology

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

September 25, 2026
Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120
AI & Technology

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

September 25, 2026
Warzone Is Adding A Button To Hide All The Goofy Skins
AI & Technology

Warzone Is Adding A Button To Hide All The Goofy Skins

September 24, 2026
How These AI Glasses Compare
AI & Technology

How These AI Glasses Compare

September 24, 2026
Next Post
Reddit Q3 Preview: Licensing The Next Growth Catalyst After Ads (NYSE:RDDT)

Reddit Q3 Preview: Licensing The Next Growth Catalyst After Ads (NYSE:RDDT)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Canada’s trade minister: ‘We’re not waiting by the phone’ as tariff war escalates

Canada’s trade minister: ‘We’re not waiting by the phone’ as tariff war escalates

September 24, 2026
Microsoft, OpenAI workers privately worried their work is ‘largest theft of labor in human history’

Microsoft, OpenAI workers privately worried their work is ‘largest theft of labor in human history’

September 18, 2026
Anthropic quietly sets up biology lab as it ramps AI drug program

Anthropic quietly sets up biology lab as it ramps AI drug program

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!