• bitcoinBitcoin(BTC)$78,056.00-1.43%
  • ethereumEthereum(ETH)$2,462.10-1.06%
  • tetherTether(USDT)$1.00-0.03%
  • binancecoinBNB(BNB)$747.870.44%
  • rippleXRP(XRP)$1.400.35%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.79-1.77%
  • tronTRON(TRX)$0.3388171.09%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,161.03-1.13%
  • HyperliquidHyperliquid(HYPE)$82.70-5.30%
  • dogecoinDogecoin(DOGE)$0.089256-1.43%
  • RainRain(RAIN)$0.0165981.00%
  • USDSUSDS(USDS)$1.00-0.02%
  • moneroMonero(XMR)$501.89-6.26%
  • whitebitWhiteBIT Coin(WBT)$79.358.65%
  • chainlinkChainlink(LINK)$12.44-4.37%
  • leo-tokenLEO Token(LEO)$9.190.47%
  • cardanoCardano(ADA)$0.219223-0.87%
  • stellarStellar(XLM)$0.188027-2.17%
  • bitcoin-cashBitcoin Cash(BCH)$254.60-2.13%
  • daiDai(DAI)$1.00-0.02%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • uniswapUniswap(UNI)$6.91-2.20%
  • USD1USD1(USD1)$1.00-0.03%
  • litecoinLitecoin(LTC)$54.68-4.25%
  • CantonCanton(CC)$0.103902-3.77%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-1.78%
  • hedera-hashgraphHedera(HBAR)$0.079894-3.02%
  • avalanche-2Avalanche(AVAX)$7.94-1.67%
  • suiSui(SUI)$0.80-2.66%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.60%
  • nearNEAR Protocol(NEAR)$2.30-2.82%
  • crypto-com-chainCronos(CRO)$0.0601054.17%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,393.58-0.24%
  • MemeCoreMemeCore(M)$1.174.57%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BittensorBittensor(TAO)$251.94-4.32%
  • okbOKB(OKB)$114.21-1.36%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.11%
  • mantleMantle(MNT)$0.63-1.76%
  • AsterAster(ASTER)$0.76-4.96%
  • aaveAave(AAVE)$128.13-3.53%
  • pax-goldPAX Gold(PAXG)$4,398.48-0.19%
  • polkadotPolkadot(DOT)$1.080.39%
  • OndoOndo(ONDO)$0.373882-3.21%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Do Video-Language Models Understand Actions? If Not, How To Fix It? Meet Paxion: A Novel Framework For Patching Action Knowledge in Video-Language Foundation Models

June 8, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Do Video-Language Models Understand Actions? If Not, How To Fix It? Meet Paxion: A Novel Framework For Patching Action Knowledge in Video-Language Foundation Models
ShareShareShareShareShare

Recent video-language models’ (VidLMs) performance on various video-language tasks has been outstanding. Such multimodal models only come with drawbacks. For example, it is shown that vision-language models have difficulty understanding compositional and order relations in images, treating images as collections of objects, and that many popular video-language benchmarks can be solved by looking at a single frame. Such restrictions imply that models’ awareness of object connections and understanding of actions, which may need many structures, may need to be improved. To test this hypothesis, they begin by defining action knowledge as a comprehension of the cause and consequence of actions in textual, visual, and temporal dimensions. 

Researchers from UIUC and UNC introduce the Action Dynamics Benchmark (ActionBench) to measure a model’s action understanding. ActionBench includes two challenging tasks: identifying (1) the original and reversed movies and (2) the video caption with the action verbs substituted by their antonyms. A baseline task for minimizing the negative effects of domain mismatch and examining potential bias in favor of objects is also included in the benchmark. The baseline challenge is for the model to distinguish between the original video subtitles and edited versions with arbitrary item replacements. 

Modern video-language foundation models perform nearly randomly on action-oriented probing tasks, but very well on object-oriented baseline tests. This demonstrates the need for action knowledge in VidLMs. Their remarkable performance on other benchmarks may be due to their object identification skills rather than their grasp of actions. They offer a unique framework called PAXION (Patching Actions) to patch current VidLMs with action knowledge while maintaining their general vision-language (VL) capabilities to remedy this weakness. The Knowledge Patcher and the Knowledge Fuser are PAXION’s two major parts. 

🚀 JOIN the fastest ML Subreddit Community

They found that the widely-used Video-Text Contrastive (VTC) aim needs to be revised, supporting earlier studies’ findings. This poses a significant barrier to patching action knowledge. To add action-aware representations to the VidLM, the Knowledge Patcher (KP), a Perceiver-based lightweight module coupled to a frozen VidLM backbone, is employed. The Discriminative Video Dynamics Modelling (DVDM) objective forces the model to learn the correlation between an action’s textual signifier, the action text (for example, the word “falling”), and the action’s visual depiction (for example, a clip of a falling book), is thus introduced. It is inspired by dynamics modeling in robotics and reinforcement learning. 

Video-Action Contrastive (VAC) and Action-Temporal Matching (ATM), two new features in DVDM, are compatible with VTC without requiring different settings. They develop discriminative tasks employing action antonyms and reversed films, focusing on learning from examples of data with major state transitions. They show that their ActionBench tasks significantly improve thanks to the interaction between the Knowledge Patcher and DVDM. They next look at how their Knowledge Patcher, which focuses on action understanding, might be included in already-existing VidLMs for jobs that need both action and object knowledge downstream. 

To do this, they offer the Knowledge Fuser (KF) component of PAXION, which utilizes cross-attention to fuse the object-centric representation from the firm backbone with the action-centric representation from the Knowledge Patcher. They demonstrate that on a variety of tasks, such as Video-Text Retrieval (SSv2-label), Video-to-Action Retrieval (SSv2-template, Temporal), and Causal-Temporal Video Question Answering (NExT-QA), the fused representation from PAXION increases both object and action knowledge. Furthermore, their research demonstrates that the Knowledge Fuser is crucial for preserving a balance between the models’ object-related comprehension and enhancing performance on downstream action and temporal-oriented tasks. 

By taking into account a zero-shot cross-domain transfer setting on the Moments-in-Time and Kinetics datasets, they additionally assess PAXION’s resilience. They discover that further assembling PAXION with the backbone model can positively transfer to new domains while boosting strength to domain changes. This is the first study to rigorously analyze action knowledge and incorporate it into video-language foundation models to the best of their ability. 

Three things are their primary contributions: 

1. They provide the Action Dynamics Benchmark, which tests the ability of video-language models to recognize actions. After analyzing three cutting-edge video-language foundation models, they need a fundamental understanding of action knowledge. 

2. They put forth the unique learning framework PAXION, which adds the missing action knowledge to foundation models of frozen video language without impairing those models’ overall vision-language skills. A Perceiver-based Knowledge Patcher and a cross-attention-based Knowledge Fuser are two of PAXION’s main building blocks. 

3. They suggest the DVDM goal, which pushes the model to encode the relationship between the action text and the proper sequencing of video frames, as an improvement over the often-used VTC loss. Numerous investigations demonstrate that PAXION with DVDM enhances the mutual comprehension of things and activities while being resilient to domain shift.


Check Out The Paper and Code. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


Check out https://aitoolsclub.com to find 100’s of Cool AI Tools

Credit: Source link

ShareTweetSendSharePin

Related Posts

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?
AI & Technology

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

September 8, 2026
What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI
AI & Technology

What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI

September 8, 2026
Motional Releases nuReasoning Dataset and Launches ECCV Challenge – Unite.AI
AI & Technology

Motional Releases nuReasoning Dataset and Launches ECCV Challenge – Unite.AI

September 8, 2026
Renault Is Building Its €17,900 Dacia Spring EV In Europe To Qualify For Local Subsidies
AI & Technology

Renault Is Building Its €17,900 Dacia Spring EV In Europe To Qualify For Local Subsidies

September 8, 2026
Next Post
Here’s Why Jim Cramer Is Focused on the Fed Next Week as June Policy Meeting Looms

Here's Why Jim Cramer Is Focused on the Fed Next Week as June Policy Meeting Looms

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Federal investigators probe Amazon cargo jet’s fiery runway crash that killed 5 in Miami – AP News

Federal investigators probe Amazon cargo jet’s fiery runway crash that killed 5 in Miami – AP News

September 7, 2026
Shark sightings on the rise this summer

Shark sightings on the rise this summer

September 4, 2026
Dyson Unveils AI-Powered Toothbrush With Built-In Camera

Dyson Unveils AI-Powered Toothbrush With Built-In Camera

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!