• bitcoinBitcoin(BTC)$84,541.000.11%
  • ethereumEthereum(ETH)$2,693.660.69%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$780.261.68%
  • rippleXRP(XRP)$1.531.94%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$117.422.48%
  • tronTRON(TRX)$0.3405880.05%
  • zcashZcash(ZEC)$1,561.222.84%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.58%
  • HyperliquidHyperliquid(HYPE)$94.440.84%
  • dogecoinDogecoin(DOGE)$0.0962924.16%
  • moneroMonero(XMR)$548.13-0.41%
  • whitebitWhiteBIT Coin(WBT)$84.68-0.08%
  • chainlinkChainlink(LINK)$13.227.78%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2480693.94%
  • RainRain(RAIN)$0.012050-1.86%
  • leo-tokenLEO Token(LEO)$8.91-1.14%
  • stellarStellar(XLM)$0.2131874.96%
  • bitcoin-cashBitcoin Cash(BCH)$337.62-2.95%
  • nearNEAR Protocol(NEAR)$4.748.50%
  • uniswapUniswap(UNI)$9.301.21%
  • litecoinLitecoin(LTC)$70.9115.94%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • avalanche-2Avalanche(AVAX)$10.471.57%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.1147035.62%
  • USD1USD1(USD1)$1.00-0.01%
  • suiSui(SUI)$1.014.75%
  • hedera-hashgraphHedera(HBAR)$0.0926202.61%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.431.26%
  • shiba-inuShiba Inu(SHIB)$0.0000064.08%
  • BittensorBittensor(TAO)$300.023.87%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0634123.33%
  • BitwayBitway(BTW)$1.066.01%
  • MemeCoreMemeCore(M)$1.221.27%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,269.61-0.37%
  • okbOKB(OKB)$119.651.44%
  • OndoOndo(ONDO)$0.5124.22%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.01%
  • aaveAave(AAVE)$146.795.50%
  • mantleMantle(MNT)$0.694.46%
  • EthenaEthena(ENA)$0.2192816.07%
  • polkadotPolkadot(DOT)$1.176.77%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

TWLV-I: A New Video Foundation Model that Constructs Robust Visual Representations for both Motion and Appearance-based Videos

August 25, 2024
in AI & Technology
Reading Time: 4 mins read
A A
TWLV-I: A New Video Foundation Model that Constructs Robust Visual Representations for both Motion and Appearance-based Videos
ShareShareShareShareShare

Language Foundation Models (LFMs) and Large Language Models (LLMs) have demonstrated their ability to handle multiple tasks efficiently with a single fixed model. This achievement has motivated the development of Image Foundation Models (IFMs) in computer vision, which aim to encode general information from images into embedding vectors. However, using these techniques poses a challenge in video analysis. One approach involves treating videos as a sequence of images, where each frame is sampled and embedded before combining; however, this approach faces challenges in capturing detailed motion and small changes between frames. It becomes difficult to understand the continuous flow of information in videos, especially when it comes to tracking object movement and minor frame-to-frame differences

The existing works tried to overcome these challenges using two main approaches based on the Vision Transformer architecture (ViT). The first approach uses distillation with high-performance IFMs like CLIP as teachers, and the second approach is based on masked modeling, where the model predicts missing information from partial input. However, both approaches have their limitations. Distillation-based methods, like UMT and InternVideo2, struggle with motion-sensitive benchmarks like Something-Something-v2 and Diving-48. The masked modeling-based methods, like V-JEPA, perform badly on appearance-centric benchmarks like Kinetics-400 and Moments-in-Time. These limitations highlight the difficulty in capturing the appearance of objects and their motion in videos.

YOU MAY ALSO LIKE

Congressman Calls for National Data Center Strategy

Trump-Xi Summit Puts Global AI Race in Focus

A team from Twelve Labs has proposed TWLV-I, a new model designed to provide embedding vectors for videos that capture appearance and motion. Even though trained only on publicly available datasets, TWLV-I shows strong performance on appearance and motion-focused action recognition benchmarks. Moreover, the model achieves state-of-the-art performance in video-centric tasks such as temporal and spatiotemporal action localization, as well as temporal action segmentation. The current evaluation methods are enhanced to analyze the TWLV-I and other Video Foundation Models (VFMs), with a new analytical approach and a technique to find the model’s ability to differentiate videos based on motion direction, independent of appearance.

TWLV-I adopts ViT architecture, available in Base with 86M parameters and Large with 307M parameters versions. The model tokenizes input videos into patches, processes them through the transformer, and pools the resulting patch-wise embeddings to obtain the overall video embedding. Moreover, the pretraining dataset contains Kinetics-710, HowTo360K, WebVid10M, and various image datasets. The training objective of TWLV-I integrates strengths from distillation-based and masked modeling-based approaches using different reconstruction target strategies. The model utilizes two frame sampling methods, (a) Uniform Embedding for shorter videos and (b) Multi-Clip Embedding for longer videos, to overcome computational constraints.

The results obtained on TWLV-I show significant performance enhancement over existing models in action recognition tasks. Based on the average top-1 accuracy of linear probing across five action recognition benchmarks and using only publicly available datasets for pretraining, TWLV-I outperforms VJEPA (ViT-L) by 4.6% points and UMT (ViT-L) by 7.7% points. This model outperforms larger models like DFN (ViT-H) by 7.2% points, V-JEPA (ViT-H) by 2.7% points, and InternVideo2 (ViT-g) by 2.8% points. Researchers also provided embedding vectors generated by TWLV-I from widely used video benchmarks and evaluation source code that can directly utilize these embeddings.

A team from Twelve Labs has proposed TWLV-I, a novel model designed to provide embedding vectors for videos that capture appearance and motion. TWLV-I proves a strong video foundation model that shows great performance in understanding motion and appearance. The TWLV-I model and its embeddings are expected to be used widely in various applications. Moreover, the evaluation and analysis methods will be actively adopted in the video foundation model domain. In the future, these methods are expected to guide research in the video understanding field, making further progress in developing more comprehensive video analysis models.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 49k+ ML SubReddit

Find Upcoming AI Webinars here


Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Congressman Calls for National Data Center Strategy
AI & Technology

Congressman Calls for National Data Center Strategy

September 24, 2026
Trump-Xi Summit Puts Global AI Race in Focus
AI & Technology

Trump-Xi Summit Puts Global AI Race in Focus

September 24, 2026
Google Takes on Apple, Microsoft With AI-Powered Laptops
AI & Technology

Google Takes on Apple, Microsoft With AI-Powered Laptops

September 24, 2026
The Global AI Race: Chips, Talent, and World Models
AI & Technology

The Global AI Race: Chips, Talent, and World Models

September 24, 2026
Next Post
Two critical tax moves to make before the end of 2024

Two critical tax moves to make before the end of 2024

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Clancy defense attorney reacts to mistrial denial

Clancy defense attorney reacts to mistrial denial

September 24, 2026
Judge to decide whether accused Charlie Kirk assassin will stand trial

Judge to decide whether accused Charlie Kirk assassin will stand trial

September 19, 2026
Live Updates: Mamdani and Trump Discuss Immigration and Affordability at Gracie Mansion – The New York Times

Live Updates: Mamdani and Trump Discuss Immigration and Affordability at Gracie Mansion – The New York Times

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!