• bitcoinBitcoin(BTC)$76,342.000.65%
  • ethereumEthereum(ETH)$2,437.771.44%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$725.001.51%
  • rippleXRP(XRP)$1.300.21%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.522.64%
  • tronTRON(TRX)$0.3352860.26%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.65%
  • zcashZcash(ZEC)$1,361.3818.07%
  • HyperliquidHyperliquid(HYPE)$78.882.23%
  • dogecoinDogecoin(DOGE)$0.0808061.11%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$501.91-0.53%
  • whitebitWhiteBIT Coin(WBT)$78.520.76%
  • RainRain(RAIN)$0.012875-8.07%
  • chainlinkChainlink(LINK)$11.122.33%
  • leo-tokenLEO Token(LEO)$8.930.72%
  • cardanoCardano(ADA)$0.1962060.33%
  • stellarStellar(XLM)$0.1822793.75%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.00-0.02%
  • bitcoin-cashBitcoin Cash(BCH)$220.880.62%
  • USD1USD1(USD1)$1.00-0.02%
  • uniswapUniswap(UNI)$6.766.34%
  • litecoinLitecoin(LTC)$51.971.72%
  • CantonCanton(CC)$0.0979867.06%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-0.14%
  • nearNEAR Protocol(NEAR)$2.6613.69%
  • avalanche-2Avalanche(AVAX)$7.512.51%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.073595-1.06%
  • suiSui(SUI)$0.724.40%
  • shiba-inuShiba Inu(SHIB)$0.0000051.54%
  • crypto-com-chainCronos(CRO)$0.0584315.47%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,298.92-0.64%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$223.933.02%
  • MemeCoreMemeCore(M)$1.110.60%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$111.510.59%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.02%
  • AsterAster(ASTER)$0.737.13%
  • BitwayBitway(BTW)$0.72-6.68%
  • aaveAave(AAVE)$122.251.18%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0589203.37%
  • pax-goldPAX Gold(PAXG)$4,300.63-0.70%
  • mantleMantle(MNT)$0.562.53%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA AI Research Proposes Language Instructed Temporal-Localization Assistant (LITA), which Enables Accurate Temporal Localization Using Video LLMs

March 31, 2024
in AI & Technology
Reading Time: 5 mins read
A A
NVIDIA AI Research Proposes Language Instructed Temporal-Localization Assistant (LITA), which Enables Accurate Temporal Localization Using Video LLMs
ShareShareShareShareShare

Large Language Models (LLMs) have proven their impressive instruction-following capabilities, and they can be a universal interface for various tasks such as text generation, language translation, etc. These models can be extended to multimodal LLMs to process language and other modalities, such as Image, video, and audio. Several recent works introduce models that specialize in processing videos. These Video LLMs preserve the instruction following capabilities of LLMs and allow users to ask various questions about a given video. However, one important missing piece in these Video LLMs is temporal localization. When prompted with the “When?” questions, these models cannot accurately localize periods and often hallucinate irrelevant information.

Three key aspects limit the temporal localization capabilities of existing Video LLMs: time representation, architecture, and data. First, existing models often represent timestamps as plain text (e.g., 01:22 or 142sec). However, given a set of frames, the correct timestamp still depends on the frame rate, which the model cannot access. This makes learning temporal localization harder. Second, the architecture of existing Video LLMs might need more temporal resolution to interpolate time information accurately. For example, Video-LLaMA only uniformly samples eight frames from the entire video, which needs to be revised for accurate temporal localization. Finally, temporal localization is largely ignored in the data used by existing Video LLMs. Data with timestamps are only a small subset of video instruction tuning data, and the accuracy of these timestamps is also not verified.

NVIDIA researchers propose Language Instructed Temporal-Localization Assistant (LITA). The three key components they have proposed are: (1) Time Representation: time tokens to represent relative timestamps and allow Video LLMs to better communicate about time than using plain text. (2) Architecture: They introduced SlowFast tokens to capture temporal information at fine temporal resolution to enable accurate temporal localization. (3) Data: They have emphasized temporal localization data for LITA. They have proposed a new task, Reasoning Temporal Localization (RTL), along with the dataset ActivityNet-RTL, to learn this task.

LITA is built on Image LLaVA due to its simplicity and effectiveness. LITA does not depend on the underlying Image LLM architecture and can be easily adapted to other base architectures. Given a video, they first uniformly select T frames and encode each frame into M tokens. T × M is a large number that usually cannot be directly processed by the LLM module. Thus, they use SlowFast pooling to reduce the T × M tokens to T + M tokens. The text tokens (prompt) are processed to convert referenced timestamps to specialized time tokens. All the input tokens are then jointly processed by the LLM module sequentially.  The model is fine-tuned with RTL data and other video tasks, such as dense video captioning and event localization. LITA learns to use time tokens instead of absolute timestamps. 

Comparing LITA with LLaMA-Adapter, Video-LLaMA, VideoChat, and Video-ChatGPT. Video-ChatGPT slightly outperforms other baselines, including VideoLLaMA-v2. LITA significantly outperforms these two existing Video LLMs in all aspects. In particular, LITA achieves a 22% improvement in the Correctness of Information (2.94 vs. 2.40) and a 36% relative improvement in Temporal Understanding (2.68 vs. 1.98). This shows that the emphasis on temporal understanding in training enables accurate temporal localization and improves LITA’s video understanding.

In conclusion, NVIDIA researchers present LITA, a game-changer in temporal localization using Video LLMs. With its unique model design, LITA introduces time tokens and SlowFast tokens, significantly improving the representation of time and the processing of video inputs. LITA demonstrates promising capabilities to answer complex temporal localization questions and substantially enhances video-based text generation compared to existing Video LLMs, even for non-temporal questions. 


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 39k+ ML SubReddit


YOU MAY ALSO LIKE

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
AI & Technology

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

September 17, 2026
House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI
AI & Technology

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

September 16, 2026
Snap Introduces A Standalone AI Assistant, Specs Intelligence
AI & Technology

Snap Introduces A Standalone AI Assistant, Specs Intelligence

September 16, 2026
Standalone AR Glasses Are Here
AI & Technology

Standalone AR Glasses Are Here

September 16, 2026
Next Post
Qualcomm: Fundamentals Deteriorate With Intensifying Competition (NASDAQ:QCOM)

Qualcomm: Fundamentals Deteriorate With Intensifying Competition (NASDAQ:QCOM)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

September 14, 2026
US consumer prices accelerate in August, push Fed closer to rate hike – Reuters

US consumer prices accelerate in August, push Fed closer to rate hike – Reuters

September 11, 2026
Baltimore Ravens vs. Indianapolis Colts Live Score and Stats – September 13, 2026 Gametracker – CBS Sports

Baltimore Ravens vs. Indianapolis Colts Live Score and Stats – September 13, 2026 Gametracker – CBS Sports

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!