• bitcoinBitcoin(BTC)$78,789.001.90%
  • ethereumEthereum(ETH)$2,525.730.77%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$723.970.41%
  • rippleXRP(XRP)$1.424.80%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$102.611.59%
  • tronTRON(TRX)$0.3411980.04%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.00%
  • zcashZcash(ZEC)$1,138.042.43%
  • HyperliquidHyperliquid(HYPE)$80.462.46%
  • dogecoinDogecoin(DOGE)$0.0843870.20%
  • RainRain(RAIN)$0.014326-6.25%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$514.99-3.70%
  • whitebitWhiteBIT Coin(WBT)$81.461.67%
  • chainlinkChainlink(LINK)$11.500.88%
  • leo-tokenLEO Token(LEO)$8.99-0.66%
  • cardanoCardano(ADA)$0.2109301.49%
  • stellarStellar(XLM)$0.1952809.00%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$224.99-0.14%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$54.10-1.48%
  • uniswapUniswap(UNI)$6.390.79%
  • CantonCanton(CC)$0.0972501.60%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.06%
  • hedera-hashgraphHedera(HBAR)$0.0775592.03%
  • avalanche-2Avalanche(AVAX)$7.572.18%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • nearNEAR Protocol(NEAR)$2.496.49%
  • shiba-inuShiba Inu(SHIB)$0.0000050.93%
  • suiSui(SUI)$0.731.80%
  • crypto-com-chainCronos(CRO)$0.0589831.19%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,314.07-0.82%
  • BittensorBittensor(TAO)$234.63-0.27%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-4.00%
  • okbOKB(OKB)$114.060.78%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.09%
  • aaveAave(AAVE)$126.980.00%
  • BitwayBitway(BTW)$0.723.90%
  • AsterAster(ASTER)$0.700.55%
  • mantleMantle(MNT)$0.570.61%
  • pax-goldPAX Gold(PAXG)$4,319.18-0.81%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0574610.85%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Panda-70M: A Large-Scale Dataset with 70M High-Quality Video-Caption Pairs

March 6, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Panda-70M: A Large-Scale Dataset with 70M High-Quality Video-Caption Pairs
ShareShareShareShareShare

The significance of computing and data size is undeniable in large-scale multimodal learning. Still, collecting data from high-quality video text is always challenging due to its temporal structure. Vision-language datasets (VLDs) like HD-VILA-100M and HowTo100M are extensively employed across various tasks, including action recognition, video understanding, VQA, and retrieval. These models are annotated by automatic speech recognition (ASR), enabling them to transcribe spoken language into text and enhance their performance. 

However, meta-information like subtitles, video descriptions, and voice-overs frequently need more precision due to being overly broad, misaligned in timing, or insufficiently descriptive of the video content. If we are using VLDs like HD-VILA-100M and HowTo100M, the subtitles usually fail to include the main content and action presented in the video. This limitation reduces the effectiveness of these datasets for multimodal training. Some datasets suffer from low resolution, ASR annotations, small-scale samples, or brief captions.

Researchers from Snap Inc., the University of California, Merced, and the University of Trento proposed Panda-70M to tackle the challenge of generating high-quality video captions. Panda-70M leverages multimodal inputs, such as textual video description, subtitles, and individual video frames, to establish a video dataset with high-quality captions by curating 3.8M high-resolution videos from publicly available HD-VILA-100M datasets. Also, to use several captioning, Panda-70M incorporates five base models – Video LLaMA, VideoChat, VideoChat Text, BLIP-2, and MiniGPT-4 as teacher models. Researchers follow established protocols and evaluate metrics like BLEU-4, ROGUE-L, METEOR, and CIDEr.

To develop Panda-70M, meticulously crafted 3.8M high-resolution videos are split into semantically consistent video clips, leveraging multiple cross-modality teacher models. Then, a fine-grained retrieval model is fine-tuned to obtain captions for each video. The models trained on the proposed data score consistently outperform others on most metrics, demonstrating their superiority in various tasks. Compared to other VLDs, Panda-70M took 8.5 seconds on average to process a video of length 166.8Khr, which is the most efficient of all VLDs.  

In visual processing, we utilize a familiar framework, similar to the one employed in Video LLaMA, to derive video representations compatible with LLM. For the text branch, a direct approach would be to input text embedding directly into the LLM. However, this poses two challenges:

  • Lengthy text prompts containing video descriptions and subtitles may overwhelm the LLM’s decision-making process, resulting in computationally intensive burdens.
  • The information derived from descriptions and subtitles can often be noisy and irrelevant to the video content.
  • Researchers also introduced text Q-former, designed to extract fixed-length text representations, thereby fostering a more seamless connection between video and text representations. This Q-former mirrors the architecture of the Query Transformer found in BLIP-2.

In conclusion, this research introduces a large-scale video dataset with caption annotations called Panda-70M. An automatic pipeline that can leverage multimodal information has been proposed to caption 70M videos and facilitate three downstream tasks: video captioning, video and text retrieval, and text-to-video generation. Despite impressive results, Panda-70M limits the content diversity within a single video and reduces average video duration; thus, it fails to build datasets with long videos. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 38k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel

You may also like our FREE AI Courses….


YOU MAY ALSO LIKE

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI
AI & Technology

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

September 14, 2026
How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error
AI & Technology

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

September 14, 2026
Temporal Raises 0M Series E at .55B Valuation to Expand Operations – Unite.AI
AI & Technology

Temporal Raises $550M Series E at $12.55B Valuation to Expand Operations – Unite.AI

September 14, 2026
What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?
AI & Technology

What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?

September 14, 2026
Next Post
California investigates whether migrant flight came from Florida

California investigates whether migrant flight came from Florida

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Hillary Clinton: We’re still learning lessons from 9/11

Hillary Clinton: We’re still learning lessons from 9/11

September 13, 2026
DEA seizes 2,000 pounds of meth hidden in cabbages

DEA seizes 2,000 pounds of meth hidden in cabbages

September 13, 2026
A Native Gemini App Is Finally Available For Windows PCs

A Native Gemini App Is Finally Available For Windows PCs

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!