• bitcoinBitcoin(BTC)$83,677.00-0.85%
  • ethereumEthereum(ETH)$2,686.230.03%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$766.81-1.13%
  • rippleXRP(XRP)$1.50-1.18%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$119.21-2.11%
  • tronTRON(TRX)$0.3347490.31%
  • zcashZcash(ZEC)$1,527.64-3.41%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$88.78-2.59%
  • dogecoinDogecoin(DOGE)$0.094306-2.41%
  • chainlinkChainlink(LINK)$15.208.01%
  • moneroMonero(XMR)$539.43-1.01%
  • whitebitWhiteBIT Coin(WBT)$83.63-0.69%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.246164-3.01%
  • RainRain(RAIN)$0.012551-0.12%
  • leo-tokenLEO Token(LEO)$9.060.55%
  • stellarStellar(XLM)$0.2247224.47%
  • nearNEAR Protocol(NEAR)$4.94-4.31%
  • bitcoin-cashBitcoin Cash(BCH)$310.02-7.16%
  • hedera-hashgraphHedera(HBAR)$0.12721536.23%
  • uniswapUniswap(UNI)$8.90-8.06%
  • litecoinLitecoin(LTC)$69.40-2.59%
  • CantonCanton(CC)$0.130540-2.19%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.44-4.45%
  • suiSui(SUI)$1.16-6.29%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.65-0.84%
  • daiDai(DAI)$1.000.03%
  • USD1USD1(USD1)$1.000.00%
  • quant-networkQuant(QNT)$242.3932.76%
  • crypto-com-chainCronos(CRO)$0.0690152.69%
  • BittensorBittensor(TAO)$302.41-6.43%
  • shiba-inuShiba Inu(SHIB)$0.000006-3.78%
  • Global DollarGlobal Dollar(USDG)$1.000.03%
  • tether-goldTether Gold(XAUT)$4,133.95-3.40%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.16-1.50%
  • EthenaEthena(ENA)$0.259940-5.66%
  • BitwayBitway(BTW)$0.95-19.84%
  • OndoOndo(ONDO)$0.52-4.63%
  • Pump.funPump.fun(PUMP)$0.00538810.77%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$118.14-2.44%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.07%
  • aaveAave(AAVE)$147.57-4.41%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Hugging Face Open-Sourced FineVision: A New Multimodal Dataset with 24 Million Samples for Training Vision-Language Models (VLMs)

September 6, 2025
in AI & Technology
Reading Time: 6 mins read
A A
Hugging Face Open-Sourced FineVision: A New Multimodal Dataset with 24 Million Samples for Training Vision-Language Models (VLMs)
ShareShareShareShareShare

Hugging Face has just released FineVision, an open multimodal dataset designed to set a new standard for Vision-Language Models (VLMs). With 17.3 million images, 24.3 million samples, 88.9 million question-answer turns, and nearly 10 billion answer tokens, FineVision position itself as one of the largest and structured publicly available VLM training datasets.

FineVision aggregates 200+ sources into a unified format, rigorously filtered for duplicates and benchmark contamination. Rated systematically across multiple quality dimensions, the dataset enables researchers and devs to construct robust training mixtures while minimizing data leakage.

YOU MAY ALSO LIKE

SpaceX Starship Lifts Off for Key Orbital Test

The Safest And Easiest Way To Debloat Windows 11

Why is FineVision Important for VLM Training?

Most state-of-the-art VLMs rely on proprietary datasets, limiting reproducibility and accessibility for the broader research community. FineVision addresses this gap by:

  • Scale and Coverage: 5 TB of curated data across 9 categories, including General VQA, OCR QA, Chart & Table reasoning, Science, Captioning, Grounding & Counting, and GUI navigation.
  • Benchmark Gains: Across 11 widely used benchmarks (e.g., AI2D, ChartQA, DocVQA, ScienceQA, OCRBench), models trained on FineVision outperform alternatives by significant margins—up to 46.3% over LLaVA, 40.7% over Cauldron, and 12.1% over Cambrian.
  • New Skill Domains: FineVision introduces data for emerging tasks like GUI navigation, pointing, and counting, expanding the capabilities of VLMs beyond conventional captioning and VQA.

How Was FineVision Built?

The curation pipeline followed a three-step process:

  1. Collection and Augmentation
    Over 200 publicly available image-text datasets were gathered. Missing modalities (e.g., text-only data) were reformatted into QA pairs. Underrepresented domains, such as GUI data, were supplemented through targeted collection.
  2. Cleaning
    • Removed oversized QA pairs (>8192 tokens).
    • Resized large images to a maximum of 2048 px while preserving aspect ratio.
    • Discarded corrupted samples.
  3. Quality Rating
    Using Qwen3-32B and Qwen2.5-VL-32B-Instruct as judges, every QA pair was rated on four axes:
    • Text Formatting Quality
    • Question-Answer Relevance
    • Visual Dependency
    • Image-Question Correspondence

    These ratings enable selective training mixtures, though ablations show that retaining all samples yields the best performance, even when lower-rated samples are included.

Comparative Analysis: FineVision vs. Existing Open Datasets

Dataset Images Samples Turns Tokens Leakage Perf. Drop After Deduplication
Cauldron 2.0M 1.8M 27.8M 0.3B 3.05% -2.39%
LLaVA-Vision 2.5M 3.9M 9.1M 1.0B 2.15% -2.72%
Cambrian-7M 5.4M 7.0M 12.2M 0.8B 2.29% -2.78%
FineVision 17.3M 24.3M 88.9M 9.5B 1.02% -1.45%

FineVision is not only one of the largest but also the least hallucinated dataset, with just 1% overlap with benchmark test sets. This ensures minimal data leakage and reliable evaluation performance.

Performance Insights

  • Model Setup: Ablations were conducted using nanoVLM (460M parameters), combining SmolLM2-360M-Instruct as the language backbone and SigLIP2-Base-512 as the vision encoder.
  • Training Efficiency: On 32 NVIDIA H100 GPUs, one full epoch (12k steps) takes ~20 hours.
  • Performance Trends:
    • FineVision models improve steadily with exposure to diverse data, overtaking baselines after ~12k steps.
    • Deduplication experiments confirm FineVision’s low leakage compared to Cauldron, LLaVA, and Cambrian.
    • Multilingual subsets, even when the backbone is monolingual, show slight performance gains, suggesting diversity outweighs strict alignment.
    • Attempts at multi-stage training (two or 2.5 stages) did not yield consistent benefits, reinforcing that scale + diversity is more critical than training heuristics.

Why FineVision Brings the New Standard?

  1. +20% Average Performance Boost: Outperforms all existing open datasets across 10+ benchmarks.
  2. Unprecedented Scale: 17M+ images, 24M+ samples, 10B tokens.
  3. Skill Expansion: GUI navigation, counting, pointing, and document reasoning included.
  4. Lowest Data Leakage: 1% contamination, compared to 2–3% in other datasets.
  5. Fully Open Source: Available on Hugging Face Hub for immediate use via the datasets library.

Conclusion

FineVision marks a significant advancement in open multimodal datasets. Its large scale, systematic curation, and transparent quality assessments create a reproducible and extensible foundation for training state-of-the-art Vision-Language Models. By reducing dependence on proprietary resources, it enables researchers and devs to build competitive systems and accelerate progress in areas such as document analysis, visual reasoning, and agentic multimodal tasks.


Check out the Dataset and Technical details. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceX Starship Lifts Off for Key Orbital Test
AI & Technology

SpaceX Starship Lifts Off for Key Orbital Test

September 28, 2026
The Safest And Easiest Way To Debloat Windows 11
AI & Technology

The Safest And Easiest Way To Debloat Windows 11

September 28, 2026
How To Improve Your Samsung Galaxy’s Battery Performance
AI & Technology

How To Improve Your Samsung Galaxy’s Battery Performance

September 28, 2026
A Modular, Repairable GPS Watch Is A Good First Step
AI & Technology

A Modular, Repairable GPS Watch Is A Good First Step

September 28, 2026
Next Post
Towing failure sends car crashing into restaurant twice

Towing failure sends car crashing into restaurant twice

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NBC Nightly News with Tom Llamas Full Episode – Aug. 19

NBC Nightly News with Tom Llamas Full Episode – Aug. 19

September 27, 2026
Brendan Hunt on the future of ‘Ted Lasso’ and how a rejection led him to get Coach Beard

Brendan Hunt on the future of ‘Ted Lasso’ and how a rejection led him to get Coach Beard

September 24, 2026
U.S. debt balloons to record-breaking  trillion

U.S. debt balloons to record-breaking $40 trillion

September 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!