• bitcoinBitcoin(BTC)$84,075.00-2.61%
  • ethereumEthereum(ETH)$2,663.97-2.80%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$765.27-2.97%
  • rippleXRP(XRP)$1.50-3.50%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$114.57-2.36%
  • tronTRON(TRX)$0.338843-0.68%
  • zcashZcash(ZEC)$1,552.33-0.30%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.20%
  • HyperliquidHyperliquid(HYPE)$93.39-2.23%
  • dogecoinDogecoin(DOGE)$0.092686-6.62%
  • moneroMonero(XMR)$549.11-3.95%
  • whitebitWhiteBIT Coin(WBT)$84.38-2.64%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.28-5.23%
  • cardanoCardano(ADA)$0.238978-3.93%
  • RainRain(RAIN)$0.012452-7.19%
  • leo-tokenLEO Token(LEO)$8.970.00%
  • stellarStellar(XLM)$0.203705-4.39%
  • bitcoin-cashBitcoin Cash(BCH)$349.997.82%
  • uniswapUniswap(UNI)$9.15-0.44%
  • nearNEAR Protocol(NEAR)$4.28-3.39%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$60.30-2.14%
  • daiDai(DAI)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$10.32-6.95%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.108126-4.05%
  • hedera-hashgraphHedera(HBAR)$0.090475-5.16%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-2.01%
  • suiSui(SUI)$0.96-4.03%
  • BittensorBittensor(TAO)$293.30-5.93%
  • shiba-inuShiba Inu(SHIB)$0.000006-5.75%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.061951-6.86%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.21-7.02%
  • BitwayBitway(BTW)$0.9916.23%
  • tether-goldTether Gold(XAUT)$4,287.45-1.11%
  • okbOKB(OKB)$118.60-2.53%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.04%
  • aaveAave(AAVE)$140.12-1.96%
  • mantleMantle(MNT)$0.65-0.81%
  • EthenaEthena(ENA)$0.2067420.35%
  • OndoOndo(ONDO)$0.417501-2.58%
  • Pump.funPump.fun(PUMP)$0.004035-9.79%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MINT-1T Dataset Released: A Multimodal Dataset with One Trillion Tokens to Build Large Multimodal Models

July 26, 2024
in AI & Technology
Reading Time: 6 mins read
A A
MINT-1T Dataset Released: A Multimodal Dataset with One Trillion Tokens to Build Large Multimodal Models
ShareShareShareShareShare

Artificial intelligence, particularly in training large multimodal models (LMMs), relies heavily on vast datasets that include sequences of images and text. These datasets enable the development of sophisticated models capable of understanding and generating multimodal content. As AI models’ capabilities advance, the need for extensive, high-quality datasets becomes even more critical, driving researchers to explore new data collection and curation methods.

A significant challenge in AI research is the need for large-scale, open-source, multimodal interleaved datasets. These datasets are essential for training models seamlessly integrating text and image data. The limited availability of such datasets hampers the development of robust and high-performing open-source models, resulting in a performance gap between open-source and proprietary models. Addressing this gap requires innovative approaches to dataset creation that can provide the necessary scale and diversity.

YOU MAY ALSO LIKE

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

Never Use ChatGPT For These Five Tasks

Existing methods for creating multimodal datasets often involve collecting and curating data from HTML documents. Notable datasets like OBELICS have been instrumental but are limited in scale and diversity, primarily sourcing data from HTML. This restriction affects the variety and richness of the data, impacting the performance and applicability of the resulting AI models. Researchers have found that datasets sourced solely from HTML documents must capture the full spectrum of required multimodal content for comprehensive model training.

Researchers from the University of Washington, Salesforce Research, Stanford University, the University of Texas at Austin, and the University of California Berkeley introduced MINT-1T, the most extensive & diverse open-source multimodal interleaved dataset to date, addressing the need for larger and more varied datasets. MINT-1T comprises one trillion text tokens and 3.4 billion images from HTML, PDFs, and ArXiv papers. This dataset represents a tenfold increase from previous datasets, significantly enhancing the data for training multimodal models. Institutions such as the University of Washington and Salesforce Research collaborated on this initiative, demonstrating a concerted effort to bridge the gap in dataset availability.

Creating the MINT-1T dataset involved an intricate process of sourcing, filtering, and deduplicating data. HTML documents were expanded to include data from earlier years, and PDFs were processed to extract readable text and images. ArXiv papers were parsed for figures and text, ensuring a comprehensive collection of multimodal content. Advanced filtering methods were employed to remove low-quality, non-English, and inappropriate content. Deduplication processes were also implemented to eliminate repetitive data, ensuring the dataset’s quality and diversity.

Experiments demonstrated that LMMs trained on the MINT-1T dataset matched and often surpassed the performance of models trained on previous leading datasets like OBELICS. Including more diverse sources in MINT-1, T resulted in better generalization and performance across various benchmarks. Notably, the dataset significantly improved performance in tasks involving visual question answering and multimodal reasoning. The researchers found that models trained on MINT-1T performed better across multiple demonstrations, highlighting the dataset’s effectiveness.

The MINT-1T dataset’s construction included detailed steps to ensure data quality and diversity. For instance, the dataset consists of 922 billion HTML tokens, 106 billion PDF tokens, and 9 billion ArXiv tokens. The filtering process involved eliminating documents with inappropriate content and non-English texts, using tools like Fasttext for language identification and NSFW detectors for image content. The deduplication process was crucial, involving Bloom filters to remove duplicate paragraphs and documents and hashing techniques to eliminate repetitive images.

In conclusion, the MINT-1T dataset addresses dataset scarcity and diversity. By introducing a larger and more varied dataset, the researchers have enabled the development of more robust and high-performing open-source multimodal models. This work highlights the importance of data diversity and scale in AI research and paves the way for future improvements and applications in multimodal AI. The dataset’s extensive scale, including one trillion text tokens and 3.4 billion images, provides a solid foundation for advancing AI capabilities.


Check out the Paper, Details, and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 47k+ ML SubReddit

Find Upcoming AI Webinars here


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age
AI & Technology

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

September 23, 2026
Never Use ChatGPT For These Five Tasks
AI & Technology

Never Use ChatGPT For These Five Tasks

September 23, 2026
Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes
AI & Technology

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

September 23, 2026
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
AI & Technology

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

September 23, 2026
Next Post
Trump on trial: Cohen says Trump directed payment to Daniels

Trump on trial: Cohen says Trump directed payment to Daniels

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
7 Ways To Get Free Movies And TV Channels On Your Smart TV

7 Ways To Get Free Movies And TV Channels On Your Smart TV

September 20, 2026
Couple married in waist-deep water in the Philippines

Couple married in waist-deep water in the Philippines

September 17, 2026
Trump says U.S. has entered oil deal with Venezuela

Trump says U.S. has entered oil deal with Venezuela

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!