• bitcoinBitcoin(BTC)$77,227.001.43%
  • ethereumEthereum(ETH)$2,470.511.90%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$748.953.71%
  • rippleXRP(XRP)$1.321.99%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$104.975.86%
  • tronTRON(TRX)$0.3358240.07%
  • zcashZcash(ZEC)$1,492.3410.49%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.15%
  • HyperliquidHyperliquid(HYPE)$86.549.49%
  • dogecoinDogecoin(DOGE)$0.0842574.39%
  • moneroMonero(XMR)$515.693.18%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$79.541.63%
  • RainRain(RAIN)$0.012660-1.56%
  • chainlinkChainlink(LINK)$11.726.05%
  • leo-tokenLEO Token(LEO)$8.89-0.57%
  • cardanoCardano(ADA)$0.2136139.58%
  • stellarStellar(XLM)$0.1877933.39%
  • uniswapUniswap(UNI)$8.6127.00%
  • bitcoin-cashBitcoin Cash(BCH)$244.7011.29%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.4630.49%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.10743810.64%
  • litecoinLitecoin(LTC)$54.444.98%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.353.27%
  • avalanche-2Avalanche(AVAX)$7.884.87%
  • hedera-hashgraphHedera(HBAR)$0.0771684.76%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • suiSui(SUI)$0.789.02%
  • shiba-inuShiba Inu(SHIB)$0.0000057.50%
  • crypto-com-chainCronos(CRO)$0.0586280.68%
  • MemeCoreMemeCore(M)$1.2814.82%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BittensorBittensor(TAO)$238.967.30%
  • tether-goldTether Gold(XAUT)$4,353.511.48%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$114.052.58%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.08%
  • aaveAave(AAVE)$133.6410.72%
  • AsterAster(ASTER)$0.754.05%
  • mantleMantle(MNT)$0.583.95%
  • Pump.funPump.fun(PUMP)$0.0040887.75%
  • polkadotPolkadot(DOT)$1.1211.52%
  • pax-goldPAX Gold(PAXG)$4,351.001.38%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Enhancing Self-Supervised Learning with Automatic Data Curation: A Hierarchical K-Means Approach

May 31, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Enhancing Self-Supervised Learning with Automatic Data Curation: A Hierarchical K-Means Approach
ShareShareShareShareShare

Self-supervised features are central to modern machine learning, typically requiring extensive human effort for data collection and curation, similar to supervised learning. Self-supervised learning (SSL) allows models to be trained without human annotations, enabling scalable data and model expansion. However, scaling efforts have sometimes resulted in subpar performance due to issues like the long-tail distribution of concepts in uncurated datasets. Successful SSL applications involve careful data curation, such as filtering internet data to match high-quality sources like Wikipedia for language models or balancing visual concepts for image models. This curation enhances robustness and performance in downstream tasks.

Researchers from FAIR at Meta, INRIA, Université Paris Saclay, and Google address the automatic curation of high-quality datasets for self-supervised pre-training. They propose a clustering-based approach to create large, diverse, balanced datasets. This method involves hierarchical k-means clustering on a vast data repository and balanced sampling from these clusters. Experiments across web images, satellite images, and text demonstrate that features trained on these curated datasets outperform those trained on uncurated data, matching or exceeding manually curated data. This approach addresses the challenge of balancing datasets to improve model performance in self-supervised learning.

✅ [Featured Article] LLMWare.ai Selected for 2024 GitHub Accelerator: Enabling the Next Wave of Innovation in Enterprise RAG with Small Specialized Language Models

SSL is crucial in modern machine learning. In natural language processing (NLP), language modeling has evolved from simple neural architectures to large-scale models, significantly advancing the field. Similarly, SSL in computer vision has progressed from pretext tasks to sophisticated joint embedding architectures, employing methods like contrastive learning, clustering, and distillation. High-quality data is essential for training state-of-the-art models. Automatic data curation techniques, such as hierarchical k-means clustering, are proposed to balance large datasets without requiring labels, improving the performance of SSL models across various domains.

The pre-training dataset must be large, diverse, and balanced to train models effectively using self-supervised learning. Balanced datasets ensure each concept is equally represented, avoiding biases toward dominant concepts. Creating such datasets involves selecting balanced subsets from large online repositories, often using clustering methods like k-means. However, standard k-means may over-represent dominant concepts. To address this, hierarchical k-means with resampling can be used, ensuring centroids follow a uniform distribution. This process, combined with specific sampling strategies, helps maintain balance across various conceptual levels in the dataset, promoting better model performance.

Four experiments were conducted to study the proposed algorithm. Initially, simulated data was used to illustrate hierarchical k-means, showing a more uniform cluster distribution than other methods. Next, web-based image data was curated, resulting in a dataset of 743 million images, and a ViT-L model was trained and evaluated on various benchmarks, demonstrating improved performance. The algorithm was then applied to curate text data for training large language models, yielding significant gains across benchmarks. Lastly, satellite images were curated for tree canopy height prediction, enhancing model performance on all evaluated datasets.

In conclusion, The study introduces an automatic data curation pipeline that generates large, diverse, and balanced training datasets for self-supervised feature learning. By successively applying k-means clustering and resampling, the method ensures uniform cluster distribution among concepts. Extensive experiments show that this pipeline enhances feature learning across web-based images, satellite imagery, and text data. The curated datasets outperform raw data and ImageNet1k in robustness but slightly lags behind the heavily curated ImageNet22k on certain benchmarks. The approach highlights the importance of data curation in self-supervised learning and suggests hierarchical k-means as a valuable alternative in various data-dependent tasks. Future work should address dataset quality, reliance on pre-trained features, and scalability. Automated dataset creation poses risks such as reinforcing biases and privacy breaches, mitigated here by blurring human faces and striving for concept balance.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 43k+ ML SubReddit | Also, check out our AI Events Platform


YOU MAY ALSO LIKE

eGPUs Do Work, But They Come With Some Notable Limitations

Google’s Revamped CC Is An AI Agent For Families And Groups

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Google’s Revamped CC Is An AI Agent For Families And Groups
AI & Technology

Google’s Revamped CC Is An AI Agent For Families And Groups

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year
AI & Technology

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Next Post
The Decline of Publicly-Listed Companies

The Decline of Publicly-Listed Companies

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Stalking suspect arrested near Kris Jenner’s L.A. home

Stalking suspect arrested near Kris Jenner’s L.A. home

September 14, 2026
An iOS 27 Bug Can Temporarily Freeze Your iPhone

An iOS 27 Bug Can Temporarily Freeze Your iPhone

September 17, 2026
Widow breaks tradition of keeping politics out of 9/11 remembrances

Widow breaks tradition of keeping politics out of 9/11 remembrances

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!