• bitcoinBitcoin(BTC)$84,465.00-2.02%
  • ethereumEthereum(ETH)$2,683.87-2.65%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$765.64-2.92%
  • rippleXRP(XRP)$1.50-4.83%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$114.94-3.07%
  • tronTRON(TRX)$0.341864-0.05%
  • zcashZcash(ZEC)$1,501.24-7.77%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.54%
  • HyperliquidHyperliquid(HYPE)$94.01-2.94%
  • dogecoinDogecoin(DOGE)$0.092575-7.88%
  • moneroMonero(XMR)$551.52-3.60%
  • whitebitWhiteBIT Coin(WBT)$84.78-2.23%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$12.35-5.42%
  • cardanoCardano(ADA)$0.238468-6.99%
  • RainRain(RAIN)$0.012244-6.42%
  • leo-tokenLEO Token(LEO)$8.96-0.16%
  • stellarStellar(XLM)$0.202331-6.56%
  • bitcoin-cashBitcoin Cash(BCH)$336.86-1.94%
  • uniswapUniswap(UNI)$9.30-7.76%
  • nearNEAR Protocol(NEAR)$4.30-2.63%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$62.04-2.24%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.32-8.23%
  • USD1USD1(USD1)$1.000.01%
  • CantonCanton(CC)$0.109900-4.63%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-4.05%
  • hedera-hashgraphHedera(HBAR)$0.090575-8.80%
  • suiSui(SUI)$0.96-6.64%
  • shiba-inuShiba Inu(SHIB)$0.000006-7.88%
  • BittensorBittensor(TAO)$287.60-9.27%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.061142-8.75%
  • BitwayBitway(BTW)$1.0417.17%
  • MemeCoreMemeCore(M)$1.22-7.07%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,289.52-1.65%
  • okbOKB(OKB)$118.89-3.41%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.23%
  • mantleMantle(MNT)$0.65-2.51%
  • aaveAave(AAVE)$139.09-5.91%
  • EthenaEthena(ENA)$0.205530-6.19%
  • OndoOndo(ONDO)$0.413498-6.66%
  • AsterAster(ASTER)$0.69-5.21%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Proposes Novel Machine Learning Algorithms for Differentially Private Partition Selection

August 23, 2025
in AI & Technology
Reading Time: 9 mins read
A A
Google AI Proposes Novel Machine Learning Algorithms for Differentially Private Partition Selection
ShareShareShareShareShare

Differential privacy (DP) stands as the gold standard for protecting user information in large-scale machine learning and data analytics. A critical task within DP is partition selection—the process of safely extracting the largest possible set of unique items from massive user-contributed datasets (such as queries or document tokens), while maintaining strict privacy guarantees. A team of researchers from MIT and Google AI Research present novel algorithms for differentially private partition selection, which is an approach to maximize the number of unique items selected from a union of sets of data, while strictly preserving user-level differential privacy

The Partition Selection Problem in Differential Privacy

At its core, partition selection asks: How can we reveal as many distinct items as possible from a dataset, without risking any individual’s privacy? Items only known to a single user must remain secret; only those with sufficient “crowdsourced” support can be safely disclosed. This problem underpins critical applications such as:

YOU MAY ALSO LIKE

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

  • Private vocabulary and n-gram extraction for NLP tasks.
  • Categorical data analysis and histogram computation.
  • Privacy-preserving learning of embeddings over user-provided items.
  • Anonymizing statistical queries (e.g., to search engines or databases).

Standard Approaches and Limits

Traditionally, the go-to solution (deployed in libraries like PyDP and Google’s differential privacy toolkit) involves three steps:

  1. Weighting: Each item receives a “score”, usually its frequency across users, with every user’s contribution strictly capped.
  2. Noise Addition: To hide precise user activity, random noise (usually Gaussian) is added to each item’s weight.
  3. Thresholding: Only items whose noisy score passes a set threshold—calculated from privacy parameters (ε, δ)—are released.

This method is simple and highly parallelizable, allowing it to scale to gigantic datasets using systems like MapReduce, Hadoop, or Spark. However, it suffers from fundamental inefficiency: popular items accumulate excess weight that doesn’t further aid privacy, while less-common but potentially valuable items often miss out because the excess weight isn’t redirected to help them cross the threshold.

Adaptive Weighting and the MaxAdaptiveDegree (MAD) Algorithm

Google’s research introduces the first adaptive, parallelizable partition selection algorithm—MaxAdaptiveDegree (MAD)—and a multi-round extension MAD2R, designed for truly massive datasets (hundreds of billions of entries).

Key Technical Contributions

  • Adaptive Reweighting: MAD identifies items with weight far above the privacy threshold, reroutes the excess weight to boost lesser-represented items. This “adaptive weighting” increases the probability that rare-but-shareable items are revealed, thus maximizing output utility.
  • Strict Privacy Guarantees: The rerouting mechanism maintains the exact same sensitivity and noise requirements as classic uniform weighting, ensuring user-level (ε, δ)-differential privacy under the central DP model.
  • Scalability: MAD and MAD2R require only linear work in dataset size and a constant number of parallel rounds, making them compatible with massive distributed data processing systems. They need not fit all data in-memory and support efficient multi-machine execution.
  • Multi-Round Improvement (MAD2R): By splitting privacy budget between rounds and using noisy weights from the first round to bias the second, MAD2R further boosts performance, allowing even more unique items to be safely extracted—especially in long-tailed distributions typical of real-world data.

How MAD Works—Algorithmic Details

  1. Initial Uniform Weighting: Each user shares their items with a uniform initial score, ensuring sensitivity bounds.
  2. Excess Weight Truncation and Rerouting: Items above an “adaptive threshold” have their excess weight trimmed and rerouted proportionally back to contributing users, who then redistribute this to their other items.
  3. Final Weight Adjustment: Additional uniform weight is added to make up for small initial allocation mistakes.
  4. Noise Addition and Output: Gaussian noise is added; items above the noisy threshold are output.

In MAD2R, the first-round outputs and noisy weights are used to refine which items should be focused on in the second round, with weight biases ensuring no privacy loss and further maximizing output utility.

Experimental Results: State-of-the-Art Performance

Extensive experiments across nine datasets (from Reddit, IMDb, Wikipedia, Twitter, Amazon, all the way to Common Crawl with nearly a trillion entries) show:

  • MAD2R outperforms all parallel baselines (Basic, DP-SIPS) on seven out of nine datasets in terms of number of items output at fixed privacy parameters.
  • On the Common Crawl dataset, MAD2R extracted 16.6 million out of 1.8 billion unique items (0.9%), but covered 99.9% of users and 97% of all user-item pairs in the data—demonstrating remarkable practical utility while holding the line on privacy.
  • For smaller datasets, MAD approaches the performance of sequential, non-scalable algorithms, and for massive datasets, it clearly wins in both speed and utility.
https://research.google/blog/securing-private-data-at-scale-with-differentially-private-partition-selection/
https://research.google/blog/securing-private-data-at-scale-with-differentially-private-partition-selection/

Concrete Example: Utility Gap

Consider a scenario with a “heavy” item (very commonly shared) and many “light” items (shared by few users). Basic DP selection overweights the heavy item without lifting the light items enough to pass the threshold. MAD strategically reallocates, increasing the output probability of the light items and resulting in up to 10% more unique items discovered compared to the standard approach.

Summary

With adaptive weighting and parallel design, the research team brings DP partition selection to new heights in scalability and utility. These advances ensure researchers and engineers can make fuller use of private data, extracting more signal without compromising individual user privacy.


Check out the Blog and Technical paper here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses
AI & Technology

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

September 23, 2026
Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips
AI & Technology

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

September 23, 2026
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
AI & Technology

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

September 23, 2026
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
AI & Technology

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

September 23, 2026
Next Post
I Lied To My Husband About Our Finances Until It Was Too Late

I Lied To My Husband About Our Finances Until It Was Too Late

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Peloton Has Made A Foldable (Treadmill)

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

September 23, 2026
Meta settles social media addiction lawsuit for  billion

Meta settles social media addiction lawsuit for $18 billion

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!